Chroma format and bit depth indication in coded video
By including complete signaling chroma format, bit depth, image width, and height in the VVC video file format, the problem of incomplete signaling in existing technologies is solved, achieving more accurate and efficient video decoding processing.
Patent Information
- Application Number
- CN202111090735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-17
- Filing Date
- 2021-09-17
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-09-17
AI Technical Summary
The existing VVC video file format lacks complete indication of chroma format, bit depth, picture width and picture height in signaling, resulting in incomplete decoder configuration and operating point information, affecting the accuracy and efficiency of video processing.
Signaling chroma format and bit depth in VvcDecoderConfigurationRecord and adding indication of picture width and height in operation point information sample group and entity group ensures these parameters are passed correctly in all cases.
It implements complete signaling of the VVC video file format, ensuring that the decoder can correctly decode and process multi-layer, multi-operation point video data, improving the accuracy and efficiency of video processing.
Smart Images

Figure CN114205603B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is filed to claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 079,910, filed on September 17, 2020, in accordance with applicable patent laws and / or the rules of the Paris Convention. The entire disclosure of the above application is incorporated by reference into this disclosure for all purposes of law. Technical Field
[0003] This patent document relates to the generation, storage and consumption of digital audio-visual media information in file format. Background Art
[0004] Digital video accounts for the largest usage of bandwidth on the Internet and other digital communications networks. Bandwidth demand for digital video usage is expected to continue to grow as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of videos or images according to file formats.
[0006] In one example aspect, a method for processing visual media data is disclosed. The method includes performing conversion between the visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule, wherein the format rule specifies whether a first element indicating whether the track contains a bitstream control corresponding to a particular output layer set, a second element indicating a chroma format for the track, and / or a third element indicating bit depth information for the track is included in a configuration record for the track.
[0007] In another example aspect, another method for processing visual media data is disclosed. The method includes performing conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule, wherein the format rule specifies whether to include a first element indicating a picture width of the track and / or a second element indicating a picture height of the track in a configuration record for the track based on (1) whether a third element indicates whether the track contains a specific bitstream corresponding to a specific output layer set and / or (2) whether the configuration record is for a single-layer bitstream, and wherein the format rule further specifies that when the first element and / or the second element are included in the configuration record for the track, the first element and / or the second element are represented in a field comprising 16 bits.
[0008] In another example aspect, another method for processing visual media data is disclosed. The method includes performing conversion between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file includes operation point records and operation point group boxes, and wherein the format rule specifies that for each operation point indicated in the visual media file, a first element indicating a chroma format, a second element indicating bit depth information, a third element indicating a maximum picture width, and / or a fourth element indicating a maximum picture height are included in the operation point record and the operation point group box.
[0009] In yet another example aspect, a video processing apparatus is disclosed. The video processing apparatus includes a processor configured to implement the above-described method.
[0010] In yet another example aspect, a method of storing visual media data to a file including one or more bitstreams is disclosed. The method corresponds to the above-described method, and further includes storing the one or more bitstreams to a non-transitory computer-readable recording medium.
[0011] In yet another example aspect, a computer-readable medium storing a bitstream is disclosed. The bitstream is generated according to the above-described method.
[0012] In yet another example aspect, a video processing apparatus for storing a bitstream is disclosed, wherein the video processing apparatus is configured to implement the above-described method.
[0013] In yet another example aspect, a computer-readable medium is disclosed, wherein a bitstream conforming to a file format generated according to the above-described method is stored.
[0014] These and other features are described throughout this document. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 is a block diagram of an example video processing system.
[0016] Figure 2 is a block diagram of a video processing apparatus.
[0017] Figure 3 is a flowchart of an example method of video processing.
[0018] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.
[0019] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0020] Figure 6is a block diagram illustrating a decoder according to some embodiments of the disclosure.
[0021] Figure 7 An example of an encoder block diagram is shown.
[0022] Figures 8-10 An example method of processing visual media data based on some implementations of the disclosed technology is shown. DETAILED DESCRIPTION
[0023] The use of section headings in this document is for convenience only and does not limit the applicability of techniques and embodiments disclosed in each section to the section alone. Also, H.266 terminology is used in some descriptions only for ease of understanding and not for limiting the scope of the disclosed technology. Thus, the technology described herein is applicable to other video codec protocols and designs as well. In this document, editorial changes to text are shown by strikeout and highlighting (including boldface italics) indicating deleted text and added text, respectively, relative to the current draft of the VVC specification or ISOBMFF file format specification.
[0024] 1. Preliminary Discussion
[0025] This document is related to video file formats. In particular, it relates to the signaling of picture format information, including chroma format, bit depth, picture width, and picture height, in a media file carrying a Versatile Video Coding (VVC) video bitstream based on the ISO Base Media File Format (ISOBMFF). These ideas can be applied to video bitstreams coded by any codec (e.g., the VVC standard), and to any video file format, e.g., the VVC video file format that is being developed, individually or in various combinations.
[0026] 2. Abbreviations
[0027] ACT Adaptive Color Transform
[0028] ALF Adaptive Loop Filter
[0029] AMVR Adaptive Motion Vector Resolution
[0030] APS Adaptive Parameter Set
[0031] AU Access Unit
[0032] AUD Access Unit Delimiter
[0033] AVC Advanced Video Coding (Rec. ITU-T H.264 | ISO / IEC 14496-10)
[0034] B Bi-prediction
[0035] BCW bi-prediction with CU-level weights
[0036] BDOF bi-directional optical flow
[0037] BDPCM block-based delta pulse code modulation
[0038] BP buffering period
[0039] CABAC context-based adaptive binary arithmetic coding
[0040] CB coded block
[0041] CBR constant bit rate
[0042] CCALF cross-component adaptive loop filter
[0043] CPB coded picture buffer
[0044] CRA clean random access
[0045] CRC cyclic redundancy check
[0046] CTB coded tree block
[0047] CTU coded tree unit
[0048] CU coding unit
[0049] CVS coded video sequence
[0050] DPB decoded picture buffer
[0051] DCI decoding capability information
[0052] DRAP dependent random access point
[0053] DU decoding unit
[0054] DUI decoding unit information
[0055] EG exponential-golomb
[0056] EGk exponential-golomb of order k
[0057] EOB end of bitstream
[0058] EOS end of sequence
[0059] FD filler data
[0060] FIFO first-in-first-out
[0061] FL fixed length
[0062] GBR green, blue, and red
[0063] GCI general constraint information
[0064] GDR gradual decoding refresh
[0065] GPM geometric partition mode
[0066] HEVC high efficiency video coding (Rec. ITU-T H.265 | ISO / IEC 23008-2)
[0067] HDR hypothetical reference decoder
[0068] HSS hypothetical stream scheduler
[0069] I intra
[0070] IBC intra block copy
[0071] IDR instantaneous decoding refresh
[0072] ILRP inter-layer reference picture
[0073] IRAP intra random access point
[0074] LFNST low-frequency non-separable transform
[0075] LPS least probable symbol
[0076] LSB least significant bit
[0077] LTRP long-term reference picture
[0078] LMCS luminance mapping and chrominance scaling
[0079] MIP matrix-based intra prediction
[0080] MPS most probable symbol
[0081] MSB most significant bit
[0082] MTS multiple transform selection
[0083] MVP motion vector prediction
[0084] NAL network abstraction layer
[0085] OLS output layer set
[0086] OP operation point
[0087] OPI operation point information
[0088] P prediction
[0089] PH picture header
[0090] POC picture order count
[0091] PPS picture parameter set
[0092] PROF prediction refinement of optical flow
[0093] PT picture timing
[0094] PU picture unit
[0095] QP quantization parameter
[0096] RADL random access decodable leading (picture)
[0097] RASL random access skipped leading (picture)
[0098] RBSP raw byte sequence payload
[0099] RGB red, green, and blue
[0100] RPL reference picture list
[0101] SAO sample adaptive offset
[0102] SAR sample aspect ratio
[0103] SEI supplemental enhancement information
[0104] SH slice header
[0105] SLI subpicture level information
[0106] SODB string of data bits
[0107] SPS sequence parameter set
[0108] STRP short-term reference picture
[0109] STSA stepping time dimension sub-layer access
[0110] TR truncated rice
[0111] VBR variable bit rate
[0112] VCL video coding layer
[0113] VPS video parameter set
[0114] VSEI general supplemental enhancement information (Rec. ITU-T H.274 | ISO / IEC 23002-7)
[0115] VUI video usability information
[0116] VVC Versatile Video Coding (Rec. ITU-T H.266 | ISO / IEC 23090-3)
[0117] 3. Introduction to Video Coding
[0118] 3.1 Video Coding Standards
[0119] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure where temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by the JVET and applied to the reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as Joint Video Expert Team (JVET). VVC is a new coding standard, which aims to reduce 50% bit rate compared to HEVC, and the standard has been finalized by the JVET at its 19th meeting, which ended on July 1, 2020.
[0120] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for a wide range of applications, including traditional uses such as television broadcast, video conferencing, or playback from storage media as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, combining and merging content from multiple coded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360° immersive media.
[0121] 3.2 File Format Standards
[0122] Media streaming applications are typically based on IP, TCP and HTTP transport methods and often rely on file formats, such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). For video formats using ISOBMFF and DASH, a video format specific file format specification, such as the AVC file format and the HEVC file format in ISO / IEC 14496-15 (“Information technology - Coding of audio-visual objects - Part 15: Carriage of Network Abstraction Layer (NAL) unit structured video in the ISO base media file format”), is required to encapsulate video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream, such as profile, level and tier, as well as many other information, need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g., for selecting appropriate media segments, for initialization at the beginning of a streaming session and for stream adaptation during a streaming session.
[0123] Similarly, for image formats using ISOBMFF, an image format specific file format specification, such as the AVC image file format and the HEVC image file format in MPEG output document N19454 (“Information technology - Coding of audio-visual objects - Part 15: Carriage of Network Abstraction Layer (NAL) unit structured video in the ISO base media file format - Amendment 2: Carriage of VVC and EVC in ISOBMFF”, July 2020), is required.
[0124] The VVC video file format, i.e., a file format based on ISOBMFF for storing VVC video content, is currently being developed by MPEG. The latest draft specification of the VVC video file format is contained in MPEG output document N19460 (“Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image file format - Amendment 3: Support for VVC, EVC, slideshow and other improvements”, July 2020).
[0125] The VVC image file format, i.e., a file format based on ISOBMFF for storing image content coded using VVC, is currently being developed by MPEG. The latest draft specification of the VVC image file format is contained in
[12] .
[0126] 3.3 Some details of the VVC video file format
[0127] 3.3.1 Decoder configuration information
[0128] 3.3.1.1 VVC decoder configuration record
[0129] 3.3.1.1.1 Definition
[0130] This section specifies decoder configuration information for ISO / IEC 23090-3 video content.
[0131] This record contains the size of the length field used in each sample to indicate the length of the NAL units it contains and parameter sets if stored in the sample entry. This record is for the outer framework (whose size is provided by the structure containing it).
[0132] This record contains a version field. The specification of this version defines version 1 of this record. Incompatible changes to the record will be indicated by a change in the version number. If the version number is not recognized, the reader should not attempt to decode the record or apply the stream of the record.
[0133] Compatible extensions to this record will extend it and will not change the configuration version code. The reader should be prepared to ignore unrecognized data definitions beyond their understanding.
[0134] A VvcPtlRecord shall be present in the decoder configuration record when the track itself contains a VVC bitstream or by resolving the'subp' track reference. If ptl_present_flag is equal to zero in the decoder configuration record of the track, the track shall have an 'oref' track reference.
[0135] The values of the syntax elements VvcPTLRecord, chroma_format_idc, and bit_depth_minus8 are valid for all parameter sets activated when the stream described by this record is decoded (referred to as "all parameter sets" in the following sentences of this paragraph). Specifically, the following restrictions apply:
[0136] The profile indicates that general_profile_idc shall indicate the profile to which the stream associated with this configuration record conforms.
[0137] NOTE 1: If the SPS is marked with different profiles, the stream can need to be inspected to determine which profile, if any, the entire stream conforms to. If the entire stream is not inspected, or the inspection indicates that there is no profile to which the entire stream conforms, the entire stream shall be split into two or more sub-streams with separate configuration records in which the rules can be satisfied.
[0138] The tier indicates that general_tier_flag shall indicate a tier equal to or greater than the highest tier indicated in all parameter sets.
[0139] Each bit in general_constraint_info can only be set if the bit is also set for all parameter sets.
[0140] level indicates that general_level_idc shall indicate a capability level equal to or greater than the highest level indicated for the highest level in all parameter sets.
[0141] The following constraint applies to chroma_format_idc:
[0142] - If the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS referenced by the NAL units of the track, chroma_format_idc shall be equal to sps_chroma_format_idc.
[0143] - Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[ output_layer_set_idx ] as defined in ISO / IEC 23090-3.
[0144] - Otherwise, chroma_format_idc shall not be present.
[0145] The following constraint applies to bit_depth_minus8:
[0146] - If the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 is the same in all SPS referenced by the NAL units of the track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0147] - Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[ output_layer_set_idx ] as defined in ISO / IEC 23090-3.
[0148] - Otherwise, bit_depth_minus8 shall not be present.
[0149] Explicit indication about chroma format and bit depth and other important format information used by VVC video elementary stream is provided in VVC decoder configuration record. Two different VVC sample entries are also required if two sequences have different color space indication in their VUI information.
[0150] There are array sets to carry initialization NAL units. NAL unit types are restricted to indicate DCI, VPS, SPS, PPS, prefix APS and prefix SEI NAL units. NAL unit types reserved in ISO / IEC 23090-3 and this specification can be defined in the future, the reader shall ignore the array with reserved or not allowed NAL unit type value.
[0151] NOTE 2: This "tolerant" behavior is designed to avoid errors and to allow the possibility of backward compatible extension of these arrays in future specifications.
[0152] NOTE 3: The NAL units carried in the sample entry are included in the access unit immediately following the reconstruction of the first sample from the reference sample entry or after the AUD and OPI NAL units (if any) at the beginning of the access unit.
[0153] The order of the array is recommended as DCI, VPS, SPS, PPS, prefix APS, prefix SEI.
[0154] 3.3.1.1.2 Syntax
[0155]
[0156]
[0157] 3.3.1.1.3 Semantics
[0158] general_profile_idc, general_tier_flag, general_sub_profile_idc, general_constraint_info, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, sublayer_level_present, and sublayer_level_idc[i] contain the matching values of the bits in the fields general_profile_idc, general_tier_flag, general_sub_profile_idc, general_constraint_info(), general_level_idc, ptl_multilayer_enabled_flag, ptl_frame_only_constraint_flag, sublayer_level_present, and sublayer_level_idc[i] defined in ISO / IEC 23090-3 for the stream to which this configuration record is applied.
[0159] avgFrameRate gives the average frame rate of the stream to which this configuration record is applied in frames per (256 seconds). The value 0 indicates that the average frame rate is not specified.
[0160] constantFrameRate equal to 1 indicates that the stream to which this configuration record is applied has a constant frame rate. The value 2 indicates that each temporal layer representation in the stream has a constant frame rate. The value 0 indicates that the stream can or can not have a constant frame rate.
[0161] numTemporalLayers greater than 1 indicates that the track to which this configuration record is applied is scalable in the temporal domain and contains a number of temporal layers (also referred to as temporal sub-layers or sub-layers in ISO / IEC 23090-3) equal to numTemporalLayers. The value 1 indicates that the track to which this configuration record is applied is not scalable in the temporal domain. The value 0 indicates that it is not known whether the track to which this configuration record is applied is scalable in the temporal domain.
[0162] lengthSizeMinusOne plus 1 indicates the byte length of the NALUnitLength field in the VVC video stream samples in the stream to which this configuration record is applied. For example, a size of one byte is indicated with the value 0. The value of this field shall be one of 0, 1, or 3, corresponding to lengths coded with 1, 2, or 4 bytes, respectively.
[0163] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular output layer set, but can contain one or more individual layers that do not form an output layer set or individual sub-layers that do not include a sub-layer with Temporalld equal to 0.
[0164] num_sub_profiles defines the number of sub-profiles indicated in the decoder configuration record.
[0165] track_ptl specifies the profile, level and tier of the output layer set represented by the VVC bitstream contained in the track.
[0166] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. As specified in ISO / IEC 23090-3, the value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided to the VVC decoder by external means for decoding the bitstream contained in the track.
[0167] chroma_format_present_flag equal to 0 specifies that chroma_format_idc is not present. chroma_format_present_flag equal to 1 specifies that chroma_format_idc is present.
[0168] chroma_format_idc indicates the chroma format applied to the track. The following constraints apply to chroma_format_idc:
[0169] - If the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 is the same in all SPS referenced by the NAL units of the track, chroma_format_idc shall be equal to sps_chroma_format_idc.
[0170] - Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[ output_layer_set_idx ] as defined in ISO / IEC 23090-3.
[0171] - Otherwise, chroma_format_idc shall not be present.
[0172] bit_depth_present_flag equal to 0 specifies that bit_depth_minus8 is not present. bit_depth_present_flag equal to 1 specifies that bit_depth_minus8 is present.
[0173] bit_depth_minus8 indicates the bit depth applied to this track. The following constraint applies to bit_depth_minus8:
[0174] - If the value of sps_bitdepth_minus8 defined in ISO / IEC 23090-3 is the same in all SPS referenced by the NAL units of the track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0175] - Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[output_layer_set_idx] defined in ISO / IEC 23090-3.
[0176] - Otherwise, bit_depth_minus8 shall not be present.
[0177] numAnumArrays indicates the number of arrays of NAL units of the indicated type.
[0178] When array_completeness is equal to 1, it indicates that all NAL units of the given type are in the following array and that no NAL unit is in the stream; when array_completeness is equal to 0, it indicates that additional NAL units of the indicated type can be in the stream; the default value and allowed values are constrained by the sample entry name.
[0179] NAL_unit_type indicates the type of NAL units (all shall be of this type) in the following array; it takes values defined in ISO / IEC 23090-2; it is limited to take one of the values indicating DCI, VPS, SPS, PPS, APS, prefix SEI, or suffix SEI NAL units.
[0180] numNalus indicates the number of NAL units of the indicated type contained in the configuration record of the stream for which this configuration record applies. SEI arrays shall contain only SEI messages of a “declarative” nature, i.e., those providing information about the entire stream. One example of such SEI can be the user data SEI.
[0181] nalUnitLength indicates the byte length of the NAL unit.
[0182] A nalUnit contains a DCI, VPS, SPS, PPS, APS, or declarative SEI NAL unit as specified in ISO / IEC 23090-3.
[0183] 3.3.2 Operation point information sample group
[0184] 3.3.2.1 Definition
[0185] An application learns about the different operation points provided by a given VVC bitstream and their composition by using the operation point information sample group ('vopi'). Each operation point is associated with an output layer set, a maximum Temporalld value, and tier, level, and class signaling. All this information is captured by the 'vopi' sample group. In addition to this information, the sample group provides inter-layer dependency information.
[0186] When a VVC bitstream has more than one VVC track and the VVC bitstream does not have an operation point entity group, the following two cases apply:
[0187] - There shall be and there shall be only one track in the VVC track of the VVC bitstream that carries the 'vopi' sample group.
[0188] - All other VVC tracks of the VVC bitstream shall have a track reference of type "oref" pointing to the track carrying the 'vopi' sample group.
[0189] For any particular sample in a given track, a temporally co-located sample in another track is defined as the sample that has the same decoding temporal domain as the particular sample. For each sample S N in a track T N with an 'oref' track reference pointing to the track T K carrying the 'vopi' sample group, the following applies:
[0190] - If there is a temporally co-located sample S k in the track T k , sample S N is associated with the same 'vopi' sample group entry as sample S k .
[0191] - Otherwise, sample S N is associated with the same 'vopi' sample group entry as the last sample in track T k preceding sample S N in the decoding temporal domain.
[0192] When a VVC bitstream refers to multiple VPSs, it can be necessary to include multiple entries of grouping type 'vopi' in the sample group description box. For the more common case where there is a single VPS, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the operation point information sample group in the sample table box instead of including it in each track fragment.
[0193] grouping_type_parameter is not defined for SampleToGroupBox with grouping type 'vopi'.
[0194] 3.3.2.2 Syntax
[0195]
[0196]
[0197]
[0198] 3.3.2.3 Semantics
[0199] num_profile_tier_level_minus1 plus 1 gives the number of the following profile, tier and level combinations and the related fields.
[0200] ptl_max_temporal_id[ i ] : gives the maximum TemporalID of the NAL units of the associated bitstream referring to the specified i-th profile, tier and level structure.
[0201] NOTE: The semantics of ptl_max_temporal_id[ i ] and max_temporal_id for an operation point given below are different, although they can have the same numeric value.
[0202] ptl[ i ] specifies the i-th profile, tier and level structure.
[0203] all_independent_layers_flag, each_layer_is_an_ols_flag, ols_mode_idc and max_tid_il_ref_pics_plus1 are defined in ISO / IEC 23090-3.
[0204] num_operating_points : gives the number of operation points for which the following information applies.
[0205] output_layer_set_idx is the index of the output layer set defining the operation point. The mapping between output_layer_set_idx and layer_id values shall be the same as specified by the VPS for the output layer set with index output_layer_set_idx.
[0206] ptl_idx: Signalled for the output layer set with index output_layer_set_idx the zero-based index of the listed profile, level and tier structure.
[0207] max_temporal_id: Gives the maximum Temporalld of the NAL units of this operation point.
[0208] NOTE: The maximum Temporalld value indicated in the layer information sample group has a different semantics than the maximum Temporalld indicated here. However, they can have the same literal value.
[0209] layer_count: This field indicates the number of necessary layers of the operation point as defined in ISO / IEC 23090-3.
[0210] layer_id: Provides the nuh_layer_id values of the layers of the operation point.
[0211] is_outputlayer: A flag indicating whether a layer is an output layer. One indicates one output layer.
[0212] frame_rate_info_flag equal to 0 indicates that there is no frame rate information for the operation point. Value 1 indicates that there is frame rate information for the operation point.
[0213] bit_rate_info_flag equal to 0 indicates that there is no bit rate information for the operation point. Value 1 indicates that there is bit rate information for the operation point.
[0214] avgFrameRate gives the average frame rate of the operation point in frames per (256 seconds). Value 0 indicates that the average frame rate is not specified.
[0215] constantFrameRate equal to 1 indicates that the stream of the operation point has a constant frame rate. Value 2 indicates that each temporal layer representation in the stream of the operation point has a constant frame rate. Value 0 indicates that the stream of the operation point can or can not have a constant frame rate.
[0216] maxBitRate gives the maximum bit rate of the stream of the operation point in bits per second over any 1 second window.
[0217] avgBitRate gives the average bit rate of the stream at the operating point in bits per second.
[0218] max_layer_count: The count of all unique layers in all operation points related to this associated base track.
[0219] layerID: nuh_layer_id of the layer for which all direct reference layers are given in the direct_ref_layerID loop below.
[0220] num_direct_ref_layers: Number of direct reference layers for the layer whose nuh_layer_id is equal to layerID.
[0221] direct_ref_layerID: nuh_layer_id of the direct reference layer.
[0222] 3.3.3 Operation Point Entity Group
[0223] 3.3.3.1 Overview
[0224] An operating point entity group is defined to provide mapping of tracks to operating points and grade level information of the operating points.
[0225] The implicit reconstruction process when aggregating track samples that map to the operation point described in this entity group does not need to remove any further NAL units to produce a conforming VVC bitstream. Tracks belonging to an operation point entity group shall have a track reference of type 'oref' pointing to the group_id indicated in the operation point entity group.
[0226] All entity_id values contained in an operating point entity group shall belong to the same VVC bitstream.When present, the OperatingPointGroupBox shall be contained in the GroupsListBox in the movie-level MetaBox, but not in the file-level or track-level MetaBox.
[0227] 3.3.3.2 Syntax
[0228]
[0229]
[0230] 3.3.3.3 Semantics
[0231] num_profile_tier_level_minus1 plus 1 gives the number of profile, grade, and level combinations and related fields below.
[0232] opeg_ptl[i] specifies the i-th tier, level and profile structure.
[0233] num_operating_points: Number of operating points for which the following information is given.
[0234] output_layer_set_idx is the index of the output layer set defining the operating point. The mapping between output_layer_set_idx and layer_id values shall be the same as specified by the VPS for the output layer set with index output_layer_set_idx.
[0235] ptl_idx: Signalled from zero for the output layer set with index output_layer_set_idx the index of the listed tier, level and profile structure.
[0236] max_temporal_id: Gives the maximum Temporalld of the NAL units of this operating point.
[0237] NOTE: The maximum Temporalld value indicated in the layer information sample group has a different semantics than the maximum Temporalld indicated here. However, they can have the same literal value.
[0238] layer_count: This field indicates the number of necessary layers for this operating point as defined in ISO / IEC 23090-3.
[0239] layer_id: Provides the nuh_layer_id values of the layers of the operating point.
[0240] is_outputlayer: Flag indicating whether a layer is an output layer. One indicates one output layer.
[0241] frame_rate_info_flag equal to 0 indicates that there is no frame rate information for the operating point. Value 1 indicates that there is frame rate information for the operating point.
[0242] bit_rate_info_flag equal to 0 indicates that there is no bit rate information for the operating point. Value 1 indicates that there is bit rate information for the operating point.
[0243] avgFrameRate gives the average frame rate of the operating point in frames / (256 seconds). Value 0 indicates that the average frame rate is not specified.
[0244] constantFrameRate equal to 1 indicates that the stream of the operation point has constant frame rate. Value 2 indicates that each timed layer's representation in the stream of the operation point has constant frame rate. Value 0 indicates that the stream of the operation point can or can not have constant frame rate.
[0245] maxBitRate gives the maximum bit rate of the stream of the operation point in any 1 second window, in bits per second.
[0246] avgBitRate gives the average bit rate of the stream of the operation point, in bits per second.
[0247] entity_count specifies the number of tracks present in the operation point.
[0248] entity_idx specifies the index of the list of entity_id belonging to the operation point in the entity group.
[0249] 4. Examples of technical problems solved by the disclosed technical solutions
[0250] The latest design of VVC video file format on the signaling of picture format information has the following problems:
[0251] 1) VvcDecoderConfigurationRecord includes optional signaling of chroma format and bit depth, but does not include signaling of picture width and picture height, and the operation point information 'vopi' sample group entry and operation point entity group 'opeg' box do not include any of these parameters.
[0252] However, when PTL is signaled at a certain location, chroma format, bit depth, and picture width and picture height should also be signaled as additional capability indication.
[0253] Note that the width and height fields of the visual sample entry are the cropped frame width and height. Therefore, unless the cropping window offset is all zero and the picture is a frame whose width and height values will not be the same as the picture width and height of the decoded picture.
[0254] Currently, the following situations can occur:
[0255] a. Single-layer bitstream contained only in one VVC track, without 'oref' track reference. Therefore, VvcPtlRecord should appear in the decoder configuration record. However, in this case, there is no signaling of part or all of chroma format, bit depth, picture width and picture height in any of the sample entry, 'vopi' sample group entry or 'opeg' entity group box.
[0256] b. A multi-layer bitstream stored in multiple tracks, the operation point information (including PTL information of each operation point) of which is stored in 'vopi' sample group entry or 'opeg' entity group box, while chroma format, bit depth, picture width and picture height are not signaled in sample entry, 'vopi' sample group entry or 'opeg' entity group box.
[0257] 2) The parameter chroma_format_idc in VvcDecoderConfigurationRecord should be used for capability indication, not for decoding, as the parameter set itself is sufficient for decoding. Even for decoding, not only chroma_format_idc in SPS is needed, but also vps_ols_dpb_chroma_format[] for multi-layer OLS. So, actually the maximum dpb_chroma_format should be signaled here, which is not the case in the current design. The same is true for the corresponding bit depth, picture width and picture height parameters.
[0258] 3) It is specified that when ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[output_layer_set_idx]. There are two problems (similar for the corresponding bit depth parameter):
[0259] a. The value of vps_ols_dpb_chroma_format[] can be different for different CVSs. So, it is needed to require that the value is the same for all VPSs, or specify that the value is equal to or greater than the maximum value.
[0260] b. The index value idx of ps_ols_dpb_chroma_format[idx] is the index of the multi-layer OLS list, so it is incorrect to use output_layer_set_idx directly, which is the index of all OLS lists.
[0261] 5. List of technical solutions
[0262] To solve the above problems, the methods summarized as follows are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these items can be applied individually or in any way combined.
[0263] 1) When ptl_present_flag is equal to 1, chroma_format_idc and bit_depth_minus8 are signaled in VvcDecoderConfigurationRecord, when ptl_present_flag is equal to 0, they are not signaled in VvcDecoderConfigurationRecord.
[0264] 2) When the VVC stream to which the configuration record applies is a single-layer bitstream, for all SPS referred to by VCL NAL units in the sample described by the current sample entry, the value of sps_chroma_format_idc shall be the same and the value of chroma_format_idc shall be equal to sps_chroma_format_idc.
[0265] 3) When the VVC stream to which the configuration record applies is a multi-layer bitstream, for all CVS described by the current sample entry, the value of chroma_format_idc shall be equal to the maximum of vps_ols_dpb_chroma_format[output_layer_set_idx] applied to the OLS identified by output_layer_set_idx.
[0266] a. Alternatively, change the above “equal to” to “equal to or greater than”.
[0267] 4) When the VVC stream to which the configuration record applies is a single-layer bitstream, for all SPS referred to by VCL NAL units in the sample described by the current sample entry, the value of sps_bitdepth_minus8 shall be the same and the value of bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0268] 5) When the VVC stream to which the configuration record applies is a multi-layer bitstream, for all CVS described by the current sample entry, the value of bit_depth_minus8 shall be equal to the maximum of vps_ols_dpb_bitdepth_minus8[output_layer_set_idx] applied to the OLS identified by output_layer_set_idx.
[0269] a. Alternatively, change the above “equal to” to “equal to or greater than”.
[0270] 6) Add signaling of picture width and picture height in VvcDecoderConfigurationRecord, as well as chroma format idc and bit depth minus 8. And both picture width and picture height fields are signaled using 16 bits.
[0271] a. Alternatively, both picture width and picture height fields are signaled using 24 bits.
[0272] b. Alternatively, both picture width and picture height fields are signaled using 32 bits.
[0273] c. Alternatively, in addition, when the VVC stream to which the configuration record applies is a single-layer bitstream, when both cropping window offsets are zero and the picture is a frame, signaling of picture width and picture height fields can be skipped.
[0274] 7) In VvcOperatingPointsRecord and OperatingPointGroupBox, add signaling of chroma format idc, bit depth minus 8, picture width and picture height for each operating point, e.g., immediately after ptl idx, with similar semantics and constraints as above when present in VvcDecoderConfigurationRecord.
[0275] 6. Embodiments
[0276] Below are some example embodiments of the invention aspects summarized in Section 5 above, which can be applied to the standard specification of VVC video file format. The changed text is based on the latest draft specification in MPEG output document N19454, “Information technology - Coding of audio-visual objects - Part 15: Network abstraction layer (NAL) unit structured video - Amendment 2: Transport of VVC and EVC in ISOBMFF”, July 2020”. Most of the added or modified relevant parts are highlighted in bold and italic, and some of the deleted parts are highlighted in double brackets (e.g., [[a]] means the deleted character “a”). There can be some other editorial changes that are not highlighted.
[0277] 6.1 First embodiment
[0278] This example applies to items 1 to 7.
[0279] 6.1.1 Decoder configuration information
[0280] 6.1.1.1 VVC decoder configuration record
[0281] 6.1.1.1.1 Definition
[0282] This clause specifies the decoder configuration information for ISO / IEC 23090-3 video content. ...
[0284] The following constraint applies to chroma_format_idc:
[0285] - If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 shall be identical in all SPS referenced by the VCL NAL units of the [[track of the]] sample in which the current sample entry description applies, and the value of chroma_format_idc shall be equal to sps_chroma_format_idc.
[0286] - Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), for all CVS to which the current sample entry description applies, [[if ptl_present_flag is equal to 1]], the value of chroma_format_idc shall be equal to the maximum value of vps_ols_dpb_chroma_format[MultiLayerOlsIdx[output_layer_set_idx]] as defined in ISO / IEC 23090-3 that applies to the OLS identified by output_layer_set_idx.
[0287] - [[Otherwise, chroma_format_idc shall not be present.]]
[0288] The following constraint applies to bit_depth_minus8:
[0289] - If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 shall be identical in all SPS referenced by the VCL NAL units of the [[track of the]] sample in which the current sample entry description applies, and the value of bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0290] - Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), for all CVS to which the current sample entry description applies, [[if ptl_present_flag is equal to 1]], the value of bit_depth_minus8 shall be equal to the maximum value of vps_ols_dpb_bitdepth_minus8[ MultiLayerOlsIdx[ output_layer_set_idx ]] defined in ISO / IEC 23090-3 that applies to the OLS identified by output_layer_set_idx.
[0291] - [[Otherwise, bit_depth_minus8 shall not be present.]]
[0292] The following constraint applies to picture_width:
[0293] - If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_pic_width_max_in_luma_samples defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the sample to which the current sample entry description applies, and the value of picture_width shall be equal to sps_pic_width_max_in_luma_samples.
[0294] - Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), for all CVS to which the current sample entry description applies, the value of picture_width shall be equal to the maximum value of vps_ols_dpb_pic_width[ MultiLayerOlsIdx[ output_layer_set_idx ]] defined in ISO / IEC 23090-3 that applies to the OLS identified by output_layer_set_idx.
[0295] The following constraint applies to picture_height:
[0296] - If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_pic_height_max_in_luma_samples defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the sample to which the current sample entry description applies, and the value of picture_height shall be equal to sps_pic_height_max_in_luma_samples.
[0297] - Otherwise (the VVC stream for which the application configuration record applies is a multi-layer bitstream), the value of picture height shall be equal to the maximum value of vps_ols_dpb_pic_height[ MultiLayerOlsIdx[ output_layer_set_idx ] ] defined in ISO / IEC 23090-3 applied to the OLS identified by output_layer_set_idx for all CVS for which the current sample entry description applies. ...
[0299] 6.1.1.1.2 Syntax
[0300]
[0301]
[0302] 6.1.1.1.3 Semantics ...
[0304] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to a particular output layer set. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular output layer set, but can contain a VVC bitstream corresponding to multiple output layer sets or can contain one or more individual layers that do not form an output layer set or individual sub-layers that do not include sub-layers with Temporalld equal to 0.
[0305] track_ptl specifies the profile, tier and level of the output layer set represented by the VVC bitstream contained in the track.
[0306] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided to the VVC decoder by external means for decoding the bitstream contained in the track as specified in ISO / IEC 23090-3. When ptl_present_flag is equal to 1 and output_layer_set_idx is not present, its value is inferred to be equal to the OLS index of the OLS that contains only the layers carried in the VVC track (after parsing the reference VVC track or VVC subpicture track, if any).
[0307] [[chroma_format_present_flag equal to 0 specifies that chroma_format_idc is not present. chroma_format_present_flag equal to 1 specifies that chroma_format_idc is present. ]]
[0308] chroma_format_idc indicates the chroma format that applies to this track. [[The following constraints apply to chroma_format_idc:
[0309] - If the value of sps_chroma_format_idc defined in ISO / IEC 23090-3 is the same in all SPSs referenced by NAL units of a track, chroma_format_idc shall be equal to sps_chroma_format_idc.
[0310] - Otherwise, if ptl_present_flag is equal to 1, chroma_format_idc shall be equal to vps_ols_dpb_chroma_format[output_layer_set_idx] defined in ISO / IEC 23090-3.
[0311] - Otherwise, chroma_format_idc will not be present.
[0312] bit_depth_present_flag equal to 0 specifies that bit_depth_minus8 is not present. bit_depth_present_flag equal to 1 specifies that bit_depth_minus8 is present.
[0313] bit_depth_minus8 indicates the bit depth that applies to this track. [[The following constraints apply to bit_depth_minus8:
[0314] - If the value of sps_bitdepth_minus8 defined in ISO / IEC 23090-3 is the same in all SPSs referenced by NAL units of a track, bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0315] - Otherwise, if ptl_present_flag is equal to 1, bit_depth_minus8 shall be equal to vps_ols_dpb_bitdepth_minus8[output_layer_set_idx] defined in ISO / IEC 23090-3.
[0316] - Otherwise, bit_depth_minus8 shall not be present.
[0317] picture_width indicates the maximum picture width, in luma samples, that applies to this track.
[0318] picture_height indicates the maximum picture height, in luma samples, that applies to this track.
[0319] numAnumArrays indicates the number of arrays of NAL units of the indicated type.
[0320] When array_completeness is equal to 1, it indicates that all NAL units of the given type are in the following array and that no additional NAL units of the indicated type are in the stream; when array_completeness is equal to 0, it indicates that additional NAL units of the indicated type can be in the stream; the value is allowed to be constrained by the sample entry name.
[0321] NAL_unit_type indicates the type of NAL units (all shall be of this type) in the following array; it takes the values defined in ISO / IEC 23090-3; it is limited to take one of the values indicating a DCI, VPS, SPS, PPS, prefix APS, or prefix SEI NAL unit.
[0322] numNalus indicates the number of NAL units of the indicated type contained in the configuration record of the stream to which this configuration record applies. SEI arrays shall contain only SEI messages of a “declarative” nature, i.e. those providing information about the entire stream. One example of such SEI can be the user data SEI.
[0323] nalUnitLength indicates the byte length of the NAL unit.
[0324] nalUnit contains a DCI, VPS, SPS, PPS, APS, or declarative SEI NAL unit as specified in ISO / IEC 23090-3.
[0325] 6.1.2 Operation point information sample group
[0326] 6.1.2.1 Definition ...
[0328] 6.1.2.2 Syntax
[0329]
[0330]
[0331] 6.1.2.3 Semantics ...
[0333] num_operating_points: Number of operating points for which the following information is given.
[0334] output_layer_set_idx is the index of the output layer set defining the operating point. The mapping between output_layer_set_idx and layer_id values shall be the same as specified by the VPS for the output layer set with index output_layer_set_idx.
[0335] ptl_idx: Signalled from zero for the output layer set with index output_layer_set_idx the index of the listed profile, level and tier structure.
[0336] chroma_format_idc indicates the chroma format applied for this operating point. The following constraint applies to chroma_format_idc:
[0337] - If this operating point contains only one layer, the value of sps_chroma_format_idc as defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operating point, and the value of chroma_format_idc shall be equal to sps_chroma_format_idc.
[0338] - Otherwise (this operating point contains more than one layer), the value of chroma_format_idc shall be equal to the value of vps_ols_dpb_chroma_format[ MultiLayerOlsIdx[ output_layer_set_idx ] ] as defined in ISO / IEC 23090-3.
[0339] bit_depth_minus8 indicates the bit depth applied for this operating point. The following constraint applies to bit_depth_minus8:
[0340] - If this operating point contains only one layer, the value of sps_bitdepth_minus8 as defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operating point, and the value of bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0341] - Otherwise (this operation point contains more than one layer), the value of picture width shall be equal to the value of vps_ols_dpb_pic_width[ MultiLayerOlsIdx[ output_layer_set_idx ] ] as defined in ISO / IEC 23090-3.
[0342] picture width indicates the maximum picture width, in luma samples, that applies to this operation point. The following constraint applies to picture width:
[0343] - If this operation point contains only one layer, the value of sps_pic_width_max_in_luma_samples as defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operation point, and the value of picture width shall be equal to sps_pic_width_max_in_luma_samples.
[0344] - Otherwise (this operation point contains more than one layer), the value of picture width shall be equal to the value of vps_ols_dpb_pic_width[ MultiLayerOlsIdx[ output_layer_set_idx ] ] as defined in ISO / IEC 23090-3.
[0345] picture height indicates the maximum picture height, in luma samples, that applies to this operation point. The following constraint applies to picture height:
[0346] - If this operation point contains only one layer, the value of sps_pic_height_max_in_luma_samples as defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operation point, and the value of picture height shall be equal to sps_pic_height_max_in_luma_samples.
[0347] - Otherwise (this operation point contains more than one layer), the value of picture height shall be equal to the value of vps_ols_dpb_pic_height[ MultiLayerOlsIdx[ output_layer_set_idx ] ] as defined in ISO / IEC 23090-3.
[0348] max_temporal_id: Gives the maximum Temporalld of the operation point NAL units.
[0349] NOTE The maximum Temporalld value indicated in the layer information sample group has a different semantics than the maximum Temporalld indicated here. However, they can have the same literal value. ...
[0351] 6.1.3 Operation point entity group
[0352] 6.1.3.1 Overview
[0353] An operation point entity group is defined to provide the mapping of tracks to operation points as well as the tier level information of the operation points.
[0354] The implicit reconstruction process of aggregating the track samples mapped to the operation points described in this entity group does not require removal of any further NAL units to produce a conforming VVC bitstream. The tracks belonging to the operation point entity group shall have track references of type 'oref' pointing to the group_id indicated in the operation point entity group.
[0355] All entity_id values contained in the operation point entity group shall belong to the same VVC bitstream. When present, the OperatingPointGroupBox shall be contained in the GroupsListBox in the movie-level MetaBox and shall not be contained in the file-level or track-level MetaBox.
[0356] 6.1.3.2 Syntax
[0357]
[0358] 6.1.3.3 Semantics ...
[0360] num_operating_points: Gives the number of operation points the following information is targeted at.
[0361] output_layer_set_idx is the index of the output layer set defining the operation point. The mapping between output_layer_set_idx and layer_id values shall be the same as specified by the VPS for the output layer set with index output_layer_set_idx.
[0362] ptl_idx: Signaled for the output layer set with index output_layer_set_idx is the zero-based index of the listed tier, level and rank structure.
[0363] chroma_format_idc indicates the chroma format applied to this operation point. The following constraint applies to chroma_format_idc:
[0364] - If this operation point contains only one layer, the value of sps_chroma_format_idc defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operation point, and the value of chroma_format_idc shall be equal to sps_chroma_format_idc.
[0365] - Otherwise (this operation point contains more than one layer), the value of chroma_format_idc shall be equal to the value of vps_ols_dpb_chroma_format[ MultiLayerOlsIdx[ output_layer_set_idx ] ] defined in ISO / IEC 23090-3.
[0366] bit_depth_minus8 indicates the bit depth applied to this operation point. The following constraint applies to bit_depth_minus8:
[0367] - If this operation point contains only one layer, the value of sps_bitdepth_minus8 defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of this operation point, and the value of bit_depth_minus8 shall be equal to sps_bitdepth_minus8.
[0368] - Otherwise (this operation point contains more than one layer), the value of bit_depth_minus8 shall be equal to the value of vps_ols_dpb_bitdepth_minus8[ MultiLayerOlsIdx[ output_layer_set_idx ] ] defined in ISO / IEC 23090-3.
[0369] picture_width indicates the maximum picture width, in luma samples, applied to this operation point. The following constraint applies to picture_width:
[0370] - If the operation point contains only one layer, the value of sps_pic_width_max_in_luma_samples defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of the operation point, and the value of picture_width shall be equal to sps_pic_width_max_in_luma_samples.
[0371] - Otherwise (this operation point contains more than one layer), the value of picture_width shall be equal to the value of vps_ols_dpb_pic_width[ MultiLayerOlsIdx[ output_layer_set_idx ] ] defined in ISO / IEC 23090-3.
[0372] picture_height indicates the maximum picture height in luma samples that applies to the operation point. The following constraint applies to picture_height:
[0373] - If the operation point contains only one layer, the value of sps_pic_height_max_in_luma_samples defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the VVC bitstream of the operation point, and the value of picture_height shall be equal to sps_pic_height_max_in_luma_samples.
[0374] - Otherwise (this operation point contains more than one layer), the value of picture_height shall be equal to the value of vps_ols_dpb_pic_height[ MultiLayerOlsIdx[ output_layer_set_idx ] ] defined in ISO / IEC 23090-3.
[0375] max_temporal_id: gives the maximum Temporalld of the NAL units of the operation point.
[0376] NOTE: The maximum Temporalld value indicated in the layer information sample group has a different semantics than the maximum Temporalld indicated here. However, they can have the same literal value. ...
[0378] Figure 1is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0379] System 1900 can include a codec component 1904, which can implement various coding or encoding methods described in this document. Codec component 1904 can reduce the average bitrate of a video from input 1902 to the output of codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via a communication connected by component 1906. Component 1908 can use a stored or communicated bitstream (or coded) representation of the video received at input 1902 to generate pixel values or displayable video sent to a display interface 1910. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it should be understood that a coding tool or operation is used at an encoder, and a decoder will perform a corresponding decoding tool or operation that reverses the result of the coding.
[0380] Examples of peripheral bus interfaces or display interfaces can include a universal serial bus (USB) or a high-definition multimedia interface (HDMI) or display port, etc. Examples of storage interfaces include SATA (serial advanced technology attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smartphones, or other devices that are capable of performing digital data processing and / or video display.
[0381] Figure 2is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more methods described herein. The device 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory (memories) 3604 can be used for storing data and code used during operation of the present techniques. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in the present document. In some embodiments, the video processing hardware 3606 can be included at least in part within the processor(s) 3602 (e.g., a graphics co-processor).
[0382] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0383] As shown in Figure 4 , the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, and can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and can be referred to as a video decoding device.
[0384] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0385] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the I / O interface 116 via the network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.
[0386] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0387] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120 which is configured to interface with an external display device.
[0388] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0389] Figure 5 is shown to illustrate that the video encoder 114 can be a component of the source device 110. Figure 4 A block diagram of an example of a video encoder 200 of the video encoder 114 in the system 100 shown in FIG. 1.
[0390] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 In an example, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the functional components of the video encoder 200. In some examples, the processor(s) can be configured to perform any or all of the techniques described in this disclosure.
[0391] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202 which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0392] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0393] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately in Figure 5 for the sake of explanation.
[0394] The partition unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0395] The mode selection unit 203 can select a type of coding mode (intra or inter), for example, based on the error results, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode in which the prediction is based on both an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select a resolution of a motion vector (e.g., sub-pixel or integer pixel precision) for the block in the case of inter prediction.
[0396] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 (except for the picture associated with the current video block).
[0397] For example, the motion estimation unit 204 and the motion compensation unit 205 can perform different operations depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0398] In some examples, the motion estimation unit 204 can perform single directional prediction for the current video block, and the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of List 0 or List 1. The motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0399] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 and also search for another reference video block for the current video block in a reference picture in list 1. The motion estimation unit 204 can then generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0400] In some examples, the motion estimation unit 204 can output a full set of motion information for a decoding process of a decoder.
[0401] In some examples, the motion estimation unit 204 can not output a full set of motion information for a current video. Instead, the motion estimation unit 204 can signal motion information for a current video block with reference to motion information for another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.
[0402] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0403] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector for the current video block and a motion vector for the indicated video block. The video decoder 300 can use the motion vector for the indicated video block and the motion vector difference to determine the motion vector for the current video block.
[0404] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0405] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0406] Residual generation unit 207 can generate residual data for a current video block by subtracting (e.g., indicated by the minus sign) a prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0407] In other examples, such as in a skip mode, there can be no residual data for the current video block, and residual generation unit 207 can not perform a subtraction operation.
[0408] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0409] Quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block after transform processing unit 208 generates the transform coefficient video blocks associated with the current video block.
[0410] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to a transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0411] Loop filtering operations can be performed to reduce video block artifacts in the video blocks after reconstruction unit 212 reconstructs the video blocks.
[0412] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0413] Figure 6 is shown to illustrate that the video decoder 300 of the video decoder 114 in the system 100 can be Figure 4 A block diagram of an example of a video decoder 300 of the video decoder 114 shown in the system 100.
[0414] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 In examples, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0415] In Figure 6 In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform a decoding pass generally reciprocal to the encoding pass described with respect to video encoder 200. Figure 5
[0416] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data and motion compensation unit 302 can determine, from the entropy decoded video data, motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, motion compensation unit 302 can determine such information by performing AMVP and merge mode.
[0417] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter used at sub-pixel precision can be included in syntax elements.
[0418] Motion compensation unit 302 can use an interpolation filter used by video encoder 200 during video block encoding to calculate interpolated values for sub-integer pixels of a reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from received syntax information and use the interpolation filter to generate a prediction block.
[0419] Motion compensation unit 302 can use some syntax information to determine the size of blocks used to encode frames and / or slices of a coded video sequence, partition information describing how to partition each macroblock of a picture of the coded video sequence, modes indicating how to encode each partition, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.
[0420] Intra-prediction unit 303 can use intra-prediction modes, e.g., received in the bitstream, to form a prediction block from spatial neighboring blocks. Inverse quantization unit 303 inverse quantizes (i.e., de-quantizes) quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transformation unit 303 applies an inverse transform.
[0421] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.
[0422] A list of preferred solutions for some embodiments is provided next.
[0423] A first set of solutions is provided below. The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, item 1).
[0424] 1. A method for processing visual media (e.g., Figure 3 ), comprising performing a conversion (702) between visual media data and a file storing a bitstream representation of the visual media data according to a format rule; wherein the format rule specifies: a first record indicating whether a profile-grade-level is indicated in the file controls whether a second record indicating a chroma format of the visual media data and / or a third record indicating a bit depth used to represent the visual media data is included in the file.
[0425] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 2, 4).
[0426] 1. A method for processing visual media, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the bitstream representation is a single-layer bitstream; wherein the format rules specify constraints on the single-layer bitstream stored in the file.
[0427] 2. The method according to solution 1, wherein the constraint is that one or more chroma format values indicated in one or more sequence parameter sets referenced by the video codec layer network abstraction layer unit included in the file sample are equal.
[0428] 3. The method according to solution 1, wherein the constraint is that one or more bit depth values indicated in one or more sequence parameter sets referenced by the video codec layer network abstraction layer unit included in the file sample are equal.
[0429] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 3, 5).
[0430] 1. A method of processing visual media, comprising performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to a format rule; wherein the bitstream representation is a multi-layer bitstream; wherein the format rule specifies a constraint on the multi-layer bitstream stored in the file.
[0431] 2. The method according to solution 1, wherein the constraint is that a value of chroma format is set equal to a maximum value of chroma format identified in sample entry descriptions of output layer sets of all coded video sequences to which the sample entry descriptions apply.
[0432] 3. The method according to solution 1, wherein the constraint is that a value of bit depth is set equal to a maximum value of bit depth identified in sample entry descriptions of output layer sets of all coded video sequences to which the sample entry descriptions apply.
[0433] 8. The method according to any of solutions 1-7, wherein the conversion comprises generating the bitstream representation of the visual media data according to the format rule and storing the bitstream representation into the file.
[0434] 9. The method according to any of solutions 1-7, wherein the conversion comprises parsing the file according to the format rule to recover the visual media data.
[0435] 10. A video decoding apparatus comprising a processor configured to implement one or more of the methods described in solutions 1 to 9.
[0436] 11. A video encoding apparatus comprising a processor configured to implement one or more of the methods described in solutions 1 to 9.
[0437] 12. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement the method described in any of solutions 1 to 9.
[0438] 13. A computer readable medium having stored thereon a bitstream generated according to any of solutions 1 to 9 in compliance with a file format.
[0439] 14. The methods, apparatuses or systems described in this document. In the solutions described herein, an encoder can comply with a format rule by generating a coded representation according to the format rule. In the solutions described herein, a decoder can use a format rule to parse syntax elements in a coded representation according to the format rule, with knowledge of the presence and absence of syntax elements, to produce a decoded video.
[0440] A second set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., items 1-5).
[0441] 1. A method of processing visual media data (e.g., method 800 as shown in Figure 8 FIG. 8), comprising performing a conversion between the visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the format rule specifies whether a first element indicates whether a track contains a bitstream corresponding to a particular output layer set and / or whether a second element indicating a chroma format of the track and / or a third element indicating bit depth information of the track is included in a configuration record of the track (802).
[0442] 2. The method according to solution 1, wherein the format rule specifies that the second element and / or the third element is included in response to the first element indicating that the track contains the bitstream corresponding to the particular output layer set.
[0443] 3. The method according to solution 1, wherein the format rule specifies that the second element and / or the third element is omitted in response to the first element indicating that the track is allowed not to contain the bitstream corresponding to the particular output layer set.
[0444] 4. The method according to solution 1, wherein the format rule further specifies a syntax constraint depending on whether a bitstream to which the configuration record is applied is a multi-layer bitstream.
[0445] 5. The method according to solution 4, wherein the format rule further specifies that, in response to the bitstream not being the multi-layer bitstream, the syntax constraint is that one or more chroma format values indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in samples of the visual media file described by a sample entry to which the configuration record is applied are equal.
[0446] 6. The method according to solution 5, wherein the format rule further specifies that the chroma format value indicated in the configuration record is equal to the one or more chroma format values.
[0447] 7. The method according to solution 4, wherein the format rule further specifies that, in response to the bitstream being the multi-layer bitstream, the syntax constraint is that a value of a chroma format indicated in the configuration record is set to be equal to a maximum value of chroma formats indicated in video parameter sets and applied to an output layer set identified by an output layer set index among all chroma format values indicated in all video parameter sets of all coded video sequences described by a sample entry to which the configuration record is applied.
[0448] 8. The method according to solution 4, wherein the format rule further specifies that, in response to the bitstream not being the multi-layer bitstream, one or more bit-depth information values indicated in one or more sequence parameter sets referred by NAL (Network Abstraction Layer) unit references included in samples of the media file to which the sample entry description of the configuration record applies are equal.
[0449] 9. The method according to solution 8, wherein the format rule further specifies that the bit-depth information value indicated in the configuration record is equal to the one or more bit-depth information values.
[0450] 10. The method according to solution 4, wherein the format rule further specifies that, in response to the bitstream being the multi-layer bitstream, the syntax constraint is that the bit-depth information value indicated in the configuration record is set equal to a maximum value of bit-depth information indicated in video parameter sets and applied to output layer sets identified by output layer set indexes among all bit-depth information values indicated in all video parameter sets of all coded video sequences to which the sample entry description of the configuration record applies.
[0451] 11. The method according to any of solutions 1-10, wherein the converting comprises generating the media file according to the format rule and storing the one or more bitstreams to the media file.
[0452] 12. The method according to any of solutions 1-10, wherein the converting comprises parsing the media file according to the format rule to reconstruct the one or more bitstreams.
[0453] 13. An apparatus for processing visual media data, the apparatus comprising a processor configured to implement a method comprising performing a conversion between visual media data and a media file comprising one or more tracks storing one or more bitstreams of visual media data according to a format rule, wherein the format rule specifies that a presence of a chroma format syntax element and / or a bit-depth syntax element and / or a syntax constraint on the chroma format syntax element and / or the bit-depth syntax element depends on (1) whether a track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether a bitstream to which a configuration record applies is a multi-layer bitstream.
[0454] 14. The apparatus according to solution 13, wherein the format rule specifies that, in case a track contains a particular bitstream corresponding to a particular output layer set, a chroma format syntax element and / or a bit-depth syntax element is included in a configuration record of the track.
[0455] 15. The apparatus according to solution 13, wherein the format rule specifies that chroma format syntax elements and / or bit depth syntax elements are omitted from the configuration record of a track in case the track is allowed to not contain a bitstream corresponding to a particular output layer set.
[0456] 16. The apparatus according to solution 13, wherein the format rule specifies that, in response to the bitstream not being a multi-layer bitstream, the syntax constraint is that values of one or more chroma format syntax elements indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in a sample of the visual media file described by the sample entry to which the configuration record applies are equal.
[0457] 17. The apparatus according to solution 13, wherein the format rule specifies that, in response to the bitstream being a multi-layer bitstream, the syntax constraint is that a value of a chroma format indicated in the configuration record is set equal to a maximum value of all chroma format values indicated in all video parameter sets of all coded video sequences described by the sample entry to which the configuration record applies and applied to an output layer set identified by an output layer set index.
[0458] 18. The apparatus according to solution 13, wherein the format rule specifies, in response to the bitstream not being a multi-layer bitstream, that the syntax constraint is that values of one or more bit depth syntax elements indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in a sample of the visual media file described by the sample entry to which the configuration record applies are equal.
[0459] 19. The apparatus according to solution 13, wherein the format rule specifies that, in response to the bitstream being a multi-layer bitstream, the syntax constraint is that a value of a bit depth syntax element is set equal to or greater than a maximum value of all bit depth information values indicated in all video parameter sets of all coded video sequences described by the sample entry to which the configuration record applies and applied to an output layer set identified by an output layer set index.
[0460] 20. A non-transitory computer-readable recording medium storing instructions causing a processor to perform a conversion between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule, wherein the format rule specifies that a presence of chroma format syntax elements and / or bit depth syntax elements and / or a syntax constraint on the chroma format syntax elements and / or the bit depth syntax elements depends on (1) whether a track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether a bitstream to which a configuration record applies is a multi-layer bitstream.
[0461] 21. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing apparatus, wherein the method comprises generating a visual media file comprising one or more tracks storing one or more bitstreams of visual media data according to a format rule, wherein the format rule specifies that a presence of and / or a syntax constraint on a chroma format syntax element and / or a bit depth syntax element depends on (1) whether a track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether a bitstream to which a configuration record applies is a multi-layer bitstream.
[0462] 22. A video processing apparatus comprising a processor configured to implement a method recited by any one or more of solutions 1 to 12.
[0463] 23. A method of storing visual media data into a file comprising one or more bitstreams, the method comprising a method recited by any one of solutions 1 to 12, and further comprising storing the one or more bitstreams to a non-transitory computer-readable recording medium.
[0464] 24. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited by any one or more of solutions 1 to 12.
[0465] 25. A computer-readable medium storing a bitstream generated according to any of the methods above.
[0466] 26. A video processing apparatus configured to implement a method recited by any one or more of solutions 1 to 12 for storing a bitstream.
[0467] 27. A computer-readable medium having stored thereon a bitstream generated according to any of solutions 1 to 12 and conforming to a file format.
[0468] 28. A method, apparatus or system described in the present document.
[0469] A third set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., item 6).
[0470] 1. A method of processing visual media data (e.g., as in Figure 9The illustrated method 900 includes performing a conversion between a visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule, wherein the format rule specifies whether a first element indicating a picture width of the track and / or a second element indicating a picture height of the track is included in a configuration record of the track is based on (1) a third element indicating whether the track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether the configuration record is for a single-layer bitstream, and
[0471] wherein the format rule further specifies that when the first element and / or the second element is included in the configuration record of the track, the first element and / or the second element is represented in a field comprising 16 bits.
[0472] 2. The method according to solution 1, wherein the format rule specifies that the first element and / or the second element is included in response to the third element indicating that the track contains the particular bitstream corresponding to the particular output layer set.
[0473] 3. The method according to solution 1, wherein the format rule specifies that the first element and / or the second element is omitted in response to the third element indicating that the track is allowed not to contain the particular bitstream corresponding to the particular output layer set.
[0474] 4. The method according to solution 1, wherein the field comprises 24 bits.
[0475] 5. The method according to solution 1, wherein the field comprises 32 bits.
[0476] 6. The method according to solution 1, wherein the format rule specifies that the first element and / or the second element is omitted in response to the bitstream being the single-layer bitstream, a cropping window offset being all zeros and a picture being a frame.
[0477] 7. The method according to solution 1,
[0478] wherein the format rule further specifies a syntax constraint on a value of the first element and / or the second element based on whether the bitstream to which the configuration record applies is a single-layer bitstream.
[0479] 8. The method according to solution 7, wherein the format rule further specifies that in response to the bitstream being the single-layer bitstream, the syntax constraint is that one or more picture width values indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in samples of a visual media file described by a sample entry to which the configuration record applies are equal.
[0480] 9. The method according to solution 8, wherein the format rule further specifies that a value of the first element storing the track of the bitstream is equal to the one or more picture width values.
[0481] 10. The method according to solution 7, wherein the format rule further specifies that, in response to the bitstream being the single-layer bitstream, the syntax constraint is that one or more picture height values indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in samples of the visual media file described by the sample entry to which the configuration record applies are equal.
[0482] 11. The method according to solution 10, wherein the format rule further specifies that a value of the second element storing the track of the bitstream is equal to the one or more picture height values.
[0483] 12. The method according to solution 7, wherein the format rule further specifies that, in response to the bitstream not being the single-layer bitstream, the syntax constraint is that a value of the first element is set equal to a maximum of picture width values indicated in all video parameter sets of all coded video sequences described by the sample entry to which the configuration record applies and applied to picture widths of output layer sets identified by output layer set indexes.
[0484] 13. The method according to solution 7, wherein the format rule further specifies that, in response to the bitstream not being the single-layer bitstream, the syntax constraint is that a value of the second element is set equal to a maximum of picture height values indicated in all video parameter sets of all coded video sequences described by the sample entry to which the configuration record applies and applied to picture heights of output layer sets identified by output layer set indexes.
[0485] 14. The method according to any of solutions 1-13, wherein the converting comprises generating the visual media file and storing the one or more bitstreams to the visual media file according to the format rule.
[0486] 15. The method according to any of solutions 1-13, wherein the converting comprises parsing the visual media file to reconstruct the one or more bitstreams according to the format rule.
[0487] 16. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising performing a conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the format rule specifies whether a first element indicating a picture width of a track and / or a second element indicating a picture height of the track is included in a configuration record of the track is based on (1) a third element indicating whether the track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether the configuration record is for a single-layer bitstream, and wherein the format rule further specifies that when the first element and / or the second element is included in the configuration record of the track, the first element and / or the second element is represented in a field comprising 16 bits.
[0488] 17. The apparatus according to solution 16, wherein the format rule specifies that the first element and / or the second element is included in response to the third element indicating that the track contains the particular bitstream corresponding to the particular output layer set.
[0489] 18. The apparatus according to solution 16, wherein the format rule specifies that the first element and / or the second element is omitted in response to the third element indicating that the track is allowed not to contain the particular bitstream corresponding to the particular output layer set.
[0490] 19. The apparatus according to solution 16, wherein the format rule further specifies a syntax constraint on a value of the first element and / or the second element based on whether a bitstream to which the configuration record is applied is a single-layer bitstream.
[0491] 20. The apparatus according to solution 19, wherein the format rule further specifies that in response to the bitstream being the single-layer bitstream, the syntax constraint is that one or more picture width values indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in samples of a visual media file described by a sample entry to which the configuration record is applied are equal.
[0492] 21. The apparatus according to solution 19, wherein the format rule further specifies that in response to the bitstream being the single-layer bitstream, the syntax constraint is that one or more picture height values indicated in one or more sequence parameter sets referenced by NAL (Network Abstraction Layer) units included in samples of a visual media file described by a sample entry to which the configuration record is applied are equal.
[0493] 22. The apparatus according to solution 19, wherein the format rule further specifies that, in response to the bitstream not being the single-layer bitstream, the syntax constraint is that a value of the first element is set equal to a maximum of picture width values indicated in video parameter sets of all coded video sequences described by a sample entry to which the configuration record applies and applied to picture widths of output layer sets identified by output layer set indexes.
[0494] 23. The apparatus according to solution 19, wherein the format rule further specifies that, in response to the bitstream not being the single-layer bitstream, the syntax constraint is that a value of the second element is set equal to a maximum of picture height values indicated in video parameter sets of all coded video sequences described by a sample entry to which the configuration record applies and applied to picture heights of output layer sets identified by output layer set indexes.
[0495] 24. A non-transitory computer-readable recording medium storing instructions causing a processor to perform a conversion between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the format rule specifies that whether a first element indicating a picture width of a track and / or a second element indicating a picture height of the track is included in a configuration record of the track is based on (1) whether a third element indicates that the track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether the configuration record is for a single-layer bitstream, and wherein the format rule further specifies that, when the first element and / or the second element is included in the configuration record of the track, the first element and / or the second element is represented in a field including 16 bits.
[0496] 25. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing apparatus, wherein the method comprises performing a conversion between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the format rule specifies that whether a first element indicating a picture width of a track and / or a second element indicating a picture height of the track is included in a configuration record of the track is based on (1) whether a third element indicates that the track contains a particular bitstream corresponding to a particular output layer set and / or (2) whether the configuration record is for a single-layer bitstream, and wherein the format rule further specifies that, when the first element and / or the second element is included in the configuration record of the track, the first element and / or the second element is represented in a field including 16 bits.
[0497] 26. A video processing apparatus comprising a processor configured to implement any one or more of the methods recited in solutions 1 to 15.
[0498] 27. A method of storing visual media data into a file comprising one or more bitstreams, the method comprising any of the methods recited in solutions 1 to 15, and further comprising storing the one or more bitstreams to a non-transitory computer readable recording medium.
[0499] 28. A computer readable medium storing program code that when executed causes a processor to implement any one or more of the methods recited in solutions 1 to 15.
[0500] 29. A computer readable medium storing a bitstream generated according to any of the methods.
[0501] 30. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement any one or more of the methods recited in solutions 1 to 15.
[0502] 31. A computer readable medium having stored thereon a bitstream generated according to any of the solutions 1 to 15 in accordance with a file format.
[0503] A fourth set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., item 7).
[0504] 1. A method of processing visual media data (e.g., method 1000 as shown in Figure 10 FIG. 1), comprising performing a conversion between the visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file comprises operation point records and operation point group boxes, and wherein the format rule specifies that, for each operation point indicated in the visual media file, a first element indicating a chroma format, a second element indicating bit depth information, a third element indicating a maximum picture width, and / or a fourth element indicating a maximum picture height are included in the operation point record and the operation point group box.
[0505] 2. The method of solution 1, wherein the format rule further specifies that the first element, the second element, the third element, and / or the fourth element immediately follow a fifth element indicating a zero-based index of a tier, level, and scale structure of output layer sets identified by output layer set indexes.
[0506] 3. The method of solution 1, wherein the format rule further specifies a syntax constraint on a value of the first element, a value of the second element, a value of the third element, and / or a value of the fourth element applied to operation points associated with a bitstream based on whether the operation point contains only a single layer.
[0507] 4. The method of solution 3, wherein the format rule further specifies that, in response to the operation point containing the single layer, the syntax constraint is that one or more chroma format values indicated in one or more sequence parameter sets referenced by NAL (network abstraction layer) units in the bitstream of the operation point are equal.
[0508] 5. The method of solution 4, wherein the format rule further specifies that the value of the first element is equal to the one or more chroma format values.
[0509] 6. The method of solution 3, wherein the format rule further specifies that, in response to the operation point containing more than one layer, the syntax constraint is that the value of the first element is set equal to a chroma format value indicated in a video parameter set and applied to output layer sets identified by output layer set indexes.
[0510] 7. The method of solution 3, wherein the format rule further specifies that, in response to the operation point containing the single layer, the syntax constraint is that one or more bit depth information values indicated in one or more sequence parameter sets referenced by NAL (network abstraction layer) units in the bitstream of the operation point are equal.
[0511] 8. The method of solution 7, wherein the format rule further specifies that the value of the second element is equal to the one or more bit depth information values.
[0512] 9. The method of solution 3, wherein the format rule further specifies that, in response to the operation point containing more than one layer, the syntax constraint is that the value of the second element is set equal to a bit depth information value indicated in a video parameter set and applied to output layer sets identified by output layer set indexes.
[0513] 10. The method of solution 3, wherein the format rule further specifies that, in response to the operation point containing the single layer, the syntax constraint is that one or more picture width values indicated in one or more sequence parameter sets referenced by NAL (network abstraction layer) units in the bitstream of the operation point are equal.
[0514] 11. The method of solution 10, wherein the format rule further specifies that the value of the third element is equal to the one or more picture width values.
[0515] 12. The method according to solution 3, wherein the format rule further specifies that, in response to the operation point containing more than one layer, the syntax constraint is that a value of the third element is set equal to a picture width value indicated in a video parameter set and applied to an output layer set identified by an output layer set index.
[0516] 13. The method according to solution 3, wherein the format rule further specifies that, in response to the operation point containing the single layer, the syntax constraint is that one or more picture height values indicated in one or more sequence parameter sets referenced by NAL (network abstraction layer) units in the bitstream of the operation point are equal.
[0517] 14. The method according to solution 13, wherein the format rule further specifies that a value of the fourth element is equal to the one or more picture height values.
[0518] 15. The method according to solution 3, wherein the format rule further specifies that, in response to the operation point containing more than one layer, the syntax constraint is that a value of the fourth element is set equal to a picture height value indicated in a video parameter set and applied to an output layer set identified by an output layer set index.
[0519] 16. The method according to any of solutions 1-15, wherein the converting comprises generating a visual media file and storing the one or more bitstreams to the visual media file according to the format rule.
[0520] 17. The method according to any of solutions 1-15, wherein the converting comprises parsing the visual media file to reconstruct the one or more bitstreams according to the format rule.
[0521] 18. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising performing a conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of visual media data according to a format rule; wherein the visual media file comprises an operation point record and an operation point group box, and wherein the format rule specifies whether, for each operation point indicated in the visual media file, a first element indicating a chroma format, a second element indicating bit depth information, a third element indicating a maximum picture width, and / or a fourth element indicating a maximum picture height are included in the operation point record and the operation point group box.
[0522] 19. The apparatus according to solution 18, wherein the format rule further specifies that the first element, the second element, the third element, and / or the fourth element immediately follow a fifth element that indicates a zero-based index of a tier, a level, and a hierarchy structure of output layer sets identified by an output layer set index.
[0523] 20. The apparatus according to solution 18, wherein based on whether the operation point contains only a single layer, the format rule further specifies a syntax constraint on a value of the first element, a value of the second element, a value of the third element, and / or a value of the fourth element that applies to operation points associated with bitstreams.
[0524] 21. A non-transitory computer-readable recording medium storing instructions causing a processor to perform a conversion between a visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file includes an operation point record and an operation point group box, and wherein the format rule specifies whether, for each operation point indicated in the visual media file, a first element indicating a chroma format, a second element indicating bit depth information, a third element indicating a maximum picture width, and / or a fourth element indicating a maximum picture height are included in the operation point record and the operation point group box.
[0525] 22. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing apparatus, wherein the method comprises generating a visual media file including one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file includes an operation point record and an operation point group box, and wherein the format rule specifies whether, for each operation point indicated in the visual media file, a first element indicating a chroma format, a second element indicating bit depth information, a third element indicating a maximum picture width, and / or a fourth element indicating a maximum picture height are included in the operation point record and the operation point group box.
[0526] 26. A video processing apparatus comprising a processor configured to implement a method recited by any one or more of solutions 1 to 17.
[0527] 27. A method of storing visual media data into a file including one or more bitstreams, the method comprising a method recited by any one of solutions 1 to 17, and further comprising storing the one or more bitstreams into a non-transitory computer-readable recording medium.
[0528] 28. A computer readable medium storing program code that, when executed, causes a processor to implement the method recited in any one or more of solutions 1 to 17.
[0529] 29. A computer readable medium storing a bitstream generated according to any of the methods.
[0530] 30. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method recited in any one or more of solutions 1 to 17.
[0531] 31. A computer readable medium having stored thereon a bitstream generated according to any of the solutions 1 to 17, in conformance with a file format.
[0532] In example solutions, the visual media data corresponds to a video or an image. In this document, the term "video processing" can refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, a bitstream representation of a current video block can correspond to bits as defined by a syntax, located at the same position or scattered at different positions within the bitstream. For example, a macroblock can be encoded according to a transform and a coded residual, and can also use bits in a header and other fields in the bitstream. Furthermore, during a conversion, a decoder can parse a bitstream according to determinations described in the above solutions, with knowledge that some fields can or can not be present. Similarly, an encoder can determine whether to include certain syntax fields, and generate a coded representation accordingly by including or excluding syntax fields from the coded representation.
[0533] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof). The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling its operation. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0534] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0535] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0536] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0537] Although the present patent document contains many details, these should not be construed as limiting the scope of any subject matter or potentially patentable content in any way. Some of the features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any appropriate subcombination. Moreover, although the features described above can be described as acting in particular combinations and even initially claimed as such, in some cases the features from a claimed combination can be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0538] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0539] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing visual media data, comprising: performing conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to format rules; wherein the format rule specifies that a first element indicating whether a track contains a specific bitstream corresponding to a specific output layer set controls whether a second element indicating a chroma format of the track and / or a third element indicating bit depth information of the track are included in a configuration record of the track, wherein the specific output layer set is formed by a plurality of output layers, and The format rule further specifies syntax constraints according to whether the bitstream to which the configuration record is applicable is a multi-layer bitstream. 2 . The method of claim 1 , wherein the format rule specifies that the second element and / or the third element are included in response to the first element indicating that the track contains the specific bitstream corresponding to the specific output layer set. 3 . The method of claim 1 , wherein the format rule specifies omitting the second element and / or the third element in response to the first element indicating that the track is allowed to not contain the specific bitstream corresponding to the specific output layer set.
4. The method of claim 1 , wherein the format rule further specifies that, in response to the bitstream not being the multi-layer bitstream, the syntax constraint is equality of one or more chroma format values indicated in one or more sequence parameter sets referenced by network abstraction layer (NAL) units included in samples of the visual media file described by a sample entry of the configuration record. 5 . The method of claim 4 , wherein the format rule further specifies that the chroma format value indicated in the configuration record is equal to the one or more chroma format values.
6. The method of claim 1 , wherein the format rule further specifies that, in response to the bitstream being the multi-layer bitstream, the syntax constraint is that the value of the chroma format indicated in the configuration record is set equal to the maximum format value among all chroma format values indicated in the video parameter set and indicated in all video parameter sets of all codec video sequences to which the sample entry description of the configuration record applies and the maximum value of the chroma format applicable to the output layer set identified by the output layer set index.
7. The method of claim 1 , wherein the format rule further specifies that, in response to the bitstream not being the multi-layer bitstream, one or more bit depth information values indicated in one or more sequence parameter sets referenced by network abstraction layer (NAL) units included in samples of the visual media file described by the sample entry of the configuration record are equal.
8. The method of claim 7, wherein the format rule further specifies that the bit depth information value indicated in the configuration record is equal to the one or more bit depth information values.
9. The method according to claim 1, wherein the format rule further specifies: in response to the bitstream being the multi-layer bitstream, the syntax constraint is that the bit depth information value indicated in the configuration record is set to be equal to the maximum value of the bit depth information indicated in the video parameter set and indicated in all video parameter sets of all codec video sequences to which the sample entry description of the configuration record applies and is applicable to the output layer set identified by the output layer set index.
10. The method of any one of claims 1-9, wherein the converting comprises generating the visual media file according to the format rules and storing the one or more bitstreams to the visual media file.
11. The method of any one of claims 1-9, wherein the converting comprises parsing the visual media file according to the format rules to reconstruct the one or more bitstreams.
12. A video processing device, comprising a processor, configured to implement the method according to any one of claims 1 to 11.
13. A method of storing visual media data in a file comprising one or more bitstreams, the method comprising the method of any one of claims 1 to 11, and further comprising storing the one or more bitstreams in a non-transitory computer-readable recording medium.
14. A computer-readable medium storing program code, which, when executed, causes a processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Design of tracks and operation point signaling in layered HEVC file format
US20160373771A1