Signaling of coded picture buffer level in video coding
By independently specifying constraints on image size, DPB size, and AU size for each layer, and defining an overall DPB size limit for multi-layer video encoding and decoding, the inconsistency problem of multi-layer video encoding and decoding in the existing VVC level definition is solved, and the consistency of the decoder's processing and storage capabilities is achieved.
Patent Information
- Application Number
- CN202080090438.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-26
- Filing Date
- 2020-12-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2040-12-27
AI Technical Summary
The existing VVC level definition contains constraints such as image size, DPB size, AU size, and number of stripes that are not applicable to multi-layer video encoding and decoding, resulting in inconsistent decoder capabilities.
By specifying constraints on image size, DPB size, and AU size independently for each layer, and defining an overall DPB size limit for multi-layer video encoding and decoding, the decoder is ensured to be able to handle multi-layer video encoding and decoding.
It achieves effective decoding of multi-layer video codecs, ensuring consistency between the decoder's processing and storage capabilities, and meeting the video codec requirements of different levels.
Smart Images

Figure CN114846805B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 953,815, filed on December 26, 2019, in a timely manner under applicable patent laws and / or the Paris Convention. The entire disclosure of the aforementioned application is incorporated by reference into the disclosure of this application for all legal purposes. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to signal picture buffer levels as part of performing video encoding or decoding.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule, wherein the bitstream includes one or more independently decodable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and wherein the rule specifies that a maximum decoded picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on a maximum allowed picture size of the one or more codec video pictures corresponding to the bitstream or the one or more bitstream portions.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video comprising one or more codec layer video sequences, wherein the bitstream conforms to a rule, and wherein the rule provides that a maximum buffer size of decoded pictures of the codec layer video sequences is constrained to be less than or equal to a maximum picture size selected from pictures in the codec layer video sequences.
[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to a rule, wherein the bitstream includes one or more independently decodable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and wherein the rule specifies at least one of a maximum allowed picture size, a maximum allowed picture width, and a maximum allowed picture height for the conversion.
[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video. Wherein the bitstream includes one or more output layer sets, wherein at least one output layer set includes a plurality of video layers. And wherein the bitstream conforms to a rule that specifies a constraint on an overall size of a decoded picture buffer for the at least one output layer set including the plurality of video layers.
[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies that one or more constraints on at least one of a coded picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, and a total number of bytes in a network abstraction layer (NAL) unit for an access unit are based on a sum of picture sizes of each of a plurality of pictures of the access unit.
[0011] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies a limit on a maximum number of slices in an access unit.
[0012] In another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement a method recited above.
[0013] In another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement a method recited above.
[0014] In another example aspect, a computer readable medium storing code is disclosed. The code embodied in the form of processor-executable code embodies one of the methods described herein.
[0015] These and other features are described in this document. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram illustrating an example video processing system in which various techniques disclosed in this document can be implemented.
[0017] Figure 2 is a block diagram of an example hardware platform for video processing.
[0018] Figure 3 shows a block diagram of an example video encoding system that can implement some embodiments of the disclosure.
[0019] Figure 4 A block diagram illustrating an example encoder in which some embodiments of the disclosure can be implemented is shown.
[0020] Figure 5 A block diagram illustrating an example decoder in which some embodiments of the disclosure can be implemented is shown.
[0021] Figure 6 A flowchart illustrating an example method of video processing is shown.
[0022] Figure 7 A flowchart illustrating an example method of video processing is shown. DETAILED DESCRIPTION
[0023] The use of section headings in this document is for convenience only and does not limit the applicability of techniques and embodiments disclosed in each section to the section itself. Also, H.266 terminology is used in some descriptions only for ease of understanding and not for limiting the scope of the disclosed techniques. Therefore, the techniques described here are applicable to other video codec protocols and designs as well.
[0024] 1. OVERVIEW
[0025] This document relates to video coding techniques. In particular, it is about defining levels of a video codec that support single-layer video coding and multi-layer video coding. It can be applied to any video coding standard or non-standard video codec that supports single-layer video coding and multi-layer video coding, such as the Versatile Video Coding (VVC) that is being developed.
[0026] 2. ABBREVIATIONS
[0027] APS Adaptation Parameter Set
[0028] AU Access Unit
[0029] AUD Access Unit Delimiter
[0030] AVC Advanced Video Coding
[0031] CLVS Coded Layer Video Sequence
[0032] CPB Coded Picture Buffer
[0033] CRA Clean Random Access
[0034] CTU Coding Tree Unit
[0035] CVS Coded Video Sequence
[0036] DPB Decoded Picture Buffer
[0037] DPS Decoding Parameter Set
[0038] EOB End Of Bitstream
[0039] EOS End Of Sequence
[0040] GDR Gradual Decoding Refresh
[0041] HEVC High Efficiency Video Coding
[0042] IDR Instantaneous Decoding Refresh
[0043] JEM Joint Exploration Model
[0044] MCTS Motion-Constrained Tile Sets
[0045] NAL Network Abstraction Layer
[0046] OLS Output Layer Set
[0047] PH Picture Header
[0048] PPS Picture Parameter Set
[0049] PU Picture Unit
[0050] RBSP Raw Byte Sequence Payload
[0051] SEI Supplemental Enhancement Information
[0052] SPS Sequence Parameter Set
[0053] VCL Video Coding Layer
[0054] VPS Video Parameter Set
[0055] VTM VVC Test Model
[0056] VUI Video Usability Information
[0057] VVC Versatile Video Coding
[0058] 3. Preliminary Discussion
[0059] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by the JVET and put into the reference software named Joint Exploration Model (JEM). The JVET meeting is held simultaneously every quarter, and the new coding standard targets 50% bitrate reduction when compared to HEVC. The new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was released at that time. As there are continuous effort on VVC standardization, new coding technologies are being adopted into the VVC standard in every JVET meeting. The working draft of VVC and test model VTM are updated after every meeting. The VVC project now targets technical completion (FDIS) in the July 2020 meeting.
[0060] 3.1 Tiers, tiers and levels
[0061] Video coding standards typically specify tiers and levels. Some video coding standards also specify tiers, e.g., HEVC and VVC under development.
[0062] Tiers, tiers and levels specify restrictions on bitstreams, and thus on the capabilities required to decode bitstreams. Tiers, tiers and levels can also be used to indicate interoperability points between various decoder implementations.
[0063] Each tier specifies a subset of restrictions and algorithmic features supported by all decoders conforming to that tier. Note that encoders are not required to use all the coding tools or features supported in a tier, while decoders conforming to a tier are required to support all the coding tools or features.
[0064] Each level of a tier specifies restrictions on the set of values bitstream syntax elements can take. The same set of tiers and levels is typically used for all tiers, but various implementations can support different tiers, and within one tier, each supported tier can support different levels. For any given tier, the levels of a tier typically correspond to specific decoder processing load and memory capabilities.
[0065] The capabilities of a video decoder conforming to a video codec specification are specified in terms of the capabilities to decode video streams conforming to the restrictions specified in the tiers, tiers and levels of the video codec specification. When expressing the capabilities of a decoder for a particular tier, the tiers and levels supported by that tier should also be expressed.
[0066] 3.2 Existing VVC level and tier definitions
[0067] In the latest VVC draft text in JVET-P2001-v14, publicly available text can be found at http: / / phenix.int-evry.fr / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P2001-v14.zip, the level definitions are as follows.
[0068] A.1.1 General tier and level restrictions
[0069] For the purpose of comparing tier capabilities, a tier with general_tier_flag equal to 0 is considered to be a lower tier than a tier with general_tier_flag equal to 1.
[0070] For the purpose of comparing level capabilities, a particular level of a specified layer is considered to be a lower level than other levels of the same layer when the value of general_level_idc or sublayer_level_idc[i] of the particular level is smaller than the values of other levels.
[0071] To express the constraints in this appendix, the following are specified:
[0072] –Let access unit n be the nth access unit in decoding order, and the first access unit be access unit 0 (i.e., the 0th access unit).
[0073] – Let picture n be the codec picture of access unit n or the corresponding decoded picture
[0074] When the specified level is other than Level 8.5, a bitstream conforming to the profile of the specified layer and level shall be subject to the following constraints for each bitstream conformance test specified in Annex C:
[0075] a) PicSizeInSamplesY shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.
[0076] b)The value of pic_width_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).
[0077] c)The value of pic_height_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).
[0078] d) The value of num_tile_columns_minus1 shall be less than MaxTileCols, and the value of num_tile_rows_minus1 shall be less than MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.
[0079] e) For VCL HRD parameters, for at least one value of i in the range 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] shall be less than or equal to CpbVclFactor*MaxCPB, where CpbSize[Htid][i] is specified in clause 7.4.6.3 based on the parameters specified in clause C.1, CpbVclFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in bits of CpbVclFactor.
[0080] f) For NAL HRD parameters, for at least one i value in the range of 0 to hrd_cpb_cnt_minusl, inclusive, CpbSize[ Htid ][ i ] shall be less than or equal to CpbNalFactor * MaxCPB, where CpbSize[ Htid ][ i ] is specified in Clause 7.4.6.3 in terms of the parameters specified in Clause C.1, CpbNalFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0081] Table A.1 specifies the constraints for each level of each layer except for level 8.5.
[0082] The layers and levels to which the bitstream conforms are indicated by the syntax elements general_tier_flag and general_level_idc, and the sublayer levels to which the sublayers conform are indicated by the syntax elements sublayer_level_idc[ i ], as follows:
[0083] - If the specified level is not level 8.5, then according to the tier constraints specified in Table A.1, general_tier_flag equal to 0 indicates conformance to the Main tier, and general_tier_flag equal to 1 indicates conformance to the High tier, and according to the layer constraints specified in Table A.1, general_tier_flag shall be equal to 0 for layers below level 4 (corresponding to the entries in Table A.1 marked with a "-").
[0084] Otherwise (the specified level is level 8.5), the requirement for bitstream conformance is that general_tier_flag shall be equal to 1, and the value 0 for general_tier_flag is reserved for future use by ITU-T | ISO / IEC, and decoders shall ignore the value of general_tier_flag.
[0085] - general_level_idc and sublayer_level_idc[ i ] shall be set equal to 30 times the level number specified in Table A.1.
[0086] Table A.1 - General tier and level constraints
[0087]
[0088] A.1.2 Tier - specified level constraints
[0089] To express the constraints in this annex, the following are specified:
[0090] - the variable fR is set equal to 1 ÷ 300.
[0091] The variable HbrFactor is defined as follows:
[0092] - If the bitstream is indicated to conform to the Main 10 profile or the Main 4:4:4 10 profile, HbrFactor is set equal to 1.
[0093] The variable BrVclFactor, representing the VCL bitrate scaling factor, is set equal to CpbVclFactor * HbrFactor.
[0094] The variable BrNalFactor, representing the NAL bitrate scaling factor, is set equal to CpbNalFactor * HbrFactor.
[0095] The variable MinCr is set equal to MinCrBase * MinCrScaleFactor ÷ HbrFactor.
[0096] When the specified level is not level 8.5, the value of max_dec_pic_buffering_minus1 [Htid] + 1 shall be less than or equal to MaxDpbSize, which is derived as follows:
[0097]
[0098] where MaxLumaPs is specified in Table A.1 and maxDpbPicBuf is equal to 8.
[0099] A bitstream conforming to the Main 10 or Main 4:4:4 10 profile of a specified layer and level shall obey the following restrictions of each bitstream conformance test specified in Annex C:
[0100] a) The nominal removal time of access unit n (n greater than 0) from the CPB, as specified in clause C.2.3, shall satisfy the following constraint: AuNominalRemovalTime [n] - AuCpbRemovalTime [n - 1] is greater than or equal to Max(PicSizeInSamplesY ÷ MaxLumaSr, fR) for the value of PicSizeInSamplesY of picture n - 1, where MaxLumaSr is the value specified to apply to picture n - 1.
[0101] b) The difference between successive output times of pictures in the DPB shall satisfy the following constraint as specified in clause C.3.3: DpbOutputlnterval[n] is greater than or equal to Max(PicSizeInSamplesY ÷ MaxLumaSr, fR) for the value of PicSizeInSamplesY of picture n, where MaxLumaSr is the value specified in Table A.2 for picture n, provided that picture n is an output picture and is not the last picture output by the bitstream.
[0102] c) The removal time of access unit 0 shall satisfy the constraint that the number of slice segments in picture 0 is less than or equal to Min(Max(1, MaxSliceSegmentsPerPicture * MaxLumaSr / MaxLumaPs * (AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) +
[0103] MaxSliceSegmentsPerPicture * PicSizeInSamplesY / MaxLumaPs), MaxSliceSegmentsPerPicture) for the value of PicSizeInSamplesY of picture 0, where MaxSliceSegmentsPerPicture, MaxLumaPs and MaxLumaSr are the values specified in Table A.1 and Table A.2, respectively, that apply to picture 0.
[0104] d) The difference between successive CPB removal times of access unit n and access unit n-1 (n greater than 0) shall satisfy the constraint that the number of slice segments in picture n is less than or equal to Min((Max(1, MaxSliceSegmentsPerPicture * MaxLumaSr / MaxLumaPs
[0105] (AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1])), MaxSliceSegmentsPerPicture) for the values of MaxSliceSegmentsPerPicture, MaxLumaPs and MaxLumaSr specified in Table A.1 and Table A.2, respectively, that apply to picture n.
[0106] e) For the parameter VCL HRD, for at least one i value in the range of 0 to hrd_cpb_cnt_minusl, inclusive, BitRate[Htid][i] shall be less than or equal to BrVclFactor * MaxBR, where BitRate[Htid][i] is specified in clause 7.4.6.3 based on the parameters specified in clause C.l, MaxBR is specified in Table A.2 in units of BrVclFactor bits / s.
[0107] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minusl, inclusive, BitRate[Htid][i] shall be less than or equal to BrNalFactor * MaxBR, where BitRate[Htid][i] is specified in clause 7.4.6.3 based on the parameters specified in clause C.l, MaxBR is specified in Table A.2 in units of BrVclFactor bits / s.
[0108] g) For the value of PicSizeInSamplesY of picture 0, the sum of the NumBytesInNalUnit variables of access unit 0 shall be less than or equal to FormatCapabilityFactor * (Max(PicSizeInSamplesY, fR * MaxLumaSr) +
[0109] MaxLumaSr * (AuCpbRemovalTime[0] - AuNominalRemovalTime[0])) ÷ MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Table A.2 and Table A.3, respectively, that apply to picture 0.
[0110] h) The sum of the NumBytesInNalUnit variables of access unit n (n greater than 0) shall be less than or equal to FormatCapabilityFactor * MaxLumaSr * (AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1]) ÷ MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Table A.2 and Table A.3, respectively, that apply to picture n.
[0111] i) For the value of pic sizes in samples y of picture 0, the removal time of access unit 0 shall satisfy the constraint that the number of tiles in picture 0 is less than or equal to Min(Max(1, MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) + MaxTileCols*MaxTileRows*PicSizeInSamplesY / MaxLumaPs, MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.l that apply to picture 0.
[0112] j) The difference between the consecutive CPB removal times of access unit n and n-1 (n is greater than 0) shall satisfy the constraint that the number of tiles in picture n is less than or equal to Min(Max(1, MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1])), MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.l that apply to picture n.
[0113] …
[0114] Table A.2 - Layer and level restrictions for video profiles
[0115]
[0116] Table A.3 - Specification of CpbVclFactor, CpbNalFactor, FormatCapabilityFactor and MinCrScaleFactor
[0117]
[0118] 4. Technical problem to be solved by the technical solution
[0119] The existing VVC design for level definition has the following problems:
[0120] 1) For the equation to derive the variable MaxDpbSize, i.e. equation (A.1) above, the variable PicSizeInSamplesY is used, which is the picture size of a particular picture. However, in VVC, the picture size can vary from one picture to another within a coded layer video sequence (CLVS). Therefore, the maximum picture size among the pictures within a CLVS should be used to derive MaxDpbSize. Furthermore, because there can be multiple layers and different layers in the bitstream of an output layer set (OLS) have different maximum picture size values in the pictures within a CLVS, the value of MaxDpbSize needs to be derived for each layer.
[0121] 2) The constraint related to DPB size specifies that the value of max_dec_pic_buffering_minus1[ ] + 1 shall be less than or equal to MaxDpbSize, which constraint only applies to single-layer bitstreams. For an output layer set (OLS) that contains multiple layers in its bitstream, because there can be multiple different instances of the max_dec_pic_buffering_minus1[ ] syntax element applied to different layers, the existing constraint does not apply. Instead, the constraint needs to be applied to each layer using the appropriate instance of the max_dec_pic_buffering_minus1[ ] syntax element.
[0122] 3) The definition of the constraints on picture size, picture width and picture height, i.e. items a, b and c in section A.1.1 above, uses the variable PicSizeInSamplesY and the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples. Again, because in VVC the picture size can vary from one picture to another within a CLVS, the values of the maximum picture size, picture width and picture height should be used. Furthermore, because there can be multiple layers and different layers in the bitstream of an OLS have different maximum picture size, picture width and picture height values in the pictures within a CLVS, the constraints need to be specified separately for each layer or each SPS that a layer refers to.
[0123] 4) Furthermore, there is a need to constrain the overall DPB size for an OLS that contains multiple layers, such that the same DPB size limit of the decoder will be in effect regardless of the number of layers. This is currently missing.
[0124] 5) The definition of the constraints on the sum of the AU's CPB removal time, nominal CPB removal time, DPB output time and NumBytesInNalUnit variables, i.e. the a, b, c, d, g and i terms in the A.1.2 section above, uses the variable PicSizeInSamplesY, which is the picture size of the particular picture. However, since an AU can contain multiple pictures, the sum of the picture sizes of all pictures in the AU should be used.
[0125] 6) The limit on the maximum number of slices per AU should be specified and used when specifying these constraints, instead of specifying the limit on the maximum number of slices per picture for each level and using this limit when specifying some constraints on the AU's CPB removal time, i.e. the c and d terms in the A.1.2 section above.
[0126] 5. Example embodiments and techniques
[0127] To solve the above problems and others, methods summarized as follows are disclosed in a list of various items. These items should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Moreover, these items can be used individually or combined in any way.
[0128] 1) To solve the first problem, the equation for deriving the variable MaxDpbSize is updated to use the maximum picture size among the pictures within the CLVS, instead of using the variable PicSizeInSamplesY, and the value of MaxDpbSize is derived for each layer.
[0129] 2) To solve the second problem, for each layer, the constraint requires that the value of a particular instance of the max_dec_pic_buffering_minus1[ ] syntax element plus 1 is less than or equal to the value of MaxDpbSize for that layer, where the particular instance of the max_dec_pic_buffering_minus1[ ] syntax element is max_dec_pic_buffering_minus1[ i ] for which i takes the maximum value in the dpb_parameters( ) syntax structure, which is determined as follows: if the layer is an output layer of an OLS, the dpb_parameters( ) syntax structure is the one that is applied to the layer when it is an output layer in the OLS; otherwise, the dpb_parameters( ) syntax structure is the dpb_parameters( ) syntax structure that is applied to the layer when it is not an output layer in the OLS.
[0130] 3) To address the third issue, the definition of the constraints on picture size, picture width and picture height use the maximum picture size, picture width and picture height instead of the variables PicSizeInSamplesY and the syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples. In addition, the constraints are specified separately for each layer or each SPS referred by a layer.
[0131] 4) To address the fourth issue, a constraint on the overall DPB size of an OLS containing multiple layers is specified.
[0132] a. In one alternative, the constraints are specified as follows:
[0133] After each AU is decoded, numDecPics is set to the number of decoded pictures in the DPB and picSizeInSamplesY[i] is set to the value of picSizeInSamplesY for the i-th decoded picture in the DPB, where i is in the range of 0 to numDecPics - 1, inclusive, The value of numDecPics * picSizeInSamplesY shall be less than or equal to maxDpbPicBuf * MaxLumaPs, where maxDpbPicBuf is equal to 8 and MaxLumaPs is specified in Table A.l.
[0134] b. In another alternative, the constraints are specified as follows:
[0135] Let numLayers be the number of layers in the OLS, maxDecBuff[i] and PicSizeMaxInSamplesY[i] be the values of max_dec_pic_buffering_minusl[maxTid] and PicSizeMaxInSamplesY, respectively, for the i-th layer, i being in the range of 0 to numLayers - 1, inclusive, where PicSizeMaxInSamplesY is equal to pic_width_max_in_luma_samples * pic_height_max_in_luma_samples for the layer and max_dec_pic_buffering_minusl[maxTid] for the layer is the value of max_dec_pic_buffering_minusl[i] in the dpb_parameters() syntax structure for the maximum value of i, which is determined as follows:
[0136] If the layer is an output layer of an OLS, the dpb_parameters( ) syntax structure is the one applied to the layer when it is an output layer in an OLS; otherwise, the dpb_parameters( ) syntax structure is the one applied to the layer when it is not an output layer in an OLS. The value of max_dec_pic shall be less than or equal to maxDpbPicBuf * MaxLumaPs, where maxDpbPicBuf is equal to 8 and MaxLumaPs is specified in Table A.l.
[0137] 5) To address the fifth issue, the sum of the picture size of all pictures in an AU is used for the definition of several constraints on the sum of the CPB removal time, nominal CPB removal time, DPB output time and NumBytesInNalUnit variables for an AU, instead of the variable PicSizeInSamplesY.
[0138] 6) To address the sixth issue, a limit on the maximum number of slices per AU is specified and used when specifying these constraints, instead of specifying a limit on the maximum number of slices per picture for each level and using this limit when specifying some constraints on the CPB removal time of an AU.
[0139] 6 Embodiments
[0140] The following are some example embodiments that can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-P2001-v14. Most of the added or modified parts are highlighted in bold double brackets. There are also some editorial nature of modifications, so they are not highlighted.
[0141] 6.1 First embodiment 6.1.1 Profiles, tiers, levels overview
[0142] Profiles, tiers and levels specify limits on bitstreams, and thus limit the capabilities required to decode a bitstream. Profiles, tiers and levels can also be used to indicate interoperability points between various decoder implementations.
[0143] Each profile specifies a subset of the limits and algorithmic features supported by all decoders conforming to the profile.
[0144] NOTE - An encoder is not required to use any particular subset of the features supported in a profile.
[0145] Each level of a tier specifies a restriction on the set of values that a syntax element of this Specification can take. All profiles typically use the same set of tier and level definitions, but individual implementations can support different tiers, and within a tier, each supported profile can support different levels. For any given profile, the levels of a tier typically correspond to specific decoder processing load and storage capabilities.
[0146] {{In this clause, the phrase like "bitstream conforms" shall be interpreted as "bitstream of OLS conforms."}}
[0147] 6.1.2 Overview of tiers and levels
[0148] For the purpose of comparing tier functionality, a tier with general_tier_flag equal to 0 is considered a lower tier than a tier with general_tier_flag equal to 1.
[0149] For the purpose of comparing level capability, a particular level of a specified tier is considered a lower level than another level of the same tier when the value of general_level_idc or sublayer_level_idc[ i ] for the particular level is less than the value for the other level.
[0150] For the purpose of expressing the constraints in this annex, the following are specified:
[0151] - Let AU n be the nth AU in decoding order, the first AU being AU 0 (i.e., the 0th AU).
[0152] - Let PicSizeMaxInSamplesY be the value of pic_width_max_in_luma_samples * pic_height_max_in_luma_samples for a particular reference SPS
[0153] When the specified level is not level 8.5, the value of MaxDpbSize is derived as follows for each tier in the OLS:
[0154]
[0155]
[0156] where MaxLumaPs is specified in Table A.1 and maxDpbPicBuf is equal to 8.
[0157] When the specified level is not level 8.5, a bitstream conforming to a profile of the specified tier and level shall obey the following constraints of each bitstream conformance test specified in Annex C:
[0158] a) {{ each reference SPS PicSizeInSamplesY}} shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.
[0159] b) {{ for each reference SPS,}} the value of pic_width_{{ max_}}in_luma_samples shall be less than or equal to Sqrt(MaxLumaPs * 8).
[0160] c) {{ for each reference SPS,}} the value of pic_height_{{ max_}}in_luma_samples shall be less than or equal to Sqrt(MaxLumaPs * 8).
[0161] d) {{ for each reference PPS,}} the value of NumTileColumns shall be less than MaxTileCols and the value of NumTileRows shall be less than MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.
[0162] e) For VCL HRD parameters, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1, inclusive, CpbSize[Htid][i] shall be less than or equal to CpbVclFactor * MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 in terms of the parameters specified in Clause C.1, CpbVclFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in units of CpbVclFactor bits.
[0163] f) For NAL HRD parameters, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1, inclusive, CpbSize[Htid][i] shall be less than or equal to CpbNalFactor * MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 in terms of the parameters specified in Clause C.1, CpbNalFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0164] g) The value of max_dec_pic_buffering_minus1[ maxTid ] + 1 for each layer in the OLS shall be less than or equal to the MaxDpbSize of that layer, where max_dec_pic_buffering_minus1[ maxTid ] is the value of max_dec_pic_buffering_minus1[ i ] in the dpb_parameters( ) syntax structure for the largest value of i, which is determined as follows: if the layer is an output layer of the OLS, the dpb_parameters( ) syntax structure is the one applied to the layer when it is an output layer in the OLS; otherwise, the dpb_parameters( ) syntax structure is applied when the layer is not an output layer in the OLS.
[0165] h) Immediately after each AU is decoded, numDecPics is set equal to the number of decoded pictures in the DPB, and picSizeInSamplesY[ i ] is set equal to the value of picSizeInSamplesY for the i-th decoded picture in the DPB, i being in the range of 0 to numDecPics - 1, inclusive, the value of max_dec_pic_buffering_minus1[ maxTid ] + 1 for each layer in the OLS shall be less than or equal to the MaxDpbSize of that layer, where max_dec_pic_buffering_minus1[ maxTid ] is the value of max_dec_pic_buffering_minus1[ i ] in the dpb_parameters( ) syntax structure for the largest value of i, which is determined as follows: if the layer is an output layer of the OLS, the dpb_parameters( ) syntax structure is the one applied to the layer when it is an output layer in the OLS; otherwise, the dpb_parameters( ) syntax structure is applied when the layer is not an output layer in the OLS.
[0166] Table A.1 specifies the constraints for each level except level 8.5.
[0167] The layers and levels to which the bitstream conforms are indicated by the syntax elements general_tier_flag and general_level_idc, and the sub-layer representations conform to the levels indicated by the syntax elements sublayer_level_idc[ i ], as follows:
[0168] - If the specified level is not level 8.5, then according to the level constraints specified in Table A.1, general_tier_flag equal to 0 indicates conformance to the main tier and general_tier_flag equal to 1 indicates conformance to the high tier, and for levels below level 4, general_tier_flag shall be equal to 0 (corresponding to the entries in Table A.1 marked with a "-"). Otherwise (the specified level is level 8.5), the requirement for bitstream conformance is that general_tier_flag shall be equal to 1, and the value 0 for general_tier_flag is reserved for future use by ITU-T | ISO / IEC and decoders shall ignore the value of general_tier_flag.
[0169] - general_level_idc and sublayer_level_idc[ i ] shall be set equal to a number equal to 30 times the level number specified in Table A.1.
[0170] Table A.1 - General tier and level restrictions
[0171]
[0172] 6.1.3 Tier - Specified level restrictions
[0173] To express the constraints in this annex, the following are specified:
[0174] - Let the variable fR be equal to 1 ÷ 300.
[0175] The variable HbrFactor is defined as follows:
[0176] - If the bitstream is indicated to conform to the Main 10 tier or the Main 4:4:4 10 tier, HbrFactor is set equal to 1.
[0177] The variable BrVclFactor, representing the VCL bitrate scaling factor, is set equal to CpbVclFactor * HbrFactor.
[0178] The variable BrNalFactor, representing the NAL bitrate scaling factor, is set equal to CpbNalFactor * HbrFactor.
[0179] The variable MinCr is set equal to MinCrBase * MinCrScaleFactor ÷ HbrFactor.
[0180] {{The variable AuSizeInSamplesY[ n ] is set equal to where numDecPics is the number of pictures in the AU n, and picSizeInSamplesY[ i ] is the picSizeInSamplesY value of the i-th picture in the AU n, i being in the range of 0 to numDecPics - 1, inclusive.}}
[0181] A bitstream conforming to the Main 10 or Main 4:4:4 10 tier of a specified layer and level shall obey the following restrictions for each bitstream conformance test specified in Annex C:
[0182] a) The nominal removal time of AU n (n greater than 0) from the CPB shall satisfy the constraint that AuNominalRemovalTime[n] - AuCpbRemovalTime[n-1] is greater than or equal to Max({{PicSizeInSamplesY[n-1]}} ÷ MaxLumaSr, fR), where MaxLumaSr is the value specified in Table A.2 that applies to AU n-1.
[0183] b) The difference between the consecutive output times of pictures of different AUs in the DPB shall satisfy the constraint that DpbOutputInterval[n] is greater than or equal to Max({{AuSizeInSamplesY[n]}} ÷ MaxLumaSr, fR), where MaxLumaSr is the value specified in Table A.2 for AU n, provided that the AU has an output picture and AU n is not the last AU with an output picture in the bitstream.
[0184] c) The CPB removal time of AU 0 shall satisfy the constraint that the number of slices in AU 0 is less than or equal to Min(Max(1,{{MaxSlicesPerAu}}*MaxLumaSr / MaxLumaPs
[0185] AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) +{{MaxSlicesPerAu}}*{{AuSizeInSamplesY[0]}} / MaxLumaPs),{{MaxSlicesPerAu}}), where {{MaxSlicesPerAu}}, MaxLumaPs and MaxLumaSr are the values specified in Table A.1 and Table A.2, respectively, that apply to AU 0.
[0186] d) The difference between the consecutive CPB removal times of AU n and AU n-1 (n greater than 0) shall satisfy the constraint that the number of slices in AU n is less than or equal to Min((Max(1,{{MaxSlicesPerAu}}*MaxLumaSr / MaxLumaPs*(AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1])),{{MaxSlicesPerAu}}), where {{MaxSlicesPerAu}}, MaxLumaPs and MaxLumaSr are the values specified in Table A.1 and Table A.2, respectively, that apply to AU n.
[0187] e) For the parameter VCL HRD, for at least one i value in the range of 0 to hrd_cpb_cnt_minusl, inclusive, BitRate[ Htid ][ i ] shall be less than or equal to BrVclFactor * MaxBR, where BitRate[ Htid ][ i ] is specified in clause 7.4.6.3 based on the parameters specified in clause C.1, and MaxBR is specified in Table A.2 in units of BrVclFactor bits / s.
[0188] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minusl, inclusive, BitRate[ Htid ][ i ] shall be less than or equal to BrNalFactor * MaxBR, where BitRate[ Htid ][ i ] is specified in clause 7.4.6.3 based on the parameters specified in clause C.1, and MaxBR is specified in Table A.2 in units of BrVclFactor bits / s.
[0189] g) The sum of the NumBytesInNalUnit variables for AU 0 shall be less than or equal to FormatCapabilityFactor * (Max({{AuSizeInSamplesY[0]}}, fR * MaxLumaSr) + MaxLumaSr * (AuCpbRemovalTime[0] - AuNominalRemovalTime[0])) ÷ MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Table A.2 and Table A.3, respectively, that apply to AU 0.
[0190]
[0191] h) The sum of the NumBytesInNalUnit variables for AU n (n greater than 0) shall be less than or equal to FormatCapabilityFactor * MaxLumaSr * (AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1]) ÷ MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Table A.2 and Table A.3, respectively, that apply to AU n.
[0192] i) The removal time of AU 0 shall satisfy the constraint that the number of tiles per picture in AU 0 is less than or equal to Min(Max(1, MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) + MaxTileCols*MaxTileRows*{{AuSizeInSamplesY[0]}} / MaxLumaPs, MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.l that apply to AU 0.
[0193] j) The difference between the consecutive CPB removal times of AU n and AU n-1 (n greater than 0) shall satisfy the constraint that the number of tiles per picture in AU n is less than or equal to Min(Max(1, MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1])), MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.l that apply to AU n.
[0194] …
[0195] Figure 1 FIG. 10 is a block diagram illustrating an example video processing system 1000 that can perform various techniques in this disclosure. Various implementations can include some or all of the components of system 1000. System 1000 can include an input 1002 to receive video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or in a compressed or encoded format. Input 1002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc. and wireless interfaces such as Wi-Fi or cellular interfaces.
[0196] The system 1000 can include a codec component 1004 that can implement various coding or encoding methods described in this document. The codec component 1004 can reduce the average bitrate of a video from an input 1002 to an output of the codec component 1004 to produce a coded representation of the video. Thus, the codec techniques are sometimes called video compression or video transcoding techniques. The output of the codec component 1004 can be stored or transmitted via a connected communication as represented by component 1006. The stored or transmitted bitstream representation (or coded representation) of the video received at the input 1002 can be used by a component 1008 to generate pixel values or displayable video that is sent to a display interface 1010. The process of generating user-viewable video from a bitstream representation is sometimes called video decompression. Also, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that encoding tools or encoding operations are used at an encoder and corresponding decoding tools or decoding operations that reverse the results of the encoding will be performed by a decoder.
[0197] Examples of peripheral bus interfaces or display interfaces can include a Universal Serial Bus (USB) or a High Definition Multimedia Interface (HDMI) or Displayport, among others. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, and so forth. The techniques described in this document can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other devices capable of performing digital data processing and / or video display.
[0198] Figure 2 is a block diagram of a video processing device 2000. The device 2000 can be used to implement one or more methods described herein. The device 2000 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, or the like. The device 2000 can include one or more processors 2002, one or more memories 2004, and video processing circuitry 2006. The processor(s) 2002 can be configured to implement one or more methods described in the present document (e.g., Figures 6-7 ). The memory(ies) 2004 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 2006 can be used to implement, in hardware circuitry, some of the techniques described in the present document.
[0199] Figure 3 A block diagram of an example video codec system 100 that can utilize the techniques of this disclosure is shown. As Figure 3As shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 generates encoded video data, which can be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110, which can be referred to as a video decoding device. Source device 110 can include a video source 112, a video encoder 114, and an input / output (VO) interface 116.
[0200] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. VO interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to destination device 120 by VO interface 116 via network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by destination device 120.
[0201] Destination device 120 can include VO interface 126, video decoder 124, and display device 122.
[0202] VO interface 126 can include a receiver and / or a modem. VO interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.
[0203] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0204] Figure 4 is a block diagram illustrating an example of a video encoder 200, which can be Figure 3 the video encoder 114 in the system 100 as shown.
[0205] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 4In an example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0206] The functional components of video encoder 200 can include partitioning unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 810, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.
[0207] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0208] Furthermore, some components (e.g., motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but are represented separately for illustrative purposes. Figure 4
[0209] Partitioning unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0210] Mode select unit 203 can select one of a plurality of coding modes (intra or inter), for example, based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode select unit 203 can select a combination of intra and inter prediction (CIIP) mode, in which prediction is based on inter prediction signaling and intra prediction signaling. In the case of inter prediction, mode select unit 203 can also select a precision of motion vectors for the block (e.g., sub-pixel or integer-pixel precision).
[0211] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples from other pictures than the picture in which the current video block is associated with, from buffer 213.
[0212] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0213] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 or list 7. Motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 7 that includes the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0214] In other examples, motion estimation unit 204 can perform bi-prediction on a current video block, motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0, and can also search for another reference video block for the current video block in a reference picture in list 7. Motion estimation unit 204 can then generate a reference index indicating the reference pictures in list 0 and list 7 that include the reference video blocks and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0215] In some examples, motion estimation unit 204 can output full motion information for a decoder's decoding process.
[0216] In some examples, motion estimation unit 204 can not output a full set of motion information for a current video. Instead, motion estimation unit 204 can signal motion information for a current video block with reference to motion information for another video block. For example, motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.
[0217] In one example, motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to video decoder 900 that the current video block has the same motion information as another video block.
[0218] In another example, the motion estimation unit 204 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0219] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of prediction signaling techniques that can be implemented by the video encoder 200 include advanced motion vector predication (AMVP) and merge mode signaling.
[0220] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0221] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0222] In other instances, the current video block can not have residual data for the current video block, such as in skip mode, and the residual generation unit 207 can not perform the subtraction operation.
[0223] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0224] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0225] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. The reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0226] After the video block is reconstructed by the reconstruction unit 212, an in-loop filtering operation can be performed to reduce video block artifacts in the video block.
[0227] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0228] Figure 5 is a block diagram illustrating an example of a video decoder 300 that can be Figure 3 the video decoder 114 in the system 100 shown.
[0229] The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 examples, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0230] In Figure 5 examples, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200 Figure 4 ).
[0231] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and merge modes.
[0232] The motion compensation unit 302 can generate a motion compensated block, can perform interpolation based on an interpolation filter. An identifier for an interpolation filter used at sub-pixel precision can be included in the syntax elements.
[0233] Motion compensation unit 302 can use an interpolation filter as used by video encoder 200 during encoding of a video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.
[0234] Motion compensation unit 302 can use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence was partitioned, modes indicating how each partition was encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information to decode the encoded video sequence.
[0235] Intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form the prediction block from spatially neighboring blocks. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0236] Reconstruction unit 306 can add the residual block to the corresponding prediction block produced by motion compensation unit 202 or intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to the decoded block to filter out blockiness artifacts. The decoded video block is then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0237] Figures 6-7 An example method is shown that can implement the above-described techniques (e.g., the techniques described above in connection with the embodiments shown). Figures 1 to 5 An example method is shown that can implement the above-described techniques (e.g., the techniques described above in connection with the embodiments shown).
[0238] Figure 6 A flowchart of an example method 600 of video processing is shown. The method 600 includes, at operation 610, performing a conversion between a video and a bitstream of the video, the bitstream being organized according to a rule into one or more access units, the rule specifying that one or more constraints on at least one of a coded picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, and a sum of a number of bytes in network abstraction layer (NAL) units for an access unit are based on a sum of picture sizes of each of a plurality of pictures of the access unit.
[0239] Figure 7A video processing example method 700 is shown. The method 700 includes, at operation 710, performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule that specifies a limit on a maximum number of slices in an access unit.
[0240] Next, a list of preferred options for some embodiments is provided.
[0241] A1. A method of video processing, comprising performing a conversion between a video and a bitstream of the video according to a rule, wherein the bitstream comprises one or more independently decodable bitstream portions each corresponding to one or more coded video pictures of the video, and wherein the rule specifies that a maximum decoding picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on a maximum allowed picture size of the one or more coded video pictures corresponding to the bitstream or the one or more bitstream portions.
[0242] A2. The method of solution 1, further comprising avoiding determining the maximum picture buffer size based on a picture size in luma samples (denoted as PicSizeInSamplesY).
[0243] A3. The method of solution 1 or 2, wherein the maximum decoding picture buffer size is denoted as MaxDpbSize, a maximum luma picture size is denoted as MaxLumaPs, the maximum allowed picture size is denoted as PicSizeMaxInSamplesY, a default maximum decoding picture buffer size is denoted as maxDpbPicBuf, and wherein MaxDpbSize is determined as follows:
[0244] if (PicSizeMaxInSamplesY <= (MaxLumaPs >> 2))
[0245] MaxDpbSize = Min(4 * maxDpbPicBuf, 16)
[0246] else if (PicSizeMaxInSamplesY <= (MaxLumaPs >> 1))
[0247] MaxDpbSize = Min(2 * maxDpbPicBuf, 16)
[0248] else if (PicSizeMaxInSamplesY <= ((3 * MaxLumaPs) >> 2))
[0249] MaxDpbSize = Min( (4 * maxDpbPicBuf) / 3, 16 )
[0250] Other
[0251] MaxDpbSize = maxDpbPicBuf.
[0252] A4. The method of any of solutions 1-3, wherein each of the one or more independently decodable bitstream portions is a coded layer video sequence (CLVS).
[0253] A5. The method of any of solutions 1-4, wherein the maximum decoded picture buffer size is derived on a per-CLVS basis.
[0254] A6. A method of video processing, comprising performing a conversion between a video and a bitstream of the video according to a rule, wherein the bitstream includes one or more independently decodable bitstream portions, each bitstream portion corresponding to one or more coded video pictures of the video, and wherein the rule specifies at least one of a maximum allowed picture size, a maximum allowed picture width, a maximum allowed picture height for the conversion.
[0255] A7. The method of solution 6, wherein each of the one or more independently decodable bitstream portions is a coded layer video sequence (CLVS).
[0256] A8. The method of solution 7, wherein one or more of the maximum allowed picture size, the maximum allowed picture width, and the maximum allowed picture height are specified on a per-CLVS basis.
[0257] A9. The method of solution 7, wherein the maximum allowed picture size, the maximum allowed picture width, and the maximum allowed picture height are specified for each sequence parameter set (SPS) associated with each of the at least one CLVS.
[0258] A10. The method of any of solutions 6-9, wherein the rule specifies that the maximum allowed picture size (denoted as PicSizeMaxInSamplesY) of each reference sequence parameter set (SPS) is less than or equal to a maximum luma picture size (denoted as MaxLumaPs).
[0259] A11. The method of any of solutions 6-10, wherein the rule specifies that the maximum allowed picture width (expressed as pic width max in luma samples) of each reference sequence parameter set (SPS) is less than or equal to Sqrt(8 x MaxLumaPs), where MaxLumaPs is a maximum luma picture size.
[0260] A12. The method of any of solutions 6-11, wherein the rule specifies that the maximum allowed picture height (expressed as pic height max in luma samples) of each reference sequence parameter set (SPS) is less than or equal to Sqrt(8 x MaxLumaPs), where MaxLumaPs is a maximum luma picture size.
[0261] A13. A method of video processing, comprising:
[0262] performing a conversion between a video and a bitstream of the video that comprises one or more coded layer video sequences, wherein the bitstream conforms to a rule, and wherein the rule specifies that a maximum buffer size of decoded pictures of the coded layer video sequence is constrained to be less than or equal to a maximum picture size selected from among pictures in the coded layer video sequence.
[0263] A14. The method of solution 13, wherein the maximum buffer size of decoded pictures is selected based on a maximum buffer size from a decoded picture buffer (DPB) parameter set.
[0264] A15. The method of solution 14, wherein the DPB parameter set corresponds to a DPB parameter set of a coded layer that is an outer layer of an output layer set (OLS).
[0265] 16. The method of solution 14, wherein the DPB parameter set corresponds to a DPB parameter set of a coded layer that is not an outer layer of an output layer set (OLS).
[0266] A17. A method of video processing, comprising performing a conversion between a video and a bitstream of the video, wherein the bitstream comprises one or more output layer sets, wherein at least one output layer set comprises a plurality of video layers, and wherein the bitstream conforms to a rule that specifies a constraint on an overall size of a decoded picture buffer for the at least one output layer set that comprises the plurality of video layers.
[0267] A18. The method of paragraph 17, wherein the decoded picture buffer comprises a plurality of decoded pictures after decoding each access unit, wherein each of the plurality of decoded pictures has a width in luma samples, and wherein the constraint specifies that a sum of the widths of the plurality of decoded pictures is less than or equal to a predetermined value.
[0268] A19. The method of paragraph 18, wherein the predetermined value is based on an index of a respective coding layer of the plurality of coding layers.
[0269] A20. The method of paragraph 17, wherein each of the plurality of coding layers is associated with each of a plurality of decoded pictures, wherein each of the plurality of decoded pictures has a maximum width, and wherein the constraint specifies that a sum of the maximum widths of the plurality of decoded pictures is less than or equal to a predetermined value.
[0270] A21. The method of paragraph 20, wherein the maximum width is a product of a maximum picture width in luma samples and a maximum picture height in luma samples of a corresponding layer of the plurality of coding layers.
[0271] A22. The method of any of paragraphs 1-21, wherein the converting comprises decoding the video from the bitstream.
[0272] A23. The method of any of paragraphs 1-21, wherein the converting comprises encoding the video into the bitstream.
[0273] A24. The method of any of paragraphs 1-21, wherein performing the converting comprises:
[0274] encoding the video into the bitstream; and storing the bitstream in a non-transitory computer-readable storage medium.
[0275] A25. A video processing apparatus comprising a processor configured to perform the method of any one or more of paragraphs 1-24.
[0276] A26. A non-transitory computer-readable storage medium configured to store a bitstream of a video generated by the method of any one or more of paragraphs Al to A24.
[0277] A27. A non-transitory computer-readable storage medium configured to store instructions that cause a processor to implement the method of any one or more of paragraphs Al to A24.
[0278] A28. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method of any one or more of paragraphs Al to A24.
[0279] A29. A method of storing a bitstream of a video, comprising: generating the bitstream from the video according to a rule; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream comprises one or more independently decodable bitstream portions, each bitstream portion corresponding to one or more coded video pictures of the video, and wherein the rule specifies that a maximum decoding picture buffer size required for decoding the bitstream or the one or more bitstream portions is determined based on a maximum allowed picture size of the one or more coded video pictures corresponding to the bitstream or the one or more bitstream portions.
[0280] A30. A method of storing a bitstream of a video, comprising: generating the bitstream from the video according to a rule; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream comprises one or more independently decodable bitstream portions, each bitstream portion corresponding to one or more coded video pictures of the video, and wherein the rule specifies at least one of a maximum allowed picture size, a maximum allowed picture width, a maximum allowed picture height for the conversion.
[0281] Next, another list of preferred embodiments is provided.
[0282] B1. A method of video processing, comprising performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies that one or more constraints on at least one of a coded picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, or a sum of a number of bytes in a network abstraction layer (NAL) unit of an access unit are based on a sum of picture sizes of each of a plurality of pictures of the access unit.
[0283] B2. The method of solution B1, wherein the sum of the number of bytes in the NAL unit is denoted as NumBytesInNalUnit.
[0284] B3. The method of solution B1 or B2, wherein the rule is independent of a current size of a decoded picture in luma samples (denoted as PicSizeInSamplesY).
[0285] B4. The method of solution B1, wherein the rule specifies that a nominal CPB removal time of an access unit (AU) satisfies a constraint:
[0286] AuNominalRemovalTime[n] - AuCpbRemovalTime[n-1]
[0287] Max(AuSizeInSamplesY[n-1] ÷ MaxLumaSr, fR),
[0288] where AuNominalRemovalTime[n] is the nominal CPB removal time of the nth AU, AuCpbRemovalTime[n-1] is the CPB removal time of the n-1th AU, AuSizeInSamplesY[n-1] is the size of the n-1th AU in samples, MaxLumaSr is the maximum luma sample rate (samples per second), fR is a variable equal to 1 ÷ 300, and where n is an integer greater than zero.
[0289] B5. The method of solution Bl, wherein the rule specifies that the difference between the DPB output times of pictures from different access units (AUs) from the DPB satisfies the constraint:
[0290] DpbOutputInterval[n] ≥ Max(AuSizeInSamplesY[n-1] ÷ MaxLumaSr, fR),
[0291] where DpbOutputInterval[n] is the DPB output time of the nth AU, MaxLumaSr is the maximum luma sample rate (samples per second), AuSizeInSamplesY[n-1] is the size of the n-1th AU in samples, fR is a variable equal to 1 ÷ 300, and where n is an integer greater than zero.
[0292] B6. The method of solution Bl, wherein the rule specifies that the sum of the number of bytes in NAL units satisfies the constraint:
[0293] NumBytesInNalUnit[0] ≤ FormatCapabilityFactor
[0294] × (Max(AuSizeInSamplesY[0] ÷ MaxLumaSr, fR × MaxLumaSr)
[0295] + MaxLumaSr × (AuCpbRemovalTime[0] - AuNominalRemovalTime[0])) ÷ MinCr,
[0296] where NumBytesInNalUnit[0] is the sum of the number of bytes in the NAL units of the first access unit (AU), AuSizeInSamplesY[0] is the size of the first AU in samples, MaxLumaSr is the maximum luma sample rate (samples per second), aucpbreaktime[0] is the CPB removal time of the first AU, AuNominalRemovalTime[0] is the nominal CPB removal time of the first AU, MinCr is the minimum coding level, and fR is a variable equal to 1 ÷ 300.
[0297] B7. A method of video processing, comprising performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies a limit on a maximum number of slices in an access unit.
[0298] B8. The method of solution B7, wherein the rule further specifies that the maximum number of slices in an access unit (denoted as MaxSlicesPerAu) is based on a level of the access unit.
[0299] B9. The method of solution B8, wherein the maximum number of slices in an access unit is determined using the following table:
[0300] Level 1 2 2.1 3 3.1 4 4.1 5 5.1 5.2 6 6.1 6.2 MaxSlicesPerAu 16 16 20 30 40 75 75 200 200 200 600 600 600
[0301] B10. The method of solution B7, wherein the rule further specifies that a constraint on coded picture buffer (CPB) removal time (denoted as AuCpbRemovalTime) for each access unit is based on the limit on the maximum number of slices in an access unit.
[0302] B11. The method of solution B10, wherein the rule specifies that the CPB removal time of the first access unit (denoted as AuCpbRemovalTime[0]) satisfies the constraint:
[0303] NumSlicesPerAu[0] ≤ Min(Max(1, MaxSlicesPerAu x MaxLumaSr / MaxLumaPs x (AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) + MaxSlicesPerAu x AuSizeInSamplesY[0] / MaxLumaPs), MaxSlicesPerAu),
[0304] where MaxSlicesPerAu is the maximum number of slices in an access unit, NumSlicesPerAu[0] is the number of slices in a first access unit (AU), MaxLumaPs is the maximum luma picture size, MaxLumaSr is the maximum luma sample rate (in samples per second), AuCpbRemovalTime[0] is the CPB removal time of the first Au, and AuNominalRemovalTime[0] is the nominal CPB removal time of the first AU.
[0305] B12. The method of any of solution B10, wherein the rule specifies that the difference between consecutive CPB removal times satisfies a constraint:
[0306] NumSlicesPerAu[n] < Min((Max(1, MaxSlicesPerAu x MaxLumaSr / MaxLumaPs x (AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1])), MaxSlicesPerAu)
[0307] where MaxSlicesPerAu is the maximum number of slices in an access unit, NumSlicesPerAu[n] is the number of slices in an nth access unit (AU), MaxLumaPs is the maximum luma picture size, MaxLumaSr is the maximum luma sample rate (in samples per second), AuCpbRemovalTime[n] is the CPB removal time of the nth AU, and AuCpbRemovalTime[n-1] is the CPB removal time of the (n-1)th AU, where n is an integer greater than zero.
[0308] B13. The method of any of solution B7, wherein the rule is independent of the maximum number of slices per picture.
[0309] B14. The method of any of solutions B1 to B13, wherein the conversion comprises decoding the video from the bitstream.
[0310] B15. The method of any of solutions B1 to B13, wherein the conversion comprises encoding the video into the bitstream.
[0311] B16. The method of any of solutions B1 to B13, wherein performing the conversion comprises encoding the video into the bitstream; and storing the bitstream in a non-transitory computer- readable storage medium.
[0312] B17. A video processing apparatus comprising a processor configured to implement a method recited by any one or more of solutions B1 to B16.
[0313] B18. A non-transitory computer-readable storage medium configured to store a bit stream of a video generated by the method described in any one or more of schemes B1 to B16.
[0314] B19. A non-transitory computer-readable storage medium configured to store instructions that cause a processor to implement the method described in any one or more of schemes B1 to B16.
[0315] B20. A video processing device for storing a bitstream, wherein the video processing device is configured to implement any one or more of the methods described in schemes B1 to B16.
[0316] B21. A method for storing a video bitstream, comprising: generating a bitstream from a video according to a rule; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream is organized into one or more access units according to the rule, and wherein the rule specifies that one or more constraints of the access unit on at least one of a codec picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, or a sum of the number of bytes in a network abstraction layer (NAL) unit is based on the sum of the picture sizes of each of a plurality of pictures of the access unit.
[0317] B22. A method for storing a bitstream of a video, comprising: generating a bitstream from a video according to a rule; and storing the bitstream in a non-temporary computer-readable storage medium, wherein the bitstream is organized into one or more access units according to the rule, and wherein the rule specifies a limit on the maximum number of slices in an access unit.
[0318] Next, another list of preferred aspects of some embodiments is provided.
[0319] P1. A video processing method, comprising performing a conversion between a video unit of a video and a codec representation of the video, wherein a maximum picture buffer size used during the conversion is determined from a maximum picture size among pictures in a codec layer of the video unit, wherein the maximum picture buffer size is specific to the codec layer.
[0320] P2. The method according to solution P1, wherein the determination of the maximum picture buffer size is independent of the variables defining the picture size associated with the codec layer.
[0321] P3. A video processing method comprising performing conversion between video units of a video and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that constraints related to a maximum buffer size for decoded pictures apply only to a single-layer codec representation.
[0322] P4. The method according to any of the previous solutions P3, wherein the format rule further provides that, in case the coded representation comprises multiple layers of video, different values of the maximum decoder buffer size apply to different layers.
[0323] P5. A method of video processing, comprising performing a conversion between a video unit of a video and a coded representation of the video, wherein the coded representation conforms to a format rule, the format rule providing that, in case the coded representation comprises multiple layers, values of a maximum picture size, a picture width and a picture height of a picture are separately specified within each coded layer of the coded representation of the video.
[0324] P6. The method according to solution P5, wherein the values are specified at sequence parameter set level.
[0325] P7. The method according to any of the previous solutions P1 to P6, wherein performing the conversion comprises encoding the video to generate the coded representation.
[0326] P8. The method according to any of the previous solutions P1 to P6, wherein performing the conversion comprises parsing and decoding the coded representation to generate the video.
[0327] P9. A video decoding apparatus comprising a processor configured to implement one or more of the methods described in solutions P1 to P8.
[0328] P10. A video encoding apparatus comprising a processor configured to implement one or more of the methods described in solutions P1 to P8.
[0329] P11. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement the method of any of the previous solutions P1 to P8.
[0330] In the present document, the term “video processing” can refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. A bitstream representation of a current video block (or simply, bitstream) can for example correspond to bits that are co-located within the bitstream or scattered in different locations within the bitstream, as defined by the syntax. For example, a macroblock can be encoded according to a transformed and encoded error residual value, and also using bits in a header and other fields in the bitstream.
[0331] Implementations of the subject matter and the functional operations described in this specification can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0332] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0333] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0334] Processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, to receive data from or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0335] Although this patent document contains many details, these details should not be interpreted as limitations on any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, in some cases one or more features in the claimed combination may be excluded from the combination, and the claimed combination may involve a subcombination or a variation of the subcombination.
[0336] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve the desired effect. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0337] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of video processing, comprising: performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies a limit on a maximum number of slices in an access unit, wherein the rule is independent of a maximum number of slices per picture.
2. The method of claim 1, wherein the rule further specifies that the maximum number of slices in the access unit (denoted as MaxSlicesPerAu) is based on a level of the access unit.
3. The method of claim 2, wherein, The following table is used to determine the maximum number of slices in the access unit. The rule further specifies that a constraint on a coded picture buffer (CPB) removal time (denoted as AuCpbRemovalTime) for each access unit is based on the limit on the maximum number of slices in the access unit.
4. The method of claim 1, wherein, The rule specifies that a CPB removal time (denoted as AuCpbRemovalTime[ 0 ]) of a first access unit satisfies a constraint:
5. The method of claim 4, wherein, NumSlicesPerAu[ 0 ] ≤ Min( Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × ( AuCpbRemovalTime[ 0 ] − AuNominalRemovalTime[ 0 ] ) +MaxSlicesPerAu × AuSizeInSamplesY[ 0 ] / MaxLumaPs ), MaxSlicesPerAu ), where MaxSlicesPerAu is the maximum number of slices of the access unit, NumSlicesPerAu[ 0 ] is a number of slices in the first access unit (AU), MaxLumaPs is a maximum luma picture size, MaxLumaSr is a maximum luma sample rate (samples per second), AuCpbRemovalTime[ 0 ] is the CPB removal time of the first access unit, and AuNominalRemovalTime[ 0 ] is a nominal CPB removal time of the first access unit. The rule specifies that a difference between consecutive CPB removal times satisfies a constraint:
6. The method of claim 4, wherein, NumSlicesPerAu[ n ] ≤ Min( (Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × ( AuCpbRemovalTime[ n ] − AuCpbRemovalTime[ n − 1 ] ) ),MaxSlicesPerAu ) where MaxSlicesPerAu is the maximum number of slices in the access unit, NumSlicesPerAu[ n ] is the number of slices of the n-th access unit (AU), MaxLumaPs is the maximum luma picture size, MaxLumaSr is the maximum luma sample rate (samples per second), AuCpbRemovalTime[ n ] is the CPB removal time of the n-th access unit, AuCpbRemovalTime[ n−1 ] is the nominal CPB removal time of the (n−1)-th access unit, and where n is an integer greater than 0.
7. The method of claim 1, wherein, The rules further provide that one or more constraints on at least one of the coded picture buffer (CPB) removal time, the nominal CPB removal time, the decoded picture buffer (DPB) output time, and the sum of the number of bytes in network abstraction layer (NAL) units for an access unit are based on a sum of picture sizes of each of a plurality of pictures of the access unit.
8. The method of claim 7, wherein, The sum of the number of bytes in the NAL units is denoted as NumBytesInNalUnit.
9. The method of claim 7 or 8, wherein, The rules are independent of a current size of a decoded picture in luma samples (denoted as PicSizeInSamplesY).
10. The method of claim 7, wherein, The rules provide that the nominal CPB removal time of an access unit satisfies a constraint: AuNominalRemovalTime[ n ] − AuCpbRemovalTime[ n−1 ] ≥ Max( AuSizeInSamplesY[ n−1 ] ÷ MaxLumaSr, fR ), where AuNominalRemovalTime[ n ] is the nominal CPB removal time of the n-th access unit, AuCpbRemovalTime[ n−1 ] is the nominal CPB removal time of the (n−1)-th access unit, AuSizeInSamplesY[ n−1 ] is the size of the (n−1)-th access unit in samples, MaxLumaSr is the maximum luma sample rate (samples per second), fR is a variable equal to 1 ÷ 300, and where n is an integer greater than 0.
11. The method of claim 7, wherein, The rules provide that a difference in DPB output times of pictures from different access units (AUs) of a DPB satisfies a constraint: DpbOutputInterval[ n ] ≥ Max( AuSizeInSamplesY[ n−1 ] ÷ MaxLumaSr, fR ), where DpbOutputInterval[ n ] is the DPB output time of the n-th access unit, MaxLumaSr is the maximum luma sample rate (samples per second), AuSizeInSamplesY[ n−1 ] is the size of the (n−1)-th access unit in samples, fR is a variable equal to 1 ÷ 300, and where n is an integer greater than 0.
12. The method of claim 7, wherein, The rule specifies that a sum of a number of bytes in the NAL units satisfies a constraint: NumBytesInNalUnit[ 0 ] ≤ FormatCapabilityFactor × ( Max( AuSizeInSamplesY[ 0 ] ÷ MaxLumaSr, fR × MaxLumaSr ) + MaxLumaSr × ( AuCpbRemovalTime[ 0 ] − AuNominalRemovalTime[ 0 ] ) ) ÷ MinCr, where NumBytesInNalUnit[ 0 ] is a sum of a number of bytes in NAL units of a first access unit (AU), AuSizeInSamplesY[ 0 ] is a size of the first access unit in samples, MaxLumaSr is a maximum luma sample rate (samples per second), AuCpbRemovalTime[ 0 ] is a CPB removal time of the first access unit, AuNominalRemovalTime[ 0 ] is a nominal CPB removal time of the first access unit, MinCr is a minimum compression level, and fR is a variable equal to 1 ÷ 300.
13. The method of any one of claims 1-8, 10-12, wherein, The conversion includes decoding the video from the bitstream.
14. The method of any one of claims 1-8, 10-12, wherein, The conversion includes encoding the video into the bitstream.
15. The method of any one of claims 1-8, 10-12, wherein, Performing the conversion includes: encoding the video into the bitstream; and storing the bitstream in a non-transitory computer-readable recording medium.
16. A video processing apparatus comprising a processor configured to perform the method of any one of claims 1-15.
17. A non-transitory computer-readable storage medium configured to store instructions that cause a processor to implement the method of any one of claims 1-15.
18. A method of storing a bitstream of a video, comprising: generating the bitstream from the video according to a rule; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the bitstream is organized into one or more access units according to the rule, and wherein the rule specifies a limit on a maximum number of slices in an access unit, wherein the rule is independent of a maximum number of slices per picture.
19. The method of Claim 18, wherein, The rule further specifies that one or more constraints on at least one of a coded picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, and a sum of a number of bytes in network abstraction layer (NAL) units for an access unit are based on a sum of picture sizes of each of a plurality of pictures of the access unit.