Video processing method and apparatus, medium, method of storing bitstream of video

By updating the MaxDpbSize calculation formula and setting DPB and CPB constraints for each layer, the problem of improper DPB and CPB management in multi-layer video encoding and decoding was solved, thus improving video encoding and decoding efficiency.

CN114846792BActive Publication Date: 2025-10-28DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080090395.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-26
Filing Date
2020-12-27
Publication Date
2025-10-28
Estimated Expiration
2040-12-27

AI Technical Summary

Technical Problem

The existing VVC level definitions do not apply to constraints such as image size, DPB size, and CPB time in multi-layer video encoding and decoding. This leads to improper management of DPB size and CPB time by the decoder, affecting the efficiency of video encoding and decoding.

Method used

By updating the calculation formula for MaxDpbSize, using the maximum image size within CLVS and the sum of the image sizes of all images in AU, constraints are set for DPB and CPB for each layer, ensuring that the DPB size and CPB time management for each layer are reasonable.

Benefits of technology

It enables effective management of DPB and CPB in multi-layer video encoding and decoding, improving video encoding and decoding efficiency and decoder processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114846792B_ABST
    Figure CN114846792B_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatuses for signaling notification at the picture buffer level as part of video encoding / decoding or decoding are described. An example method for video processing includes: performing a conversion between video and a video bitstream according to rules, wherein the bitstream comprises one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and the rules specify that the maximum decoded picture buffer size required to decode the bitstream or one or more bitstream portions is determined based on the maximum permissible picture size of the one or more codec video pictures corresponding to the bitstream or one or more bitstream portions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] Pursuant to applicable patent law and / or the Paris Convention, this application promptly claims priority and interest in U.S. Provisional Application No. 62 / 953,815, filed December 26, 2019. For all legal purposes, the entire disclosure of the aforementioned application is incorporated herein by reference. Technical Field

[0003] This patent document relates to image and video encoding / decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to send image buffer levels as part of the video encoding or decoding process.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to rules. The bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and the rules specify that the maximum decoded picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on the maximum allowed picture size of the one or more codec video pictures corresponding to the bitstream or the one or more bitstream portions.

[0007] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising video units and a bitstream of a video comprising one or more codec layer video sequences, wherein the bitstream conforms to a rule, and wherein the rule specifies that the maximum buffer size of the decoded images of the codec layer video sequences is constrained to be less than or equal to the maximum image size selected from the images in the codec layer video sequences.

[0008] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a bitstream of the video according to rules. The bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video images of the video, and the rules specify at least one of a maximum allowed image size, a maximum allowed image width, and a maximum allowed image height for the conversion.

[0009] In another example, a different video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video. The bitstream includes one or more output layer sets, wherein at least one output layer set includes multiple video layers. Furthermore, the bitstream conforms to a rule that specifies a constraint on the overall size of a decoded image buffer comprising the at least one output layer set including the multiple video layers.

[0010] In another example, a different video processing method is disclosed. This method includes performing a conversion between video and a video bitstream, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies one or more constraints on at least one of the following for an access unit: the sum of the codec picture buffer (CPB) removal time, the nominal CPB removal time, the decode picture buffer (DPB) output time, or the sum of the number of bytes in a Network Abstraction Layer (NAL) unit, based on the sum of the picture sizes of each of a plurality of pictures in that access unit.

[0011] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies a limit on the maximum number of stripes in the access unit.

[0012] In another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0013] In another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0014] In another example, a computer-readable medium storing code is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0015] These and other features are described in this document. Attached Figure Description

[0016] Figure 1This is a block diagram illustrating an example video processing system in which various technologies disclosed in this document can be implemented.

[0017] Figure 2 This is a block diagram of an example hardware platform used for video processing.

[0018] Figure 3 A block diagram of an example video encoding system that can implement some embodiments of the present disclosure is shown.

[0019] Figure 4 A block diagram of an example encoder that can implement some embodiments of the present disclosure is shown.

[0020] Figure 5 A block diagram of an example decoder that can implement some embodiments of the present disclosure is shown.

[0021] Figures 6-9 A flowchart of an example method for video processing is shown. Detailed Implementation

[0022] The use of chapter headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.

[0023] 1. Overview

[0024] This document relates to video codec technology. Specifically, it concerns the levels of video codecs that support single-layer and multi-layer video codecs. It can be applied to any standard or non-standard video codec that supports single-layer and multi-layer video codecs, such as the under-development Multi-Functional Video Codec (VVC).

[0025] 2. Abbreviations

[0026] APS Adaptive Parameter Set

[0027] AU Access Unit

[0028] AUD (Access Unit Delimiter)

[0029] AVC (Advanced Video Coding)

[0030] CLVS (Coded Layer Video Sequence)

[0031] CPB Coded Picture Buffer

[0032] CRA Clean Random Access

[0033] CTU (Coding Tree Unit)

[0034] CVS (Coded Video Sequence)

[0035] DPB Decoded Picture Buffer

[0036] DPS Decoding Parameter Set

[0037] End of EOB bitstream

[0038] End of EOS sequence

[0039] GDR Gradual Decoding Refresh

[0040] HEVC (High Efficiency Video Coding)

[0041] Instantaneous Decoding Refresh (IDR)

[0042] JEM Joint Exploration Model

[0043] MCTS Motion-Constrained Tile Sets

[0044] NAL (Network Abstraction Layer)

[0045] OLS Output Layer Set

[0046] PH Picture Header

[0047] PPS Image Parameter Set

[0048] PU Picture Unit

[0049] RBSP Raw Byte Sequence Payload

[0050] SEI Supplemental Enhancement Information

[0051] SPS Sequence Parameter Set

[0052] VCL (Video Coding Layer)

[0053] VPS Video Parameter Set

[0054] VTM VVC Test Model

[0055] VUI Video Usability Information

[0056] VVC (Versatile Video Coding)

[0057] 3. Preliminary Discussion

[0058] Video coding standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture, utilizing time prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new coding standards is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.

[0059] 3.1 Grades, Levels, and Classes

[0060] Video codec standards typically specify levels and grades. Some video codec standards also define layers, such as HEVC and the developing VVC.

[0061] Grades, layers, and levels specify limitations on the bitstream and thus limit the capabilities required to decode it. Grades, layers, and levels can also be used to indicate interoperability points between different decoder implementations.

[0062] Each grade specifies a subset of constraints and algorithmic features supported by all decoders that conform to that grade. Note that the encoder does not need to use all the codec tools or features supported in the grade, while the decoder that conforms to the grade needs to support all the codec tools or features.

[0063] Each level of a layer specifies a set of restrictions on the values ​​that bitstream syntax elements can take. All grades typically use the same set of layer and level definitions, but different implementations can support different layers, and within a layer, each supported grade can support different levels. For any given grade, the level of the layer typically corresponds to the specific decoder processing load and storage capacity.

[0064] The capabilities of a video decoder that conforms to a video codec specification are defined by its ability to decode video streams that conform to the constraints of the tiers, layers, and levels specified in the video codec specification. When expressing the capabilities of a decoder for a specific tier, the layers and levels supported by that tier should also be expressed.

[0065] 3.2 Existing VVC Levels and Layer Definitions

[0066] The latest VVC draft text in JVET-P2001-v14 is available here: http: / / phenix.int-evry.fr / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P2001-v14.zip, with the following level definitions.

[0067] A.1.1 General Layer and Level Restrictions

[0068] For the purpose of comparing layer capabilities, a layer with general_tier_flag equal to 0 is considered a lower layer than a layer with general_tier_flag equal to 1.

[0069] For the purpose of comparing level capabilities, when the value of general_level_idc or sublayer_level_idc[i] for a specific level is less than the values ​​of other levels, the specific level of the specified layer is considered to be a lower level than other levels in the same layer.

[0070] To express the constraints in this appendix, the following are specified:

[0071] – Let access unit n be the nth access unit in the decoding order, and the first access unit be access unit 0 (i.e. the 0th access unit).

[0072] – Let image n be the encoded or decoded image of access unit n, or the corresponding decoded image.

[0073] When the specified level is not level 8.5, bitstreams conforming to the specified tier and level shall comply with the following constraints for each bitstream conformance test as specified in Appendix C:

[0074] a) PicSizeInSamplesY should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.

[0075] b) The value of pic_width_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).

[0076] c) The value of pic_height_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).

[0077] d) The value of num_tile_columns_minus1 should be less than MaxTileCols, and the value of num_tile_rows_minus1 should be less than MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.

[0078] e) For the VCL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] shall be less than or equal to CpbVclFactor*MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 according to the parameter specified in Clause C.1, CpbVclFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in CpbVclFactor bits.

[0079] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] shall be less than or equal to CpbNalFactor*MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 according to the parameter specified in Clause C.1, CpbNalFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in CpbNalFactor bits.

[0080] Table A.1 specifies the restrictions for each level of each layer, except for level 8.5.

[0081] The layer and level of the bitstream conformance are indicated by the syntax elements general_tier_flag and general_level_idc, and the sublayer level is indicated by the syntax element sublayer_level_idc[i], as follows:

[0082] – If the specified level is not level 8.5, then according to the level constraints specified in Table A.1, a general_tier_flag of 0 indicates compliance with the Main tier, and a general_tier_flag of 1 indicates compliance with the Hightier. According to the layer constraints specified in Table A.1, for layers below level 4, general_tier_flag should be equal to (corresponding to the entry marked with "-" in Table A.1). Otherwise (if the specified level is level 8.5), the requirement for bitstream consistency is that general_tier_flag should be equal to 1, and the value of general_tier_flag 0 is reserved for future use by ITU-T|ISO / IEC, and the decoder should ignore the value of general_tier_flag.

[0083] –general_level_IDC and sublayer_level_idc[i] should be set to 30 times the number of levels specified in Table A.1.

[0084] Table A.1 - General Layer and Level Restrictions

[0085]

[0086] A.1.2 Tier - Specified Level Restrictions

[0087] To express the constraints in this appendix, the following are specified:

[0088] – Let variable fR equal 1 ÷ 300.

[0089] The variable HbrFactor is defined as follows:

[0090] – If the bitstream is indicated to conform to main 10 or main 4:4:4 10, then HbrFactor is set to equal to 1.

[0091] The variable BrVclFactor, representing the VCL bit rate scaling factor, is set to be equal to CpbVclFactor * HbrFactor.

[0092] The variable BrNalFactor, representing the NAL bit rate scaling factor, is set to be equal to CpbNalFactor * HbrFactor.

[0093] The variable MinCr is set to equal MinCrBase * MinCrScaleFactor ÷ HbrFactor.

[0094] When the specified level is not level 8.5, the value of max_dec_pic_buffering_minus1[Htid]+1 should be less than or equal to MaxDpbSize, as deduced below:

[0095]

[0096] MaxLumaPs is specified in Table A.1, and maxDpbPicBuf is equal to 8.

[0097] Bitstreams conforming to the specified tier and level of Master 10 or Master 4:4:4 10 tiers shall comply with the following constraints for each bitstream conformance test as specified in Appendix C:

[0098] a) As specified in Clause C.2.3, the nominal removal time for removing access unit n (n greater than 0) from the CPB shall satisfy the following constraint: for the value of PicSizeInSamplesY of picture n-1, AuNominalRemovalTime[n]-AuCpbRemovalTime[n-1] is greater than or equal to Max(PicSizeInSamplesY÷MaxLumaSr,fR), where MaxLumaSr is the value specified for picture n-1.

[0099] (b) As specified in Clause C.3.3, the time difference between consecutive outputs of images in the DPB shall satisfy the following constraint: for the PicSizeInSamplesY value of image n, DpbOutputInterval[n] is greater than or equal to Max(PicSizeInSamplesY÷MaxLumaSr,fR), where MaxLumaSr is the value specified for image n in Table A.2, provided that image n is the output image and not the last image of the bitstream output.

[0100] c) For the value of PicSizeInSamplesY for image 0, the removal time of access unit 0 should satisfy the constraint that the number of slices in image 0 is less than or equal to Min(Max(1,MaxSliceSegmentsPerPicture*MaxLumaSr / MaxLumaPs*AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxSliceSegmentsPerPicture*PicSizeInSamplesY / MaxLumaPs), where MaxSliceSegmentsPerPicture, MaxLumaPs, and MaxLumaSr are the values ​​specified in Tables A.1 and A.2 respectively for image 0.

[0101] d) The difference in consecutive CPB removal times between access unit n and access unit n-1 (n > 0) should satisfy the condition that the number of strip segments in image n is less than or equal to Min((Max(1,MaxSliceSegmentsPerPicture*MaxLumaSr / MaxLumaPs*).

[0102] (AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxSliceSegmentsPerPicture), where MaxSliceSegmentsPerPicture, MaxLumaPs, and MaxLumaSr are the values ​​specified in Tables A.1 and A.2, respectively, applicable to picture n.

[0103] e) For parameter VCL HRD, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] shall be less than or equal to BrVclFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.

[0104] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1, BitRate[Htid][i] shall be less than or equal to BrNalFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.

[0105] g) For the value of PicSizeInSamplesY for image 0, the sum of the NumBytesInNalUnit variables in access unit 0 should be less than or equal to FormatCapabilityFactor*(Max(PicSizeInSamplesY,fR*MaxLumaSr)+

[0106] MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0]))÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values ​​specified in Tables A.2 and A.3 for image 0, respectively.

[0107] h) The sum of the NumBytesInNalUnit variables of the access unit n (n is greater than 0) should be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values ​​specified in Tables A.2 and A.3 for the image n, respectively.

[0108] i) For the value of picsizesinsamplesy for image 0, the removal time of access unit 0 should satisfy the constraint that the number of pieces in image 0 is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxTileCols*MaxTileRows*PicSizeInSamplesY / MaxLumaPs), where MaxTileCols and MaxTileRows are the values ​​specified in Table A.1 applicable to image 0.

[0109] j) The difference between the consecutive CPB removal times of access units n and n-1 (n is greater than 0) should satisfy the constraint that the number of tiles in picture n is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values ​​specified in Table A.1 applicable to picture n.

[0110]

[0111] Table A.2 - Video Qualification Layers and Level Restrictions

[0112]

[0113] Table A.3 - Specifications of CpbVclFactor, CpbNalFactor, FormatCapabilityFactor andMinCrScaleFactor

[0114]

[0115] 4. Example of a technical solution: The technical problem to be solved

[0116] The existing VVC design with defined levels has the following problems:

[0117] 1) To derive the equation for the variable MaxDpbSize, i.e., equation (A.1) above, we use the variable PicSizeInSamplesY, which is the image size of a specific image. However, in VVC, the image size can vary from one image to another within a codec layer video sequence (CLVS). Therefore, the maximum image size within the images in the CLVS should be used to derive MaxDpbSize. Furthermore, because there can be multiple layers, and different layers in the OLS bitstream have different maximum image size values ​​within the images in the CLVS, it is necessary to derive the value of MaxDpbSize for each layer.

[0118] 2) The constraint related to DPB size specifies that the value of max_dec_pic_buffering_minus1[]+1 should be less than or equal to MaxDpbSize. This constraint applies only to single-layer bitstreams. For output layer sets (OLS) containing multiple layers in their bitstreams, the existing constraint does not apply because multiple different instances of the max_dec_pic_buffering_minus1[] syntax element can be applied to different layers. Instead, an appropriate instance of the max_dec_pic_buffering_minus1[] syntax element needs to be used to constrain each layer.

[0119] 3) The constraints on image size, width, and height are defined as items a, b, and c in section A.1.1 above, using the variable `PicSizeInSamplesY` and the syntax elements `pic_width_in_luma_samples` and `pic_height_in_luma_samples`. Similarly, since the image size within a CLVS can vary from one image to another in VVC, the maximum image size, width, and height values ​​should be used. Furthermore, because there can be multiple layers, and different layers in the OLS bitstream have different maximum image size, width, and height values ​​within the CLVS images, constraints need to be specified individually for each layer or each SPS referenced by the layer.

[0120] 4) In addition, there is a need to constrain the overall DPB size for OLS containing multiple layers, so that the same DPB size constraint of the decoder will take effect regardless of the number of layers. This is currently missing.

[0121] 5) The AU's constraints on the sum of CPB removal time, nominal CPB removal time, DPB output time, and the NumBytesInNalUnit variable are defined as items a, b, c, d, g, and i in section A.1.2 above. The variable PicSizeInSamplesY is used, which is the image size of a specific image. However, since an AU can contain multiple images, the sum of the image sizes of all images in the AU should be used.

[0122] 6) A limit should be specified for the maximum number of stripes per AU and used when specifying these constraints, rather than specifying a limit for the maximum number of stripes per image for each level and using this limit when specifying some constraints for the CPB removal time for the AU (i.e., items c and d in section A.1.2 above).

[0123] 5. Example Implementations and Technologies

[0124] To address the aforementioned and other issues, the following summarized methods are published in a list of various projects. These projects should be considered as examples for explaining general concepts, rather than for interpreting them in a narrow way. Furthermore, these projects can be used individually or in combination in any way.

[0125] 1) To address the first issue, the equation for the exported variable MaxDpbSize is updated to use the maximum image size in the CLVS image instead of the variable PicSizeInSamplesY, and a value for MaxDpbSize is exported for each layer.

[0126] 2) To address the second problem, for each layer, the constraint requires that the value of a specific instance of the `max_dec_pic_buffering_minus1[]` syntax element plus 1 be less than or equal to the value of `MaxDpbSize` for that layer. The specific instance of the `max_dec_pic_buffering_minus1[]` syntax element is the `max_dec_pic_buffering_minus1[i]` where `i` is the maximum value in the `dpb_parameters()` syntax structure. This is determined as follows: if the layer is an output layer of the OLS, then the `dpb_parameters()` syntax structure is the syntax structure applied to the layer when it is an output layer in the OLS; otherwise, the `dpb_parameters()` syntax structure is the `dpb_parameters()` syntax structure applied to the layer when it is not an output layer in the OLS.

[0127] 3) To address the third issue, the constraints on image size, width, and height are defined using the maximum image size, width, and height, instead of the variables `PicSizeInSamplesY` and the syntax elements `pic_width_in_luma_samples` and `pic_height_in_luma_samples`. Furthermore, constraints are specified separately for each layer or each SPS referenced by a layer.

[0128] 4) To address the fourth problem, constraints are specified on the overall DPB size of OLS containing multiple layers.

[0129] a. In an alternative solution, the constraints are specified as follows:

[0130] After each AU is decoded, immediately let numDecPics be the number of decoded images in DPB, and let picSizeInSamplesY[i] be the value of picSizeInSamplesY for the i-th decoded image in DPB, where i is in the range from 0 to numDecPics-1 (inclusive). The value should be less than or equal to maxDpbPicBuf*MaxLumaPs, where maxDpbPicBuf is equal to 8 and MaxLumaPs is specified in Table A.1.

[0131] b. In another alternative, the constraints are specified as follows:

[0132] Let numLayers be the number of layers in OLS, and maxDecBuff[i] and PicSizeMaxInSamplesY[i] be the values ​​of max_dec_pic_buffering_minus1[maxTid] and PicSizeMaxInSamplesY for the i-th layer, respectively, where i is in the range from 0 to numLayers-1 (inclusive). PicSizeMaxInSamplesY is equal to pic_width_max_in_luma_samples * pic_height_max_in_luma_samples for that layer, and the layer's max_dec_pic_buffering_minus1[maxTid] is the value of max_dec_pic_buffering_minus1[i] with the maximum value of i in the dpb_parameters() syntax structure, determined as follows:

[0133] If the layer is an output layer of OLS, the dpb_parameters() syntax structure is the syntax structure applied to the layer when the layer is an output layer in OLS; otherwise, the dpb_parameters() syntax structure is the syntax structure applied to the layer when the layer is not an output layer in OLS. The value should be less than or equal to maxDpbPicBuf*MaxLumaPs, where maxDpbPicBuf is equal to 8 and MaxLumaPs is specified in Table A.1.

[0134] 5) To address the fifth issue, the definitions of several constraints for the sum of the CPB removal time, nominal CPB removal time, DPB output time, and NumBytesInNalUnit variables in the AU are based on the sum of the image sizes of all images in the AU, instead of the variable PicSizeInSamplesY.

[0135] 6) To address the sixth issue, specify a maximum number of stripes per AU and use that limit when specifying these constraints, instead of specifying a maximum number of stripes per image for each level, and use that limit when specifying some constraints on the CPB removal time for AUs.

[0136] 6 Examples

[0137] The following are some example implementations that can be applied to the VVC specification. The revised text is based on the latest VVC text in JVET-P2001-v14. Most of the relevant additions or modifications are marked in bold double brackets. There are also some editorial changes, which are not highlighted.

[0138] 6.1 First Implementation Example 6.1.1 Overview of Grades, Layers, and Levels

[0139] Grades, layers, and levels define limitations on the bitstream and thus limit the capabilities required to decode it. Grades, layers, and levels can also be used to indicate interoperability points between different decoder implementations.

[0140] Each tier specifies a subset of the constraints and algorithmic features supported by all decoders that conform to that tier.

[0141] Note – The encoder does not need to use any specific subset of features supported in the grade.

[0142] Each level of a layer specifies a set of restrictions on the values ​​that the syntax elements of this specification can take. All grades typically use the same set of layer and level definitions, but different implementations may support different layers, and within a layer, each supported grade may support different levels. For any given grade, the level of the layer typically corresponds to the specific decoder processing load and storage capacity.

[0143] In this clause, phrases like “bitstream conformance” should be interpreted as “bitstream conformance of OLS”.

[0144] 6.1.2 Overview of Layer and Level Display

[0145] For the purpose of comparison layer functionality, a layer with general_tier_flag equal to 0 is considered a lower layer than a layer with general_tier_flag equal to 1.

[0146] For the purpose of comparing level capabilities, when the value of general_level_idc or sublayer_level_idc[i] for a specific level is less than the values ​​of other levels, the specific level of the specified layer is considered to be a lower level than other levels in the same layer.

[0147] To express the constraints in this appendix, the following are specified:

[0148] – Let AU n be the nth AU in the decoding order, and the first AU be AU 0 (i.e. the 0th AU).

[0149] – For a specific reference SPS, let PicSizeMaxInSamplesY be the value of pic_width_max_in_luma_samples * pic_height_max_in_luma_samples.

[0150] When the specified level is not level 8.5, the value of MaxDpbSize for each level in OLS is derived as follows:

[0151]

[0152]

[0153] MaxLumaPs is specified in Table A.1, and maxDpbPicBuf is equal to 8.

[0154] When the specified level is not level 8.5, bitstreams conforming to the specified tier and level shall comply with the following constraints for each bitstream conformance test as specified in Appendix C:

[0155] a) {{PicSizeInSamplesY for each reference SPS}} should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.

[0156] b) For each reference SPS, the value of pic_width_max_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).

[0157] c) For each reference SPS, the value of pic_height_max_in_luma_samples should be less than or equal to Sqrt(MaxLumaPs*8).

[0158] d) For each reference PPS, the value of NumTileColumns should be less than MaxTileCols, and the value of NumTileRows should be less than MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.

[0159] e) For the VCL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] shall be less than or equal to CpbVclFactor*MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 according to the parameter specified in Clause C.1, CpbVclFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in CpbVclFactor bits.

[0160] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] shall be less than or equal to CpbNalFactor*MaxCPB, where CpbSize[Htid][i] is specified in Clause 7.4.6.3 according to the parameter specified in Clause C.1, CpbNalFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in CpbNalFactor bits.

[0161] g) For each layer in OLS, the value of max_dec_pic_buffering_minus1[maxTid]+1 should be less than or equal to the MaxDpbSize of that layer, where max_dec_pic_buffering_minus1[maxTid] is the value of max_dec_pic_buffering_minus1[i] when i takes the maximum value in the dpb_parameters() syntax structure, which is determined as follows: if the layer is an output layer of OLS, then the dpb_parameters() syntax structure is the syntax structure applied to the layer when the layer is an output layer in OLS; otherwise, when the layer is not an output layer in OLS, the dpb_parameters() syntax structure will be applied.

[0162] h) After each AU is decoded, immediately let numDecPics be the number of decoded images in the DPB, and let picSizeInSamplesY[i] be the value of picSizeInSamplesY for the i-th decoded image in the DPB, where i is in the range from 0 to numDecPics-1 (inclusive). The value should be less than or equal to maxDpbPicBuf * MaxLumaPs, where maxDpbPicBuf equals 8, and MaxLumaPs is specified in Table A.1.

[0163] Table A.1 specifies the restrictions for each level except for level 8.5.

[0164] The layer and level of the bitstream conformance are indicated by the syntax elements general_tier_flag and general_level_idc, and the sublayer level is indicated by the syntax element sublayer_level_idc[i], as follows:

[0165] – If the specified level is not level 8.5, then according to the level restrictions specified in Table A.1, a general_tier_flag of 0 indicates compliance with the primary layer, a general_tier_flag of 1 indicates compliance with higher layers, and for levels lower than level 4, general_tier_flag should be equal to 0 (corresponding to the entries marked with "-" in Table A.1). Otherwise (if the specified level is level 8.5), the requirement for bitstream consistency is that general_tier_flag should be equal to 1, and the value of general_tier_flag 0 is reserved for future use by ITU-T|ISO / IEC, and the decoder should ignore the value of general_tier_flag.

[0166] –general_level_IDC and sublayer_level_idc[i] should be set to values ​​equal to 30 times the number of levels specified in Table A.1.

[0167] Table A.1 - General Layer and Level Restrictions

[0168]

[0169] 6.1.3 Tier - Specified Level Restrictions

[0170] To express the constraints in this appendix, the following are specified:

[0171] – Let variable fR equal 1 ÷ 300.

[0172] The variable HbrFactor is defined as follows:

[0173] – If the bitstream is indicated to conform to main 10 or main 4:4:4 10, then HbrFactor is set to equal to 1.

[0174] The variable BrVclFactor, representing the VCL bit rate scaling factor, is set to be equal to CpbVclFactor * HbrFactor.

[0175] The variable BrNalFactor, representing the NAL bit rate scaling factor, is set to be equal to CpbNalFactor * HbrFactor.

[0176] The variable MinCr is set to equal MinCrBase * MinCrScaleFactor ÷ HbrFactor.

[0177] The variable AuSizeInSamplesY[n] is set to equal to Where numDecPics is the number of images in AU n, and picSizeInSamplesY[i] is the picSizeInSamplesY value of the i-th image in AU n, where i is in the range from 0 to numDecPics-1 (inclusive).

[0178] Bitstreams conforming to the specified tier and level of Master 10 or Master 4:4:4 10 tiers shall comply with the following constraints for each bitstream conformance test as specified in Appendix C:

[0179] a) As specified in Clause C.2.3, the nominal removal time for removing AU n (n greater than 0) from the CPB shall satisfy the following constraint: AuNominalRemovalTime[n] - AuCpbRemovalTime[n-1] is greater than or equal to Max({{PicSizeInSamplesY[n-1]}} ÷ MaxLumaSr,fR), where MaxLumaSr is the value specified in Table A.2 applicable to AU n-1.

[0180] b) As specified in Clause C.3.3, the difference in consecutive output times of images from different AUs in the DPB shall satisfy the following constraint: DpbOutputInterval[n] is greater than or equal to Max({{AuSizeInSamplesY[n]}}MaxLumaSr,fR), where MaxLumaSr is the value specified for AU n in Table A.2, provided that the AU has an output image and AU n is not the last AU in the bitstream with an output image.

[0181] c) The CPB removal time of AU 0 should satisfy the condition that the number of stripes in AU 0 is less than or equal to Min(Max(1,{{MaxSlicesPerAu}}*MaxLumaSr / MaxLumaPs*).

[0182] The constraints are: AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+{{MaxSlicesPerAu}}*{{AuSizeInSamplesY[0]}} / MaxLumaPs),{{MaxSlicesPerAu}}), where {{MaxSlicesPerAu}},MaxLumaPs and MaxLumaSr are the values ​​applicable to AU 0 specified in Tables A.1 and A.2, respectively.

[0183] d) The difference between consecutive CPB removal times of AU n and AU n-1 (n > 0) should satisfy the constraint that the number of stripes in AU n is less than or equal to Min((Max(1,{{MaxSlicesPerAu}}*MaxLumaSr / MaxLumaPs*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),{{MaxSlicesPerAu}}), where {{MaxSlicesPerAu}}, MaxLumaPs, and MaxLumaSr are the values ​​applicable to AU n specified in Tables A.1 and A.2, respectively.

[0184] e) For parameter VCL HRD, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] shall be less than or equal to BrVclFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.

[0185] f) For the NAL HRD parameter, for at least one i value in the range of 0 to hrd_cpb_cnt_minus1, BitRate[Htid][i] shall be less than or equal to BrNalFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.

[0186] g) The sum of the NumBytesInNalUnit variables of AU 0 should be less than or equal to FormatCapabilityFactor*(Max({{AuSizeInSamplesY[0]}},fR*MaxLumaSr)

[0187] +MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0]))÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values ​​specified in Tables A.2 and A.3 for AU 0, respectively.

[0188] h) The sum of the NumBytesInNalUnit variables for AU n (n greater than 0) should be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values ​​specified in Tables A.2 and A.3 for AU n, respectively.

[0189] i) The removal time of AU 0 should satisfy the constraint that the number of pieces of each picture in AU 0 is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxTileCols*MaxTileRows*{{AuSizeInSamplesY[0]}} / MaxLumaPs),MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values ​​specified in Table A.1 applicable to AU 0.

[0190] j) The difference between consecutive CPB removal times of AU n and AU n-1 (n > 0) should satisfy the constraint that the number of tiles for each picture in AU n is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values ​​specified in Table A.1 applicable to AU n.

[0191]

[0192] Figure 1 This is a block diagram of an example video processing system 1000 that can perform various techniques described in this disclosure. Various implementations may include some or all of the components of system 1000. System 1000 may include input 1002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values) or in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Networking (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0193] System 1000 may include an encoding / decoding component 1004 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 1004 can reduce the average bit rate of the video from input 1002 to the output of encoding / decoding component 1004 to produce an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1006, the output of encoding / decoding component 1004 can be stored or transmitted via connected communication. The stored or transmitted bitstream representation (or encoded / decoded representation) of the video received at input 1002 can be used by component 1008 to generate pixel values ​​or displayable video, which is then sent to display interface 1010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding result are performed by the decoder.

[0194] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, and so on. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0195] Figure 2 This is a block diagram of a video processing apparatus 2000. Apparatus 2000 can be used to implement one or more methods described herein. Apparatus 2000 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing circuitry 2006. Processor 2002 can be configured to implement one or more methods described in this document (e.g., ...). Figures 6-9 Multiple memories 2004 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 2006 can be used to implement some of the techniques described in this document in hardware circuitry.

[0196] Figure 3 A block diagram of an example video encoding / decoding system 100 that can utilize the technology of the present invention is shown. Figure 3As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0197] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. Encoded pictures are encoded and decoded representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage media / server 130b for access by destination device 120.

[0198] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0199] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.

[0200] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVM) standard, and other current and / or further standards.

[0201] Figure 4 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 3 The video encoder 114 in the system 100 shown.

[0202] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 4In the example, the video encoder 200 includes multiple functional components. The techniques described in this invention can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0203] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 810, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0204] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0205] Furthermore, some components (e.g., motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for interpretative purposes... Figure 4 It is represented separately in the instance.

[0206] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0207] The mode selection unit 203 may, for example, select one of multiple encoding / decoding modes (intra-frame or inter-frame) based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signaling and intra-frame prediction signaling. In the case of inter-frame prediction, the mode selection unit 203 may also select the precision of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0208] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0209] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0210] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and may search for a reference video block for the current video block in the reference images of list 0 or list 7. Motion estimation unit 204 may then generate a reference index indicating the reference images in list 0 or list 7, which includes the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0211] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 7. Motion estimation unit 204 can then generate a reference index and a motion vector, the reference index indicating the reference images in lists 0 and 7 containing the reference video block, and the motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0212] In some examples, the motion estimation unit 204 can output complete motion information for the decoder's decoding processing.

[0213] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block to another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0214] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 900 that the current video block has the same motion information as another video block.

[0215] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0216] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0217] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0218] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0219] In other instances, the current video block may not have residual data, such as in skip mode, and the residual generation unit 207 may not be able to perform subtraction operations.

[0220] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0221] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0222] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in buffer 213.

[0223] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0224] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0225] Figure 5 This is a block diagram illustrating an example of a video decoder 300. The video decoder 300 can be... Figure 3 The video decoder 114 in the system 100 shown.

[0226] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 5 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0227] exist Figure 5 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 4 The decoding process is the inverse of the encoding process described.

[0228] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-coded video data, motion compensation unit 302 can determine motion information, which includes motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 may determine this information, for example, by executing AMVP and merge modes.

[0229] The motion compensation unit 302 can generate motion compensation blocks and perform interpolation based on an interpolation filter. The syntax element may contain identifiers for the interpolation filter used at sub-pixel precision.

[0230] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of a video block, to calculate interpolated values ​​for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate a prediction block.

[0231] The motion compensation unit 302 may use some syntax information to determine the size of the blocks of frames and / or stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0232] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0233] The reconstruction unit 306 can add the residual block to the corresponding predicted block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0234] Figures 6-9 The above technical solutions are shown (e.g., Figures 1 to 5 Example method of the embodiment shown.

[0235] Figure 6 A flowchart of an example method 600 for video processing is shown. Method 600 includes, in operation 610, performing a conversion between a video and a bitstream of the video according to a rule, the bitstream comprising one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and the rule specifying that the maximum decoded picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on the maximum allowed picture size of the one or more codec video pictures corresponding to the bitstream or the one or more bitstream portions.

[0236] Figure 7A flowchart of an example method 700 for video processing is shown. Method 700 includes, in operation 710, performing a conversion between a video comprising video units and a bitstream of the video comprising one or more codec layer video sequences, the bitstream conforming to a rule that stipulates that the maximum buffer size of the decoded images of the codec layer video sequences is constrained to be less than or equal to the maximum image size selected from the images of the codec layer video sequences.

[0237] Figure 8 A flowchart of an example method 800 for video processing is shown. Method 800 includes, in operation 810, performing a conversion between a video and a bitstream of the video according to a rule, the bitstream comprising one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more encoded video pictures of the video, and the rule specifying at least one of a maximum allowed picture size, a maximum allowed picture width, and a maximum allowed picture height for the conversion.

[0238] Figure 9 A flowchart of an example method 900 for video processing is shown. Method 900 includes, in operation 910, performing a conversion between a video and a bitstream of the video, the bitstream comprising one or more output layer sets, at least one output layer set comprising multiple video layers, and the bitstream conforming to a rule that specifies a constraint on the overall size of the decoded image buffer of the at least one output layer set comprising the multiple video layers.

[0239] The following is a list of preferred embodiments.

[0240] A1. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to rules, wherein the bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and wherein the rules specify that the maximum decoded picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on the maximum allowed picture size of the one or more codec video pictures corresponding to the bitstream or the one or more bitstream portions.

[0241] A2. The method according to Scheme 1 further includes: avoiding determining the maximum image buffer size based on the image size in the brightness samples (denoted as PicSizeInSamplesY).

[0242] A3. According to the method described in Scheme 1 or 2, the maximum decoded image buffer size is represented as MaxDpbSize, the maximum luminance image size is represented as MaxLumaPs, the maximum allowed image size is represented as PicSizeMaxInSamplesY, the default maximum decoded image buffer size is represented as maxDpbPicBuf, and MaxDpbSize is determined as follows:

[0243] If (PicSizeMaxInSamplesY≤(MaxLumaPs>>2))

[0244] MaxDpbSize=Min(4*maxDpbPicBuf,16)

[0245] Otherwise, if (PicSizeMaxInSamplesY≤(MaxLumaPs>>1))

[0246] MaxDpbSize=Min(2*maxDpbPicBuf,16)

[0247] Other than (PicSizeMaxInSamplesY≤((3*MaxLumaPs)>>2))

[0248] MaxDpbSize=Min((4*maxDpbPicBuf) / 3,16)

[0249] other

[0250] MaxDpbSize = maxDpbPicBuf.

[0251] A4. The method according to any one of schemes 1-3, wherein each of the one or more independently decodeable bitstream portions is a codec layer video sequence (CLVS).

[0252] A5. The method according to any one of schemes 1-4, wherein the maximum decoded image buffer size is based on each CLVS exported.

[0253] A6. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to rules, wherein the bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more encoded / decoded video images of the video, and wherein the rules specify at least one of a maximum allowed image size, a maximum allowed image width, and a maximum allowed image height for the conversion.

[0254] A7. The method according to Scheme 6, wherein each of the one or more independently decodeable bitstream portions is a codec layer video sequence (CLVS).

[0255] A8. The method according to Scheme 7, wherein one or more of the maximum allowed image size, the maximum allowed image width, and the maximum allowed image height are specified based on each CLVS.

[0256] A9. The method according to Scheme 7, wherein the maximum allowed image size, the maximum allowed image width, and the maximum allowed image height are specified for each sequence parameter set (SPS) associated with each of the at least one CLVS.

[0257] A10. The method according to any one of schemes 6-9, wherein the rule specifies that the maximum permissible image size (denoted as PicSizeMaxInSamplesY) of each reference sequence parameter set (SPS) is less than or equal to the maximum luminance image size (denoted as MaxLumaPs).

[0258] A11. The method according to any one of schemes 6-10, wherein the rule specifies that the maximum permissible image width (represented as pic_width_max_in_luma_samples) of each reference sequence parameter set (SPS) is less than or equal to Sqrt(8×MaxLumaPs), where MaxLumaPs is the maximum luminance image size.

[0259] A12. The method according to any one of schemes 6-11, wherein the rule specifies that the maximum permissible image height (denoted as pic_height_max_in_luma_samples) of each reference sequence parameter set (SPS) is less than or equal to Sqrt(8×MaxLumaPs), where MaxLumaPs is the maximum luminance image size.

[0260] A13. A video processing method, comprising:

[0261] Perform a conversion between a video and a bitstream of the video comprising one or more codec layer video sequences, wherein the bitstream conforms to a rule, and wherein the rule specifies that the maximum buffer size of the decoded images of the codec layer video sequences is constrained to be less than or equal to the maximum image size selected from the images in the codec layer video sequences.

[0262] A14. The method according to Scheme 13, wherein the maximum buffer size of the decoded image is selected based on the maximum buffer size from the decoded image buffer (DPB) parameter set.

[0263] A15. The method according to Scheme 14, wherein the DPB parameter set corresponds to the DPB parameter set of the codec layer of the output layer as the output layer set (OLS).

[0264] 16. The method according to Scheme 14, wherein the DPB parameter set corresponds to the DPB parameter set of the codec layer of the output layer that is not an output layer set (OLS).

[0265] A17. A video processing method comprising: performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets, wherein at least one output layer set includes a plurality of video layers, and wherein the bitstream conforms to a rule specifying a constraint on the overall size of a decoded image buffer of the at least one output layer set including the plurality of video layers.

[0266] A18. The method according to Scheme 17, wherein the decoded image buffer includes a plurality of decoded images after decoding each access unit, wherein each of the plurality of decoded images has a width in a luminance sample, and wherein the constraint specifies that the sum of the widths of the plurality of decoded images is less than or equal to a predetermined value.

[0267] A19. The method according to Scheme 18, wherein the predetermined value is based on the index of the corresponding codec layer among the plurality of codec layers.

[0268] A20. The method according to Scheme 17, wherein each of the plurality of codec layers is associated with each of the plurality of decoded images, wherein each of the plurality of decoded images has a maximum width, and wherein the constraint specifies that the sum of the maximum widths of the plurality of decoded images is less than or equal to a predetermined value.

[0269] A21. The method according to Scheme 20, wherein the maximum width is the product of the maximum image width and the maximum image height among the luminance samples of the corresponding layers in the plurality of encoding and decoding layers.

[0270] A22. The method according to any one of schemes 1-21, wherein the conversion includes decoding the video from the bitstream.

[0271] A23. The method according to any one of claims 1-21, wherein the conversion includes encoding the video into the bitstream.

[0272] A24. The method according to any one of schemes 1-21, wherein performing the conversion includes:

[0273] The video is encoded into the bitstream; and the bitstream is stored in a non-transitory computer-readable storage medium.

[0274] A25. A video processing apparatus including a processor configured to perform one or more of the methods described in Schemes 1-24.

[0275] A26. A non-transitory computer-readable storage medium configured to store a bitstream of video generated by the method described in any one or more of schemes A1 to A24.

[0276] A27. A non-transitory computer-readable storage medium configured to store instructions for a processor to implement one or more of the methods described in embodiments A1 to A24.

[0277] A28. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to be the method described in any one or more of embodiments A1 to A24.

[0278] A29. A method for storing a bitstream of video, comprising: generating a bitstream from video according to rules; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more codec video pictures of the video, and wherein the rules specify that the maximum decoded picture buffer size required to decode the bitstream or the one or more bitstream portions is determined based on the maximum permissible picture size of the one or more codec video pictures corresponding to the bitstream or the one or more bitstream portions.

[0279] A30. A method for storing a bitstream of video, comprising: generating a bitstream from the video according to rules; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream includes one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more coded video images of the video, and wherein the rules specify at least one of a maximum permissible image size, a maximum permissible image width, and a maximum permissible image height for the conversion.

[0280] The following is another list of preferred embodiments.

[0281] B1. A video processing method comprising performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to a rule, and wherein the rule specifies one or more constraints on at least one of the following for the access units: a codec picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, or the sum of the number of bytes in a Network Abstraction Layer (NAL) unit, based on the sum of the picture sizes of each of a plurality of pictures in the access unit.

[0282] B2. According to the method of scheme B1, the sum of the number of bytes in the NAL unit is represented as NumBytesInNalUnit.

[0283] B3. The method according to scheme B1 or B2, wherein the rule is independent of the current size of the decoded image in the luminance samples (denoted as PicSizeInSamplesY).

[0284] B4. According to the method of scheme B1, wherein the rule stipulates that the nominal CPB removal time of the access unit (AU) satisfies the following constraint:

[0285] AuNominalRemovalTime[n]-AuCpbRemovalTime[n-1]

[0286] ≥Max(AuSizeInSamplesY[n-1]÷MaxLumaSr,fR),

[0287] Where AuNominalRemovalTime[n] is the nominal CPB removal time of the nth AU, AuCpbRemovalTime[n-1] is the CPB removal time of the (n-1)th AU, AuSizeInSamplesY[n-1] is the size of the (n-1)th AU in units of samples, MaxLumaSr is the maximum luminance sampling rate (number of samples per second), and fR is a variable equal to 1 ÷ 300, where n is an integer greater than 0.

[0288] B5. According to the method of scheme B1, where the rule stipulates that the difference in DPB output time of images from different access units (AUs) of DPB satisfies the constraint:

[0289] DpbOutputInterval[n]≥Max(AuSizeInSamplesY[n-1]÷MaxLumaSr,fR),

[0290] Where DpbOutputInterval[n] is the DPB output time of the nth AU, MaxLumaSr is the maximum luminance sampling rate (samples per second), AuSizeInSamplesY[n-1] is the size of the (n-1)th AU in the sample points, and fR is a variable equal to 1÷300, where n is a positive integer.

[0291] B6. According to the method of scheme B1, the rule stipulates that the sum of the number of bytes in the NAL unit satisfies the constraint condition:

[0292] NumBytesInNalUnit[0]≤FormatCapabilityFactor

[0293] ×(Max(AuSizeInSamplesY[0]÷MaxLumaSr,fR×MaxLumaSr)

[0294] +MaxLumaSr×(AuCpbRemovalTime[0]-AuNominalRemovalTime[0]))÷MinCr,

[0295] Where NumBytesInNalUnit[0] is the sum of the number of bytes in the NAL unit of the first access unit (AU), AuSizeInSamplesY[0] is the size of the first AU in the samples, MaxLumaSr is the maximum luminance sampling rate (samples per second), aucpbreaktime[0] is the CPB removal time of the first AU, AuNominalRemovalTime[0] is the nominal CPB removal time of the first AU, MinCr is the minimum compression level, and fR is a variable equal to 1÷300.

[0296] B7. A video processing method comprising performing a conversion between a video and a bitstream of the video, wherein the bitstream is organized into one or more access units according to rules, and wherein the rules specify a limit on the maximum number of stripes in the access units.

[0297] B8. According to the method of scheme B7, the rule further specifies the maximum number of stripes in an access unit (denoted as MaxSlicesPerAu) based on the access unit level.

[0298] B9. Following the method in scheme B8, the maximum number of stripes in an access cell is determined using the following table:

[0299] level 1 2 2.1 3 3.1 4 4.1 5 5.1 5.2 6 6.1 6.2 MaxSlicesPerAu 16 16 20 30 40 75 75 200 200 200 600 600 600

[0300] B10. According to the method of scheme B7, the rule further specifies the constraint on the codec picture buffer (CPB) removal time (denoted as AuCpbRemovalTime) for each access unit based on the limit on the maximum number of stripes in the access unit.

[0301] B11. According to the method of scheme B10, where the rule stipulates that the CPB removal time of the first access unit (denoted as AuCpbRemovalTime[0]) satisfies the constraint:

[0302] NumSlicesPerAu[0]≤Min(Max(1,MaxSlicesPerAu×MaxLumaSr / MaxLumaPs×(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxSlicesPerAu×AuSizeInSamplesY[0] / MaxLumaPs),MaxSlicesPerAu),

[0303] Wherein, MaxSlicesPerAu is the maximum number of slices in the access unit, NumSlicesPerAu[0] is the number of slices in the first access unit (AU), MaxLumaPs is the maximum luminance image size, MaxLumaSr is the maximum luminance sampling rate (in samples per second), AuCpbRemovalTime[0] is the CPB removal time of the first Au, and AuNominalRemovalTime[0] is the nominal CPB removal time of the first AU.

[0304] B12. According to the method of scheme B10, where the rule stipulates that the difference between consecutive CPB removal times satisfies the following constraint:

[0305] NumSlicesPerAu[n]≤Min((Max(1,MaxSlicesPerAu×MaxLumaSr / MaxLumaPs×(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxSlicesPerAu)

[0306] Where MaxSlicesPerAu is the maximum number of slices in the access unit, NumSlicesPerAu[n] is the number of slices in the nth access unit (AU), MaxLumaPs is the maximum luminance image size, MaxLumaSr is the maximum luminance sampling rate (in samples per second), AuCpbRemovalTime[n] is the CPB removal time of the nth AU, and AuCpbRemovalTime[n-1] is the CPB removal time of the (n-1)th AU, where n is a positive integer.

[0307] B13. The method of scheme B7, where the rule is independent of the maximum number of stripes per image.

[0308] B14. The method according to any one of schemes B1 to B13, wherein the conversion includes decoding video from the bitstream.

[0309] B15. The method according to any one of schemes B1 to B13, wherein the conversion includes encoding the video into a bitstream.

[0310] B16. The method according to any one of schemes B1 to B13, wherein performing the conversion includes encoding the video into a bitstream; and storing the bitstream in a non-transitory computer-readable storage medium.

[0311] B17. A video processing apparatus, including a processor configured to implement any one or more of the methods described in schemes B1 to B16.

[0312] B18. A non-transitory computer-readable storage medium configured to store a bitstream of video generated by the method described in any one or more of schemes B1 to B16.

[0313] B19. A non-transitory computer-readable storage medium configured to store instructions that cause a processor to implement one or more of the methods described in schemes B1 to B16.

[0314] B20. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement any one or more of the methods described in schemes B1 to B16.

[0315] B21. A method for storing a video bitstream, comprising: generating a bitstream from a video according to rules; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream is organized into one or more access units according to the rules, and wherein the rules specify one or more constraints on at least one of the following: a codec picture buffer (CPB) removal time, a nominal CPB removal time, a decoded picture buffer (DPB) output time, or the sum of the number of bytes in a network abstraction layer (NAL) unit, based on the sum of the picture sizes of each of a plurality of pictures in the access unit.

[0316] B22. A method for storing a bitstream of video, comprising: generating a bitstream from the video according to rules; and storing the bitstream in a non-transitory computer-readable storage medium, wherein the bitstream is organized into one or more access units according to rules, and wherein the rules specify a limit on the maximum number of stripes in the access units.

[0317] The following is another list of preferred embodiments.

[0318] P1. A video processing method comprising performing a conversion between video units of a video and a codec representation of the video, wherein a maximum picture buffer size used during the conversion is determined from the maximum picture size in a picture of a codec layer of the video unit, wherein the maximum picture buffer size is specific to the codec layer.

[0319] P2. According to the method of scheme P1, the determination of the maximum image buffer size is independent of the variable defining the image size associated with the encoding and decoding layers.

[0320] P3. A video processing method comprising performing a conversion between video units of a video and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies constraints related to the maximum buffer size of the decoded image that apply only to a single-layer codec representation.

[0321] P4. According to the method of scheme P3, where the format rule further specifies that, in the case of multi-layer video encoding and decoding representation, different values ​​of the maximum decoder buffer size apply to different layers.

[0322] P5. A video processing method, comprising performing a conversion between video units of a video and a video codec representation, wherein the codec representation conforms to a format rule that specifies, in the case where the codec representation comprises multiple layers, values ​​for a maximum image size, an image width, and an image height are individually defined within each codec layer of the video codec representation.

[0323] P6. Following the method in scheme P5, these values ​​are specified at the sequence parameter set level.

[0324] P7. The method according to any one of schemes P1 to P6, wherein performing the conversion includes encoding the video to generate the codec representation.

[0325] P8. The method according to any one of schemes P1 to P6, wherein performing the conversion includes parsing and decoding the codec representation to generate the video.

[0326] P9. A video decoding apparatus, including a processor configured to implement one or more of the methods described in schemes P1 to P8.

[0327] P10. A video encoding apparatus, including a processor configured to implement one or more of the methods described in schemes P1 to P8.

[0328] P11. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of schemes P1 to P8.

[0329] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation (or simply, the bitstream) of the current video block may, for example, correspond to bits that are commonly located or scattered at different locations within the bitstream, as defined in the syntax. For example, macroblocks may be encoded based on error residuals from the transform and encoding, and may also utilize bits in the header and other fields in the bitstream.

[0330] The subject matter and functional operations described in this patent document can be implemented as various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this specification and their equivalents, or as a combination of one or more of them. The subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transient computer-readable medium for execution by or control of the operation of a data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that realizes a machine-readable propagating signal, or a combination of one or more of these. The term "data processing device" encompasses all devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the device may also contain code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these. The propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver device.

[0331] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple harmonizing files (e.g., files storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one location or distributed across multiple locations and interconnected via a communication network.

[0332] The processes and logic flows described in this specification can be executed by one or more programmable processors to execute one or more computer programs, thereby performing functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0333] Processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0334] Although this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as a description of features that may be specific to particular embodiments of a particular invention. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations of sub-combinations.

[0335] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or to perform all shown operations to achieve the desired effect. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0336] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A video processing method, comprising: Perform the conversion between the video and the video bitstream according to the first rule. The bitstream comprises one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more encoded / decoded video images of the video. The first rule stipulates that the maximum decoded image buffer size required to decode the bitstream or one or more bitstream portions is determined based on the first maximum allowed image size of the one or more codec video images corresponding to the bitstream or one or more bitstream portions; The method further includes: Avoid determining the maximum decoded image buffer size based on the image size in the luminance samples (denoted as PicSizeInSamplesY).

2. The method as described in claim 1, wherein, The maximum decoded image buffer size is denoted as MaxDpbSize, the maximum luminance image size is denoted as MaxLumaPs, the first maximum allowed image size is denoted as PicSizeMaxInSamplesY, and the default maximum decoded image buffer size is denoted as maxDpbPicBuf, wherein MaxDpbSize is determined as follows: If (PicSizeMaxInSamplesY ≤ (MaxLumaPs >> 2)) MaxDpbSize = Min( 4 * maxDpbPicBuf, 16 ) Otherwise, if (PicSizeMaxInSamplesY ≤ (MaxLumaPs >> 1)) MaxDpbSize = Min( 2 * maxDpbPicBuf, 16 ) Otherwise, if (PicSizeMaxInSamplesY ≤ ((3 * MaxLumaPs) >> 2)) MaxDpbSize = Min( ( 4 * maxDpbPicBuf ) / 3, 16 ) otherwise MaxDpbSize = maxDpbPicBuf.

3. The method as described in claim 1 or 2, wherein, Each of the one or more independently decodeable bitstream portions is a codec layer video sequence (CLVS).

4. The method as described in claim 1 or 2, wherein, The maximum decoded image buffer size is based on the output of each CLVS.

5. The method of claim 1, wherein, The first rule also specifies at least one of the second maximum allowed image size, maximum allowed image width, and maximum allowed image height for the transformation.

6. The method of claim 5, wherein, Each of the one or more independently decodeable bitstream portions is a codec layer video sequence (CLVS).

7. The method of claim 6, wherein, One or more of the second maximum allowed image size, the maximum allowed image width, and the maximum allowed image height are specified based on each CLVS.

8. The method of claim 6, wherein, Specify the second maximum allowed image size, the maximum allowed image width, and the maximum allowed image height for each sequence parameter set (SPS) associated with each of at least one CLVS.

9. The method according to any one of claims 5-8, wherein, The first rule stipulates that the second maximum permissible image size (denoted as PicSizeMaxInSamplesY) for each reference sequence parameter set (SPS) is less than or equal to the maximum luminance image size (denoted as MaxLumaPs).

10. The method according to any one of claims 5-8, wherein, The first rule stipulates that the maximum permissible image width (denoted as pic_width_max_in_luma_samples) of each reference sequence parameter set (SPS) is less than or equal to Sqrt(8 × MaxLumaPs), where MaxLumaPs is the maximum luminance image size.

11. The method according to any one of claims 5-8, wherein, The first rule stipulates that the maximum permissible image height (denoted as pic_height_max_in_luma_samples) for each reference sequence parameter set (SPS) is less than or equal to Sqrt(8 × MaxLumaPs), where MaxLumaPs is the maximum luminance image size.

12. The method according to any one of claims 1-2 and 5-8, wherein, The conversion includes decoding the video from the bitstream.

13. The method according to any one of claims 1-2 and 5-8, wherein, The conversion includes encoding the video into the bitstream.

14. The method according to any one of claims 1-2 and 5-8, wherein, Performing the conversion includes: Encode the video into the bitstream; and The bit stream is stored in a non-transitory computer-readable recording medium.

15. The method of claim 1, further comprising: Perform conversion between the video and a bitstream of the video, comprising one or more codec layers. Wherein, the bitstream comprising one or more codec layer video sequences conforms to the second rule, and The second rule stipulates that the maximum buffer size of the decoded images of the one or more codec layer video sequences is constrained to be less than or equal to the maximum image size selected from the images of the one or more codec layer video sequences.

16. The method of claim 15, wherein, The maximum buffer size for decoding the image is selected based on the maximum buffer size from the DPB parameter set of the decoded image buffer.

17. The method of claim 16, wherein, The DPB parameter set corresponds to the DPB parameter set of the codec layer of the output layer, which is the output layer set (OLS).

18. The method of claim 16, wherein, The DPB parameter set corresponds to the DPB parameter set of the output layer's encoder / decoder layer, which is not the output layer set (OLS).

19. The method according to any one of claims 15-18, wherein, The conversion between the video and a bitstream of the video comprising one or more codec layers includes decoding the video from the bitstream comprising one or more codec layers.

20. The method according to any one of claims 15-18, wherein, The conversion between the video and a bitstream of the video comprising one or more codec layers includes encoding the video into the bitstream comprising one or more codec layers.

21. The method according to any one of claims 15-18, wherein, Performing the conversion between the video and a bitstream of the video comprising one or more codec layers includes: The video is encoded into a bitstream comprising one or more codec layers; as well as The bitstream comprising one or more codec layers of video sequences is stored in a non-transitory computer-readable recording medium.

22. The method of claim 1, further comprising: Perform conversion between the video and the bitstream of the video, including one or more output layer sets. At least one output layer set includes multiple video layers, and The bitstream comprising one or more output layer sets conforms to a third rule, which specifies a constraint on the overall size of the decoded image buffer of the at least one output layer set comprising the plurality of video layers.

23. The method of claim 22, wherein, The decoded image buffer includes a plurality of decoded images after decoding each access unit, wherein each of the plurality of decoded images has a width in a luminance sample, and wherein the constraint specifies that the sum of the widths of the plurality of decoded images is less than or equal to a predetermined value.

24. The method of claim 23, wherein, The predetermined value is based on the index of the corresponding codec layer among multiple codec layers.

25. The method of claim 22, wherein, Each of the plurality of codec layers is associated with each of the plurality of decoded images, wherein each of the plurality of decoded images has a maximum width, and wherein the constraint specifies that the sum of the maximum widths of the plurality of decoded images is less than or equal to a predetermined value.

26. The method of claim 25, wherein, The maximum width is the product of the maximum image width and the maximum image height among the luminance samples of the corresponding layer in the plurality of encoding and decoding layers.

27. The method according to any one of claims 22-26, wherein, The conversion between the video and the bitstream of the video including one or more output layer sets includes decoding the video from the bitstream including one or more output layer sets.

28. The method according to any one of claims 22-26, wherein, The conversion between the video and the bitstream of the video including one or more output layer sets includes encoding the video into the bitstream including one or more output layer sets.

29. The method according to any one of claims 22-26, wherein, Performing the conversion between the video and the bitstream of the video, which includes one or more output layer sets, includes: The video is encoded into a bitstream comprising one or more output layer sets; as well as The bitstream comprising one or more output layer sets is stored in a non-transitory computer-readable recording medium.

30. A video processing apparatus including a processor configured to perform the method as claimed in any one of claims 1-29.

31. A non-transitory computer-readable storage medium configured to store instructions for causing a processor to perform the method of any one of claims 1-29.

32. A method for storing a video bitstream, comprising: Generate a bitstream from the video according to the first rule; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The bitstream comprises one or more independently decodeable bitstream portions, each bitstream portion corresponding to one or more encoded / decoded video images of the video. The first rule stipulates that the maximum decoded image buffer size required to decode the bitstream or one or more bitstream portions is determined based on the first maximum allowed image size of the one or more codec video images corresponding to the bitstream or one or more bitstream portions; The method further includes: Avoid determining the maximum decoded image buffer size based on the image size in the luminance samples (denoted as PicSizeInSamplesY).

33. The method of claim 32, wherein, The first rule also specifies at least one of the second maximum allowed image size, maximum allowed image width, and maximum allowed image height for generating the bitstream.

Citation Information

Patent Citations

  • An apparatus, a method and a computer program for omnidirectional video

    EP3422724A1

  • Level definitions for multi-layer video codecs

    WO2015142694A1