Constraints on the number of slices in a coded video picture
Patent Information
- Application Number
- CN202180041428.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-08
- Filing Date
- 2021-06-08
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2041-06-08
Smart Images

Figure CN115943627B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] Pursuant to the applicable patent law and / or rules of the Paris Convention, this application promptly claims priority and interest in U.S. Provisional Patent Application No. 63 / 036,321, filed June 8, 2020. For all purposes required by law, the entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] The patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques for implementing constraints used by video encoders and decoders to perform video encoding or decoding.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of a video comprising one or more images containing one or more stripes, wherein the bitstream is organized into a plurality of access units (AUs) AU0 to AUn based on a format rule, where n is a positive integer, and wherein the format rule specifies a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) during decoding and the number of stripes in each of the plurality of AUs.
[0007] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images containing one or more slices and a bitstream of the video, wherein the bitstream is organized into a plurality of Access Units (AUs) AU0 to AUn based on a format rule, where n is a positive integer, and wherein the format rule specifies a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) and the number of slices in each of the plurality of AUs.
[0008] In yet another example, another video processing method is disclosed. The method includes: performing a conversion between a video and a video bitstream comprising one or more images containing one or more stripes, wherein the conversion conforms to a rule, wherein the bitstream is organized into one or more access units, wherein the rule specifies a constraint on the number of decoded images stored in a decoded image buffer (DPB), wherein each decoded image (i) is marked for reference, (ii) has a flag indicating that the decoded image is output, and (iii) has an output time later than the current image's decoding time.
[0009] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0010] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0011] In yet another example, a computer-readable medium on which code is stored is disclosed. This code implements one of the methods described herein in the form of processor-executable code.
[0012] These and other features will be described in this document. Attached Figure Description
[0013] Figure 1 This is a block diagram illustrating an example video processing system in which various techniques disclosed herein can be implemented.
[0014] Figure 2 This is a block diagram of an example hardware platform used for video processing.
[0015] Figure 3 This is a block diagram illustrating an example video codec system that can implement some embodiments of the present disclosure.
[0016] Figure 4 This is a block diagram illustrating an example of an encoder that can implement some embodiments of the present disclosure.
[0017] Figure 5 This is a block diagram illustrating examples of decoders that can implement some embodiments of the present disclosure.
[0018] Figures 6-8 A flowchart of an example method for video processing is shown. Detailed Implementation
[0019] The use of chapter headings in this document is for ease of understanding and not to limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.
[0020] 1. Introduction
[0021] This document relates to video codec technology. Specifically, it defines levels and bitstream consistency for video codecs that support both single-level and multi-level video codecs. This document can be applied to any standard or non-standard video codec that supports both single-level and multi-level video codecs, such as the Versatile Video Codec (VVC) currently under development.
[0022] 2. Abbreviation
[0023] APS Adaptive Parameter Set
[0024] AU Access Unit
[0025] AUD Access Unit Delimiter
[0026] AVC Advanced Video Codec
[0027] CLVS codec layer video sequence
[0028] CPB encoded image buffer
[0029] CRA (Clean Random Access)
[0030] CTU (Codec Tree Unit)
[0031] CVS codec video sequence
[0032] DPB Decoding Image Buffer
[0033] DPS Decoding Parameter Set
[0034] EOB End of Bitstream
[0035] End of EOS sequence
[0036] GCI General Constraint Information
[0037] GDR Progressive Decoding Refresh
[0038] HEVC High-Efficiency Video Encoding and Decoding
[0039] HRD Hypothetical Reference Decoder
[0040] IDR Instant Decoding and Refresh
[0041] JEM Joint Exploration Model
[0042] MCTS Motion Constraint Pieces
[0043] NAL Network Abstraction Layer
[0044] OLS Output Layer Set
[0045] PH image header
[0046] PPS Image Parameter Set
[0047] PTL profile, tier, and level
[0048] PU Image Unit
[0049] RRP reference image resampling
[0050] RBSP raw byte sequence payload
[0051] SEI Supplemental Enhancement Information
[0052] SH strip header
[0053] SPS Sequence Parameter Set
[0054] SVC Scalable Video Codec
[0055] VCL (Video Codec Layer)
[0056] VPS Video Parameter Set
[0057] VTM VVC Test Model
[0058] VUI Video Availability Information
[0059] VVC Multi-Functional Video Encoding and Decoding
[0060] 3. Preliminary Discussion
[0061] Video coding standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture, utilizing time prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Universal Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and the VTM test model are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.
[0062] 3.1. Changes in image resolution within a sequence
[0063] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC allows changing the image resolution within a sequence at locations where IRAP images are not encoded; IRAP images are always intra-frame encoded and decoded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.
[0064] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, the same as in motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.
[0065] Other aspects of the VVC design that support this feature differ from HEVC include: i) Picture resolution and the corresponding consistency window are signaled in the PPS instead of the SPS, where the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (the slot in the DPB used to store a decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.
[0066] 3.2. Scalable Video Codec (SVC) in General and VVC
[0067] Scalable video codec (SVC, sometimes referred to as scalability in video codec) refers to video codec using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below the intermediate layer (such as a base layer or any intermediate enhancement layer) and simultaneously serve as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0068] In SVC, parameters used by the encoder or decoder are grouped into parameter sets based on the codec level (e.g., video level, sequence level, picture level, stripe level, etc.), and these sets may be utilized at that codec level. For example, parameters that can be utilized by one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters that can be utilized by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.
[0069] Because VVC supports Reference Picture Resampling (RPR), it's possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without requiring any additional signal processing-level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires a higher level of syntax changes (compared to no scalability support). Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions to AVC and HEVC), VVC's scalability is designed to be as friendly as possible to single-layer decoder designs. The decoding capability of a multi-layer bitstream is specified as if there were only one layer in the bitstream. For example, decoding capabilities such as the DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams don't require many changes to decode multi-layer bitstreams. Compared to the multi-layer extension designs of AVC and HEVC, HLS is significantly simplified at the expense of some flexibility. For example, IRAPU requires a picture of every layer present in CVS.
[0070] 3.3. Parameter Set
[0071] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All versions of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0072] The Sequence-Level Prefix (SPS) is designed to carry sequence-level header information, while the Picture-Level Prefix (PPS) is designed to carry infrequently changing picture-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or picture, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving error resilience.
[0073] The VPS was introduced to carry sequence-level header information shared by all layers in a multi-layer bitstream.
[0074] The purpose of APS is to carry such image-level or stripe-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.
[0075] 3.4. Grades, Levels, and Classes
[0076] Video codec standards typically specify quality levels and grades. Some video coding standards (e.g., HEVC and the developing VVC) also specify layers.
[0077] Grades, levels, and tiers specify limitations on the bitstream and thus limit the capabilities required to decode it. Grades, levels, and tiers can also be used to indicate points of interoperability between various decoder implementations.
[0078] Each grade specifies a subset of algorithmic features and constraints that should be supported by all decoders conforming to that grade. Note that the encoder does not need to use all codec tools or features supported in the grade, while decoders conforming to the grade need to support all codec tools or features.
[0079] Each level of a hierarchy specifies a set of restrictions on the values that can be taken for bitstream syntax elements. The same set of hierarchy and level definitions is typically used for all grades, but various implementations may support different hierarchies, and within a hierarchy, different levels may be supported for each supported grade. For any given grade, the level of a hierarchy generally corresponds to the specific decoder's processing payload and memory capabilities.
[0080] The capabilities of a video decoder that conforms to a video codec specification are defined by its ability to decode video streams that conform to the constraints of the tiers, levels, and grades specified in the video codec specification. When expressing the capabilities of a specific tier decoder, the tiers and grades supported by that tier should also be expressed.
[0081] 3.5. Existing VVC hierarchy and level definitions
[0082] In the latest VVC draft text of JVET-S0152-v5, the hierarchy and level are defined as follows.
[0083] A.4.1 General Hierarchy and Level Restrictions
[0084] For the purpose of comparing hierarchical capabilities, a hierarchy with general_tier_flag equal to 0 (i.e., the primary hierarchy) is considered a lower hierarchy than a hierarchy with general_tier_flag equal to 1 (i.e., the higher hierarchy).
[0085] For the purpose of comparing level capabilities, when the value of general_level_idc or sublayer_level_idc[i] of a specific level of a specified layer is less than the value of general_level_idc or sublayer_level_idc[i] of other levels, that specific level is considered to be a lower level than other levels of the same layer.
[0086] To express the constraints in this appendix, the following are specified:
[0087] – Let AU n be the nth AU in the decoding order, where the first AU is AU 0 (i.e., the 0th AU).
[0088] – For an OLS with an OLS index TargetOlsIdx, the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, and PicSizeMaxInSamplesY, along with the applicable dpb_parameters() syntax structure, are deduced as follows:
[0089] If NumLayersInOls[TargetOlsIdx] equals 1, then PicWidthMaxInSamplesY is set to equal sps_pic_width_max_in_luma_samples, PicHeightMaxInSamplesY is set to equal sps_pic_height_max_in_luma_samples, and PicSizeMaxInSamplesY is set to equal PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, where sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are found in the SPS referenced by the layer in OLS, and the applicable dpb_parameters() is also found in that SPS.
[0090] Otherwise (if NumLayersInOls[TargetOlsIdx] is greater than 1), then PicWidthMaxInSamplesY is set to equal vps_ols_dpb_pic_width[MultiLayerOlsIdx[TargetOlsIdx]], PicHeightMaxInSamplesY is set to equal vps_ols_dpb_pic_height[MultiLayerOlsIdx[TargetOlsIdx]], PicSizeMaxInSamplesY is set to equal PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, and the applicable dpb_parameters() syntax structure is identified by vps_ols_dpb_params_idx[MultiLayerOlsIdx[TargetOlsIdx]] found in the VPS.
[0091] When the specified level is not level 15.5, bitstreams conforming to the specified tier and grade at that level should comply with the following constraints for each bitstream conformance test as specified in Appendix C:
[0092] a) PicSizeMaxInSamplesY should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.
[0093] b) The value of PicWidthMaxInSamplesY should be less than or equal to Sqrt(MaxLumaPs*8).
[0094] c) The value of PicHeightMaxInSamplesY should be less than or equal to Sqrt(MaxLumaPs*8).
[0095] d) For each referenced PPS, the value of NumTileColumns should be less than or equal to MaxTileCols, and the value of NumTileRows should be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.
[0096] e) For the VCL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] should be less than or equal to CpbVclFactor*MaxCPB, wherein CpbSize[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, CpbVclFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in bits of CpbVclFactor.
[0097] f) For the NAL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), CpbSize[Htid][i] should be less than or equal to CpbNalFactor*MaxCPB, wherein CpbSize[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, CpbNalFactor is specified in Table A.3, and MaxCPB is specified in Table A.1 in bits of CpbNalFactor.
[0098] Table A.1 specifies the restrictions for each level in each hierarchy, except for level 15.5.
[0099] The tier and level that the bitstream conforms to are indicated by the syntax elements general_tier_flag and general_level_idc, and the level that the sublayer indicates conforms to is indicated by the syntax element sublayer_level_idc[i], as follows:
[0100] – If the specified level is not level 15.5, then according to the tier constraints specified in Table A.1, a general_tier_flag of 0 indicates compliance with the primary tier, a general_tier_flag of 1 indicates compliance with the higher tier, and for levels below level 4, general_tier_flag should be equal to 0 (corresponding to the items marked with "-" in Table A.1). Otherwise (if the specified level is level 15.5), the bitstream consistency requirement is that general_tier_flag should be equal to 1, and the value 0 for general_tier_flag is reserved for future use by ITU-T|ISO / IEC, and the decoder should ignore the value of general_tier_flag.
[0101] –general_level_idc and sublayer_level_idc[i] should be set to the general_level_idc value equal to the level number specified in Table A.1.
[0102] Table A.1 - General Hierarchy and Level Restrictions
[0103]
[0104] A.4.2 Specific level restrictions
[0105] To express the constraints in this appendix, the following are specified:
[0106] – Suppose that the variable fR is set to equal 1 ÷ 300.
[0107] The variable HbrFactor is defined as follows:
[0108] – If the bitstream is indicated to conform to main 10 or main 4:4:4 10, then HbrFactor is set to equal to 1.
[0109] The variable BrVclFactor, representing the VCL bit rate scaling factor, is set to be equal to CpbVclFactor * HbrFactor.
[0110] The variable BrNalFactor, representing the NAL bit rate scaling factor, is set to be equal to CpbNalFactor * HbrFactor.
[0111] The variable MinCr is set to be equal to MinCrBase * MinCrScaleFactor ÷ HbrFactor.
[0112] When the specified level is not level 15.5, the value of max_dec_pic_buffering_minus1[Htid]+1 should be less than or equal to MaxDpbSize, which is deduced as follows:
[0113] if(PicSizeMaxInSamplesY<=(MaxLumaPs>>2))
[0114] MaxDpbSize=Min(4*maxDpbPicBuf,16)
[0115] else if(PicSizeMaxInSamplesY<=(MaxLumaPs>>1))
[0116] MaxDpbSize=Min(2*maxDpbPicBuf,16)
[0117] else if(PicSizeMaxInSamplesY<=((3*MaxLumaPs)>>2))
[0118] MaxDpbSize=Min((4*maxDpbPicBuf) / 3,16)
[0119] else
[0120] MaxDpbSize=maxDpbPicBuf (A.1)
[0121] In Table A.1, MaxLumaPs is specified, maxDpbPicBuf is equal to 8, and max_dec_pic_buffering_minus1[Htid] can be found or derived from the applicable dpb_parameters() syntax structure.
[0122] Let numDecPics be the number of images in AU n. The variable AuSizeMaxInSamplesY[n] is set to be equal to PicSizeMaxInSamplesY * numDecPics.
[0123] Bitstreams conforming to the specified hierarchy and level (primary 10 or primary 4:4:4 10) should comply with the following constraints specified in Appendix C for each bitstream conformance test:
[0124] a) As specified in Clause C.2.3, the nominal removal time of AU n (n greater than 0) from the CPB shall satisfy the constraint that AuNominalRemovalTime[n] - AuCpbRemovalTime[n-1] is greater than or equal to Max(AuSizeMaxInSamplesY[n-1] ÷ MaxLumaSr, fR), where MaxLumaSr is the value specified in Table A.2 applicable to AU n-1.
[0125] b) As specified in Clause C.3.3, the difference between consecutive output times of images from different AUs in the DPB shall satisfy the constraint that DpbOutputInterval[n] is greater than or equal to Max(AuSizeMaxInSamplesY[n]÷MaxLumaSr,fR), where MaxLumaSr is the value of AU n specified in Table A.2, assuming that AU n has the images being output, and AU n is not the last AU in the bitstream with the images being output.
[0126] c) For the value of PicSizeMaxInSamplesY for image 0, the removal time of AU 0 should satisfy the constraint that the number of stripes in each image of AU 0 is less than or equal to Min(Max(1,MaxSlicesPerPicture*MaxLumaSr / MaxLumaPs*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxSlicesPerPicture*PicSizeMaxInSamplesY / MaxLumaPs),MaxSlicesPerPicture), where MaxSlicesPerPicture, MaxLumaPs and MaxLumaSr are the values specified in Table A.1 and Table A.2 for AU 0, respectively.
[0127] d) The difference between consecutive CPB removal times of AU n and AU n-1 (n > 0) should satisfy that the number of stripes in each picture of AU n is less than or equal to Min((Max(1,MaxSlicesPerPicture*MaxLumaSr / MaxLumaPs*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxSlicesPerPicture), where MaxSlicesPerPicture, MaxLumaPs, and MaxLumaSr are the values specified in Tables A.1 and A.2 applicable to AU n.
[0128] e) For the VCL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] should be less than or equal to BrVclFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.
[0129] f) For the NAL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] should be less than or equal to BrNalFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, and MaxBR is specified in Table A.2 in BrNalFactor bits / second.
[0130] g) The sum of the NumBytesInNalUnit variables for AU 0 should be less than or equal to FormatCapabilityFactor*(Max(AuSizeMaxInSamplesY[0],fR*MaxLumaSr)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0]))÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 for AU 0, respectively.
[0131] h) The sum of the NumBytesInNalUnit variables for AU n (n greater than 0) should be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 for AUn, respectively.
[0132] i) The removal time of AU 0 should satisfy the constraint that the number of tiles in each picture in AU 0 is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxTileCols*MaxTileRows*AuSizeMaxInSamplesY[0] / MaxLumaPs),MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.1 applicable to AU 0.
[0133] j) The difference between consecutive CPB removal times of AU n and AU n-1 (n > 0) should satisfy the constraint that the number of tiles in each picture in AU n is less than or equal to Min(Max(1,MaxTileCols*MaxTileRows*120*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxTileCols*MaxTileRows), where MaxTileCols and MaxTileRows are the values specified in Table A.1 applicable to AU n.
[0134] Table A.2 - Video Qualification Levels and Restrictions
[0135]
[0136] Table A.3 - Specifications for CpbVclFactor, CpbNalFactor, FormatCapabilityFactor, and MinCrScaleFactor
[0137]
[0138] 3.6. Existing VVC bitstream consistency definition
[0139] In the latest VVC draft text of JVET-S0152-v5, bitstream consistency is defined as follows.
[0140] A.4 Bitstream Consistency
[0141] Encoded and decoded data bitstreams conforming to this specification should meet all the requirements specified in this clause.
[0142] Bitstreams should be constructed in accordance with the syntax, semantics and constraints specified in this specification, except for those in this appendix.
[0143] The first encoded / decoded image in the bitstream should be an IRAP image (i.e., an IDR image or a CRA image) or a GDR image.
[0144] The HRD tests whether the bitstream conforms to the specifications specified in Clause C.1.
[0145] Let currPicLayerId be equal to the nuh_layer_id of the current image.
[0146] For each current image, let the variables maxPicOrderCnt and minPicOrderCnt be set to the maximum and minimum values of PicOrderCntVal for the following images, respectively, where nuh_layer_id equals currPicLayerId:
[0147] - Current image.
[0148] - The previous image in the decoding order has both TemporalId and ph_non_ref_pic_flag equal to 0, and is not a RASL or RADL image.
[0149] - The STRP referenced by all entries in RefPicList[0] and all entries in RefPicList[1] of the current image.
[0150] -PictureOutputFlag equals 1, AuCpbRemovalTime[n] is less than AuCpbRemovalTime[currPic] and DpbOutputTime[n] is greater than or equal to AuCpbRemovalTime[currPic] for all images n, where currPic is the current image.
[0151] For each of the bitstream consistency tests, all of the following conditions should be met:
[0152] 1. For each AU n (n greater than 0) associated with a BP SEI message, let the variable deltaTime90k[n] be specified as follows:
[0153] deltaTime90k[n]=90000*(AuNominalRemovalTime[n]-AuFinalArrivalTime[n-1]) (C.17)
[0154] The value of InitCpbRemovalDelay[Htid][ScIdx] is constrained as follows:
[0155] - If cbr_flag[ScIdx] equals 0, then the following condition should be true:
[0156] InitCpbRemovalDelay[Htid][ScIdx]<=Ceil(deltaTime90k[n])
[0157] (C.18)
[0158] Otherwise (cbr_flag[ScIdx] equals 1), the following condition should be true:
[0159] Floor(deltaTime90k[n])<=InitCpbRemovalDelay[Htid][ScIdx]<=Ceil(deltaTime90k[n]) (C.19)
[0160] Note 1 – The exact number of bits in the CPB at the removal time of each AU or DU may depend on which BP SEI message is chosen to initialize the HRD. The encoder must take this into account to ensure that all specified constraints are met regardless of which BP SEI message is chosen to initialize the HRD, since the HRD can be initialized at any of the BP SEI messages.
[0161] 2. CPB overflow is defined as the situation where the total number of bits in the CPB exceeds the size of the CPB. The CPB should never overflow.
[0162] 3. When low_delay_hrd_flag[Htid] equals 0, CPB should never underflow. CPB underflow is specified as follows:
[0163] - If DecodingUnitHrdFlag equals 0, then CPB underflow is specified as a condition: for at least one n value, the nominal CPB removal time AuNominalRemovalTime[n] of AU n is less than the final CPB arrival time AuFinalArrivalTime[n] of AU n.
[0164] - Otherwise (DecodingUnitHrdFlag equals 1), CPB underflow is specified as conditional: for at least one m value, the nominal CPB removal time DuNominalRemovalTime[m] of DU m is less than the final CPB arrival time DuFinalArrivalTime[m] of DU m.
[0165] 4. When DecodingUnitHrdFlag equals 1, low_delay_hrd_flag[Htid] equals 1, and the nominal removal time of DU m of AU n is less than the final CPB arrival time of DU m (i.e., DuNominalRemovalTime[m] < DuFinalArrivalTime[m]), the nominal removal time of AU n should be less than the final CPB arrival time of AU n (i.e., AuNominalRemovalTime[n] < AuFinalArrivalTime[n]).
[0166] 5. The nominal removal time of an AU from the CPB (starting from the second AU in decoding order) should satisfy the constraints on AuNominalRemovalTime[n] and AuCpbRemovalTime[n] expressed in Clauses A.4.1 and A.4.2.
[0167] 6. For each current picture, after invoking the process of removing pictures from the DPB as specified in Clause C.3.2, the number of decoded pictures in the DPB (including all pictures n marked as "for reference", or pictures n where PictureOutputFlag equals 1 and CpbRemovalTime[n] is less than CpbRemovalTime[currPic] where currPic is the current picture) should be less than or equal to max_dec_pic_buffering_minus1[Htid].
[0168] 7. When prediction is required, all reference pictures should be present in the DPB. Each picture with PictureOutputFlag equal to 1 should be present in the DPB at its DPB output time, unless it has been removed from the DPB by one of the processes specified in Clause C.3 before its output time.
[0169] 8. For each current picture that is not a CLVSS picture, the value of maxPicOrderCnt - minPicOrderCnt should be less than MaxPicOrderCntLsb / 2.
[0170] 9. The value of DpbOutputInterval[n] given by Equation C.16 (i.e., the difference between the output time of a picture and the output time of the first picture after it in output order with PictureOutputFlag equal to 1) should satisfy the constraints expressed in Clause A.4.1 for the profile, tier, and level specified in the bitstream using the decoding processes specified in Clauses 2 to 9.
[0171] 10. For each current image, when bp_du_cpb_params_in_pic_timing_sei_flag equals 1, let tmpCpbRemovalDelaySum be derived as follows:
[0172] tmpCpbRemovalDelaySum=0
[0173] for(i=0; i <pt_num_decoding_units_minus1;i++)
[0174] tmpCpbRemovalDelaySum+=
[0175] pt_du_cpb_removal_delay_increment_minus1[i][Htid]+1(C.20)
[0176] The value of ClockSubTick*tmpCpbRemovalDelaySum should be equal to the difference between the nominal CPB removal time of the current AU and the nominal CPB removal time of the first DU in the current AU according to the decoding order.
[0177] 11. For any two images m and n in the same CVS, when DpbOutputTime[m] is greater than DpbOutputTime[n], the PicOrderCntVal of image m should be greater than the PicOrderCntVal of image n.
[0178] Note 2 – All images from earlier CVSs in decoding order are output before any images from later CVSs in decoding order. Within any given CV, the output images are output in ascending order of PicOrderCntVal.
[0179] 12. The DPB output time derived for all images in any given AU should be the same.
[0180] 4. The technical problem solved by the disclosed technical solution
[0181] The existing level-defined VVC design has the following problems:
[0182] 1) The definitions of the two level limits for the relationship between CPB removal time and the number of stripes for AU 0 and AU n (n>0) (i.e., items c and d in Clause A.4.2 (tier-specific level limits)) are based on MaxSlicesPerPicture and the maximum picture size. However, MaxSlicesPerPicture is defined as a picture-level limit, while CPB removal time is an AU-level parameter.
[0183] 2) The definitions of the two level limits for the relationship between CPB removal time and the number of slices for AU 0 and AU n (n>0) (i.e., items i and j in Clause A.4.2 (tier-specific level limits)) are based on MaxTileCols*MaxTileRows and the maximum AU size. However, similar to the above, MaxTileCols and MaxTileRows are defined as picture-level limits, while CPB removal time is an AU-level parameter.
[0184] 3) The sixth constraint in Clause C.4 (Bitstream Consistency) is as follows: For each current picture, after invoking the process specified in Clause C.3.2 to remove a picture from the DPB, the number of decoded pictures in the DPB (including all pictures n marked "for reference", or pictures n where PictureOutputFlag equals 1 and CpbRemovalTime[n] is less than CpbRemovalTime[currPic] (where currPic is the current picture)) should be less than or equal to max_dec_pic_buffering_minus1[Htid]. The part "CpbRemovalTime[n] is less than CpbRemovalTime[currPic]" describes the condition that the decoding time of a decoded picture in the DPB is less than the decoding time of the current picture. However, all decoded pictures in the DPB are always decoded earlier than the current picture, so CpbRemovalTime[n] in the context is always less than CpbRemovalTime[currPic].
[0185] 5. Examples of solutions and implementation methods
[0186] To address the aforementioned and other issues, methods outlined below are disclosed. These items should be considered as examples for explaining general concepts, and not interpreted in a narrow sense. Furthermore, these items can be used individually or in combination in any way.
[0187] 1) To address the first issue, the definition of the two levels of constraints on the relationship between the CPB removal time and the number of stripes for AU 0 and AU n (n>0) (i.e., items c and d in clause A.4.2 of the latest VVC draft) is changed from being based on MaxSlicesPerPicture and the maximum picture size to being based on MaxSlicesPerPicture*(the number of pictures in the AU) and the maximum AU size.
[0188] 2) To address the second issue, the definition of the two levels of constraints on the relationship between CPB removal time and the number of slices for AU 0 and AU n (n>0) (i.e., items i and j in clause A.4.2 of the latest VVC draft) is changed from being based on MaxTileCols*MaxTileRows and the maximum AU size to being based on MaxTileCols*MaxTileRows*(the number of pictures in the AU) and the maximum AU size.
[0189] 3) To address the third issue, the sixth constraint in Clause C.4 of the latest VVC draft is changed to impose a constraint on the number of decoded pictures stored in the DPB that are marked "for reference", have PictureOutputFlag equal to 1, and whose output time is later than the decoding time of the current picture.
[0190] a. In one example, in constraint 6 of clause C.4 of the latest VVC draft, "or PictureOutputFlag equals 1 and CpbRemovalTime[n] is less than CpbRemovalTime[currPic]" is changed to "or PictureOutputFlag equals 1 and DpbOutputTime[n] is greater than CpbRemovalTime[currPic]".
[0191] 6. Example
[0192] The following are some example embodiments of the invention summarized in this section, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-S0152-v5. Most relevant additions or modifications are indicated in bold, underlined, and italic, such as "Using A..." The deleted parts are indicated in italics and enclosed in bold double brackets, for example, "based on..." "B". There may be other editable changes, which are not highlighted.
[0193] 6.1. Example 1
[0194] This example applies to projects 1 to 3 and their sub-projects.
[0195] A.4.2 Specific level restrictions
[0196] To express the constraints in this appendix, the following are specified:
[0197] - Suppose that the variable fR is set to equal 1 ÷ 300.
[0198] The variable HbrFactor is defined as follows:
[0199] - If the bitstream is indicated to conform to main 10 or main 4:4:4 10, then HbrFactor is set to equal to 1.
[0200] The variable BrVclFactor, representing the VCL bit rate scaling factor, is set to be equal to CpbVclFactor * HbrFactor.
[0201] The variable BrNalFactor, representing the NAL bit rate scaling factor, is set to be equal to CpbNalFactor * HbrFactor.
[0202] The variable MinCr is set to be equal to MinCrBase * MinCrScaleFactor ÷ HbrFactor.
[0203] When the specified level is not level 15.5, the value of max_dec_pic_buffering_minus1[Htid]+1 should be less than or equal to MaxDpbSize, which is deduced as follows:
[0204] if(PicSizeMaxInSamplesY<=(MaxLumaPs>>2))
[0205] MaxDpbSize=Min(4*maxDpbPicBuf,16)
[0206] else if(PicSizeMaxInSamplesY<=(MaxLumaPs>>1))
[0207] MaxDpbSize=Min(2*maxDpbPicBuf,16)
[0208] else if(PicSizeMaxInSamplesY<=((3*MaxLumaPs)>>2))
[0209] MaxDpbSize=Min((4*maxDpbPicBuf) / 3,16)
[0210] else
[0211] MaxDpbSize=maxDpbPicBuf (A.1)
[0212] In Table A.1, MaxLumaPs is specified, maxDpbPicBuf is equal to 8, and max_dec_pic_buffering_minus1[Htid] can be found or derived from the applicable dpb_parameters() syntax structure.
[0213] set up This represents the number of images in AU n. The variable AuSizeMaxInSamplesY[n] is set to equal to
[0214] Bitstreams conforming to the specified hierarchy and level (primary 10 or primary 4:4:4 10) should comply with the following constraints specified in Appendix C for each bitstream conformance test:
[0215] k) The nominal removal time of AU n (n greater than 0) from the CPB as specified in Clause C.2.3 shall satisfy the constraint that AuNominalRemovalTime[n] - AuCpbRemovalTime[n-1] is greater than or equal to Max(AuSizeMaxInSamplesY[n-1] ÷ MaxLumaSr, fR), where MaxLumaSr is the value specified in Table A.2 applicable to AU n-1.
[0216] l) As specified in Clause C.3.3, the difference between consecutive output times of images from different AUs from the DPB shall satisfy the constraint that DpbOutputInterval[n] is greater than or equal to Max(AuSizeMaxInSamplesY[n]÷MaxLumaSr,fR), where MaxLumaSr is the value specified in Table A.2 for AU n, assuming that AU n has images as outputs and AU n is not the last AU in the bitstream to have images as outputs.
[0217] m) The removal time of AU 0 should meet the requirements of AU 0. The number of stripes is less than or equal to The constraints are defined as follows: MaxSlicesPerPicture, MaxLumaPs, and MaxLumaSr are the values specified in Tables A.1 and A.2 that apply to AU 0.
[0218] The difference between the consecutive CPB removal times of AUn and AUn-1 (n > 0) should satisfy the condition in AUn. The number of stripes is less than or equal to Among them, MaxSlicesPerPicture, MaxLumaPs, and MaxLumaSr are the values applicable to AUn specified in Tables A.1 and A.2.
[0219] o) For the VCL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] should be less than or equal to BrVclFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, and MaxBR is specified in Table A.2 in BrVclFactor bits / second.
[0220] p) For the NAL HRD parameter, for at least one value of i in the range of 0 to hrd_cpb_cnt_minus1 (inclusive), BitRate[Htid][i] should be less than or equal to BrNalFactor*MaxBR, where BitRate[Htid][i] is specified in Clause 7.4.6.3 based on the parameter selected as specified in Clause C.1, and MaxBR is specified in Table A.2 in BrNalFactor bits / second.
[0221] q) The sum of the NumBytesInNalUnit variables of AU 0 should be less than or equal to FormatCapabilityFactor*(Max(AuSizeMaxInSamplesY[0],fR*MaxLumaSr)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0]))÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 for AU 0, respectively.
[0222] The sum of the NumBytesInNalUnit variables for r)AU n (n greater than 0) should be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])÷MinCr, where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 for AUn, respectively.
[0223] s) The removal time of AU 0 should satisfy the conditions in AU 0. The number of pieces is less than or equal to The constraints, wherein MaxTileCols and MaxTileRows are the values specified in Table A.1 applicable to AU 0.
[0224] The difference between the consecutive CPB removal times of AUn and AUn-1 (n > 0) should satisfy the condition in AUn. The number of pieces is less than or equal to The constraints, wherein MaxTileCols and MaxTileRows are the values specified in Table A.1 applicable to AU n.
[0225] ...
[0226] C.4 Bitstream Consistency
[0227] ...
[0228] 6. For each current image, after invoking the procedure for removing images from the DPB as specified in Clause C.3.2, the decoded images in the DPB (including all images n marked "for reference"), or The number of images (where currPic is the current image) should be less than or equal to max_dec_pic_buffering_minus1[Htid].
[0229] ...
[0230] Figure 1This is a block diagram illustrating an example video processing system 1000, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0231] System 1000 may include codec component 1004, which may implement various codec or encoding methods described in this document. Codec component 1004 may reduce the average bit rate of the video from input 1002 to the output of codec component 1004 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1006, the output of codec component 1004 may be stored or transmitted via connected communication. Component 1008 may use the stored or communicated bitstream (or codec) representation of the video received at input 1002 to generate pixel values or displayable video sent to display interface 1010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of the encoding result, will be performed by the decoder.
[0232] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0233] Figure 2 This is a block diagram of a video processing apparatus 2000. Apparatus 2000 can be used to implement one or more methods described herein. Apparatus 2000 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The processors (multiple) 2002 may be configured to implement the methods described herein (e.g., in...). Figure 6One or more methods are depicted (as shown in Figure 9). Memory (multiple memories) 2004 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 2006 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, hardware 2006 may be partially or wholly located within one or more processors 2002 (e.g., graphics processors).
[0234] Figure 3 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein. Figure 3 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0235] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems used to generate video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0236] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0237] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, which is configured to interface with an external display device.
[0238] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Multi-Function Video Coding (VVM) standard, and other current and / or further standards.
[0239] Figure 4 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 3 The video encoder 114 in the system 100 shown.
[0240] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 4 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0241] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0242] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0243] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 4 The examples are shown separately.
[0244] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0245] The mode selection unit 203 can, for example, select one of the coding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0246] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.
[0247] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0248] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating a reference image in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0249] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates a reference image in list 0 or list 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0250] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding processing.
[0251] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 can signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0252] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0253] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0254] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0255] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0256] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0257] In other examples, there may be no residual data for the current video block, for example, in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[0258] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0259] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0260] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is then stored in buffer 213.
[0261] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.
[0262] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.
[0263] Figure 5 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 3 The video decoder 114 in the system 100 shown.
[0264] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 5 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0265] exist Figure 5 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform functions typically associated with the video encoder 200. Figure 4 The decoding process is the inverse of the encoding process described.
[0266] Entropy decoding unit 301 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-coded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine this information, for example, by executing AMVP and Merge modes.
[0267] Motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The syntax elements can include identifiers of the interpolation filters to be used with sub-pixel precision.
[0268] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate a prediction block.
[0269] The motion compensation unit 302 can use some syntax information to determine the size of the blocks of (multiple) frames and / or (multiple) strips used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0270] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0271] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0272] Figures 6-8 It shows that it can be done in, for example Figures 1-5 The embodiments shown are example methods for implementing the above-described technical solution.
[0273] Figure 6 A flowchart of an example method 600 for video processing is shown. Method 600 includes, at operation 610, performing a conversion between a video and a bitstream of video comprising one or more pictures containing one or more stripes, the bitstream being organized into a plurality of access units (AUs) AU0 to AUn based on a format rule specifying a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) during decoding and the number of stripes in each of the plurality of AUs, and n being a positive integer.
[0274] Figure 7 A flowchart of an example method 700 for video processing is shown. Method 700 includes, in operation 710, performing a conversion between a video and a bitstream comprising one or more pictures containing one or more slices, based on format rules. The bitstream is organized into a plurality of access units (AUs) AU0 to AUn based on format rules specifying a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) and the number of slices in each of the plurality of AUs, where n is a positive integer.
[0275] Figure 8 A flowchart of an example method 800 for video processing is shown. Method 800 includes, at operation 810, performing a conversion between a video and a video bitstream comprising one or more images containing one or more stripes, the bitstream being organized into one or more access units, the conversion conforming to a rule specifying a constraint on the number of decoded images stored in a decoded image buffer, wherein each decoded image (i) is marked for reference, (ii) has a flag indicating that the decoded image is output, and (iii) has an output time later than the current image's decoding time.
[0276] The following provides an example of preferred solutions for some embodiments.
[0277] A1. A method for processing video data, comprising: performing a conversion between a video and a bitstream of video comprising one or more pictures containing one or more stripes, wherein the bitstream is organized into a plurality of access units (AUs) AU0 to AUn based on a format rule, wherein n is a positive integer, wherein the format rule specifies a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) during decoding and the number of stripes in each of the plurality of AUs.
[0278] A2. The method of solution A1, wherein the relationship is based on (i) the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit.
[0279] A3. The method of solution A1, wherein the format rule specifies that the removal time of the first access unit AU 0 among multiple access units satisfies the constraint.
[0280] A4. The method of solution A3, wherein the constraint specifies that the number of slices in AU 0 is less than or equal to Min(Max(1,MaxSlicesPerAu×MaxLumaSr / MaxLumaPs×(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxSlicesPerAu×AuSizeMaxInSamplesY[0] / MaxLumaPs),MaxSlicesPerAu), where MaxSlicesPerAu is the maximum number of slices per access unit, MaxLumaSr is the maximum luminance sampling rate, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the maximum size of the luminance sample of the decoded image of the reference sequence parameter set (SPS).
[0281] A5. Solution A4's method, where the values of MaxLumaPs and MaxLumaSr are selected from the values corresponding to AU 0.
[0282] A6. The method of solution A1, wherein the format rule specifies that the difference between the removal times of two consecutive access units AUn-1 and AUn satisfies the constraint.
[0283] A7. The method of solution A6, wherein the constraint specifies that the number of stripes in AU n is less than or equal to Min((Max(1,MaxSlicesPerAu×MaxLumaSr / MaxLumaPs×(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxSlicesPerAu), where MaxSlicesPerAu is the maximum number of stripes per access unit, MaxLumaSr is the maximum luminance sampling rate, MaxLumaPs is the maximum luminance image size, and AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit.
[0284] A8. The method of solution A7, wherein the values of MaxSlicesPerAu, MaxLumaPs, and MaxLumaSr are selected from the values corresponding to AU n.
[0285] A9. A method for processing video data, comprising: performing a conversion between a video and a bitstream of video comprising one or more pictures containing one or more slices, wherein the bitstream is organized into a plurality of access units (AUs) AU0 to AUn based on a format rule, wherein n is a positive integer, and wherein the format rule specifies a relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) and the number of slices in each of the plurality of AUs.
[0286] A10. The method of solution A9, wherein the relationship is based on (i) the product of the maximum number of tile columns (MaxTileCols) per image, the maximum number of tile rows (MaxTileRows) per image, and the number of images in the access cell, and (ii) the maximum size of the access cell.
[0287] A11. The method of solution A9, wherein the format rule specifies that the removal time of the first access unit AU 0 among a plurality of access units satisfies the constraint.
[0288] A12. The method of solution A11, wherein the constraint specifies that the number of slices in AU 0 is less than or equal to Min(Max(1,MaxTilesPerAu×120×(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])+MaxTilesPerAu×AuSizeMaxInSamplesY[0] / MaxLumaPs),MaxTilesPerAu), where MaxTilesPerAu is the maximum number of slices per access unit, MaxLumaPs is the maximum luminance picture size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the maximum size of the decoded picture with luminance samples of the reference sequence parameter set (SPS).
[0289] A13. The method of solution A12, wherein the value of MaxTilesPerAu is selected from the value corresponding to AU 0.
[0290] A14. The method of solution A9, wherein the format rule specifies that the difference between the removal times of two consecutive access units AUn-1 and AUn satisfies the constraint.
[0291] A15. The method of solution A14, wherein the constraint specifies that the number of slices in AU n is less than or equal to Min(Max(1,MaxTilesPerAu×120×(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])),MaxTilesPerAu), where MaxTilesPerAu is the maximum number of slices per access unit, and AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit.
[0292] A16. The method of solution A15, wherein the value of MaxTilesPerAu is selected from the values corresponding to AU n.
[0293] A17. A method of any one of solutions A1 to A16, wherein the conversion includes decoding the video from the bitstream.
[0294] A18. The method of any one of solutions A1 to A16, wherein the conversion includes encoding the video into the bitstream.
[0295] A19. A method for storing a bitstream representing a video to a computer-readable recording medium, comprising: generating a bitstream from the video according to the method described in any one or more of solutions A1 to A16, and storing the bitstream in the computer-readable recording medium.
[0296] A20. A video processing apparatus, including a processor configured to implement the method described in any one or more of solutions A1 to A19.
[0297] A21. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method described in any one or more of solutions A1 to A19.
[0298] A22. A computer-readable medium for storing a bitstream generated according to any one or more of solutions A1 to 19.
[0299] A23. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions A1 to A19.
[0300] The following provides another example of preferred solutions for some embodiments.
[0301] B1. A method for processing video data, comprising: performing a conversion between a video and a video bitstream comprising one or more pictures containing one or more stripes, wherein the conversion conforms to a rule, wherein the bitstream is organized into one or more access units, wherein the rule specifies a constraint on the number of decoded pictures stored in a decoded picture buffer (DPB), wherein each decoded picture (i) is marked for reference, (ii) has a flag indicating that the decoded picture is output, and (iii) has an output time later than the decoding time of the current picture.
[0302] B2. The method of solution B1, wherein the number of decoded images is less than or equal to the maximum required size of the DPB in units of image storage buffer minus 1.
[0303] B3. The method of solution B1, wherein the flag is PictureOutputFlag, the output time of the m-th image is DpbOutputTime[m], and the decoding time of the m-th image is AuCpbRemovalTime[m].
[0304] B4. The method of any one of solutions B1 to B3, wherein the conversion includes decoding the video from the bitstream.
[0305] B5. The method of any one of solutions B1 to B3, wherein the conversion includes encoding the video into the bitstream.
[0306] B6. A method for storing a bitstream representing a video to a computer-readable recording medium, comprising: generating a bitstream from the video according to one or more of the methods described in solutions B1 to B3, and storing the bitstream in the computer-readable recording medium.
[0307] B7. A video processing apparatus, including a processor configured to implement the method described in any one or more of solutions B1 to B6.
[0308] B8. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method described in any one or more of solutions B1 to B6.
[0309] B9. A computer-readable medium for storing a bit stream generated according to any one or more of solutions B1 to B6.
[0310] B10. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions B1 to B6.
[0311] The following provides another example of preferred solutions for some embodiments.
[0312] P1. A video processing method comprising: performing a conversion between a video comprising one or more images containing one or more stripes and a bitstream representation of the video, wherein the bitstream representation is organized into one or more access units according to a format rule, wherein the format rule specifies a relationship between one or more syntax elements in the bitstream representation and the removal time of one or more access units from a codec image buffer.
[0313] P2. Solution P1's method, where the rule specifies two levels of restrictions on the relationship between the product of the maximum number of stripes per image and the number of images in the access unit, and the removal time of the maximum allowed size of the access unit.
[0314] P3. Solution P1's method, where the rule specifies two levels of restrictions on the relationship between the removal time and a first parameter whose value is MaxTileCols*MaxTileRows* (the number of pictures in the access cell) and a second parameter whose value is the maximum allowed size of one or more access cells.
[0315] P4. A video processing method comprising: performing a conversion between a video comprising one or more pictures containing one or more stripes and a bitstream representation of the video, wherein the conversion conforms to a rule, wherein the bitstream representation is organized into one or more access units, wherein the rule specifies a constraint on the number of decoded pictures stored in a decoded picture buffer, the decoded pictures being marked for reference and PictureOutputFlag being equal to 1, and having an output time later than the current picture's decoding time.
[0316] P5. A method of any one of solutions P1 to P4, wherein performing the conversion includes encoding the video to generate a codec representation.
[0317] P6. A method from any of solutions P1 to P4, wherein performing the conversion includes parsing and decoding the encoded / decoded representation to generate a video.
[0318] P7. A video decoding apparatus, including a processor configured to implement the methods described in any one or more of solutions P1 to P6.
[0319] P8. A video encoding apparatus, including a processor configured to implement the methods described in any one or more of solutions P1 to P6.
[0320] P9. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the methods described in any one or more of solutions P1 to P6.
[0321] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation (or simply bitstream) of the current video block can, for example, correspond to bits that are commonly located or scattered at different locations within the bitstream. For example, a macroblock can be encoded based on the error residual values from the transformation and encoding, and also using bits from the header and other fields in the bitstream.
[0322] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a combination of machine-readable storage devices, machine-readable storage substrates, memory devices, substances that implement machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0323] Computer programs (also referred to as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suited to a computing environment. A computer program does not necessarily correspond to a document in a documenting system. A program may be stored as a portion of a document containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single document dedicated to the program in question, or in multiple collaborative documents (e.g., a document storing portions of one or more modules, subroutines, or code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one place or distributed across multiple locations and interconnected via a communication network.
[0324] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0325] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented or incorporated therein by dedicated logic circuitry.
[0326] While this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular technology. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0327] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0328] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A method for processing video data, comprising: Perform the conversion between the video and the video bitstream. The bitstream is organized into multiple access units (AUs) from AU0 to AUn, where n is a positive integer, and each AU includes one or more images. The conversion is performed according to the first format rule, and Wherein, the first format rule specifies a first constraint related to the removal time of AU 0, wherein the first constraint is based on (i) a value equal to the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit; and The first format rule also specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
2. The method according to claim 1, wherein, The first constraint specifies that the number of stripes in AU 0 is less than or equal to: Min( Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × (AuCpbRemovalTime[ 0 ] − AuNominalRemovalTime[ 0 ] ) + MaxSlicesPerAu ×AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxSlicesPerAu ), Where MaxSlicesPerAu is the maximum number of slices per access unit, and MaxLumaSr is the maximum luminance sampling rate.
3. The method according to claim 2, wherein, The values of MaxLumaPs and MaxLumaSr are selected from the values corresponding to AU 0.
4. The method according to claim 1, wherein, The first format rule also specifies a second constraint related to the difference between the removal times of consecutive codec picture buffers (CPBs) of AU n and AU n-1, wherein the second constraint is based on a value equal to the product of the maximum number of stripes per picture and the number of pictures in the access unit, wherein the second constraint specifies that the number of stripes in AU n is less than or equal to Min( (Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × ( AuCpbRemovalTime[ n ] − AuCpbRemovalTime[ n − 1 ] ) ),MaxSlicesPerAu ), where MaxSlicesPerAu is the maximum number of stripes per access unit and MaxLumaSr is the maximum luminance sampling rate.
5. The method according to claim 4, wherein, The values of MaxSlicesPerAu, MaxLumaPs, and MaxLumaSr are selected from the values corresponding to AU n.
6. The method according to claim 1, wherein, The first format rule also specifies a fourth constraint related to the difference between consecutive CPB removal times of AU n and AU n-1, wherein the fourth constraint is based on a value equal to the product of the maximum number of tile columns (MaxTileCols) per image, the maximum number of tile rows (MaxTileRows) per image, and the number of images in the access unit.
7. The method according to claim 1, wherein, The value of MaxTilesPerAu is selected from the value corresponding to AU 0.
8. The method according to claim 6, wherein, The fourth constraint specifies that the number of slices in AU n is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ n ] −AuCpbRemovalTime[ n − 1 ] ) ), MaxTilesPerAu ).
9. The method according to claim 8, wherein, The value of MaxTilesPerAu is selected from the values corresponding to AU n.
10. The method according to claim 1, wherein, The conversion includes decoding the video from the bitstream.
11. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
12. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform the conversion between the video and the video bitstream. in, The bitstream is organized into multiple access units (AUs) from AU0 to AUn, where n is a positive integer, and each AU includes one or more images. The conversion is performed according to the first format rule, and Wherein, the first format rule specifies a first constraint related to the removal time of AU 0, wherein the first constraint is based on (i) a value equal to the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit; and The first format rule also specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
13. The apparatus according to claim 12, wherein, The first constraint specifies that the number of stripes in AU 0 is less than or equal to: Min( Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × (AuCpbRemovalTime[ 0 ] − AuNominalRemovalTime[ 0 ] ) + MaxSlicesPerAu ×AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxSlicesPerAu ), Where MaxSlicesPerAu is the maximum number of slices per access unit, and MaxLumaSr is the maximum luminance sampling rate.
14. The apparatus according to claim 12, wherein, The first format rule also specifies a second constraint related to the difference between the removal times of consecutive codec picture buffers (CPBs) of AU n and AU n-1, wherein the second constraint is based on a value equal to the product of the maximum number of stripes per picture and the number of pictures in the access unit, wherein the second constraint specifies that the number of stripes in AU n is less than or equal to: Min( (Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × (AuCpbRemovalTime[ n ] − AuCpbRemovalTime[ n − 1 ] ) ), MaxSlicesPerAu ), Where MaxSlicesPerAu is the maximum number of slices per access unit, MaxLumaSr is the maximum luminance sampling rate, MaxLumaPs is the maximum luminance image size, and AuCpbRemovalTime[m] is the CPB removal time for the m-th access unit.
15. The apparatus according to claim 12, in, The first format rule also specifies a fourth constraint related to the difference between consecutive CPB removal times of AU n and AU n-1, wherein the fourth constraint is based on a value equal to the product of the maximum number of tile columns (MaxTileCols) per image, the maximum number of tile rows (MaxTileRows) per image, and the number of images in the access unit.
16. The apparatus according to claim 15, wherein, The fourth constraint specifies that the number of slices in AU n is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ n ] −AuCpbRemovalTime[ n − 1 ] ) ), MaxTilesPerAu ).
17. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform the conversion between the video and the video bitstream. in, The bitstream is organized into multiple access units (AUs) from AU0 to AUn, where n is a positive integer, and each AU includes one or more images. The conversion is performed according to the first format rule, and Wherein, the first format rule specifies a first constraint related to the removal time of AU 0, wherein the first constraint is based on (i) a value equal to the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit; and The first format rule also specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
18. A computer-readable storage medium storing a bitstream and computer-readable instructions thereon, the computer-readable instructions, when executed by a processor, causing the processor to generate a bitstream of video. in, The bitstream is organized into multiple access units (AUs) from AU0 to AUn, where n is a positive integer, and each AU includes one or more images. The bitstream is generated according to a first format rule, and Wherein, the first format rule specifies a first constraint related to the removal time of AU 0, wherein the first constraint is based on (i) a value equal to the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit; and The first format rule also specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
19. A method for storing a video bitstream, comprising: Generate the bitstream of the video; as well as The bitstream is stored in a computer-readable recording medium. The bitstream is organized into multiple access units (AUs) from AU0 to AUn, where n is a positive integer, and each AU includes one or more images. The bitstream is generated according to a first format rule, and Wherein, the first format rule specifies a first constraint related to the removal time of AU 0, wherein the first constraint is based on (i) a value equal to the product of the maximum number of stripes per image and the number of images in the access unit, and (ii) the maximum size of the access unit; and The first format rule also specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
20. A method for processing video data, comprising: Perform conversion between a video containing one or more images of one or more stripes and the bitstream of said video. The bitstream is organized into multiple access units (AUs) AU0 to AUn based on format rules. Where n is a positive integer, The format rule specifies the relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) during decoding and the number of stripes in each of the plurality of AUs. The formatting rule specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
21. The method according to claim 20, wherein, The format rule specifies that the removal time of the first access unit AU 0 among the plurality of access units satisfies the first constraint.
22. The method according to claim 21, wherein, The first constraint specifies that the number of stripes in AU 0 is less than or equal to Min( Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × ( AuCpbRemovalTime[ 0 ] − AuNominalRemovalTime[ 0 ] ) + MaxSlicesPerAu × AuSizeMaxInSamplesY[0 ] / MaxLumaPs ), MaxSlicesPerAu ), where MaxSlicesPerAu is the maximum number of stripes per access unit, and MaxLumaSr is the maximum luminance sampling rate.
23. The method according to claim 22, wherein, The values of MaxLumaPs and MaxLumaSr are selected from the values corresponding to AU 0.
24. The method of claim 20, wherein, The format rules also specify that the difference between the removal times of two consecutive access units AUn−1 and AUn satisfies a second constraint.
25. The method according to claim 24, wherein, The second constraint specifies that the number of stripes in AU n is less than or equal to Min( (Max( 1, MaxSlicesPerAu × MaxLumaSr / MaxLumaPs × ( AuCpbRemovalTime[ n ] − AuCpbRemovalTime[ n − 1 ] ) ), MaxSlicesPerAu ), where MaxSlicesPerAu is the maximum number of stripes per access unit, and MaxLumaSr is the maximum luminance sampling rate.
26. The method of claim 25, wherein, The values of MaxSlicesPerAu, MaxLumaPs, and MaxLumaSr are selected from the values corresponding to AU n.
27. A method for processing video data, comprising: Perform conversion between a video containing one or more images (one or more slices) and the bitstream of said video. The bitstream is organized into multiple access units (AUs) AU0 to AUn based on format rules. Where n is a positive integer, The format rule specifies the relationship between the removal time of each of the plurality of AUs from the codec picture buffer (CPB) and the number of medium pieces in each of the plurality of AUs. The formatting rule specifies a third constraint related to the removal time of AU 0, wherein the third constraint specifies that the number of slices in AU 0 is less than or equal to: Min( Max( 1, MaxTilesPerAu × 120 × ( AuCpbRemovalTime[ 0 ] −AuNominalRemovalTime[ 0 ] ) + MaxTilesPerAu × AuSizeMaxInSamplesY[ 0 ] / MaxLumaPs ), MaxTilesPerAu ), Where MaxTilesPerAu is the maximum number of slices in each access unit, MaxLumaPs is the maximum luminance image size, AuCpbRemovalTime[m] is the CPB removal time of the m-th access unit, AuNominalRemovalTime[m] is the nominal CPB removal time of the m-th access unit, and AuSizeMaxInSamplesY[m] is the product of the maximum size of the decoded image in luminance samples and the number of images in the access unit.
28. The method according to claim 27, wherein, The value of MaxTilesPerAu is selected from the value corresponding to AU 0.
29. The method according to claim 27, wherein, The format rules also specify that the difference between the removal times of two consecutive access units AUn−1 and AUn satisfies the fourth constraint.
30. The method according to claim 29, wherein, The fourth constraint specifies that the number of slices in AU n is less than or equal to Min(Max(1, MaxTilesPerAu × 120 × (AuCpbRemovalTime[n] −AuCpbRemovalTime[n − 1])), MaxTilesPerAu).
31. The method according to claim 30, wherein, The value of MaxTilesPerAu is selected from the values corresponding to AU n.
32. The method according to any one of claims 20-31, wherein, The conversion includes decoding the video from the bitstream.
33. The method according to any one of claims 20-31, wherein, The conversion includes encoding the video into the bitstream.
34. A method for storing a bitstream representing video to a computer-readable recording medium, comprising: The method according to any one of claims 20-31 generates the bitstream from the video; as well as The bitstream is stored in the computer-readable recording medium.
35. A video processing apparatus comprising a processor configured to implement the method of any one of claims 20-33.
36. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method of any one of claims 20-33.
37. A computer-readable storage medium having stored thereon a bit stream and computer instructions, the computer instructions, when executed by a processor, causing the processor to perform the method according to any one of claims 20-31 to generate the bit stream.