Parameters are signaled at the sub-picture level in the video bitstream
By signaling the parameters at the sub-picture level in the video bitstream, the encoding and decoding behavior of the sub-picture is controlled, and the problem of large bandwidth occupancy in the video encoding and decoding process is solved, and more efficient encoding and decoding is achieved, which is suitable for existing and future video encoding and decoding standards.
Patent Information
- Application Number
- CN202080089302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-25
- Filing Date
- 2020-12-25
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2040-12-25
AI Technical Summary
Existing video encoding and decoding technologies occupy a large amount of bandwidth in the Internet and digital communication networks. With the increase in connected user equipment, bandwidth demand continues to grow, and it is difficult for the existing technology to effectively optimize the video encoding and decoding process to reduce bandwidth demand.
Using a sub-picture-based encoding and decoding method, the encoding and decoding behavior of the sub-picture is controlled by signaling parameters at the sub-picture level in the video bit stream, including indicating the existence and boundary processing of the sub-picture in the sequence parameter set, and the encoding and decoding process is optimized to reduce loop filtering operations and improve encoding efficiency.
Through sub-picture-level signaling notification parameters optimization, video decoding quality is improved, bandwidth requirements are reduced, encoding efficiency is improved, and it is applicable to existing video encoding and codec standards such as HEVC and future video encoding and codec standards.
Smart Images

Figure CN115362677B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 128124, filed on December 25, 2019, in accordance with applicable Patent Laws and / or the Paris Convention. The entire disclosure of the aforementioned application is incorporated herein by reference and made a part of the disclosure of this application for all legal purposes. Technical Field
[0003] This application document relates to video and image encoding and decoding technology. Background Art
[0004] Despite advances in video compression, digital video still accounts for the largest usage of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] Methods, systems, and apparatus for signaling parameters at a sub-picture level in a video bitstream are described. The disclosed techniques can be used by video or image decoder or encoder embodiments in which sub-picture based encoding or decoding is performed.
[0006] In one example aspect, a method of video processing is disclosed. The method includes performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the bitstream includes a parameter set that controls codec behavior of a sub-picture of the one or more sub-pictures associated with an identification (ID) of the sub-picture.
[0007] In another example aspect, a method of video processing is disclosed. The method includes performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein a current parameter set is configured to control a codec behavior of at least one of the one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies signaling in the bitstream a default parameter set corresponding to the current parameter set before signaling in the bitstream a difference between the current parameter set and a default parameter set.
[0008] In yet another example aspect, a method of video processing is disclosed. The method includes performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the bitstream comprises a parameter set, the parameter set comprising a first control parameter and a second control parameter for controlling codec properties of the sub-picture, and wherein the format rule specifies whether or how the first control parameter is overridden by the second control parameter for decoding.
[0009] In yet another example aspect, a method of video processing is disclosed. The method includes performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein one or more first flags corresponding to each of the one or more sub-pictures are included in a sequence parameter set (SPS), wherein each first flag of the one or more first flags indicates whether constraint information is signaled for the sub-picture corresponding to each first flag, and wherein the constraint information indicates a codec tool that is not applied to the corresponding sub-picture on a codec layer video sequence (CLVS).
[0010] In yet another example aspect, the above-described method may be implemented by a video encoder device comprising a processor.
[0011] In yet another example aspect, the above-described method may be implemented by a video decoder device comprising a processor.
[0012] In yet another example aspect, the methods may be implemented in the form of processor-executable instructions and stored on a computer-readable program medium.
[0013] These and other aspects are described further throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Examples of temporal motion vector prediction (TMVP) and region constraints in sub-block TMVP are shown.
[0015] Figure 2 An example of a hierarchical motion estimation scheme is shown.
[0016] Figure 3 An example of a picture with 18 by 12 luma CTUs is shown, which is partitioned into 12 slices and 3 raster scan strips.
[0017] Figure 4 An example of a picture with 18 by 12 luma CTUs is shown, which is partitioned into 24 slices and 9 rectangular strips.
[0018] Figure 5An example of a picture divided into 4 slices, 11 tiles, and 4 rectangular strips is shown.
[0019] Figure 6 is a block diagram illustrating an example video processing system in which the various techniques disclosed herein may be implemented.
[0020] Figure 7 is a block diagram of an example hardware platform for video processing.
[0021] Figure 8 is a block diagram illustrating an example video codec system in which some embodiments of the present disclosure can be implemented.
[0022] Figure 9 is a block diagram illustrating an example of an encoder capable of implementing some embodiments of the present disclosure.
[0023] Figure 10 is a block diagram illustrating an example of a decoder capable of implementing some embodiments of the present disclosure.
[0024] Figure 11-14 A flow chart illustrating an example method of video processing is shown. DETAILED DESCRIPTION
[0025] This document provides various techniques that decoders of image or video bitstreams can use to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" as used here includes both sequences of pictures (traditionally called video) and individual images. In addition, video encoders can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0026] The section headings used in this document are for ease of understanding and do not limit the embodiments and techniques to the corresponding sections. Thus, embodiments from one section can be combined with embodiments from other sections.
[0027] 1. Summary
[0028] This application relates to video codec technology. Specifically, this application relates to palette codecs, which use a primary color representation in video codecs. This technology can be applied to existing video codec standards, such as HEVC, as well as to a pending standard (Multi-Function Video Codec). It may also be applicable to future video codec standards or codecs.
[0029] 2. Preliminary Discussion
[0030] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards [1,2]. Since H.262, video codec standards have been based on a hybrid video codec architecture that uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of reducing bitrate by 50% compared to HEVC.
[0031] The latest version of the VVC draft, Versatile Video Codec (Draft 4), can be found at:
[0032] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=5755
[0033] The latest reference software for VVC, called VTM, can be found at:
[0034] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0
[0035] 2.1 TMVP in VVC and Region Constraints in Sub-Block TMVP
[0036] Figure 1 Shows the area constraints in TMVP and sub-block TMVP
[0037] like Figure 1 As shown, in TMVP and sub-block TMVP, the constrained time-domain MV can only be taken from the collocated CTU plus a column of 4×4 blocks.
[0038] 2.2 Sub-images proposed in JVET-O0141
[0039] This paper proposes a sub-picture based VVC codec design. This proposal is a follow-up to JVET-N0826, but is now based on the flexible slice and brick based sharding approach adopted at the 14th JVET conference.
[0040] The proposal is outlined as follows:
[0041] 1) Pictures can be divided into sub-pictures.
[0042] 2) An indication in the SPS indicating the presence of a sub-picture, along with other sequence-level information of the sub-picture.
[0043] 3) Whether a sub-picture is considered as a picture in the decoding process (excluding loop filtering operations) can be controlled by the bitstream.
[0044] 4) Whether to disable loop filtering across sub-picture boundaries can be controlled by the bitstream of each sub-picture. The DBF, SAO, and ALF processes are updated to control loop filtering operations across sub-picture boundaries.
[0045] 5) For simplicity, as a starting point, the sub-picture width, height, horizontal offset and vertical offset in units of luma samples are signaled in the SPS. The sub-picture boundaries are constrained to be slice boundaries.
[0046] 6) The coding_tree_unit() syntax is slightly updated to specify that sub-pictures are treated as pictures in the decoding process (excluding loop filtering operations), and the decoding process is updated to the following:
[0047] (Advanced) Derivation of temporal luminance motion vector prediction
[0048] Luminance sample bilinear interpolation process
[0049] Luminance sample 8-tap interpolation filtering process
[0050] Chroma sample interpolation process
[0051] 7) The sub-picture ID is explicitly specified in the SPS and included in the slice group header to enable extraction of sub-picture sequences without changing the VCL NAL units.
[0052] 8) The output sub-picture set (OSPS) is proposed to specify the canonical extraction and consistency points of sub-pictures and their sets.
[0053] 2.3 Sub-pictures in the Versatile Video Codec (Draft 6)
[0054] ■Sequence Parameter Set RBSP Syntax
[0055]
[0056] subpics_present_flag equal to 1 indicates that sub-picture parameters are present in the SPS RBSP syntax. subpics_present_flag equal to 0 indicates that sub-picture parameters are not present in the SPS RBSP syntax.
[0057] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag equal to 1 in the RBSP of the SPS.
[0058] max_subpics_minus1 plus 1 specifies the maximum number of subpictures that may be present in the CVS. max_subpics_minus1 should be in the range 0 to 254. The value 255 is reserved for future use by ITU-T | ISO / IEC.
[0059] subpic_grid_col_width_minus1 plus 1 specifies the width of each element of the sub-picture identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits.
[0060] The variable NumSubPicGridCols is derived as follows:
[0061] NumSubPicGridCols=
[0062] (pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*4+3) / (subpic_grid_col_width_minus1*4+4) (7-5)
[0063] subpic_grid_row_height_minus1 plus 1 specifies the height of each element of the sub-picture identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.
[0064] The variable NumSubPicGridRows is derived as follows:
[0065] NumSubPicGridRows=
[0066] (pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*4+3) / (subpic_grid_row_height_minus1*4+4) (7-6)
[0067] subpic_grid_idx[i][j] specifies the sub-picture index of the grid position (i, j). The length of the syntax element is Ceil(Log2(max_subpics_minus1+1)) bits.
[0068] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]], and NumSubPics are derived as follows:
[0069]
[0070]
[0071] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th sub-picture of each codec picture in the CVS is treated as a picture in the decoding process that does not include loop filtering operations.
[0072] subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each codec picture in the CVS is not to be treated as a picture in the decoding process that does not include loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to 0.
[0073] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations can be performed across the boundaries of the i-th sub-picture of each coded picture in the CVS.
[0074] loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that loop filtering operations are not performed across the boundaries of the i-th sub-picture of each codec picture in the CVS. When not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0075] The requirements for bitstream conformance are that the following constraints apply:
[0076] – For any two sub-pictures subpicA and subpicB, when the index of subpicA is less than the index of subpicB, in decoding order, any coded NAL unit of subPicA will be after any coded NAL unit of subPicB.
[0077] – The shape of the sub-pictures shall be such that when decoded, the entire left and the entire top border of each sub-picture shall consist of the picture border, or of the borders of previously decoded sub-pictures.
[0078] The list CtbToSubPicIdx[ctbAddrRs] specifies the conversion from CTB addresses in the picture raster scan to sub-picture indices for ctbAddrRs in the range 0 to PicSizeInCtbsY-1 (inclusive), which is derived as follows:
[0079]
[0080]
[0081] num_bricks_in_slice_minus1, when present, specifies the number of bricks in the slice minus 1. The value of num_bricks_in_slice_minus1 shall be in the range of 0 to NumBricksInPic-1, inclusive. When rect_slice_flag is equal to 0 and single_brick_per_slice_flag is equal to 1, the value of num_bricks_in_slice_minus1 is inferred to be equal to 0. When
[0082] When single_brick_per_slice_flag is equal to 1, the value of num_bricks_in_slice_minus1 is inferred to be equal to 0.
[0083] The variable NumBricksInCurrSlice specifies the number of bricks in the current slice, and SliceBrickIdx[i] specifies the brick index of the i-th brick in the current slice, which is derived as follows:
[0084]
[0085] The variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[0086]
[0087] Derivation process of temporal luminance motion vector prediction
[0088] The inputs to this process are:
[0089] – The luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,
[0090] –The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0091] –The variable cbHeight specifies the height of the current codec block in units of luminance samples.
[0092] – Reference index refIdxLX, where X is 0 or 1.
[0093] The output of this process is:
[0094] –1 / 16 fractional sample accuracy motion vector prediction mvLXCol,
[0095] – Availability flag availableFlagLXCol.
[0096] The variable currCb specifies the current luma codec block at luma location (xCb, yCb).
[0097] The variables mvLXCol and availableFlagLXCol are derived as follows:
[0098] – If slice_temporal_mvp_enabled_flag is equal to 0 or (cbWidth*cbHeight) is less than or equal to 32, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[0099] – Otherwise (slice_temporal_mvp_enabled_flag is equal to 1), apply the following ordered steps:
[0100] 1. The derivation of the collocated motion vector at the bottom right and the positions of the bottom and right boundary samples is as follows:
[0101] xColBr=xCb+cbWidth (8-421)
[0102] yColBr=yCb+cbHeight (8-422)
[0103] rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicRightBoundaryPos:pic_width_in_luma_samples-1(8-423)
[0104] botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos:pic_height_in_luma_samples-1(8-424)
[0105] – If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos, and xColBr is less than or equal to rightBoundaryPos, then the following applies:
[0106] – The variable colCb specifies the luma codec block covering the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the collocated picture specified by ColPic.
[0107] – The luma position (xColCb, yColCb) is set equal to the top left luma sample of the collocated luma codec block specified by colCb relative to the top left luma sample of the collocated picture specified by ColPic.
[0108] – Invoke the derivation process of the collocated motion vector specified in clause 8.5.2.12 with currCb, colCb, (xColCb, yColCb), refIdxLX and sbFlag set to 0 as input and the output assigned to mvLXCol and availableFlagLXCol.
[0109] Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[0110] …
[0111] Luminance sample bilinear interpolation process
[0112] The inputs to this process are:
[0113] – Luminance position in full sample units (xInt L ,yInt L ),
[0114] – Luma position in fractional samples (xFrac L ,yFrac L ),
[0115] – Luminance reference sample array refPicLX L .
[0116] The output of this process is the predicted luminance sample value predSampleLX L
[0117] The variables shift1, shift2, shift3, shift4, offset1, offset2, and offset3 are derived as follows:
[0118] shift1 = BitDepth Y -6 (8-453)
[0119] offset1=1<<(shift1-1) (8-454)
[0120] shift2=4 (8-455)
[0121] offset2=1<<(shift2-1) (8-456)
[0122] shift3=10-BitDepth Y (8-457)
[0123] shift4=BitDepth Y -10 (8-458)
[0124] offset4=1<<(shift4-1) (8-459)
[0125] The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples.
[0126] Luma interpolation filter coefficient fb for each 1 / 16 fractional sample position p L [p] is equal to xFrac L or yFrac L , specified in Table 8-10.
[0127] ■For i=0..1, the luminance position in units of full samples (xInt i ,yInt i ) is derived as follows:
[0128] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following conditions apply: xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i)(8-460)
[0129] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i)(8-461)
[0130] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following conditions apply: xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?
[0131] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,(xInt L +i)):(8-462)
[0132] xInt L +i)
[0133] yInt i =Clip3(0,picH-1,yInt L +i) (8-463)
[0134] …
[0135] Derivation process of sub-block-based temporal merging candidates
[0136] The inputs to this process are:
[0137] – The luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,
[0138] –The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0139] –The variable cbHeight specifies the height of the current codec block in units of luma samples.
[0140] – The availability flag of the adjacent codec unit availableFlagA1,
[0141] – the reference index of the adjacent codec unit refIdxLXA1,
[0142] – The prediction list of adjacent codec units uses the flag predFlagLXA1, where X is 0 or 1
[0143] – The motion vector mvLXA1 of the adjacent codec unit with 1 / 16 fractional sample accuracy, where X is 0 or 1.
[0144] The output of this process is:
[0145] – Availability flag availableFlagSbCol,
[0146] – The number of luminance codec sub-blocks in the horizontal direction numSbX and the number in the vertical direction numSbY,
[0147] – reference indexes refIdxL0SbCol and refIdxL1SbCol,
[0148] – Luma motion vectors mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx] with 1 / 16 fractional sample accuracy, xSbIdx = 0..numSbX–1, ySbIdx = 0..numSbY-1,
[0149] – Prediction list utilization flag pred flag 0sbcol[xs bidx][ySbIdx] and pred flag 1sbcol[xs bidx][ySbIdx], xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1.
[0150] The derivation of the availability flag availableFlagSbCol is as follows.
[0151] – If one or more of the following conditions are true, availableFlagSbCol is set equal to 0.
[0152] –slice_temporal_mvp_enabled_flag is equal to 0.
[0153] –sps_sbtmvp_enabled_flag is equal to 0.
[0154] –cbWidth is less than 8.
[0155] –cbHeight is less than 8.
[0156] – Otherwise, the following ordered steps apply:
[0157] 1. The position of the upper left sample point (xCtb, yCtb) of the luma coding tree block containing the current codec block and the position of the lower right center sample point (xCtr, yCtr) of the current luma codec block are derived as follows:
[0158] xCtb=(xCb>>CtuLog2Size) <CtuLog2Size (8-542)
[0159] yCtb=(yCb>>CtuLog2Size)< <CtuLog2Size (8-543)
[0160] xCtr=xCb+(cbWidth / 2) (8-544)
[0161] yCtr=yCb+(cbHeight / 2) (8-545)
[0162] 2. The luma position (xColCtrCb, yColCtrCb) is set equal to the top left luma sample of the collocated luma coding block covering the position given by (xCtr, yCtr) inside ColPic, relative to the top left luma sample of the collocated picture specified by ColPic.
[0163] 3. Invoke the derivation process of the sub-block based temporal merging basic motion data specified in clause 8.5.5.4, with position (xCtb, yCtb), position (xColCtrCb, yColCtrCb), availability flag availableFlagA1, prediction list utilization flag predFlagLXA1, reference index refIdxLXA1 and motion vector mvLXA1 as input, where X is 0 and 1, and with motion vector ctrMvLX and the prediction list utilization flag ctrPredFlagLX of the collocated block and temporal motion vector tempMv as output, where X is 0 and 1.
[0164] 4. The variable availableFlagSbCol is derived as follows:
[0165] – If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, availableFlagSbCol is set equal to 0.
[0166] – Otherwise, set availableFlagSbCol equal to 1.
[0167] When availableFlagSbCol is equal to 1, the following applies:
[0168] – The variables numSbX, numSbY, sbWidth, sbHeight and refIdxLXSbCol are derived as follows:
[0169] numSbX=cbWidth>>3 (8-546)
[0170] numSbY=cbHeight>>3 (8-547)
[0171] sbWidth=cbWidth / numSbX (8-548)
[0172] sbHeight=cbHeight / numSbY (8-549)
[0173] refIdxLXSbCol=0 (8-550)
[0174] – For xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY-1, the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list utilization flag
[0175] The derivation of predFlagLXSbCol[xSbIdx][ySbIdx] is as follows:
[0176] – The luminance position (xSb, ySb) of the upper left sample of the current codec sub-block relative to the upper left luminance sample of the current picture is derived as follows:
[0177] xSb=xCb+xSbIdx*sbWidth+sbWidth / 2 (8-551)
[0178] ySb=yCb+ySbIdx*sbHeight+sbHeight / 2 (8-552)
[0179] The position of the collocated sub-block within ColPic (xColSb, yColSb) is derived as follows.
[0180] – The following applies to:
[0181] yColSb=Clip3(yCtb,
[0182] Min(CurPicHeightInSamplesY-1,yCtb+(1< <CtbLog2SizeY)-1), (8-553)
[0183] ySb+(tempMv[1]>>4))
[0184] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following conditions apply:
[0185] xColSb=Clip3(xCtb,
[0186] Min(SubPicRightBoundaryPos,xCtb+(1< <CtbLog2SizeY)+3), (8-554)
[0187] xSb+(tempMv[0]>>4))
[0188] – Otherwise (sub pic_treated_as_pic_flag[sub picidx] is equal to 0), the following applies:
[0189] xColSb=Clip3(xCtb,
[0190] Min(CurPicWidthInSamplesY-1,xCtb+(1< <CtbLog2SizeY)+3), (8-555)
[0191] xSb+(tempMv[0]>>4))
[0192] …
[0193] Derivation process of basic motion data based on sub-block temporal merging
[0194] The inputs to this process are:
[0195] – contains the position of the top left sample of the luma coding tree block of the current codec block (xCtb, yCtb),
[0196] – The position of the upper left sample of the collocated luma codec block covering the lower right center sample (xColCtrCb, yColCtrCb).
[0197] – The availability flag of the adjacent codec unit availableFlagA1,
[0198] – the reference index of the adjacent codec unit refIdxLXA1,
[0199] – The prediction list of adjacent codec units uses the flag predFlagLXA1,
[0200] – Motion vector mvLXA1 of the adjacent codec unit with 1 / 16 fractional sample accuracy.
[0201] The output of this process is:
[0202] – motion vectors ctrMvL0 and ctrMvL1,
[0203] – The prediction list utilizes the flags ctrPredFlagL0 and ctrPredFlagL1,
[0204] – Temporal motion vector tempMv.
[0205] The variable tempMv is set as follows:
[0206] tempMv[0]=0 (8-558)
[0207] tempMv[1]=0 (8-559) The variable currPic specifies the current picture.
[0208] When availableFlagA1 equals TRUE, the following applies:
[0209] – If all of the following conditions are TRUE, then tempMv is set equal to mvL0A1:
[0210] –predFlagL0A1 is equal to 1,
[0211] –DiffPicOrderCnt(ColPic, RefPicList[0][refIdxL0A1]) is equal to 0,
[0212] – Otherwise, if all of the following conditions are true, then tempMv is set equal to mvL1A1:
[0213] –slice_type is equal to B,
[0214] –predFlagL1A1 is equal to 1,
[0215] –DiffPicOrderCnt(ColPic, RefPicList[1][refIdxL1A1]) is equal to 0.
[0216] The location (xColCb, yColCb) of the collocated block within ColPic is derived as follows.
[0217] – The following applies to:
[0218] yColCb=Clip3(yCtb,
[0219] Min(CurPicHeightInSamplesY-1,yCtb+(1< <CtbLog2SizeY)-1), (8-560)
[0220] yColCtrCb+(tempMv[1]>>4))
[0221] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following conditions apply:
[0222] xColCb=Clip3(xCtb,
[0223] Min(SubPicRightBoundaryPos,xCtb+(1< <CtbLog2SizeY)+3), (8-561)
[0224] xColCtrCb+(tempMv[0]>>4))
[0225] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following applies:
[0226] xColCb=Clip3(xCtb,
[0227] Min(CurPicWidthInSamplesY-1,xCtb+(1< <CtbLog2SizeY)+3), (8-562)
[0228] xColCtrCb+(tempMv[0]>>4))
[0229] Luminance sample interpolation filtering process
[0230] The inputs to this process are:
[0231] – Luminance position in full sample units (xInt L , yInt L ),
[0232] – Luma position in fractional samples (xFrac L ,yFrac L ),
[0233] – Luminance position in units of full samples (xSbInt L ,ySbInt L ), specifies the upper left sample of the boundary block used for reference sample filling relative to the upper left luminance sample of the reference picture,
[0234] – Luminance reference sample array refPicLX L ,
[0235] – Half-sample interpolation filter index hpelIfIdx,
[0236] – variable sbWidth specifies the width of the current sub-block,
[0237] – Variable sbHeight specifies the current sub-block height,
[0238] –Specify the luminance position (xSb, ySb) of the upper left sample of the current sub-block relative to the upper left luminance sample of the current picture,
[0239] The output of this process is the predicted luminance sample value predSampleLX L
[0240] The variables shift1, shift2, and shift3 are derived as follows:
[0241] – The variable shift1 is set equal to Min(4, BitDepth Y -8), the variable shift2 is set to 6, and the variable shift3 is set to Max(2, 14-BitDepth Y ).
[0242] – The variable picW is set equal to pic_width_in_luma_samples and the variable picH is set equal to pic_height_in_luma_samples.
[0243] Luma interpolation filter coefficient f for each 1 / 16 fractional sample position p L [p] is equal to xFrac L or yFrac L , which is derived as follows:
[0244] – If MotionModelIdc[xSb][ySb] is greater than 0, and sbWidth and sbHeight are both equal to 4, then the brightness interpolation filter coefficient f L [p] is specified in Table 8-12.
[0245] – Otherwise, the luminance interpolation filter coefficient fL [p] Specified in Table 8-11 depending on hpelIfIdx.
[0246] For i=0..7, the luminance position in units of full samples (xInt i 、yInt i ) is derived as follows:
[0247] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following conditions apply:
[0248] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3)(8-771)
[0249] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772)
[0250] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following conditions apply:
[0251] xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?
[0252] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3): (8-773)
[0253] xInt L +i-3)
[0254] yInt i =Clip3(0,picH-1,yInt L +i-3)
[0255] (8-774)
[0256] …
[0257] Chroma sample interpolation process
[0258] The inputs to this process are:
[0259] – Chroma position in units of full sample points (xInt C , yInt C ),
[0260] – Chroma position in 1 / 32 fractional samples (xFrac C ,yFrac C ),
[0261] – chroma position in full sample units (xSbIntC, ySbIntC), which specifies the top left sample of the boundary block used for reference sample filling relative to the top left chroma sample of the reference picture,
[0262] – variable sbWidth specifies the width of the current sub-block,
[0263] – Variable sbHeight specifies the current sub-block height,
[0264] – Chroma reference sample array refPicLX C .
[0265] The output of this process is the predicted chroma sample value predSampleLX C
[0266] The variables shift1, shift2, and shift3 are derived as follows:
[0267] – Variable shift1 is set equal to Min(4, B BitDepth C -8), the variable shift2 is set to 6, and the variable shift3 is set to Max(2, 14-BitDepth C ).
[0268] –Variable picW C Set equal to pic_width_in_luma_samples / SubWidthC, variable picH C Set equal to pic_height_in_luma_samples / SubHeightC.
[0269] Chroma interpolation filter coefficient f for each 1 / 32 fractional sample position p C [p] is equal to xFrac C or yFrac C , are specified in Table 8-13.
[0270] The variable xOffset is set equal to (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC.
[0271] For i = 0..3, the chroma position in units of full samples (xInt i ,yInt i ) is derived as follows:
[0272] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following conditions apply: xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (8-785)
[0273] yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (8-786)
[0274] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following conditions apply: xInt i =Clip3(0,picW C -1,
[0275] sps_ref_wraparound_enabled_flag? ClipH(xOffset,picW C ,xInt C +i-1):(8-787)
[0276] xInt C +i-1)
[0277] yInt i =Clip3(0,picH C -1,yInt C +i-1)
[0278] (8-788)
[0279] 2.4 GOP-based temporal filter for encoder only (JCTVC-AI0023)
[0280] JCTVC-AI0023 proposes an encoder-only temporal filter. This filtering is performed at the encoder as a preprocessing step. Source pictures preceding and following the selected picture to be encoded are read and a block-based motion compensation method is applied to these source pictures relative to the selected picture. Temporal filtering is then performed on the samples in the selected picture using the motion-compensated sample values.
[0281] The overall filtering strength is set depending on the temporal sublayer and QP of the selected picture. Only pictures of temporal sublayers 0 and 1 are filtered, with pictures of layer 0 being filtered with a stronger filter than pictures of layer 1. The per-sample filtering strength is adjusted based on the difference between the sample value in the selected picture and the collocated sample in the motion compensated picture, so that small differences between the motion compensated picture and the selected picture are filtered more strongly than large differences.
[0282] GOP-based temporal filter
[0283] The temporal filter is introduced directly after reading the picture and before encoding. The following is a more detailed description of the steps.
[0284] Step 1: Read the image from the encoder
[0285] Step 2: If the picture is low enough in the codec hierarchy, it is filtered before encoding. Otherwise, the picture is encoded without filtering. RA pictures with POC% 8 == 0 and LD pictures with POC% 4 == 0 are filtered. AI pictures are never filtered.
[0286] The overall filter strength is set for RA according to the following equation.
[0287]
[0288] Where n is the number of images read.
[0289] For the LD case, use s o (n)=0.95.
[0290] Step 3: Read the two pictures before and / or after the selected picture (hereinafter referred to as the original picture). In edge cases, for example, if it is the first picture or close to the last picture, only read the available pictures.
[0291] Step 4: For each 8×8 image block, estimate the motion before and after the read image relative to the original image.
[0292] Use a hierarchical motion estimation scheme, and Figure 2 Layers L0, L1 and L2 are shown in FIG. By comparing all read images and original images (i.e. Figure 1The subsampled image is generated by averaging each 2×2 block of L1 in
[15] . L2 is obtained from L1 using the same subsampling method.
[0293] Figure 2 An example of different layers of hierarchical motion estimation is shown. L0 is the original resolution. L1 is a subsampled version of L0. L2 is a subsampled version of L1.
[0294] First, motion estimation is performed for each 16×16 block in L2. The squared difference is calculated for each selected motion vector, and the motion vector corresponding to the minimum difference is selected. The selected motion vector is then used as the initial value when estimating motion in L1. The same operation is then performed for motion estimation in L0. As a final step, sub-pixel motion is estimated for each 8×8 block by using an interpolation filter on L0.
[0295] Using the VTM 6-tap interpolation filter:
[0296] 0:0,0,64,0,0,0
[0297] 1:1,-3,64,4,-2,0
[0298] 2:1,-6,62,9,-3,1
[0299] 3:2,-8,60,14,-5,1
[0300] 4:2,-9,57,19,-7,2
[0301] 5:3,-10,53,24,-8,2
[0302] 6:3,-11,50,29,-9,2
[0303] 7:3,-11,44,35,-10,3
[0304] 8:1,-7,38,38,-7,1
[0305] 9:3,-10,35,44,-11,3
[0306] 10:2,-9,29,50,-11,3
[0307] 11:2,-8,24,53,-10,3
[0308] 12:2,-7,19,57,-9,2
[0309] 13:1,-5,14,60,-8,2
[0310] 14:1,-3,9,62,-6,1
[0311] 15:0,-2,4,64,-3,1
[0312] Step 5: Apply motion compensation to the pictures before and after the original picture based on the best matching motion of each block. That is, make the sample coordinates of the original picture in each block have the best matching coordinates in the reference picture.
[0313] Step 6: Process the samples of the luma and chroma channels one by one as described in the following steps.
[0314] Step 7: Calculate the new sample value I using the following formula n .
[0315]
[0316] Among them I o is the sample value of the original sample point, I r (i) is the intensity of the corresponding sample of motion compensated picture i, and w r (i, a) is the weight of the motion compensated picture i when the number of available motion compensated pictures is a.
[0317] In the brightness channel, the weight w r (i,a) is defined as follows:
[0318]
[0319] in
[0320] s l =0.4
[0321]
[0322]
[0323] For all other cases of i and a: s r (i,a)=0.3
[0324] σ l (QP) = 3*(QP-10)
[0325] ΔI(i)=I r (i)-I o
[0326] For the chroma channel, the weight w r (i,a) is defined as follows:
[0327]
[0328] where s c =0.55 and σc =30.
[0329] Step 8: Apply the filter to the current sample. The resulting sample value is stored separately.
[0330] Step 9: Encode the filtered image.
[0331] 2.5. Image Segmentation (Slices, Bricks, Strips) in JVET-O2001-vE
[0332] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a series of CTUs covering a rectangular area of a picture.
[0333] A slice is divided into one or more bricks, and each brick consists of multiple CTU rows within the slice.
[0334] A slice that is not split into multiple bricks is also called a brick. However, a brick that is a proper subset of a slice is not called a slice.
[0335] A strip contains either multiple slices of an image, or multiple tiles of a slice.
[0336] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0337] Two striping modes are supported: raster scan striping and rectangular striping. In raster scan striping, a strip consists of a sequence of slices from the image's raster scan. In rectangular striping, a strip consists of multiple bricks of the image, which together form a rectangular region of the image. Bricks within a rectangular strip are arranged in the strip's brick raster scan order.
[0338] Figure 3 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0339] Figure 4 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0340] Figure 5 An example of a picture partitioned into slices, bricks, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the upper left slice contains 1 brick, the upper right slice contains 5 bricks, the lower left slice contains 2 bricks, and the lower right slice contains 3 bricks), and 4 rectangular strips.
[0341] Picture Parameter Set RBSP Syntax
[0342]
[0343]
[0344]
[0345]
[0346] single_tile_in_pic_flag equal to 1 specifies that there is only one slice in each picture of the referenced PPS.
[0347] single_tile_in_pic_flag equal to 0 specifies that there is more than one slice in each picture of the reference PPS.
[0348] Note – If there is no further bricking within a slice, then the entire slice is called a brick. When an image contains only a single slice without further bricking, it is called a single brick.
[0349] The bitstream conformance requirement is that the value of single_tile_in_pic_flag should be the same for all PPSs referenced by a coded picture within a CVS.
[0350] uniform_tile_spacing_flag equal to 1 specifies that slice column boundaries and slice row boundaries are uniformly distributed across the picture and is signaled using the syntax elements tile_cols_width_minus1 and tile_rows_height_minus1. uniform_tile_spacing_flag equal to 0 specifies that slice column boundaries and slice row boundaries may or may not be uniformly distributed across the picture and is signaled using the syntax elements num_tile_columns_minus1 and num_tile_rows_minus1 and a list of syntax element pairs tile_column_width_minus1[i] and tile_row_height_minus1[i]. When not present, the value of uniform_tile_spacing_flag is inferred to be equal to 1.
[0351] tile_cols_width_minus1 plus 1 specifies the width of the tile columns in the picture, excluding the rightmost tile column, in units of CTBs when uniform_tile_spacing_flag is equal to 1. The value of tile_cols_width_minus1 shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. When not present, the value of tile_cols_width_minus1 is inferred to be equal to PicWidthInCtbsY-1.
[0352] tile_rows_height_minus1 plus 1 specifies the height of the tile rows in the picture, excluding the bottom tile row, in units of CTBs when uniform_tile_spacing_flag is equal to 1. The value of tile_rows_height_minus1 shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. When not present, the value of tile_rows_height_minus1 is inferred to be equal to PicHeightInCtbsY-1.
[0353] num_tile_columns_minus1 plus 1 specifies the number of tile columns used to partition the picture when uniform_tile_spacing_flag is equal to 0. The value of num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_columns_minus1 is inferred as specified in clause 6.5.1.
[0354] num_tile_rows_minus1 plus 1 specifies the number of tile rows used to partition the picture when uniform_tile_spacing_flag is equal to 0. The value of num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_rows_minus1 is inferred as specified in clause 6.5.1.
[0355] The variable NumTilesInPic is set equal to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1).
[0356] When single_tile_in_pic_flag is equal to 0, NumTilesInPic should be greater than 1.
[0357] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of CTB.
[0358] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of CTB.
[0359] brick_splitting_present_flag equal to 1 specifies that one or more slices of a picture that references a PPS may be split into two or more bricks. brick_splitting_present_flag equal to 0 specifies that slices of a picture that does not reference a PPS are split into two or more bricks.
[0360] num_tiles_in_pic_minus1 plus 1 specifies the number of tiles in each picture of the referenced PPS. The value of num_tiles_in_pic_minus1 shall be equal to NumTilesInPic – 1. When not present, the value of num_tiles_in_pic_minus1 is inferred to be equal to NumTilesInPic – 1.
[0361] brick_split_flag[i] equal to 1 specifies that the i-th slice is split into two or more bricks. brick_split_flag[i] equal to 0 specifies that the i-th slice is not split into two or more bricks. When absent, the value of brick_split_flag[i] is inferred to be equal to 0. [Ed.(HD / YK): SPS-dependent PPS parsing is introduced by adding the syntax condition "if(RowHeight[i]>1". The same is true for uniform_brick_spacing_flag[i].]
[0362] uniform_brick_spacing_flag[i] equal to 1 specifies that the horizontal brick boundaries are uniformly distributed over the i-th slice and is signaled using the syntax element brick_height_minus1[i].
[0363] uniform_brick_spacing_flag[i] equal to 0 specifies that the horizontal brick boundaries may or may not be uniformly distributed across the i-th slice, and is signaled using the list of syntax elements num_brick_rows_minus2[i] and brick_row_height_minus1[i][j]. When not present, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1.
[0364] brick_height_minus1[i] plus 1 specifies the height of the brick row in CTB units, excluding the bottom brick, in the i-th slice when uniform_brick_spacing_flag[i] is equal to 1. If present, the value of brick_height_minus1 shall be in the range of 0 to RowHeight[i]-2, inclusive. When not present, the value of brick_height_minus1[i] is inferred to be equal to RowHeight[i]-1.
[0365] num_brick_rows_minus2[i] plus 2 specifies the number of bricks used to split the i-th slice when uniform_brick_spacing_flag[i] is equal to 0. When present, the value of num_brick_rows_minus2[i] shall be in the range of 0 to RowHeight[i]-2, inclusive. If brick_split_flag[i] is equal to 0, the value of num_brick_rows_minus2[i] is inferred to be equal to -1. Otherwise, when uniform_brick_spacing_flag[i] is equal to 1, the value of num_brick_rows_minus2[i] is inferred as specified in clause 6.5.1.
[0366] brick_row_height_minus1[i][j] plus 1 specifies the height of the j-th brick in the i-th slice in CTB units when uniform_tile_spacing_flag is equal to 0.
[0367] The following variables are derived, and the values of num_tile_columns_minus1 and num_tile_rows_minus1 are derived when uniform_tile_spacing_flag is equal to 1, and the value of num_brick_rows_minus2[i] is derived when uniform_brick_spacing_flag[i] is equal to 1 for each i in the range from 0 to NumTilesInPic-1 (inclusive), by invoking the CTB raster and brick scan conversion process specified in clause 6.5.1:
[0368] – List RowHeight[j] specifies the height of the j-th tile row in CTB units, where j ranges from 0 to num_tile_rows_minus1, including 0 and num_tile_rows_minus1.
[0369] – List CtbAddrRsToBs[ctbAddrRs] specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the brick scan. The range of ctbAddrRs is from 0 to PicSizeInCtbsY-1, including 0 and PicSizeInCtbsY-1.
[0370] – List CtbAddrBsToRs[ctbAddrBs] specifies the conversion from CTB addresses in the brick scan to CTB addresses in the CTB raster scan of the picture. The range of ctbAddrRs is from 0 to PicSizeInCtbsY-1, including 0 and PicSizeInCtbsY-1.
[0371] – List BrickId[ctbAddrBs] specifies the conversion from CTB address in brick scan to brick ID, ctbAddrBs ranges from 0 to PicSizeInCtbsY-1, including 0 and PicSizeInCtbsY-1,
[0372] – The list NumCtusInBrick[brickIdx] specifies the conversion from brick index to the number of CTUs in the brick, brickIdx ranges from 0 to NumBricksInPic–1, inclusive.
[0373] – List FirstCtbAddrBs[brickIdx] specifies the translation from brick ID to CTB address in brick scan of the first CTB in the brick, brickIdx ranges from 0 to NumBricksInPic–1, inclusive.
[0374] single_brick_per_slice_flag equal to 1 specifies that each slice of the referenced PPS includes one brick. single_brick_per_slice_flag equal to 0 specifies that a slice of the referenced PPS may include more than one brick. When not present, the value of single_brick_per_slice_flag is inferred to be equal to 1.
[0375] rect_slice_flag equal to 0 specifies that the bricks within each slice are in raster scan order and that slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the bricks within each slice cover a rectangular area of the picture and that slice information is signaled in the PPS. When brick_splitting_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1. When not present, rect_slice_flag is inferred to be equal to 1.
[0376] num_slices_in_pic_minus1 plus 1 specifies the number of slices in each picture of the referenced PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to NumBricksInPic-1, inclusive. When not present and single_brick_per_slice_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to NumBricksInPic-1.
[0377] bottom_right_brick_idx_length_minus1 plus 1 specifies the number of bits used to represent the syntax element bottom_right_brick_idx_delta[i].
[0378] The value of bottom_right_brick_idx_length_minus1 should be in the range of 0 to Ceil(Log2(NumBricksInPic))-1, inclusive.
[0379] bottom_right_brick_idx_delta[i] specifies the difference between the brick index of the bottom right corner of the i-th slice and the brick index of the bottom right corner of the (i–1)-th slice, when i is greater than 0. bottom_right_brick_idx_delta[0] specifies the brick index of the bottom right corner of the 0-th slice. When single_brick_per_slice_flag is equal to 1, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 1. The value of bottomRightBrickIdx[num_slices_in_pic_minus1] is inferred to be equal to NumBricksInPic–1. The length of the bottom_right_brick_idx_delta[i] syntax element is bottom_right_brick_idx_length_minus1+1 bits.
[0380] brick_idx_delta_sign_flag[i] plus 1 indicates the positive sign of bottom_right_brick_idx_delta[i]. sign_bottom_right_brick_idx_delta[i] equal to 0 indicates the negative sign of bottom_right_brick_idx_delta[i].
[0381] The bitstream conformance requirement is that a slice should consist of either multiple complete slices, or a contiguous sequence of complete bricks of just one slice.
[0382] The variables TopLeftBrickIdx[i], BottomRightBrickIdx[i], NumBricksInSlice[i], and BricksToSliceMap[j] specify the brick index of the brick at the top left corner of the i-th stripe, the brick index of the brick at the bottom right corner of the i-th stripe, the number of bricks in the i-th stripe, and the brick-to-strip mapping, which are derived as follows:
[0383]
[0384]
[0385] Common Strip Header Semantics
[0386] When present, the value of each of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and lice_temporal_mvp_enabled_flag shall be the same in all slice headers of a codec picture.
[0387] The variable CuQpDeltaVal specifies the difference between the luma quantization parameter of the codec unit containing cu_qp_delta_abs and its prediction, and is set equal to 0. The variable CuQpOffset Cb 、CuQpOffset Cr and CuQpOffset CbCr Specifies the Qp' of the codec unit containing cu_chroma_qp_offset_flag Cb 、Qp' Cr and Qp' CbCr The corresponding values of the quantization parameters are the values to be used when these variables are set equal to 0.
[0388] slice_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id for the PPS being used. The value of slice_pic_parameter_set_id shall be in the range of 0 to 63, inclusive.
[0389] The bitstream conformance requirement is that the value of TemporalId of the current picture should be greater than or equal to the value of TemporalId of the PPS whose pps_pic_parameter_set_id is equal to slice_pic_parameter_set_id.
[0390] slice_address specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0.
[0391] If rect_slice_flag is equal to 0, the following applies:
[0392] – The stripe address is the brick ID specified by equation (7-59).
[0393] – The length of slice_address is Ceil(Log2(NumBricksInPic)) bits.
[0394] The value of –slice_address should be in the range of 0 to NumBricksInPic-1, inclusive.
[0395] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0396] – Stripe address is the stripe ID of the stripe.
[0397] – The length of slice_address is signalled_slice_id_length_minus1+1 bits.
[0398] – If signalled_slice_id_flag is equal to 0, the value of slice_address shall be in the range of 0 to num_slices_in_pic_minus1, inclusive. Otherwise, the value of slice_address shall be in the range of 0 to 2 (signalled_slice_id_length_minus1+1) -1, including 0 and 2 (signalled _slice_id_length_minus1+1) -1.
[0399] The requirements for bitstream conformance are that the following constraints apply:
[0400] – The value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.
[0401] – When rect_slice_flag is equal to 0, the slices of the picture will be arranged in ascending order of their slice_address values.
[0402] – The shape of a picture slice shall be such that, when decoded, each tile shall have its entire left and entire top border consisting of either the picture border or the borders of the previously decoded tile(s).
[0403] num_bricks_in_slice_minus1, if present, specifies the number of bricks in the slice minus 1. The value of num_bricks_in_slice_minus1 shall be in the range of 0 to NumBricksInPic-1, inclusive. The value of num_bricks_in_slice_minus1 is inferred to be equal to 0 when rect_slice_flag is equal to 0 and single_brick_per_slice_flag is equal to 1. The value of num_bricks_in_slice_minus1 is inferred to be equal to 0 when single_brick_per_slice_flag is equal to 1.
[0404] The variable NumBricksInCurrSlice specifies the number of bricks in the current strip, and SliceBrickIdx[i] specifies the brick index of the i-th brick in the current strip, which is derived as follows:
[0405]
[0406] The variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[0407]
[0408]
[0409] 2.6 Syntax and Semantics in JVET-P2001-V8
[0410] Sequence Parameter Set RBSP Syntax
[0411]
[0412]
[0413]
[0414]
[0415]
[0416]
[0417]
[0418]
[0419] Picture Parameter Set RBSP Syntax
[0420]
[0421]
[0422]
[0423]
[0424]
[0425] Image header RBSP syntax
[0426]
[0427]
[0428]
[0429]
[0430]
[0431]
[0432]
[0433] subpics_present_flag equal to 1 indicates that sub-picture parameters are present in the SPS RBSP syntax. subpics_present_flag equal to 0 indicates that sub-picture parameters are not present in the SPS RBSP syntax.
[0434] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPSs.
[0435] sps_num_subpics_minus1 specifies the number of sub-pictures plus 1. sps_num_subpics_minus1 should be in the range of 0 to 254. When not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.
[0436] subpic_ctu_top_left_x[i] specifies the horizontal position of the top-left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0437] subpic_ctu_top_left_y[i] specifies the vertical position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0438] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.
[0439] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.
[0440] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th sub-picture of each codec picture in the CVS is treated as a picture in the decoding process that does not include loop filtering operations. subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th sub-picture of each codec picture in the CVS is treated as a picture in the decoding process that does not include loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to 0.
[0441] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that loop filtering operations may be performed across the boundaries of the i-th sub-picture of each codec picture in the CVS. loop_filter_cross_subpic_enabled_flag[i] equal to 0 specifies that loop filtering operations are not performed across the boundaries of the i-th sub-picture of each codec picture in the CVS. When not present, the value of loop_filter_cross_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0442] The requirements for bitstream conformance are that the following constraints apply:
[0443] – For any two sub-pictures subpicA and subpicB, when the index of subpicA is less than the index of subpicB, in decoding order, any codec NAL unit of subPicA will be after any codec NAL unit of subPicB.
[0444] – The shape of sub-pictures shall be such that, when decoded, the entire left and the entire top border of each sub-picture shall consist of the picture border, or of the borders of previously decoded sub-pictures.
[0445] sps_subpic_id_present_flag equal to 1 specifies that sub-picture ID mapping is present in the SPS. sps_subpic_id_present_flag equal to 0 specifies that sub-picture ID mapping is not present in the SPS.
[0446] sps_subpic_id_signaling_present_flag equal to 1 specifies that sub-picture Id mapping is signaled in the SPS. sps_subpic_id_signaling_present_flag equal to 0 specifies that sub-picture Id mapping is not signaled in the SPS. When not present, the value of sps_subpic_id_signaling_present_flag is inferred to be equal to 0.
[0447] sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.
[0448] sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. When not present and sps_subpic_id_present_flag is equal to 0, for each i in the range of 0 to sps_num_subpics_minus1, inclusive, the value of sps_subpic_id[i] is inferred to be equal to i
[0449] ph_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS being used. The value of ph_pic_parameter_set_id shall be in the range of 0 to 63, inclusive.
[0450] The bitstream conformance requirement is that the value of TemporalId in the picture header should be greater than or equal to the value of TemporalId of the PPS with pps_pic_parameter_set_id equal to ph_pic_parameter_set_id.
[0451] ph_subpic_id_signaling_present_flag equal to 1 specifies that sub-picture ID mapping is signaled in the picture header. ph_subpic_id_signaling_present_flag equal to 0 indicates that sub-picture ID mapping is not signaled in the picture header.
[0452] ph_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element ph_subpic_id[i]. The value of pic_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.
[0453] A bitstream conformance requirement is that the value of ph_subpic_id_len_minus1 shall be the same for all picture headers referenced by coded pictures in the CVS.
[0454] ph_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the ph_subpic_id[i] syntax element is ph_subpic_id_len_minus1+1 bits.
[0455] The derivation of the list SubpicIdList[i] is as follows:
[0456] The list SubpicIdList[i] is derived as follows:
[0457]
[0458] Deblocking filtering process
[0459] Overview
[0460] The input of this process is the reconstructed image before deblocking, that is, the array recPicture L , and when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[0461] The output of this process is the modified reconstructed image after deblocking, that is, the array recPicture L , and when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[0462] First, vertical edges in the image are filtered. Then, horizontal edges in the image are filtered using the samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTBs of each CTU are processed separately on a codec unit basis. Vertical edges of a codec block within a codec unit are filtered starting with the edge on the left side of the codec block and proceeding through the edges in their geometric order toward the right side of the codec block. Horizontal edges of a codec block within a codec unit are filtered starting with the edge at the top of the codec block and proceeding through the edges in their geometric order toward the bottom of the codec block.
[0463] NOTE – Although the filtering process is specified on a picture basis in this specification, the filtering process can also be implemented on a codec unit basis with equivalent results, as long as the decoder correctly respects the processing dependency order to produce the same output values.
[0464] The deblocking filtering process is applied to all codec sub-block edges and transform block edges of the picture, except for the following types of edges:
[0465] – edges on the image border,
[0466] – an edge that coincides with the boundary of a sub-picture with loop_filter_cross_subpic_enabled_flag[SubPicIdx] equal to 0,
[0467] – When PPS_loop_filter_cross_virtual_boundaries_disabled_flag is equal to 1, edges that coincide with the virtual boundaries of the picture,
[0468] – edges that coincide with tile boundaries when loop_filter_cross_tiles_enabled_flag is equal to 0,
[0469] – edges that coincide with slice boundaries when loop_filter_cross_slices_enabled_flag is equal to 0,
[0470] – an edge that coincides with the upper or left border of a slice for which slice_deblocking_filter_disabled_flag is equal to 1,
[0471] –slice_deblocking_filter_disabled_flag is equal to 1 for the edges within the slice,
[0472] – edges that do not correspond to the boundaries of the 4×4 sample grid of the luma component,
[0473] – edges that do not correspond to the boundaries of the 8×8 sample grid of the chroma components,
[0474] – edges with intra_bdpcm_luma_flag equal to 1 on either side of an edge within the luma component,
[0475] – edges with intra_bdpcm_chroma_flag equal to 1 on both sides of an edge within a chroma component,
[0476] – Edges of chroma sub-blocks that are not edges of the associated transform unit.
[0477] …
[0478] Deblocking filtering process in one direction
[0479] The inputs to this process are:
[0480] – specifies the variable treeType that currently processes the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA),
[0481] –When treeType is equal to DUAL_TREE_LUMA, the reconstructed picture before removing the block, that is, the array recPicture L ,
[0482] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr ,
[0483] – The variable edgeType specifies whether to filter vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR).
[0484] The output of this process is the modified reconstructed image after deblocking, namely:
[0485] – When treeType is equal to DUAL_TREE_LUMA, the array recPicture L ,
[0486] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr .
[0487] The variables firstCompIdx and lastCompIdx are derived as follows:
[0488] firstCompIdx=(treeType==DUAL_TREE_CHROMA)? 1:0 (8-1010)
[0489] lastCompIdx=(treeType==DUAL_TREE_LUMA||ChromaArrayType==0)? 0:2(8-1011)
[0490] For each codec unit and each codec block of each color component of the codec unit indicated by the color component index cIdx, with the codec block width nCbW, the codec block height nCbH, and the position of the top left sample point of the codec block (xCb, yCb), cIdx ranges from firstCompIdx to lastCompIdx, including first CompIdx and lastCompIdx, when cIdx is equal to 0, or when cIdx is not equal to 0 and edgeType is equal to EDGE_VER and xCb%8 is equal to 0, or when cIdx is not equal to 0 and edgeType is equal to EDGE_HOR and yCb%8 is equal to 0, filter the edges by the following ordered steps:
[0491] 1. The derivation of the variable filterEdgeFlag is as follows:
[0492] – If edgeType is equal to EDGE_VER and one or more of the following conditions are true, filterEdgeFlag is set equal to 0:
[0493] – The left border of the current codec block is the left border of the picture.
[0494] – The left boundary of the current codec block is the left or right boundary of the sub-picture, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] is equal to 0.
[0495] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[0496] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[0497] – The left boundary of the current codec block is one of the vertical virtual boundaries of the picture, and VirtualBoundariesDisabledFlag is equal to 1.
[0498] Otherwise, if edgeType is equal to EDGE_HOR, and one or more of the following conditions are true, then the variable filterEdgeFlag is set equal to 0:
[0499] – The top boundary of the current luma codec block is the top boundary of the picture.
[0500] – The top boundary of the current codec block is the top or bottom boundary of the sub-picture, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] is equal to 0.
[0501] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[0502] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[0503] – The top boundary of the current codec block is one of the horizontal virtual boundaries of the picture, and VirtualBoundariesDisabledFlag is equal to 1.
[0504] – Otherwise, filterEdgeFlag is set equal to 1.
[0505] 2.7 TPM, HMVP, and GEO
[0506] TPM (Triangle Prediction Mode) in VVC divides a block into two triangles with different motion information.
[0507] The HMVP (history-based motion vector prediction) in VVC maintains a motion information table for motion vector prediction. This table is updated after decoding an inter-coded block, but not if the inter-coded block is TPM-coded.
[0508] GEO, proposed in JVET-P0884, is an extension of TPM. Using GEO, a block can be split into two partitions using a straight line, which may or may not be triangular.
[0509] 2.8 ALF, CC-ALF, and Virtual Boundaries
[0510] The ALF (Adaptive Loop Filter) in VVC is applied after the picture is decoded to improve the picture quality.
[0511] The use of virtual boundaries (VB) in VVC makes ALF easier to design in hardware. With VB, ALF is performed in an ALF processing unit bounded by two ALF virtual boundaries.
[0512] The CC-ALF proposed in JVET-P1008 filters the chrominance samples by referring to the information of the luma samples.
[0513] 2.9 SEI of Subgraphs in JVET-P2001-v14
[0514] D.2.8 Sub-picture level information SEI message syntax
[0515]
[0516] D.3.8 Sub-picture level information SEI message semantics
[0517] When testing the conformance of an extracted bitstream containing sub-pictures according to Annex A, the sub-picture level information SEI message contains information about the level of conformance of the sub-pictures in the bitstream.
[0518] When a sub-picture level information SEI message is present in any picture of a CLVS, the sub-picture level information SEI message shall be present in the first picture of the CLVS. Sub-picture level information SEI messages continue from the current picture to the current layer in decoding order until the end of the CLVS. All sub-picture level information SEI messages applicable to the same CLVS shall have the same content.
[0519] sli_seq_parameter_set_id indicates and shall be equal to the sps_seq_parameter_set_id of the SPS referenced by the codec picture associated with the sub-picture level information SEI message.The value of sli_seq_parameter_set_id shall be equal to the value of pps_seq_parameter_set_id in the PPS referenced by the ph_pic_parameter_set_id of the codec picture associated with the sub-picture level information SEI message.
[0520] A bitstream conformance requirement is that, when a sub-picture level information SEI message is present for CLVS, the value of subpic_treated_as_pic_flag[i] shall be equal to 1 for every value of i in the range of 0 to sps_num_subpics_minus1, inclusive.
[0521] num_ref_levels_minus1 plus 1 specifies the number of reference levels signaled for each of the sps_num_subpics_minus1+1 sub-pictures.
[0522] explicit_fraction_present_flag equal to 1 specifies that the syntax element ref_level_fraction_minus1[i] is present. explicit_fraction_present_flag equal to 0 specifies that the syntax element ref_level_fraction_minus1[i] is not present.
[0523] ref_level_idc[i] indicates that each sub-picture conforms to the level specified in Annex A. The bitstream shall not contain values of ref_level_idc other than those specified in Annex A. Other values of ref_level_idc[i] are reserved for future use by ITU-T | ISO / IEC. A requirement for bitstream conformance is that, for any value of k greater than i, the value of ref_level_idc[i] shall be less than or equal to ref_level_idc[k].
[0524] ref_level_fraction_minus1[i][j] plus 1 specifies the fraction of the level constraint associated with ref_level_idc[i], the j-th sub-picture of ref_level_idc[i] conforming to the specification of clause A.4.1.
[0525] The variable SubPicSizeY[j] is set equal to (subpic_width_minus1[j]+1)*(subpic_height_minus1[j]+1).
[0526] When not present, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256*SubPicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])–1.
[0527] The variable RefLevelFraction[i][j] is set equal to ref_level_fraction_minus1[i][j]+1.
[0528] The variables SubPicNumTileCols[j] and SubPicNumTileRows[j] are derived as follows:
[0529]
[0530]
[0531] The variables SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j] are derived as follows:
[0532] SubPicCpbSizeVcl[i][j]=
[0533] Floor(CpbVclFactor*MaxCPB*RefLevelFraction[i][j]÷256)(D.6)
[0534] SubPicCpbSizeNal[i][j]=
[0535] Floor(CpbNalFactor*MaxCPB*RefLevelFraction[i][j]÷256)(D.7)
[0536] where MaxCPB is derived from ref_level_idc[i] as specified in clause A.4.2.
[0537] NOTE 1 - When extracting a sub-picture, the resulting bitstream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j].
[0538] The bitstream conformance requirement is that a bitstream resulting from the extraction of the j-th subpicture and conforming to the profile with general_tier_flag equal to 0 and level equal to ref_level_idc[i] shall obey the following constraints for each bitstream conformance test specified in Annex C, where j ranges from 0 to sps_num_subpics_minus1, inclusive, and i ranges from 0 to num_ref_level_minus1, inclusive:
[0539] –Ceil(256*SubPicSizeY[i]÷RefLevelFraction[i][j]) shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.
[0540] –The value of Ceil(256*(subpic_width_minus1[i]+1)÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0541] –The value of Ceil(256*(subpic_height_minus1[i]+1)÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0542] – The value of SubPicNumTileCols[j] shall be less than or equal to MaxTileCols, and the value of SubPicNumTileRows[j] shall be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.
[0543] For any sub-picture set containing one or more sub-pictures and consisting of multiple sub-pictures in the sub-picture index list SubPicSetIndices and the sub-picture set NumSubPicInSet, the level information of the sub-picture set is derived.
[0544] Variable for the total level score relative to the reference level ref_level_idc[i]
[0545] SubPicSetAccLevelFraction[i] and variables of sub-picture sets
[0546] The derivation of SubPicSetCpbSizeVcl[i][j] and SubPicSetCpbSizeNal[i][j] is as follows:
[0547]
[0548] The value of the sub-picture set sequence level indicator SubPicSetLevelIdc is derived as follows:
[0549]
[0550] The MaxTileCols and MaxTileRows for ref_level_idc[i] are specified in Table A.1.
[0551] A sub-picture set bitstream conforming to a profile with general_tier_flag equal to 0 and level equal to SubPicSetLevelIdc shall obey the following constraint C for each bitstream conformance test specified in Appendix C:
[0552] – For VCL HRD parameters, SubPicSetCpbSizeVcl[i] shall be less than or equal to CpbVclFactor*MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbVclFactor bits.
[0553] – For NAL HRD parameters, SubPicSetCpbSizeVcl[i] shall be less than or equal to CpbNalFactor*MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0554] NOTE 2 - When extracting a sub-picture set, the resulting bitstream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubPicCpbSizeVcl[i][j] and SubPicSetCpbSizeNal[i][j].
[0555] 3. Examples of technical problems solved by the disclosed embodiments
[0556] 1. Some designs in VVC violate sub-image constraints.
[0557] a. The TMVP among the candidates constructed by affine construction can obtain MV in the collocated picture outside the range of the current sub-picture.
[0558] b. When deriving gradients in Bidirectional Optical Flow (BDOF) and Prediction Refinement Optical Flow (PROF), it is necessary to extract two extended rows and two extended columns of integer reference samples. These reference samples may be outside the range of the current sub-picture.
[0559] c. When deriving chroma residual scaling factors in luma-mapped chroma scaling (LMCS), the reconstructed luma samples accessed may be outside the range of the current sub-picture.
[0560] d. When deriving luma intra prediction modes, reference samples for intra prediction, reference samples for CCLM, neighbor block availability for spatial neighbor candidates for merge / AMVP / CIIP / IBC / LMCS, quantization parameters, CABAC initialization, ctxInc derivation using the left and above syntax elements, and ctxInc for the mtt_split_cu_vertical_flag syntax element, neighboring blocks may be outside the range of the current sub-picture. The representation of sub-pictures may result in sub-pictures with incomplete CTUs. The CTU partitioning and CU splitting processes may need to take into account incomplete CTUs.
[0561] 2. The syntax elements signaled related to sub-pictures can be arbitrarily large, which may lead to overflow problems.
[0562] 3. The representation of sub-images may result in non-rectangular sub-images.
[0563] 4. Currently, sub-pictures and sub-picture grids are defined in units of 4 samples. The length of the syntax element depends on the picture height divided by 4. However, since pic_width_in_luma_samples and pic_height_in_luma_samples should currently be integer multiples of Max(8,MinCbSizeY), it may be necessary to define the sub-picture grid in units of 8 samples.
[0564] 5.SPS syntax, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples may need to be restricted to no less than 8.
[0565] 6. Reference picture resampling / scalability and the interaction between sub-pictures are not considered in the current design.
[0566] 7. In temporal filtering, it may be necessary to span samples from different sub-pictures.
[0567] 8. When a stripe is signaled, in some cases the information can be inferred without signaling.
[0568] 9. It is possible that all defined strips may not cover the entire picture or sub-picture.
[0569] 10. The IDs of two sub-images can be the same.
[0570] 11. pic_width_max_in_luma_samples / CtbSizeY may be equal to 0, resulting in meaningless Log2() operation.
[0571] 12. ID in PH is more preferred than that in PPS, but less preferred than that in SPS, which is inconsistent.
[0572] 13. log2_transform_skip_max_size_minus2 in PPS is parsed based on sps_transform_skip_enabled_flag in SPS, resulting in parsing dependencies.
[0573] 14. loop_filter_cross_subpic_enabled_flag for deblocking only considers the current sub-picture, not the adjacent sub-pictures.
[0574] 15. In applications, sub-pictures are designed to provide flexibility so that areas at the same position in a sequence of pictures can be decoded or extracted independently. The area may have some special requirements. For example, it may be a region of interest (ROI) that requires high quality. In another example, it can be used as a track for quickly browsing the video. In yet another example, it can provide a low-resolution, low-complexity, and low-power bitstream that can be fed to end users who are sensitive to complexity. All these applications may require that areas of the sub-picture should be encoded with a configuration different from other parts. However, in the current VVC, there is no mechanism to independently configure sub-pictures.
[0575] 4. Examples of technical solutions
[0576] The following are examples that should be considered as illustrative of the general concepts. These entries should not be construed in a narrow sense. Additionally, these entries may be combined in any way.
[0577] Hereinafter, a temporal filter is used to represent a filter that requires samples from other pictures (e.g., the filter proposed in JCTVC - AI0023).
[0578] Max(x, y) returns the larger one of x and y.
[0579] Min(x, y) yields the smaller one of x and y.
[0580] 1. Assume that the top - left coordinates of the required sub - picture are (xTL, yTL) and the bottom - right coordinates of the required sub - picture are (xBR, yBR). Then, the position (referred to as position RB) in the picture to obtain a temporal MV predictor for generating an affine motion candidate (e.g., a constructed affine merge candidate) must be within the required sub - picture.
[0581] a. In one example, the required sub - picture is the sub - picture covering the current block.
[0582] b. In one example, if the position RB with coordinates (x, y) is outside the required sub - picture, the temporal MV predictor is considered unavailable.
[0583] i. In one example, if x > xBR, then the position RB is outside the required sub - picture.
[0584] ii. In one example, if y > yBR, then the position RB is outside the required sub - picture.
[0585] iii. In one example, if x < xTL, then the position RB is outside the required sub - picture.
[0586] iv. In one example, if y < yTL, then the position RB is outside the required sub - picture.
[0587] c. In one example, if the position RB is outside the required sub - picture, a replacement for RB is utilized.
[0588] i. Alternatively, in addition, the replacement position should be within the required sub - picture.
[0589] d. In one example, the position RB is clipped to within the required sub - picture.
[0590] i. In one example, x is clipped to x = Min(x, xBR).
[0591] ii. In one example, y is clipped to y = Min(y, yBR).
[0592] iii. In one example, x is clipped to x = Max(x, xTL).
[0593] iv. In one example, y is clipped to y = Max(y, yTL).
[0594] e. In one example, the position RB can be the lower right position within the corresponding block of the current block in the juxtaposed picture.
[0595] f. The proposed method can be used for other coding and decoding tools that require accessing motion information from a picture different from the current picture.
[0596] g. In one example, whether to apply the above method (e.g., the position RB must be within the required subpicture (e.g., as required in 1.a and / or 1.b)) can depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index of the subpicture covering the current block.
[0597] 2. Assuming the upper left coordinates of the required subpicture are (xTL, yTL) and the lower right coordinates of the required subpicture are (xBR, yBR), the position (referred to as position S) for obtaining integer samples in a reference not used during the interpolation process must be within the required subpicture.
[0598] a. In one example, the required subpicture is the subpicture covering the current block.
[0599] b. In one example, if the position S with coordinates (x, y) is outside the required subpicture, the reference sample is considered unavailable.
[0600] [[ID=
[0605] i. In one example, x is clipped to x = Min(x, xBR).
[0606] ii. In one example, y is clipped to y = Min(y, yBR).
[0607] iii. In one example, x is clipped to x = Max(x, xTL).
[0608] iv. In one example, y is clipped to y = Max(y, yTL).
[0609] d. In one example, whether the position S must be within the required sub - picture (e.g., as required in 2.a and / or 2.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub - picture index of the sub - picture covering the current block.
[0610] e. In one example, the extracted integer samples are used to generate gradients in BDOF and / or PORF.
[0611] 3. Assume that the upper - left corner coordinates of the required sub - picture are (xTL, yTL) and the lower - right corner coordinates of the required sub - picture are (xBR, yBR). The position (referred to as position R) where the reconstructed luminance sample value is extracted
[0612] can be within the required sub - picture.
[0613] a. In one example, the required sub - picture is the sub - picture covering the current block.
[0614] b. In one example, if the position R with coordinates (x, y) is outside the required sub - picture, the reference samples are considered unavailable.
[0615] i. In one example, if x > xBR, then the position R is outside the required sub - picture.
[0616] ii. In one example, if y > yBR, then the position R is outside the required sub - picture.
[0617] iii. In one example, if x < xTL, then the position R is outside the required sub - picture.
[0618] iv. In one example, if y < yTL, then the position R is outside the required sub - picture.
[0619] c. In one example, the position R is cropped into the desired sub - picture.
[0620] i. In one example, x is cropped to x = Min(x, xBR).
[0621] ii. In one example, y is cropped to y = Min(y, yBR).
[0622] iii. In one example, x is cropped to x = Max(x, xTL).
[0623] iv. In one example, y is cropped to y = Max(y, yTL).
[0624] d. In one example, whether the position R must be within the desired sub - picture (e.g., as required in 4.a and / or 4.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub - picture index of the sub - picture covering the current block.
[0625] e. In one example, the obtained luminance samples are used to derive the scaling factors for the chrominance component(s) in LMCS.
[0626] 4. Assuming that the top - left coordinates of the desired sub - picture are (xTL, yTL) and the bottom - right coordinates of the desired sub - picture are (xBR, yBR), the position (referred to as position N) for picture boundary checking for signaling of BT / TT / QT partitioning, BT / TT / QT depth derivation, and / or CU partitioning flag must be within the desired sub - picture.
[0627] a. In one example, the desired sub - picture is the sub - picture covering the current block.
[0628] b. In one example, if the position N with coordinates (x, y) is outside the desired sub - picture, the reference samples are considered unavailable.
[0629] i. In one example, if x > xBR, the position N is outside the desired sub - picture.
[0630] ii. In one example, if y > yBR, the position N is outside the desired sub - picture.
[0631] iii. In one example, if x < xTL, the position N is outside the desired sub - picture.
[0632] iv. In one example, if y < yTL, the position N is outside the desired sub - picture.
[0633] c. In one example, position N is cropped to the desired sub-picture.
[0634] i. In one example, x is clipped to x=Min(x,xBR).
[0635] ii. In one example, y is clipped to y=Min(y,yBR).
[0636] iii. In one example, x is clipped to x=Max(x, xTL).
[0637] iv. In one example, y is clipped to y=Max(y,yTL).
[0638] d. In one example, whether position N must be in the required sub-picture (e.g., as required in 5.a and / or 5.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / slice header / slice group header. For example, the syntax element may be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub-picture index of the sub-picture covering the current block.
[0639] 5. The History-Based Motion Vector Prediction (HMVP) table can be reset before decoding a new sub-picture in a picture.
[0640] a. In one example, the HMVP table for IBC codec can be reset
[0641] b. In one example, the HMVP table for inter-frame coding and decoding can be reset
[0642] c. In one example, the HMVP table for intra-frame encoding and decoding can be reset
[0643] 6. The sub-picture syntax element may be defined in units of N (eg, N=8, 32, etc.) samples.
[0644] a. In one example, the width of each element of the sub-picture identifier grid is in units of N samples.
[0645] b. In one example, the height of each element of the sub-picture identifier grid is in units of N samples.
[0646] c. In one example, N is set to the width and / or height of the CTU.
[0647] 7. The syntax elements of picture width and picture height may be restricted to be no less than K (K>=8).
[0648] a. In one example, the image width may need to be limited to no less than 8.
[0649] b. In one example, the image height may need to be limited to no less than 8.
[0650] 8. The conforming bitstream shall satisfy that sub-picture coding and adaptive resolution conversion (ARC) / dynamic resolution conversion (DRC) / reference picture resampling (RPR) are not allowed to be enabled for a video unit (eg, sequence).
[0651] a. In one example, signaling to enable sub-picture codec may not be allowed.
[0652] Under the conditions of ARC / DRC / RPR.
[0653] i. In one example, when sub-pictures are enabled, such as subpics_present_flag is equal to 1, pic_width_in_luma_samples is equal to max_width_in_luma_sample for all pictures for which this SPS is valid.
[0654] b. Alternatively, both sub-picture coding and ARC / DRC / RPR may be enabled for one video unit (eg, sequence).
[0655] i. In one example, the conforming bitstream will satisfy that the downsampled sub-picture due to ARC / DRC / RPR will still be of the form of K CTUs in width and M CTUs in height, where K and M are both integers.
[0656] ii. In one example, the conforming bitstream will satisfy the requirement for sub-pictures that are not located at picture boundaries (e.g., right and / or bottom boundaries) since the downsampled sub-pictures of ARC / DRC / RPR will still be in the form of K CTUs in width and M CTUs in height, where K and M are both integers.
[0657] iii. In one example, the CTU size can be adaptively changed based on the picture resolution.
[0658] 1) In one example, the maximum CTU size can be signaled in the SPS. For each picture with a lower resolution, the CTU size can be changed accordingly based on the reduced resolution.
[0659] 2) In one example, the CTU size may be signaled in the SPS and PPS and / or sub-picture level.
[0660] 9. The syntax elements subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 may be constrained.
[0661] a. In one example, subpic_grid_col_width_minus1 must not be greater than (or must be less than) T1.
[0662] b. In one example, subpic_grid_row_height_minus1 must not be greater than (or must be less than) T2.
[0663] c. In one example, in a conforming bitstream, subpic_grid_col_width_minus1 and / or subpic_grid_row_height_minus1 must follow constraints such as item 3.a or 3.b
[0664] d. In one example, T1 in 3.a and / or T2 in 3.b may depend on the profile / level / tier of the video codec standard.
[0665] e. In one example, T1 in 3.a may depend on the picture width.
[0666] i. For example, T1 is equal to pic_width_max_in_luma_samples / 4 or pic_width_max_in_luma_samples / 4+Off. Off can be 1, 2, -1, -2, etc.
[0667] f. In one example, T2 in 3.b may depend on the picture width.
[0668] i. For example, T2 is equal to pic_height_max_in_luma_samples / 4 or pic_height_max_in_luma_samples / 4-1+Off. Off can be 1, 2, -1, -2, etc.
[0669] 10. The constraint requires that the boundary between two sub-pictures must be the boundary between two CTUs.
[0670] a. In other words, a CTU cannot be covered by more than one sub-picture.
[0671] b. In one example, the unit of subpic_grid_col_width_minus1 can be CTU width (e.g., 32, 64, 128) instead of 4 as in VVC. The sub-picture grid width should be (subpic_grid_col_width_minus1+1)*CTU width.
[0672] c. In one example, the unit of subpic_grid_col_height_minus1 can be CTU height (e.g., 32, 64, 128) instead of 4 as in VVC. The sub-picture grid height should be (subpic_grid_col_height_minus1+1)*CTU height.
[0673] d. In one example, in a conforming bitstream, if the sub-picture scheme is applied, the constraints must be met.
[0674] 11. Constraints require that the shape of the sub-image must be rectangular.
[0675] a. In one example, in a conforming bitstream, if the sub-picture scheme is applied, the constraints must be met.
[0676] b. A sub-picture may only contain rectangular slices. For example, in a conforming bitstream, if the sub-picture scheme is applied, the constraints must be met.
[0677] 12. The constraint requires that two sub-images cannot overlap.
[0678] a. In one example, in a conforming bitstream, if the sub-picture scheme is applied, the constraints must be met.
[0679] b. Alternatively, the two sub-pictures may overlap each other.
[0680] 13. The constraint requires that any position in the image must be covered by one and only one sub-image.
[0681] a. In one example, in a conforming bitstream, if the sub-picture scheme is applied, the constraints must be met.
[0682] b. Alternatively, a sample may not belong to any sub-picture.
[0683] c. Alternatively, a sample may belong to more than one sub-picture.
[0684] 14. It is possible to constrain the positions and / or sizes of sub-pictures defined in the SPS mapped to each resolution present in the same sequence to obey the above constraints.
[0685] a. In one example, the width and height of a sub-picture defined in an SPS that maps to a resolution present in the same sequence should be an integer multiple of N (e.g., 8, 16, 32) luma samples.
[0686] b. In one example, sub-pictures may be defined for certain layers and mapped to other layers.
[0687] i. For example, a sub-picture may be defined for the layer with the highest resolution in the sequence.
[0688] ii. For example, a sub-picture may be defined for the layer with the lowest resolution in the sequence.
[0689] iii. It may be signaled in the SPS / VPS / PPS / slice header for which layer the sub-picture is defined.
[0690] c. In one example, when both sub-pictures and different resolutions are applied, all resolutions (eg, width and / or height) may be integer multiples of a given resolution.
[0691] d. In one example, the width and / or height of the sub-picture defined in the SPS can be an integer multiple of the CTU size (e.g., M).
[0692] e. Alternatively, sub-pictures and different resolutions in a sequence may not be allowed simultaneously.
[0693] 15. Sub-images can be applied only to certain layers
[0694] a. In one example, the sub-pictures defined in the SPS may only be applied to the layer with the highest resolution in the sequence.
[0695] b. In one example, the sub-picture defined in the SPS may only be applied to the layer with the lowest temporal id in the sequence.
[0696] c. One or more syntax elements in the SPS / VPS / PPS may indicate to which layer(s) a sub-picture may apply.
[0697] d. One or more syntax elements in the SPS / VPS / PPS may indicate to which layer(s) the sub-picture cannot be applied.
[0698] 16. In one example, the location and / or dimensions of the sub-picture may be signaled without using subpic_grid_idx.
[0699] a. In one example, the top left position of the sub-picture may be signaled.
[0700] b. In one example, the bottom right position of the sub-picture may be signaled.
[0701] c. In one example, the width of the sub-picture may be signaled.
[0702] d. In one example, the height of the sub-picture may be signaled.
[0703] 17. For temporal filters, when performing temporal filtering of samples, only samples within the same sub-picture to which the current sample belongs can be used. The required samples may be in the same picture to which the current sample belongs, or in other pictures.
[0704] 18. In one example, whether and / or how to apply a partitioning method (such as QT, horizontal BT, vertical BT, horizontal TT, vertical TT, or no partitioning, etc.) may depend on whether the current block (or partition) crosses one or more boundaries of a sub-picture.
[0705] a. In one example, when picture boundaries are replaced by sub-picture boundaries, the picture boundary processing method used for segmentation in VVC can also be applied.
[0706] b. In one example, whether to parse a syntax element (e.g., a flag) indicating a partitioning method (such as QT, horizontal BT, vertical BT, horizontal TT, vertical TT, or no partitioning, etc.) may depend on whether the current block (or partition) crosses one or more boundaries of a sub-picture.
[0707] 19. Instead of dividing a picture into multiple sub-pictures where each sub-picture is independently coded, it is proposed to divide the picture into at least two sets of sub-regions, the first set including multiple sub-pictures and the second set including all remaining samples.
[0708] a. In one example, the samples in the second set are not in any sub-picture.
[0709] b. Alternatively, in addition, the second set may be encoded / decoded based on the information of the first set.
[0710] c. In one example, a default value may be used to mark whether a sample point / MxK sub-region belongs to the second set.
[0711] i. In one example, the default value may be set equal to (max_subpics_minus1+K), where K is an integer greater than 1.
[0712] ii. A default value may be assigned to subpic_grid_idx[i][j] to indicate that the grid belongs to the second set.
[0713] 20. It is proposed that the syntax element subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0714] a. For example, the constraint requires that in a conforming bitstream, subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0715] b. For example, the codeword for encoding and decoding subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0716] 21. It is proposed that any integer from 0 to max_subpics_minus1 must be equal to at least one subpic_grid_idx[i][j].
[0717] 22. Before decoding a new sub-picture in a picture, the IBC virtual buffer can be reset.
[0718] a. In one example, all samples in the IBC virtual buffer may be reset to -1.
[0719] 23. Before decoding a new sub-picture in a picture, the palette entry list can be reset.
[0720] a. In one example, before decoding a new sub-picture in a picture, the PredictorPaletteSize may be set equal to 0.
[0721] 24. Whether to signal the information of the stripes (eg, the number of stripes and / or the range of the stripes) may depend on the number of slices and / or the number of bricks.
[0722] a. In one example, if the number of tiles in a picture is 1, num_slices_in_pic_minus1 is not signaled and is inferred to be 0.
[0723] b. In one example, if the number of tiles in a picture is 1, information of the slices (eg, the number of slices and / or the range of the slices) may not be signaled.
[0724] c. In one example, if the number of bricks in a picture is 1, the number of slices can be inferred to be 1. And the slice covers the entire picture. In one example, if the number of bricks in a picture is 1, single_brick_per_slice_flag is not signaled and is inferred to be 1.
[0725] i. Alternatively, if the number of bricks in the picture is 1, then single_brick_per_slice_flag must be 1.
[0726] d. The exemplary grammar design is as follows:
[0727]
[0728]
[0729] 25. Whether slice_address is signaled may be independent of whether the slice is signaled as rectangular (eg, whether rect_slice_flag is equal to 0 or 1).
[0730] a. The exemplary grammar design is as follows:
[0731]
[0732] 26. When slices are signaled as rectangles, whether slice_address is signaled may depend on the number of slices.
[0733]
[0734] 27. Whether num_bricks_in_slice_minus1 is signaled may depend on slice_address and / or the number of bricks in the picture.
[0735] a. The exemplary grammar design is as follows:
[0736]
[0737] 28. Whether the loop_filter_across_bricks_enabled_flag is signaled may depend on the number of slices and / or the number of bricks.
[0738] a. In one example, if the number of bricks is less than 2, loop_filter_across_bricks_enabled_flag is not signaled.
[0739] b. The exemplary grammar design is as follows:
[0740]
[0741] 29. The bitstream consistency requirement is that all slices of a picture must cover the entire picture.
[0742] a. This requirement must be met when the slice is signaled as rectangular (eg, rect_slice_flag is equal to 1).
[0743] 30. The bitstream consistency requirement is that all slices of a sub-picture must cover the entire sub-picture.
[0744] a. This requirement must be met when the slice is signaled as rectangular (eg, rect_slice_flag is equal to 1).
[0745] 31. A bitstream conformance requirement is that a slice cannot overlap with more than one sub-picture.
[0746] 32. A bitstream conformance requirement is that a slice cannot overlap with more than one sub-picture.
[0747] 33. A bitstream conformance requirement is that a tile cannot overlap with more than one sub-picture.
[0748] In the following discussion, a basic unit block (BUB) of dimension CW×CH is a rectangular area. For example, a BUB may be a codec tree block (CTB).
[0749] 34. In one example, the number of sub-pictures (denoted as N) can be signaled.
[0750] a. If sub-pictures are used (eg, subpics_present_flag is equal to 1), then at least two sub-pictures in a picture may be required on a conforming bitstream.
[0751] b. Alternatively, N minus d (ie, Nd) may be signaled, where d is an integer such as 0, 1 or 2.
[0752] c. For example, Nd may be encoded and decoded using a fixed-length codec, such as u(x).
[0753] i. In one example, x can be a fixed number, such as 8.
[0754] ii. In one example, x or x-dx may be signaled before signaling Nd, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the conforming bitstream.
[0755] iii. In one example, x can be exported on the fly.
[0756] 1) For example, x can be derived as a function of the total number of BUBs in the picture (denoted as M). For example, x = Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0757] 2) M can be derived as M = Ceiling(W / CW)×Ceiling(H / CH), where W and H represent the width and height of the image, and CW and CH represent the width and height of the BUB.
[0758] d. For example, Nd can be encoded and decoded using a unary codec or a truncated unary codec.
[0759] e. In one example, the maximum allowed value of Nd may be a fixed number.
[0760] Alternatively, the maximum allowed value of Nd can be derived as a function of the total number of BUBs in the picture (denoted as M). For example, x = Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0761] 35. In one example, a sub-picture may be signaled by an indication of one or more of its selected position (eg, top left / top right / bottom left / bottom right position) and / or its width and / or its height.
[0762] a. In one example, the top left position of a sub-picture may be signaled at a granularity of a basic unit block (BUB) of dimension CW×CH.
[0763] i. For example, the BUB-wise column index (denoted as Col) of the top left BUB of the sub-picture may be signaled.
[0764] 1) For example, Col-d can be signaled, where d is an integer such as 0, 1 or 2.
[0765] a) Alternatively, d may be equal to Col of the previously coded sub-picture plus d1, where d1 is an integer such as -1, 0, or 1.
[0766] b) The symbol of Col-d may be signaled.
[0767] ii. For example, the BUB-wise row index (denoted as Row) of the top left BUB of the sub-picture may be signaled.
[0768] 1) For example, Row-d can be signaled, where d is an integer such as 0, 1 or 2.
[0769] a) Alternatively, d may be equal to the Row of the previously coded sub-picture plus d1, where d1 is an integer such as -1, 0, or 1.
[0770] b) The symbol of Row-d may be signaled.
[0771] iii. The row / column index (labeled as Row) mentioned above can be represented by a codec tree block (CTB) unit. For example, the x or y coordinate relative to the upper left position of the picture can be divided by the CTB size and signaled.
[0772] iv. In one example, whether to signal the position of a sub-picture may depend on the sub-picture index.
[0773] 1) In one example, for the first sub-picture within a picture, the top left position may not be signaled.
[0774] a) Alternatively, furthermore, the top left position may be inferred, for example to (0, 0).
[0775] 2) In one example, for the last sub-picture within a picture, the top left position may not be signaled.
[0776] a) The top left position can be inferred from the information of the previously signaled sub-picture.
[0777] b. In one example, the indication of the width / height / selected position of the sub-picture may be signaled using truncated unary / truncated binary / unary / fixed length / Kth EG codec (eg, K=0, 1, 2, 3).
[0778] c. In one example, the width of a sub-picture may be signaled with a granularity of a BUB of dimension CW×CH.
[0779] i. For example, the number of columns of BUBs in a sub-picture (denoted as W) may be signaled.
[0780] ii. For example, Wd may be signaled, where d is an integer such as 0, 1 or 2.
[0781] 1) Alternatively, d may be equal to W of the previously coded sub-picture plus d1, where d1 is an integer such as -1, 0, or 1.
[0782] 2) The sign of Wd may be signaled.
[0783] d. In one example, the height of the sub-picture can be signaled with the granularity of a BUB of dimension CW×CH.
[0784] i. For example, the number of rows of BUBs in a sub-picture (denoted as H) may be signaled.
[0785] ii. For example, Hd may be signaled, where d is an integer such as 0, 1 or 2.
[0786] 1) Alternatively, d may be equal to H of the previously coded sub-picture plus d1, where d1 is an integer such as -1, 0, or 1.
[0787] 2) The symbol of Hd can be signaled.
[0788] e. In one example, Col-d may be encoded and decoded using a fixed-length codec, such as u(x).
[0789] i. In one example, x can be a fixed number, such as 8.
[0790] ii. In one example, x or x-dx may be signaled before Col-d is signaled, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the conforming bitstream.
[0791] iii. In one example, x can be exported on the fly.
[0792] 1) For example, x can be derived as a function of the total number of BUB columns in the picture (denoted as M). For example, x=Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0793] 2) M can be derived as M=Ceiling(W / CW), where W represents the width of the picture and CW represents the width of the BUB.
[0794] f. In one example, fixed-length encoding and decoding can be used to encode and decode Row-d, such as u(x).
[0795] i. In one example, x can be a fixed number, such as 8.
[0796] ii. In one example, x or x-dx may be signaled before signaling Row-d, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the conforming bitstream.
[0797] iii. In one example, x can be exported on the fly.
[0798] 1) For example, x can be derived as a function of the total number of BUB rows in the picture (denoted as M). For example, x=Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0799] 2) M can be derived as M=Ceiling(H / CH), where H represents the height of the picture and CH represents the height of the BUB.
[0800] g. In one example, Wd can be encoded and decoded using a fixed-length codec, such as u(x).
[0801] i. In one example, x can be a fixed number, such as 8.
[0802] ii. In one example, x or x-dx may be signaled before signaling Wd, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the conforming bitstream.
[0803] iii. In one example, x can be exported on the fly.
[0804] 1) For example, x can be derived as a function of the total number of BUB columns in the picture (denoted as M). For example, x=Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0805] 2) M can be derived as M=Ceiling(W / CW), where W represents the width of the picture and CW represents the width of the BUB.
[0806] h. In one example, Hd can be encoded and decoded using a fixed-length codec, such as u(x).
[0807] i. In one example, x can be a fixed number, such as 8.
[0808] ii. In one example, x or x-dx may be signaled before signaling Hd, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the conforming bitstream.
[0809] iii. In one example, x can be exported on the fly.
[0810] 1) For example, x can be derived as a function of the total number of BUB rows in the picture (denoted as M). For example, x=Ceil(log2(M+d0))+d1, where d0 and d1 are two integers such as -2, -1, 0, 1, 2, etc.
[0811] 2) M can be derived as M=Ceiling(H / CH), where H represents the height of the picture and CH represents the height of the BUB.
[0812] i. Col-d and / or Row-d may be signaled for all sub-pictures.
[0813] i. Alternatively, Col-d and / or Row-d may not be signaled for all sub-pictures.
[0814] 1) If the number of sub-pictures is less than 2 (equal to 1), Col-d and / or Row-d may not be signaled.
[0815] 2) For example, for the first sub-picture, Col-d and / or Row-d may not be signaled (eg, sub-picture index (or sub-picture ID) equals 0) a) When they are not signaled, they may be inferred to be 0.
[0816] 3) For example, for the last sub-picture (eg, the sub-picture index (or sub-picture ID) is equal to NumSubPics-1), Col-d and / or Row-d may not be signaled.
[0817] a) When they are not signaled, they can be inferred from the signaled position and dimensions of the sub-picture.
[0818] j. Wd and / or Hd may be signaled for all sub-pictures.
[0819] i. Alternatively, Wd and / or Hd may not be signaled for all sub-pictures.
[0820] 1) If the number of sub-pictures is less than 2 (equal to 1), Wd and / or Hd may not be signaled.
[0821] 2) For example, for the last sub-picture (eg, the sub-picture index (or sub-picture ID) is equal to NumSubPics-1), Wd and / or Hd may not be signaled.
[0822] a) When they are not signaled, they can be inferred from the signaled position and dimensions of the sub-picture.
[0823] k. In the above bullet points, a BUB may be a codec tree block (CTB).
[0824] 36. In one example, the information of the sub-picture should be signaled after the information of the CTB size (eg, log2_ctu_size_minus5) has been signaled.
[0825] 37. subpic_treated_as_pic_flag[i] may not be signaled for each sub-picture. Instead, one subpic_treated_as_pic_flag is signaled for all sub-pictures to control whether the sub-picture is treated as a picture.
[0826] 38. The loop_filter_across_subpic_enabled_flag may not be signaled for each sub-picture
[0827] [i] Instead, for all sub-pictures, a loop_filter_across_subpic_enabled_flag is signaled to control whether the loop filter can be applied across sub-pictures.
[0828] 39. subpic_treated_as_pic_flag[i] (subpic_treated_as_pic_flag) and / or loop_filter_across_subpic_enabled_flag[i] (loop_filter_across_subpic_enabled_flag) may be conditionally signaled.
[0829] a. In one example, if the number of sub-pictures is less than 2 (equal to 1), subpic_treated_as_pic_flag[i] and / or loop_filter_across_subpic_enabled_flag[i] may not be signaled.
[0830] 40. When using sub-pictures, RPR can be applied.
[0831] a. In one example, when sub-pictures are used, the scaling ratios in the RPR may be constrained to a limited set, such as {1:1, 1:2 and / or 2:1}, or {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1}, {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1, 1:8 and / or 8:1}.
[0832] b. In one example, if the resolutions of picture A and picture B are different, the CTB size of picture A and the CTB size of picture B may be different.
[0833] c. In one example, suppose a sub-image SA of dimensions SAW×SAH is in image A, a sub-image SB of dimensions SBW×SBH is in image B, SA corresponds to SB, and the scaling ratios between images A and B in the horizontal and vertical directions are Rw and Rh, then
[0834] i.SAW / SBW or SBW / SAW should be equal to Rw.
[0835] ii. SAH / SBH or SBH / SAH should be equal to Rh.
[0836] 41. When sub-pictures are used (eg, sub_pics_present_flag is true), the sub-picture index (or sub-picture ID) may be signaled in the slice header, and the slice address is interpreted as an address in the sub-picture rather than an address in the entire picture.
[0837] 42. If the first sub-picture and the second sub-picture are not the same sub-picture, the sub-picture ID of the first sub-picture must be different from the sub-picture ID of the second sub-picture.
[0838] a. In one example, in a conforming bitstream, if i is not equal to j, then sps_subpic_id[i] must not be equal to sps_subpic_id[j]
[0839] b. In one example, in a conforming bitstream, if i is not equal to j, then pps_subpic_id[i] must not be equal to pps_subpic_id[j].
[0840] c. In one example, in a conforming bitstream, if i is not equal to j, then ph_subpic_id[i] must not be equal to ph_subpic_id[j].
[0841] d. In one example, in a conforming bitstream, if i is not equal to j, then SubpicIdList[i] must not be equal to SubpicIdList[j]
[0842] e. In one example, the difference expressed as D[i] may be signaled, where D[i] is equal to X_subpic_id[i]-X_subpic_id[iP].
[0843] i. For example, X can be sps, pps, or ph.
[0844] ii. For example, P is equal to 1.
[0845] iii. For example, i>P.
[0846] iv. For example, D[i] must be greater than 0.
[0847] v. For example, D[i]-1 can be signaled.
[0848] 43. It is proposed that the length of the syntax element specifying the horizontal or vertical position of the top left CTU (e.g., subpic_ctu_top_left_x or subpic_ctu_top_left_y) can be derived as Ceil(Log2(SS)) bits, where SS must be greater than 0.
[0849] a. In one example, when the syntax element specifies the horizontal position of the top left CTU (eg, subpic_ctu_top_left_x), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0850] b. In one example, when the syntax element specifies the vertical position of the top left CTU (eg, subpic_ctu_top_left_y), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0851] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0852] 44. It is proposed that the length of the syntax element specifying the horizontal or vertical position of the top left CTU of the sub-picture (e.g., subpic_ctu_top_left_x or subpic_ctu_top_left_y) can be derived as Ceil(Log2(SS)) bits, where SS must be greater than 0.
[0853] a. In one example, when the syntax element specifies the horizontal position of the top left CTU of the sub-picture (eg, subpic_ctu_top_left_x), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0854] b. In one example, when the syntax element specifies the vertical position of the top left CTU of the sub-picture (eg, subpic_ctu_top_left_y), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0855] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0856] 45. It is proposed that the default value of the syntax element specifying the width or height of the sub-picture (e.g., subpic_width_minus1 or subpic_height_minus1) (which may be supplemented by an offset P such as 1) can be derived as Ceil(Log2(SS))-P, where SS must be greater than 0.
[0857] a. In one example, when the syntax element specifies a default width of the sub-picture (eg, subpic_width_minus1) (an offset P may be added), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0858] b. In one example, when the syntax element specifies a default height of the sub-picture (eg, subpic_height_minus1) (an offset P may be added), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0859] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0860] 46. If it is determined that the sub-picture ID information should be signaled, it is proposed that the sub-picture ID information should be signaled in at least one of the SPS, PPS, and picture header.
[0861] a. In one example, if sps_subpic_id_present_flag is equal to 1, then at least one of sps_subpic_id_signalling_present_flag, pps_subpic_id_signalling_present_flag, and ph_subpic_id_signalling_present_flag shall be equal to 1 in the conforming bitstream.
[0862] 47. It is proposed that if the information of the sub-picture ID is not signaled in any of the SPS, PPS and picture header, but it is determined that the information should be signaled, a default ID should be assigned.
[0863] a. In one example, if ps_subpic_id_signalling_present_flag, pps_subpic_id_signalling_present_flag, and ph_subpic_id_signalling_present_flag are all equal to 0, and sps_subpic_id_present_flag is equal to 1, then SubpicIdList[i] should be set equal to i+P, where P is an offset such as 0. An exemplary description is as follows:
[0864]
[0865] 48. It is proposed that if the information of the sub-picture ID is signaled in the corresponding PPS, it is not signaled in the picture header.
[0866] a. The exemplary grammar design is as follows,
[0867]
[0868]
[0869] b. In one example, if the sub-picture ID is signaled in the SPS, the sub-picture ID is set according to the information of the sub-picture ID signaled in the SPS; otherwise, if the sub-picture ID is signaled in the PPS, the sub-picture ID is set according to the information of the sub-picture ID signaled in the PPS; otherwise, if the sub-picture ID is signaled in the picture header, the sub-picture ID is set according to the information of the sub-picture ID signaled in the picture header. An exemplary description is as follows,
[0870]
[0871] c. In one example, if the sub-picture ID is signaled in the picture header, the sub-picture ID is set according to the information of the sub-picture ID signaled in the picture header; otherwise, if the sub-picture ID is signaled in the PPS, the sub-picture ID is set according to the information of the sub-picture ID signaled in the PPS; otherwise, if the sub-picture ID is signaled in the SPS, the sub-picture ID is set according to the information of the sub-picture ID signaled in the SPS. An exemplary description is as follows,
[0872]
[0873]
[0874] 49. It is proposed that the deblocking process on an edge E should be dependent on determining whether loop filtering is enabled across the sub-picture boundary on both sides of the edge (denoted as P-side and Q-side) (e.g., determined by loop_filter_across_subpic_enabled_flag). The P-side refers to the side in the current block, while the Q-side refers to the side in the adjacent block, which may belong to different sub-pictures. In the following discussion, it is assumed that the P-side and the Q-side belong to two different sub-pictures.
[0875] loop_filter_across_subpic_enabled_flag[P]=0 / 1 means that loop filtering is not allowed / enabled on the sub-picture boundary including the P-side sub-picture.
[0876] loop_filter_across_subpic_enabled_flag[Q]=0 / 1 indicates that loop filtering is not allowed / enabled on the sub-picture boundary including the Q-side sub-picture.
[0877] a. In one example, if loop_filter_across_subpic_enabled_flag[P] is equal to 0 or loop_filter_across_subpic_enabled_flag[Q] is equal to 0, E is not filtered.
[0878] b. In one example, if loop_filter_across_subpic_enabled_flag[P] is equal to 0 and loop_filter_across_subpic_enabled_flag[Q] is equal to 0, E is not filtered.
[0879] c. In one example, whether to filter the two sides of E is controlled separately.
[0880] i. For example, if and only if loop_filter_across_subpic_enabled_flag[P] is equal to 1, the P side of E is filtered.
[0881] ii. For example, if and only if loop_filter_across_subpic_enabled_flag[Q] is equal to 1, the Q side of E is filtered.
[0882] 50. It is proposed that the signaling / parsing of the syntax element SE (specifying the maximum block size for transform skipping) in the PPS (such as log2_transform_skip_max_size_minus2) should be decoupled from any syntax element in the SPS (such as sps_transform_skip_enabled_flag).
[0883] a. Example syntax changes are as follows:
[0884]
[0885]
[0886] b. Alternatively, the SE may be signaled in the SPS, such as:
[0887]
[0888] c. Alternatively, the SE may be signaled in the picture header, such as:
[0889]
[0890] 51. Whether and / or how the HMVP table (or named list / storage / mapping etc.) is updated after decoding the first block may depend on whether the first block was decoded using GEO codec.
[0891] a. In one example, if the first block is decoded using the GEO codec, the HMVP table may not be updated after decoding the first block.
[0892] b. In one example, if the first block is decoded using the GEO codec, the HMVP table may be updated after decoding the first block.
[0893] i. In one example, the HMVP table may be updated with motion information of one partition divided by GEO.
[0894] ii. In one example, the HMVP table may be updated using motion information of multiple partitions into which the GEO is divided.
[0895] 52. In CC-ALF, luma samples outside the current processing unit (e.g., an ALF processing unit bounded by two ALF virtual boundaries) are excluded from filtering the chroma samples in the corresponding processing unit.
[0896] a. Filling luma samples outside the current processing unit can be used to filter the chroma samples in the corresponding processing unit.
[0897] i. Any filling method disclosed herein can be used to fill luma samples.
[0898] b. Alternatively, luma samples outside the current processing unit can be used to filter chroma samples in the corresponding processing unit.
[0899] Signaling of sub-picture level parameters
[0900] 53. It is proposed that a parameter set that controls the sub-picture encoding and decoding behavior can be signaled in association with the sub-picture. In other words, for each sub-picture, a parameter set can be signaled. The parameter set may include:
[0901] a. For inter and / or intra slices / pictures, the quantization parameter (QP) or QP increment of the luma component in the sub-picture.
[0902] b. For inter and / or intra slices / pictures, the quantization parameter (QP) or QP increment for the chroma components in the sub-picture.
[0903] c. Reference picture list management information.
[0904] d. Inter-frame and / or intra-frame slice / picture CTU size.
[0905] e. Minimum CU size for inter and / or intra slices / pictures.
[0906] f. Maximum TU size of inter-frame and / or intra-frame slices / pictures.
[0907] g. Maximum / minimum quadtree (QT) partition size for inter-frame and / or intra-frame slices / pictures.
[0908] h. Maximum / minimum quadtree (QT) partitioning depth for inter-frame and / or intra-frame slices / pictures.
[0909] i. Maximum / minimum binary tree (BT) partition size for inter and / or intra slices / pictures.
[0910] j. Inter-frame and / or intra-frame slice / picture maximum / minimum binary tree (BT) partitioning depth.
[0911] k. Maximum / minimum ternary tree (TT) partition size for inter-frame and / or intra-frame slices / pictures.
[0912] l. Maximum / minimum ternary tree (TT) partition depth for inter-frame and / or intra-frame slices / pictures.
[0913] m. Maximum / minimum multi-tree (MTT) partition size for inter-frame and / or intra-frame slices / pictures.
[0914] n. Maximum / minimum multi-tree (MTT) partitioning depth for inter-frame and / or intra-frame slices / pictures.
[0915] o. Codec control tools (including on / off control and / or setting control), including: (see JVET-P2001-v14 for abbreviations).
[0916] i. Weighted prediction
[0917] ii.SAO
[0918] iii.ALF
[0919] iv. Transformation skip
[0920] v.BDPCM
[0921] vi. Joint Cb-Cr Residual Codec (JCCR)
[0922] vii. Reference surround
[0923] viii.TMVP
[0924] ix.sbTMVP
[0925] x.AMVR
[0926] xi.BDOF
[0927] xii.SMVD
[0928] xiii.DMVR
[0929] xiv.MMVD
[0930] xv.ISP
[0931] xvi.MRL
[0932] xvii.MIP
[0933] xviii.CCLM
[0934] xix.CCLM Collocation Chroma Control
[0935] xx.MTS for intra-frame and / or inter-frame
[0936] xxi.MTS for inter-frame
[0937] xxii.SBT
[0938] xxiii. Maximum size of SBT
[0939] xxiv. Affine
[0940] xxv. Affine type
[0941] xxvi. Color Palette
[0942] xxvii.BCW
[0943] xxviii.IBC
[0944] xxix.CIIP
[0945] xxx.Triangle-based motion compensation
[0946] xxxi.LMCS
[0947] p. Any other parameters that have the same meaning as those in the VPS / SPS / PPS / picture header / slice header, but control sub-pictures.
[0948] 54. A flag may be signaled first to indicate whether all sub-pictures share the same parameters.
[0949] a. Alternatively, furthermore, if the parameters are shared, there is no need to signal multiple parameter sets for different sub-pictures.
[0950] b. Alternatively, further signaling of multiple parameter sets for different sub-pictures is required if the parameters are not shared.
[0951] 55. Parametric prediction coding and decoding between different sub-pictures can be applied.
[0952] a. In one example, the difference between two values of the same syntax element of two sub-pictures can be encoded and decoded.
[0953] 56. A default parameter set may be signaled first, and then the differences from the default parameter set may be further signaled.
[0954] a. Alternatively, in addition, a flag may be signaled first to indicate whether the parameter sets of all sub-pictures are the same as the parameter sets in the default set.
[0955] 57. In one example, the parameter set that controls the encoding and decoding behavior of the sub-picture can be signaled in the SPS or PPS or picture header.
[0956] a. Alternatively, the parameter set controlling the sub-picture encoding and decoding behavior may be signaled in a SEI message (eg, the sub-picture level information SEI message defined in JVET-P2001-v14) or a VUI message.
[0957] 58. In this example, the parameter set that controls the encoding and decoding behavior of the sub-picture can be signaled in association with the sub-picture ID.
[0958] 59. In one example, a video unit different from VPS / SPS / PPS / picture header / slice header (called SPPS, sub-picture parameter set) can be signaled, which includes a parameter set that controls the encoding and decoding behavior of the sub-picture.
[0959] a. In one example, the SPPS_index associated with the SPPS is signaled.
[0960] b. In one example, SPPS_index is signaled for a sub-picture to indicate the SPPS associated with the sub-picture.
[0961] 60. In one example, a first control parameter in a parameter set that controls the codec behavior of a sub-picture may override or be overridden by a second control parameter in the parameter set, but control the same codec behavior. For example, an on / off control flag for a codec tool such as BDOF in a parameter set for a sub-picture may override or be overridden by an on / off control flag for a codec tool outside the parameter set.
[0962] a. The second control parameter outside of this parameter set can be in the VPS / SPS / PPS / picture header / slice header.
[0963] 61. When any of the above examples apply, syntax elements associated with a slice / tile / sub-picture depend on parameters associated with the sub-picture containing the current slice, rather than on parameters associated with the picture / sequence.
[0964] 62. Constraint In a conforming bitstream, a first control parameter in a parameter set that controls the codec behavior of a sub-picture must be the same as a second control parameter outside the parameter set, but controlling the same codec behavior.
[0965] 63. In one example, a first flag is signaled in the SPS, one flag per sub-picture, and the first flag specifies whether a general_constraint_info() syntax structure is signaled for the sub-picture associated with the first flag. When present for a sub-picture, the general_constraint_info() syntax structure indicates that no tools apply to the sub-picture on the CLVS.
[0966] a. Alternatively, a general_constraint_info() syntax structure is signaled for each sub-picture.
[0967] b. Alternatively, the second flag is signaled in the SPS, only once, and specifies whether the first flag is present or not for each sub-picture in the SPS.
[0968] 64. In one example, an SEI message or a VUI parameter is specified to indicate that certain codec tools are not applied or are applied in a specific manner to a set of one or more sub-pictures in a CLVS (i.e., a codec slice of a sub-picture set), so that when the sub-picture set is extracted and decoded (e.g., decoded by a mobile device), the decoding complexity is relatively low, and therefore the power consumption of the decoding is relatively low.
[0969] a. Alternatively, the same information can be signaled in the DPS, VPS, SPS or a separate NAL unit.
[0970] 5. Examples
[0971] In the following examples, text deleted from the VVC specification is enclosed in bold double brackets, for example, [[a]] indicates that "a" has been deleted.
[0972] 5.1 Example 1: Affine-Constructed Merge Candidate Sub-Image Constraints (Solution 1)
[0973] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[0974] 8.5.5.6 Derivation of Affine Control Point Motion Vector Merge Candidates for Construction
[0975] The inputs to this process are:
[0976] –Specify the luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,
[0977] – Two variables cbWidth and cbHeight specifying the width and height of the current brightness codec block,
[0978] – Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3,
[0979] – Sample point locations (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2), and (xNbB3, yNbB3).
[0980] The output of this process is:
[0981] – Availability flag availableFlagConstK of the constructed affine control point motion vector merge candidate, where K = 1..6,
[0982] – reference index refIdxLXConstK, where K=1..6, X is 0 or 1,
[0983] - Prediction list uses flags predFlagLXConstK, where K = 1..6, X is 0 or 1,
[0984] – Affine motion model index motionModelIdcConstK, where K = 1..6,
[0985] – Bidirectional prediction weight index bcwIdxConstK, where K = 1..6,
[0986] – The constructed affine control point motion vector cpMvLXConstK[cpIdx], where cpIdx=0..2, K=1..6, and X is 0 or 1.
[0987] …
[0988] The fourth (juxtaposed lower right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list utilization flag predFlagLXCorner[3] and availability flag availableFlagCorner[3] are derived as follows, where X is 0 and 1:
[0989] – The reference index refIdxLXCorner[3] of the temporal merge candidate is set equal to 0, where X is 0 or 1.
[0990] – The variables mvLXCol and availableFlagLXCol are derived as follows, where X is 0 or 1:
[0991] – If slice_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[0992] – Otherwise (slice_temporal_mvp_enabled_flag is equal to 1), the following applies:
[0993] xColBr=xCb+cbWidth (8-601)
[0994] yColBr=yCb+cbHeight (8-602)
[0995] rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?
[0996] SubPicRightBoundaryPos:pic_width_in_luma_samples-1
[0997] botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?
[0998] SubPicBotBoundaryPos:pic_height_in_luma_samples-1
[0999] – If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos, and xColBr is less than or equal to rightBoundaryPos, then the following applies:
[1000] – The variable colCb specifies the overlay
[1001] The luma codec block at the modification position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) is located in the collocated picture specified by ColPic.
[1002] – The luma position (xColCb, yColCb) is set equal to the top left sample of the collocated luma codec specified by colCb relative to the top left sample of the collocated picture specified by ColPic.
[1003] – Invoke the derivation process of the collocated motion vector specified in clause 8.5.2.12 with as input currCb, colCb, (xColCb, yColCb), refIdxLXCorner[3] and sbFlag set equal to 0, and assign the output to mvLXCol and availableFlagLXCol.
[1004] – Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[1005] …
[1006] 5.2 Example 2: Affine-Constructed Merge Candidate Sub-Image Constraints (Solution 2)
[1007] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[1008] 8.5.5.6 Derivation of Affine Control Point Motion Vector Merge Candidates for Construction
[1009] The inputs to this process are:
[1010] –Specify the luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,
[1011] – Two variables cbWidth and cbHeight specifying the width and height of the current brightness codec block,
[1012] – Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3,
[1013] – Sample point locations (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2), and (xNbB3, yNbB3).
[1014] The output of this process is:
[1015] – Availability flag availableFlagConstK of the constructed affine control point motion vector merge candidate, where K = 1..6,
[1016] – reference index refIdxLXConstK, where K=1..6, X is 0 or 1,
[1017] – prediction list utilization flag predFlagLXConstK, where K=1..6, X is 0 or 1, – affine motion model index motionModelIdcConstK, where K=1..6,
[1018] – Bidirectional prediction weight index bcwIdxConstK, where K = 1..6,
[1019] – The constructed affine control point motion vector cpMvLXConstK[cpIdx], where cpIdx=0..2, K=1..6, and X is 0 or 1.
[1020] …
[1021] The fourth (juxtaposed lower right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list utilization flag predFlagLXCorner[3] and availability flag availableFlagCorner[3] are derived as follows, where X is 0 and 1:
[1022] – The reference index refIdxLXCorner[3] of the temporal merge candidate is set equal to 0, where X is 0 and 1.
[1023] – The variables mvLXCol and availableFlagLXCol are derived as follows, where X is 0 and 1:
[1024] – If slice_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[1025] – Otherwise (slice_temporal_mvp_enabled_flag is equal to 1), the following applies:
[1026] ColBr=xCb+cbWidth (8-601)
[1027] yColBr=yCb+cbHeight (8-602)
[1028] rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?
[1029] SubPicRightBoundaryPos:pic_width_in_luma_samples-1
[1030] botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?
[1031] SubPicBotBoundaryPos:pic_height_in_luma_samples–1
[1032] xColBr=Min(rightBoundaryPos,xColBr)
[1033] yColBr=Min(botBoundaryPos,yColBr)
[1034] – If yCb>>CtbLog2SizeY equals yColBr>>CtbLog2SizeY, [[yColBr is less than pic_height_in_luma_samples, xColBr is less than
[1035] pic_width_in_luma_samples, if applicable]]:
[1036] – The variable colCb specifies that the luma codec block covering the modification position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) is within the collocated picture specified by ColPic.
[1037] – The luma position (xColCb, yColCb) is set equal to the top left sample of the collocated luma codec specified by colCb relative to the top left sample of the collocated picture specified by ColPic.
[1038] – Invoke the derivation process of the collocated motion vector specified in clause 8.5.2.12 with as input currCb, colCb, (xColCb, yColCb), refIdxLXCorner[3] and sbFlag set equal to 0, and assign the output to mvLXCol and availableFlagLXCol.
[1039] – Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.
[1040] …
[1041] 5.3 Example 3: Extracting integer samples under sub-image constraints
[1042] 8.5.6.3.3 Luminance integer sample acquisition process
[1043] The inputs to this process are:
[1044] – Luminance position in full sample units (xInt L , yInt L ),
[1045] – Luminance reference sample array refPicLXL,
[1046] The output of this process is the predicted luminance sample value predSampleLX L
[1047] The variable shift is set equal to Max(2,14-BitDepth Y ).
[1048] The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples.
[1049] The luminance position (xInt, yInt) in units of all samples is derived as follows:
[1050] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:
[1051] xInt=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt)
[1052] yInt=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt)
[1053] -otherwise:
[1054] xInt=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?(8-782)
[1055] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L ):xInt L )
[1056] yInt=Clip3(0,picH-1,yInt L )
[1057] (8-783)
[1058] Predicted brightness sample value predSampleLX L The derivation is as follows:
[1059] predSampleLX L =refPicLX L [xInt][yInt]< <shift3 (8-784)
[1060] 5.4 Example 4: Derivation of the variable invAvgLuma in LMCS chroma residual scaling
[1061] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[1062] 8.7.5.3 Image Reconstruction Using the Scaling Process of Luma-Dependent Chroma Residuals of Chroma Samples
[1063] The inputs to this process are:
[1064] – The chroma position (xCurr, yCurr) of the upper left chroma sample of the current chroma transform block relative to the upper left chroma sample of the current picture,
[1065] – variable nCurrSw that specifies the width of the chroma transform block,
[1066] – variable nCurrSh that specifies the chroma transform block height,
[1067] – Variable tuCbfChroma that specifies the codec block flag of the current chroma transform block,
[1068] –Specify the (nCurrSw)×(nCurrSh) array predSamples of the chroma prediction samples of the current block,
[1069] –Specify the (nCurrSw)×(nCurrSh) array resSamples of the chroma residual samples of the current block,
[1070] The output of this process is the reconstructed chroma picture sample array recSamples.
[1071] The variable sizeY is set equal to Min(CtbSizeY,64).
[1072] For i = 0..nCurrSw 1, j = 0..nCurrSh 1, the reconstructed chroma picture samples recSamples are derived as follows:
[1073] –…
[1074] – Otherwise, the following applies:
[1075] –…
[1076] –The variable currPic specifies the array of reconstructed luminance samples in the current picture.
[1077] – For the derivation of the variable varScale, the following ordered steps are applied:
[1078] 1. The variable invAvgLuma is derived as follows:
[1079] – The derivation of the array recLuma[i] and the variable cnt is as follows, where i=0..(2*sizeY-1):
[1080] –The variable cnt is set equal to 0.
[1081] – The derivation of variables rightBoundaryPos and botBoundaryPos is as follows: rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? ...SubPicRightBoundaryPos:pic_width_in_luma_samples-1botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? .........SubPicBotBoundaryPos:pic_height_in_luma_samples–1
[1082] – When availL is equal to TRUE, the array recLuma[i] (where i = 0..sizeY-1) is set equal to where i = 0..sizeY–1, and cnt is set equal to sizeY
[1083] – When availT is equal to TRUE, the array recLuma[cnt+i] (where i = 0..sizeY–1) is set equal to Where i = 0..sizeY–1, and cnt is set equal to (cnt+sizeY)
[1084] – The variable invAvgLuma is derived as follows:
[1085] – If cnt is greater than 0, then the following applies:
[1086] invAvgLuma=Clip1 Y ((+(cnt>>1))>>Log2(cnt))(8-1013)
[1087] – Otherwise (cnt equals 0), the following applies:
[1088] invAvgLuma=1<<(BitDepth Y –1) (8-1014)
[1089] 5.5 Example 5: Example of defining sub-picture elements in units of N (such as N=8 or 32) samples other than 4 samples
[1090] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[1091] 7.4.3.3 Sequence Parameter Set RBSP Semantics
[1092] subpic_grid_col_width_minus1 plus 1 specifies the width of each element of the sub-picture identifier grid in units of [[4]]N samples. The length of the syntax element is
[1093] Ceil(Log2(pic_width_max_in_luma_samples / [[4]]N))) bits.
[1094] The variable NumSubPicGridCols is derived as follows:
[1095]
[1096] subpic_grid_row_height_minus1 plus 1 specifies the height of each element of the sub-picture identifier grid in units of 4 samples. The length of the syntax element is
[1097] Ceil(Log2(pic_height_max_in_luma_samples / [[4]]N)) bits.
[1098] The variable NumSubPicGridRows is derived as follows:
[1099] NumSubPicGridRows=
[1100] (pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*[[4]]N+N-1) /
[1101] (subpic_grid_row_height_minus1*[[4+3]]N+N-1)
[1102] 7.4.7.1 Common Strip Header Semantics
[1103] The variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[1104]
[1105]
[1106] 5.6 Example 6: Limiting the image width and image height to be equal to or greater than 8
[1107] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[1108] 7.4.3.3 Sequence Parameter Set RBSP Semantics
[1109] pic_width_max_in_luma_samples specifies the maximum width in units of luma samples of each decoded picture that references the SPS. pic_width_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of [[MinCbSizeY]]Max(8,MinCbSizeY).
[1110] pic_height_max_in_luma_samples specifies the maximum height in luma samples of each decoded picture that references the SPS. pic_height_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of [[MinCbSizeY]]Max(8,MinCbSizeY).
[1111] 5.7 Example 7: Sub-picture Boundary Check for Signaling of BT / TT / QT Split, BT / TT / QT Depth Derivation, and / or CU Split Flag
[1112] The working draft specified in JVET-O2001-v14 is subject to change as follows.
[1113] 6.4.2 Allowable binary partitioning procedures
[1114] The variable allowBtSplit is derived as follows:
[1115] –…
[1116] – Otherwise, if all of the following conditions are true, allowBtSplit will be set equal to FALSE
[1117] –btSplit equals SPLIT_BT_VER
[1118] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos+1:pic_height_in_luma_samples.
[1119] – Otherwise, if all of the following conditions are true, allowBtSplit will be set equal to FALSE
[1120] –btSplit equals SPLIT_BT_VER
[1121] –cbHeight is greater than MaxTbSizeY
[1122] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]?
[1123] SubPicRightBoundaryPos+1:pic_width_in_luma_samples
[1124] – Otherwise, if all of the following conditions are true, allowBtSplit will be set equal to FALSE
[1125] –btSplit is equal to SPLIT_BT_HOR
[1126] –cbWidth is greater than MaxTbSizeY
[1127] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos+1:pic_height_in_luma_samples.
[1128] – Otherwise, if all of the following conditions are true, allowBtSplit will be set equal to FALSE
[1129] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]?
[1130] SubPicRightBoundaryPos+1:pic_width_in_luma_samples
[1131] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos+1:pic_height_in_luma_samples.
[1132] –cbWidth is greater than minQtSize
[1133] – Otherwise, if all of the following conditions are true, allowBtSplit will be set equal to FALSE
[1134] –btSplit is equal to SPLIT_BT_HOR
[1135] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]?
[1136] SubPicRightBoundaryPos+1:pic_width_in_luma_samples
[1137] –y0+cbHeight is less than or equal to [[pic_height_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]?
[1138] SubPicBotBoundaryPos+1:pic_height_in_luma_samples.
[1139] 6.4.3 Allowed ternary division processes
[1140] The derivation of the variable allowTtSplit is as follows:
[1141] – allowTtSplit is set equal to FALSE if one or more of the following conditions are true:
[1142] –cbSize is less than or equal to 2*MinTtSizeY
[1143] –cbWidth is greater than Min(MaxTbSizeY,maxTtSize)
[1144] –cbHeight is greater than Min(MaxTbSizeY,maxTtSize)
[1145] –mttDepth is greater than or equal to maxMttDepth
[1146] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]?
[1147] SubPicRightBoundaryPos+1:pic_width_in_luma_samples
[1148] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos+1:pic_height_in_luma_samples.
[1149] –treeType is equal to DUAL_TREE_CHROMA, and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32
[1150] –treeType equals DUAL_TREE_CHROMA, modeType equals MODE_TYPE_INTRA
[1151] – Otherwise, allowTtSplit is set to TRUE.
[1152] 7.3.8.2 Codec Tree Unit Syntax
[1153]
[1154] 7.3.8.4 Codec Tree Syntax
[1155]
[1156]
[1157]
[1158] 5.8 Example 8: Example of defining a sub-image
[1159]
[1160]
[1161] 5.9 Example 9: Example of defining a sub-image
[1162]
[1163]
[1164] 5.10 Example 10: Example of Defining a Sub-Picture
[1165]
[1166]
[1167] 5.11 Example 11: Example of defining a sub-image
[1168]
[1169]
[1170] NumSubPics=num_subpics_minus2+2.
[1171] 5.12 Example: Considering Deblocking of Sub-Pictures
[1172] Deblocking filtering process
[1173] Overview
[1174] The input of this process is the reconstructed image before deblocking, that is, the array recPicture L , and when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1175] The output of this process is the modified reconstructed picture after deblocking, namely the array recPictureL, and when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1176] First, vertical edges in the image are filtered. Then, horizontal edges in the image are filtered using the samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTBs of each CTU are processed separately on a codec unit basis. Vertical edges of the codec blocks in a codec unit are filtered starting at the edge on the left side of the codec block and proceeding through the edges in their geometric order towards the right side of the codec block. Horizontal edges of the codec blocks in a codec unit are filtered starting at the edge at the top of the codec block and proceeding through the edges in their geometric order towards the bottom of the codec block.
[1177] NOTE – Although the filtering process is specified on a picture basis in this specification, the filtering process can also be implemented on a codec unit basis with equivalent results, as long as the decoder correctly respects the processing dependency order to produce the same output values.
[1178] The deblocking filtering process is applied to all codec sub-block edges and transform block edges of the picture, except for the following types of edges:
[1179] – edges on the image border,
[1180] –[[When loop_filter_cross_sub pic_enabled_flag[SubPicIdx] is equal to 0, the edge that coincides with the boundary of the sub-picture,]]
[1181] – When PPS_loop_filter_cross_virtual_boundaries_disabled_flag is equal to 1, edges that coincide with the virtual boundaries of the picture,
[1182] – When loop_filter_cross_tiles_enabled_flag is equal to 0, edges that coincide with tile boundaries,
[1183] – When loop_filter_cross_slices_enabled_flag is equal to 0, edges that coincide with slice boundaries,
[1184] – When slice_deblocking_filter_disabled_flag is equal to 1, the edge that coincides with the upper or left border of the slice,
[1185] –slice_deblocking_filter_disabled_flag is equal to 1, inner edge of the strip,
[1186] – edges that do not correspond to the boundaries of the 4×4 sample grid of the luma component,
[1187] – edges that do not correspond to the boundaries of the 8×8 sample grid of the chroma components,
[1188] – edges with intra_bdpcm_luma_flag equal to 1 on both sides of the edge in the luma component,
[1189] – edges with intra_bdpcm_chroma_flag equal to 1 on either side of the edge in the chroma components,
[1190] – Edges of chroma sub-blocks that are not edges of the associated transform unit.
[1191] …
[1192] Deblocking filtering process in one direction
[1193] The inputs to this process are:
[1194] – specifies the variable treeType that currently processes the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA),
[1195] –When treeType is equal to DUAL_TREE_LUMA, the reconstructed picture before removing the block, that is, the array recPicture L ,
[1196] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr ,
[1197] – The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are to be filtered.
[1198] The output of this process is the modified reconstructed image after deblocking, namely:
[1199] – When treeType is equal to DUAL_TREE_LUMA, the array recPicture L
[1200] – When ChromaArrayType is not equal to 0 and treeType is equal to
[1201] DUAL_TREE_CHROMA, array recPicture Cb and recPicture Cr .
[1202] The variables firstCompIdx and lastCompIdx are derived as follows:
[1203] firstCompIdx=(treeType==DUAL_TREE_CHROMA)? 1:0
[1204] (8-1010)
[1205] lastCompIdx=(treeType==DUAL_TREE_LUMA||ChromaArrayType==0)? 0:2 (8-1011)
[1206] For each codec unit and each codec block of each color component of the codec unit indicated by the color component index cIdx, having a codec block width nCbW, a codec block height nCbH, and a position of the top left sample point (xCb, yCb) of the codec block, where cIdx ranges from firstCompIdx to lastCompIdx, inclusive, when cIdx is equal to 0, or when cIdx is not equal to 0 and edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when cIdx is not equal to 0 and edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, edge filtering is performed by the following ordered steps:
[1207] 2. The derivation of the variable filterEdgeFlag is as follows:
[1208] – If edgeType is equal to EDGE_VER and one or more of the following conditions are true, filterEdgeFlag is set equal to 0:
[1209] – The left border of the current codec block is the left border of the picture.
[1210] –[[The left border of the current codec block is the left or right border of a sub-picture, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] is equal to 0. ]]
[1211] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[1212] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[1213] – The left boundary of the current codec block is one of the vertical virtual boundaries of the picture, and VirtualBoundariesDisabledFlag is equal to 1.
[1214] Otherwise, if edgeType is equal to EDGE_HOR and one or more of the following conditions are true, then the variable filterEdgeFlag is set equal to 0:
[1215] – The top boundary of the current luma codec block is the top boundary of the picture.
[1216] –[[The top boundary of the current codec block is the top or bottom boundary of a sub-picture, and loop_filter_cross_sub pic_enabled_flag[SubPicIdx] is equal to 0. ]]
[1217] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[1218] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[1219] – The top boundary of the current codec block is one of the horizontal virtual boundaries of the picture, and VirtualBoundariesDisabledFlag is equal to 1.
[1220] – Otherwise, filterEdgeFlag is set equal to 1.
[1221] …
[1222] Use a short filter to filter the luminance samples
[1223] Inputs to this process include:
[1224] – Sample value p i and q i , where i = 0..3,
[1225] –p i and q i The position of (xP i , yP i ) and (xQ i , yQ i ), where i = 0..2,
[1226] – variable dE,
[1227] – The variables dEp and dEq contain the decision to filter the samples p1 and q1 respectively,
[1228] –Variable t C .
[1229] The output of this process is:
[1230] – the number of filtered samples nDp and nDq,
[1231] –Filtered sample value p i 'and q j', where i = 0..nDp-1, j = 0..nDq–1.
[1232] Depending on the value of dE, the following procedure applies
[1233] – If the variable dE is equal to 2, then both nDp and nDq are set equal to 3 and the following strong filtering is applied:
[1234] p0′=Clip3(p0-3*t C ,p0+3*t C ,(p2+2*p1+2*p0+2*q0+q1+4)>>3)(8-1150)
[1235] p1′=Clip3(p1-2*t C ,p1+2*t C ,(p2+p1+p0+q0+2)>>2)(8-1151)
[1236] p2′=Clip3(p2-1*t C ,p2+1*t C ,(2*p3+3*p2+p1+p0+q0+4)>>3)(8-1152)
[1237] q0′=Clip3(q0-3*t C ,q0+3*t C ,(p1+2*p0+2*q0+2*q1+q2+4)>>3)(8-1153)
[1238] q1′=Clip3(q1-2*t C ,q1+2*t C ,(p0+q0+q1+q2+2)>>2)(8-1154)
[1239] q2′=Clip3(q2-1*t C ,q2+1*t C ,(p0+q0+q1+3*q2+2*q3+4)>>3)(8-1155)
[1240] – Otherwise, both nDp and nDq are set to 0 and the following weak filtering is applied:
[1241] – Apply the following:
[1242] Δ=(9*(q0-p0)-3*(q1-p1)+8)>>4 (8-1156)
[1243] –When Abs(Δ) is less than t C *10, apply the following sequential steps:
[1244] – The filtered sample values p0' and q0' are specified as follows:
[1245] Δ=Clip3(-t C ,t C ,Δ) (8-1157)
[1246] p0′=Clip1(p0+Δ) (8-1158)
[1247] q0′=Clip1(q0-Δ) (8-1159)
[1248] – When dEp is equal to 1, the filtered sample value p1' is specified as follows:
[1249] Δp=Clip3(-(t C >>1),t C >>1,(((p2+p0+1)>>1)-p1+Δ)>>1) (8-1160)
[1250] p1′=Clip1(p1+Δp)
[1251] (8-1161)
[1252] – When dEq is equal to 1, the filtered sample value q1' is specified as follows:
[1253] Δq=Clip3(-(t C >>1),t C >>1,(((q2+q0+1)>>1)-q1-Δ)>>1) (8-1162)
[1254] q1′=Clip1(q1+Δq)
[1255] (8-1163)
[1256] –nDp is set equal to dEp+1, and nDq is set equal to dEq+1.
[1257] When nDp is greater than 0 and pred_mode_plt_flag of the codec unit including the codec block containing sample p0 is equal to 1, nDp is set to 0
[1258] When nDq is greater than 0 and pred_mode_plt_flag of the codec unit including the codec block containing sample q0 is equal to 1, nDq is set to 0
[1259] When nDp is greater than 0 and loop_filter_across_subpic_enabled_flag[subPicIdxP] is equal to 0, nDp is set equal to 0, where subPicIdxP is the sub-picture index of the sub-picture containing sample p0.
[1260] When nDq is greater than 0 and loop_filter_across_subpic_enabled_flag[subPicIdxQ] is equal to 0, nDq is set equal to 0, where subPicIdxQ is the sub-picture index of the sub-picture containing sample q0.
[1261] Filtering the luminance samples using a long filter
[1262] The inputs to this process are:
[1263] – variables maxFilterLengthP and maxFilterLengthQ,
[1264] – Sample point p i and q j , where i = 0..maxFilterLengthP and j = 0..maxFilterLengthQ,
[1265] –p i and q j Position (xP i , yP i ) and (xQ j , yQ j ), where i = 0..maxFilterLengthP-1 and j = 0..maxFilterLengthQ-1,
[1266] –Variable t C .
[1267] The output of this process is:
[1268] –Filtered sample value p i 'and q j ', where i=0..maxFilterLengthP-1 and j=0..maxFilterLengthQ-1.
[1269] The derivation of the variable refMiddle is as follows:
[1270] – If maxFilterLengthP is equal to maxFilterLengthQ and maxFilterLengthP is equal to 5, then the following applies:
[1271] refMiddle=(p4+p3+2*(p2+p1+p0+q0+q1+q2)+q3+q4+8)>>4
[1272] (8-1164)
[1273] – Otherwise, if maxFilterLengthP is equal to maxFilterLengthQ and
[1274] If maxFilterLengthP is not equal to 5, the following applies:
[1275] refMiddle=(p6+p5+p4+p3+p2+p1+2*(p0+q0)+q1+q2+q3+q4+q5+q6+8)>>4 (8-1165)
[1276] – Otherwise, if one of the following conditions is true,
[1277] – maxFilterLengthQ is equal to 7 and maxFilterLengthP is equal to 5,
[1278] – maxFilterLengthQ is equal to 5 and maxFilterLengthP is equal to 7,
[1279] Apply the following:
[1280] refMiddle=(p5+p4+p3+p2+2*(p1+p0+q0+q1)+q2+q3+q4+q5+8)>>4(8-1166)
[1281] – Otherwise, if one of the following conditions is true,
[1282] – maxFilterLengthQ is equal to 5 and maxFilterLengthP is equal to 3,
[1283] – maxFilterLengthQ is equal to 3 and maxFilterLengthP is equal to 5, the following applies:
[1284] refMiddle=(p3+p2+p1+p0+q0+q1+q2+q3+4)>>3 (8-1167)
[1285] – Otherwise, if maxFilterLengthQ is equal to 7 and maxFilterLengthP is equal to 3, then the following applies:
[1286] refMiddle=(2*(p2+p1+p0+q0)+p0+p1+q1+q2+q3+q4+q5+q6+8)>>4 (8-1168)
[1287] – Otherwise, the following applies:
[1288] refMiddle=(p6+p5+p4+p3+p2+p1+2*(q2+q1+q0+p0)+q0+q1+8)>>4 (8-1169)
[1289] The variables refP and refQ are derived as follows:
[1290] refP=(p maxFilterLengtP +p maxFilterLengthP-1 +1)>>1 (8-1170)
[1291] refQ=(q maxFilterLengtQ +q maxFilterLengthQ-1 +1)>>1 (8-1171)
[1292] variable f i and t C PD i The definitions are as follows:
[1293] – If maxFilterLengthP is equal to 7, then the following applies:
[1294] f 0..6 ={59,50,41,32,23,14,5}
[1295] (8-1172)
[1296] t C PD 0..6 ={6,5,4,3,2,1,1}
[1297] (8-1173)
[1298] – Otherwise, if maxFilterLengthP is equal to 5, then the following applies:
[1299] f 0..4 ={58,45,32,19,6} (8-1174)
[1300] t C PD 0..4 ={6,5,4,3,2} (8-1175)
[1301] – Otherwise, the following applies:
[1302] f 0..2={53,32,11} (8-1176)
[1303] t C PD 0..2 ={6,4,2} (8-1177)
[1304] variable g j and t C QD j The definitions are as follows:
[1305] – If maxFilterLengthQ is equal to 7, then the following applies:
[1306] g 0..6 ={59,50,41,32,23,14,5}
[1307] (8-1178)
[1308] t C QD 0..6 ={6,5,4,3,2,1,1} (8-1179)
[1309] – Otherwise, if maxFilterLengthQ is equal to 5, then the following applies:
[1310] g 0..4 ={58,45,32,19,6} (8-1180)
[1311] t C QD 0..4 ={6,5,4,3,2} (8-1181)
[1312] – Otherwise, the following applies:
[1313] g 0..2 ={53,32,11} (8-1182)
[1314] t C QD 0..2 ={6,4,2} (8-1183)
[1315] Filtered sample value p i 'and q j ' is derived as follows, where i = 0..maxFilterLengthP-1 and j = 0..maxFilterLengthQ-1:
[1316] p i ′=Clip3(p i -(t C *t C PD i )>>1,pi +(t C *t C PD i )>>1,(refMiddle*f i +refP*(64-f i )+32)>>6) (8-1184)
[1317] q j ′=Clip3(q j -(t C *t C QD j )>>1,q j +(t C *t C QD j )>>1,(refMiddle*g j +refQ*(64-g j )+32)>>6) (8-1185)
[1318] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i=0..maxFilterLengthP-1.
[1319] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The corresponding input sample value q j Instead, where j = 0..maxFilterLengthQ-1.
[1320] When loop_filter_across_subpic_enabled_flag[subPicIdxP] is equal to 0, the filtered sample value p i 'The corresponding input sample value p i Instead, where subPicIdxP is the sub-picture index of the sub-picture containing sample p0, i=0..maxFilterLengthP-1.
[1321] When loop_filter_across_subpic_enabled_flag[subPicIdxQ] is equal to 0, the filtered sample value q i 'The corresponding input sample value q jInstead, where subPicIdxQ is the sub-picture index of the sub-picture containing sample q0, j = 0..maxFilterLengthQ-1.
[1322] Filtering process of chroma samples
[1323] This procedure is called only if ChromaArrayType is not equal to 0.
[1324] The inputs to this process are:
[1325] – variable maxFilterLength,
[1326] – Chroma sample value p i and q i , where i = 0..maxFilterLengthCbCr,
[1327] –p i and q i The chromaticity position (xP i , yP i ) and (xQ i , yQ i ), where i = 0..maxFilterLengthCbCr-1,
[1328] –Variable t C .
[1329] The output of this process is the filtered sample value p i 'and q i ',in
[1330] i=0..maxFilterLengthCbCr-1.
[1331] Filtered sample value p i 'and q i ' is derived as follows, where i = 0..maxFilterLengthCbCr-1:
[1332] – If maxFilterLengthCbCr is equal to 3, the following strong filtering is applied:
[1333] p0′=Clip3(p0-t C ,p0+t C ,(p3+p2+p1+2*p0+q0+q1+q2+4)>>3) (8-1186)
[1334] p1′=Clip3(p1-t C ,p1+t C,(2*p3+p2+2*p1+p0+q0+q1+4)>>3) (8-1187)
[1335] p2′=Clip3(p2-t C ,p2+t C ,(3*p3+2*p2+p1+p0+q0+4)>>3)(8-1188)
[1336] q0′=Clip3(q0-t C ,q0+t C ,(p2+p1+p0+2*q0+q1+q2+q3+4)>>3) (8-1189)
[1337] q1′=Clip3(q1-t C ,q1+t C ,(p1+p0+q0+2*q1+q2+2*q3+4)>>3) (8-1190)
[1338] q2′=Clip3(q2-t C ,q2+t C ,(p0+q0+q1+2*q2+3*q3+4)>>3)(8-1191)
[1339] – Otherwise, apply the following weak filtering:
[1340] Δ=Clip3(-t C ,t C ,((((q0-p0)<<2)+p1-q1+4)>>3))(8-1192)
[1341] p0′=Clip1(p0+Δ) (8-1193)
[1342] q0′=Clip1(q0-Δ) (8-1194)
[1343] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead,
[1344] i=0..maxFilterLengthCbCr-1.
[1345] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i'The defined input sample value q i Instead, where i=0..maxFilterLengthCbCr-1:
[1346] When loop_filter_across_subpic_enabled_flag[subPicIdxP] is equal to 0, the filtered sample value p i 'The corresponding input sample value p i Instead, where subPicIdxP is the sub-picture index of the sub-picture containing sample p0, i = 0..maxFilterLengthCbCr-1.
[1347] When loop_filter_across_subpic_enabled_flag[subPicIdxQ] is equal to 0, the filtered sample value q i 'The corresponding input sample value q i Instead, where subPicIdxQ is the sub-picture index of the sub-picture containing sample q0, i = 0..maxFilterLengthCbCr-1:
[1348] Figure 6 6 is a block diagram illustrating an example video processing system 6000 in which the various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 6000. System 6000 may include an input 6002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10 bit multi-component pixel values, or may be in a compressed or encoded format. Input 6002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or a cellular interface.
[1349] System 6000 may include a codec component 6004 that can implement the various codecs or encoding methods described in this document. The codec component 6004 can reduce the average bit rate of the video from the input 6002 to the output of the codec component 6004 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. As represented by component 6006, the output of the codec component 6004 can be stored or sent via a connected communication. Component 6008 can use the stored or transmitted bitstream (or encoding) representation of the video received at input 6002 to generate pixel values or displayable video sent to display interface 6010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the encoding and decoding results will be performed by the decoder.
[1350] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[1351] Figure 7 7000 is a block diagram of a video processing apparatus 7000. The apparatus 7000 may be used to implement one or more methods described herein. The apparatus 7000 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 7000 may include one or more processors 7002, one or more memories 7004, and video processing hardware 7006. The processor 7002 may be configured to implement the present document (e.g., Figure 11-14 ) . Memory 7004 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 7006 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, hardware 7006 can be partially or completely within processor 7002, such as a graphics processor.
[1352] Figure 8 is a block diagram illustrating an example video encoding and decoding system 100 that may utilize the techniques of this disclosure. Figure 8As shown, video codec system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data, which may be referred to as a video encoding device. Destination device 120 may decode the encoded video data generated by source device 110, which may be referred to as a video decoding device. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[1353] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. Associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 via the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[1354] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[1355] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain coded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the coded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, with the destination device 120 being configured to interface with an external display device.
[1356] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.
[1357] Figure 9 is a block diagram illustrating an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown.
[1358] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[1359] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[1360] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.
[1361] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated but are not shown here for explanation purposes. Figure 9 In the example, it is represented separately.
[1362] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[1363] The mode selection unit 203 may, for example, select a codec mode - intra or inter - based on the error result, and provide the resulting intra or inter codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[1364] To perform inter-frame prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from buffer 213. Motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from buffer 213.
[1365] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[1366] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[1367] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[1368] In some examples, motion estimation unit 204 may output complete motion information for use in the decoding process of a decoder.
[1369] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for the neighboring video block.
[1370] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[1371] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[1372] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[1373] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[1374] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[1375] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[1376] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[1377] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[1378] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block for storage in buffer 213.
[1379] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[1380] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[1381] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown.
[1382] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 10 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[1383] exist Figure 10 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those generally performed for the video encoder 200 ( Figure 9 ) is a decoding process that is the inverse of the encoding process described.
[1384] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information from the entropy-decoded video data. The motion compensation unit 302 can determine this information, for example, by performing AMVP and merge modes.
[1385] The motion compensation unit 302 may generate a motion compensated block and may perform interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[1386] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used by video encoder 20 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to produce a prediction block.
[1387] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[1388] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[1389] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[1390] Figure 11-14 It is shown that it is possible to Figures 6-10 The illustrated embodiment is an exemplary method for implementing the above technical solution.
[1391] Figure 11 A flowchart of an example method 1100 for video processing is shown. The method 1100 includes performing, at operation 1110, conversion between a video including a picture and a bitstream of the video, the picture including one or more sub-pictures, the bitstream conforming to a format rule, the format rule specifying that the bitstream includes a parameter set, the parameter set controlling a codec behavior of a sub-picture of the one or more sub-pictures associated with an identifier (ID) of the sub-picture.
[1392] Figure 12A flowchart of an example method 1200 for video processing is shown. The method 1200 includes performing, at operation 1210, conversion between a video including a picture and a bitstream of the video, the picture including one or more sub-pictures, a current parameter set configured to control a codec behavior of at least one of the one or more sub-pictures, and the bitstream conforming to a format rule, the format rule specifying that a default parameter set corresponding to the current parameter set be signaled in the bitstream before a difference between the current parameter set and the default parameter set is signaled in the bitstream.
[1393] Figure 13 A flowchart of an example method 1300 for video processing is shown. The method 1300 includes: at operation 1310, performing conversion between a video including a picture and a bitstream of the video, the picture including one or more sub-pictures, the bitstream including a parameter set, the parameter set including a first control parameter and a second control parameter for controlling codec properties of the sub-picture, and the bitstream conforming to a format rule, the format rule specifying whether or how the first control parameter is overwritten by the second control parameter for decoding.
[1394] Figure 14 A flowchart of an example method 1400 for video processing is shown. The method 1400 includes performing conversion between a video including a picture and a bitstream of the video at operation 1410, the picture including one or more sub-pictures, one or more first flags corresponding to each of the one or more sub-pictures being included in a sequence parameter set (SPS), each of the one or more first flags indicating whether constraint information is signaled for the sub-picture corresponding to each first flag, and the constraint information indicating a codec tool not applied to the corresponding sub-picture on a codec layer video sequence (CLVS).
[1395] A list of solutions preferred by some embodiments is provided next.
[1396] 1. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the bitstream comprises a parameter set, the parameter set controlling encoding and decoding behavior of the sub-pictures in the one or more sub-pictures associated with an identifier (ID) of the sub-picture.
[1397] 2. The method of solution 1, wherein the parameter set comprises at least one of: a quantization parameter (QP) or a QP increment of a luma component of a sub-picture, a QP or a QP increment of a chroma component of a sub-picture, reference picture list management information, a codec tree unit (CTU) size of a picture, a minimum codec unit (CU) size of a picture, a maximum transform unit (TU) size of a picture, a maximum quadtree (QT) partition size of a picture, a minimum QT partition size of a picture, a maximum QT partition depth of a picture, a minimum QT partition depth of a picture, a maximum binary tree (BT) partition size of a picture, a minimum BT partition size of a picture, a maximum BT partition depth of a picture, a minimum BT partition depth of a picture, a maximum ternary tree (TT) partition size of a picture, a minimum TT partition size of a picture, a maximum TT partition depth of a picture, a minimum TT partition depth of a picture, a maximum multitree (MT) partition size of a picture, a minimum MT partition size of a picture, a maximum MT partition depth of a picture, and a minimum MT partition depth of a picture, and the parameter set controls one or more codec tools.
[1398] 3. The method according to solution 2, wherein the one or more codec tools include at least one of the following: weighted prediction, sample adaptive offset (SAO), adaptive loop filtering (ALF), transform skip, block differential pulse coding modulation (BDPCM), joint Cb-Cr residual (JCCR) codec, reference surround, temporal motion vector prediction (TMVP), sub-block temporal motion vector prediction (sbTMVP), adaptive motion vector resolution (AMVR), bidirectional optical flow (BDOF), symmetric motion vector difference (SMVD), decoder-side motion vector refinement (DMVR), using motion vector Merge of Quantity Difference (MMVD), Intra Sub-Partitioning (ISP) mode, (MRL), Matrix-based Intra Prediction (MIP), Cross-Component Linear Model (CCLM), CCLM collocated chroma control, Multiple Transform Set (MTS) for Intra and / or Inter, MTS for Inter, Sub-Block Transform (SBT), SBT maximum size, Affine Codec, Affine Type Codec, Palette Codec, Bidirectional Prediction with CU Weights (BCW), Intra Block Copy (IBC), Combined Inter-Intra Prediction (CIIP), Triangle-based Motion Compensation, and Luma Mapping and Chroma Transform (LMCS).
[1399] 4. The method according to any one of solutions 1 to 3, wherein the bitstream includes a single flag indicating that the parameter set for each of the one or more sub-pictures is the same.
[1400] 5. The method of solution 4, wherein for one or more sub-pictures, only a single copy of the parameter set is signaled in the bitstream.
[1401] 6. The method of solution 1, wherein performing the conversion comprises applying predictive coding to parameter sets of at least two sub-pictures of the one or more sub-pictures.
[1402] 7. The method according to solution 6, wherein the difference between two values of a syntax element of two different sub-pictures is encoded and decoded.
[1403] 8. The method of solution 1, wherein the parameter set is signaled in a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header.
[1404] 9. The method of solution 1, wherein the parameter set is signaled in a Supplemental Enhancement Information (SEI) message or a Video Usage Information (VUI) message.
[1405] 10. The method of solution 9, wherein the SEI message is a sub-picture level information SEI message.
[1406] 11. A method according to solution 1, wherein the parameter set is signaled in a sub-picture parameter set (SPPS), and the sub-picture parameter set is different from the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), picture header and slice header.
[1407] 12. The method of solution 11, wherein an SPPS index associated with the SPPS is signaled in the bitstream.
[1408] 13. The method of solution 11, wherein an SPPS index indicating the SPPS associated with the corresponding sub-picture is signaled in the bitstream.
[1409] 14. A method according to any one of solutions 1 to 13, wherein the bitstream includes syntax elements associated with a slice, a tile, a tile or a sub-picture, and wherein the syntax elements depend on a parameter set of a sub-picture comprising the current slice.
[1410] 15. A method according to any one of solutions 1 to 13, wherein a first control parameter of a parameter set controls a codec behavior, and wherein, based on a consistency rule, a second control parameter of the parameter set controlling the codec behavior is the same as the first control parameter.
[1411] 16. The method of solution 1, wherein an indication regarding application of one or more codec tools in a Codec Layer Video Sequence (CLVS) is signaled in the bitstream.
[1412] 17. The method of solution 16, wherein the indication is signaled in a Supplemental Enhancement Information (SEI) message or a Video Usage Information (VUI) message.
[1413] 18. The method of solution 16, wherein the indication is signaled in a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), or a standalone network abstraction layer (NAL) unit.
[1414] 19. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein a current parameter set is configured to control encoding and decoding behavior of at least one of the one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a default parameter set corresponding to the current parameter set is signaled in the bitstream before a difference between the current parameter set and the default parameter set is signaled in the bitstream.
[1415] 20. The method of solution 19, wherein the format rule further specifies that after the default parameter set, the difference between the default parameter set and the current parameter set is signaled.
[1416] 21. The method of solution 9, wherein, before the default parameter set, a flag indicating that the encoding and decoding behavior of each of the one or more sub-pictures is controlled by the default parameter set is signaled in the bitstream.
[1417] 22. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein the bitstream conforms to a format rule, wherein the bitstream comprises a parameter set, the parameter set comprising a first control parameter and a second control parameter for controlling encoding and decoding properties of the sub-picture, and wherein the format rule specifies whether or how the first control parameter is overwritten by the second control parameter for decoding.
[1418] 23. The method of solution 22, wherein the second control parameter is signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, or a slice header.
[1419] 24. A video processing method, comprising: performing conversion between a video including a picture and a bitstream of the video, the picture including one or more sub-pictures, wherein one or more first flags corresponding to each of the one or more sub-pictures are included in a sequence parameter set (SPS), wherein each first flag of the one or more first flags indicates whether constraint information is signaled for the sub-picture corresponding to each first flag, and wherein the constraint information indicates a codec tool that is not applied to the corresponding sub-picture on a codec layer video sequence (CLVS).
[1420] 25. The method of solution 24, wherein a second flag indicating whether each of the one or more first flags is signaled is signaled in the bitstream.
[1421] 26. The method of solution 24 or 25, wherein the constraint information comprises a general_constraint_info() syntax structure.
[1422] 27. The method of any one of solutions 1 to 26, wherein converting comprises decoding the video from a bitstream.
[1423] 28. The method of any one of solutions 1 to 26, wherein converting comprises encoding the video into a bitstream.
[1424] 29. A method of writing a bitstream representing a video to a computer-readable recording medium, comprising: generating a bitstream from a video according to the method of any one of Solutions 1 to 26; and writing the bitstream to the computer-readable recording medium.
[1425] 30. A method for storing a bitstream of a video, comprising: performing a conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures; generating a bitstream from a current block; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the bitstream complies with a format rule, wherein the format rule specifies that the bitstream includes a parameter set that controls the encoding and decoding behavior of a sub-picture in one or more sub-pictures associated with an identifier (ID) of the sub-picture.
[1426] 31. A video processing device comprising a processor configured to implement the method as described in any one or more of solutions 1 to 30.
[1427] 32. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to implement the method as described in any one or more of solutions 1 to 30.
[1428] 33. A computer-readable medium storing a bitstream generated according to any one of solutions 1 to 30.
[1429] 34. A video processing device for storing a bitstream, wherein the video processing device is configured to implement the method as described in any one or more of solutions 1 to 30.
[1430] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this application document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or any combination thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible, non-volatile computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or any combination thereof. The term "data processing unit" or "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or a plurality of processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[1431] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers, located at one site or distributed across multiple sites and interconnected by a communications network.
[1432] The processes and logic flows described in this application document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and the apparatus can also be implemented as, special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[1433] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[1434] While this patent document contains many specifics, they should not be construed as limitations on the scope of any invention or the claims, but rather as descriptions of features for particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable subcombination. Furthermore, while the features described above may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features in a claim combination may be removed from the combination, and a claim combination may be directed to a subcombination or variations of a subcombination.
[1435] Likewise, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[1436] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, wherein the bitstream complies with the format rules, wherein the format rule specifies that the bitstream includes a parameter set, the parameter set controlling encoding and decoding behavior of a sub-picture of the one or more sub-pictures associated with an identification (ID) of the sub-picture; The bitstream includes a single flag indicating that the parameter set for each of the one or more sub-pictures is the same.
2. The method according to claim 1, wherein The parameter set includes at least one of the following: a quantization parameter QP or a QP increment of a luma component of the sub-picture, a QP or a QP increment of a chroma component of the sub-picture, reference picture list management information, a codec tree unit (CTU) size of the picture, a minimum codec unit (CU) size of the picture, a maximum transform unit (TU) size of the picture, a maximum quadtree (QT) partition size of the picture, a minimum QT partition size of the picture, a maximum QT partition depth of the picture, a minimum QT partition depth of the picture, a maximum binary tree (BT) partition size of the picture, a minimum BT partition size of the picture, a maximum BT partition depth of the picture, a minimum BT partition depth of the picture, a maximum ternary tree (TT) partition size of the picture, a minimum TT partition size of the picture, a maximum TT partition depth of the picture, a minimum TT partition depth of the picture, a maximum multitree (MT) partition size of the picture, a minimum MT partition size of the picture, a maximum MT partition depth of the picture, and a minimum MT partition depth of the picture, and the parameter set controls one or more codec tools.
3. The method according to claim 2, wherein: The one or more codec tools include at least one of the following: weighted prediction, sample adaptive offset SAO, adaptive loop filtering ALF, transform skip, block differential pulse coding modulation BDPCM, joint Cb-Cr residual JCCR codec, reference surround, temporal motion vector prediction TMVP, sub-block temporal motion vector prediction sbTMVP, adaptive motion vector resolution AMVR, bidirectional optical flow BDOF, symmetric motion vector difference SMVD, decoder-side motion vector refinement DMVR, Merge with motion vector difference (MMVD), intra-frame sub-partitioning ISP mode, MRL, matrix-based intra-frame prediction MIP, cross-component linear model CCLM, CCLM collocated chroma control, multiple transform set MTS for intra and / or inter-frame, MTS for inter-frame, sub-block transform SBT, SBT maximum size, affine codec, affine type codec, palette codec, bidirectional prediction BCW with CU weights, intra-frame block copy IBC, combined inter-frame-intra prediction CIIP, triangle-based motion compensation, and luma mapping and chroma transform LMCS.
4. The method according to claim 1, wherein For the one or more sub-pictures, only a single copy of the parameter set is signaled in the bitstream.
5. The method according to claim 1, wherein Performing the conversion includes applying predictive coding to parameter sets for at least two of the one or more sub-pictures.
6. The method according to claim 5, wherein: The difference between two values of a syntax element of two different sub-pictures is encoded and decoded.
7. The method according to claim 1, wherein The parameter sets are signaled in a sequence parameter set SPS, a picture parameter set PPS or a picture header.
8. The method according to claim 1, wherein The parameter set is signaled in a Supplemental Enhancement Information SEI message or a Video Usage Information VUI message.
9. The method according to claim 8, wherein The SEI message is a sub-picture level information SEI message.
10. The method according to claim 1, wherein The parameter set is signaled in a sub-picture parameter set SPPS, which is different from the video parameter set VPS, sequence parameter set SPS, picture parameter set PPS, picture header and slice header.
11. The method according to claim 10, wherein: An SPPS index associated with the SPPS is signaled in the bitstream.
12. The method according to claim 10, wherein: An SPPS index indicating the SPPS associated with the corresponding sub-picture is signaled in the bitstream.
13. The method according to any one of claims 1 to 12, wherein The bitstream comprises syntax elements associated with a slice, a tile, a tile, or a sub-picture, and wherein the syntax elements depend on a parameter set of a sub-picture comprising a current slice.
14. The method according to any one of claims 1 to 12, wherein A first control parameter of the parameter set controls a codec behavior, and wherein, based on a consistency rule, a second control parameter of the parameter set that controls the codec behavior is the same as the first control parameter.
15. The method according to claim 1, wherein An indication of application of one or more codec tools in a codec layer video sequence CLVS is signaled in the bitstream.
16. The method according to claim 15, wherein The indication is signaled in a Supplemental Enhancement Information SEI message or a Video Usage Information VUI message.
17. The method according to claim 15, wherein: The indication is signaled in a decoder parameter set DPS, a video parameter set VPS, a sequence parameter set SPS or an independent network abstraction layer NAL unit.
18. The method according to any one of claims 1 to 12, wherein The converting includes decoding the video from the bitstream.
19. The method according to any one of claims 1 to 12, wherein The converting includes encoding the video into the bitstream.
20. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, in, The current parameter set is configured to control the encoding and decoding behavior of at least one sub-picture of the one or more sub-pictures, wherein the bitstream complies with the format rules, The format rule specifies that the default parameter set corresponding to the current parameter set is signaled in the bitstream before a difference between the current parameter set and the default parameter set is signaled in the bitstream.
21. The method according to claim 20, wherein The format rule further specifies that after the default parameter set, the difference between the default parameter set and the current parameter set is signaled.
22. The method according to claim 20, wherein Before the default parameter set, a flag indicating that a coding behavior of each of the one or more sub-pictures is controlled by the default parameter set is signaled in the bitstream.
23. The method according to any one of claims 20 to 22, wherein The converting includes decoding the video from the bitstream.
24. The method according to any one of claims 20 to 22, wherein The converting includes encoding the video into the bitstream.
25. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, in, The bitstream complies with the format rules, The bitstream includes a parameter set, the parameter set including a first control parameter and a second control parameter for controlling the encoding and decoding properties of the sub-picture, and The format rule specifies whether or how the first control parameter is overwritten by the second control parameter for decoding.
26. The method according to claim 25, wherein The second control parameter is signaled in a video parameter set VPS, a sequence parameter set SPS, a picture parameter set PPS, a picture header, or a slice header.
27. The method according to claim 25 or 26, wherein The converting includes decoding the video from the bitstream.
28. The method according to claim 25 or 26, wherein The converting includes encoding the video into the bitstream.
29. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures, in, One or more first flags corresponding to each of the one or more sub-pictures are included in a sequence parameter set SPS, wherein each first flag of the one or more first flags indicates whether constraint information is signaled for the sub-picture corresponding to each first flag, and wherein the constraint information indicates a codec tool that is not applied to the corresponding sub-picture on the codec layer video sequence CLVS; wherein the bitstream complies with a format rule, the format rule specifying that the bitstream includes a parameter set, the parameter set controlling encoding and decoding behavior of a sub-picture of the one or more sub-pictures associated with an identifier (ID) of the sub-picture; The bitstream includes a single flag indicating that the parameter set for each of the one or more sub-pictures is the same.
30. The method according to claim 29, wherein A second flag is signaled in the bitstream indicating whether each of the one or more first flags is signaled.
31. The method according to claim 29 or 30, wherein The constraint information includes a general_constraint_info() syntax structure.
32. The method according to claim 29 or 30, wherein The converting includes decoding the video from the bitstream.
33. The method according to claim 29 or 30, wherein The converting includes encoding the video into the bitstream.
34. A method of writing a bitstream representing a video to a computer-readable recording medium, comprising: Generating a bitstream from a video according to the method of any one of claims 1 to 31; and The bit stream is written to a computer-readable recording medium.
35. A method for storing a bitstream of a video, comprising: performing conversion between a video comprising a picture and a bitstream of the video, the picture comprising one or more sub-pictures; generating the bitstream from the current block; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein the bitstream complies with the format rules, wherein the format rule specifies that the bitstream includes a parameter set, the parameter set controlling encoding and decoding behavior of a sub-picture of the one or more sub-pictures associated with an identification (ID) of the sub-picture; The bitstream includes a single flag indicating that the parameter set for each of the one or more sub-pictures is the same.
36. A video processing apparatus comprising a processor configured to implement the method of any one or more of claims 1 to 35.
37. A computer readable medium having stored thereon instructions which, when executed, cause a processor to implement the method of any one or more of claims 1 to 35.
38. A computer readable medium having stored therein a bitstream generated according to any one of claims 1 to 35.
39. A video processing device for storing a bit stream, wherein: The video processing device is configured to implement the method according to any one or more of claims 1 to 35.
Citation Information
Patent Citations
Methods and apparatuses for encoding and decoding video
CN107105302A
Tile alignment signaling and conformance constraints
US20160165247A1