Subpicture boundary filtering in video coding
By employing sub-image-based encoding and decoding techniques and temporal motion vector prediction in video encoding and decoding, and optimizing the filtering operation at sub-image boundaries, the problem of video quality degradation in existing technologies is solved, achieving more efficient video encoding and decoding results.
Patent Information
- Application Number
- CN202180009033.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-13
- Filing Date
- 2021-01-13
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-01-13
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively utilize information from juxtaposed images for filtering when processing sub-image boundaries, leading to a decline in video quality.
By employing sub-image-based encoding and decoding techniques during video encoding and decoding, and using regional constraints and temporal motion vector prediction, the filtering operations within the juxtaposed images are optimized, including the interpolation process of luminance and chrominance samples, thereby achieving accurate processing of sub-image boundaries.
It improves video decoding quality, enhances video coding efficiency and compression performance, especially in HEVC and future video codec standards, achieving higher bit rate reduction and image clarity.
Smart Images

Figure CN115280768B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] In accordance with the applicable Patent Law and / or the Paris Convention, this application promptly claims priority and interest in International Patent Application No. PCT / CN2020 / 071863, filed on January 13, 2020. For all legal purposes, the entire disclosure of the foregoing application is incorporated herein by reference as a part of this application disclosure. Technical Field
[0003] This document covers video and image encoding and decoding technologies. Background Technology
[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] The disclosed techniques can be used by video or image decoder or encoder embodiments, wherein sub-image-based encoding or decoding is performed.
[0006] In one example aspect, a video processing method is disclosed. The method includes a conversion between a current video block in a current frame of a video and the bitstream of that video, determining how to modify the y-coordinate yColSb of a juxtaposed sub-block within a juxtaposed image of the current frame based on whether the sub-frame is considered an image. The juxtaposed image is one of one or more reference images of the current frame. The method also includes performing the conversion based on this determination.
[0007] In another example, a video processing method is disclosed. The method includes a conversion between a current image of a video comprising at least two sub-images and a bitstream of the video, determining, based on information from the two sub-images, how to apply a filtering operation to a region covering the boundary between the two sub-images. The method also includes performing the conversion according to the determination.
[0008] In another example, a video processing method is disclosed. The method includes, for a video block in a first video region of a video, determining whether the location of a temporal motion vector prediction value determined by the transformation between the video block and the bitstream representation of the current video block using an affine mode is within a second video region; and performing a transformation based on this determination.
[0009] In another example, a different video processing method is disclosed. This method includes, for a video block in a first video region of the video, determining whether the location of an integer sample in a reference image extracted for the conversion between the bitstream representation of the video block and the current video block is within a second video region, wherein the reference image is not used for the interpolation process during the conversion; and performing the conversion based on this determination.
[0010] In another example, a different video processing method is disclosed. This method includes, for a video block in a first video region of a video, determining whether the location of the reconstructed luminance sample value extracted for the conversion between the video block and the current video block bitstream representation is within a second video region; and performing a conversion based on this determination.
[0011] In another example, a different video processing method is disclosed. This method includes, for a video block within a first video region of the video, determining whether the location of a video block partitioning-related check, depth derivation, or partition flag signaling notification performed during the conversion between the video block and the bitstream representation of the current video block is within a second video region; and performing a conversion based on this determination.
[0012] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a codec representation of that video, wherein the one or more video images comprise one or more video blocks, and the codec representation conforms to the codec syntax requirements that the conversion does not use sub-picture encoding / decoding within video units and dynamic precision conversion encoding / decoding tools or reference picture resampling tools.
[0013] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein the one or more video images comprise one or more video blocks, and the codec representation conforms to the codec syntax requirement that a first syntax element `subpic_grid_idx[i][j]` is not greater than a second syntax element `max_subpics_minus1`.
[0014] In another example, a different video processing method is disclosed. This method includes performing a conversion between a first video region of the video and a codec representation of the video, wherein a set of parameters defining the codec characteristics of the first video region is included at the first video region level in the codec representation.
[0015] In yet another example, the above method can be implemented by a video encoder device that includes a processor.
[0016] In yet another example, the above method can be implemented by a video decoder device that includes a processor.
[0017] In yet another example, these methods can be implemented as processor-executable instructions and stored on a computer-readable program medium.
[0018] These and other aspects are further described in this document. Attached Figure Description
[0019] Figure 1 Examples of temporal motion vector prediction (TMVP) and region constraints in sub-block TMVP are shown.
[0020] Figure 2 An example of a graded motion estimation scheme is shown.
[0021] Figure 3 This is a block diagram of an example hardware platform used to implement the technologies described in this document.
[0022] Figure 4 This is a flowchart of an example method for video processing.
[0023] Figure 5 An example of an image with an 18x12 brightness CTU is shown, which is divided into 12 slices and 3 raster scan strips (informative).
[0024] Figure 6 An example of an image with an 18x12 brightness CTU is shown, which is divided into 24 slices and 9 rectangular strips (informative).
[0025] Figure 7 An example of an image is shown, divided into 4 slices, 11 tiles, and 4 rectangular strips (informative).
[0026] Figure 8 An example of a block encoded in palette mode is shown.
[0027] Figure 9 An example of using a predicted value palette to signal palette entries is shown.
[0028] Figure 10 Examples of horizontal and vertical traversal scans are shown.
[0029] Figure 11 An example of encoding and decoding a palette index is shown.
[0030] Figure 12 An example of a Merge estimation region (MER) is shown.
[0031] Figure 13This is a block diagram illustrating an example video processing system in which various techniques disclosed herein can be implemented.
[0032] Figure 14 This is a block diagram illustrating an example video encoding / decoding system.
[0033] Figure 15 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0034] Figure 16 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0035] Figure 17 This is a flowchart representation of a video processing method based on this technology.
[0036] Figure 18 This is a flowchart representation of another method for video processing according to the present technology. Detailed Implementation
[0037] This document provides various techniques that decoders of image or video bitstreams can use to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" used in this document includes both sequences of images (traditionally referred to as video) and individual images. Furthermore, video encoders may also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0038] The chapter headings used in this document are for ease of understanding and do not limit the embodiments and techniques to the corresponding chapters. Thus, embodiments from one section can be combined with embodiments from other sections.
[0039] 1. Preliminary Discussion
[0040] This document relates to video codec technology. Specifically, it relates to palette codecs that use primary color-based representations in video codecs. It can be applied to existing video codec standards, such as HEVC, or upcoming standards (General Video Codec). It can also be applied to future video codec standards or video codecs.
[0041] 2. Introduction to Video Encoding and Decoding
[0042] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards [1, 2]. Since H.262, video codec standards have been based on hybrid video codec architectures, which utilize temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, VCEG (Q6 / 16) and ISO / IEC JTC1SC29 / WG11 (MPEG) established the Joint Video Experts Group (JVET), which is dedicated to the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.
[0043] 2.1 Region constraints in TMVP and sub-block TMVP in VVC
[0044] Figure 1 Example region constraints are shown in TMVP and sub-block TMVP. In TMVP and sub-block TMVP, as... Figure 1 As shown, the time-domain MV can only be obtained from the juxtaposed CTU plus a column of 4×4 blocks.
[0045] 2.2 Sub-image Example
[0046] In some embodiments, a sub-image-based encoding / decoding technique based on a flexible tiling method can be implemented. An overview of the sub-image-based encoding / decoding technique includes the following:
[0047] 1) An image can be divided into sub-images.
[0048] 2) In SPS, indicate the presence of sub-images and other sequence-level information about the sub-images.
[0049] 3) Whether sub-images are treated as images during the decoding process (excluding in-loop filtering operations) can be controlled by the bitstream.
[0050] 4) Whether to disable loop filtering across sub-image boundaries is controlled by the bitstream of each sub-image. The DBF, SAO, and ALF processes are updated to control the loop filtering operation across sub-image boundaries.
[0051] 5) For simplicity, as a starting point, the width, height, horizontal offset, and vertical offset of sub-images are expressed in units of luminance samples in SPS. Sub-image boundaries are constrained to strip boundaries.
[0052] 6) By slightly updating the coding_tree_unit() syntax, sub-images are treated as images during the decoding process (excluding in-loop filtering operations), and the following decoding process is updated:
[0053] – Derivation of (Advanced) Temporal Luminance Motion Vector Prediction
[0054] –Biolinear interpolation process for brightness samples
[0055] – Brightness sample 8-tap interpolation filtering process
[0056] – Colorimetric sample point interpolation process
[0057] 7) The sub-image ID is explicitly specified in SPS and included in the slice group header so that the sub-image sequence can be extracted without changing the VCL NAL unit.
[0058] 8) Propose an Output Subpicture Set (OSPS) to specify the canonical extraction and consistency points of the subpictures and their sets.
[0059] 2.3 Example Sub-images in General Video Encoding and Decoding
[0060] Sequence Parameter Set (RBSP) Syntax
[0061]
[0062]
[0063] A subpics_present_flag value of 1 indicates that a subpics parameter exists in the SPS RBSP syntax. A subpics_present_flag value of 0 indicates that a subpics parameter does not exist in the SPS RBSP syntax.
[0064] Note 2: When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of subpics of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPS.
[0065] Incrementing `max_subpics_minus1` by 1 specifies the maximum number of subpicks that may exist in CVS. `max_subpics_minus1` should be in the range of 0 to 254. The value 255 is reserved for future use by ITU-T|ISO / IEC.
[0066] `subpic_grid_col_width_minus1` incremented by 1 specifies the width of each element of the subpicture identifier grid, in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits.
[0067] The derivation of the variable NumSubPicGridCols is as follows:
[0068] NumSubPicGridCols=(pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*4+3) / (subpic_grid_col_width_minus1*4+4) (7-5)
[0069] `subpic_grid_row_height_minus1` incremented by 1 specifies the height of each element of the subpick identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.
[0070] The derivation of the variable NumSubPicGridRows is as follows:
[0071] NumSubPicGridRows=(pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*4+3) / (subpic_grid_row_height_minus1*4+4) (7-6)
[0072] `subpic_grid_idx[i][j]` specifies the subpick index at grid position (i,j). The length of the syntax element is Ceil(Log2(max_subpics_minus1+1)) bits.
[0073] The derivation of variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]], and NumSubPics is as follows:
[0074]
[0075]
[0076] A subpic_treated_as_pic_flag[i] equal to 1 indicates that the i-th subpic of each encoded / decoded image in CVS is considered as an image in the decoding process excluding loop filtering operations. A subpic_treated_as_pic_flag[i] equal to 0 indicates that the i-th subpic of each encoded / decoded image in CVS is not considered as an image in the decoding process excluding loop filtering operations. When it does not exist, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to 0.
[0077] A loop filter_across_subpic_enabled_flag[i] equal to 1 indicates that loop filtering can be performed across the boundary of the i-th subpic in each codec image of the CVS. A loop filter_cross_subpic_enabled_flag[i] equal to 0 indicates that loop filtering is not performed across the boundary of the i-th subpic in each codec image of the CVS. When it does not exist, the value of loop filter_cross_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0078] One requirement for bitstream consistency is the application of the following constraints:
[0079] For any two subpicks, subpicA and subpicB, if the index of subpicA is less than the index of subpicB, then, in the order of decoding, any NAL unit of subpicA will be after any NAL unit of subpicB.
[0080] – The shape of the sub-image should be such that, when decoded, the entire left and entire top boundaries of each sub-image should include the image boundary or consist of the boundaries of the previously decoded sub-images.
[0081] The list CtbToSubPicIdx[ctbAddrRs] specifies the conversion from the CTB address based on the image raster scan to the sub-image index, where ctbAddrRs ranges from 0 to PicSizeInCtbsY–1, inclusive, and its derivation is as follows:
[0082]
[0083]
[0084] `num_bricks_in_slice_minus1`, if present, decrements the number of bricks in the specified strip by 1. The value of `num_bricks_in_slice_minus1` should be in the range of 0 to `NumBricksInPic-1`, inclusive. When `rect_slice_flag` is 0 and `single_brick_per_slice_flag` is 1, the value of `num_bricks_in_slice_minus1` is inferred to be 0. When `single_brick_per_slice_flag` is 1, the value of `num_bricks_in_slice_minus1` is inferred to be 0.
[0085] The variable NumBricksInCurrSlice specifies the number of tiles in the current slice, and SliceBrickIdx[i] specifies the tile index of the i-th tile in the current slice, which is derived as follows:
[0086]
[0087] The derivation of variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos is as follows:
[0088]
[0089] Derivation of temporal brightness motion vector prediction
[0090] The input to this process is:
[0091] – The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.
[0092] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0093] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0094] – Refer to the index refIdxLX, where X is 0 or 1.
[0095] The output of this process is:
[0096] Motion vector prediction with a precision of -1 / 16 fractional sample points, mvLXCol
[0097] –Availability flag: availableFlagLXCol.
[0098] The variable currCb specifies the current luminance codec block at the luminance position (xCb, yCb).
[0099] The derivation of variables mvLXCol and availableFlagLXCol is as follows:
[0100] – If slice_temporal_MVP_enabled_flag is equal to 0 or (cbWidth*cbHeight) is less than or equal to 32, then both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0101] Otherwise (slice_temporal_MVP_enabled_flag equals 1), apply the following ordered steps:
[0102] 1. The derivation of the juxtaposed motion vector in the lower right corner and the positions of the bottom and right boundary sample points is as follows:
[0103] xColBr=xCb+cbWidth (8-421)
[0104] yColBr=yCb+cbHeight (8-422)
[0105] rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?
[0106] SubPicRightBoundaryPos:pic_width_in_luma_samples-1 (8-423)
[0107] botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos:pic_height_in_luma_samples-1 (8-424)
[0108] 2. If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, yColBr is less than or equal to bottomBoundaryPos, and xColBr is less than or equal to rightBoundaryPos, then the following applies:
[0109] – The variable colCb specifies the luminance codec block that overrides the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the juxtaposed image specified by ColPic.
[0110] – The luminance position (xColCb, yColCb) is set to be equal to the top left sample of the juxtaposed luminance codec module specified by ColCb relative to the top left sample of the juxtaposed image specified by ColPic.
[0111] – The derivation of the juxtaposed motion vectors as specified in Clause 8.5.2.12 is invoked, where currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag are set to 0 as inputs, and the outputs are assigned to mvLXCol and availableFlagLXCol.
[0112] Otherwise, both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0113] …
[0114] Brightness sample bilinear interpolation process
[0115] The input to this process is:
[0116] – Brightness position in units of the entire sample (xInt) L ,yInt L ),
[0117] – Brightness position in fractional samples (xFrac) L ,yFrac L ),
[0118] –Luminance reference sample array refPicLX L .
[0119] The output of this process is the predicted luminance sample value, predSampleLX. L
[0120] The derivation of variables shift1, shift2, shift3, shift4, offset1, offset2, and offset3 is as follows:
[0121] shift1 = BitDepth Y -6 (8-453)
[0122] offset1=1<<(shift1-1) (8-454)
[0123] shift2 = 4 (8-455)
[0124] offset2=1<<(shift2-1) (8-456)
[0125] shift3 = 10-BitDepth Y (8-457)
[0126] shift4 = BitDepth Y -10 (8-458)
[0127] offset4=1<<(shift4-1) (8-459)
[0128] The variable picW is set to equal pic_width_in_luma_samples, and the variable picH is set to equal pic_height_in_luma_samples.
[0129] The brightness interpolation filter coefficients fb at each 1 / 16 fractional sample location p L [p] equals xFrac L or yFrac L As specified in Table 8-10.
[0130] For i = 0..1, the brightness position (xInt) in units of the entire sample point. i yInt i The derivation of ) is as follows:
[0131] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0132] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i) (8-460)
[0133] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i) (8-461)
[0134] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[0135] xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?
[0136]
[0137] yInt i =Clip3(0,picH-1,yInt) L +i) (8-463)
[0138] …
[0139] Derivation of the Temporal Merge Candidate Based on Sub-Blocks
[0140] The input to this process is:
[0141] – The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.
[0142] – The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0143] The variable cbHeight specifies the height of the current codec block in units of luminance samples.
[0144] –Availability flag for adjacent codec units, availableFlagA1
[0145] – The reference index of adjacent encoding / decoding units is refIdxLXA1.
[0146] – The prediction list of adjacent codec units uses the flag predFlagLXA1, where X is 0 or 1.
[0147] – Motion vector mvLXA1 of adjacent codec units with a precision of 1 / 16 fractional sample point, where X is 0 or 1.
[0148] The output of this process is:
[0149] –Availability flag availableFlagSbCol
[0150] – The number of luminance codec sub-blocks in the horizontal direction, numSbX, and the number of luminance codec sub-blocks in the vertical direction, numSbY.
[0151] –Refer to the indices refIdxL0SbCol and refIdxL1SbCol.
[0152] The brightness motion vectors mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx] with a precision of –1 / 16 fractional samples, xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY–1.
[0153] – Predict the utilization flags predFlagL0SbCol[xSbIdx][ySbIdx] and predFlagL1SbCol[xSbIdx][ySbIdx], xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY–1.
[0154] The derivation of the availability flag availableFlagSbCol is as follows.
[0155] – AvailableFlagSbCol is set to 0 if one or more of the following conditions are true.
[0156] –slice_temporal_mvp_enabled_flag equals 0.
[0157] –sps_sbtmvp_enabled_flag equals 0.
[0158] –cbWidth is less than 8.
[0159] –cbHeight is less than 8.
[0160] Otherwise, apply the following ordered steps:
[0161] 1. The derivation of the positions (xCtb, yCtb) of the top-left sample point of the current luma code block and the position (xCtr, yCtr) of the bottom-right center sample point of the current luma code block is as follows:
[0162] xCtb=(xCb>>CtuLog2Size)< <CtuLog2Size (8-542)
[0163] yCtb=(yCb>>CtuLog2Size)< <CtuLog2Size (8-543)
[0164] xCtr=xCb+(cbWidth / 2) (8-544)
[0165] yCtr=yCb+(cbHeight / 2) (8-545)
[0166] 2. The luminance position (xColCtrCb, yColCtrCb) is set to be equal to the top left sample of the juxtaposed luminance codec module covering the position given by (xCtr, yCtr) within ColPic relative to the top left luminance sample of the juxtaposed image specified by ColPic.
[0167] 3. Invoke the derivation process of the basic motion data of the temporal Merge based on sub-blocks specified in Clause 8.5.5.4, taking the positions (xCtb, yCtb), (xColCtrCb, yColCtrCb), availability flag availableFlagA1, prediction list utilization flag predFlagLXA1, reference index refIdxLXA1, and motion vector mvLXA1 as inputs, where X is 0 and 1, and taking the motion vector ctrMvLX, the prediction list utilization flag ctrPredFlagLX of the juxtaposed block, and the temporal vector tempMv as outputs, where X is 0 and 1.
[0168] 4. The derivation of the variable availableFlagSbCol is as follows:
[0169] – If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, then availableFlagSbCol is set to equal to 0.
[0170] Otherwise, availableFlagSbCol is set to 1.
[0171] When availableFlagSbCol equals 1, the following applies:
[0172] The derivation of variables numSbX, numSbY, sbWidth, sbHeight, and refIdxLXSbCol is as follows:
[0173] numSbX=cbWidth>>3 (8-546)
[0174] numSbY=cbHeight>>3 (8-547)
[0175] sbWidth=cbWidth / numSbX (8-548)
[0176] sbHeight=cbHeight / numSbY (8-549)
[0177] refIdxLXSbCol=0 (8-550)
[0178] – For xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, the derivation of the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list using the flag predFlagLXSbCol[xSbIdx][ySbIdx] is as follows:
[0179] – Specifies the brightness position (xSb, ySb) of the top-left sample of the current encoding / decoding sub-block relative to the top-left luminance sample of the current image. Its derivation is as follows:
[0180] xSb=xCb+xSbIdx*sbWidth+sbWidth / 2 (8-551)
[0181] ySb=yCb+ySbIdx*sbHeight+sbHeight / 2 (8-552)
[0182] The derivation of the positions (xColSb, yColSb) of the juxtaposed sub-blocks within –ColPic is as follows.
[0183] –Applicable to the following:
[0184]
[0185] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0186]
[0187] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[0188]
[0189] …
[0190] Derivation process of temporal Merge basic motion data based on sub-blocks
[0191] The input to this process is:
[0192] – The position of the top left sample of the luminance code tree block containing the current code block (xCtb, yCtb), – The position of the top left sample of the juxtaposed luminance code block covering the bottom right center sample (xColCtrCb, yColCtrCb).
[0193] –Availability flag for adjacent codec units, availableFlagA1
[0194] – The reference index of adjacent encoding / decoding units is refIdxLXA1.
[0195] – The prediction list of adjacent codec units uses the flag predFlagLXA1.
[0196] – Motion vector mvLXA1 with 1 / 16 fractional sample precision of adjacent codec units.
[0197] The output of this process is:
[0198] – Motion vectors ctrMvL0 and ctrMvL1
[0199] – The prediction list uses the flags ctrPredFlagL0 and ctrPredFlagL1.
[0200] –Time-domain motion vector tempMv.
[0201] The variable tempMv is set as follows:
[0202] tempMv[0]=0 (8-558)
[0203] tempMv[1]=0 (8-559)
[0204] The variable currPic specifies the current image.
[0205] When availableFlagA1 equals TRUE, the following applies:
[0206] – If all of the following conditions are true, then tempMv is set to equal mvL0A1:
[0207] –predFlagL0A1 equals 1.
[0208] –DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A1]) equals 0,
[0209] Otherwise, if all of the following conditions are true, then tempMv is set to equal mvL1A1:
[0210] –slice_type equals B,
[0211] –predFlagL1A1 equals 1,
[0212] –DiffPicOrderCnt(ColPic,RefPicList[1][refIdxL1A1]) equals 0.
[0213] The derivation of the positions (xColCb, yColCb) of the juxtaposed blocks within ColPic is as follows.
[0214] –Applicable to the following:
[0215]
[0216] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0217]
[0218] – Otherwise, if subpic_treated_as_pic_flag[SubPicIdx] equals 0, then the following applies:
[0219]
[0220] …
[0221] Brightness sample interpolation filtering process
[0222] The input to this process is:
[0223] – Brightness position in units of the entire sample (xInt) L yInt L ),
[0224] – Brightness position in fractional samples (xFrac) L yFrac L ),
[0225] – Brightness position in units of the entire sample (xSbInt) L ySbInt L This specifies the top-left sample point of the boundary block used for reference sample filling relative to the top-left brightness sample point of the reference image.
[0226] –Luminance reference sample array refPicLX L ,
[0227] – Half-sample interpolation filter index hpelIfIdx,
[0228] – The variable sbWidth specifies the width of the current child block.
[0229] – Specifies the height of the current child block, sbHeight.
[0230] – Specifies the brightness position (xSb, ySb) of the top-left sample point of the current sub-block relative to the top-left luminance sample point of the current image.
[0231] The output of this process is the predicted luminance sample value, predSampleLX. L
[0232] The derivation of variables shift1, shift2, and shift3 is as follows:
[0233] – Set the variable shift1 to equal Min(4, BitDepth) Y -8), variable shift2 is set to equal to 6, and variable shift3 is set to equal to Max(2, 14-BitDepth). Y ).
[0234] – Set the variable picW to equal pic_width_in_luma_samples, and set the variable picH to equal pic_height_in_luma_samples.
[0235] The brightness interpolation filter coefficients f at each 1 / 16 fractional sample location p L [p] equals xFrac L or yFrac L The derivation is as follows:
[0236] – If MotionModelIdc[xSb][ySb] is greater than 0, and sbWidth and sbHeight are both equal to 4, then the luminance interpolation filter coefficient f L [p] is specified in Table 8-12.
[0237] Otherwise, the brightness interpolation filter coefficients f L [p] is specified in Table 8-11, which depends on hpelIfIdx.
[0238] For i = 0..7, the brightness position (xInt) in units of the entire sample point. i yInt i The derivation of ) is as follows:
[0239] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following conditions apply:
[0240] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3) (8-771)
[0241] yInt i=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772)
[0242] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following conditions apply:
[0243] xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?
[0244]
[0245] yInt i =Clip3(0,picH-1,yInt) L +i-3)
[0246] (8-774)
[0247] …
[0248] Chromaticity sample interpolation process
[0249] The input to this process is:
[0250] - Chromaticity position in units of the entire sample (xInt) C yInt C ),
[0251] – Chromaticity position in units of 1 / 32 fractional samples (xFrac) C yFrac C ),
[0252] – Chromaticity position in units of full sample points (xSbIntC, ySbIntC), which specifies the top-left sample point of the boundary block used for reference sample point filling relative to the top-left chromaticity sample point of the reference image.
[0253] – The variable sbWidth specifies the width of the current child block.
[0254] – Specifies the height of the current child block, sbHeight.
[0255] –Color reference sample array refPicLX C .
[0256] The output of this process is the predicted chromaticity sample value, predSampleLX. C
[0257] The derivation of variables shift1, shift2, and shift3 is as follows:
[0258] – Set the variable shift1 to equal Min(4, B BitDepth) C -8), variable shift2 is set to equal to 6, and variable shift3 is set to equal to Max(2, 14-BitDepth). C ).
[0259] – picW C Set to equal pic_width_in_luma_samples / SubWidthC, variable picH C Set it to equal pic_height_in_luma_samples / SubHeightC.
[0260] The chromaticity interpolation filter coefficients f at each 1 / 32 fractional sample location p C [p] equals xFrac C or yFrac C , is specified in Table 8-13.
[0261] The variable xOffset is set to equal to (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC.
[0262] For i = 0..3, the chromaticity position (xInt) in units of the entire sample point. i ,yInt i The derivation of ) is as follows:
[0263] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following conditions apply:
[0264] xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (8-785)
[0265] yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (8-786)
[0266] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following conditions apply:
[0267]
[0268] yInt i =Clip3(0,picH C -1,yInt C +i-1)
[0269] (8-788)
[0270] 2.4 Example: Encoder-only GOP-based temporal filter
[0271] In some embodiments, a temporal filter for the encoder only can be implemented. As a preprocessing step, filtering is performed at the encoder end. Source images before and after the selected image to be encoded are read, and a block-based motion compensation method relative to the selected image is applied to these source images. Temporal filtering is then performed on the samples in the selected image using the motion-compensated sample values.
[0272] The overall filtering strength is set based on the temporal sublayer and QP of the selected image. Only images in temporal sublayers 0 and 1 are filtered, and images in layer 0 are filtered with a stronger filter than images in layer 1. The filtering strength per sample is adjusted based on the difference between the sample values in the selected image and the juxtaposed samples in the motion-compensated image, so that small differences between the motion-compensated image and the selected image are filtered more strongly than larger differences.
[0273] GOP-based time-domain filters
[0274] A temporal filter is introduced directly after image reading and before encoding. The steps are described in more detail below.
[0275] Operation 1: Read the image by the encoder
[0276] Operation 2: If the image is low enough in the encoding / decoding hierarchy, it is filtered before encoding. Otherwise, the image is encoded without filtering. RA images with POC%8==0 and LD images with POC%4==0 are filtered. AI images are never filtered.
[0277] The overall filter strength is set for RA according to the following formula.
[0278]
[0279] Where n is the number of images read.
[0280] For the LD case, use so (n) = 0.95.
[0281] Operation 3: Read the two images before and / or after the selected image (hereinafter referred to as the original image). In edge cases, such as if it is the first image or close to the last image, only the available images are read.
[0282] Operation 4: For each 8×8 image block, estimate the motion before and after reading the image relative to the original image.
[0283] Using a hierarchical motion estimation scheme, and Figure 2 Layers L0, L1, and L2 are shown in the diagram. This is achieved by examining all read images and the original image (i.e., ...). Figure 1 The subsampled image is generated by averaging each 2×2 block of L1. L2 is obtained from L1 using the same subsampling method.
[0284] Figure 2 Examples of different layers in hierarchical motion estimation are shown. L0 is the original precision. L1 is a subsampled version of L0. L2 is a subsampled version of L1.
[0285] First, motion estimation is performed for each 16×16 block in L2. The squared difference is calculated for each selected motion vector, and the motion vector corresponding to the minimum difference is chosen. Then, when estimating motion in L1, the chosen motion vector is used as the initial value. The same operation is then performed for motion estimation in L0. As a final step, sub-pixel motion for each 8×8 block is estimated using an interpolation filter on L0.
[0286] Using a VTM 6-tap interpolation filter:
[0287]
[0288]
[0289] Step 5: Apply motion compensation to the images before and after the original image based on the best-matching motion for each block. That is, ensure that the sample coordinates of the original image in each block have the best-matching coordinates in the reference image.
[0290] Step 6: Process the samples of the luminance and chroma channels one by one as described in the following steps.
[0291] Operation 7: Calculate the new sample value I using the following formula. n .
[0292]
[0293] Where I o It is the sample value of the original sample point, Ir (i) is the intensity of the corresponding sample point of motion-compensated image i, and w r (i,a) is the weight of motion-compensated image i when the number of available motion-compensated images is a.
[0294] In the luminance channel, the weight w r (i, a) is defined as follows:
[0295]
[0296] in
[0297] s l =0.4
[0298]
[0299]
[0300] For all other cases of i and a: s r (i,a)=0.3
[0301] σ l (QP) = 3 * (QP - 10)
[0302] ΔI(i)=I r (i)-I o
[0303] For the chroma channel, the weight w r (i, a) is defined as follows:
[0304]
[0305] Where s c =0.55 and σ c =30.
[0306] Operation 8: Apply the filter to the current sample. The resulting sample values are stored separately.
[0307] Operation 9: Encode the filtered image.
[0308] 2.5 Example Image Segmentation (Pieces, Patches, Strips)
[0309] In some embodiments, the image is divided into one or more slice rows and one or more slice columns. A slice is a series of CTUs that cover a rectangular area of the image.
[0310] The slice is divided into one or more tiles, and each tile includes multiple CTU rows within the slice.
[0311] A slice that is not divided into multiple tiles is also called a tile. However, a tile that is a proper subset of a slice is not called a slice.
[0312] A strip or multiple slices containing images, or multiple tiles containing slices.
[0313] A sub-image contains one or more stripes that collectively cover a rectangular area of the image.
[0314] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices raster scanned according to the image's slices. In rectangular stripe mode, the stripe contains multiple tiles of the image, which together form a rectangular area of the image. The tiles within the rectangular stripe are arranged according to the raster scan order of the stripe's tiles.
[0315] Figure 5 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.
[0316] Figure 6 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0317] Figure 7 An example is shown where an image is divided into slices, tiles, and rectangular strips, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the top left slice contains 1 tile, the top right slice contains 5 tiles, the bottom left slice contains 2 tiles, and the bottom right slice contains 3 tiles) and 4 rectangular strips.
[0318] Image Parameter Set RBSP Syntax
[0319]
[0320]
[0321]
[0322]
[0323] The single_tile_in_pic_flag value being 1 indicates that there is only one tile in each picture of the reference PPS.
[0324] A single_tile_in_pic_flag value of 0 indicates that each picture in the reference PPS contains more than one tile.
[0325] Note – If there are no further tile divisions within a piece, the entire piece is called a tile. When an image contains only a single piece without further tile divisions, it is called a single tile.
[0326] The requirement for bitstream consistency is that the value of single_tile_in_pic_flag should be the same for all PPS referenced by the CVS in-process encoding and decoding images.
[0327] A `uniform_tile_spacing_flag` of 1 specifies that tile column and row boundaries are uniformly distributed across the image, and is signaled using the syntax elements `tile_cols_width_minus1` and `tile_rows_height_minus1`. A `uniform_tile_spacing_flag` of 0 specifies that tile column and row boundaries may be uniformly or non-uniformly distributed across the image, and is signaled using the syntax elements `num_tile_columns_minus1` and `num_tile_rows_minus1`, as well as a list of syntax elements `tile_column_width_minus1[i]` and `tile_row_height_minus1[i]`. When `uniform_tile_spacing_flag` does not exist, its value is inferred to be 1.
[0328] Incrementing `tile_cols_width_minus1` by 1 specifies the width, in CTB, of all tile columns in the image except the rightmost one, when `uniform_tile_spacing_flag` is 1. The value of `tile_cols_width_minus1` should be between 0 and `PicWidthInCtbsY-1`, inclusive. If it does not exist, the value of `tile_cols_width_minus1` is inferred to be equal to `PicWidthInCtbsY-1`.
[0329] Incrementing `tile_rows_height_minus1` by 1 specifies the height of all tile rows in the image, excluding the bottom tile rows, in CTB units when `uniform_tile_spacing_flag` is equal to 1. The value of `tile_rows_height_minus1` should be between 0 and `PicHeightInCtbsY-1`, inclusive. If it does not exist, the value of `tile_rows_height_minus1` is inferred to be equal to `PicHeightInCtbsY-1`.
[0330] `num_tile_columns_minus1` incremented by 1 specifies the number of tile columns to segment the image when `uniform_tile_spacing_flag` equals 0. The value of `num_tile_columns_minus1` should be in the range of 0 to `PicWidthInCtbsY-1`, inclusive. If `single_tile_in_pic_flag` equals 1, then the value of `num_tile_columns_minus1` is inferred to be equal to 0. Otherwise, when `uniform_tile_spacing_flag` equals 1, the value of `num_tile_columns_minus1` is inferred as specified in Clause 6.5.1.
[0331] `num_tile_rows_minus1` incremented by 1 specifies the number of tile rows to segment the image when `uniform_tile_spacing_flag` equals 0. The value of `num_tile_rows_minus1` should be in the range of 0 to `PicHeightInCtbsY-1`, inclusive. If `single_tile_in_pic_flag` equals 1, then the value of `num_tile_rows_minus1` is inferred to be 0. Otherwise, when `uniform_tile_spacing_flag` equals 1, the value of `num_tile_rows_minus1` is inferred as specified in Clause 6.5.1.
[0332] The variable NumTilesInPic is set to equal to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1).
[0333] When single_tile_in_pic_flag equals 0, NumTilesInPic should be greater than 1.
[0334] tile_column_width_minus1[i] incremented by 1 specifies the width of the i-th tile column, in CTB units.
[0335] tile_row_height_minus1[i] incremented by 1 specifies the height of the i-th tile row, in CTB units.
[0336] A `brick_splitting_present_flag` value of 1 indicates that one or more slices of a reference PPS image can be split into two or more tiles. A `brick_splitting_present_flag` value of 0 indicates that slices of images without a reference PPS are split into two or more tiles.
[0337] `num_tiles_in_pic_minus1` plus 1 specifies the number of tiles in each picture of the reference PPS. The value of `num_tiles_in_pic_minus1` should be equal to `NumTilesInPic – 1`. If it does not exist, the value of `num_tiles_in_pic_minus1` is inferred to be equal to `NumTilesInPic – 1`.
[0338] A `brick_split_flag[i]` value of 1 indicates that the i-th piece is split into two or more tiles. A `brick_split_flag[i]` value of 0 indicates that the i-th piece is not split into two or more tiles. When it does not exist, the value of `brick_split_flag[i]` is inferred to be 0. [Ed.(HD / YK): SPS-dependent PPS parsing is introduced by adding the syntax condition "if(RowHeight[i]>1)". The same applies to `uniform_brick_spacing_flag[i]`.]
[0339] The uniform_brick_spacing_flag[i] being equal to 1 indicates that the horizontal tile boundaries are uniformly distributed on the i-th tile, and the syntax element brick_height_minus1[i] is used for signaling notification.
[0340] The uniform_brick_spacing_flag[i] being equal to 0 indicates that the horizontal tile boundaries may or may not be uniformly distributed on the i-th tile, and is signaled using a list of syntax elements num_brick_rows_minus2[i] and brick_row_height_minus1[i][j]. When it does not exist, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1.
[0341] Incrementing `brick_height_minus1[i]` by 1 specifies the height of the tile row in the i-th slice, excluding the bottom tile, in CTB units, when `uniform_brick_spacing_flag[i]` equals 1. If present, the value of `brick_height_minus1` should be in the range of 0 to `RowHeight[i] - 2`, inclusive. If absent, the value of `brick_height_minus1[i]` is inferred to be equal to `RowHeight[i] - 1`.
[0342] `num_brick_rows_minus2[i]` plus 2 specifies the number of tiles to split the i-th slice when `uniform_brick_spacing_flag[i]` equals 0. When present, the value of `num_brick_rows_minus2[i]` should be in the range of 0 to `RowHeight[i] - 2`, inclusive. If `brick_split_flag[i]` equals 0, then the value of `num_brick_rows_minus2[i]` is inferred to be -1. Otherwise, when `uniform_brick_spacing_flag[i]` equals 1, the value of `num_brick_rows_minus2[i]` is inferred as specified in Clause 6.5.1.
[0343] The value of brick_row_height_minus1[i][j] plus 1 specifies the height of the j-th tile in the i-th tile when uniform_tile_spacing_flag is equal to 0, in CTB units.
[0344] Derive the following variables, and when uniform_tile_spacing_flag equals 1, derive the values of num_tile_columns_minus1 and num_tile_rows_minus1, and for each i in the range from 0 to NumTilesInPic-1 (inclusive), when uniform_brick_spacing_flag[i] equals 1, infer the value of num_brick_rows_minus2[i] by calling the CTB raster and tile scan conversion procedure specified in Item 6.5.1:
[0345] – The list RowHeight[j] specifies the height of the j-th tile row in CTB, where j ranges from 0 to num_tile_rows_minus1, including the end value.
[0346] The list CtbAddrRsToBs[ctbAddrRs] specifies the conversion from CTB addresses in the image's CTB raster scan to CTB addresses in the tile scan, where ctbAddrRs ranges from 0 to PicSizeInCtbsY-1, including the end value.
[0347] The list CtbAddrBsToRs[ctbAddrBs] specifies the conversion from CTB addresses in a tile scan to CTB addresses in a picture raster scan, where ctbAddrRs ranges from 0 to PicSizeInCtbsY-1, including the end value.
[0348] – The list BrickId[ctbAddrBs] specifies the conversion from CTB addresses in a tile scan to tile IDs, where ctbAddrBs ranges from 0 to PicSizeInCtbsY-1, including end values.
[0349] The list NumCtusInBrick[brickIdx] specifies the conversion from tile index to the number of CTUs in the tile, where brickIdx ranges from 0 to NumBricksInPic–1, including the end value.
[0350] The list FirstCtbAddrBs[brickIdx] specifies the translation from the tile ID to the CTB address in the first CTB in the tile during a tile scan. The range of brickIdx is from 0 to NumBricksInPic–1, including the end value.
[0351] A single_brick_per_slice_flag value of 1 indicates that each slice of the reference PPS includes one tile. A single_brick_per_slice_flag value of 0 indicates that a slice of the reference PPS may include more than one tile. When it does not exist, the value of single_brick_per_slice_flag is inferred to be equal to 1.
[0352] A `rect_slice_flag` value of 0 indicates that the tiles within each slice are in the raster scan order and that slice information is not signaled in the PPS. A `rect_slice_flag` value of 1 indicates that the tiles within each slice cover a rectangular area of the image and that slice information is signaled in the PPS. When `brick_splitting_present_flag` is 1, the value of `rect_slice_flag` should be 1. If it does not exist, `rect_slice_flag` is inferred to be 1.
[0353] `num_slices_in_pic_minus1` incremented by 1 specifies the number of slices in each image of the reference PPS. The value of `num_slices_in_pic_minus1` should be in the range of 0 to `NumBricksInPic-1`, inclusive. When it does not exist and `single_brick_per_slice_flag` is equal to 1, the value of `num_slices_in_pic_minus1` is inferred to be equal to `NumBricksInPic-1`.
[0354] The increment of 1 in bottom_right_brick_idx_length_minus1 specifies the number of bits used to represent the syntax element bottom_right_brick_idx_delta[i].
[0355] The value of bottom_right_brick_idx_length_minus1 should be in the range of 0 to Ceil(Log2(NumBricksInPic))-1, inclusive.
[0356] `bottom_right_brick_idx_delta[i]`, when `i` is greater than 0, specifies the difference between the tile index of the bottom right corner of the `i`th strip and the tile index of the bottom right corner of the `(i-1)`th strip. `bottom_right_brick_idx_delta[0]` specifies the tile index of the bottom right corner of the `0`th strip. When `single_brick_per_slice_flag` equals 1, the value of `bottom_right_brick_idx_delta[i]` is inferred to be equal to 1. The value of `BottomRightBrickIdx[num_slices_in_pic_minus1]` is inferred to be equal to `NumBricksInPic-1`. The length of the `bottom_right_brick_idx_delta[i]` syntax element is `bottom_right_brick_idx_length_minus1+1` bits.
[0357] Increasing `brick_idx_delta_sign_flag[i]` by 1 indicates a positive sign for `bottom_right_brick_idx_delta[i]`. Setting `sign_bottom_right_brick_idx_delta[i]` to 0 indicates a negative sign for `bottom_right_brick_idx_delta[i]`.
[0358] The requirement for bitstream consistency is that a stripe should either consist of multiple complete slices or a continuous sequence of complete tiles from a single slice.
[0359] The variables TopLeftBrickIdx[i], BottomRightBrickIdx[i], NumBricksInSlice[i], and BricksToSliceMap[j] specify the tile index of the top-left corner tile in the i-th strip, the tile index of the bottom-right corner tile in the i-th strip, the number of tiles in the i-th strip, and the mapping from tiles to stripes. Their derivation is as follows:
[0360]
[0361] General strip header semantics
[0362] When present, the value of each of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and slice_temporal_mvp_enabled_flag should be the same in all slice headers for encoding and decoding images.
[0363] The variable CuQpDeltaVal specifies the difference between the luminance quantization parameter of the codec unit containing cu_qp_delta_abs and its prediction; this variable is set to 0. The variable CuQpOffset... Cb CuQpOffset Cr and CuQpOffset CbCr Specify the Qp' in the codec unit containing cu_chroma_qp_offset_flag. Cb 、Qp' Cr and Qp' CbCr The values to be used when quantizing the corresponding values of the parameters are all set to 0.
[0364] `slice_pic_parameter_set_id` specifies the value of `pps_pic_parameter_set_id` for the PPS being used. The value of `slice_pic_parameter_set_id` should be in the range of 0 to 63, inclusive.
[0365] The requirement for bitstream consistency is that the TemporalId value of the current image should be greater than or equal to the TemporalId value of the PPS whose pps_pic_parameter_set_id is equal to slice_pic_parameter_set_id.
[0366] `slice_address` specifies the slice address. If it does not exist, the value of `slice_address` is inferred to be 0.
[0367] If rect_slice_flag equals 0, then the following applies:
[0368] – The stripe address is the tile ID specified by equation (7-59).
[0369] The length of –slice_address is Ceil(Log2(NumBricksInPic)) bits.
[0370] The value of –slice_address should be in the range of 0 to NumBricksInPic-1, inclusive.
[0371] Otherwise (rect_slice_flag equals 1), the following applies:
[0372] – The stripe address is the stripe ID of the stripe.
[0373] The length of –slice_address is signalled_slice_id_length_minus1+1 bits.
[0374] – If `signalled_slice_id_flag` equals 0, then the value of `slice_address` should be in the range of 0 to `num_slices_in_pic_minus1`, inclusive. Otherwise, the value of `slice_address` should be in the range of 0 to 2. (signalled _slice_id_length_minus1+1) The range is -1, including the endpoints.
[0375] Bitstream consistency requirements are based on the following constraints:
[0376] The value of –slice_address should not be equal to the slice_address value of any other codec strip NAL unit of the same codec image.
[0377] – When rect_slice_flag equals 0, the image slices will be sorted in ascending order of their slice_address values.
[0378] – The shape of the image stripes should be such that, when decoded, each tile should have a full left boundary and a full top boundary consisting of the image boundary or the boundaries of the previously decoded (multiple) tiles.
[0379] `num_bricks_in_slice_minus1`, if present, specifies the number of tiles in the slice minus 1. The value of `num_bricks_in_slice_minus1` should be in the range of 0 to `NumBricksInPic-1`, inclusive. When `rect_slice_flag` is equal to 0 and `single_brick_per_slice_flag` is equal to 1, the value of `num_bricks_in_slice_minus1` is inferred to be 0. When `single_brick_per_slice_flag` is equal to 1, the value of `num_bricks_in_slice_minus1` is inferred to be 0.
[0380] The variable NumBricksInCurrSlice specifies the number of tiles in the current slice, and SliceBrickIdx[i] specifies the tile index of the i-th tile in the current slice, which is derived as follows:
[0381]
[0382]
[0383] The derivation of variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos is as follows:
[0384]
[0385] 2.6 Example Syntax and Semantics
[0386] Sequence Parameter Set (RBSP) Syntax
[0387]
[0388]
[0389]
[0390]
[0391]
[0392]
[0393]
[0394]
[0395] Image Parameter Set RBSP Syntax
[0396]
[0397]
[0398]
[0399]
[0400]
[0401] Image header RBSP syntax
[0402]
[0403]
[0404]
[0405]
[0406]
[0407]
[0408]
[0409]
[0410] A subpics_present_flag value of 1 indicates that a subpics parameter exists in the SPS RBSP syntax. A subpics_present_flag value of 0 indicates that a subpics parameter does not exist in the SPS RBSP syntax.
[0411] Note 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of subpics of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPSs.
[0412] Incrementing sps_num_subpics_minus1 by 1 specifies the number of subpicks. sps_num_subpics_minus1 should be in the range of 0 to 254. If it does not exist, the value of sps_num_subpics_minus1 is inferred to be equal to 0.
[0413] `subpic_ctu_top_left_x[i]` specifies the horizontal position of the top-left corner (CTU) of the i-th subpicture, in units of `CtbSizeY`. The length of the syntax element is... Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) Bit. When not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0414] `subpic_ctu_top_left_y[i]` specifies the vertical position of the top-left corner (CTU) of the i-th subpicture, in units of `CtbSizeY`. The length of the syntax element is... Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) Bit. When not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0415] `subpic_width_minus1[i]` incremented by 1 specifies the width of the i-th subpicture, in units of `CtbSizeY`. The length of this syntax element is `Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY))` bits. If it does not exist, the value of `subpic_width_minus1[i]` is inferred to be equal to... Ceil(pic_width_max_in_luma_samples / CtbSizeY) -1.
[0416] `subpic_height_minus1[i]` incremented by 1 specifies the height of the i-th subpicture, in units of `CtbSizeY`. The length of this syntax element is `Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY))` bits. If it does not exist, the value of `subpic_height_minus1[i]` is inferred to be equal to... Ceil(pic_height_max_in_luma_samples / CtbSizeY) -1.
[0417] A subpic_treated_as_pic_flag[i] equal to 1 indicates that the i-th subpic of each encoded / decoded image in CVS is considered as an image in the decoding process excluding loop filtering operations. A subpic_treated_as_pic_flag[i] equal to 0 indicates that the i-th subpic of each encoded / decoded image in CVS is not considered as an image in the decoding process excluding loop filtering operations. When it does not exist, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to 0.
[0418] A loop filter_across_subpic_enabled_flag[i] equal to 1 indicates that loop filtering can be performed across the boundary of the i-th subpic of each codec image in the CVS. A loop filter_cross_subpic_enabled_flag[i] equal to 0 indicates that loop filtering is not performed across the boundary of the i-th subpic of each codec image in the CVS. When it does not exist, the value of loop filter_cross_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0419] Bitstream consistency requirements are based on the following constraints:
[0420] For any two subpicks, subpicA and subpicB, when the index of subpicA is less than the index of subpicB, in terms of decoding order, any codec NAL unit of subpicA will be after any codec NAL unit of subpicB.
[0421] – The shape of the sub-image should be such that, when decoded, the entire left and entire top boundaries of each sub-image include the image boundary, or include the boundary of the previously decoded sub-image.
[0422] A value of 1 for `sps_subpic_id_present_flag` indicates that a subpicture ID mapping exists in the SPS. A value of 0 for `sps_subpic_id_present_flag` indicates that a subpicture ID mapping does not exist in the SPS.
[0423] `sps_subpic_id_signaling_present_flag` equal to 1 specifies that signaling notifications are made for the subpicture ID mapping in SPS. `sps_subpic_id_signaling_present_flag` equal to 0 specifies that signaling notifications are not made for the subpicture ID mapping in SPS. When it does not exist, the value of `sps_subpic_id_signaling_present_flag` is inferred to be equal to 0.
[0424] Increasing 1 in sps_subpic_id_len_minus1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 should be in the range of 0 to 15, inclusive.
[0425] `sps_subpic_id[i]` specifies the subpick ID of the i-th subpick. The length of the `sps_subpic_id[i]` syntax element is `sps_subpic_id_len_minus1+1` bits. When `sps_subpic_id_present_flag` does not exist and is equal to 0, for each `i` in the range 0 to `sps_num_subpics_minus1`, including the end value, the value of `sps_subpic_id[i]` is inferred to be equal to `i`.
[0426] ph_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id for the PPS being used. The value of ph_pic_parameter_set_id should be in the range of 0 to 63, inclusive.
[0427] The requirement for bitstream consistency is that the TemporalId value of the image header should be greater than or equal to the TemporalId value of the PPS whose pps_pic_parameter_set_id is equal to ph_pic_parameter_set_id.
[0428] A value of 1 for `ph_subpic_id_signaling_present_flag` specifies that the subpick ID mapping is signaled in the picture header. A value of 0 for `ph_subpic_id_signaling_present_flag` indicates that the subpick ID mapping is not signaled in the picture header.
[0429] Incrementing 1 by ph_subpic_id_len_minus1 specifies the number of bits used to represent the syntax element ph_subpic_id[i]. The value of pic_subpic_id_len_minus1 should be in the range of 0 to 15, inclusive.
[0430] The requirement for bitstream consistency is that the value of ph_subpic_id_len_minus1 should be the same for all image headers referenced by the encoded images in CVS.
[0431] ph_subpic_id[i] specifies the subpick ID of the i-th subpick. The length of the ph_subpic_id[i] syntax element is ph_subpic_id_len_minus1+1 bits.
[0432] The comprehension of the list SubpicIdList[i] is as follows:
[0433]
[0434] Deblocking filtering process
[0435] Overview
[0436] The input to this process is the reconstructed image before the removal of blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[0437] The output of this process is the reconstructed image after removing the blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[0438] First, the vertical edges in the image are filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the image are filtered. Vertical and horizontal edges in the CTB of each CTU are processed separately on a codec unit basis. Vertical edges of the codec blocks in a codec unit are filtered starting from the left edge of the codec block and proceeding geometrically towards the right edge of the codec block. Horizontal edges of the codec blocks in a codec unit are filtered, starting from the top edge of the codec block and proceeding geometrically towards the bottom edge of the codec block.
[0439] Note – Although the filtering process is specified on an image-based basis in this specification, the filtering process can also be implemented on an encoding / decoding unit-based basis with equivalent results, provided that the decoder correctly considers the processing dependency order to produce the same output value.
[0440] The deblocking filtering process is applied to all encoded and decoded sub-block edges and transform block edges of the image, except for the following types of edges:
[0441] – Edges on the image boundary
[0442] – Edges of subpicks where loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0 The edge where the boundaries overlap
[0443] – When PPS_loop_filter_cross_virtual_boundaries_disabled_flag equals 1, the edges that coincide with the virtual boundaries of the image.
[0444] – Edges that coincide with tile boundaries when loop_filter_cross_tiles_enabled_flag is equal to 0
[0445] – Edges that coincide with the slice boundary when loop_filter_cross_slices_enabled_flag equals 0
[0446] – Edges that coincide with the top or left boundary of a slice where slice_deblocking_filter_disabled_flag is equal to 1.
[0447] Edges within the slice where –slice_deblocking_filter_disabled_flag equals 1
[0448] – Edges that do not correspond to the 4×4 sample grid boundary of the brightness component.
[0449] – Edges that do not correspond to the boundaries of the 8×8 sample grid for the chromaticity components
[0450] – Edges on both sides of the edge within the luminance component where intra_bdpcm_luma_flag is equal to 1.
[0451] – Edges on both sides of the edge within the chroma component where intra_bdpcm_chroma_flag is equal to 1.
[0452] – The edge of a chromatic sub-block that is not the edge of a correlated transform unit.
[0453] …
[0454] Deblocking filtering process in one direction
[0455] The input to this process is:
[0456] – Specifies the variable `treeType` to indicate whether the current processing is for the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA).
[0457] – When treeType equals DUAL_TREE_LUMA, reconstruct the image before removing the blocks.
[0458] That is, the array recPicture L ,
[0459] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr ,
[0460] – The variable edgeType specifies whether to filter vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR).
[0461] The output of this process is the reconstructed image after removing the blocks, i.e.:
[0462] – When treeType equals DUAL_TREE_LUMA, the array recPicture L ,
[0463] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr .
[0464] The derivation of variables firstCompIdx and lastCompIdx is as follows:
[0465] firstCompIdx=(treeType==DUAL_TREE_CHROMA)? 1:0 (8-1010)
[0466] lastCompIdx=(treeType==DUAL_TREE_LUMA||ChromaArrayType==0)? 0:2(8-1011)
[0467] For each codec unit and each codec block of each color component indicated by the color component index cIdx, having a codec block width nCbW, a codec block height nCbH, and the position (xCb, yCb) of the top-left sample point of the codec block, cIdx ranges from firstCompIdx to lastCompIdx, including first compidx and lastCompIdx. When cIdx equals 0, or when cIdx is not equal to 0 and edgeType equals EDGE_VER and xCb%8 equals 0, or when cIdx is not equal to 0 and edgeType equals EDGE_HOR and yCb%8 equals 0, the edges are filtered by the following ordered steps:
[0468] 1. The derivation of the variable filterEdgeFlag is as follows:
[0469] – If edgeType equals EDGE_VER, and one or more of the following conditions are true, then filterEdgeFlag is set to 0:
[0470] – The left boundary of the current encoding / decoding block is the left boundary of the image.
[0471] – The left boundary of the current codec block is either the left or right boundary of the sub-image, and loop_filter_cross_ subpic_enabled_flag[SubPicIdx] is equal to 0.
[0472] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[0473] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[0474] – The left boundary of the current codec block is one of the vertical virtual boundaries of the image, and VirtualBoundariesDisabledFlag is equal to 1.
[0475] Otherwise, if edgeType equals EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set to 0:
[0476] – The top boundary of the current luminance codec block is the top boundary of the image.
[0477] – The top boundary of the current codec block is either the top or bottom boundary of the sub-image, and loop_filter_ cross_subpic_enabled_flag[SubPicIdx] is equal to 0.
[0478] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[0479] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[0480] – The top boundary of the current codec block is one of the horizontal virtual boundaries of the image, and VirtualBoundariesDisabledFlag is equal to 1.
[0481] Otherwise, filterEdgeFlag is set to 1.
[0482] 2.7 Examples of TPM, HMVP, and GEO
[0483] In VVC, TPM (triangular prediction mode) divides a block into two triangles with different motion information.
[0484] In VVC, HMVP (History-based Motion Vector Prediction) maintains a table of motion information used for motion vector prediction. This table is updated after a block of inter-frame codecs is decoded, but it is not updated if the block is TPM codec.
[0485] GEO (geometry partition mode) is an extension of TPM. Using GEO, a block can be divided into two partitions using straight lines. These partitions can be triangular or not.
[0486] 2.8 ALF, CC-ALF, and Virtual Boundaries
[0487] In VVC, the ALF (Adaptive Loop-Filter) is applied after the image is decoded to improve image quality.
[0488] The use of Virtual Boundary (VB) in VVC makes ALF easier to design in hardware. Using VB, ALF runs within an ALF processing unit defined by two ALF virtual boundaries.
[0489] CC-ALF (cross-component ALF) filters chrominance samples by referencing information from luminance samples.
[0490] 2.9 Example SEI of Sub-images
[0491] D.2.8 Sub-picture level information SEI message syntax
[0492]
[0493] D.3.8 Sub-picture level information SEI message semantics
[0494] When testing the consistency of the extracted bitstream containing sub-pictures according to Appendix A, the Subpicture Level Information (SEI) message contains information about the level of conformity of the sub-pictures in the bitstream.
[0495] When a Sub-Picture Level Information (SEI) message exists in any picture within a CLVS, the SEI message will exist in the first picture of the CLVS. Sub-Picture Level Information (SEI) messages continue from the current picture to the current layer in decoding order until the end of the CLVS. All Sub-Picture Level Information (SEI) messages applicable to the same CLVS should have the same content.
[0496] The `sli_seq_parameter_set_id` indicates and should be equal to the `sps_seq_parameter_set_id` of the SPS, which is referenced by the codec image associated with the subpicture level information (SEI) message. The value of `sli_seq_parameter_set_id` should be equal to the value of `pps_seq_parameter_set_id` in the PPS referenced by the `ph_pic_parameter_set_id` of the codec image associated with the subpicture level information (SEI) message.
[0497] The requirement for bitstream consistency is that when a subpic-level information (SEI) message exists for CLVS, the value of subpic_treated_as_pic_flag[i] should be equal to 1 for each i value in the range from 0 to sps_num_subpics_minus1, including end values.
[0498] num_ref_levels_minus1 plus 1 specifies the number of reference levels for each signaling notification in the sps_num_subpics_minus1+1 subpics.
[0499] An explicit_fraction_present_flag of 1 indicates that the syntax element ref_level_fraction_minus1[i] exists. An explicit_fraction_present_flag of 0 indicates that the syntax element ref_level_fraction_minus1[i] does not exist.
[0500] ref_level_idc[i] indicates that each sub-picture conforms to the level specified in Appendix A. The bitstream should not contain values of ref_level_idc other than those specified in Appendix A. Other values of ref_level_idc[i] are reserved for future use by ITU-T|ISO / IEC. The requirement for bitstream consistency is that for any k value greater than i, the value of ref_level_idc[i] should be less than or equal to ref_level_idc[k].
[0501] The increment of 1 in ref_level_fraction_minus1[i][j] specifies the score of the level constraint associated with ref_level_idc[i], where the j-th sub-picture of ref_level_idc[i] conforms to the specification of Clause A.4.1.
[0502] The variable SubPicSizeY[j] is set to equal to (subpic_width_minus1[j]+1)*(subpic_height_minus1[j]+1).
[0503] When it does not exist, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256*SubPicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])–1.
[0504] The variable RefLevelFraction[i][j] is set to equal ref_level_fraction_minus1[i][j]+1.
[0505] The derivation of variables SubPicNumTileCols[j] and SubPicNumTileRows[j] is as follows:
[0506]
[0507]
[0508] The derivation of variables SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j] is as follows:
[0509] SubPicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*RefLevelFraction[i][j]÷256) (D.6)
[0510] SubPicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*RefLevelFraction[i][j]÷256) (D.7)
[0511] MaxCPB is derived from ref_level_idc[i], as specified in Clause A.4.2.
[0512] Note 1 – When extracting sub-pictures, the resulting bitstream has a CpbSize greater than or equal to SubPicCpbSizeVcl[i][j] and SubPicCpbSizeNal[i][j] (indicated or inferred in SPS).
[0513] The requirement for bitstream consistency is that the bitstream generated from the extraction of the j-th subgraph and conforming to a configuration file with general_tier_flag equal to 0 and level equal to ref_level_idc[i] should comply with the following constraints for each bitstream consistency test specified in Appendix C, where j ranges from 0 to sps_num_subpics_minus1, inclusive, and i ranges from 0 to num_ref_level_minus1, inclusive:
[0514] –Ceil(256*SubPicSizeY[i]÷RefLevelFraction[i][j]) should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1.
[0515] The value of –Ceil(256*(subpic_width_minus1[i]+1)÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0516] The value of –Ceil(256*(subpic_height_minus1[i]+1)÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0517] The value of SubPicNumTileCols[j] should be less than or equal to MaxTileCols, and the value of SubPicNumTileRows[j] should be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1.
[0518] For any set of sub-images that contains one or more sub-images and consists of multiple sub-images in the sub-image index list SubPicSetIndices and the sub-image set NumSubPicInSet, the level information of the sub-image set is derived.
[0519] The variable representing the total level score relative to the reference level ref_level_idc[i].
[0520] SubPicSetAccLevelFraction[i] and the variables of the sub-picture set
[0521] The derivation of SubPicSetCpbSizeVcl[i][j] and SubPicSetCpbSizeNal[i][j] is as follows:
[0522]
[0523] The derivation of the value of the sub-picture set sequence level indicator SubPicSetLevelIdc is as follows:
[0524]
[0525] The MaxTileCols and MaxTileRows used for ref_level_idc[i] are specified in Table A.1.
[0526] Subpicture set bitstreams that conform to a profile with general_tier_flag equal to 0 and a level equal to SubPicSetLevelIdc should comply with the following constraints C for each bitstream conformance test specified in Appendix C:
[0527] – For the VCL HRD parameter, SubPicSetCpbSizeVcl[i] should be less than or equal to CpbVclFactor*MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in CpbVclFactor bits in Table A.1.
[0528] – For the NAL HRD parameter, SubPicSetCpbSizeVcl[i] should be less than or equal to CpbNalFactor*MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in CpbNalFactor bits in Table A.1.
[0529] Note 2 – When extracting sub-picture sets, the resulting bitstream has a CpbSize greater than or equal to SubPicCpbSizeVcl[i][j] and SubPicSetCpbSizeNal[i][j] (indicated or inferred in SPS).
[0530] 2.10. Palette Mode
[0531] 2.10.1 The concept of palette mode
[0532] The basic idea behind palette mode is that pixels in a CU are represented by a small, representative set of color values. This set is called the palette. Samples outside the palette can be indicated by signaling followed by an escape term for (potentially quantized) component values. These pixels are called escape pixels. Palette mode is as follows: Figure 10 As shown. Figure 10 As shown, for each pixel with three color components (luminance and two chrominance components), an index of the color palette is established, and the block can be reconstructed based on the values established in the color palette.
[0533] 2.10.2 Encoding and Decoding of Palette Entries
[0534] For palette coding blocks, the following key aspects are introduced:
[0535] 1. Construct the current color palette based on the predicted color palette and new entries for signaling notification of the current color palette, if they exist.
[0536] 2. Divide the current sample / pixel into two categories: one category (Category 1) includes the sample / pixel in the current color palette, and the other category (Category 2) includes the sample / pixel outside the current color palette.
[0537] A. For samples / pixels in the second category, quantization is applied (at the encoder) to the sample / pixel, and signaling is used to notify the quantization value; and dequantization is applied (at the decoder).
[0538] 2.10.2.1 Predicted Value Palette
[0539] For the encoding of palette entries, the predicted palette is maintained and updated after the palette encoding block is decoded.
[0540] 2.10.2.1.1 Initialization of the Predicted Value Palette
[0541] The prediction palette is initialized at the beginning of each strip and each slice. Signaling in the SPS informs the maximum size of the palette and the prediction palette. In HEVC-SCC, the `palette_predictor_initializer_present_flag` is introduced in the PPS. When this flag is 1, signaling in the bitstream informs the entries used to initialize the prediction palette.
[0542] Depending on the value of `palette_predictor_initializer_present_flag`, the size of the predictor palette is either reset to 0 or initialized using the predictor palette initialization value entry notified by signaling in the PPS. In HEVC-SCC, a predictor palette initializer of size 0 is enabled to allow explicit disabling of predictor palette initialization at the PPS level.
[0543] The corresponding syntax, semantics, and decoding process are defined as follows:
[0544] 7.3.2.2.3 Sequence Parameter Set Screen Content Encoding and Decoding Extended Syntax
[0545]
[0546] A palette_mode_enabled_flag value of 1 indicates that the palette mode decoding process can be used within an intra-frame block. A palette_mode_enabled_flag value of 0 indicates that the palette mode decoding process is not applied. When it does not exist, the value of palette_mode_enabled_flag is inferred to be 0.
[0547] Palette_max_size specifies the maximum allowed palette size. When it does not exist, the value of palette_max_size is inferred to be 0.
[0548] `delta_palette_max_predictor_size` specifies the difference between the maximum allowed palette prediction size and the maximum allowed palette size. When it does not exist, the value of `delta_palette_max_predictor_size` is inferred to be 0. The derivation of the variable `PaletteMaxPredictorSize` is as follows:
[0549] PaletteMaxPredictorSize=palette_max_size+delta_palette_max_predictor_size (0-57)
[0550] One requirement for bitstream consistency is that when palette_max_size equals 0, the value of delta_palette_max_predictor_size should also equal 0.
[0551] A value of 1 for `sps_palette_predictor_initializer_present_flag` specifies that the sequence palette predictions are initialized using `sps_palette_predictor_initializer_present_flag`. A value of 0 for `sps_palette_predictor_initializer_present_flag` specifies that entries in the sequence palette predictions are initialized to 0. When `sps_palette_predictor_initializer_present_flag` is not present, its value is inferred to be 0.
[0552] One requirement for bitstream consistency is that when palette_max_size equals 0, the value of sps_palette_predictor_initializer_present_flag should be equal to 0.
[0553] The increment of sps_num_palette_predictor_initializer_minus1 specifies the number of entries in the sequence palette prediction initialization setting.
[0554] One requirement for bitstream consistency is that the value of sps_num_palette_predictor_initializer_minus1 plus 1 should be less than or equal to PaletteMaxPredictorSize.
[0555] `sps_palette_predictor_initializers[comp][i]` specifies the value of the `comp`th component of the `i`th palette entry in the SPS, which is used to initialize the `PredictorPaletteEntries` array. For values of `i` ranging from 0 to `sps_num_palette_predictor_initializer_minus1`, the value of `sps_palette_predictor_initializers[0][i]` should be between 0 and (1 < 1). <BitDepth Y Within the range of 0 to 1, including the endpoints, the values of sps_palette_predictor_initializers[1][i] and sps_palette_predictor_initializers[2][i] should be between 0 and (1 < 1). <BitDepth C Within the range of )-1, including the endpoints.
[0556] 7.3.2.3.3 Image Parameter Set Image Content Encoding Extended Syntax
[0557]
[0558] `pps_palette_predictor_initializer_present_flag` equal to 1 specifies that the initial values for the palette predictions used to reference the PPS image are derived based on the initial values specified by the PPS. `pps_palette_predictor_initializer_flag` equal to 0 specifies that the initial values for the palette predictions used to reference the PPS image are inferred to be equal to the initial values specified by the active SPS. When it does not exist, the value of `pps_palette_predictor_initializer_flag` is inferred to be 0.
[0559] One requirement for bitstream consistency is that the value of pps_palette_predictor_initializer_flag should be equal to 0 when palette_max_size is equal to 0 or palette_mode_enabled_flag is equal to 0.
[0560] pps_num_palette_predictor_initializer specifies the number of entries in the image palette prediction initialization settings.
[0561] One requirement for bitstream consistency is that the value of pps_num_palette_predictor_initializer should be less than or equal to PaletteMaxPredictorSize.
[0562] The palette prediction variable is initialized as follows:
[0563] – If the codec tree unit is the first codec tree unit in the chip, then the following applies:
[0564] – Call the initialization process of the palette prediction value variable
[0565] Otherwise, if entropy_coding_sync_enabled_flag equals 1 and CtbAddrInRs%PicWidthInCtbsY equals 0 or TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs-1]], then the following applies:
[0566] – Using the position (x0, y0) of the top-left luminance sample of the current codec block, the position (xNbT, yNbT) of the top-left luminance sample of the spatially adjacent block T is derived as follows:
[0567] (xNbT,yNbT)=(x0+CtbSizeY,y0-CtbSizeY) (0-58)
[0568] – Call the availability derivation process of the blocks in the z-scan order, taking the position (xCurr, yCurr) set to equal to (x0, y0) and the adjacent position (xNbY, yNbY) set to equal to (xNbT, yNbT) as input, and assign the output to availableFlagT.
[0569] – The synchronization process for calling context variables, Rice parameter initialization states, and palette prediction value variables is as follows:
[0570] – If availableFlagT equals 1, then the synchronization process for the context variables, Rice parameter initialization state, and palette prediction value variables is invoked, with TableStateIdxWpp, TableMpsValWpp, TableStatCoeffWpp, PredictorPaletteSizeWpp, and TablePredictorPaletteEntriesWpp as inputs.
[0571] Otherwise, the following applies:
[0572] – Call the initialization process of the palette prediction value variable.
[0573] Otherwise, if CtbAddrInRs equals slice_segment_address and dependent_slice_segment_flag equals 1, then the synchronization process for initializing the context variables and Rice parameter is invoked, with TableStateIdxDs, TableMpsValDs, TableStatCoeffDs, PredictorPaletteSizeDs, and TablePredictorPaletteEntriesDs as input.
[0574] Otherwise, the following applies:
[0575] – Call the initialization process of the palette prediction value variable.
[0576] 9.3.2.3 Initialization process for palette prediction value entries
[0577] The output of this process is the initialized palette prediction variables PredictorPaletteSize and PredictorPaletteEntries.
[0578] The derivation of the variable numComps is as follows:
[0579] numComps=(ChromaArrayType==0)? 1:3 (0-59)
[0580] – If pps_palette_predictor_initializer_present_flag equals 1, then the following applies:
[0581] –PredictorPaletteSize is set to equal to pps_num_palette_predictor_initializer.
[0582] The derivation of the `PredictorPaletteEntries` array is as follows:
[0583] for(comp=0;comp <numComps;comp++)
[0584] for(i=0; i <PredictorPaletteSize;i++) (0-60)
[0585] PredictorPaletteEntries[comp][i]=pps_palette_predictor_initializers[comp][i]
[0586] – Otherwise (pps_palette_predictor_initializer_present_flag equals 0), if sps_palette_predictor_initializer_present_flag equals 1, then the following applies:
[0587] –PredictorPaletteSize is set to equal to sps_num_palette_predictor_initializer_minus1 plus 1.
[0588] The derivation of the `PredictorPaletteEntries` array is as follows:
[0589] for(comp=0;comp <numComps;comp++)
[0590] for(i=0; i <PredictorPaletteSize;i++) (0-61)
[0591] PredictorPaletteEntries[comp][i]=sps_palette_predictor_initializers[comp][i]
[0592] Otherwise (pps_palette_predictor_initializer_present_flag equals 0 and sps_palette_predictor_initializer_present_flag equals 0), PredictorPaletteSize is set to 0.
[0593] 2.10.2.1.2 Use of the Predicted Value Palette
[0594] For each entry in the palette prediction values, signaling notifies a reuse flag to indicate whether it is part of the current palette. This is in Figure 9 As shown in the diagram, a reuse flag is delivered using a zero-run length codec. Subsequently, the number of new palette entries is signaled using Exponential Golomb (EG) code, i.e., EG-0. Finally, the component values of the new palette entries are signaled.
[0595] 2.10.2.2 Updating the Predicted Value Palette
[0596] The predicted color palette is updated through the following steps:
[0597] 1. Before decoding the current block, there exists a prediction palette, represented by PltPred0.
[0598] 2. Construct the current palette table by first inserting an entry from PltPred0, and then inserting a new entry for the current palette.
[0599] 3. Construct PltPred1:
[0600] A. First, add these to the current color palette table (possibly including those from PltPred0).
[0601] B. If not full, add unreferenced entries from PltPred0 based on the ascending entry index.
[0602] 2.10.3 Palette Index Encoding and Decoding
[0603] like Figure 15 As shown, the palette index is encoded and decoded using horizontal and vertical traversal scans. The scan order is explicitly signaled in the bitstream using the `palette_transpose_flag`. For the remainder of this section, it is assumed that the scan is horizontal.
[0604] The palette index is encoded using two palette sample modes: "COPY_LEFT" and "COPY_ABOVE". In "COPY_LEFT" mode, the palette index is assigned to the decoding index. In "COPY_ABOVE" mode, the palette index of the sample in the previous row is copied. For both "COPY_LEFT" and "COPY_ABOVE" modes, signaling informs the run-length value, which specifies the number of subsequent samples that are also encoded and decoded using the same mode.
[0605] In palette mode, the index value of the escaped sample is the number of palette entries. Furthermore, when an escaped symbol is part of the downstream process in "COPY_LEFT" or "COPY_ABOVE" mode, the escaped component value is signaled for each escaped symbol. The encoding and decoding of the palette index is as follows: Figure 16 As shown.
[0606] This grammatical sequence is performed as follows: First, the number of index values for the CU is signaled. Next, the actual index values for the entire CU are signaled using truncated binary encoding / decoding. Both the index number and index values are encoded / decoded in bypass mode. This groups the bypass libraries associated with the indexes together. Then, the palette sample mode (if necessary) and run lengths are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and encoded / decoded in bypass mode. The binaryization of the escape samples is EG encoding with third order, i.e., EG-3.
[0607] The additional syntax element `last_run_type_flag` is signaled after the index value. This syntax element, combined with the number of indices, eliminates the need for signaling the run value corresponding to the last run in the block.
[0608] In HEVC-SCC, the palette mode also supports 4:2:2, 4:2:0, and monochrome chroma formats. For all chroma formats, the signaling notifications for palette entries and palette indices are almost identical. In non-monochrome formats, each palette entry includes three components. In monochrome formats, each palette entry includes one component. For subsampled chroma directions, chroma samples are associated with a luminance sample index divisible by 2. After reconstructing the palette index for the CU, if a sample has only one associated component, only the first component of the palette entry is used. The only difference in signaling notifications is the number of escaped component values. For each escaped sample, the number of escaped component values in the signaling notification may differ depending on the number of components associated with that sample.
[0609] In addition, there is an index adjustment process in the palette index encoding and decoding. When signaling notifies the palette index, the left-adjacent index or the upper-adjacent index should be different from the current index. Therefore, by eliminating one possibility, the range of the current palette index can be reduced by 1. Afterwards, the index is represented using truncated binary (TB) binary representation.
[0610] The relevant text for this section is shown below, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.
[0611] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is the index of an array represented by CurrentPaletteEntries. The array indices xC and yC specify the position (xC, yC) of the sample relative to the top-left luminance sample of the image. The value of PaletteIndexMap[xC][yC] should be in the range of 0 to MaxPaletteIndex, inclusive.
[0612] The derivation of the variable adjustedRefPaletteIndex is as follows:
[0613]
[0614]
[0615] When CopyAboveIndicesFlag[xC][yC] equals 0, the derivation of the variable CurrPaletteIndex is as follows:
[0616] if(CurrPaletteIndex>=adjustedRefPaletteIndex)
[0617] CurrPaletteIndex++
[0618] 2.10.3.1 Decoding process of palette encoding blocks
[0619] 1. Read the forecast information to mark which entries in the forecast value palette will be reused;
[0620] (palette_predictor_run)
[0621] 2. Read the new palette entry for the current block.
[0622] a)num_signalled_palette_entries
[0623] b)new_palette_entries
[0624] 3. Construct CurrentPaletteEntries based on a) and b).
[0625] 4. Read the escape symbol presence flag: palette_escape_val_present_flag to deduce MaxPaletteIndex.
[0626] 5. How many samples are not encoded using copy mode / run-length mode?
[0627] a)num_palette_indices_minus1
[0628] b) For each sample that is not encoded using copy mode / run-length mode, encode palette_idx_idc in the current plt table.
[0629] 2.11 Merge Estimation Region (MER)
[0630] HEVC employs a Merge Candidate List (MER). The way the Merge candidate list is constructed introduces dependencies between adjacent blocks. Especially in embedded encoder implementations, the motion estimation phases for adjacent blocks are typically executed in parallel, or at least pipelined, to increase throughput. This isn't a major issue for AMVP, as the MVP is only used for differential encoding and decoding of the MVs found through motion search. However, the motion estimation phase for Merge patterns typically only involves constructing the candidate list and deciding which candidate to select based on a cost function. Due to the aforementioned dependencies between adjacent blocks, Merge candidate lists for adjacent blocks cannot be generated in parallel, becoming a bottleneck in parallel encoder designs. Therefore, a parallel Merge estimation level is introduced in HEVC, indicating the region where the Merge candidate list can be independently derived by checking if a candidate block is located within the Merge Estimation Region (MER). Candidate blocks within the same MER are not included in the Merge candidate list. Therefore, their motion data does not need to be available during list construction. When this level is, for example, 32, all prediction units in a 32×32 region can construct the Merge candidate list in parallel because none of the Merge candidates in the same 32×32 MER have been inserted into the list. Figure 12 The illustration shows an example of a CTU partition with seven CUs and ten PUs. All potential Merge candidates for the first PU 0 are available because they are outside the first 32×32 MER.
[0631] For the second MER, when the merge estimates within that MER should be independent, the merge candidate lists for PUs 2-6 cannot include motion data from these PUs. Therefore, for example, when looking at PU 5, there are no available merge candidates, and they are not inserted into the merge candidate list. In this case, the merge list for PU 5 only includes temporal candidates (if available) and zero MV candidates. To allow the encoder to compromise between parallelism and encoding / decoding efficiency, the parallel merge estimation level is adaptive and signaled as log2_parallel_Merge_level_minus2 in the picture parameter set. The following MER sizes are allowed: 4×4 (parallel merge estimation is not possible), 8×8, 16×16, 32×32, and 64×64. The higher degree of parallelization enabled by larger MERs excludes more potential candidates from the merge candidate list. On the other hand, this reduces encoding / decoding efficiency. Another modification to the merge list construction begins to increase throughput when the merge estimation region is larger than a 4×4 block. For a CU with an 8×8 brightness CB, only a single Merge candidate list is used for all PUs within the CU.
[0632] 3. Examples of technical problems solved by the disclosed embodiments
[0633] (1) Some designs can violate sub-image constraints.
[0634] A. The TMVP in the affine construction candidate can obtain the MV in the juxtaposed image outside the range of the current sub-image.
[0635] B. When deriving gradients in Bi-Directional Optical Flow (BDOF) and Prediction Refinement Optical Flow (PROF), it is necessary to extract integer reference samples for two extended rows and two extended columns. These reference samples may be outside the range of the current sub-image.
[0636] C. When deriving the chroma residual scaling factor in luma mapping chroma scaling (LMCS), the reconstructed luma samples accessed may be outside the range of the current sub-image.
[0637] D. When deriving the lumen intra-prediction mode, intra-prediction reference samples, CCLM reference samples, spatial neighbor block availability of Merge / AMVP / CIIP / IBC / LMCS spatial neighbor candidates, quantization parameters, CABAC initialization process, ctxInc derivation using left and top syntax elements, and ctxInc for the syntax element mtt_split_cu_vertical_flag, neighbor blocks may be outside the range of the current sub-image. Sub-image representation may result in sub-images with incomplete CTUs. CTU segmentation and CU partitioning processes may need to account for incomplete CTUs.
[0638] (2) The syntax elements related to sub-images in signaling notifications can be arbitrarily large, which may lead to overflow issues.
[0639] (3) The representation of sub-images may result in non-rectangular sub-images.
[0640] (4) Currently, sub-pictures and sub-picture grids are defined in units of 4 samples. And the length of the syntax element depends on the picture height divided by 4. However, since the current pic_width_in_luma_samples and pic_height_in_luma_samples should be integer multiples of Max(8,MinCbSizeY), it may be necessary to define sub-picture grids in units of 8 samples.
[0641] (5) SPS syntax, pic_width_max_in_luma_samples and pic_height_max_in_luma_samples may need to be limited to no less than 8.
[0642] (6) The current design does not take into account the reference image resampling / scalability and the interaction between sub-images.
[0643] (7) In temporal filtering, it may be necessary to cross samples from different sub-images.
[0644] (8) When signaling notification stripes are used, information can be inferred in some cases without signaling notification.
[0645] (9) It is possible that none of the defined stripes can cover the entire image or sub-image.
[0646] (10) The IDs of the two sub-images can be the same.
[0647] (11) pic_width_max_in_luma_samples / CtbSizeY may be equal to 0, resulting in meaningless Log2() calculation.
[0648] (12) ID in PH is worse than ID in PPS, but not as good as ID in SPS, which is inconsistent.
[0649] (13) log2_transform_skip_max_size_minus2 in PPS is parsed based on sps_transform_skip_enabled_flag in SPS, resulting in a parsing dependency.
[0650] (14) The loop_filter_cross_subpic_enabled_flag used for deblocking only considers the current subpic and not the adjacent subpic.
[0651] (15) In the derivation of temporal merging candidates based on sub-blocks, when subpic_treatment_as_pic_flag[SubPicIdx] equals 1, the range of the ordinate of the juxtaposed sub-blocks in the juxtaposed image is not limited, and the sample points outside the sub-image can be referenced.
[0652] (16) In applications, subpicks are designed to provide flexibility, allowing regions at the same location in a sequence of pictures to be decoded or extracted independently. These regions may have specific requirements. For example, they might require a high-quality Region of Interest (ROI). In another example, they could be used to quickly navigate video tracks. In yet another example, they could provide a low-precision, low-complexity, and low-power bitstream that can be fed to complexity-sensitive end users. All these applications may require that regions of a subpick be encoded with a different configuration than other parts. However, current VVC lacks a mechanism for independently configuring subpicks.
[0653] 4. Example technologies and implementation examples
[0654] Below are detailed examples that should be considered as explanations of general concepts. These items should not be interpreted in a narrow sense. Furthermore, these items can be combined in any way. In the following text, a temporal filter is used to represent a filter that requires samples from other images. Max(x,y) yields the larger of x and y. Min(x,y) yields the smaller of x and y.
[0655] 1. Assuming the top-left corner coordinates of the desired sub-image are (xTL, yTL) and the bottom-right corner coordinates are (xBR, yBR), the location (called position RB) where the temporal MV prediction value is obtained in the image to generate affine motion candidates (e.g., the constructed affine Merge candidate) must be in the desired sub-image.
[0656] a. In one example, the required sub - picture is the sub - picture that covers the current block.
[0657] b. In one example, if the position RB with coordinates (x, y) is outside the required sub - picture, the time - domain MV prediction value is considered unavailable.
[0658] i. In one example, if x > xBR, the position RB is outside the required sub - picture.
[0659] ii. In one example, if y > yBR, the position RB is outside the required sub - picture.
[0660] iii. In one example, if x < xTL, the position RB is outside the required sub - picture.
[0661] iv. In one example, if y < yTL, the position RB is outside the required sub - picture.
[0662] c. In one example, if the position RB is outside the required sub - picture, a replacement for RB is utilized.
[0663] i. Alternatively, in addition, the replacement position should be within the required sub - picture.
[0664] d. In one example, the position RB is cropped into the required sub - picture.
[0665] i. In one example, x is cropped to x = Min(x, xBR).
[0666] ii. In one example, y is cropped to y = Min(y, yBR).
[0667] iii. In one example, x is cropped to x = Max(x, xTL).
[0668] iv. In one example, y is cropped to y = Max(y, yTL).
[0669] e. In one example, the position RB can be the lower - right position within the corresponding block of the current block in the juxtaposed picture.
[0670] f. The proposed method can be used for other codec tools that need to access motion information from a picture different from the current picture.
[0671] g. In one example, whether to apply the above method (e.g., the position RB must be within the required sub - picture (e.g., as required in 1.a and / or 1.b)) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub - picture index of the sub - picture covering the current block.
[0672] 2. Assume that the upper - left corner coordinates of the required sub - picture are (xTL, yTL) and the lower - right corner coordinates of the required sub - picture are (xBR, yBR). Then the position (referred to as position S) where the integer sample points are extracted from the reference not used during the interpolation process must be within the required sub - picture.
[0673] a. In one example, the required sub - picture is the sub - picture covering the current block.
[0674] b. In one example, if the position S with coordinates (x, y) is outside the required sub - picture, the reference sample points are considered unavailable.
[0675] i. In one example, if x > xBR, then the position S is outside the required sub - picture.
[0676] ii. In one example, if y > yBR, then the position S is outside the required sub - picture.
[0677] [[ID=1⑧]]iii. In one example, if x < xTL, then the position S is outside the required sub - picture.
[0678] iv. In one example, if y < yTL, then the position S is outside the required sub - picture.
[0679] c. In one example, the position S is cropped to the required sub - picture.
[0680] i. In one example, x is cropped to x = Min(x, xBR).
[0681] ii. In one example, y is cropped to y = Min(y, yBR).
[0682] iii. In one example, x is cropped to x = Max(x, xTL).
[0683] iv. In one example, y is cropped to y = Max(y, yTL).
[0684] d. In one example, whether the position S must be within the required sub-picture (e.g., as required in 2.a and / or 2.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub-picture index of the sub-picture covering the current block.
[0685] e. In one example, the extracted integer samples are used to generate gradients in the BDOF and / or PORF.
[0686] 3. Assuming that the upper left corner coordinates of the required sub-picture are (xTL, yTL) and the lower right corner coordinates of the required sub-picture are (xBR, yBR), the position (referred to as position R) where the reconstructed luminance sample value is extracted can be within the required sub-picture.
[0687] a. In one example, the required sub-picture is the sub-picture covering the current block.
[0688] b. In one example, if the position R with coordinates (x, y) is outside the required sub-picture, the reference samples are considered unavailable.
[0689] i. In one example, if x > xBR, then the position R is outside the required sub-picture.
[0690] ii. In one example, if y > yBR, then the position R is outside the required sub-picture.
[0691] iii. In one example, if x < xTL, then the position R is outside the required sub-picture.
[0692] iv. In one example, if y < yTL, then the position R is outside the required sub-picture.
[0693] c. In one example, the position R is clipped to the required sub-picture.
[0694] i. In one example, x is clipped to x = Min(x, xBR).
[0695] ii. In one example, y is clipped to y = Min(y, yBR).
[0696] iii. In one example, x is clipped to x = Max(x, xTL).
[0697] iv. In one example, y is clipped to y = Max(y, yTL).
[0698] d. In one example, whether the position R must be within the required sub - picture (e.g., as required in 3.a and / or 3.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / strip header / slice group header. For example, the syntax element can be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the sub - picture index of the sub - picture covering the current block.
[0699] e. In one example, the obtained luma samples are used to derive the scaling factors for the chroma component(s) in LMCS.
[0700] 4. Assuming that the upper - left coordinates of the required sub - picture are (xTL, yTL) and the lower - right coordinates of the required sub - picture are (xBR, yBR), then the location (referred to as location N) where the picture boundary check for BT / TT / QT partitioning, the depth derivation for BT / TT / QT, and / or the signaling of the CU partitioning flag must be within the required sub - picture.
[0701] a. In one example, the required sub - picture is the sub - picture covering the current block.
[0702] b. In one example, if the location N with coordinates (x, y) is outside the required sub - picture, the reference samples are considered unavailable.
[0703] i. In one example, if x > xBR, then the location N is outside the required sub - picture.
[0704] ii. In one example, if y > yBR, then the location N is outside the required sub - picture.
[0705] iii. In one example, if x < xTL, then the location N is outside the required sub - picture.
[0706] iv. In one example, if y < yTL, then the location N is outside the required sub - picture.
[0707] c. In one example, the location N is clipped to the required sub - picture.
[0708] i. In one example, x is clipped to x = Min(x, xBR).
[0709] ii. In one example, y is clipped to y = Min(y, yBR).
[0710] iii. In one example, x is clipped to x = Max(x, xTL).
[0711] iv. In one example, y is clipped to y = Max(y, yTL).
[0712] d. In one example, whether position N must be in the desired subpicture (e.g., as required in 4.a and / or 4.b) may depend on one or more syntax elements in the signaling notification in the VPS / DPS / SPS / PPS / APS / strip header / piece group header. For example, the syntax element could be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index that covers the subpicture of the current block.
[0713] 5. The history-based motion vector prediction (HMVP) table can be reset before decoding a new sub-image in an image.
[0714] a. In one example, the HMVP table used for IBC encoding / decoding can be reset.
[0715] b. In one example, the HMVP table used for inter-frame encoding and decoding can be reset.
[0716] c. In one instance, the HMVP table used for intra-frame encoding / decoding can be reset.
[0717] 6. Sub-image syntax elements can be defined in units of N (e.g., N = 8, 32, etc.).
[0718] a. In one example, the width of each element of the sub-image identifier grid is in units of N samples.
[0719] b. In one example, the height of each element of the sub-image identifier grid is in units of N samples.
[0720] c. In one example, N is set to the width and / or height of the CTU.
[0721] 7. The syntax elements for image width and image height can be restricted to be no less than K (K>=8).
[0722] a. In one example, the image width may need to be limited to no less than 8.
[0723] b. In one example, the image height may need to be limited to no less than 8.
[0724] 8. A consistent bitstream should not allow subpicture encoding / decoding and adaptive resolution conversion (ARC) / dynamic resolution conversion (DRC) / reference picture resampling (RPR) to be enabled for a video unit (e.g., a sequence).
[0725] a. In one example, signaling notifications that enable subpicture encoding / decoding can be made under conditions where ARC / DRC / RPR are not allowed.
[0726] i. In one example, when subpics are enabled, such as subpics_present_flag equals 1, and for all pictures valid for that SPS, pic_width_in_luma_samples equals max_width_in_luma_sample.
[0727] b. Alternatively, both subpicture encoding / decoding and ARC / DRC / RPR can be enabled for a video unit (e.g., a sequence).
[0728] i. In one example, the consistent bitstream will satisfy the condition that the downsampled sub-pictures generated by ARC / DRC / RPR will still be in the form of a width of K CTUs and a height of M CTUs, where K and M are integers.
[0729] ii. In one example, the consistent bitstream will satisfy that for sub-pictures not located at picture boundaries (e.g., right and / or bottom boundaries), the downsampled sub-pictures generated by ARC / DRC / RPR will still be in the form of a width of K CTUs and a height of M CTUs, where K and M are integers.
[0730] iii. In one example, the CTU size can be adaptively changed based on the image resolution.
[0731] 1) In one example, the maximum CTU size can be signaled in the SPS. For each image with lower precision, the CTU size can be changed accordingly based on the reduced precision.
[0732] 2) In one example, the CTU size can be signaled at the SPS and PPS and / or sub-picture levels.
[0733] 9. You can constrain the syntax elements subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1.
[0734] a. In one example, subpic_grid_col_width_minus1 must not be greater than (or must be less than) T1.
[0735] b. In one example, subpic_grid_row_height_minus1 must not be greater than (or must be less than) T2.
[0736] c. In one example, within a consistent bitstream, subpic_grid_col_width_minus1 and / or subpic_grid_row_height_minus1 must adhere to constraints such as those specified in item 3.a or 3.b.
[0737] d. In one example, T1 in 3.a and / or T2 in 3.b may depend on the grade / level / hierarchy of the video codec standard.
[0738] e. In one example, T1 in 3.a can depend on the image width.
[0739] i. For example, T1 is equal to pic_width_max_in_luma_samples / 4 or pic_width_max_in_luma_samples / 4+Off. Off can be 1, 2, -1, -2, etc.
[0740] f. In one example, T2 in 3.b can depend on the image width.
[0741] i. For example, T2 equals pic_height_max_in_luma_samples / 4 or pic_height_max_in_luma_samples / 4-1+Off. Off can be 1, 2, -1, -2, etc.
[0742] 10. The boundary between two sub-images must be the boundary between two CTUs.
[0743] a. In other words, a CTU cannot be covered by more than one sub-image.
[0744] b. In one example, the unit of subpic_grid_col_width_minus1 can be the CTU width (e.g., 32, 64, 128), instead of 4 as in VVC. The subpic grid width should be (subpic_grid_col_width_minus1 + 1) * CTU width.
[0745] c. In one example, the unit of subpic_grid_col_height_minus1 can be CTU height (e.g., 32, 64, 128), instead of 4 as in VVC. The subpic grid height should be (subpic_grid_col_height_minus1 + 1) * CTU height.
[0746] d. In one example, in a consistent bitstream, if a subpicture scheme is applied, constraints must be satisfied.
[0747] 11. The shape of the constrained sub-image must be rectangular.
[0748] a. In one example, in a consistent bitstream, if a subpicture scheme is applied, constraints must be met.
[0749] b. Subpicks can contain only rectangular stripes. For example, in a consistent bitstream, if a subpick scheme is applied, constraints must be met.
[0750] 12. Constrain that two sub-images cannot overlap.
[0751] a. In one example, in a consistent bitstream, if a subpicture scheme is applied, constraints must be met.
[0752] b. Alternatively, two sub-images can overlap each other.
[0753] 13. Constrain that any location in the image must be covered by one and only one sub-image.
[0754] a. In one example, in a consistent bitstream, if a subpicture scheme is applied, constraints must be met.
[0755] b. Alternatively, a sample point may not belong to any sub-image.
[0756] c. Alternatively, a sample point can belong to more than one sub-image.
[0757] 14. The position and / or size of sub-pictures defined in each precision of the SPS that exists in the same sequence can be constrained to conform to the above constraints.
[0758] a. In one example, the width and height of the sub-images defined in the precision SPS that are mapped to exist in the same sequence should be an integer multiple of N (e.g., 8, 16, 32) luminance samples.
[0759] b. In one example, sub-images can be defined for certain layers, and these sub-images can be mapped to other layers.
[0760] i. For example, a sub-image can be defined for the layer with the highest precision in the sequence.
[0761] ii. For example, a sub-image can be defined for the layer with the lowest precision in the sequence.
[0762] iii. The signaling notification in the SPS / VPS / PPS / strip header can specify which layer the sub-image is defined for.
[0763] c. In one example, when sub-images and different precisions are applied, all precisions (e.g., width and / or height) can be integer multiples of a given precision.
[0764] d. In one example, the width and / or height of a sub-image defined in the SPS can be an integer multiple of the CTU size (e.g., M).
[0765] e. Alternatively, sub-images and different precisions in a sequence may not be allowed simultaneously.
[0766] 15. Sub-images can be applied to only one (or some) layers.
[0767] a. In one example, the sub-images defined in SPS can be applied only to the layer with the highest precision in the sequence.
[0768] b. In one example, the sub-image defined in SPS can be applied only to the layer in the sequence that has the lowest temporal ID.
[0769] c. One or more syntax elements in SPS / VPS / PPS can indicate which layer(s) a sub-image can be applied to.
[0770] d. One or more syntax elements in SPS / VPS / PPS can be used to indicate which layer(s) a sub-image cannot be applied to.
[0771] 16. In one example, the location and / or dimensions of a subpic can be signaled without using subpic_grid_idx.
[0772] a. In one example, signaling can be used to notify the top left position of a sub-image.
[0773] b. In one example, signaling can be used to notify the bottom right position of a sub-image.
[0774] c. In one example, the width of a sub-image can be signaled.
[0775] d. In one example, the height of a sub-image can be signaled.
[0776] 17. For temporal filters, when performing temporal filtering on samples, only samples within the same sub-image to which the current sample belongs can be used. The required sample may be in the same image to which the current sample belongs, or it may be in other images.
[0777] 18. In one example, whether and / or how to apply a segmentation method (such as QT, horizontal BT, vertical BT, horizontal TT, vertical TT, or no segmentation, etc.) may depend on whether the current block (or partition) spans one or more boundaries of the sub-image.
[0778] a. In one example, when the image boundary is replaced by the sub-image boundary, the image boundary handling method used for segmentation in VVC can also be applied.
[0779] b. In one example, whether to parse syntax elements (such as flags) representing segmentation methods (such as QT, horizontal BT, vertical BT, horizontal TT, vertical TT, or no segmentation) may depend on whether the current block (or partition) spans one or more boundaries of the sub-image.
[0780] 19. Instead of dividing an image into multiple sub-images and encoding and decoding each sub-image independently, the proposed approach is to divide the image into at least two sets of sub-regions: the first set includes multiple sub-images, and the second set includes all remaining samples.
[0781] a. In one example, the samples in the second set are not in any sub-picture.
[0782] b. Alternatively, the second set can be encoded / decoded based on information from the first set.
[0783] c. In one example, a default value can be used to indicate whether a sample point / M×K sub-region belongs to the second set.
[0784] i. In one example, the default value can be set to equal to (max_subpics_minus1+K), where K is an integer greater than 1.
[0785] ii. A default value can be assigned to subpic_grid_idx[i][j] to indicate that the grid belongs to the second set.
[0786] 20. It is proposed that the syntax element subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0787] a. For example, the constraint requires that in a consistent bitstream, subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0788] b. For example, the codewords of encoding and decoding subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1.
[0789] 21. Propose that any integer from 0 to max_subpics_minus1 must be equal to at least one subpic_grid_idx[i][j].
[0790] 22. The IBC virtual buffer can be reset before decoding a new sub-image in an image.
[0791] a. In one example, all samples in the IBC virtual buffer can be reset to -1.
[0792] 23. Before decoding a new sub-image in an image, you can reset the palette entry list.
[0793] a. In one example, PredictorPaletteSize can be set to 0 before decoding a new sub-image in an image.
[0794] 24. Whether signaling notifications are sent about information about the stripes (e.g., the number of stripes and / or the extent of the stripes) may depend on the number of slices and / or tiles.
[0795] a. In one example, if the number of tiles in the picture is 1, then num_slices_in_pic_minus1 is not signaled and is inferred to be 0.
[0796] b. In one example, if the number of bricks in the image is 1, then information about the stripes (e.g., the number of stripes and / or the range of stripes) can be communicated without signaling.
[0797] c. In one example, if the number of tiles in the image is 1, then the number of stripes can be inferred as 1. And the stripes cover the entire image. In another example, if the number of tiles in the image is 1, then `single_brick_per_slice_flag` is not signaled and is inferred as 1.
[0798] i. Alternatively, if the number of tiles in the image is 1, then single_brick_per_slice_flag must be 1.
[0799] d. An example syntax design is as follows:
[0800]
[0801] 25. Whether the slice_address is signaled can be independent of whether the slice is signaled as a rectangle (e.g., whether rect_slice_flag is equal to 0 or 1).
[0802] a. An example syntax design is as follows:
[0803] if([[rect_slice_flag||]]NumBricksInPic>1) slice_address u(v)
[0804] 26. When a stripe is signaled as a rectangle, whether to signal the slice_address can depend on the number of stripes.
[0805]
[0806]
[0807] 27. Whether to signal num_bricks_in_slice_minus1 can depend on slice_address and / or the number of tiles in the image.
[0808] a. An example syntax design is as follows:
[0809]
[0810] 28. Whether to signal the loop_filter_across_bricks_enabled_flag can depend on the number of slices and / or tiles.
[0811] a. In one example, if the number of tiles is less than 2, no signaling is sent to loop_filter_across_bricks_enabled_flag.
[0812] b. An example syntax design is as follows:
[0813]
[0814] 29. The requirement for bitstream consistency is that all stripes of an image must cover the entire image.
[0815] a. This requirement must be met when the slice is signaled to be rectangular (e.g., rect_slice_flag equals 1).
[0816] 30. The requirement for bitstream consistency is that all stripes of a sub-image must cover the entire sub-image.
[0817] a. This requirement must be met when the slice is signaled to be rectangular (e.g., rect_slice_flag equals 1).
[0818] 31. The requirement for bitstream consistency is that a stripe cannot overlap with more than one sub-image.
[0819] 32. The requirement for bitstream consistency is that a slice cannot overlap with more than one sub-picture.
[0820] 33. The requirement for bitstream consistency is that a tile cannot overlap with more than one sub-picture.
[0821] In the following discussion, a basic unit block (BUB) with dimensions CW×CH is a rectangular region. For example, a BUB can be a coding tree block (CTB).
[0822] 34. In one example, the number of sub-images (denoted as N) can be signaled.
[0823] a. If subpics are used (e.g., subpics_present_flag equals 1), then at least two subpics may be required in the picture for a consistent bitstream.
[0824] b. Alternatively, N minus d (i.e., Nd) can be signaled, where d is an integer such as 0, 1, or 2.
[0825] c. For example, Nd can be encoded or decoded using fixed-length encoding and decoding, such as u(x).
[0826] i. In one example, x can be a fixed number, such as 8.
[0827] ii. In one example, signaling notification of x or x-dx can precede signaling notification of Nd, where dx is an integer such as 0, 1, or 2. The signaled x can be no greater than the maximum value in the consistent bitstream.
[0828] iii. In one example, x can be exported on the fly.
[0829] 1) For example, x can be derived from the total number of BUBs in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0830] 2) M can be derived as M = Ceiling(W / CW) × Ceiling(H / CH), where W and H represent the width and height of the image, and CW and CH represent the width and height of the BUB.
[0831] d. For example, unary encoding / decoding or truncated unary encoding / decoding can be used to encode / decode Nd.
[0832] e. In one example, the maximum allowed value for Nd can be a fixed number.
[0833] i. Alternatively, the maximum allowed value of Nd can be derived from the total number of BUBs in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0834] 35. In one example, a sub-image can be signaled by one or more selected locations (e.g., top left / top right / bottom left / bottom right positions) and / or an indication of its width and / or its height.
[0835] a. In one example, the top-left position of a sub-image can be signaled at the granularity of a basic unit block (BUB) with dimensions CW×CH.
[0836] i. For example, signaling can be used to inform about the column index (denoted as Col) of the top-left BUB of a sub-image.
[0837] 1) For example, Col-d can be signaled, where d is an integer, such as 0, 1 or 2.
[0838] a) Alternatively, d can be equal to Col of the previously encoded / decoded sub-image plus d1, where d1 is an integer such as -1, 0, or 1.
[0839] b) The symbol of Col-d can be signaled.
[0840] ii. For example, signaling can be used to notify the row index (denoted as Row) of the top-left BUB of a sub-image.
[0841] 1) For example, Row-d can be signaled, where d is an integer, such as 0, 1 or 2.
[0842] a) Alternatively, d can be equal to the Row of the previously encoded / decoded sub-image plus d1, where d1 is an integer such as -1, 0, or 1.
[0843] b) The symbol of Row-d can be signaled.
[0844] iii. The row / column indexes mentioned above (labeled as Row) can be represented in code-decode tree block (CTB) units. For example, the x or y coordinates relative to the top left position of the image can be divided by the CTB size and signaled.
[0845] iv. In one example, whether signaling notifies the location of a sub-image can depend on the sub-image index.
[0846] 1) In one example, for the first sub-image within an image, the top left position may not be notified by signaling.
[0847] a) Alternatively, the top-left position can be inferred, for example, as (0, 0).
[0848] 2) In one example, for the last sub-image within an image, the top left position may not require signaling notification.
[0849] a) The top left position can be inferred from the information of the sub-image in the previous signaling notification.
[0850] b. In one example, the width / height / selected position indication of a sub-image can be signaled using truncated unary / truncated binary / unary / fixed length / Kth EG codec (e.g., K = 0, 1, 2, 3).
[0851] c. In one example, the width of a sub-image can be signaled using a BUB with dimensions CW×CH.
[0852] i. For example, the number of columns (denoted as W) of BUB in a sub-image can be signaled.
[0853] ii. For example, Wd can be signaled, where d is an integer such as 0, 1, or 2.
[0854] 1) Alternatively, d can be equal to W of the previously encoded / decoded sub-image plus d1, where d1 is an integer such as -1, 0, or 1.
[0855] 2) The symbol of Wd can be signaled.
[0856] d. In one example, the height of a sub-image can be signaled using a BUB with dimensions CW×CH.
[0857] i. For example, the row number of BUB in a sub-image can be signaled (denoted as H).
[0858] ii. For example, Hd can be signaled, where d is an integer, such as 0, 1 or 2.
[0859] 1) Alternatively, d can be equal to H of the previously encoded / decoded sub-image plus d1, where d1 is an integer such as -1, 0, or 1.
[0860] 2) The symbol of Hd can be signaled.
[0861] e. In one example, a fixed-length codec can be used to encode and decode Col-d, such as u(x).
[0862] i. In one example, x can be a fixed number, such as 8.
[0863] ii. In one example, signaling can precede signaling to Col-d to x or x-dx, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value in the coherent bitstream.
[0864] iii. In one example, x can be exported on the fly.
[0865] 1) For example, x can be derived from the total number of BUB columns in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0866] 2) M can be derived as M = Ceiling(W / CW), where W represents the width of the image and CW represents the width of the BUB.
[0867] f. In one example, a fixed-length codec can be used to encode and decode Row-d, such as u(x).
[0868] i. In one example, x can be a fixed number, such as 8.
[0869] ii. In one example, signaling x or x-dx can precede signaling Row-d, where dx is an integer such as 0, 1, or 2. The signaled x can be no greater than the maximum value in the consistent bitstream.
[0870] iii. In one example, x can be exported on the fly.
[0871] 1) For example, x can be derived from the total number of BUB rows in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0872] 2) M can be derived as M = Ceiling(H / CH), where H represents the height of the image and CH represents the height of the BUB.
[0873] g. In one example, Wd can be encoded or decoded using a fixed-length codec, such as u(x).
[0874] i. In one example, x can be a fixed number, such as 8.
[0875] ii. In one example, signaling x or x-dx can be notified before signaling Wd, where dx is an integer such as 0, 1, or 2. The signaled x can be no greater than the maximum value in the consistent bitstream.
[0876] iii. In one example, x can be exported on the fly.
[0877] 1) For example, x can be derived from the total number of BUB columns in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0878] 2) M can be derived as M = Ceiling(W / CW), where W represents the width of the image and CW represents the width of the BUB.
[0879] h. In one example, Hd can be encoded and decoded using a fixed-length codec, such as u(x).
[0880] i. In one example, x can be a fixed number, such as 8.
[0881] ii. In one example, signaling notification of x or x-dx can precede signaling notification of Hd, where dx is an integer such as 0, 1, or 2. The signaled notification of x can be no greater than the maximum value in the consistent bitstream.
[0882] iii. In one example, x can be exported on the fly.
[0883] 1) For example, x can be derived from the total number of BUB rows in the image (denoted as M). For example, x = Ceil(log2(M+d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc.
[0884] 2) M can be derived as M = Ceiling(H / CH), where H represents the height of the image and CH represents the height of the BUB.
[0885] i. Signaling notifications can be sent to all sub-pictures, including Col-d and / or Row-d.
[0886] i. Alternatively, it is not necessary to signal Col-d and / or Row-d for all sub-pictures.
[0887] 1) If the number of sub-images is less than 2 (equal to 1), then no signaling is required to notify Col-d and / or Row-d.
[0888] 2) For example, for the first sub-image (e.g., the sub-image index (or sub-image ID) is equal to 0), it is possible not to signal Col-d and / or Row-da. When they are not signaled, they can be inferred to be 0.
[0889] 3) For example, for the last sub-picture (e.g., the sub-picture index (or sub-picture ID) is equal to NumSubPics-1), no signaling notification may be given to Col-d and / or Row-d.
[0890] a) When they are not signaled, they can be inferred from the location and dimensions of the sub-images that have been signaled.
[0891] j. Signaling notifications can be sent to all sub-images for Wd and / or Hd.
[0892] i. Alternatively, signaling notifications to Wd and / or Hd may not be required for all sub-pictures.
[0893] 1) If the number of sub-images is less than 2 (equal to 1), then no signaling is required to notify Wd and / or Hd.
[0894] 2) For example, for the last sub-picture (e.g., the sub-picture index (or sub-picture ID) is equal to NumSubPics-1), no signaling notification may be required for Wd and / or Hd.
[0895] a) When they are not signaled, they can be inferred from the location and dimensions of the sub-images that have been signaled.
[0896] k. In the bullet points above, BUB can be a codec tree block (CTB).
[0897] 36. In one example, the information about the sub-image should be signaled after the CTB size (e.g., log2_ctu_size_minus5) has already been signaled.
[0898] 37. It is not necessary to signal subpic_treated_as_pic_flag[i] for each subpic. Instead, for all subpics, signal a subpic_treated_as_pic_flag to control whether a subpic is treated as a picture.
[0899] 38. It is not necessary to signal loop_filter_across_subpic_enabled_flag[i] for each subpic. Instead, for all subpic, signal a loop_filter_across_subpic_enabled_flag to control whether the loop filter can be applied across subpic.
[0900] 39. Conditional signaling notifications can be sent to subpic_treated_as_pic_flag[i] (subpic_treated_as_pic_flag) and / or loop_filter_across_subpic_enabled_flag[i] (loop_filter_across_subpic_enabled_flag).
[0901] a. In one example, if the number of subpics is less than 2 (equal to 1), then signaling to subpic_treated_as_pic_flag[i] and / or loop_filter_across_subpic_enabled_flag[i] is not required.
[0902] 40. When using sub-images, RPR can be applied.
[0903] a. In one example, when using sub-images, the scaling ratio in the RPR can be constrained to a finite set, such as {1:1, 1:2 and / or 2:1}, or {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1}, {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1, 1:8 and / or 8:1}.
[0904] b. In one example, if image A and image B have different resolutions, then the CTB size of image A and the CTB size of image B can be different.
[0905] c. In one example, suppose a sub-image SA of dimension SAW×SAH is in image A, and a sub-image SB of dimension SBW×SBH is in image B, with SA corresponding to SB. The scaling ratios between image A and image B along the horizontal and vertical directions are Rw and Rh, respectively.
[0906] i.SAW / SBW or SBW / SAW should be equal to Rw.
[0907] ii.SAH / SBH or SBH / SAH should be equal to Rh.
[0908] 41. When using subpicks (e.g., when sub_pics_present_flag is true), the subpick index (or subpick ID) can be signaled in the stripe header, and the stripe address is interpreted as an address within the subpick rather than an address within the entire image.
[0909] 42. If the first sub-image and the second sub-image are not the same sub-image, then the sub-image ID of the first sub-image must be different from the sub-image ID of the second sub-image.
[0910] a. In one example, in a consistent bitstream, if i is not equal to j, then sps_subpic_id[i] must not be equal to sps_subpic_id[j].
[0911] b. In one example, in a consistent bitstream, if i is not equal to j, then pps_subpic_id[i] must not be equal to pps_subpic_id[j].
[0912] c. In one example, in a consistent bitstream, if i is not equal to j, then ph_subpic_id[i] must not be equal to ph_subpic_id[j].
[0913] d. In one example, in a consistent bitstream, if i is not equal to j, then SubpicIdList[i] must not be equal to SubpicIdList[j].
[0914] e. In one example, the signaling notification can be represented as the difference D[i], where D[i] equals X_subpic_id[i] - X_subpic_id[iP].
[0915] i. For example, X can be sps, pps, or ph.
[0916] ii. For example, P equals 1.
[0917] iii. For example, i>P.
[0918] iv. For example, D[i] must be greater than 0.
[0919] v. For example, D[i]-1 can be notified by signaling.
[0920] 43. It is proposed that the length of a syntax element (e.g., subpic_ctu_top_left_x or subpic_ctu_top_left_y) specifying the horizontal or vertical position of the top-left CTU can be deduced as Ceil(Log2(SS)) bits, where SS must be greater than 0. Here, the Ceil() function returns the smallest integer value greater than or equal to the input value.
[0921] a. In one example, when the syntax element specifies the horizontal position of the top-left CTU (e.g., subpic_ctu_top_left_x), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0922] b. In one example, when the syntax element specifies the vertical position of the top-left CTU (e.g., subpic_ctu_top_left_y), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0923] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0924] 44. A syntax element (e.g., subpic_ctu_top_left_x or subpic_ctu_top_left_y) specifying the horizontal or vertical position of the top-left CTU of a subpicture can be deduced as a Ceil(Log2(SS)) bit, where SS must be greater than 0. Here, the Ceil() function returns the smallest integer value greater than or equal to the input value.
[0925] a. In one example, when the syntax element specifies the horizontal position of the top-left CTU of a subpicture (e.g., subpic_ctu_top_left_x), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0926] b. In one example, when the syntax element specifies the vertical position of the top-left CTU of the subpicture (e.g., subpic_ctu_top_left_y), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0927] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0928] 45. The default value (which can be plus an offset P such as 1) for the syntax element that specifies the width or height of a subpicture (e.g., subpic_width_minus1 or subpic_height_minus1) can be derived as Ceil(Log2(SS)) - P, where SS must be greater than 0. Here, the Ceil() function returns the smallest integer value that is greater than or equal to the input value.
[0929] a. In one example, when the syntax element specifies the default width of the subpicture (e.g., subpic_width_minus1) (with an offset P added), SS = (pic_width_max_in_luma_samples + RR) / CtbSizeY.
[0930] b. In one example, when the syntax element specifies the default height of the subpicture (e.g., subpic_height_minus1) (with an offset P added), SS = (pic_height_max_in_luma_samples + RR) / CtbSizeY.
[0931] c. In one example, RR is a non-zero integer, such as CtbSizeY-1.
[0932] 46. It is proposed that if it is determined that the information of the sub-image ID should be signaled, then the information of the sub-image ID should be signaled in at least one of the SPS, PPS and image header.
[0933] a. In one example, if sps_subpic_id_present_flag is equal to 1, then at least one of sps_subpic_id_signalling_present_flag, pps_subpic_id_signalling_present_flag, and ph_subpic_id_signalling_present_flag must be equal to 1 in the consistent bitstream.
[0934] 47. It is proposed that if there is no signaling notification of the sub-image ID in any of the SPS, PPS, and image headers, but it is determined that the signaling should notify of that information, then a default ID should be assigned.
[0935] a. In one example, if ps_subpic_id_signalling_present_flag, pps_subpic_id_signalling_present_flag, and ph_subpic_id_signalling_present_flag are all equal to 0, and sps_subpic_id_present_flag is equal to 1, then SubpicIdList[i] should be set to i+P, where P is an offset such as 0. An exemplary description follows:
[0936] for(i=0;i<=sps_num_subpics_minus1;i++)
[0937] SubpicIdList[i]=sps_subpic_id_present_flag?
[0938] (sps_subpic_id_signalling_present_flag? sps_subpic_id[i]:
[0939] (ph_subpic_id_signalling_present_flag?ph_subpic_id[i]:
[0940] (pps_subpic_id_signalling_present_flag?pps_subpic_id[i]:i) )):i
[0941] 48. It was proposed that if the information of the sub-image ID is notified by signaling in the corresponding PPS, then they should not be notified by signaling in the image header.
[0942] a. An example syntax design is as follows.
[0943]
[0944]
[0945] b. In one example, if the sub-image ID is signaled in the SPS, the sub-image ID is set according to the information of the sub-image ID signaled in the SPS; otherwise, if the sub-image ID is signaled in the PPS, the sub-image ID is set according to the information of the sub-image ID signaled in the PPS; otherwise, if the sub-image ID is signaled in the image header, the sub-image ID is set according to the information of the sub-image ID signaled in the image header. An exemplary description follows.
[0946] for(i=0;i<=sps_num_subpics_minus1;i++)
[0947] SubpicIdList[i]=sps_subpic_id_present_flag?
[0948] (sps_subpic_id_signalling_present_flag? sps_subpic_id[i]:
[0949] (pps_subpic_id_signalling_present_flag?pps_subpic_id[i]:
[0950] (ph_subpic_id_signalling_present_flag?ph_subpic_id[i]:i))):i
[0951] c. In one example, if the sub-image ID is signaled in the image header, the sub-image ID is set according to the information of the sub-image ID signaled in the image header; otherwise, if the sub-image ID is signaled in the PPS, the sub-image ID is set according to the information of the sub-image ID signaled in the PPS; otherwise, if the sub-image ID is signaled in the SPS, the sub-image ID is set according to the information of the sub-image ID signaled in the SPS. An exemplary description follows.
[0952] for(i=0;i<=sps_num_subpics_minus1;i++)
[0953] SubpicIdList[i]=sps_subpic_id_present_flag?
[0954] (ph_subpic_id_signalling_present_flag?ph_subpic_id[i]:
[0955] (pps_subpic_id_signalling_present_flag?pps_subpic_id[i]:
[0956] (sps_subpic_id_signalling_present_flag?sps_subpic_id[i]:i))):i
[0957] 49. It is proposed that the deblocking process on edge E should depend on determining whether loop filtering is allowed on the subpicture boundaries on both sides of the edge (denoted as the P side and the Q side) (e.g., determined by loop_filter_across_subpic_enabled_flag). The P side represents the side in the current block, while the Q side represents the side in the adjacent block, which may belong to different subpictures. In the following discussion, it is assumed that the P side and the Q side belong to two different subpictures. loop_filter_across_subpic_enabled_flag[P] = 0 / 1 means that loop filtering is not allowed / allowed on the subpicture boundaries of the subpicture containing the P side. loop_filter_across_subpic_enabled_flag[Q] = 0 / 1 means that loop filtering is not allowed / allowed on the subpicture boundaries of the subpicture containing the Q side.
[0958] a. In one example, if loop_filter_across_subpic_enabled_flag[P] equals 0 or loop_filter_across_subpic_enabled_flag[Q] equals 0, then E is not filtered.
[0959] b. In one example, if loop_filter_across_subpic_enabled_flag[P] equals 0 and loop_filter_across_subpic_enabled_flag[Q] equals 0, then E is not filtered.
[0960] c. In one example, whether or not filtering is applied to both sides of E is controlled separately.
[0961] i. For example, the P side of E is filtered if and only if loop_filter_across_subpic_enabled_flag[P] is equal to 1.
[0962] ii. For example, the Q side of E is filtered if and only if loop_filter_across_subpic_enabled_flag[Q] is equal to 1.
[0963] 50. It is proposed that the signaling / parsing of syntax elements SE (such as log2_transform_skip_max_size_minus2) in PPS that specify the maximum block size for transformation skipping should be decoupled from any syntax elements (such as sps_transform_skip_enabled_flag) in SPS.
[0964] a. An example syntactic change is as follows:
[0965]
[0966] b. Alternatively, the SE can be notified via signaling in the SPS, such as:
[0967]
[0968] c. Alternatively, the SE can be notified via signaling in the image header, such as:
[0969]
[0970] 51. Whether and / or how the HMVP table (or named list / store / map, etc.) is updated after decoding the first block may depend on whether the first block was encoded and decoded using GEO.
[0971] a. In one example, if the first block is encoded and decoded using GEO, the HMVP table does not need to be updated after decoding the first block.
[0972] b. In one example, if the first block is encoded and decoded using GEO, the HMVP table can be updated after decoding the first block.
[0973] i. In one example, motion information from a partition divided by GEO can be used to update the HMVP table.
[0974] ii. In one example, motion information from multiple partitions divided by GEO can be used to update the HMVP table.
[0975] 52. In CC-ALF, luminance samples outside the current processing unit (e.g., an ALF processing unit defined by two ALF virtual boundaries) are excluded from filtering chrominance samples in the corresponding processing unit.
[0976] a. Fill luminance samples outside the current processing unit can be used to filter chrominance samples in the corresponding processing unit.
[0977] i. Any filling method disclosed in this document can be used to fill luminance samples.
[0978] b. Alternatively, luminance samples outside the current processing unit can be used to filter chrominance samples in the corresponding processing unit.
[0979] Signaling notification of sub-image level parameters
[0980] 53. A parameter set for signaling control of the encoding and decoding behavior of a sub-image is proposed, which can be associated with the sub-image. That is, for each sub-image, a parameter set can be signaled. This parameter set may include:
[0981] a. For inter-frame and / or intra-frame stripes / pictures, the quantization parameter (QP) or QP increment of the luminance component in the sub-picture.
[0982] b. For inter-frame and / or intra-frame stripes / pictures, the quantization parameter (QP) or QP increment of the chroma components in the sub-picture.
[0983] c. Refer to the image list for information management.
[0984] d. Inter-frame and / or intra-frame strip / image CTU size.
[0985] e. Minimum CU size for inter-frame and / or intra-frame stripes / pictures.
[0986] f. Maximum TU size for inter-frame and / or intra-frame stripes / pictures.
[0987] g. Maximum / minimum quadtree (Qual-Tree, QT) partitioning size for inter-frame and / or intra-frame stripes / pictures.
[0988] h. Maximum / minimum quadtree (QT) partitioning depth between frames and / or within frames, for stripes / pictures.
[0989] i. Maximum / minimum binary tree (BT) partitioning size for inter-frame and / or intra-frame stripes / pictures.
[0990] j. Maximum / minimum binary-tree (BT) partitioning depth between frames and / or within frames / strips / pictures.
[0991] k. Maximum / minimum ternary-tree (TT) partitioning size for inter-frame and / or intra-frame stripes / pictures.
[0992] l. Maximum / minimum ternary tree (TT) partitioning depth between frames and / or within frames, for stripes / pictures.
[0993] m. Maximum / minimum multi-tree (MTT) partitioning size for inter-frame and / or intra-frame stripes / pictures.
[0994] n. Maximum / minimum multi-tree (MTT) partitioning depth between frames and / or within frames / strips / pictures.
[0995] o. Control codec tools (including on / off control and / or setting control), including: (see JVET-P2001-v14 for abbreviations).
[0996] i. Weighted prediction
[0997] ii.SAO
[0998] iii.ALF
[0999] iv. Transformation skip
[1000] v.BDPCM
[1001] vi. Joint Cb-Cr Residual Encoding and Decoding (JCCR)
[1002] vii. Reference Encirclement
[1003] viii.TMVP
[1004] ix.sbTMVP
[1005] x.AMVR
[1006] xi.BDOF
[1007] xii.SMVD
[1008] xiii.DMVR
[1009] xiv.MMVD
[1010] xv.ISP
[1011] xvi.MRL
[1012] xvii.MIP
[1013] xviii.CCLM
[1014] xix.CCLM juxtaposition color control
[1015] xx. Intra-frame and / or inter-frame MTS
[1016] xxi. Inter-frame MTS
[1017] xxii.SBT
[1018] xxiii.SBT maximum size
[1019] xxiv. Affine
[1020] xxv. Affine type
[1021] xxvi. Palette
[1022] xxvii.BCW
[1023] xxviii.IBC
[1024] xxix.CIIP
[1025] xxx. Triangle-based motion compensation
[1026] xxxi.LMCS
[1027] p. Any other parameter that has the same meaning as the parameters in VPS / SPS / PPS / Image header / Strip header, but controls the sub-image.
[1028] 54. A flag can be sent first to indicate whether all sub-images share the same parameters.
[1029] a. Alternatively, if the parameters are shared, it is not necessary to notify multiple parameter sets for different sub-image signaling.
[1030] b. Alternatively, if the parameters are not shared, further signaling is required to notify multiple parameter sets of different sub-images.
[1031] 55. Predictive encoding and decoding of parameters between different sub-images can be applied.
[1032] a. In one example, the difference between two values of the same syntax element of two sub-images can be encoded and decoded.
[1033] 56. You can first signal the default parameter set. Then you can further signal the difference between the default value and the default value.
[1034] a. Alternatively, a flag can be signaled first to indicate whether the parameter set of all sub-pictures is the same as the parameter set in the default set.
[1035] 57. In one example, a set of parameters controlling the encoding and decoding behavior of a sub-picture can be signaled in the SPS, PPS, or picture header.
[1036] a. Alternatively, the set of parameters controlling the subpicture encoding / decoding behavior can be signaled in an SEI message (e.g., a subpicture level information SEI message defined in JVET-P2001-v14) or a VUI message.
[1037] 58. In this example, a set of parameters that control the encoding and decoding behavior of a sub-image can be signaled in association with the sub-image ID.
[1038] 59. In one example, signaling can be used to notify a video unit (called SPPS, Subpicture Parameter Set) that is different from the VPS / SPS / PPS / picture header / strip header, which includes a set of parameters that control the encoding and decoding behavior of the subpicture.
[1039] a. In one example, the signaling notification is associated with the SPPS_index.
[1040] b. In one example, the SPPS_index is used for sub-picture signaling to indicate the SPPS associated with the sub-picture.
[1041] 60. In one example, a first control parameter in the parameter set that controls the encoding / decoding behavior of a sub-image can override, or be overridden by, a second control parameter in the same parameter set, but control the same encoding / decoding behavior. For example, an on / off control flag for an encoding / decoding tool such as BDOF in the parameter set of the sub-image can override, or be overridden by, an on / off control flag for an encoding / decoding tool outside the parameter set.
[1042] a. A second control parameter, outside of this parameter set, can be found in VPS / SPS / PPS / Image header / Strip header.
[1043] 61. When applying any of the above examples, the syntax elements associated with a strip / piece / tile / sub-picture depend on the parameters associated with the sub-picture containing the current strip, rather than on the parameters associated with the picture / sequence.
[1044] 62. In a consistent bitstream, the first control parameter in the parameter set that controls the encoding and decoding behavior of a sub-image must be the same as the second control parameter outside the parameter set, but control the same encoding and decoding behavior.
[1045] 63. In one example, signaling notification in SPS uses a first flag, one flag per subpicture, and the first flag specifies whether the subpicture signaling notification associated with the first flag is a general_constraint_info() syntax structure. When a subpicture exists, the general_constraint_info() syntax structure indicates that no tools are applied to the subpicture on CLVS.
[1046] a. Alternatively, a general_constraint_info() syntax structure can be used for each sub-image signaling notification.
[1047] b. Alternatively, the second flag is signaled in the SPS only once, and the second flag specifies whether the first flag exists or does not exist in the SPS for each sub-picture.
[1048] 64. In one example, an SEI message or a VUI parameter is specified to indicate that certain codec tools are not applied or are applied in a specific way to a set of one or more sub-pictures in CLVS (i.e., the codec stripe of the sub-picture set), such that when the sub-picture set is extracted and decoded (e.g., decoded by a mobile device), the decoding complexity is relatively low, and therefore the power consumption of decoding is relatively low.
[1049] a. Alternatively, the same information can be signaled in a DPS, VPS, SPS, or a separate NAL unit.
[1050] Palette Encoding / Decoding
[1051] 65. The maximum number of palette sizes and / or plt prediction sizes can be limited to m*N, for example, N = 8, where m is an integer.
[1052] The value of am or m+offset can be signaled as the first syntax element, where offset is an integer, such as 0.
[1053] i. The first syntax element can be binary-coded using unary encoding, exponential Golomb encoding, Rice encoding, or fixed-length encoding.
[1054] Merge Estimated Region (MER)
[1055] 66. The dimensions of a signalable MER can depend on the maximum or minimum CU or CTU dimensions. The term "dimension" in this document can refer to width, height, width and height, or width × height.
[1056] a. In one example, signaling can be used to notify S-Delta or MS, where S is the size of the MER. Delta and S are integers that depend on the maximum or minimum CU or CTU size.
[1057] For example:
[1058] i. Delta can be the minimum CU or CTU size.
[1059] ii.M can be the maximum CU or CTU size.
[1060] iii. Delta can be the minimum CU or CTU size plus an offset, where the offset is an integer, such as 1 or -1.
[1061] iv.M can be the maximum CU or CTU size plus an offset, where the offset is an integer, such as 1 or -1.
[1062] 67. In a consistent bitstream, the size of the MER may be limited depending on the maximum or minimum size of the CU or CTU. The term "size" in this document may refer to width, height, width and height, or width × height.
[1063] a. For example, the size of the MER is not allowed to be greater than or equal to the maximum CU size or CTU size.
[1064] b. For example, the size of the MER is not allowed to be larger than the maximum CU size or CTU size.
[1065] c. For example, the size of the MER is not allowed to be less than or equal to the minimum CU size or CTU size.
[1066] d. For example, the size of the MER must not be smaller than the minimum CU or CTU size.
[1067] 68. The size of MER can be signaled through indexes.
[1068] The size of a.MER can be mapped to an index using a 1-1 mapping.
[1069] 69. The size of MER or its index can be encoded or decoded using unary codes, exponential Golomb codes, Rice codes, or fixed-length codes.
[1070] Consider adjusting the sample point coordinates of sub-images
[1071] 70. Whether and / or how to clip the vertical coordinate of the juxtaposed sub-blocks within the juxtaposed picture (denoted by yColSb) may depend on subpic_treated_as_pic_flag[SubPicIdx] of the current sub-picture.
[1072] a. When subpic_treated_as_pic_flag[SubPicIdx] of the current sub-picture is equal to 1, the vertical coordinate of the juxtaposed sub-blocks within the juxtaposed picture (denoted by yColSb) is clipped.
[1073] b. In one example, Clip3(T1, T2, yColSb) is used to modify the vertical coordinate of the juxtaposed sub-blocks.
[1074] i. In one example, yColSb can be derived using ySb+tempMv[1], where ySb represents the vertical coordinate of the basic position, which can be the lower-right center sample or the upper-left position of the current coded / decoded sub-block, and tempMv[1] is the offset, which can be derived from the spatially adjacent coded units.
[1075] ii. In one example, T1 is equal to yCtb, and T2 is equal to Min(SubPicBotBoundaryPos, yCtb+(1<<CtbLog2SizeY)-1)), where SubPicBotBoundaryPos represents the vertical coordinate of the bottom boundary of the current sub-picture, which can be equal to Min(pic_height_max_in_luma_samples-1, (subpic_ctu_top_left_y[SubPicIdx]+subpic_height_minus1[SubPicIdx]+1)*CtbSizeY-1).
[1076] ALF and SAO at the boundary between sub-images
[1077] 71. Whether and / or how to apply ALF and / or SAO to the samples in the region covering the boundary between two sub-pictures may depend on the information of the two sub-pictures.
[1078] a. Whether and / or how to apply ALF and / or SAO to the samples in the first sub-picture and the samples in the region covering the boundary between the first sub-picture and the second sub-picture may depend on the information of the second sub-picture.
[1079] b. In one example, if the loop filter is disabled on the subpicture boundary in at least one of the two subpictures (e.g., loop_filter_cross_subpic_enabled_flag equals 0), then ALF and / or SAO are not applied to the samples in the region covering the boundary between the two subpictures.
[1080] c. In one example, ALF and / or SAO are not applied to samples in the region covering the boundary between two subpictures only if the loop filter is disabled at the subpicture boundary between the two subpictures (e.g., loop_filter_cross_subpic_enabled_flag equals 0).
[1081] d. If the boundary is the horizontal boundary between the top sub-image A with bottom row coordinate y0 and the bottom sub-image B with top row coordinate y0+1, then the region includes the sample points in the two sub-images between rows y0-M and y0+1+N, including rows y0-M and y0+1+N, where M and N are integers.
[1082] e. If the boundary is the vertical boundary between the left sub-image A with the rightmost column coordinate equal to x0 and the right sub-image B with the leftmost column coordinate equal to x0+1, then the region includes the sample points in the two sub-images between columns x0-M and x0+1+N, including columns x0-M and x0+1+N, where M and N are integers.
[1083] f. The m and N in the above bullet points can be determined as follows:
[1084] iM and / or N can depend on the color format and / or color components.
[1085] ii. M and / or N can be fixed numbers, such as 1, 2, 3, or 4.
[1086] iii. M and N can be the same.
[1087] iv. M and N can be different.
[1088] v. For SAO and ALF, M and / or N can be set differently.
[1089] vi. Signaling notifications can be sent to M and / or N, such as in VPS / SPS / DPS / PPS / APS / Sequence Header / Picture Header / Slice Header / CTU / CU.
[1090] vii.M can be set to the number of rows (or columns) of samples in sub-image A used in ALF or SAO to filter samples in sub-image B.
[1091] viii.N can be set to the number of rows (or columns) of samples in sub-image B used in ALF or SAO to filter samples in sub-image A.
[1092] 5. Examples
[1093] In the following embodiments, newly added text is in bold italics, and deleted text is marked with "[[]]".
[1094] 5.1 Example 1: Sub-image constraints of Merge candidates constructed by affine model
[1095] 8.5.5.6 Derivation of the Merge candidate for the affine control point motion vector used in construction
[1096] The input to this process is:
[1097] – Specifies the brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.
[1098] – Two variables, cbWidth and cbHeight, specify the width and height of the current luma codec block.
[1099] – Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3,
[1100] – Sample locations (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2) and (xNbB3, yNbB3).
[1101] The output of this process is:
[1102] – The availability flag availableFlagConstK for the candidate affine control point motion vector Merge, where K = 1..6,
[1103] – Refer to the index refIdxLXConstK, where K = 1..6, and X is 0 or 1.
[1104] – The prediction list uses the flag predFlagLXConstK, where K = 1..6 and X is 0 or 1.
[1105] – The affine motion model index is motionModelIdcConstK, where K = 1..6.
[1106] – Bidirectional predictive weighted index bcwIdxConstK, where K = 1..6,
[1107] – The constructed affine control point motion vector is cpMvLXConstK[cpIdx], where cpIdx = 0..2, K = 1..6, and X is 0 or 1.
[1108] …
[1109] The derivation of the fourth (juxtaposed lower right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list using flag predFlagLXCorner[3] and availability flag availableFlagCorner[3] is as follows, where X is 0 and 1:
[1110] – The reference index refIdxLXCorner[3] of the time-domain Merge candidate is set to 0, where X is 0 or 1.
[1111] The derivation of variables mvLXCol and availableFlagLXCol is as follows, where X is 0 or 1:
[1112] – If slice_temporal_mvp_enabled_flag is equal to 0, then set both components of mvLXCol to equal to 0 and set availableFlagLXCol to equal to 0.
[1113] Otherwise (slice_temporal_mvp_enabled_flag equals 1), the following applies:
[1114] xColBr=xCb+cbWidth (8-601)
[1115] yColBr=yCb+cbHeight (8-602)
[1116]
[1117] –If yCb>>CtbLog2SizeY equals yColBr>>CtbLog2SizeY.
[1118] – The variable colCb specifies the coverage by
[1119] ((xColBr>>3)<<3,(yColBr>>3)<<3) The luminance codec block at the given modified position is located within the juxtaposed picture specified by ColPic.
[1120] – The luminance position (xColCb, yColCb) is set to be equal to the top-left sample of the juxtaposed luminance codec module specified by colCb relative to the top-left sample of the juxtaposed image specified by ColPic.
[1121] – Call the derivation process of the juxtaposed motion vector specified in Clause 8.5.2.12, with currCb, colCb, (xColCb, yColCb), refIdxLXCorner[3] and sbFlag set to 0 as input, and assign the output to mvLXCol and availableFlagLXCol.
[1122] Otherwise, both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[1123] …
[1124] 5.2 Example 2: Sub-image constraints of Merge candidates constructed by affine model
[1125] 8.5.5.6 Derivation of the Merge candidate for the affine control point motion vector used for construction. The input to this process is:
[1126] – Specifies the brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.
[1127] – Two variables, cbWidth and cbHeight, specify the width and height of the current luma codec block.
[1128] – Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3,
[1129] – Sample locations (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2) and (xNbB3, yNbB3).
[1130] The output of this process is:
[1131] – The availability flag availableFlagConstK for the candidate affine control point motion vector Merge, where K = 1..6,
[1132] – Refer to the index refIdxLXConstK, where K = 1..6, and X is 0 or 1.
[1133] – The prediction list uses the flag predFlagLXConstK, where K = 1..6 and X is 0 or 1. – The affine motion model index motionModelIdcConstK, where K = 1..6.
[1134] – Bidirectional predictive weighted index bcwIdxConstK, where K = 1..6,
[1135] – The constructed affine control point motion vector is cpMvLXConstK[cpIdx], where cpIdx = 0..2, K = 1..6, and X is 0 or 1.
[1136] …
[1137] The derivation of the motion vector cpMvLXCorner[3] of the fourth (the lower right corner of the juxtaposition), the reference index refIdxLXCorner[3], the prediction list using the flag predFlagLXCorner[3], and the availability flag availableFlagCorner[3] is as follows, where X is 0 and 1:
[1138] – The reference index refIdxLXCorner[3] of the time-domain Merge candidate is set to 0, where X is 0 and 1.
[1139] The derivation of variables mvLXCol and availableFlagLXCol is as follows, where X is 0 and 1:
[1140] – If slice_temporal_mvp_enabled_flag is equal to 0, then set both components of mvLXCol to equal to 0 and set availableFlagLXCol to equal to 0.
[1141] Otherwise (slice_temporal_mvp_enabled_flag equals 1), the following applies:
[1142] ColBr=xCb+cbWidth (8-601)
[1143] yColBr=yCb+cbHeight (8-602)
[1144]
[1145] – If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, [[yColBr is less than pic_height_in_luma_samples, xColBr is less than pic_width_in_luma_samples, then this applies]]:
[1146] The variable colCb specifies the luminance codec block that covers the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the juxtaposed picture specified by ColPic.
[1147] – The luminance position (xColCb, yColCb) is set to be equal to the top-left sample of the juxtaposed luminance codec module specified by colCb relative to the top-left sample of the juxtaposed image specified by ColPic.
[1148] – Call the derivation process of the juxtaposed motion vector specified in Clause 8.5.2.12, with currCb, colCb, (xColCb, yColCb), refIdxLXCorner[3] and sbFlag set to 0 as input, and assign the output to mvLXCol and availableFlagLXCol.
[1149] Otherwise, both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[1150] …
[1151] 5.3 Example 3: Extracting Integer Samples under Sub-Image Constraints
[1152] 8.5.6.3.3.1 Process of acquiring integer samples of brightness
[1153] The input to this process is:
[1154] – Brightness position in units of the entire sample (xInt) L yInt L ),
[1155] –Luminance reference sample array refPicLXL,
[1156] The output of this process is the predicted luminance sample value, predSampleLX. L
[1157] The variable shift is set to equal Max(2, 14-BitDepth). Y ).
[1158] The variable picW is set to equal pic_width_in_luma_samples, and the variable picH is set to equal pic_height_in_luma_samples.
[1159] The brightness positions (xInt, yInt) per unit of the entire sample are derived as follows:
[1160]
[1161] xInt=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag? (8-782)
[1162] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L ):xInt L )
[1163] yInt = Clip3(0, picH-1, yInt) L )
[1164] (8-783)
[1165] Predicted brightness sample value predSampleLX L The derivation is as follows:
[1166] predSampleLX L =refPicLX L [xInt][yInt]< <shift3 (8-784)
[1167] 5.4 Example 4: Derivation of the variable invAvgLuma in LMCS chromatic residual scaling
[1168] 8.7.5.3 Image Reconstruction Using Luminance-Correlated Chromaticity Residual Scaling Processing of Chromaticity Samples
[1169] The input to this process is:
[1170] – The chromaticity position (xCurr, yCurr) of the top-left chromaticity sample of the current chromaticity transform block relative to the top-left chromaticity sample of the current image.
[1171] – The variable nCurrSw specifies the width of the chroma transform block.
[1172] – The variable nCurrSh specifies the height of the chroma transform block.
[1173] – Specifies the codec block flags for the current chroma transform block.
[1174] – Specifies the (nCurrSw)×(nCurrSh) array predSamples for the chromaticity prediction samples of the current block.
[1175] – Specifies the (nCurrSw)×(nCurrSh) array resSamples for the chromaticity residual samples of the current block.
[1176] The output of this process is a reconstructed array of chroma image samples, recSamples.
[1177] The variable sizeY is set to equal Min(CtbSizeY,64).
[1178] For i = 0..nCurrSw 1, j = 0..nCurrSh 1, the reconstructed chromaticity image samples
[1179] The derivation of recSamples is as follows:
[1180] –…
[1181] – Otherwise, apply the following:
[1182] –…
[1183] – The variable curripic specifies an array of reconstructed brightness samples in the current image.
[1184] – For the derivation of the variable varScale, the following ordered steps are applied:
[1185] 1. The derivation of the variable invAvgLuma is as follows:
[1186] The derivation of array recLuma[i] and variable cnt is as follows, where i = 0..(2*sizeY-1):
[1187] – The variable cnt is set to equal to 0.
[1188]
[1189] – When availL equals TRUE, the array recLuma[i] (where i = 0..sizeY-1) is set to equal to Where i = 0..sizeY – 1, and cnt is set to equal sizeY.
[1190] – When availT equals TRUE, the array recLuma[cnt+i] (where i = 0..sizeY–1) is set to equal to Where i = 0..sizeY – 1, and cnt is set to equal to (cnt + sizeY).
[1191] The derivation of the variable invAvgLuma is as follows:
[1192] – If cnt is greater than 0, then apply the following:
[1193] invAvgLuma = Clip1 Y ((+(cnt>>1))>>Log2(cnt)) (8-1013)
[1194] Otherwise (cnt equals 0), apply the following:
[1195] invAvgLuma=1<<(BitDepth Y –1) (8-1014)
[1196] 5.5 Example 5: An example of defining sub-image elements using N (e.g., N=8 or 32) samples in addition to the 4 samples.
[1197] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics
[1198] `subpic_grid_col_width_minus1` plus 1 specifies the width of each element in the subpic identifier grid, in order to... The unit is sample points. The length of a syntax element is... Bit.
[1199] The derivation of the variable NumSubPicGridCols is as follows:
[1200]
[1201] Add 1 to specify the height of each element in the sub-image identifier grid, in units of 4 samples. The length of the syntax element is...
[1202] Bit.
[1203] The derivation of the variable NumSubPicGridRows is as follows:
[1204]
[1205] 7.4.7.1 General Strip Header Semantics
[1206] The derivation of variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos is as follows:
[1207]
[1208]
[1209] 5.6 Example 6: Limit the image width and image height to equal to or greater than 8.
[1210] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics
[1211] Specifies the maximum width in luminance samples for each decoded image from the reference SPS. `pic_width_max_in_luma_samples` should not be equal to 0 and should be [[MinCbSizeY]]. Integer multiples of.
[1212] Specifies the maximum height in luma samples for each decoded image from the reference SPS. `pic_height_max_in_luma_samples` should not be equal to 0 and should be [[MinCbSizeY]]. Integer multiples of.
[1213] 5.7 Example 7: Sub-image boundary inspection for signaling of BT / TT / QT partitioning, BT / TT / QT depth derivation and / or CU partitioning flags
[1214] 6.4.2 Permissible Binary Partitioning Processes
[1215] The derivation of the variable allowBtSplit is as follows:
[1216] –…
[1217] Otherwise, allowBtSplit will be set to FALSE if all of the following conditions are true.
[1218] –btSplit equals Split_BT_VER
[1219] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]
[1220] Otherwise, allowBtSplit will be set to FALSE if all of the following conditions are true.
[1221] –btSplit equals Split_BT_VER
[1222] –cbHeight is greater than MaxTbSizeY
[1223] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]
[1224] Otherwise, allowBtSplit will be set to FALSE if all of the following conditions are true.
[1225] –btSplit equals Split_BT_HOR
[1226] –cbWidth is greater than MaxTbSizeY
[1227] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]
[1228] Otherwise, allowBtSplit will be set to FALSE if all of the following conditions are true.
[1229] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]
[1230] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]
[1231] –cbWidth is greater than minQtSize
[1232] Otherwise, allowBtSplit will be set to FALSE if all of the following conditions are true.
[1233] –btSplit equals Split_BT_HOR
[1234] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]
[1235] –y0+cbHeight is less than or equal to [[pic_height_in_luma_samples]]
[1236] 6.4.3 Permissible Ternary Partitioning Procedures
[1237] The derivation of the variable allowTtSplit is as follows:
[1238] – If one or more of the following conditions are true, allowTtSplit will be set to equal FALSE:
[1239] –cbSize is less than or equal to 2*MinTtSizeY
[1240] –cbWidth is greater than Min(MaxTbSizeY,maxTtSize)
[1241] –cbHeight is greater than Min(MaxTbSizeY,maxTtSize)
[1242] –mttDepth is greater than or equal to maxMttDepth
[1243] –x0+cbWidth is greater than [[pic_width_in_luma_samples]]
[1244]
[1245] –y0+cbHeight is greater than [[pic_height_in_luma_samples]]
[1246] –treeType equals DUAL_TREE_CHROMA, and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32
[1247] –treeType equals DUAL_TREE_CHROMA, modeType equals MODE_TYPE_INTRA
[1248] Otherwise, allowTtSplit is set to TRUE.
[1249] 7.3.8.2 Encoder-decoder tree unit syntax
[1250]
[1251] 7.3.8.4 Encoder-decoder tree syntax
[1252]
[1253]
[1254] 5.8 Example 8: Example of defining a sub-image
[1255]
[1256]
[1257] 5.9 Example 9: Example of defining a sub-image
[1258]
[1259]
[1260] 5.10 Example 10: Example of defining a sub-image
[1261]
[1262]
[1263] 5.11 Example 11: Example of defining a sub-image
[1264]
[1265]
[1266]
[1267] 5.12 Implementation Example: Considering the removal of blocks from sub-images
[1268] 8.8.3 Deblocking Filtering Process
[1269] 8.8.3.1 Overview
[1270] The input to this process is the reconstructed image before the removal of blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1271] The output of this process is the reconstructed image after removing the blocks, i.e., the array recPictureL, and the array recPicture when ChromaArrayType is not equal to 0.Cb and recPicture Cr .
[1272] First, the vertical edges in the image are filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the image are filtered. Vertical and horizontal edges in the CTB of each CTU are processed separately on a codec unit basis. Vertical edges of the codec blocks in a codec unit are filtered starting from the left edge of the codec block and proceeding geometrically towards the right edge of the codec block. Horizontal edges of the codec blocks in a codec unit are filtered, starting from the top edge of the codec block and proceeding geometrically towards the bottom edge of the codec block.
[1273] Note – Although the filtering process is specified on an image-based basis in this specification, the filtering process can also be implemented on an encoding / decoding unit-based basis with equivalent results, provided that the decoder correctly considers the processing dependency order to produce the same output value.
[1274] The deblocking filtering process is applied to all encoded and decoded sub-block edges and transform block edges of the image, except for the following types of edges:
[1275] – Edges on the image boundary
[1276] – [[Edges that coincide with the boundary of the subpic when loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0]]
[1277] – When PPS_loop_filter_cross_virtual_boundaries_disabled_flag equals 1, the edges that coincide with the virtual boundaries of the image.
[1278] – When loop_filter_cross_tiles_enabled_flag equals 0, edges coinciding with tile boundaries
[1279] – When loop_filter_cross_slices_enabled_flag equals 0, the edge coinciding with the slice boundary.
[1280] – When slice_deblocking_filter_disabled_flag equals 1, the edge coinciding with the top or left boundary of the slice.
[1281] The inner edge of the slice when –slice_deblocking_filter_disabled_flag is equal to 1
[1282] – Edges that do not correspond to the 4×4 sample grid boundary of the brightness component.
[1283] – Edges that do not correspond to the boundaries of the 8×8 sample grid for the chromaticity components
[1284] – Edges in the luminance component where intra_bdpcm_luma_flag is equal to 1 on both sides.
[1285] – Edges in the chroma component where intra_bdpcm_chroma_flag is equal to 1 on both sides.
[1286] – The edge of a chromatic sub-block that is not the edge of a correlated transform unit.
[1287] …
[1288] Deblocking filtering process in one direction
[1289] The input to this process is:
[1290] – Specifies the variable `treeType` to indicate whether the current processing is for the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA).
[1291] – When treeType equals DUAL_TREE_LUMA, the reconstructed image before removing the blocks, i.e., the array recPicture. L ,
[1292] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPicture Cr ,
[1293] – Specifies the variable edgeType to filter whether the edge is vertical (EDGE_VER) or horizontal (EDGE_HOR).
[1294] The output of this process is the reconstructed image after removing the blocks, i.e.:
[1295] – When treeType equals DUAL_TREE_LUMA, the array recPicture L
[1296] – When ChromaArrayType is not equal to 0 and treeType is equal to DUAL_TREE_CHROMA, the array recPicture Cb and recPictureCr .
[1297] The derivation of variables firstCompIdx and lastCompIdx is as follows:
[1298] firstCompIdx=(treeType==DUAL_TREE_CHROMA)? 1:0 (8-1010)
[1299] lastCompIdx=(treeType==DUAL_TREE_LUMA||ChromaArrayType==0)? 0:2(8-1011)
[1300] For each codec unit and each codec block of each color component indicated by the color component index cIdx, having a codec block width nCbW, a codec block height nCbH, and the position (xCb, yCb) of the top-left sample point of the codec block, where cIdx ranges from firstCompIdx to lastCompIdx, inclusive of first compidx and lastCompIdx, the edges are filtered by the following ordered steps when cIdx equals 0, or when cIdx is not equal to 0 and edgeType equals EDGE_VER and xCb%8 equals 0, or when cIdx is not equal to 0 and edgeType equals EDGE_HOR and yCb%8 equals 0:
[1301] 2. The derivation of the variable filterEdgeFlag is as follows:
[1302] – If edgeType equals EDGE_VER, and one or more of the following conditions are true, then filterEdgeFlag is set to 0:
[1303] – The left boundary of the current encoding / decoding block is the left boundary of the image.
[1304] – [[The left boundary of the current codec block is either the left or right boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1305] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[1306] – The left boundary of the current codec block is the left boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[1307] – The left boundary of the current codec block is one of the vertical virtual boundaries of the image, and VirtualBoundariesDisabledFlag is equal to 1.
[1308] Otherwise, if edgeType equals EDGE_HOR, and one or more of the following conditions are true, then the variable filterEdgeFlag is set to 0:
[1309] – The top boundary of the current luminance codec block is the top boundary of the image.
[1310] – [[The top boundary of the current codec block is either the top or bottom boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1311] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_tiles_enabled_flag is equal to 0.
[1312] – The top boundary of the current codec block is the top boundary of the slice, and loop_filter_cross_slices_enabled_flag is equal to 0.
[1313] – The top boundary of the current codec block is one of the horizontal virtual boundaries of the image, and VirtualBoundariesDisabledFlag is equal to 1.
[1314] Otherwise, filterEdgeFlag is set to 1.
[1315] …
[1316]
[1317] The inputs to this process include:
[1318] –Sample value p i and q i Where i = 0..3,
[1319] -p i and q i Location, (xP) i yP i ) and (xQ i yQ i ), where i = 0..2,
[1320] – Variable dE,
[1321] – Variables dEp and dEq contain the decisions regarding filtering sample points p1 and q1, respectively.
[1322] – Variable t C .
[1323] The output of this process is:
[1324] – The number of filtered samples, nDp and nDq
[1325] –Filtered sample value p i 'and q j ', where i = 0..nDp-1, j = 0..nDq–1.
[1326] Depending on the value of dE, apply the following procedure.
[1327] – If variable dE equals 2, then nDp and nDq are both set to equal 3, and the following strong filtering is applied:
[1328] p0′=Clip3(p0-3*t C p0+3*t C ,(p2+2*p1+2*p0+2*q0+q1+4)>>3) (8-1150)
[1329] p1′=Clip3(p1-2*t C p1+2*t C ,(p2+p1+p0+q0+2)>>2) (8-1151)
[1330] p2′=Clip3(p2-1*t C p2+1*t C ,(2*p3+3*p2+p1+p0+q0+4)>>3) (8-1152)
[1331] q0′=Clip3(q0-3*t C ,q0+3*t C ,(p1+2*p0+2*q0+2*q1+q2+4)>>3) (8-1153)
[1332] q1′=Clip3(q1-2*t C ,q1+2*t C ,(p0+q0+q1+q2+2)>>2) (8-1154)
[1333] q2′=Clip3(q2-1*t C ,q2+1*t C,(p0+q0+q1+3*q2+2*q3+4)>>3) (8-1155)
[1334] Otherwise, both nDp and nDq are set to 0, and the following weak filter is applied:
[1335] – Apply the following:
[1336] Δ=(9*(q0-p0)-3*(q1-p1)+8)>>4 (8-1156)
[1337] – When Abs(Δ) is less than t C When *10 is reached, the following ordered steps shall be applied:
[1338] – The filtered sample values p0' and q0' are specified as follows:
[1339] Δ=Clip3(-t C ,t C ,Δ) (8-1157)
[1340] p0′=Clip1(p0+Δ) (8-1158)
[1341] q0′=Clip1(q0-Δ) (8-1159)
[1342] – When dEp equals 1, the filtered sample value p1' is specified as follows:
[1343] Δp=Clip3(-(t C >>1),t C >>1,(((p2+p0+1)>>1)-p1+Δ)>>1) (8-1160)
[1344] p1′=Clip1(p1+Δp) (8-1161)
[1345] – When dEq equals 1, the filtered sample value q1' is specified as follows:
[1346] Δq=Clip3(-(t C >>1),t C >>1,(((q2+q0+1)>>1)-q1-Δ)>>1) (8-1162)
[1347] q1′=Clip1(q1+Δq) (8-1163)
[1348] –nDp is set to equal to dEp+1, and nDq is set to equal to dEq+1. nDp is set to equal to 0 when nDp is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample p0 is equal to 1.
[1349] When nDq is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample q0 is equal to 1, nDq is set to 0.
[1350]
[1351] The input to this process is:
[1352] – Variables maxFilterLengthP and maxFilterLengthQ,
[1353] –Sample p i and q j Where i = 0..maxFilterLengthP and j = 0..maxFilterLengthQ,
[1354] -p i and q j Location (xP) i yP i ) and (xQ j yQ j ), where i = 0..maxFilterLengthP-1 and j = 0..maxFilterLengthQ-1,
[1355] – Variable t C .
[1356] The output of this process is:
[1357] –Filtered sample values p i 'and q j ', where i = 0..maxFilterLengthP-1 and j = 0..maxFilterLengthQ-1.
[1358] The derivation of the variable refMiddle is as follows:
[1359] – If maxFilterLengthP equals maxFilterLengthQ and maxFilterLengthP equals 5, then apply the following:
[1360] refMiddle=(p4+p3+2*(p2+p1+p0+q0+q1+q2)+q3+q4+8)>>4 (8-1164)
[1361] Otherwise, if maxFilterLengthP equals maxFilterLengthQ and
[1362] If maxFilterLengthP is not equal to 5, then the following applies:
[1363] refMiddle=(p6+p5+p4+p3+p2+p1+2*(p0+q0)+q1+q2+q3+q4+q5+q6+8)>>4 (8-1165)
[1364] Otherwise, if one of the following conditions is true,
[1365] –maxFilterLengthQ equals 7 and maxFilterLengthP equals 5.
[1366] –maxFilterLengthQ equals 5 and maxFilterLengthP equals 7.
[1367] Apply the following:
[1368] refMiddle=(p5+p4+p3+p2+2*(p1+p0+q0+q1)+q2+q3+q4+q5+8)>>4 (8-1166)
[1369] Otherwise, if one of the following conditions is true,
[1370] –maxFilterLengthQ equals 5 and maxFilterLengthP equals 3.
[1371] –maxFilterLengthQ equals 3 and maxFilterLengthP equals 5.
[1372] Apply the following:
[1373] refMiddle=(p3+p2+p1+p0+q0+q1+q2+q3+4)>>3 (8-1167)
[1374] Otherwise, if maxFilterLengthQ equals 7 and maxFilterLengthP equals 3, then
[1375] Apply the following:
[1376] refMiddle=(2*(p2+p1+p0+q0)+p0+p1+q1+q2+q3+q4+q5+q6+8)>>4 (8-1168)
[1377] – Otherwise, apply the following:
[1378] refMiddle=(p6+p5+p4+p3+p2+p1+2*(q2+q1+q0+p0)+q0+q1+8)>>4 (8-1169)
[1379] The derivation of variables refP and refQ is as follows:
[1380] refP = (p maxFilterLengtP +p maxFilterLengthP-1 +1)>>1 (8-1170)
[1381] refQ=(q maxFilterLengtQ +q maxFilterLengthQ-1 +1)>>1 (8-1171)
[1382] variable f i and t C PD i The definition is as follows:
[1383] – If maxFilterLengthP equals 7, then apply the following:
[1384] f 0..6 ={59,50,41,32,23,14,5} (8-1172)
[1385] t C PD 0..6 ={6,5,4,3,2,1,1} (8-1173)
[1386] Otherwise, if maxFilterLengthP equals 5, then apply the following:
[1387] f 0..4 ={58,45,32,19,6} (8-1174)
[1388] t C PD 0..4 ={6,5,4,3,2} (8-1175)
[1389] – Otherwise, apply the following:
[1390] f 0..2 ={53,32,11} (8-1176)
[1391] t CPD 0..2 ={6,4,2} (8-1177)
[1392] variable g j and t C QD j The definition is as follows:
[1393] – If maxFilterLengthQ equals 7, then apply the following:
[1394] g 0..6 ={59,50,41,32,23,14,5}
[1395] (8-1178)
[1396] t C QD 0..6 ={6,5,4,3,2,1,1} (8-1179)
[1397] Otherwise, if maxFilterLengthQ equals 5, then apply the following:
[1398] g 0..4 ={58,45,32,19,6} (8-1180)
[1399] t C QD 0..4 ={6,5,4,3,2} (8-1181)
[1400] – Otherwise, apply the following:
[1401] g 0..2 ={53,32,11} (8-1182)
[1402] t C QD 0..2 ={6,4,2} (8-1183)
[1403] Filtered sample value p i 'and q j The derivation of ' is as follows, where i = 0..maxFilterLengthP-1 and j = 0..maxFilterLengthQ-1:
[1404] p i =Clip3(p i -(t C *t C PD i )>>1,p i +(t C *t C PD i)>>1,(refMiddle*f i +refP*(64-f i (+32)>>6) (8-1184)
[1405] q j =Clip3(q) j -(t C *t C QD j )>>1,q j +(t C *t C QD j )>>1,(refMiddle*g j +refQ*(64-g j )+32)>>6) (8-1185)
[1406] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthP-1.
[1407] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The corresponding input sample value q j Instead, where j = 0..maxFilterLengthQ-1.
[1408]
[1409]
[1410] This procedure will only be invoked if ChromaArrayType is not equal to 0.
[1411] The input to this process is:
[1412] – Variable maxFilterLength,
[1413] –Color sample value p i and q i Where i = 0..maxFilterLengthCbCr,
[1414] -p i and q i chromaticity position (xP)i yP i ) and (xQ i yQ i ), where i = 0..maxFilterLengthCbCr-1,
[1415] – Variable t C .
[1416] The output of this process is the filtered sample value p. i 'and q i ', where i = 0..maxFilterLengthCbCr-1.
[1417] Filtered sample value p i 'and q i The derivation of ' is as follows, where i = 0..maxFilterLengthCbCr-1:
[1418] – If maxFilterLengthCbCr equals 3, then apply the following strong filter:
[1419] p0′=Clip3(p0-t C ,p0+t C ,(p3+p2+p1+2*p0+q0+q1+q2+4)>>3) (8-1186)
[1420] p1′=Clip3(p1-t C p1+t C ,(2*p3+p2+2*p1+p0+q0+q1+4)>>3) (8-1187)
[1421] p2′=Clip3(p2-t C p2+t C ,(3*p3+2*p2+p1+p0+q0+4)>>3) (8-1188)
[1422] q0′=Clip3(q0-t C ,q0+t C ,(p2+p1+p0+2*q0+q1+q2+q3+4)>>3) (8-1189)
[1423] q1′=Clip3(q1-t C ,q1+t C ,(p1+p0+q0+2*q1+q2+2*q3+4)>>3) (8-1190)
[1424] q2′=Clip3(q2-t C,q2+t C ,(p0+q0+q1+2*q2+3*q3+4)>>3) (8-1191)
[1425] Otherwise, apply the following weak filtering:
[1426] Δ=Clip3(-t C ,t C ,((((q0-p0)<<2)+p1-q1+4)>>3)) (8-1192)
[1427] p0′=Clip1(p0+Δ) (8-1193)
[1428] q0′=Clip1(q0-Δ) (8-1194)
[1429] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthCbCr-1.
[1430] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The defined input sample value q i Instead, where i = 0..maxFilterLengthCbCr-1:
[1431]
[1432] 5.13 Implementation Example: Considering Block Removal from Sub-images (Solution #2)
[1433] 8.8.3 Deblocking Filtering Process
[1434] 8.8.3.1 Overview
[1435] The input to this process is the reconstructed image before the removal of blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1436] The output of this process is the reconstructed image after removing the blocks, i.e., the array recPicture. LAnd when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1437] …
[1438] The deblocking filtering process is applied to all encoded and decoded sub-block edges and transform block edges of the image, except for the following types of edges:
[1439] – Edges on the image boundary
[1440] – [[Edges that coincide with the boundaries of subpicks where loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0]]
[1441]
[1442] – When VirtualBoundariesDisabledFlag equals 1, the edge that coincides with the virtual boundary of the image.
[1443] –…
[1444] 8.8.3.2 Deblocking Filtering Process in One Direction
[1445] The input to this process is:
[1446] – Specifies the variable `treeType` to indicate whether the current processing is for the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA).
[1447] …
[1448] 3. The derivation of the variable filterEdgeFlag is as follows:
[1449] – If edgeType equals EDGE_VER, and one or more of the following conditions are true, then filterEdgeFlag is set to 0:
[1450] – The left boundary of the current encoding / decoding block is the left boundary of the image.
[1451] – [[The left boundary of the current codec block is either the left or right boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1452]
[1453] –…
[1454] Otherwise, if edgeType equals EDGE_HOR, and one or more of the following conditions are true, then the variable filterEdgeFlag is set to 0:
[1455] – The top boundary of the current luminance codec block is the top boundary of the image.
[1456] – [[The top boundary of the current codec block is either the top or bottom boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1457]
[1458] 5.14 Implementation Example: Considering Block Removal from Sub-images (Solution #3)
[1459] 8.8.3 Deblocking Filtering Process
[1460] 8.8.3.1 Overview
[1461] The input to this process is the reconstructed image before the removal of blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1462] The output of this process is the reconstructed image after removing the blocks, i.e., the array recPicture. L And when ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1463] …
[1464] The deblocking filtering process is applied to all encoded and decoded sub-block edges and transform block edges of the image, except for the following types of edges:
[1465] – Edges on the image boundary
[1466] – [[Edges that coincide with the boundaries of subpicks where loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0]]
[1467] – When VirtualBoundariesDisabledFlag equals 1, the edge that coincides with the virtual boundary of the image.
[1468] –…
[1469] 8.8.3.2 Deblocking Filtering Process in One Direction
[1470] The input to this process is:
[1471] – Specifies the variable `treeType` to indicate whether the current processing is for the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA).
[1472] –…
[1473] 4. The derivation of the variable filterEdgeFlag is as follows:
[1474] – If edgeType equals EDGE_VER, and one or more of the following conditions are true, then filterEdgeFlag is set to 0:
[1475] – The left boundary of the current encoding / decoding block is the left boundary of the image.
[1476] – [[The left boundary of the current codec block is either the left or right boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1477] –…
[1478] Otherwise, if edgeType equals EDGE_HOR, and one or more of the following conditions are true, then the variable filterEdgeFlag is set to 0:
[1479] – The top boundary of the current luminance codec block is the top boundary of the image.
[1480] – [[The top boundary of the current codec block is either the top or bottom boundary of the subpic, and loop_filter_cross_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1481] –…
[1482] 8.8.3.6.6 Filtering process for luminance samples using a short filter
[1483] …
[1484] When nDp is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample p0 is equal to 1, nDp is set to 0.
[1485] When nDq is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample q0 is equal to 1, nDq is set to 0.
[1486]
[1487] 8.8.3.6.7 Filtering process for luminance samples using a long filter
[1488] …
[1489] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthP-1.
[1490] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The corresponding input sample value q j Instead, where j = 0..maxFilterLengthQ-1.
[1491]
[1492] 8.8.3.6.9 Filtering process for chromaticity samples
[1493] …
[1494] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthP-1.
[1495] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The defined input sample value q i Instead, where i = i = 0..maxFilterLengthQ–1.
[1496]
[1497]
[1498] 5.15 Implementation Example: Considering Block Removal from Sub-images (Solution #4)
[1499] 8.8.3.6.6 Filtering process for luminance samples using a short filter
[1500] When nDp is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample p0 is equal to 1, nDp is set to 0.
[1501] When nDq is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample q0 is equal to 1, nDq is set to 0:
[1502]
[1503] 8.8.3.6.7 Filtering process for luminance samples using a long filter
[1504] …
[1505] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthP-1.
[1506] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The defined input sample value q i Instead, where i = i = 0..maxFilterLengthQ–1.
[1507]
[1508]
[1509] 8.8.3.6.9 Filtering process for chromaticity samples
[1510] …
[1511] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p iInstead, where i = 0..maxFilterLengthP-1.
[1512] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The defined input sample value q i Instead, where i = i = 0..maxFilterLengthQ–1.
[1513]
[1514] 5.16 Example: Derivation of Temporal Merging Candidates Based on Sub-Blocks
[1515] 8.5.5.3 Derivation of Sub-Block-Based Temporal Merging Candidates
[1516] …
[1517] The derivation of the positions (xColSb, yColSb) of the juxtaposed sub-blocks within –ColPic is as follows.
[1518] –[[Applicable to the following:]]
[1519]
[1520]
[1521] yColSb=Clip3(yCtb,Min(pic_height_in_luma_samples-1,yCtb+(1< <CtbLog2SizeY)-1),(735)ySb+tempMv[1])
[1522] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1523] xColSb=Clip3(xCtb,Min(SubPicRightBoundaryPos,xCtb+(1< <CtbLog2SizeY)+3),(736)xSb+tempMv[0])
[1524] – Otherwise (subpic_treated_as_pic_flag[subpicidx] equals 0), the following applies:
[1525] xColSb=Clip3(xCtb,Min(pic_width_in_luma_samples-1,xCtb+(1< <CtbLog2SizeY)+3),(737)xSb+tempMv[0])
[1526] …
[1527] 8.5.5.4 Derivation of Sub-Block-Based Temporal Merging Fundamental Motion Data
[1528] …
[1529] The derivation of the positions (xColSb, yColSb) of the juxtaposed sub-blocks within ColPic is as follows.
[1530] –[[Applicable to the following:]]
[1531]
[1532] yColCb=Clip3(yCtb,Min(pic_height_in_luma_samples-1,yCtb+(1< <CtbLog2SizeY)-1),(742)yColCtrCb+tempMv[1])
[1533] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1534] xColCb=Clip3(xCtb,Min(SubPicRightBoundaryPos,xCtb+(1< <CtbLog2SizeY)+3),(743)xColCtrCb+tempMv[0])
[1535] – Otherwise (subpic_treated_as_pic_flag[subpicidx] equals 0), the following applies:
[1536] xColCb=Clip3(xCtb,Min(pic_width_in_luma_samples-1,xCtb+(1< <CtbLog2SizeY)+3),(744)xColCtrCb+tempMv[0])
[1537] Figure 3This is a block diagram of a video processing apparatus 300. Apparatus 300 can be used to implement one or more methods described herein. Apparatus 300 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 300 may include one or more processors 312, one or more memories 314, and video processing hardware 316. Processor 312 can be configured to implement one or more methods described in this document. Memory 314 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 316 can be used to implement some of the techniques described in this document in hardware circuitry.
[1538] Figure 4 This is a flowchart of a method 400 for processing video. Method 400 includes, for a video block in a first video region of the video, determining (402) whether the location of a temporal motion vector prediction value determined by the conversion between the video block and the bitstream representation of the current video block using an affine mode is within a second video region, and performing a conversion (404) based on this determination.
[1539] In some embodiments, the following solutions may be implemented as preferred solutions.
[1540] The following solutions can be implemented in conjunction with other technologies described in the projects listed in the previous sections (e.g., Project 1).
[1541] 1. A video processing method, comprising: for a video block in a first video region of a video, determining whether the location of a temporal motion vector prediction value determined by a conversion between the video block and a bitstream representation of the current video block using an affine mode is within a second video region; and performing a conversion based on the determination.
[1542] 2. The method according to Solution 1, wherein the video block is covered by a first region and a second region.
[1543] 3. The method according to any one of solutions 1-2, wherein if the location of the temporal motion vector prediction value is outside the second video region, the temporal motion vector prediction value is marked as unavailable and is not used in the conversion.
[1544] The following solutions can be implemented in conjunction with other technologies described in the projects listed in the previous sections (e.g., Project 2).
[1545] 4. A video processing method, comprising: for a video block in a first video region of a video, determining whether the position of an integer sample in a reference image extracted for conversion between the bitstream representation of the video block and the current video block is in a second video region, wherein the reference image is not used for an interpolation process during the conversion; and performing the conversion based on the determination.
[1546] 5. The method according to Solution 4, wherein the video block is covered by a first region and a second region.
[1547] 6. According to the method of any one of solutions 4-5, in the case where the sample is located outside the second video region, the sample is marked as unavailable and is not used in the conversion.
[1548] The following solutions can be implemented in conjunction with other technologies described in the projects listed in the previous sections (e.g., Project 3).
[1549] 7. A video processing method, comprising: for a video block in a first video region of a video, determining whether the location of a reconstructed luminance sample value extracted by a conversion between the bitstream representation of the video block and the current video block is within a second video region; and performing a conversion based on the determination.
[1550] 8. The method according to solution 7, wherein the brightness sample is covered by a first region and a second region.
[1551] 9. The method according to any one of solutions 7-8, wherein if the location of the luminance sample is outside the second video region, the luminance sample is marked as unavailable and is not used in the conversion.
[1552] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 4).
[1553] 10. A video processing method, comprising: for a video block in a first video region of a video, determining whether the location of a segmentation-related check, depth derivation, or segmentation flag signaling of the video block is within a second video region during a conversion between the video block and a bitstream representation of the current video block; and performing a conversion based on the determination.
[1554] 11. The method according to solution 10, wherein the location is covered by a first region and a second region.
[1555] 12. The method according to any one of solutions 10-11, wherein, in the case where the location is outside the second video region, the luminance sample is marked as unavailable and is not used in the conversion.
[1556] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 8).
[1557] 13. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a codec representation of the video, the video images comprising one or more video blocks, wherein the codec representation conforms to the codec syntax requirements of the conversion not using sub-image encoding / decoding within video units and dynamic precision conversion encoding / decoding tools or reference image resampling tools.
[1558] 14. The method according to solution 13, wherein the video unit corresponds to a sequence of one or more video images.
[1559] 15. The method according to any one of solutions 13-14, wherein the dynamic precision conversion encoding / decoding tool includes an adaptive precision conversion encoding / decoding tool.
[1560] 16. The method according to any one of solutions 13-14, wherein the dynamic precision conversion encoding / decoding tool includes a dynamic precision conversion encoding / decoding tool.
[1561] 17. The method according to any one of solutions 13-16, wherein the encoding / decoding representation indicates that the video unit conforms to the encoding / decoding syntax requirements.
[1562] 18. The method according to solution 17, wherein the encoding / decoding indicates that the video unit uses sub-picture encoding / decoding.
[1563] 19. The method according to solution 17, wherein the encoding / decoding indicates that the video unit uses a dynamic precision conversion encoding / decoding tool or a reference image resampling tool.
[1564] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 10).
[1565] 20. The method according to any one of solutions 1-19, wherein the second video region comprises video sub-pictures, and wherein the boundary between the second video region and the other video region is also the boundary between two codec tree units.
[1566] 21. The method according to any one of solutions 1-19, wherein the second video region comprises video sub-pictures, and wherein the boundary between the second video region and the other video region is also the boundary between two codec tree units.
[1567] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 11).
[1568] 22. The method according to any one of solutions 1-21, wherein the first video region and the second video region have a rectangular shape.
[1569] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 12).
[1570] 23. The method according to any one of solutions 1-22, wherein the first video region and the second video region do not overlap.
[1571] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 13).
[1572] 24. The method according to any one of solutions 1-23, wherein the video image is divided into video regions such that pixels in the video image are covered by one and only one video region.
[1573] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 15).
[1574] 25. The method according to any one of solutions 1-24, wherein the video image is divided into a first video region and a second video region because the video image is located in a specific layer of the video sequence.
[1575] The following solutions can be implemented in conjunction with other technologies described in the projects listed in previous sections (e.g., Project 10).
[1576] 26. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a codec representation of the video, the video images comprising one or more video blocks, wherein the codec representation conforms to the codec syntax requirement that a first syntax element subpic_grid_idx[i][j] is not greater than a second syntax element max_subpics_minus1.
[1577] 27. The method according to solution 26, wherein the codeword representing the first syntax element is not greater than the codeword representing the second syntax element.
[1578] 28. The method according to any one of solutions 1-27, wherein the first video region includes a video sub-picture.
[1579] 29. The method according to any one of solutions 1-28, wherein the second video region includes a video sub-picture.
[1580] 30. The method according to any one of solutions 1 to 29, wherein the conversion includes encoding the video into a codec representation.
[1581] 31. The method according to any one of solutions 1 to 29, wherein the conversion includes decoding the encoding / decoding representation to generate pixel values of the video.
[1582] 32. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 31.
[1583] 33. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 31.
[1584] 34. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 31.
[1585] 35. The methods, apparatus or systems described in this document.
[1586] Figure 13 This is a block diagram illustrating an example video processing system 1300 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1300. System 1300 may include an input 1302 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1302 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[1587] System 1300 may include codec component 1304, which can implement the various codec or encoding methods described in this document. Codec component 1304 can reduce the average bit rate of the video from input 1302 to the output of codec component 1304 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1306, the output of codec component 1304 can be stored or transmitted via connected communication. Component 1308 can use the stored or transmitted bitstream (or encoded) representation of the video received at input 1302 to generate pixel values or displayable video sent to display interface 1310. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of the encoded decryption results, will be performed by the decoder.
[1588] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[1589] Figure 14 This is a block diagram illustrating an example video encoding / decoding system 100 that can utilize the techniques disclosed herein.
[1590] like Figure 14 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.
[1591] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[1592] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems used to generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[1593] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[1594] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.
[1595] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.
[1596] Figure 15 This is a block diagram illustrating an example of a video encoder 200. The video encoder 200 can be... Figure 14 The video encoder 114 in the system 100 shown.
[1597] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 9 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[1598] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (including a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[1599] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[1600] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretative purposes... Figure 9 The example is shown separately.
[1601] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[1602] The mode selection unit 203 can, for example, select a codec mode—intra-frame or inter-frame—based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the codec block for use as a reference image. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the precision of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[1603] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.
[1604] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[1605] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[1606] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in list 0 or list 1 that contain the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[1607] In some examples, the motion estimation unit 204 can output complete motion information for the decoder's decoding processing.
[1608] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[1609] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[1610] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[1611] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[1612] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[1613] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[1614] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[1615] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[1616] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[1617] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in buffer 213.
[1618] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[1619] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[1620] Figure 16 This is a block diagram illustrating an example of a video decoder 300. The video decoder 300 can be... Figure 14 The video decoder 114 in the system 100 shown.
[1621] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 10 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[1622] exist Figure 16 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 15 The decoding process is the inverse of the encoding process described.
[1623] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list index, and other motion information, from the entropy-decoded video data. The motion compensation unit 302 can determine this information, for example, by executing AMVP and Merge modes.
[1624] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The syntax elements can include identifiers of the interpolation filters to be used with sub-pixel precision.
[1625] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of video blocks, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[1626] The motion compensation unit 302 can use some syntax information to determine the size of the blocks of frames and / or stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame encoded block, and other information for decoding the encoded video sequence.
[1627] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[1628] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form the decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.
[1629] Figure 17 This is a flowchart representation of a video processing method according to the present technology. Method 1700 includes, in operation 1710, a conversion between a current video block in a current frame of the video and the bitstream of the video, determining how to modify the y-coordinate yColSb of a juxtaposed sub-block within a juxtaposed image of the current frame based on whether the sub-frame is considered a frame. The juxtaposed image is one of one or more reference images of the current frame. Method 1700 further includes, in operation 1720, performing the conversion based on the determination.
[1630] In some embodiments, the bitstream includes a syntax element that includes a flag indicating whether a subpicture is considered a picture for transformation. In some embodiments, in the case where the flag indicates that the subpicture is considered a picture for transformation, the method includes cropping the vertical coordinate of the collocated sub-blocks for transformation. In some embodiments, the vertical coordinate yColSb of the collocated sub-blocks is modified by a function Clip3(T1, T2, yColSb), where T1 and T2 are real numbers. In some embodiments, yColSb is determined based on ySb + tempMv[1]. ySb represents the vertical coordinate of the base position, which is the lower right or upper left position of the current sub-block corresponding to the collocated sub-block, and tempMv[1] is an offset determined based on the spatially adjacent coded and decoded units of the current video block. In some embodiments, T1 is equal to yCtb, and T2 is equal to Min(SubPicBotBoundaryPos, yCtb+(1<<CtbLog2SizeY)-1)). yCtb represents the vertical coordinate of the coded tree block, SubPicBotBoundaryPos represents the vertical coordinate of the bottom boundary of the current subpicture, and CtbLog2SizeY represents the vertical dimension of the coded tree block. In some embodiments, SubPicBotBoundaryPos is equal to Min(pic_height_max_in_luma_samples-1, (subpic_ctu_top_left_y[SubPicIdx]+subpic_height_minus1[SubPicIdx]+1)*CtbSizeY-1), where pic_height_max_in_luma_samples indicates the maximum height of the current picture, subpic_ctu_top_left_y indicates the upper left vertical position of the current subpicture, subpic_height_minus1 indicates the height of the current subpicture, and CtbSizeY indicates the vertical dimension of the coded tree block.
[1631] Figure 18 is a flowchart representation of a video processing method according to the present technology. Method 1800 includes, in operation 1810, for the transformation between the current picture of a video including at least two subpictures and the bitstream of the video, determining a way to apply a filtering operation to a region covering the boundary between the two subpictures based on information of the two subpictures. Method 1800 further includes, in operation 1820, performing the transformation according to the determination.
[1632] In some embodiments, the filtering operation includes an adaptive loop filtering operation or a sample adaptive offset (SAO) filtering operation. In some embodiments, the method of applying the filtering operation to the region and the first sub-image of the two sub-images is determined based on information from the second sub-image of the two sub-images. In some embodiments, if the filtering operation is disabled on at least one sub-image boundary of the two sub-images, the filtering operation is not applied to the region. In some embodiments, the boundary is a horizontal boundary. The first sub-image is located above the second sub-image having a bottom row coordinate y0, the second sub-image having a top row coordinate y0+1, and the region includes samples located between rows y0-M and y0+1+N, where M and N are integers. In some embodiments, the boundary is a vertical boundary. The first sub-image is located to the left of the second sub-image having a rightmost column x0, the second sub-image having a leftmost column x0+1, and the region includes samples located between columns x0-M and x0+1+N, where M and N are integers. In some embodiments, at least M or N is determined based on the color format or color components of the video. In some embodiments, at least M or N is a fixed number. In some embodiments, M and N are the same. In some embodiments, M and N are different. In some embodiments, M and N are different for the filtering operation. In some embodiments, signaling in the bitstream indicates at least M or N. In some embodiments, M is equal to the number of rows or columns of samples in the first sub-image used for filtering samples in the second sub-image. In some embodiments, N is equal to the number of rows or columns of samples in the second sub-image used for filtering samples in the first sub-image.
[1633] In some embodiments, the conversion generates video from a bitstream. In some embodiments, the conversion generates a bitstream from video.
[1634] In one example, a method for storing a bitstream of video includes a conversion between a current video block in a current frame of the video and the bitstream of the video, determining how to modify the y-coordinate yColSb of a juxtaposed sub-block within a juxtaposed image of the current frame based on whether the sub-frame is considered a frame. The juxtaposed image is one of one or more reference images of the current frame. The method also includes generating a bitstream of video from the current video block based on this determination and storing the bitstream in a non-transitory computer-readable recording medium.
[1635] In another example, a method for storing a bitstream of video includes a conversion between a current image of a video comprising at least two sub-images and the bitstream of that video, determining, based on information from the two sub-images, how to apply a filtering operation to a region covering the boundary between the two sub-images. The method also includes generating a bitstream of video based on this determination and storing the bitstream in a non-transitory computer-readable recording medium.
[1636] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from video blocks to a bitstream representation of the video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on that decision or determination.
[1637] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode when converting video blocks into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that it has not been modified using a video processing tool or mode enabled based on the decision or determination.
[1638] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this application can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-volatile computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that influences machine-readable propagated signals, or one or more of these. The terms "data processing unit" or "data processing apparatus" include all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computer groups. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[1639] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed and executed on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.
[1640] The processes and logic flows described in this application can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuits, and the apparatus can also be implemented as special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[1641] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as one or more of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by, or merged into, special-purpose logic circuitry.
[1642] While this patent document contains numerous details, it should not be construed as limiting the scope of any invention or claim, but rather as a description of features of specific embodiments of a particular invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in certain circumstances, one or more features from a combination of claims may be removed from the combination, and a combination of claims may refer to a sub-combination or a variation of a sub-combination.
[1643] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence shown to perform such operations, or all the described operations, in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[1644] Only some implementations and examples are described. Other implementations, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A video processing method, comprising: For the conversion between the current video block in the current image of the video and the bitstream of the video, the method of modifying the ordinate yColSb of the juxtaposed sub-block within the juxtaposed image of the current image is determined based on whether the sub-image is considered an image, where, The collocated picture is one of one or more reference pictures of the current picture; And Performing the conversion based on the determination, wherein the vertical coordinate yColSb of the collocated sub-block is modified by the function Clip3(T1, T2, yColSb), where T1 and T2 are real numbers, where T1 is equal to yCtb, and T2 is equal to Min(SubPicBotBoundaryPos, yCtb + (1 << CtbLog2SizeY) - 1)), where yCtb represents the vertical coordinate of the coding tree block, where SubPicBotBoundaryPos represents the vertical coordinate of the bottom boundary of the current sub-picture, and where CtbLog2SizeY represents the vertical dimension of the coding tree block.
2. The method according to claim 1, wherein, The bitstream includes a syntax element, and the syntax element includes a flag indicating whether the sub-picture is regarded as the picture of the conversion.
3. The method according to claim 2, wherein, If the flag indicates that the sub-picture is regarded as the picture of the conversion, the method includes cropping the vertical coordinate of the converted collocated sub-block.
4. The method according to claim 1, wherein, yColSb is determined based on ySb + tempMv[1], where ySb represents the vertical coordinate of the basic position, the basic position is the lower right position or the upper left position of the current sub-block corresponding to the collocated sub-block, and where tempMv[1] is an offset determined based on the spatially adjacent coding units of the current video block.
5. The method according to claim 1, wherein, SubPicBotBoundaryPos is equal to Min(pic_height_max_in_luma_samples - 1, (subpic_ctu_top_left_y[SubPicIdx] + subpic_height_minus1[SubPicIdx] + 1) * CtbSizeY - 1), where pic_height_max_in_luma_samples indicates the maximum height of the current picture, subpic_ctu_top_left_y indicates the upper left vertical position of the current sub-picture, subpic_height_minus1 indicates the height of the current sub-picture, and CtbSizeY indicates the vertical dimension of the coding tree block.
6. The method according to claim 1, further comprising: For the conversion between the current picture of the video including at least two sub-pictures and the bitstream of the video, determining a manner of applying a filtering operation to a region covering the boundary between the two sub-pictures based on the information of the two sub-pictures.
7. The method according to claim 6, wherein, The filtering operation includes an adaptive loop filtering operation or a sample adaptive offset (SAO) filtering operation.
8. The method according to claim 6, wherein, The manner of applying the filtering operation to the region and the first sub-picture of the two sub-pictures is determined based on the information of the second sub-picture of the two sub-pictures.
9. The method according to claim 6, wherein, When the filtering operation is disabled on the sub-picture boundary in at least one of the two sub-pictures, the filtering operation is not applied to the region.
10. The method according to claim 6, wherein, The boundary is a horizontal boundary, where a first sub-picture is above a second sub-picture having a bottom row coordinate y0, where the second sub-picture has a top row coordinate y0 + 1, and where the region includes samples located between rows y0 - M and y0 + 1 + N, M and N being integers.
11. The method according to claim 6, wherein, The boundary is a vertical boundary, where a first sub-picture is to the left of a second sub-picture having a rightmost column x0, where the second sub-picture has a leftmost column x0 + 1, and where the region includes samples located between columns x0 - M and x0 + 1 + N, M and N being integers.
12. The method according to claim 10 or 11, wherein, At least M or N is determined based on the color format or color component of the video.
13. The method of claim 10 or 11, wherein, At least M or N is a fixed number.
14. The method of claim 10 or 11, wherein, M is the same as N.
15. The method of claim 10 or 11, wherein, M is different from N.
16. The method according to claim 10 or 11, wherein, M and N are different for the filtering operation.
17. The method according to claim 10 or 11, wherein, At least M or N is signaled in the bitstream.
18. The method according to claim 10 or 11, wherein, M is equal to the number of rows or columns of samples in the first sub-picture used to filter the samples in the second sub-picture.
19. The method according to claim 10 or 11, wherein, N is equal to the number of rows or columns of samples in the second sub-picture used to filter the samples in the first sub-picture.
20. The method according to any one of claims 1 to 11, wherein, The conversion generates the video from the bitstream.
21. The method according to any one of claims 1 to 11, wherein, The conversion generates the bitstream from the video.
22. A method for storing a video bitstream, comprising: For the conversion between the current video block in the current image of the video and the bitstream of the video, the method of modifying the ordinate yColSb of the juxtaposed sub-block within the juxtaposed image of the current image is determined based on whether the sub-image is considered an image, where, The collocated picture is one of one or more reference pictures of the current picture; Based on the determination, generating the bitstream of the video from a current video block; And Storing the bitstream in a non-transitory computer-readable recording medium, where the vertical coordinate yColSb of the collocated sub-block is modified by a function Clip3(T1, T2, yColSb), where T1 and T2 are real numbers, where T1 is equal to yCtb, and T2 is equal to Min(SubPicBotBoundaryPos, yCtb + (1 << CtbLog2SizeY) - 1)), where yCtb represents the vertical coordinate of a coding tree block, where SubPicBotBoundaryPos represents the vertical coordinate of the bottom boundary of the current sub-picture, and where CtbLog2SizeY represents the vertical dimension of the coding tree block.
23. The method according to claim 22, further comprising: For the conversion between a current picture of the video including at least two sub-pictures and the bitstream of the video, determining a way of applying a filtering operation to a region covering a boundary between the two sub-pictures based on information of the two sub-pictures.
24. A video processing apparatus, comprising a processor configured to implement the method according to any one or more of claims 1 to 23.
25. A computer-readable medium having code stored thereon, which when executed, causes a processor to implement the method according to any one or more of claims 1 to 23.
26. A computer-readable medium storing a bitstream generated according to any one of claims 1 to 23.
Citation Information
Patent Citations
Method and apparatus for sub-picture-based image encoding / decoding, and method for transmitting bitstream
CN114450943A