Limiting slice width in video coding
By employing adaptive resolution changes and reference image resampling techniques, the problem of low resolution adjustment efficiency in video encoding and decoding is solved, achieving more efficient video processing and seamless switching, making it suitable for multi-party video conferencing and streaming media applications.
Patent Information
- Application Number
- CN202080090129.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-03
- Filing Date
- 2020-12-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2040-12-28
AI Technical Summary
Existing video encoding and decoding technologies suffer from low efficiency, high resource consumption, and increased latency when processing videos of different resolutions. This is especially true in multi-party video conferencing and streaming media applications, where it is difficult to achieve flexible resolution adjustment and switching.
Adaptive Resolution Change (ARC) technology is employed to achieve the conversion between video blocks and bitstreams through Reference Picture Resampling (RPR) and consistency window management. This includes the selective application of interpolation filters and deblocking filters, and the dynamic adjustment of the encoding and decoding process based on the relationship between resolution and size.
It improves the efficiency and flexibility of the video encoding and decoding process, reduces resource consumption and latency, and supports resolution adjustment for active speakers in multi-party video conferences and quick start and seamless switching in streaming media.
Smart Images

Figure CN115176462B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] In accordance with the applicable Patent Law and / or the Paris Convention, this application promptly claims priority and interest in PCT application PCT / CN2019 / 129069, filed on December 27, 2019, and PCT application PCT / CN2020 / 083315, filed on April 3, 2020. For all legal purposes, the entire disclosure of the foregoing applications is incorporated herein by reference as a part of this application disclosure. Technical Field
[0003] This patent document relates to video encoding and decoding technologies, equipment, and systems. Background Technology
[0004] Currently, efforts are underway to improve the performance of existing video codec technologies to provide better compression rates or to offer video encoding and decoding schemes that allow for lower complexity or parallelization. Industry experts have recently proposed several new video codec tools, which are currently being tested to determine their effectiveness. Summary of the Invention
[0005] This paper describes devices, systems, and methods related to digital video encoding and decoding, particularly those related to motion vector management. The described methods can be applied to existing video encoding and decoding standards (e.g., High Efficiency Video Codec (HEVC) or Universal Video Codec) as well as future video encoding and decoding standards or codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between blocks of video within a video unit and a bitstream of video. The bitstream conforms to a rule specifying the maximum dimension of the video unit, wherein the dimension of the video unit includes the width, height, or number of samples within the video unit. The maximum dimension is known prior to the conversion.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current image of the video and a bitstream of the video. The current image is associated with one or more reference images such that (1) the dimensions of the current image are PW×PH, (2) the scaling window dimensions of the current image are SW×WH, (3) the scaling window dimensions of the reference images are SW'×SH', and (4) the maximum permissible dimensions Wmax and Hmax of the images satisfy constraints.
[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0009] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0010] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0011] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0012] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0013] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a picture of a video and a bitstream of the video. The bitstream conforms to a rule that specifies signaling quantization parameter information in a picture header and excludes in a picture parameter set.
[0014] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a coded representation of the current video block, wherein, during the conversion, a deblocking filter is selectively applied if a resolution and / or a size of a reference picture and a resolution and / or a size of the current video block are not the same, wherein a strength of the deblocking filter is set according to a rule related to the resolution and / or the size of the reference picture relative to the resolution and / or the size of the current video block.
[0015] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a coded representation of the current video block, wherein, during the conversion, a reference picture of the current video block is resampled according to a rule that is based on a dimension of the current video block.
[0016] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a coded representation of the current video block, wherein, during the conversion, a deblocking filter is selectively applied if a resolution and / or a size of a reference picture and a resolution and / or a size of the current video block are not the same, wherein a strength of the deblocking filter is set according to a rule related to the resolution and / or the size of the reference picture relative to the resolution and / or the size of the current video block.
[0017] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a plurality of video blocks and a coded representation of the plurality of video blocks, wherein, during the conversion, a first consistency window for a first video block is defined and a second consistency window for a second video block is defined, and wherein a ratio of a width and / or a height of the first consistency window and the second consistency window is according to a rule that is based at least on a consistent bitstream.
[0018] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a plurality of video blocks and a coded representation of the plurality of video blocks, wherein, during the conversion, a first consistency window for a first video block is defined and a second consistency window for a second video block is defined, and wherein a ratio of a width and / or a height of the first consistency window and the second consistency window is according to a rule that is based at least on a consistent bitstream.
[0019] In another representative aspect, a method of video processing is disclosed. The method includes determining, for a video segment of a video and a coded representation of the video, that the coded representation excludes syntax elements related to P / B slices as the video segment includes all I slices, and performing a conversion according to the determination.
[0020] Furthermore, in a representative aspect, an apparatus for a video system is disclosed, the apparatus comprising a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to implement any or more of the disclosed methods.
[0021] In one representative aspect, a video decoding apparatus is disclosed, which includes a processor configured to implement the methods disclosed herein.
[0022] In one representative aspect, a video encoding apparatus is disclosed, which includes a processor configured to implement the methods disclosed herein.
[0023] In addition, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product including program code for performing any one or more of the disclosed methods.
[0024] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, description and claims. Attached Figure Description
[0025] Figure 1 An example of the sub-block motion vector (VSB) and motion vector difference is shown.
[0026] Figure 2 An example of a 16×16 video block divided into 16 4×4 regions is shown.
[0027] Figure 3A An example of a specific location in the sample points is shown.
[0028] Figure 3B Another example of a specific location in the sample points is shown.
[0029] Figure 3C Another example of a specific location within the sample points is shown.
[0030] Figure 4A An example showing the location of the current sample point and its reference sample point is provided.
[0031] Figure 4B Another example is shown showing the location of the current sample point and its reference sample point.
[0032] Figure 5 This is a block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding techniques described in this document.
[0033] Figure 6 A flowchart of an example method for video encoding and decoding is shown.
[0034] Figure 7An example of decoder-side motion vector refinement is shown.
[0035] Figure 8 This example demonstrates a flow chart for cascading DMVR and BDOF procedures in VTM 5.0. The DMVR SAD operation and the BDOF SAD operation are different and not shared.
[0036] Figure 9 This is a block diagram of an example video processing system that can implement the disclosed technology.
[0037] Figure 10 This is a block diagram illustrating an example video encoding / decoding system.
[0038] Figure 11 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0039] Figure 12 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0040] Figure 13 This is a flowchart representation of the video processing method based on this technology.
[0041] Figure 14 This is a flowchart representation of another video processing method based on this technology.
[0042] Figure 15 This is a flowchart representation of another video processing method based on this technology.
[0043] Figure 16 This is a flowchart representation of another video processing method based on this technology.
[0044] Figure 17 This is a flowchart representation of another video processing method based on this technology. Detailed Implementation
[0045] 1. Video encoding and decoding in HEVC / H.265
[0046] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). In April 2018, a joint video expert team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to develop the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.
[0047] 2. Summary
[0048] 2.1 Adaptive Resolution Change (ARC)
[0049] AVC and HEVC do not have the ability to change resolution without introducing IDR or Intra-Random Access Point (IRAP) images; such a capability can be represented as Adaptive Resolution Change (ARC). Several use cases or application scenarios will benefit from the ARC feature, including the following:
[0050] - Rate Adaptation in Video Calls and Conferencing: To adapt codec video to changing network conditions, the encoder can adapt to reduced bandwidth by encoding lower-resolution images when network conditions worsen and available bandwidth becomes lower. Currently, image resolution changes can only be made after IRAP images; this presents several problems. Reasonably high-quality IRAP images will be significantly larger than inter-frame codec images and correspondingly more complex to decode: this will consume time and resources. This is also problematic if the resolution change is requested by the decoder for load reasons. It may also disrupt low-latency buffering conditions, forcing audio resynchronization and increasing end-to-end latency of the stream, at least temporarily. This can result in a poor user experience.
[0051] - Active Speaker Changes in Multi-Party Video Conferencing: In multi-party video conferencing, it's common to display the active speaker at a larger video size compared to the video used for the remaining participants. When the active speaker changes, it may also be necessary to adjust the image resolution for each participant. The need for ARC (Automatic Representation) features becomes more important when such changes in the active speaker occur frequently.
[0052] - Fast Start in Streaming: In streaming applications, it's common for the application to buffer images of a certain length before starting to display. Starting the bitstream at a lower resolution allows the application to have enough images in the buffer, thus starting to display faster.
[0053] Adaptive Streaming Switching in Streaming Media: Dynamic Adaptive Streaming (DASH) based on the HTTP specification includes a feature called `@mediaStreamStructureId`. This enables switching between different representations at random access points in an open GOP with an undecodeable leading picture (e.g., a CRA picture associated with a RASL picture in HEVC). When two different representations of the same video have different bitrates but the same spatial resolution, and simultaneously have the same `@mediaStreamStructureId` value, switching between these two representations can be performed at the CRA picture associated with the RASL picture, and the RASL picture associated with the switch at the CRA picture can be decoded at acceptable quality, thus achieving seamless switching. With the help of ARC, the `@mediaStreamStructureId` feature can also be used for switching between DASH representations with different spatial resolutions.
[0054] ARC is also known as Dynamic Resolution Conversion.
[0055] ARC can be viewed as a special case of reference image resampling (RPR) (e.g., H.263 Appendix P).
[0056] 2.2 Resampling of reference images in Appendix P of H.263
[0057] This mode describes an algorithm that warps a reference image before using it for prediction. It can be used to resample a reference image with a source format different from the image being predicted. It can also be used for global motion estimation or rotational motion estimation by warping the shape, size, and position of the reference image. The syntax includes the warping parameters to be used and the resampling algorithm. The simplest level of operation for the reference image resampling mode is resampling with an implicit factor of 4, where only an FIR filter needs to be applied to the upsampling and downsampling processes. In this case, no additional signaling overhead is required when the size of the new image (indicated in the image header) differs from the size of the previous image, as its purpose is understandable.
[0058] 2.3 Consistency Window in VVC
[0059] In VVC, the consistency window is defined as a rectangle. Samples within the consistency window belong to the image of interest. Samples outside the consistency window are discarded during output.
[0060] When applying a consistency window, the scaling ratio in the RPR is derived based on the consistency window.
[0061] Image Parameter Set RBSP Syntax
[0062]
[0063] `pic_width_in_luma_samples` specifies the width of each decoded image, in units of luminance samples, referencing the PPS. `pic_width_in_luma_samples` should not be equal to 0, should be an integer multiple of `Max(8, MinCbSizeY)`, and should be less than or equal to `pic_width_max_in_luma_samples`.
[0064] When subpics_present_flag equals 1, the value of pic_width_in_luma_samples should be equal to pic_width_max_in_luma_samples.
[0065] `pic_height_in_luma_samples` specifies the height of each decoded image, in units of luminance samples, relative to the PPS. `pic_height_in_luma_samples` should not be equal to 0 and should be an integer multiple of `Max(8, MinCbSizeY)`, and should be less than or equal to `pic_height_max_in_luma_samples`.
[0066] When subpics_present_flag equals 1, the value of pic_height_in_luma_samples should be equal to pic_height_max_in_luma_samples.
[0067] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference images of the current image referencing this PPS, respectively. Bitstream consistency requirements satisfy all of the following conditions:
[0068]
[0069] A conformance_window_flag value of 1 indicates that the conformance window offset parameter immediately follows in SPS. A conformance_window_flag value of 0 indicates that the conformance window offset parameter does not exist.
[0070] Based on the rectangular region specified in the image coordinates used for output, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset define the image samples in the CVS output during the decoding process. When conformance_window_flag equals 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.
[0071] The consistent cropping window contains luminance samples with horizontal image coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical image coordinates (including end values) from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1).
[0072] The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) should be less than pic_height_in_luma_samples.
[0073] The derivation of variables PicOutputWidthL and PicOutputHeightL is as follows:
[0074] PicOutputWidthL=pic_width_in_luma_samples- (7-43)
[0075] SubWidthC*(conf_win_right_offset+conf_win_left_offset)
[0076] PicOutputHeightL=pic_height_in_pic_size_units- (7-44)
[0077] SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)
[0078] When ChromaArrayType is not equal to 0, the corresponding specified sample points of the two chroma arrays are sample points with image coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the image coordinates of the specified brightness sample point.
[0079] NOTE - The conformance clipping window offset parameters only apply to the output. All internal decoding processes will be applied to the unclipped picture size. Figure 1
[0080] Let ppsA and ppsB be any two PPSs referencing the same SPS. The requirement for bitstream consistency is that when the values of pic_width_in_luma_samples and pic_height_in_luma_samples of ppsA and ppsB are the same, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset of ppsA and ppsB are also the same.
[0081] 2.4 Reference Image Resampling (RPR)
[0082] In some embodiments, ARC is also referred to as Reference Image Resampling (RPR). With RPR, TMVP is disabled if the juxtaposed image has a different resolution than the current image. Furthermore, BDOF and DMVR are disabled when the reference image has a different resolution than the current image.
[0083] To handle general mediating (MC) when the reference image has a different resolution than the current image, the interpolation part is defined as follows:
[0084] 8.5.6.3 Fractional Sample Interpolation Process
[0085] 8.5.6.3.1 Overview
[0086] The input for this process is:
[0087] – Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image.
[0088] – The variable sbWidth specifies the width of the current encoding / decoding sub-block.
[0089] – The variable sbHeight specifies the height of the current encoding / decoding sub-block.
[0090] – Motion vector offset mvOffset
[0091] – Refined motion vector refMvLX,
[0092] – The selected reference image sample array refPicLX,
[0093] – Half-sample interpolation filter index hpelIfIdx,
[0094] – Bidirectional optical flow flag bdofFlag
[0095] – The variable cIdx specifies the color component index of the current block.
[0096] The output of this process is:
[0097] – The array predSamplesLX represents the predicted sample values (sbWidth+brdExtSize)x(sbHeight+brdExtSize).
[0098] The predicted block border extension size brdExtSize is derived as follows:
[0099] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&sps_affine_prof_enabled_flag))? 2:0(8-752)
[0100] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[0101] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[0102] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[0103] – If cIdx equals 0, then the following applies:
[0104] - Define the scaling factor and its fixed-point representation as:
[0105] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[0106] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL(8-754)
[0107] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[0108] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L ) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[0109] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L=0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[0110] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[0111] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)
[0112] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)
[0113] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp(8-757)
[0114] refy L =((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6(8-758)
[0115] – The variable xInt is derived as follows L yInt L xFrac L and yFrac L :
[0116] xInt L =refx L >>4 (8-759)
[0117] yInt L =refyL >>4 (8-760)
[0118] xFrac L =refx L &15 (8-761)
[0119] yFrac L =refy L &15 (8-762)
[0120] – If bdofFlag is true or (sps_affine_prof_enabled_flag is true and inter_affine_flag[xSb][ySb] is true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L With >>3)-1) and refPicLX as inputs, the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[0121] -x L It equals 0.
[0122] -x L It equals sbWidth+1.
[0123] –y L It equals 0.
[0124] –y L It equals sbHeight + 1.
[0125] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt L With refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[0126] Otherwise (cIdx is not equal to 0), then the following applies:
[0127] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[0128] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[0129] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[0130] – Let (refxSb) C ,refySb C ) and (refx C ,refy C ) represents the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows: C ,refySb C refx C and refy C :
[0131] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp(8-763)
[0132] refx C =((Sign(refxSb) C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5(8-764)
[0133] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp(8-765)
[0134] refyC =((Sign(refySb) C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5(8-766)
[0135] – The variable xInt is derived as follows C yInt C xFrac C and yFrac C :
[0136] xInt C =refx C >>5 (8-767)
[0137] yInt C =refy C >>5 (8-768)
[0138] xFrac C =refy C &31 (8-769)
[0139] yFrac C =refy C &31 (8-770)
[0140] – The predicted sample values predSamplesLX[xC][yC] are derived by invoking the procedure specified in Clause 8.5.6.3.4 with (xIntC,yIntC), (xFracC,yFracC), (xSbIntC,ySbIntC), sbWidth, sbHeight, and refPicLX as inputs.
[0141] 8.5.6.3.2 Brightness Sample Interpolation and Filtering Process
[0142] The input for this process is:
[0143] – Brightness position in units of the entire sample (xInt) L ,yInt L ),
[0144] – Brightness position in fractional samples (xFrac) L ,yFrac L ),
[0145] – Brightness position in units of the entire sample (xSbInt) L ,ySbInt LThis specifies the top-left sample of the bounding block used for filling the reference sample relative to the top-left brightness sample of the reference image.
[0146] –Luminance reference sample array refPicLX L ,
[0147] – Half-sample interpolation filter index hpelIfIdx,
[0148] – The variable sbWidth specifies the width of the current child block.
[0149] – The variable sbHeight specifies the height of the current child block.
[0150] – Brightness position (xSb, ySb), specifies the top-left sample of the current sub-block relative to the top-left brightness sample of the current image.
[0151] The output of this process is the predicted luminance sample value, predSampleLX. L
[0152] The variables shift1, shift2, and shift3 are derived as follows:
[0153] – Set the variable shift1 to equal Min(4, BitDepth) Y -8), set the variable shift2 to equal 6, and set the variable shift3 to equal Max(2, 14-BitDepth). Y ).
[0154] – Set the variable picW to equal pic_width_in_luma_samples and the variable picH to equal pic_height_in_luma_samples.
[0155] According to the following derivation, it equals xFrac. L or yFrac L The brightness interpolation filter coefficient f at each 1 / 16 fractional sample location p L [p]:
[0156] – If MotionModelIdc[xSb][ySb] is greater than 0, and sbWidth and sbHeight are both equal to 4, then the luminance interpolation filter coefficients f L [p] is specified in Table 2.
[0157] Otherwise, depending on hpelIfIdx, the luminance interpolation filter coefficients f L [p] is specified in Table 1.
[0158] For i = 0..7, the brightness position (xInt) in units of the entire sample is derived as follows. i ,yInt i ):
[0159] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0160] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3) (8-771)
[0161] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772)
[0162] Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), then the following applies:
[0163] xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?
[0164] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3):(8-773)
[0165] xInt L +i-3)
[0166] yInt i =Clip3(0,picH-1,yInt) L +i-3) (8-774)
[0167] For i = 0..7, further modify the brightness position in units of full sample points as follows:
[0168] xInt i =Clip3(xSbInt) L -3,xSbInt L +sbWidth+4,xInt i (8-775)
[0169] yInt i =Clip3(ySbInt) L -3,ySbInt L +sbHeight+4,yInt i (8-776)
[0170] The predicted luminance sample value predSampleLX is derived as follows. L :
[0171] –If xFrac L and yFrac L Both are equal to 0, so according to the following derivation, predSampleLX L Value:
[0172] predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3(8-777)
[0173] – Otherwise, if xFrac L Not equal to 0 and yFrac L If it equals 0, then according to the following derivation, predSampleLX L Value:
[0174]
[0175] – Otherwise, if xFrac L It equals 0, and yFrac L If it is not equal to 0, then according to the following derivation, predSampleLX L Value:
[0176]
[0177] – Otherwise, if xFrac L Not equal to 0 and yFrac L If it is not equal to 0, then according to the following derivation, predSampleLX L Value:
[0178] – The sample array temp[n] is derived as follows, where n = 0..7:
[0179]
[0180] – The predicted luminance sample value predSampleLX is derived as follows L :
[0181]
[0182] Table 8-11 specifies the brightness interpolation filter coefficients fL[p] for each 1 / 16 fractional sample location p.
[0183]
[0184] Table 8-12 shows the luminance interpolation filter coefficients f for each 1 / 16 fractional sample position p in the affine motion mode. L [p] provisions
[0185]
[0186] 8.5.6.3.3 Luminance Integer Sample Retrieval Process
[0187] The input for this process is:
[0188] – Brightness position in units of the entire sample (xInt) L ,yInt L ),
[0189] –Luminance reference sample array refPicLX L ,
[0190] The output of this process is the predicted luminance sample value, predSampleLX. L
[0191] Set the variable shift to equal Max(2, 14-BitDepth). Y ).
[0192] Set the variable picW to equal pic_width_in_luma_samples, and set the variable picH to equal pic_height_in_luma_samples.
[0193] The brightness position (xInt, yInt) in units of the entire sample point is derived as follows:
[0194] xInt=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?(8-782)
[0195] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L ):xInt L )
[0196] yInt = Clip3(0, picH-1, yInt) L (8-783)
[0197] The predicted luminance sample value predSampleLX is derived as follows. L :
[0198] predSampleLX L =refPicLX L [xInt][yInt]< <shift3 (8-784)
[0199] 8.5.6.3.4 Chromaticity Sample Interpolation Process
[0200] The input for this process is:
[0201] - Chromaticity position in units of the entire sample (xInt) C ,yInt C ),
[0202] – Chromaticity position in units of 1 / 32 fractional samples (xFrac) C ,yFrac C ),
[0203] – The chromaticity position (xSbIntC, ySbIntC) in units of full sample points is defined as the top-left sample point of the bounding block used for filling the reference sample point relative to the top-left chromaticity sample point of the reference image.
[0204] – The variable sbWidth specifies the width of the current child block.
[0205] – The variable sbHeight specifies the height of the current child block.
[0206] –Color reference sample array refPicLX C .
[0207] The output of this process is the predicted chromaticity sample value, predSampleLX. C
[0208] The variables shift1, shift2, and shift3 are derived as follows:
[0209] – Set the variable shift1 to equal Min(4, BitDepth) C -8), set the variable shift2 to equal 6, and set the variable shift3 to equal Max(2, 14-BitDepth). C ).
[0210] – Change variable picW CSet it to equal pic_width_in_luma_samples / SubWidthC, and set the variable picH C Set it to equal pic_height_in_luma_samples / SubHeightC.
[0211] Table 3 specifies that xFrac is equal to C or yFrac C The chromaticity interpolation filter coefficients f at each 1 / 32 fractional sample location p C [p]
[0212] Set the variable xOffset to equal (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC.
[0213] For i = 0..3, the chromaticity position (xInt) in units of the entire sample point is derived as follows. i ,yInt i ):
[0214] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[0215] xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (8-785)
[0216] yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (8-786)
[0217] Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), then the following applies:
[0218] xInt i =Clip3(0,picW C -1,sps_ref_wraparound_enabled_flag? ClipH(xOffset,picW C ,xIntC +i-1):(8-787)
[0219] xInt C +i-1)
[0220] yInt i =Clip3(0,picH C -1,yInt C +i-1) (8-788)
[0221] For i = 0..3, further modify the chromaticity position (xInt) in units of the full sample point as follows. i ,yInt i ):
[0222] xInt i =Clip3(xSbIntC-1,xSbIntC+sbWidth+2,xInt i (8-789)
[0223] yInt i =Clip3(ySbIntC-1,ySbIntC+sbHeight+2,yInt i (8-790)
[0224] The predicted chromaticity sample value predSampleLX is derived as follows. C :
[0225] –If xFrac C and yFrac C Both are equal to 0, so according to the following derivation, predSampleLX C Value:
[0226] predSampleLX C =refPicLX C [xInt1][yInt1]< <shift3 (8-791)
[0227] – Otherwise, if xFrac C Not equal to 0 and yFrac C If it equals 0, then according to the following derivation, predSampleLX C Value:
[0228]
[0229] – Otherwise, if xFrac C It equals 0, and yFracC If it is not equal to 0, then according to the following derivation, predSampleLX C Value:
[0230]
[0231] – Otherwise, if xFrac C Not equal to 0 and yFrac C If it is not equal to 0, then according to the following derivation, predSampleLX C Value:
[0232] – The sample array temp[n] is derived as follows, where n = 0..3:
[0233]
[0234] – The predicted chromaticity sample value predSampleLX is derived as follows. C :
[0235]
[0236] Table 8-13 specifies the chromaticity interpolation filter coefficients fC[p] for each 1 / 32 fractional sample location p.
[0237]
[0238]
[0239] 2.5 Affine Motion Compensation Prediction Based on Refined Sub-blocks
[0240] The techniques disclosed in this paper include a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the prediction samples are refined by adding the difference derived from the optical flow equation; this is called prediction refinement using optical flow (PROF). The proposed method enables pixel-level inter-frame prediction without increasing memory access bandwidth.
[0241] To achieve better motion compensation granularity, this contribution proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the brightness prediction samples are refined by adding the difference derived from the optical flow equation. The proposed PROF (Prediction Refinement with Optical Flow) is described in the following four steps.
[0242] Step 1) Perform sub-block-based affine motion compensation to generate sub-block predictions I(i,j).
[0243] Step 2) Calculate the spatial gradient g of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. x (i,j) and g y (i,j).
[0244] g x (i,j)=I(i+1,j)-I(i-1,j)
[0245] g y (i,j)=I(i,j+1)-I(i,j-1)
[0246] Sub-block prediction extends by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, pixels on the extended boundaries are copied from the nearest integer pixel location in the reference image. This avoids additional interpolation for the filled regions.
[0247] Step 3) Refinement of brightness prediction calculated using the optical flow equation (denoted as ΔI).
[0248] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)
[0249] Among them, such as Figure 1 As shown, ΔMV (denoted as Δv(i,j)) is the difference between the pixel MV (denoted as v(i,j)) calculated for the sample point position (i,j) and the sub-block MV of the sub-block to which pixel (i,j) belongs.
[0250] Since the affine model parameters and pixel positions relative to the sub-block center do not change between sub-blocks, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the sub-block center, then Δv(x,y) can be obtained from the following equation:
[0251]
[0252] For a 4-parameter affine model
[0253]
[0254]
[0255] For a 6-parameter affine model
[0256]
[0257] Among them, (v 0x ,v0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the motion vectors of the upper left, upper right, and lower left control points, and w and h are the width and height of the CU.
[0258] Step 4) Finally, refine the brightness prediction and add it to the sub-block prediction I(i,j). Generate the final prediction I′ according to the following equation.
[0259] I′(i,j)=I(i,j)+ΔI(i,j)
[0260] The following describes some details:
[0261] a) How PROF derives the gradient
[0262] In some embodiments, gradients are computed for each sub-block (4×4 sub-blocks in VTM-4.0) of each reference list. For each sub-block, the nearest integer sample from the reference block is retrieved to fill the four side outlines of the sample.
[0263] Assume the MV of the current sub-block is (MVx, MVy). Then, calculate the fractional part according to (FracX, FracY) = (MVx & 15, MVy & 15). Calculate the integer part according to (IntX, IntY) = (MVx >> 4, MVy >> 4). Derive the offset (OffsetX, OffsetY) as follows:
[0264] OffsetX = FracX > 7? 1:0;
[0265] OffsetY = Fracy > 7? 1:0;
[0266] Assume the top-left coordinate of the current sub-block is (xCur, yCur) and the dimensions of the current sub-block are W×H. Then calculate (xCor0, yCor0), (xCor1, yCor1), (xCor2, yCor2), and (xCor3, yCor3) as follows:
[0267] (xCor0,yCor0)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY-1);
[0268] (xCor1,yCor1)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY+H);
[0269] (xCor2, yCor2) = (xCur + IntX + OffsetX - 1, yCur + IntY + OffsetY);
[0270] (xCor3, yCor3) = (xCur + IntX + OffsetX + W, yCur + IntY + OffsetY);
[0271] Assume PredSample[x][y] (where x = 0..W - 1, y = 0..H - 1) stores the prediction samples for the sub - block. Then fill in the samples according to the following derivation:
[0272] PredSample[x][-1] = (Ref(xCor0 + x, yCor0) << Shift0) – Rounding, where x = -1..W;
[0273] PredSample[x][H] = (Ref(xCor1 + x, yCor1) << Shift0) – Rounding, where x = -1..W;
[0274] PredSample[-1][y] = (Ref(xCor2, yCor2 + y) << Shift0) – Rounding, where y = 0..H - 1;
[0275] PredSample[W][y] = (Ref(xCor3, yCor3 + y) << Shift0) - Rounding, where y = 0..H - 1;
[0276] Where, Rec represents the reference picture. Rounding is an integer, which is equal to 2 in the exemplary PROF implementation 13 . Shift0 = Max(2, (14 - BitDepth)); PROF attempts to improve the accuracy of the gradient, which is different from BIO in VTM - 4.0, the latter outputs the gradient with the same accuracy as the input luminance samples.
[0277] Calculate the gradient in PROF as follows:
[0278] Shift1 = Shift0 - 4.
[0279] gradientH[x][y] = (predSamples[x + 1][y] - predSample[x - 1][y]) >> Shift1
[0280] gradientV[x][y]=(predSample[x][y+1]-predSample[x][y-1])>>Shift1
[0281] It should be noted that predSamples[x][y] maintains the precision of the interpolation.
[0282] b) How PROF derives Δv
[0283] Δv can be described as follows (denoted as dMvH[posX][posY] and dMvV[posX][posY], where posX = 0..W-1 and posY = 0..H-1).
[0284] Assume the current block has dimensions cbWidth×cbHeight, the number of control point motion vectors is numCpMv, and the control point motion vector is cpMvLX[cpIdx], where cpIdx = 0..numCpMv-1 and X is 0 or 1, representing two reference lists.
[0285] The variables log2CbW and log2CbH are derived as follows:
[0286] log2CbW = Log2(cbWidth)
[0287] log2CbH = Log2(cbHeight)
[0288] The variables mvScaleHor, mvScaleVer, dHorX, and dVerX are derived as follows:
[0289] mvScaleHor = cpMvLX[0][0] << 7
[0290] mvScaleVer = cpMvLX[0][1] << 7
[0291] dHorX=(cpMvLX[1][0]-cpMvLX[0][0])<<(7-log2CbW)
[0292] dVerX=(cpMvLX[1][1]-cpMvLX[0][1])<<(7-log2CbW)
[0293] The variables dHorY and dVerY are derived as follows:
[0294] - If numCpMv equals 3, then the following applies:
[0295] dHorY=(cpMvLX[2][0]-cpMvLX[0][0])<<(7-log2CbH)
[0296] dVerY=(cpMvLX[2][1]-cpMvLX[0][1])<<(7-log2CbH)
[0297] - Otherwise (numCpMv equals 2), the following applies:
[0298] dHorY = -dVerX
[0299] dVerY=dHorX
[0300] The variables qHorX, qVerX, qHorY, and qVerY are derived as follows:
[0301] qHorX=dHorX<<2;
[0302] qVerX = dVerX << 2;
[0303] qHorY = dHorY << 2;
[0304] qVerY = dVerY << 2;
[0305] The following calculations are performed for dMvH[0][0] and dMvV[0][0]:
[0306] dMvH[0][0]=((dHorX+dHorY)<<1)-((qHorX+qHorY)<<1);
[0307] dMvV[0][0]=((dVerX+dVerY)<<1)-((qVerX+qVerY)<<1);
[0308] The following derivation applies to dMvH[xPos][0] and dMvV[xPos][0], where xPos ranges from 1 to W-1:
[0309] dMvH[xPos][0]=dMvH[xPos-1][0]+qHorX;
[0310] dMvV[xPos][0]=dMvV[xPos-1][0]+qVerX;
[0311] For yPos from 1 to H-1, the following applies:
[0312] dMvH[xPos][yPos]=dMvH[xPos][yPos-1]+qHorY, where xPos=0..W-1
[0313] dMvV[xPos][yPos]=dMvV[xPos][yPos-1]+qVerY, where xPos=0..W-1
[0314] Finally, the right shift of dMvH[xPos][yPos] and dMvV[xPos][yPos] of posM = 0..W-1 and posY = 0..H-1 is:
[0315] dMvH[xPos][yPos]=SatShift(dMvH[xPos][yPos],7+2-1);
[0316] dMvV[xPos][yPos]=SatShift(dMvV[xPos][yPos],7+2-1);
[0317] SatShift(x,n) and Shift(x,n) are defined as follows:
[0318]
[0319] Shift(x,n) = (x + offset0) >> n
[0320] In one example, offset0 and / or offset1 are set to (1 <<n)> >1.
[0321] c) How PROF derives ΔI
[0322] For a position (posX, posY) within a sub-block, its corresponding Δv(i,j) is represented as (dMvH[posX][posY], dMvV[posX][posY]). Its corresponding gradient is represented as (gradientH[posX][posY], gradientV[posX][posY]).
[0323] Then, ΔI(posX,posY) is derived as follows.
[0324] (dMvH[posX][posY], dMvV[posX][posY]) is clipped to:
[0325] dMvH[posX][posY]=Clip3(-32768,32767,dMvH[posX][posY]);
[0326] dMvV[posX][posY]=Clip3(-32768,32767,dMvV[posX][posY]);
[0327] ΔI(posX,posY)=dMvH[posX][posY]×gradientH[posX][posY]+dMvV[posX][posY]×gradientV[posX][posY];
[0328] ΔI(posX,posY)=Shift(ΔI(posX,posY),1+1+4);
[0329] ΔI(posX,posY)=Clip3(-(2 13 -1),2 13 -1,ΔI(posX,posY));
[0330] d) How PROF derives I'
[0331] If the current block is not encoded or decoded using bidirectional prediction or weighted prediction.
[0332] I'(posX,posY)=Shift((I(posX,posY)+ΔI(posX,posY)),Shift0),
[0333] I'(posX,posY)=ClipSample(I'(posX,posY)),
[0334] ClipSample trims the sample values to the valid output sample values. Then, I'(posX, posY) is output as the inter-frame prediction value.
[0335] Otherwise (if the current block is encoded / decoded using bidirectional prediction or weighted prediction), I'(posX, posY) will be stored and used to generate inter-frame predictions based on other predictions and / or weighting values.
[0336] 2.6 Example of a strip header
[0337]
[0338]
[0339]
[0340]
[0341]
[0342]
[0343] 2.7 Example Sequence Parameter Set
[0344]
[0345]
[0346]
[0347]
[0348]
[0349] 2.8 Example Image Parameter Set
[0350]
[0351]
[0352]
[0353]
[0354] 2.9 Example Adaptive Parameter Set
[0355]
[0356]
[0357]
[0358]
[0359]
[0360]
[0361] 2.10 Example Image Header
[0362] In some embodiments, the image header is designed to have the following properties:
[0363] 1. The temporal ID and layer ID of the NAL unit in the image header are the same as the temporal ID and layer ID of the layer access unit containing the image header.
[0364] 2. The image header NAL unit should precede the NAL unit of the first strip containing its associated image. This establishes the association between the image header and the strip of the image associated with that image header without requiring signaling notification in the image header and referencing the image header ID from the strip header.
[0365] 3. The image header NAL unit should follow the image-level parameter set or a higher level, such as DPS, VPS, SPS, PPS, etc. Therefore, this requires that those parameter sets are not repeated / do not exist within the image or access unit.
[0366] 4. The image header contains information about the image type of its associated image. Image types can be used to define the following (not an exhaustive list).
[0367] a. The image is an IDR image.
[0368] b. The image is a CRA image.
[0369] c. The image is a GDR image.
[0370] d. The image is a non-IRAP, non-GDR image, and contains only I-bands.
[0371] e. The image is a non-IRAP, non-GDR image, and can only contain P-bands and I-bands.
[0372] f. The image is a non-IRAP, non-GDR image and contains any one of the B band, P band, and / or I band.
[0373] 5. Move the signaling of the picture-level syntax element in the strip header to the picture header.
[0374] 6. Signaling notifications are sent in the stripe header to non-picture-level syntax elements, which are typically identical across all stripes of the same picture in the picture header. When those syntax elements are not present in the picture header, they can be signaled in the stripe header.
[0375] In some implementations, the concept of a forced image header is used to send once per image as the first VCL NAL unit of the image. It is also proposed to move syntax elements currently in the stripe header to this image header. Functionally, syntax elements that only need to be sent once per image can be moved to the image header instead of being sent multiple times for a given image; for example, syntax elements in the stripe header are sent once per stripe. Moving stripe header syntax elements is constrained to be identical within the image.
[0376] The syntax elements have been constrained to be identical across all stripes of the image. It can be inferred that moving these fields to the image header, so that each image is signaled only once, instead of each stripe, avoids unnecessary redundant bit transmissions without altering the functionality of these syntax elements.
[0377] 1. In some implementations, the following semantic constraints exist:
[0378] When present, the value of each of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and slice_temporal_mvp_enabled_flag should be the same across all slice headers for encoding and decoding the image. Therefore, each of these syntax elements can be moved to the image header to avoid unnecessary redundant bits.
[0379] In this contribution, `recovery_poc_cnt` and `no_output_of_prior_pics_flag` are not moved to the image header. Their presence in the strip header depends on a conditional check of the strip header's `nal_unit_type`, therefore, if you expect these syntax elements to be moved to the image header, it is recommended that you investigate them.
[0380] 2. In some implementations, the following semantic constraints exist:
[0381] When present, the value of slice_lmcs_aps_id should be the same for all slices of the image.
[0382] When present, the value of slice_scaling_list_aps_id should be the same for all slices of the image. Therefore, each of these syntax elements can be moved to the image header to avoid unnecessary redundant bits.
[0383] In some embodiments, syntax elements are not currently constrained to be identical across all stripes of the image. It is recommended to evaluate the intended use of these syntax elements to determine which syntax elements can be moved to the image header to simplify the overall VVC design, as it is claimed that handling a large number of syntax elements in each strip header has a complexity impact.
[0384] 1. It is proposed to move the following syntax elements to the picture header. Currently, there are no restrictions on them having different values on different stripes, but it is claimed that transmitting them in each stripe header offers no / minor benefit and incurs little codec penalty, as their intended use will change at the picture level:
[0385] a.six_minus_max_num_merge_cand
[0386] b.five_minus_max_num_subblock_merge_cand
[0387] c.slice_fpel_mmvd_enabled_flag
[0388] d.slice_disable_bdof_dmvr_flag
[0389] e.max_num_merge_cand_minus_max_num_triangle_cand
[0390] f.slice_six_minus_max_num_ibc_merge_cand
[0391] 2. It is proposed to move the following syntax elements to the picture header. Currently, there are no restrictions on them having different values on different stripes, but it is claimed that transmitting them in each stripe header offers no / minor benefit and incurs little codec penalty, as their intended use will change at the picture level:
[0392] a.partition_constraints_override_flag
[0393] b.slice_log2_diff_min_qt_min_cb_luma
[0394] c.slice_max_mtt_hierarchy_depth_luma
[0395] d.slice_log2_diff_max_bt_min_qt_luma
[0396] e.slice_log2_diff_max_tt_min_qt_luma
[0397] f.slice_log2_diff_min_qt_min_cb_chroma
[0398] g.slice_max_mtt_hierarchy_depth_chroma
[0399] h.slice_log2_diff_max_bt_min_qt_chroma
[0400] i.slice_log2_diff_max_tt_min_qt_chroma
[0401] The conditional check "slice_type==I" associated with some of these syntax elements has been removed along with the move to the image header.
[0402] 3. It is proposed to move the following syntax elements to the picture header. Currently, there are no restrictions on them having different values on different stripes, but it is claimed that transmitting them in each stripe header offers no / minor benefit and incurs little codec penalty because their intended use will change at the picture level:
[0403] a.mvd_l1_zero_flag
[0404] The conditional check "slice_type==B" associated with some of these syntax elements has been removed along with the move to the image header.
[0405] 4. It is proposed to move the following syntax elements to the picture header. Currently, there are no restrictions on them having different values on different stripes, but it is claimed that transmitting them in each stripe header offers no / minor benefit and incurs little codec penalty because their intended use will change at the picture level:
[0406] a.dep_quant_enabled_flag
[0407] sign_data_hiding_enabled_flag4.
[0408] 2.10.1 Example Syntax Table
[0409] 7.3.2.8 Image Header RBSP Syntax
[0410] [Table starts on the next page]
[0411]
[0412]
[0413]
[0414]
[0415] 2.11 Example DMVR
[0416] Decoder-side motion vector refinement (DMVR) utilizes bidirectional matching (BM) to derive the motion information of the current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference images. The BM method calculates the distortion between two candidate blocks in reference image lists L0 and L1. Figure 7As shown, the SAD between red blocks is calculated based on each MV candidate around the initial MV. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. The cost function used in the matching process is the row-subsampled SAD (sum of absolute differences). Figure 8 The example is shown below.
[0417] In VTM 5.0, when encoding and decoding a CU using the regular merge / skip mode and bidirectional prediction, DMVR is employed at the decoder to refine the motion vector (MV) for the codec unit (CU). In terms of display order, one reference image precedes the current image, and another reference image follows the current image. The temporal distance between the current image and one reference image is equal to the temporal distance between the current image and another reference image. Equal weights are selected using bidirectional prediction (BCW) with CU weights. When DMVR is applied, a luma codec block (CB) is divided into several independently processed sub-blocks of size min(cbWidth, 16) × min(cbHeight, 16). DMVR refines the MV of each sub-block by minimizing the SAD between the 10-bit L0 and L1 prediction samples of the 1 / 2 subsampled generated by bidirectional interpolation. For each sub-block, an integer ΔMV search is first performed using SAD around the initial MV (i.e., the MV of the selected regular merge / skip candidate), followed by fractional ΔMV derivation to obtain the final MV.
[0418] When the CU uses bidirectional predictive encoding / decoding, BDOF refines the CU's luminance prediction samples, with one reference image preceding the current image and another following it in display order, and BCW selecting equal weights. Eight-tap interpolation is used to generate initial L0 and L1 prediction samples based on the input MV (e.g., the final MV of the DMVR in the case of DMVR enabled). Next, a two-stage early termination process is performed. The first early termination occurs at the sub-block level, and the second at the 4×4 block level, and is checked if the first early termination does not occur. At each level, the SAD between the fully sampled 14-bit L0 and L1 prediction samples in each sub-block / 4×4 block is first calculated. If the SAD is less than a threshold, BDOF is not applied to the sub-block / 4×4 block. Otherwise, BDOF parameters are derived and used to generate the final luminance sample prediction for each 4×4 block. In BDOF, the sub-block size is the same as in DMVR, i.e., min(cbWidth, 16) × min(cbHeight, 16).
[0419] When the CU is encoded and decoded in the regular merge / skip mode, in terms of display order, one reference image precedes the current image and another reference image follows the current image. The temporal distance between the current image and one reference image is equal to the temporal distance between the current image and another reference image. BCW selects equal weights, and this applies to both DMVR and BDOF. The flow of the cascaded DMVR and BDOF process is as follows: Figure 8 As shown.
[0420] Motion compensation in RPR The flowchart of the cascading DMVR and BDOF processes in VTM5.0 is shown. The DMVR SAD operation and the BDOF SAD operation are different and not shared.
[0421] To reduce latency and operations in this critical path, the latest VVC working draft has been revised to reuse sub-block SAD computed in DMVR for early termination of sub-blocks in BDOF when both DMVR and BDOF are applied.
[0422] SAD calculation is defined as follows:
[0423]
[0424] The two variables nSbW and nSbH specify the width and height of the current sub-block, and the two (nSbW+4)×(nSbH+4) arrays pL0 and pL1 contain the predicted samples of L0 and L1 respectively, as well as the integer sample offsets (dX, dY) in the prediction list L0.
[0425] To mitigate the adverse consequences of DMVR refinement uncertainty, a method is proposed that favors the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial (or original) MV candidate is reduced by 1 / 4 of the SAD value. That is, when dX and dY in the above equation are both equal to 0, the value of SAD is modified as follows:
[0426] sad = sad - (sad >> 2)
[0427] When the SAD value is less than the threshold (2 * sub-block width * sub-block height), BDOF is no longer required.
[0428] 2.12 Chromaticity QP Mapping Table
[0429] In some embodiments, a chroma QP mapping table is defined to define how to obtain the chroma QP from the luminance QP using the following syntax and semantics.
[0430]
[0431] A `same_qp_table_for_chroma` value of 1 specifies that only one chroma QP map table is signaled. This table applies the Cb and Cr residuals and, when `sps_joint_cbcr_enabled_flag` is 1, also applies the joint Cb-Cr residual. A `same_qp_table_for_chroma` value of 1 specifies that the chroma QP map table is signaled in SPS, with two tables for Cb and Cr, and an additional table for the joint Cb-Cr when `sps_joint_cbcr_enabled_flag` is 1. When `same_qp_table_for_chroma` is not present in the bitstream, its value is inferred to be 1.
[0432] The increment 26 in `qp_table_start_minus26[i]` specifies the starting luma and chroma QP used to describe the i-th chroma QP mapping table. The value of `qp_table_start_minus26[i]` should be in the range of -26 - QpBdOffset to 36 (inclusive). If `qp_table_start_minus26[i]` is not present in the bitstream, its value is inferred to be 0.
[0433] Incrementing 1 to `num_points_in_qp_table_minus1[i]` specifies the number of points used to describe the i-th chroma QP map. The value of `num_points_in_qp_table_minus1[i]` should be in the range of 0 to 63 + QpBdOffset (inclusive). If `num_points_in_qp_table_minus1[i]` is not present in the bitstream, its value is inferred to be 0.
[0434] delta_qp_in_val_minus1[i][j] specifies the delta value used to derive the input coordinates of the j-th pivot point of the i-th chroma QP mapping table. When delta_qp_in_val_minus1[0][j] does not exist in the bitstream, the value of delta_qp_in_val_minus1[0][j] is inferred to be equal to 0.
[0435] delta_qp_diff_val[i][j] specifies the delta value used to derive the output coordinates of the j-th pivot point of the i-th chromaticity QP mapping table.
[0436] For i = 0..numQpTables–1, the i-th chroma QP mapping table ChromaQpTable[i] is derived as follows:
[0437]
[0438]
[0439] When same_qp_table_for_chroma equals 1, ChromaQpTable[1][k] and ChromaQpTable[2][k] are set to be equal to ChromaQpTable[0][k], where k ranges from -QpBdOffset to 63 (inclusive).
[0440] The bitstream consistency requirement is that, for i in the range of 0 to numQpTables–1 (inclusive) and j in the range of 0 to num_points_in_qp_table_minus1[i]+1 (inclusive), the values of qpInVal[i][j] and qpOutVal[i][j] should be in the range of -QpBdOffset to 63 (inclusive).
[0441] 3. Deficiencies of existing implementation methods
[0442] During motion vector refinement, DMVR and BIO do not involve the original signal, which can lead to inaccurate motion information in the codec blocks. Furthermore, DMVR and BIO sometimes use fractional motion vectors after motion refinement, while screen video typically has integer motion vectors, further complicating the current motion information and degrading codec performance.
[0443] When applying RPR in VVC, RPR (ARC) may have the following issues:
[0444] 1. Using RPR, the interpolation filters may be different for adjacent samples in a block, which is undesirable in SIMD (Single Instruction Multiple Data) implementations.
[0445] 2. The enclosed area does not consider RPR.
[0446] 3. It is important to note that "the consistent crop window offset parameter only applies to the output. All internal decoding processes apply to the uncropped image size." However, when applying RPR, these parameters can be used during the decoding process.
[0447] 4. When deriving the reference sample location, RPR only considers the ratio between the two consistency windows. However, the difference in the upper left offset between the two consistency windows should also be considered.
[0448] 5. The ratio between the width / height of the reference image and the width / height of the current image is constrained in VVC. However, the ratio between the width / height of the consistency window of the reference image and the width / height of the consistency window of the current image is not constrained.
[0449] 6. Not all syntax elements in the image header are processed correctly.
[0450] 7. In current VVC, for TPM and GEO prediction modes, chroma blending weights are derived regardless of the chroma sample location type of the video sequence. For example, in TPM / GEO, if the chroma weights are derived from the luma weights, it may be necessary to downsample the luma weights to match the chroma signal sampling. Chroma downsampling is typically applied when the chroma sample location type is 0, which is widely used in ITU-R BT.601 or ITU-R BT.709. However, if different chroma sample location types are used, this can lead to misalignment between the chroma samples and the downsampled luma samples, potentially degrading encoding / decoding performance.
[0451] 8. It is important to note that the SAD calculation / SAD threshold does not take into account the effect of bit depth. Therefore, for higher bit depths (e.g., 14 or 16-bit input sequences), the early termination threshold may be too small.
[0452] 9. For non-RPR cases, an AMVR (alternate interpolation filter / switchable interpolation filter) with 1 / 2 pixel MV accuracy is applied to a 6-tap motion compensation filter, but an 8-tap filter is applied to other cases (e.g., 1 / 16 pixel). However, for RPR cases, the same interpolation filter applies to all cases, regardless of MV / MVD accuracy. Therefore, signaling for the 1 / 2 pixel case (alternate interpolation filter / switchable interpolation filter) is a waste of bits.
[0453] 10. The decision on whether to allow split tree partitioning depends on the encoding / decoding image resolution, not the output image resolution.
[0454] 11. Using SMVD / MMVD without considering the RPR case. These methods are based on the assumption that symmetric MVD is applied to two reference images. However, this assumption is incorrect when the output images have different resolutions.
[0455] 12. Paired merge candidates are generated by averaging the two MVs of two merge candidates from the same list of reference images. However, averaging is meaningless when the two reference images associated with two merge candidates have different resolutions.
[0456] 13. If all stripes in the current image are I (intra-frame) stripes, it may not be necessary to encode or decode some syntax elements in the image header that are related to inter-frame stripes. Conditional signaling of these can save syntax overhead, especially for low-resolution sequences that utilize all intra-frame encoding and decoding.
[0457] 14. In the current VVC, there are no restrictions on the dimensions of slices / strips. Adding appropriate restrictions would facilitate parallel processing by real-time software / hardware decoders, especially for ultra-high resolution sequences where each frame may be larger than 4K / 8K.
[0458] 15. The syntax element qp_table_start_minus26 of the chroma QP mapping table can be improved.
[0459] 4 Example Technologies and Embodiments
[0460] The embodiments described in detail below are intended as examples to explain general concepts. These embodiments should not be interpreted narrowly. Furthermore, these embodiments can be combined in any way.
[0461] In addition to DMVR and BIO mentioned below, the methods described below can also be applied to other decoder motion information derivation techniques.
[0462] The motion vector is represented as (mv_x, mv_y), where mv_x is the horizontal component and mv_y is the vertical component.
[0463] In this disclosure, the resolution (or dimension, or width / height, or size) of an image can refer to the resolution (or dimension, or width / height, or size) of the encoded / decoded image, or it can refer to the resolution (or dimension, width / height, or size) of the consistency window within the encoded / decoded image. In one example, the resolution (or dimension, or width / height, or size) of an image can be referenced to parameters associated with the RPR (Reference Image Resampling) process, such as the scaling window / phase offset window. In another example, the resolution (or dimension, or width / height, or size) of an image is related to the resolution of the associated output image.
[0464] Figure 2
[0465] 1. When the resolution of the reference image is different from that of the current image, or when the width and / or height of the reference image is greater than that of the current image, the same horizontal and / or vertical interpolation filters can be used to generate the predicted value of the sample set (at least two samples) of the current block.
[0466] a. In one example, the collection could include all sample points in the region of the block.
[0467] i. For example, the block can be divided into S non-overlapping M×N rectangles. Each M×N rectangle is a set. In such a case... Figure 3A In the example shown, a 16×16 block can be divided into 16 4×4 rectangles, each rectangle being a set.
[0468] ii. For example, a row with N samples is a set. N is an integer no greater than the block width. In one example, N is 4 or 8 or the block width.
[0469] iii. For example, a column with N samples is a set. N is an integer no greater than the block height. In one example, N is 4 or 8 or the block height.
[0470] iv. M and / or N can be predefined or derived on the fly based on block-dimensional / encoding / decoding information or signaling notification.
[0471] b. In one example, samples in a set can have the same MV (represented as shared MV).
[0472] c. In one example, samples in a set can have MVs with the same horizontal component (represented as shared horizontal components).
[0473] d. In one example, samples in a set can have MVs with the same vertical component (represented as shared vertical component).
[0474] e. In one example, samples in a set can have MVs with the same fractional part (represented as shared fractional level components) having a horizontal component.
[0475] i. For example, suppose the MV of the first sample point is (MV1x, MV1y) and the MV of the second sample point is (MV2x, MV2y), then MV1x & (2 M -1) equals MV2x&(2 M -1), where M represents the precision of MV. For example, M = 4.
[0476] f. In one example, samples in the group may have MVs with the same fractional part having a vertical component (represented as shared fractional vertical components).
[0477] i. For example, suppose the MV of the first sample is (MV1x, MV1y) and the MV of the second sample is (MV2x, MV2y), then MV1y & (2M-1) should satisfy MV2y & (2M-1), where M represents the precision of MV. For example, M = 4.
[0478] g. In one example, for samples in the set to be predicted, the current image and a reference image (e.g., derived in section 8.5.6.3.1 of JVET-O2001-v14, refx) can be used as a reference. L ,refy L The resolution of MV is used to deduce the result. b The motion vector is represented. Then, the MV can be... b Further modifications (e.g., rounding / truncating / clipping to MV') are made to meet requirements such as those listed above, and MV' will be used to derive the predicted sample points for that sample point.
[0479] i. In one example, MV' has the same characteristics as MV b The same integer part, and the fractional part of MV' is set to the shared fractional horizontal and / or vertical components.
[0480] ii. In one example, MV' is set to have shared fractional horizontal and / or vertical components and is closest to MV. b .
[0481] h. Shared motion vectors (and / or shared horizontal components and / or shared vertical components and / or shared fractional vertical components and / or shared fractional vertical components) can be set as the motion vectors (and / or horizontal components and / or vertical components and / or fractional vertical components and / or fractional vertical components) of specific samples in the set.
[0482] i. For example, a particular sample point may be located at the corner of a rectangular set, such as... Figure 3A The symbols “A”, “B”, “C”, and “D” are shown in the diagram.
[0483] ii. For example, a particular sample point may be located at the center of a rectangular set, such as Figure 3B The symbols “E”, “F”, “G”, and “H” are shown.
[0484] iii. For example, a particular sample point may be located at the end of a set of row or column shapes, such as Figure 3C and Figure 3B The “A” and “D” shown are shown.
[0485] iv. For example, a particular sample point may be located in the middle of a set of row or column shapes, such as Figure 3C and Figures 3A-3C The “B” and “C” shown are shown.
[0486] v. In one example, the motion vector of a specific sample point could be the MV mentioned in entry g. b .
[0487] i. Shared motion vectors (and / or shared horizontal components and / or shared vertical components and / or shared fractional vertical components and / or shared fractional vertical components) can be set as the motion vectors (and / or horizontal components and / or vertical components and / or fractional vertical components and / or fractional vertical components) of virtual samples located at different positions compared to all samples in this set.
[0488] i. In one example, the virtual sample is not in the set, but is located in a region that covers all samples in the set.
[0489] 1) Alternatively, the virtual sample point is located outside the region that covers all samples in the set, for example, at the lower right corner of the region.
[0490] ii. In one example, the MV of the virtual sample is derived in the same way as the real sample, but at a different location.
[0491] iii. Figure 3A The “V” in the diagram shows three examples of virtual samples.
[0492] j. A function that can set shared MV (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) to multiple sample points and / or virtual sample points' MV (and / or horizontal component and / or vertical component and / or fractional vertical component and / or fractional vertical component).
[0493] i. For example, the shared MV (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) can be set as the average MV (and / or horizontal component and / or vertical component and / or fractional vertical component and / or fractional vertical component) of the following samples: all or part of the samples in the set; or Figure 3A The sample points “E”, “F”, “G”, “H”; or Figure 3A The sample points “E” and “H” in the sample; or Figure 3A The sample points “A”, “B”, “C”, and “D”; or Figure 3B Sample points “A” and “D” in the sample; or Figure 3A Sample points “B” and “C” in the sample; or Figure 3C Sample points “A” and “D” in the sample; Figure 3C Sample points “B” and “C” in the sample; or Interaction between RPR and other coding tools Sample points “A” and “D” in the sample.
[0494] 2. It proposes that when the resolution of the reference image is different from the resolution of the current image, or when the width and / or height of the reference image is greater than the width and / or height of the current image, only integer MV is allowed to perform the motion compensation process to derive the predicted block of the current block.
[0495] a. In one example, the decoded motion vector used for the sample to be predicted is rounded to an integer MV before being used.
[0496] b. In one example, the decoded motion vector used for the sample to be predicted is rounded to the nearest integer MV of the decoded motion vector.
[0497] c. In one example, the decoded motion vector used for the sample to be predicted is rounded to the integer MV that is closest to the decoded motion vector in the horizontal direction.
[0498] d. In one example, the decoded motion vector used for the sample to be predicted is rounded to the integer MV that is closest to the decoded motion vector in the vertical direction.
[0499] 3. Motion vectors used in the motion compensation process of samples in the current block (e.g., shared MV / shared horizontal or vertical or fractional component / MV' mentioned in the above entries) can be stored in the decoded image buffer and used for motion vector prediction of subsequent blocks in the current / different images.
[0500] a. Alternatively, motion vectors used in the motion compensation process for samples in the current block (e.g., shared MV / shared horizontal or vertical or fractional component / MV' mentioned in the above entries) are not allowed to be used for the prediction of motion vectors in subsequent blocks in the current / different images.
[0501] i. In one example, the decoded motion vector (e.g., the MV in the above entry) can be... b It is used for motion vector prediction of subsequent blocks in the current / different images.
[0502] b. In one example, the motion vector used in the motion compensation process of the samples in the current block can be used in the filtering process (e.g., deblocking filter / SAO / ALF).
[0503] i. Alternatively, the decoded motion vector (e.g., the MV in the above entry) can be used during the filtering process. b ).
[0504] c. In one example, such an MV can be derived at the sub-block level and such an MV can be stored for each sub-block.
[0505] 4. An interpolation filter for deriving the prediction block of the current block during motion compensation is proposed, which can be selected based on whether the resolution of the reference image is different from the resolution of the current image, or whether the width and / or height of the reference image is greater than the width and / or height of the current image.
[0506] a. In one example, an interpolation filter with fewer taps can be applied when condition A is met, where condition A depends on the dimensions of the current image and / or the reference image.
[0507] i. In one example, condition A is that the resolution of the reference image is different from the resolution of the current image.
[0508] ii. In one example, condition A is that the width and / or height of the reference image is greater than the width and / or height of the current image.
[0509] iii. In one example, condition A is W1>a*W2 and / or H1>b*H2, where (W1, H1) represents the width and height of the reference image, and (W2, H2) represents the width and height of the current image, and a and b are two coefficients, for example, a = b = 1.5.
[0510] iv. In one example, condition A may also depend on whether bidirectional prediction is used.
[0511] v. In one example, a 1-tap filter is applied. In other words, the output is the unfiltered integer pixels as the interpolation result.
[0512] vi. In one example, a bilinear filter is applied when the resolution of the reference image is different from the resolution of the current image.
[0513] vii. In one example, when the resolution of the reference image is different from that of the current image, or when the width and / or height of the reference image is greater than that of the current image, a 4-tap filter or a 6-tap filter is applied.
[0514] 1) A 6-tap filter can also be used for affine motion compensation.
[0515] 2) A 4-tap filter can also be used for interpolation of chromaticity samples.
[0516] b. In one example, padding samples are used to perform interpolation when the resolution of the reference image is different from that of the current image, or when the width and / or height of the reference image is greater than that of the current image.
[0517] c. Whether and / or how the methods disclosed in Item 4 are applied may depend on the color components.
[0518] i. For example, these methods are only applied to the luminance component.
[0519] d. Whether and / or how the methods disclosed in Item 4 are applied may depend on the direction of the interpolation filter.
[0520] i. For example, these methods are only applicable to horizontal filtering.
[0521] ii. For example, these methods are only applicable to vertical filtering.
[0522] 5. A two-stage process for predicting block generation is proposed when the resolution of the reference image is different from that of the current image, or when the width and / or height of the reference image is greater than that of the current image.
[0523] a. In the first stage, virtual reference blocks are generated by upsampling or downsampling regions in the reference image, depending on the width and / or height of the current image and the reference image.
[0524] b. In the second stage, prediction samples are generated from the virtual reference block by applying interpolation filtering, independent of the width and / or height of the current and reference images.
[0525] 6. As in some embodiments, it is proposed that the calculation of the top-left coordinates (xSbIntL, ySbIntL) of the bounding block used for reference sample filling can be derived depending on the width and / or height of the current image and the reference image.
[0526] a. In one example, the brightness position in the full sample cell is modified as follows:
[0527] xInt i =Clip3(xSbInt) L -Dx,xSbInt L +sbWidth+Ux,xInt i ),
[0528] yInt i =Clip3(ySbInt) L -Dy,ySbInt L +sbHeight+Uy,yInt i ),
[0529] Where Dx and / or Dy and / or Ux and / or Uy can depend on the width and / or height of the current image and the reference image.
[0530] b. In one example, the chromaticity position in the full sample cell is modified as follows:
[0531] xInti = Clip3(xSbInt) C -Dx,xSbInt C +sbWidth+Ux,xInti)
[0532] yInti = Clip3(ySbInt) C -Dy,ySbInt C+sbHeight+Uy,yInti)
[0533] Where Dx and / or Dy and / or Ux and / or Uy can depend on the width and / or height of the current image and the reference image.
[0534] 7. Instead of storing / using motion vectors for blocks based on a reference image with the same resolution as the current image, it is proposed to use true motion vectors, which take into account the resolution difference.
[0535] a. Alternatively, when using motion vectors to generate prediction patches, it is not necessary to base the predictions on the current image and a reference image (e.g., refx). L refy L The resolution of the motion vector can be further modified to change the motion vector.
[0536] Constraints on RPR
[0537] 8. Whether / how to apply a filtering process (e.g., deblocking filter) may depend on the resolution of the reference image and / or the resolution of the current image.
[0538] a. In one example, in addition to the motion vector difference, the boundary strength (BS) setting in the deblocking filter can account for the resolution difference.
[0539] i. In one example, the difference in motion vectors scaled according to the resolution of the current and reference images can be used to determine the boundary strength.
[0540] b. In one example, compared to using the same resolution for both blocks, if the resolution of at least one reference image of block A is different (or less or greater than) the resolution of an indicative reference image of block B, the strength of the deblocking filter used for the boundary between blocks A and B can be set differently (e.g., increased / decreased).
[0541] c. In one example, if the resolution of at least one reference image of block A is different from (or less than or greater than) the resolution of at least one reference image of block B, then the boundary between block A and block B is marked as to be filtered (e.g., BS is set to 2).
[0542] d. In one example, if the resolution of at least one reference image of block A and / or block B is different from (or less than or greater than) the resolution of the current image, the strength of the deblocking filter used for the boundary between block A and block B can be set differently (e.g., increased / decreased) compared to the case where the reference image and the current image use the same resolution.
[0543] e. In one example, if at least one reference image of at least one of two blocks has a resolution different from the current image resolution, the boundary between the two blocks is marked as to be filtered (e.g., BS is set to 2).
[0544] 9. When sub-images exist, the consistent bitstream may need to satisfy the requirement that the reference image must have the same resolution as the current image.
[0545] a. Alternatively, when the reference image has a different resolution than the current image, the current image must not contain any sub-images.
[0546] b. Alternatively, for sub-images within the current image, reference images with different resolutions are not permitted as the current image.
[0547] i. Alternatively, reference image management can be invoked to exclude reference images that have different resolutions.
[0548] 10. In one example, sub-images can be defined separately for images with different resolutions (e.g., how to divide an image into multiple sub-images).
[0549] In one example, if the reference image has a different resolution than the current image, the corresponding sub-image in the reference image can be derived by scaling and / or offsetting the sub-images of the current image.
[0550] 11. PROF (Predictive Refinement Using Optical Flow) can be enabled when the reference image has a different resolution than the current image.
[0551] a. In one example, a set of MVs (denoted as MV) can be generated from the sample set. g This can be used for motion compensation, as described in Item 1. On the other hand, MV (denoted as MV) can be derived for each sample point. p And can convert MV p and MV g The difference between them (e.g., corresponding to Δv used in PROF) and the gradient (e.g., the spatial gradient of the motion-compensated block) are used to derive the prediction refinement.
[0552] b. In one example, MV p It can have the same as MV g Different levels of precision. For example, MV p It can have a precision of 1 / N pixels (N>0), where N=32, 64, etc.
[0553] c. In one example, MV g It can have a precision different from the internal MV precision (e.g., 1 / 16 pixel).
[0554] d. In one example, prediction refinement is added to the prediction block to generate a refined prediction block.
[0555] e. In one example, such a method can be applied to each prediction direction.
[0556] f. In one example, such a method can be applied only to the case of one-way prediction.
[0557] g. In one example, such a method can be applied to one-way forecasting and / or two-way forecasting.
[0558] h. In one example, this method can only be applied if the reference image has a different resolution than the current image.
[0559] 12. It is proposed that when the resolution of the reference image is different from the resolution of the current image, only one MV can be used for block / sub-block to perform motion compensation process to derive the predicted block of the current block.
[0560] a. In one example, the unique MV used for a block / sub-block can be defined as a function (e.g., average) of all MVs associated with each sample point within the block / sub-block.
[0561] b. In one example, the unique MV used for a block / sub-block can be defined as the selected MV associated with a selected sample point (e.g., the center sample point) within the block / sub-block.
[0562] c. In one example, only one MV can use a 4×4 block or a sub-block (e.g., 4×1).
[0563] d. In one example, BIO can be further applied to compensate for the loss of accuracy due to block-based motion vectors.
[0564] 13. When the width and / or height of the reference image differs from the width and / or height of the current image, a non-signaling notification can be applied to any block-based motion vector-based lazy mode.
[0565] a. In one example, motion vectors may not be signaled, and the motion compensation process is approximately the same as a pure resolution change in a still image.
[0566] b. In one example, motion vectors at the image / piece / brick / CTU level can be signaled only, and the relevant blocks can use the motion vectors when the resolution changes.
[0567] 14. For blocks coded using affine prediction mode and / or non - affine prediction mode, when the width and / or height of the reference picture is different from the width and / or height of the current picture, PROF can be applied to approximate motion compensation.
[0568] a. In one example, when the width and / or height of the reference picture is different from the width and / or height of the current picture, PROF can be enabled.
[0569] b. In one example, a set of affine motions can be generated by combining the indicated motion and resolution scaling and used by PROF.
[0570] 15. When the width and / or height of the reference picture is different from the width and / or height of the current picture, interleaved prediction (e.g., as proposed in JVET - K0102) can be applied to approximate motion compensation.
[0571] a. In one example, the resolution change (scaling) is represented as an affine motion and interleaved motion prediction can be applied.
[0572] 16. When the width and / or height of the current picture is different from the width and / or height of the IRAP picture in the same IRAP period, LMCS and / or chroma residual scaling can be disabled.
[0573] a. In one example, when LMCS is disabled, strip - level flags such as slice_lmcs_enabled_flag, slice_lmcs_aps_id, and slice_chroma_residual_scale_flag may not be signaled and are inferred as 0.
[0574] b. In one example, when chroma residual scaling is disabled, strip - level flags such as slice_chroma_residual_scale_flag may not be signaled and are inferred as 0.
[0575] b. In one example, when chroma residual scaling is disabled, strip - level flags such as slice_chroma_residual_scale_flag may not be signaled and are inferred as 0.
[0576] Conformance window dependencies
[0577] 17. RPR can be applied to coded blocks with block dimension constraints.
[0578] a. In one example, for a coded block of M×N, where M is the block width and N is the block height, when M*N < T or M*N <= T (such as T = 256), RPR may not be used.
[0579] b. In one example, when M < K (or M <= K) (such as K = 16) and / or N < L (or N <= L) (such as L = 16), RPR may not be used.
[0580] 18. Bitstream consistency may be added to limit the ratio between the width and / or height of the active reference picture (or its consistency window) and the width and / or height of the current picture. Assume refPicW and refPicH represent the width and height of the reference picture, and curPicW and curPicH represent the width and height of the current picture.
[0581] a. In one example, when (refPicW ÷ curPicW) is an integer, the reference picture may be marked as an active reference picture.
[0582] i. Alternatively, when (refPicW ÷ curPicW) is a fraction, the reference picture may be marked as unavailable.
[0583] b. In one example, when (refPicW ÷ curPicW) = (X * n), where X represents a fraction, such as X = 1 / 2, and n represents an integer, such as n = 1, 2, 3, 4..., the reference picture may be marked as an active reference picture.
[0584] i. In one example, when (refPicW ÷ curPicW) is not equal to (X * n), the reference picture may be marked as unavailable.
[0585] 19. For an M×N block, whether to enable and / or how to enable the encoding / decoding tool (e.g., the bidirectional prediction / full triangle prediction mode (TPM) / fusion process in TPM) may depend on the resolution of the reference picture (or its consistency window) and / or the resolution of the current picture (or its consistency window).
[0586] a. In one example, M * N < T or M * N <= T (such as T = 64).
[0587] b. In one example, M < K (or M <= K) (such as K = 16) and / or N < L (or N <= L) (such as L = 16).
[0588] c. In one example, when the width / height of at least one reference picture is different from the current picture, the encoding / decoding tool is not allowed.
[0589] i. In one example, when the width / height of at least one reference picture of the block is greater than the width / height of the current picture, the encoding / decoding tool is not allowed.
[0590] d. In one example, when the width / height of each reference image in a block differs from the width / height of the current image, the use of encoding / decoding tools is not permitted.
[0591] i. In one example, encoding / decoding tools are not allowed when the width / height of each reference image is greater than the width / height of the current image.
[0592] e. Alternatively, when the use of codec tools is not permitted, a single MV can be used as a one-way prediction for motion compensation.
[0593] Downsampling filter type used in TPM / GEO for chroma blending mask generation
[0594] 20. Signal the consistent clipping window offset parameters (e.g., conf_win_left_offset) with N-pixel precision instead of 1-pixel precision, where N is a positive integer greater than 1.
[0595] a. In one example, the actual offset can be derived by multiplying the offset of the signaling notification by N.
[0596] b. In one example, N is set to 4 or 8.
[0597] 21. A proposal was made to apply the consistency cropping window offset parameter not only to the output. Some internal decoding processes may depend on the size of the cropped image (e.g., the resolution of the consistency window in the image).
[0598] 22. It is proposed that when the width and / or height of the images represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are the same, the consistent cropping window offset parameters in the first video unit (e.g., PPS) and the second video unit can be different.
[0599] 23. It is proposed that when the width and / or height of the images represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are different, the consistent cropping window offset in the consistent bitstream should be the same in the first video unit (e.g., PPS) and the second video unit.
[0600] a. It is proposed that regardless of whether the width and / or height of the images represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are the same, the consistent cropping window offset parameter in the consistent bitstream should be the same in the first video unit (e.g., PPS) and the second video unit.
[0601] 24. Assume that the width and height of the consistency window defined in the first video unit (e.g., PPS) are represented as W1 and H1, respectively. The width and height of the consistency window defined in the second video unit (e.g., PPS) are represented as W2 and H2, respectively. The top-left position of the consistency window defined in the first video unit (e.g., PPS) is represented as X1 and Y1. The top-left position of the consistency window defined in the second video unit (e.g., PPS) is represented as X2 and Y2. The width and height of the codec / decoder images (e.g., pic_width_in_luma_samples and pic_height_in_luma_samples) defined in the first video unit (e.g., PPS) are represented as PW1 and PH1, respectively. The width and height of the codec / decoder images defined in the second video unit (e.g., PPS) are represented as PW2 and PH2.
[0602] a. In one example, in a consistent bitstream, W1 / W2 should be equal to X1 / X2.
[0603] i. Alternatively, in a consistent bitstream, W1 / X1 should be equal to W2 / X2.
[0604] ii. Alternatively, in a consistent bitstream, W1*X2 should be equal to W2*X1.
[0605] b. In one example, in a consistent bitstream, H1 / H2 should be equal to Y1 / Y2.
[0606] i. Alternatively, in a consistent bitstream, H1 / Y1 should be equal to H2 / Y2.
[0607] ii. Alternatively, in a consistent bitstream, H1*Y2 should be equal to H2*Y1.
[0608] c. In one example, in a consistent bitstream, PW1 / PW2 should be equal to X1 / X2.
[0609] i. Alternatively, in a consistent bitstream, PW1 / X1 should be equal to PW2 / X2.
[0610] ii. Alternatively, in a consistent bitstream, PW1*X2 should be equal to PW2*X1.
[0611] d. In one example, PH1 / PH2 should be equal to Y1 / Y2 in a consistent bitstream.
[0612] i. Alternatively, in a consistent bitstream, PH1 / Y1 should be equal to PH2 / Y2.
[0613] ii. Alternatively, in a consistent bitstream, PH1*Y2 should be equal to PH2*Y1.
[0614] e. In one example, in a consistent bitstream, PW1 / PW2 should be equal to W1 / W2.
[0615] i. Alternatively, in a consistent bitstream, PW1 / W1 should be equal to PW2 / W2.
[0616] ii. Alternatively, in a consistent bitstream, PW1*W2 should be equal to PW2*W1.
[0617] fF. In one example, PH1 / PH2 should equal H1 / H2 in a consistent bitstream.
[0618] i. Alternatively, in a consistent bitstream, PH1 / H1 should be equal to PH2 / H2.
[0619] ii. Alternatively, in a consistent bitstream, PH1*H2 should be equal to PH2*H1.
[0620] g. In a consistent bitstream, if PW1 is greater than PW2, then W1 must be greater than W2.
[0621] h. In a consistent bitstream, if PW1 is less than PW2, then W1 must be less than W2.
[0622] i. In a consistent bitstream, (PW1-PW2)*(W1-W2) must be no less than 0.
[0623] j. In a consistent bitstream, if PH1 is greater than PH2, then H1 must be greater than H2.
[0624] k. In a consistent bitstream, if PH1 is less than PH2, then H1 must be less than H2.
[0625] l. In a consistent bitstream, (PH1-PH2)*(H1-H2) must be no less than 0.
[0626] m. In a consistent bitstream, if PW1>=PW2, then W1 / W2 must not be greater than (or less than) PW1 / PW2.
[0627] In a consistent bitstream, if PH1 >= PH2, then H1 / H2 must not be greater than (or less than) PH1 / PH2.
[0628] 25. Assume the width and height of the consistency window for the current image are denoted as W and H, respectively. The width and height of the consistency window for the reference image are denoted as W' and H', respectively. Then, the consistent bitstream should comply with at least one of the following constraints.
[0629] aW*pw>=W'; pw is an integer, such as 2.
[0630] bW*pw>W'; pw is an integer, such as 2.
[0631] c.W'*pw'>=W; pw' is an integer, such as 8.
[0632] d.W'*pw'>W; pw' is an integer, such as 8.
[0633] eH*ph>=H'; ph is an integer, for example, 2.
[0634] fH*ph>H'; ph is an integer, such as 2.
[0635] g.H'*ph'>=H; ph' is an integer, such as 8.
[0636] h.H'*ph'>H; ph' is an integer, such as 8.
[0637] i. In one example, pw is equal to pw'.
[0638] j. In one example, ph equals ph'.
[0639] k. In one example, pw equals ph.
[0640] l. In one example, pw' is equal to ph'.
[0641] m. In one example, when W and H represent the width and height of the current image, respectively, the consistent bitstream needs to satisfy the above sub-entries. W' and H' represent the width and height of the reference image.
[0642] 26. A partial signaling notification consistency window parameter was proposed.
[0643] a. In one example, the top-left sample in the consistency window of the image is the same as the top-left sample in the image.
[0644] b. For example, do not signal conf_win_left_offset defined in VVC and infer it to be zero.
[0645] c. For example, the conf_win_top_offset defined in the non - belief signaling notice in VVC is inferred as zero.
[0646] 27. It is proposed that the derivation of the position of the reference sample points (e.g., (refx L , refy L )) defined in VVC) can depend on the upper - left position of the consistency window of the current picture and / or the reference picture (e.g., (conf_win_left_offset, conf_win_top_offset) defined in VVC). Figure 4 shows an example of deriving the sample point position according to VVC (a) and the proposed method (b). The dashed - line rectangle represents the consistency window.
[0647] a. In one example, the dependence exists only when the width and / or height of the current picture is different from the width and / or height of the reference picture.
[0648] b. In one example, the derivation of the horizontal position of the reference sample point (e.g., refx L ) defined as in VVC) can depend on the left position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_left_offset defined as in VVC).
[0649] i. In one example, calculate the horizontal position of the current sample point relative to the upper - left position of the consistency window in the current picture (denoted as xSb’), and use it to derive the position of the reference sample point.
[0650] 1) For example, calculate xSb’ = xSb-(conf_win_left_offset << Prec), and use it to derive the position of the reference sample point, where xSb represents the horizontal position of the current sample point in the current picture. conf_win_left_offset represents the horizontal position of the upper - left sample point in the consistency window of the current picture. Prec represents the precision of xSb and xSb’, where (xSb >> Prec) can show the actual horizontal coordinate of the current sample point relative to the current picture. For example, Prec = 0 or Prec = 4.
[0651] ii. In one example, calculate the horizontal position of the reference sample point relative to the upper - left position of the consistency window in the reference picture (denoted as Rx’).
[0652] 1) The calculation of Rx’ can depend on xSb’ and / or the motion vector and / or the resampling ratio.
[0653] iii. In one example, calculate the horizontal position of the relative position of the reference sample point in the reference picture according to Rx’ (denoted as Rx).
[0654] 1) For example, calculate Rx = Rx’ + (conf_win_left_offset_ref << Prec), where conf_win_left_offset_ref represents the horizontal position of the top-left sample in the consistency window of the reference picture. Prec presents the precision of Rx and Rx’. For example, Prec = 0 or Prec = 4.
[0655] iv. In one example, Rx can be directly calculated based on xSb’ and / or the motion vector and / or the resampling ratio. In other words, the two steps of deriving Rx’ and Rx are combined into one-step calculation.
[0656] v. Whether and / or how to use the left position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_left_offset defined in VVC) can depend on the color component and / or the color format.
[0657] 1) For example, conf_win_left_offset can be modified to conf_win_left_offset = conf_win_left_offset * SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example, for the luminance component, SubWidthC is equal to 1. When the color format is 4:2:0 or 4:2:2, for the chrominance component, SubWidthC is equal to 2.
[0658] 2) For example, conf_win_left_offset can be modified to conf_win_left_offset = conf_win_left_offset / SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example, for the luminance component, SubWidthC is equal to 1. When the color format is 4:2:0 or 4:2:2, for the chrominance component, SubWidthC is equal to 2.
[0659] c. In one example, the derivation of the vertical position of the reference sample (e.g., refy defined in VVC) L can depend on the top position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_top_offset defined in VVC).
[0660] i. In one example, calculate the vertical position of the current sample relative to the top-left position of the consistency window in the current picture (denoted as ySb’), and use it to derive the position of the reference sample.
[0661] 1) For example, calculate ySb’ = ySb - (conf_win_top_offset << Prec) and use it to derive the position of the reference sample, where ySb represents the vertical position of the current sample in the current picture. conf_win_top_offset represents the vertical position of the top-left sample in the consistency window of the current picture. Prec represents the precision of ySb and ySb’. For example, Prec = 0 or Prec = 4.
[0662] ii. In one example, calculate the vertical position of the reference sample relative to the top-left position of the consistency window in the reference picture (denoted as Ry’).
[0663] 1) The calculation of Ry’ can depend on ySb’ and / or the motion vector and / or the resampling ratio.
[0664] iii. In one example, calculate the vertical position of the reference sample relative to the reference picture based on Ry’ (denoted as Ry).
[0665] 1) For example, calculate Ry = Ry’ + (conf_win_top_offset_ref << Prec), where conf_win_top_offset_ref represents the vertical position of the top-left sample in the consistency window of the reference picture. Prec represents the precision of Ry and Ry’. For example, Prec = 0 or Prec = 4.
[0666] iv. In one example, Ry can be directly calculated based on ySb’ and / or the motion vector and / or the resampling ratio. In other words, the two steps of deriving Ry’ and Ry are combined into one-step calculation.
[0667] v. Whether and / or how to use the top position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_top_offset defined in VVC) can depend on the color component and / or the color format.
[0668] 1) For example, conf_win_top_offset can be modified to conf_win_top_offset = conf_win_top_offset * SubHeightC, where SubHeightC defines the vertical sampling step of the color component. For example, for the luminance component, SubHeightC is equal to 1. When the color format is 4:2:0, for the chrominance component, SubHeightC is equal to 2.
[0669] 2) For example, `conf_win_top_offset` can be modified to `conf_win_top_offset = conf_win_top_offset / SubHeightC`, where `SubHeightC` defines the vertical sampling step of the color components. For example, for the luminance component, `SubHeightC` equals 1. When the color format is 4:2:0, for the chrominance component, `SubHeightC` equals 2.
[0670] 28. A proposal is made to crop the integer portion of the horizontal coordinates of the reference sample points to [minW, maxW]. Assume the width and height of the consistency window in the reference image are represented by W and H, respectively. The width and height of the consistency window in the reference image are represented by W' and H', respectively. The top-left position of the consistency window in the reference image is represented as (X0, Y0).
[0671] a. In one example, minW equals 0.
[0672] b. In one example, minW equals X0.
[0673] c. In one example, maxW equals W-1.
[0674] d. In one example, maxW equals W'-1.
[0675] e. In one example, maxW equals X0 + W' - 1.
[0676] f. In one instance, minW and / or maxW can be modified based on the color format and / or color components.
[0677] i. For example, change minW to minW*SubC.
[0678] ii. For example, change minW to minW / SubC.
[0679] iii. For example, change maxW to maxW*SubC.
[0680] iv. For example, change maxW to maxW / SubC.
[0681] v. In one example, for the luminance component, SubC equals 1.
[0682] vi. In one example, when the color format is 4:2:0, SubC equals 2 for the chromaticity component.
[0683] vii. In one example, when the color format is 4:2:2, SubC equals 2 for the chromaticity component.
[0684] viii. In one example, when the color format is 4:4:4, SubC equals 1 for the chromaticity component.
[0685] g. In one example, whether and / or how to perform pruning may depend on the dimensions of the current image (or its consistency window) and the dimensions of the reference image (or its consistency window).
[0686] i. In one example, cropping is performed only if the dimensions of the current image (or its consistency window) are different from the dimensions of the reference image (or its consistency window).
[0687] 29. A proposal is made to crop the integer part of the vertical coordinates of the reference sample points to [minH, maxH]. Assume the width and height of the consistency window of the reference image are represented by W and H, respectively. The width and height of the consistency window of the reference image are represented by W' and H', respectively. The upper left position of the consistency window in the reference image is represented as (X0, Y0).
[0688] a. In one example, minH equals 0.
[0689] b. In one instance, minH equals Y 0.
[0690] c. In one example, maxH equals H-1.
[0691] d. In one example, maxH equals H'-1.
[0692] e. In one example, maxH equals Y0 + H' - 1.
[0693] f. In one instance, minH and / or maxH can be modified based on the color format and / or color components.
[0694] i. For example, change minH to minH*SubC.
[0695] ii. For example, change minH to minH / SubC.
[0696] iii. For example, change maxH to maxH*SubC.
[0697] iv. For example, change maxH to maxH / SubC.
[0698] v. In one example, for the luminance component, SubC equals 1.
[0699] vi. In one example, when the color format is 4:2:0, SubC equals 2 for the chromaticity component.
[0700] vii. In one example, when the color format is 4:2:2, SubC equals 1 for the chromaticity component.
[0701] viii. In one example, when the color format is 4:4:4, SubC equals 1 for the chromaticity component.
[0702] g. In one example, whether and / or how to perform pruning may depend on the dimensions of the current image (or its consistency window) and the dimensions of the reference image (or its consistency window).
[0703] i. In one example, cropping is performed only if the dimensions of the current image (or its consistency window) are different from the dimensions of the reference image (or its consistency window).
[0704] In the following discussion, if two syntax elements have equivalent functionality but can be used for signaling notification in different video units (e.g., VPS / SPS / PPS / strip header / image header, etc.), the first syntax element is said to "correspond" to the second syntax element.
[0705] 30. It is proposed that syntax elements be signaled in the first video unit (e.g., image header or PPS), and that corresponding syntax elements not be signaled in the second video unit at a higher level (e.g., SPS) or a lower level (e.g., strip header).
[0706] a. Alternatively, the first syntax element can be signaled in the first video unit (e.g., image header or PPS), and the corresponding second syntax element (e.g., strip header) can be signaled in the lower-level second video unit.
[0707] i. Alternatively, a signaling notification indicator can be included in the second video unit to indicate whether a signaling notification is subsequently sent to the second syntax element.
[0708] ii. In one example, if signaling notifies a second syntax element, the stripe associated with the second video unit (e.g., the stripe header) may follow the instructions of the second syntax element instead of the instructions of the first syntax element.
[0709] iii. In the first video unit, a signaling notification may be made to an indicator associated with the first syntax element to indicate whether a signaling notification may be made to the second syntax element in any stripe (or other video unit) associated with the first video unit.
[0710] b. Alternatively, the first syntax element (e.g., VPS / SPS / PPS) can be signaled in the first video unit at a higher level, and the corresponding second syntax element (e.g., image header) can be signaled in the second video unit.
[0711] i. Alternatively, a signaling notification indicator can be used to indicate whether a second syntax element will subsequently be signaled.
[0712] ii. In one example, if signaling notifies the second syntax element, the picture associated with the second video unit (which can be segmented into strips) can follow the instructions of the second syntax element instead of the instructions of the first syntax element.
[0713] c. The first syntax element in the image header may have the same functionality as the second syntax element in the stripe header as specified in Section 2.6 (e.g., but not limited to slice_temporal_mvp_enabled_flag, cabac_init_flag, six_minus_max_num_merge_cand, five_minus_max_num_subblock_merge_cand, slice_fpel_mmvd_enabled_flag, slice_disable_bdof_dmvr_flag, max_num_merge_cand_minus_max_num_triangle_cand, slice_fpel_mmvd_enabled_flag, slice_six_minus_max_num_ibc_merge_cand, slice_joint_cbcr_sign_flag, slice_qp_delta, etc.), but can control all stripes of the image.
[0714] d. The first syntax element in the SPS as specified in Section 2.6 may have the same functionality as the second syntax element in the picture header (e.g., but not limited to sps_bdof_dmvr_slice_present_flag, sps_mmvd_enabled_flag, sps_isp_enabled_flag, sps_mrl_enabled_flag, sps_mip_enabled_flag, sps_cclm_enabled_flag, sps_mts_enabled_flag, etc.), but may only control the associated picture (which may be divided into multiple stripes).
[0715] e. The first syntax element in the PPS as specified in Section 2.7 may have the same functionality as the second syntax element in the picture header (e.g., but not limited to entropy_coding_sync_enabled_flag, entry_point_offsets_present_flag, cabac_init_present_flag, rpl1_idx_present_flag, etc.), but may only control the associated picture (which may be segmented into multiple stripes).
[0716] 31. The syntax elements for signaling notifications in the image header are decoupled from other syntax elements for signaling notifications or derivations in SPS / VPS / DPS.
[0717] 32. Enable / disable indications for DMVR and BDOF can be separately signaled in the picture header, instead of being controlled by the same flag (e.g., pic_disable_bdof_dmvr_flag).
[0718] 33. Signaling notifications can be included in the image header to indicate whether PROF / Across Component ALF / Inter-frame prediction using Geometric Segmentation (GEO) is enabled / disabled.
[0719] a. Alternatively, the PROF enable flag in the SPS can be used as a conditional signaling notification of the PROF enable / disable instruction in the picture header.
[0720] b. Alternatively, the signaling notification in the picture header indicating whether to enable / disable cross-component ALF (CCALF) can be conditionally sent based on the CCALF enable flag in the SPS.
[0721] c. Alternatively, the signaling notification in the picture header indicating whether GEO is enabled or disabled can be conditionally sent based on the GEO enable flag in the SPS.
[0722] d. Alternatively, the signaling notification in the strip header may conditionally indicate whether PROF / ALF / inter-frame prediction using geometric segmentation (GEO) is enabled / disabled, based on those syntax elements in the picture header rather than in the SPS.
[0723] 34. The prediction type of strips / blocks / pieces (or other video units smaller than the image) in the same image can be indicated by signaling in the image header.
[0724] a. In one example, the signaling in the picture header can indicate whether all stripes / blocks / pieces (or other video units smaller than the picture) are intra-frame codecs (e.g., all I stripes).
[0725] i. Alternatively, if the indication tells that all stripes within the image are I stripes, then the stripe type may not be notified in the stripe header signaling.
[0726] b. Alternatively, the image header may include a signaling indication of whether at least one of the stripes / blocks / pieces (or other video units smaller than the image) is not an intra-frame codec (e.g., at least one non-I stripe).
[0727] c. Alternatively, the image header can be signaled to indicate whether all stripes / blocks / pieces (or other video units smaller than the image) have the same prediction type (e.g., I / P / B stripes).
[0728] i. Alternatively, the signaling notification of the strip type may be omitted from the strip header.
[0729] ii. Alternatively, in addition, conditional signaling may be used to indicate which tools are permitted for a particular forecast type (e.g., DMVR / BDOF / TPM / GEO are permitted only for B-strips; dual-tree is permitted only for I-strips).
[0730] d. Alternatively, the signaling for enabling / disabling tools may also depend on the prediction type of the instruction mentioned in the above sub-entries.
[0731] i. Alternatively, the instructions for enabling / disabling the tool can be derived from the instructions for the prediction type mentioned in the above sub-entries.
[0732] 35. In this disclosure (items 1-29), the term "consistency window" may be replaced by other terms such as "scaling window". The scaling window may be signaled differently from the consistency window and used to derive the scaling ratio and / or upper-left offset, which are used to derive the reference sample location for RPR.
[0733] a. In one example, the scaling window can be constrained by the consistency window. For example, in a consistent bitstream, the scaling window must be contained within the consistency window.
[0734] 36. Whether and / or how the maximum allowed block size for transform-skip-decode can be indicated by signaling may depend on the maximum block size for transform-decode.
[0735] a. Alternatively, in a consistent bitstream, the maximum block size for transform-skip-decode blocks cannot be greater than the maximum block size for transform-decode blocks.
[0736] 37. Whether and how signaling notifications indicate whether Joint Cb-Cr Residual (JCCR) encoding / decoding is enabled (e.g., sps_joint_cbcr_enabled_flag) can depend on the color format (e.g., 4:0:0, 4:
[0737] (e.g., 2:0)
[0738] a. For example, if the color format is 4:0:0, an indication to enable Joint Cb-Cr Residuals (JCCR) can be provided without signaling. An exemplary syntax design is as follows:
[0739]
[0740]
[0741] Determining the use of coding tool X
[0742] 38. The type of downsampling filter used for the hybrid weight derivation of chroma samples can be notified at the video unit level (e.g., SPS / VPS / PPS / picture header / subpicture / strip / strip header / piece / brick / CTU / VPDU level).
[0743] a. In one example, a high-level flag can be signaled to switch between different chroma format types of the content.
[0744] i. In one example, a high-level flag can be signaled to switch between chroma format type 0 and chroma format type 2.
[0745] ii. In one example, a signaling notification flag can be used to specify whether the luminance weights of the left and right top and bottom samples in the TPM / GEO prediction mode co-occur with the top left luminance weight (e.g., chroma sample position type 0).
[0746] iii. In one example, a signaling flag can be used to specify whether the lower left upper and lower sampled luminance samples in the TPM / GEO prediction module are co-located with the upper left luminance sample in the horizontal direction, but are vertically displaced by 0.5 units relative to the upper left luminance sample (e.g., sample position type 2).
[0747] b. In one example, the type of downsampling filter can be notified by signaling for 4:2:0 chroma format and / or 4:2:2 chroma format.
[0748] c. In one example, a signaling notification flag can be used to specify the type of chroma downsampling filter used for TPM / GEO prediction.
[0749] i. In one example, a signaling flag can be used to indicate whether downsampling filter A or downsampling filter B is used for chroma weight derivation in TPM / GEO prediction mode. 39. The type of downsampling filter used for hybrid weight derivation of chroma samples can be derived at the video unit level (e.g., SPS / VPS / PPS / picture header / subpicture / strip / strip header / piece / brick / CTU / VPDU level).
[0750] a. In one example, a lookup table can be defined to specify the correspondence between chroma subsampling filter types and the chroma format types of the content.
[0751] 40. When the chromaticity sample point location types are different, the specified downsampling filter can be used for TPM / GEO prediction mode.
[0752] a. In one example, under a certain chroma sample location type (e.g., chroma sample location type 0), the chroma weights of TPM / GEO can be subsampled from the co-located top-left luminance weights.
[0753] b. In one example, under a certain chroma sample location type (e.g., chroma sample location type 0 or 2), a specified X-tap filter (where X is a constant, e.g., X = 6 or 5) can be used for chroma weighted subsampling in the TPM / GEO prediction mode.
[0754] 41. In a video unit (e.g., SPS, PPS, picture header, strip header, etc.), a signaling instruction can be made to the first syntax element (e.g., a flag) to indicate whether all blocks (strips / pictures) have Multi-Transform Selection (MTS) disabled.
[0755] a. Based on the first syntax element, conditionally signal a second syntax element indicating how MTS (e.g., enabling MTS / disabling MTS / implicit MTS / explicit MTS) should be applied on intra-frame codec blocks (strips / pictures). For example, the second syntax element is signaled only if the first syntax element indicates that MTS is not disabled on all blocks (strips / pictures).
[0756] b. Based on the first syntax element, conditionally signal a third syntax element indicating how MTS (e.g., enabling MTS / disabling MTS / implicit MTS / explicit MTS) should be applied on inter-frame codec blocks (strips / pictures). For example, the third syntax element is signaled only if the first syntax element indicates that MTS is not disabled on all blocks (strips / pictures).
[0757] c. An example syntax design is as follows:
[0758]
[0759] d. Conditionally signal third syntax elements based on whether Subblock Transform (SBT) is applied. An example syntax design is as follows:
[0760]
[0761] e. An example syntax design is as follows:
[0762]
[0763]
[0764] Restrictions on slice / tile related dimensions
[0765] 42. Determining whether and / or how to enable codec tool X may depend on one or more reference images and the width and / or height of the current image.
[0766] a. The width and / or height of one or more reference images and / or the current image may be modified to make a determination.
[0767] b. Consider that the image can be defined by a consistency window or a scaling window.
[0768] i. Consider that the image can be the entire image.
[0769] c. In one example, whether and / or how to enable codec tool X may depend on the image's width minus one or more horizontal offsets and / or the image's height minus the vertical offset.
[0770] i. In one example, the horizontal offset can be defined as scaling_win_left_offset.
[0771] ii. In one example, the vertical offset can be defined as scaling_win_top_offset.
[0772] iii. In one example, the horizontal offset can be defined as (scaling_win_right_offset + scaling_win_left_offset).
[0773] iv. In one example, the vertical offset can be defined as (scaling_win_bottom_offset + scaling_win_top_offset).
[0774] v. In one example, the horizontal offset can be defined as SubWidthC*(scaling_win_right_offset+scaling_win_left_offset).
[0775] vi. In one example, the vertical offset can be defined as SubHeightC*(scaling_win_bottom_offset+scaling_win_top_offset).
[0776] d. In one example, if at least one of the two considered reference images has a different resolution (width or height) than the current image, then the codec tool X is disabled.
[0777] i. Alternatively, if at least one of the two output reference images has a dimension (width or height) greater than the dimension (width or height) of the current image, then the codec tool X is disabled.
[0778] e. In one example, if a considered reference image used for the reference image list L has a different resolution than the current image, then the codec tool X is disabled for the reference image list L.
[0779] i. Alternatively, if a considered reference image used for the reference image list L has dimensions (width or height) greater than the current image, then the encoding / decoding tool X is disabled for the reference image list L.
[0780] f. In one example, if two considered reference images in two lists of reference images have different resolutions, the encoding / decoding tools can be disabled.
[0781] i. Alternatively, the signaling can conditionally notify the codec tool of instructions based on the resolution.
[0782] ii. Alternatively, signaling instructions from the codec tool can be skipped.
[0783] g. In one example, if two considered reference images of two merge candidates are used to derive the first pair of merge candidates from at least one list of reference images, the encoding / decoding tools can be disabled, for example, the first pair of merge candidates can be marked as unavailable.
[0784] i. Alternatively, if two considered reference images of two merge candidates are used to derive the first pair of merge candidates from the list of two reference images, the encoding / decoding tools can be disabled, for example, the first pair of merge candidates can be marked as unavailable.
[0785] h. In one example, the decoding process of the encoding / decoding tool can be modified to take into account the dimensions of the image.
[0786] i. In one example, the derivation of the MVD of another list of reference images in SMVD (e.g., List 1) can be based on the resolution difference (e.g., scaling factor) of at least one of the two target SMVD reference images.
[0787] ii. In one example, the derivation of pairwise merge candidates can be based on the resolution difference (e.g., scaling factor) of at least one of the two reference images associated with the two reference images, for example, by applying a linear weighted average instead of equal weights.
[0788] i. In one example, X could be:
[0789] i.DMVR / BDOF / PROF / SMVD / MMVD / other codec tools for refining motion / prediction on the decoder side
[0790] ii. TMVP / Other codec tools that rely on temporal motion information
[0791] iii. MTS or other transform codec tools
[0792] iv.CC-ALF
[0793] v.TPM
[0794] vi.GEO
[0795] vii. Switchable interpolation filter (e.g., an alternative interpolation filter for half-pixel motion compensation)
[0796] viii. The mixing process in TPM / GEO / other codec tools divides a block into multiple segments.
[0797] ix. Encoding and decoding tools that rely on information stored in images different from the current image.
[0798] x. Paired merge candidates (paired merge candidates are not generated when certain resolution-related conditions are not met)
[0799] xi. Bidirectional prediction with CU-level weights (BCW)
[0800] xii. Weighted Prediction
[0801] xiii. Affine prediction
[0802] xiv. Adaptive Motion Vector Resolution (AMVR)
[0803] 43. Whether and / or how signaling notification is used for encoding / decoding tools may depend on the width and / or height of one or more reference images and / or the current image.
[0804] a. The width and / or height of one or more reference images and / or the current image may be modified to make a determination.
[0805] b. Consider that the image can be defined by a consistency window or a scaling window as defined in JVET-P0590.
[0806] xv. Consider that the image can be the entire image.
[0807] c. In one example, X could be Adaptive Motion Vector Resolution (AMVR).
[0808] d. In one example, X could be the merge(MMVD) method that utilizes MV difference.
[0809] xvi. In one example, the construction of the symmetric motion vector difference reference index may depend on the image resolution / indication of the RPR of different reference images.
[0810] e. In one example, X could be a symmetric MVD (SMVD) method.
[0811] f. In one example, X can be QT / BT / TT or other splitting types.
[0812] g. In one example, X can be a bidirectional prediction with CU-level weights (BCW).
[0813] h. In one example, X could be a weighted prediction.
[0814] i. In one example, X could be an affine prediction.
[0815] j. In one example, whether signaling notifications indicate the use of half-pixel motion vector precision / switchable interpolation filters can depend on resolution information / whether RPR is enabled for the current block.
[0816] k. In one example, the signaling for amvr_precision_idx can depend on the resolution information / whether RPR is enabled for the current block.
[0817] l. In one example, the signaling of sym_mvd_flag / mmvd_merge_flag may depend on the resolution information / whether RPR is enabled for the current block.
[0818] m. When the width and / or height of one or more reference images differ from the current output image, the consistency bitstream should meet 1 / 2 pixel MV and / or MVD precision (e.g., alternative interpolation filters / switchable interpolation filters) which is not allowed.
[0819] 44. Propose that blocks in RPR can still enable AMVR with 1 / 2 pixel MV and / or MVD accuracy (or alternative interpolation filters / switchable interpolation filters).
[0820] a. Alternatively, different interpolation filters can be applied to blocks with 1 / 2 pixel or other precision.
[0821] 45. The condition check for the same / different resolution in the above items can be replaced by adding a flag to the reference image and checking the flags associated with the reference image.
[0822] a. In one example, a procedure can be invoked during the reference image list construction process to set a flag to true or false (e.g., to indicate whether the reference image is an RPR case or a non-RPR case).
[0823] xvii. For example, the following can be applied:
[0824] fRefWidth is set to equal to PicOutputWidthL of the reference image RefPicList[i][j] in units of luminance samples, where PicOutputWidthL represents the width of the considered image of the reference image.
[0825] fRefWidth is set to equal PicOutputHeightL of the reference image RefPicList[i][j] in units of luminance samples, where PicOutputHeightL represents the height of the reference image in the consideration image.
[0826] RefPicScale[i][j][0]=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL, where
[0827] PicOutputWidthL represents the width of the current image as a consideration.
[0828] RefPicScale[i][j][1] = ((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL, where PicOutputWidthL represents the height of the current image considering the image.
[0829] RefPicIsScaled[i][j]=(RefPicScale[i][j][0]!=(1<<14))||
[0830] (RefPicScale[i][j][1]!=(1<<14))
[0831] Where RefPicList[i][j] represents the j-th reference image in the reference image list i.
[0832] b. In one example, when RefPicIsScaled[0][refIdxL0] is not equal to 0 or RefPicIsScaled[1][refIdxL1] is not equal to 0, the codec tool X (e.g., DMVR / BDOF / SMVD / MMVD / SMVD / PROF / those tools mentioned in the above entries) can be disabled.
[0833] c. In one example, when both RefPicIsScaled[0][refIdxL0] and RefPicIsScaled[1][refIdxL1] are not equal to 0, the codec tool X (e.g., DMVR / BDOF / SMVD / MMVD / SMVD / PROF / those tools mentioned in the above entries) can be disabled.
[0834] d. In one example, when RefPicIsScaled[0][refIdxL0] is not equal to 0, the codec tool X (e.g., PROF or those tools mentioned in the above entries) can be disabled for reference image list 0.
[0835] e. In one example, when RefPicIsScaled[1][refIdxL1] is not equal to 0, the codec tool X (e.g., PROF or those tools mentioned in the above entries) can be disabled for reference image list 1.
[0836] 46. The SAD and / or threshold used in BDOF / DMVR can depend on the bit depth.
[0837] f. In one example, the calculated SAD value may be shifted by a function of bit depth before being used for comparison with a threshold.
[0838] g. In one example, the calculated SAD value can be directly compared to a modified threshold, which can depend on a function of the bit depth.
[0839] 47. If the slice_type value of all slices in an image is equal to I (I slice), then the syntax elements related to the P / B slices do not need to be signaled in the image header.
[0840] a. In one example, multiple syntax elements can be added to the image header to indicate whether the slice_type of all slices included in a particular image is equal to I(I_slice).
[0841] i. In one example, the first syntax element can signal the notification in the picture header. The second syntax element, which determines whether and / or how the notification signals / interprets the strip type information in the strip header associated with the picture header, can depend on the first syntax element.
[0842] 1) In one example, depending on the first syntax element, the second syntax element can be notified without signaling and inferred to be of stripe type.
[0843] 2) In one example, the second syntax element can be signaled, but the consistency requirement stipulates that the second syntax element must be one of several given values according to the first syntax element.
[0844] 3) Alternatively, the first syntax element can be signaled in the AU delimiter RBSP associated with the stripe.
[0845] ii. In one example, a new syntax element (e.g., pic_all_X_slices_flag) can signal in the picture header to indicate whether only X slices are allowed for the picture, or whether all slices in the picture are X slices. For example, X can be I, P, or B.
[0846] 1) In one example, if the associated image header indicates that all stripes are I-stripes, then no signaling notification of stripe type information is given in the stripe header and the stripe is inferred to be an I-strip.
[0847] iii. In one example, a new syntax element (e.g., ph_pic_type) can be signaled in the image header to indicate the image type.
[0848] 1) For example, if ph_pic_type equals I-image (e.g., equals 0), then only slice_type equal to I is allowed in images.
[0849] 2) For example, if ph_pic_type is equal to a non--I image (e.g., 1 or 2), then slice_type in the image can be equal to I and / or P and / or B.
[0850] b. In one example, the syntax element pic_type in the AU delimiter RBSP can be used to indicate whether all stripes of a particular image are equal to I.
[0851] c. In one example, if all stripes in an image are I stripes, then for all stripes in this image, the slice_type syntax element in the slice header can be inferred as an I stripe without signaling notification (e.g., 2).
[0852] i. In one example, if a syntax element (such as, but not limited to, pic_all_I_slices_flag, ph_pic_type, pic_type) indicates that all slices included in a particular picture are equal to I, a bitstream constraint can be added to specify that the slice_type of each slice in a particular picture should be equal to the I slice.
[0853] ii. Alternatively, if a syntax element (e.g., but not limited to pic_all_I_slices_flag, ph_pic_type, pic_type) indicates that all slices contained in a particular picture are equal to I, a bitstream constraint can be added to specify that P or B slices are not allowed in a particular picture.
[0854] d. If the stripe / picture is a W-stripe / W-picture, then one or more syntax elements (represented as the set of syntax elements X specified below) are allowed for non-W stripes in the stripe header / picture header without signaling. For example, W can be I, and non-W can be B or P. In another example, W can be B, and non-W can be I or P.
[0855] i. In one example, if all stripes in the image are W stripes, then no signaling is required to notify the image header of the set of syntax elements X that are allowed for non-W stripes.
[0856] ii. Alternatively, based on whether all stripes in the image are W stripes, conditional signaling may be used to notify some syntax elements in the image header (such as the set X specified below).
[0857] iii. The set of syntax elements X can be one or more of the following.
[0858] 1) In one example, X can be a reference image-related syntax element in the image header, such as, but not limited to, pic_rpl_present_flag, pic_rpl_sps_flag, pic_rpl_idx, pic_poc_lsb_lt, pic_delta_poc_msb_present_flag, pic_delta_poc_msb_cycle_lt... If the stripe / image is signaled or inferred to be a non-inter-frame stripe / image, then X may not be signaled and may be inferred to be unused.
[0859] 2) In one example, X can be an inter-slice related syntax element in the image header, such as, but not limited to, pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice... If the slice / image is signaled or inferred to be non-inter-slice / image, then X may not be signaled and may be inferred to be unused.
[0860] 3) In one example, X can be an inter-frame prediction related syntax element in the image header, such as, but not limited to, pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_six_minus_max_num_merge_cand, pic_five_minus_max_num_subblock_merge_cand, pic_fpel_mmvd_enabled_flag, pic_disable_bdof_flag, pic_disable_dmvr_flag, pic_disable_prof_flag, pic_max_num_merge_cand_minus_max_num_triangle_cand…
[0861] If a stripe / picture is signaled or inferred to be a non-inter-frame stripe / picture, then X may not be signaled and may be inferred to be unused.
[0862] 4) In one example, X can be a bidirectional prediction-related syntax element in the image header, such as, but not limited to, pic_disable_bdof_flag, pic_disable_dmvr_flag, mvd_l1_zero_flag…
[0863] If a stripe / picture is signaled or inferred to be a non-B stripe / picture, then X may not be signaled and may be inferred to be unused.
[0864] Chroma QP mapping table
[0865] 48. The maximum slice width can be specified in the specification.
[0866] a. For example, the maximum panel width can be defined as the maximum luminance panel width in CTB units.
[0867] b. In one example, within a video unit (e.g., SPS, PPS, picture header, strip header, etc.), a new syntax element can be signaled to indicate the maximum slice width allowed in the current sequence / picture / strip / sub.
[0868] c. In one example, the maximum luminance pane width or the maximum luminance pane width in CTB can be signaled.
[0869] d. In one example, the maximum brightness pane width can be fixed at N (e.g., N = 1920 or 4096, etc.).
[0870] e. In one example, different maximum slice widths can be specified in different configuration files / levels / layers.
[0871] 49. Maximum strip / sub-image / piece dimensions (e.g., width and / or size, and / or length, and / or height) can be specified in the specification.
[0872] f. For example, the maximum strip / sub-image / piece dimension can be defined as the maximum brightness dimension in CTB units.
[0873] g. For example, the size of a strip / sub-image / piece can be defined as the number of CTBs in the strip.
[0874] h. In one example, within a video unit (e.g., SPS, PPS, picture header, strip header, etc.), a new syntax element(s) can be signaled to indicate the maximum allowed brightness strip / subpicture / piece dimension in the current sequence / picture / strip / subpicture.
[0875] i. In one example, for a rectangular strip / sub-image, the maximum brightness strip / sub-image width / height can be signaled or the maximum brightness strip / sub-image width / height can be expressed in CTB units.
[0876] j. In one example, for a raster scan strip, the maximum brightness strip size (e.g., width * height) or the maximum brightness strip size in CTB can be signaled.
[0877] k. In one example, for both rectangular strips / sub-images and raster scan strips, the maximum brightness strip / sub-image size (e.g., width * height) or the maximum brightness strip / sub-image size in CTB can be signaled.
[0878] l. In one example, for a raster scan strip, the maximum brightness strip length or the maximum brightness strip length in CTB can be signaled.
[0879] m. In one example, the width / height of the maximum brightness stripe / subimage can be fixed at N (e.g., N = 1920 or 4096, etc.).
[0880] In one example, the maximum brightness stripe / sub-image size (e.g., width * height) can be fixed at N (e.g., N = 2073600 or 83388608, etc.).
[0881] o. In one example, different maximum stripe / sub-image dimensions can be specified in different configuration files / levels / layers.
[0882] p. In one example, the maximum number of strips / sub-pictures / pieces to be divided into in a picture / sub-picture can be specified in the specification.
[0883] i. The maximum number of signals can be notified.
[0884] ii. The maximum number of configuration files / levels / layers can be different.
[0885] 50. Assume the width and height of the current image are denoted as PW and PH, respectively; the width and height of the current image scaling window are denoted as SW and SH, respectively; the width and height of the reference image scaling window are denoted as SW' and SH', respectively; and the maximum allowed image width and height are denoted as Wmax and Hmax, respectively. For convenience, let... and At least one of the following constraints should follow a consistent bitstream. These constraints should not be interpreted narrowly. For example, the constraint (a / b) >= (c / d), where a, b, c, and d are integers greater than 0, can also be interpreted as (a / b) - (c / d) >= 0, or a*d >= c*b, or a*dc*d >= 0.
[0886] qa×Wmax×SW-(b×SW′+c×SW)×PW+offw≥0, where a, b, c, and offw are integers. For example, a=135, b=128, c=7, and offw=0.
[0887] rd×Hmax×SH-(e×SH′+f×SH)×PH+offh≥0, where d, e, and f are integers. For example, d=135, e=128, f=7, and offh=0.
[0888] sa×Wmax×SW-b×SW′×PW+offw≥0, where a and b are integers. For example, a=1, b=1 and offw=0.
[0889] td×Hmax×SH-e×SH′×PH+offh≥0, where d and e are integers. For example, d=1, e=1 and offh=0.
[0890] u. Where Lw, Bw, and offw are integers. For example, Lw = 7, Bw = 128, and offw = 0.
[0891] v. Where Lh, Bh, and offh are integers. For example, Lh = 7, Bh = 128, and offh = 0.
[0892] w.rw≤a*Rw+offw, where a and offw are integers. For example, a=1 and offw=0.
[0893] x.rh≤b*Rh+offh, where b and offh are integers. For example, b=1 and offh=0.
[0894] y.rw≤a*Qw+offw, where a and offw are integers. For example, a=1 and offw=0.
[0895] z.rh≤b*Qh+offh, where b and offh are integers. For example, b=1 and offh=0.
[0896] 51. Signaling information related to QP (e.g., delta QP) can be provided in the image header but not in the PPS.
[0897] a. In one example, QP-related information (e.g., delta QP) is specified for a specific codec tool, such as Adaptive Color Transformation (ACT).
[0898] 52. A slice type is proposed that can be signaled in a slice or in an associated first video unit (VU) (such as a PPS or picture header) that can be associated with multiple slices. (For example, the slice type may include at least one of I-slice, P-slice or B-slice).
[0899] b. Propose that a first syntax element (SE) (e.g., a flag named ph_same_slice_type_flag) can be signaled in the first VU to indicate whether all codec slices associated with the first VU have the same slice type.
[0900] c. Propose that when all slices have the same slice type, the second SE (e.g., an SE named ph_slice_type) can be signaled in the first VU to indicate the slice type of all slices associated with the first VU.
[0901] i. In one example, when B-banding is not allowed at a higher level (e.g., in SPS), the second SE cannot indicate the B-banding type.
[0902] d. Propose a conditional signaling notification to the second SE based on the first SE.
[0903] i. For example, the signaling notifies the second SE only when the first SE indicates that all codec slices associated with the first VU have the same slice type (e.g., ph_same_slice_type_flag equals 1).
[0904] e. In one example, the second SE is binarized or interpreted in the same way as the SE indicating the slice type (e.g., slice_type) in the signaling notification in the slice header.
[0905] f. Propose that, based on the first and / or second SE, the SE indicating the slice type (e.g., slice_type) can be conditionally signaled in the slice header.
[0906] i. In one example, the SE used to indicate the slice type (e.g., slice_type) is notified in the slice header signaling only when the first SE indicates that the slice associated with the first VU has a different slice type (e.g., ph_same_slice_type_flag equals 0).
[0907] g. Propose that, based on the first and / or second SE, when it is not present, the SE indicating the stripe type (e.g., slice_type) in the stripe header can be inferred as the default value.
[0908] i. If it does not exist, infer that the value of slice_type is equal to (ph_same_slice_type_flag?ph_slice_type:(ph_inter_slice_allowed_flag?1:2)).
[0909] h. One or more syntax elements may be signaled in the SPS and / or PPS and / or PH to indicate whether slice_type is allowed to be X in the current video unit (e.g., CLVS / a set of pictures / pictures, such as X equals I or P or B, e.g., as in the embodiment in Section 5.13).
[0910] ii. Additionally, when a syntax element in CLVS that disallows B-slices (named sps_b_slice_allowed_flag equal to 0) is used, semantic constraints are added to the PPS syntax element pps_weighted_bipred_flag. For example, the semantic changes to pps_weighted_bipred_flag are as follows (the added parts are shown in bold italics with underlines):
[0911] A `pps_weighted_bipred_flag` value of 0 specifies that explicit weighted predictions should not be applied to the B-strips of the reference PPS. A `pps_weighted_bipred_flag` value of 1 specifies that explicit weighted predictions should be applied to the B-strips of the reference PPS. When `pps_weighted_bipred_flag` is equal to or equal to 0... When the value is 0, the value of pps_weighted_bipred_flag should be 0.
[0912] i. Propose that, based on a first SE and / or a second SE, one or more SEs may be conditionally signaled in a first VU to constrain the stripe type of the stripe associated with the first VU.
[0913] i. For example, signaling notifies one or more SEs only when the first SE indicates that the codec slice associated with the first VU has a different slice type (e.g., ph_same_slice_type_flag equals 0).
[0914] ii. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to that of an I-slice. For example, the SE could be the ph_inter_slice_allowed_flag in the picture header.
[0915] iii. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to B-slice or P-slice. For example, the SE may be ph_intra_slice_allowed_flag in the picture header.
[0916] iv. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to P-slice or I-slice. For example, the SE may be ph_b_slice_allowed_flag in the picture header.
[0917] v. For example, when ph_same_slice_type_flag equals 1, the syntax elements ph_inter_slice_allowed_flag and / or ph_intra_slice_allowed_flag and / or ph_X_slice_allowed_flag (indicating whether slice_type equal to X is allowed for all slices in the current picture, e.g., X equals I, P, or B) are not signaled and their values can be inferred, e.g., as in the embodiment in 5.13. For example, when ph_same_slice_type_flag equals 0 and the syntax element specifies that slices B are not allowed (e.g., ph_b_slice_allowed_flag equals 0), the syntax element ph_intra_slice_allowed_flag may not be signaled and is inferred to be a value (e.g., 1), e.g., as in the embodiment in 5.13.
[0918] 1) Additionally, when ph_same_slice_type_flag equals 0, ph_intra_slice_allowed_flag is sent after ph_b_slice_allowed_flag and the signaling of ph_intra_slice_allowed_flag is conditional on ph_b_slice_allowed_flag.
[0919] vi. For example, when ph_same_slice_type_flag equals 0, specify whether inter-slice syntax elements (e.g., ph_inter_slice_allowed_flag) are allowed to be unsigned, for example, as in the embodiment in Section 5.13.
[0920] vii. Additionally, whether the following inter-frame related syntax elements are signaled in the PH depends on the value of ph_same_slice_type_flag and / or the value of ph_slice_type (e.g., ph_log2_diff_min_qt_min_cb_inter_slice, ph_max_mtt_hierarchy_depth_inter_slice, ph_log2_diff_max_bt_min_qt_inter_slice, ph_log2_diff_max_tt_min_qt_inter_slice, ph_cu_qp_delta_subdiv_inter_slice). ph_cu_chroma_qp_offset_subdiv_inter_slice,ph_temporal_mvp_enabled_flag,ph_collocated_from_l0_flag,ph_collocated_ref_idx,mvd_l1_zero_flag,ph_fpel_mmvd_enabled_flag,ph_disable_bdof_flag,ph_disable_dmvr_flag,ph_disable_prof_flag, and the syntax structure pred_weight_table(), for example, as in the embodiment in Section 5.13.
[0921] 1. For example, when ph_same_slice_type_flag equals 1 or ph_slice_type is not equal to 2, the inter-frame associated syntax elements in the above PH can be signaled. Otherwise, their values can be inferred.
[0922] viii. Additionally, whether the slice_type syntax element in the signaling notification SH can depend on the value of ph_same_slice_type_flag, for example, the syntax of slice_type changes as follows (the added parts are shown in bold italics with underlined text, and the deleted parts are shown in [[]]):
[0923]
[0924] The semantic changes of slice_type are as follows:
[0925] If it does not exist, then the value of slice_type is inferred to be equal to... [[2]].
[0926] j. Propose that, based on the first SE and / or the second SE, when the second SE is not present, one or more SEs for constraining the stripe type associated with the first VU can be inferred as default values.
[0927] i. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to that of the I-slice. For example, the SE may be the ph_inter_slice_allowed_flag in the picture header.
[0928] 2. Example 1: When it does not exist, ph_inter_slice_allowed_flag is inferred to be equal to (!ph_same_slice_type_flag||(ph_slice_type<2?1:0)).
[0929] 3. Example 2: When it does not exist, ph_inter_slice_allowed_flag is inferred to be equal to (ph_same_slice_type_flag?(ph_slice_type<2?1:0):1).
[0930] 4. Example 3: When it does not exist, ph_inter_slice_allowed_flag is inferred to be equal to (ph_slice_type<2?1:0).
[0931] ii. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to B-slice or P-slice. For example, the SE may be ph_intra_slice_allowed_flag in the picture header.
[0932] 1. Example 1: When it does not exist, the value of ph_intra_slice_allowed_flag is inferred to be equal to (!ph_same_slice_type_flag||(ph_slice_type==2?1:0)).
[0933] 2. Example 2: When it does not exist, the value of ph_intra_slice_allowed_flag is inferred to be equal to (ph_same_slice_type_flag?(ph_slice_type==2?1:0):1).
[0934] iii. Those one or more SEs may include an SE indicating whether all codec slices associated with the first VU have a slice type equal to P-slice or I-slice. For example, the SE may be ph_b_slice_allowed_flag in the picture header.
[0935] 1. Example 1: When it does not exist, the value of ph_b_slice_allowed_flag is inferred to be equal to (ph_same_slice_type_flag?(ph_slice_type==0?1:0):(sps_b_slice_allowed_flag&&ph_inter_slice_allowed_flag)).
[0936] slice_type
[0937] 53. The starting point of the chroma QP mapping table can be sent directly.
[0938] k. In one example, qp_table_start_minus26 can be replaced by qp_table_start without being differentially encoded or decoded.
[0939] i. Alternatively, ue(v) can be used instead of se(v) to signal syntax elements.
[0940] 5. Additional Examples
[0941] The text changes are shown below in bold italic font with underlines.
[0942] 5.1 Examples of Constraints on Consistency Windows
[0943] Based on the rectangular region specified in the image coordinates used for output, conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset define the image samples in the CVS output during the decoding process. When conformance_window_flag equals 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.
[0944] The consistent cropping window contains luminance samples with horizontal image coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical image coordinates (including end values) from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1).
[0945] The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) should be less than pic_height_in_luma_samples.
[0946] The derivation of variables PicOutputWidthL and PicOutputHeightL is as follows:
[0947] PicOutputWidthL=pic_width_in_luma_samples- (7-43)
[0948] SubWidthC*(conf_win_right_offset+conf_win_left_offset)
[0949] PicOutputHeightL=pic_height_in_pic_size_units- (7-44)
[0950] SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)
[0951] When ChromaArrayType is not equal to 0, the corresponding specified sample points of the two chroma arrays are sample points with image coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the image coordinates of the specified brightness sample point.
[0952]
[0953] 5.2 Example 1 of the derivation of reference sample point location
[0954] 8.5.6.3.1 Overview
[0955] …
[0956] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[0957] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[0958]
[0959] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[0960] – If cIdx equals 0, then the following applies:
[0961] - Define the scaling factor and its fixed-point representation as:
[0962] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[0963] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL(8-754)
[0964] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[0965] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L ) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[0966] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L=0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[0967] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[0968]
[0969] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)
[0970]
[0971] refy L =((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)
[0972]
[0973] – The variable xInt is derived as follows L yInt L xFrac L and yFrac L :
[0974] xInt L =refx L >>4 (8-759)
[0975] yInt L =refy L >>4 (8-760)
[0976] xFrac L =refx L &15 (8-761)
[0977] yFrac L =refy L &15 (8-762)
[0978] If bdofFlag equals true or (sps_affine_prof_enabled_flag equals true and inter_affine_flag[xSb][ySb] equals true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L With >>3)-1) and refPicLX as inputs, the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[0979] -x L It equals 0.
[0980] -x L It equals sbWidth+1.
[0981] –y L It equals 0.
[0982] –y L It equals sbHeight + 1.
[0983] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt L With refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[0984] Otherwise (cIdx is not equal to 0), then the following applies:
[0985] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[0986] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[0987] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[0988] – Let (refxSb) C ,refySb C ) and (refx C ,refy C ) represents the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows: C ,refySb C refx C and refy C :
[0989]
[0990] refx C =((Sign(refxSb) C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)
[0991]
[0992] refy C =((Sign(refySb) C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5(8-766)
[0993]
[0994] – The variable xInt is derived as follows C yInt C xFrac C and yFrac C :
[0995] xInt C =refx C >>5 (8-767)
[0996] yInt C =refy C >>5 (8-768)
[0997] xFrac C =refy C &31 (8-769)
[0998] yFrac C =refy C &31 (8-770)
[0999] 5.3 Example 2 of the derivation of reference sample point positions
[1000] 8.5.6.3.1 Overview
[1001] …
[1002] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[1003] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[1004]
[1005] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[1006] – If cIdx equals 0, then the following applies:
[1007] - Define the scaling factor and its fixed-point representation as:
[1008] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[1009] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)
[1010] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[1011] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L ) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[1012] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[1013] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[1014]
[1015]
[1016] – The variable xInt is derived as follows L yInt L xFrac L and yFrac L :
[1017] xInt L =refx L >>4 (8-759)
[1018] yInt L =refy L >>4 (8-760)
[1019] xFrac L =refx L &15 (8-761)
[1020] yFrac L =refy L &15 (8-762)
[1021] If bdofFlag equals true or (sps_affine_prof_enabled_flag equals true and inter_affine_flag[xSb][ySb] equals true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L With >>3)-1) and refPicLX as inputs, the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[1022] -x L It equals 0.
[1023] -x L It equals sbWidth+1.
[1024] –y L It equals 0.
[1025] –y L It equals sbHeight + 1.
[1026] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt LWith refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[1027] Otherwise (cIdx is not equal to 0), then the following applies:
[1028] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[1029] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[1030] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[1031] – Let (refxSb) C ,refySb C ) and (refx C ,refyC) is the chromaticity position pointed to by the motion vector (mvLX[0],mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows C ,refySb C refx C and refy C :
[1032]
[1033] refx C =((Sign(refxSb) C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5(8-764)
[1034]
[1035] refy C =((Sign(refySb) C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5(8-766)
[1036]
[1037] – The variable xInt is derived as follows C yInt C xFrac C and yFrac C :
[1038] xInt C =refx C >>5 (8-767)
[1039] yInt C =refy C >>5 (8-768)
[1040] xFrac C =refy C &31 (8-769)
[1041] yFrac C =refy C &31 (8-770)
[1042] 5.4 Example 3 of the derivation of reference sample point positions
[1043] 8.5.6.3.1 Overview
[1044] …
[1045] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[1046] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[1047]
[1048] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[1049] – If cIdx equals 0, then the following applies:
[1050] - Define the scaling factor and its fixed-point representation as:
[1051] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[1052] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL(8-754)
[1053] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[1054] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L ) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[1055] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[1056] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[1057]
[1058]
[1059] – The variable xInt is derived as follows L yInt L xFrac L and yFrac L :
[1060] xInt L =refx L >>4 (8-759)
[1061] yInt L =refy L >>4 (8-760)
[1062] xFrac L =refx L &15 (8-761)
[1063] yFrac L =refy L &15 (8-762)
[1064] If bdofFlag equals true or (sps_affine_prof_enabled_flag equals true and inter_affine_flag[xSb][ySb] equals true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L With >>3)-1) and refPicLX as inputs, the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[1065] -x L It equals 0.
[1066] -x L It equals sbWidth+1.
[1067] –y L It equals 0.
[1068] –y L It equals sbHeight + 1.
[1069] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt L With refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[1070] Otherwise (cIdx is not equal to 0), then the following applies:
[1071] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[1072] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[1073] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[1074] – Let (refxSb) C ,refySb C ) and (refx C ,refy C ) represents the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows: C ,refySb C refx C and refy C :
[1075]
[1076] – The variable xInt is derived as follows CyInt C xFrac C and yFrac C :
[1077] xInt C =refx C >>5 (8-767)
[1078] yInt C =refy C >>5 (8-768)
[1079] xFrac C =refy C &31 (8-769)
[1080] yFrac C =refy C &31 (8-770)
[1081] 5.5 Example 1 of reference sample point position cutting
[1082] 8.5.6.3.1 Overview
[1083] The input for this process is:
[1084] – Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image.
[1085] – The variable sbWidth specifies the width of the current encoding / decoding sub-block.
[1086] – The variable sbHeight specifies the height of the current encoding / decoding sub-block.
[1087] – Motion vector offset mvOffset
[1088] – Refined motion vector refMvLX,
[1089] – The selected reference image sample array refPicLX,
[1090] – Half-sample interpolation filter index hpelIfIdx,
[1091] – Bidirectional optical flow flag bdofFlag
[1092] – The variable cIdx specifies the color component index of the current block.
[1093] The output of this process is:
[1094] – The array predSamplesLX represents the predicted sample values (sbWidth+brdExtSize)x(sbHeight+brdExtSize).
[1095] The predicted block boundary extension size brdExtSize is derived as follows:
[1096] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&sps_affine_prof_enabled_flag))? 2:0(8-752)
[1097] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[1098] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[1099] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[1100] – If cIdx equals 0, then the following applies:
[1101] - Define the scaling factor and its fixed-point representation as:
[1102] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[1103] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL(8-754)
[1104] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[1105] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[1106] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[1107] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[1108] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)
[1109] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)
[1110] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp(8-757)
[1111] refy L =((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)
[1112] – The variable xInt is derived as follows L yInt L xFracL and yFrac L :
[1113]
[1114] xFrac L =refx L &15 (8-761)
[1115] yFrac L =refy L &15 (8-762)
[1116] – If bdofFlag is true or (sps_affine_prof_enabled_flag is true and inter_affine_flag[xSb][ySb] is true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L >>
[1117] 3)-1) and refPicLX are inputs, and the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[1118] -x L It equals 0.
[1119] -x L It equals sbWidth+1.
[1120] –y L It equals 0.
[1121] –y L It equals sbHeight + 1.
[1122] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt LWith refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[1123] Otherwise (cIdx is not equal to 0), then the following applies:
[1124] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[1125] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[1126] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[1127] – Let (refxSb) C ,refySb C ) and (refx C ,refy C ) represents the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows: C ,refySb C refx C and refy C :
[1128] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)
[1129] refx C =((Sign(refxSb) C )*((Abs(refxSb C)+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)
[1130] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp(8-765)
[1131] refy C =((Sign(refySb) C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5(8-766)
[1132] – The variable xInt is derived as follows C yInt C xFrac C and yFrac C :
[1133]
[1134] xFrac C =refy C &31 (8-769)
[1135] yFrac C =refy C &31 (8-770)
[1136] – The predicted sample values predSamplesLX[xC][yC] are derived by invoking the procedure specified in Clause 8.5.6.3.4 with (xIntC,yIntC), (xFracC,yFracC), (xSbIntC,ySbIntC), sbWidth, sbHeight, and refPicLX as inputs.
[1137] 5.6 Example 2 of reference sample point position cutting
[1138] 8.5.6.3.1 Overview
[1139] The input for this process is:
[1140] – Luminance position (xSb, ySb), specifies the top-left sample of the current codec sub-block relative to the top-left luminance sample of the current image.
[1141] – The variable sbWidth specifies the width of the current encoding / decoding sub-block.
[1142] – The variable sbHeight specifies the height of the current encoding / decoding sub-block.
[1143] – Motion vector offset mvOffset
[1144] – Refined motion vector refMvLX,
[1145] – The selected reference image sample array refPicLX,
[1146] – Half-sample interpolation filter index hpelIfIdx,
[1147] – Bidirectional optical flow flag bdofFlag
[1148] – The variable cIdx specifies the color component index of the current block.
[1149] The output of this process is:
[1150] – The array predSamplesLX represents the predicted sample values (sbWidth+brdExtSize)x(sbHeight+brdExtSize).
[1151] The predicted block boundary extension size brdExtSize is derived as follows:
[1152] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&sps_affine_prof_enabled_flag))? 2:0(8-752)
[1153] Set the variable fRefWidth to be equal to the PicOutputWidthL of the reference image, measured in brightness samples.
[1154] Set the variable fRefHeight to be equal to the PicOutputHeightL of the reference image, measured in luminance samples.
[1155]
[1156] Set the motion vector mvLX to be equal to (refMvLX-mvOffset).
[1157] – If cIdx equals 0, then the following applies:
[1158] - Define the scaling factor and its fixed-point representation as:
[1159] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)
[1160] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL(8-754)
[1161] – Let (xIntL, yIntL) be the brightness position in units of full samples, and (xFracL, yFracL) be the offset in units of 1 / 16 samples. These variables are used only in this clause to specify the fractional sample positions within the reference sample array refPicLX.
[1162] – The top-left coordinate (xSbInt) of the bounding block used for reference point filling. L ,ySbInt L ) is set to equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).
[1163] – For each luminance sample location (x) within the predicted luminance sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), and the corresponding predicted brightness sample value predSamplesLX[x] is derived as follows. L ][y L ]:
[1164] -Let(refxSb) L ,refySb L ) and (refx L ,refy L ) represents the brightness position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in units of 1 / 16 samples. The variable refxSb is derived as follows: L refx L ,refySb L and refy L :
[1165] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)
[1166] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)
[1167] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757)
[1168] refy L =((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)
[1169] – The variable xInt is derived as follows L yInt L xFrac L and yFrac L :
[1170]
[1171] xFrac L =refx L &15 (8-761)
[1172] yFrac L =refy L &15 (8-762)
[1173] – If bdofFlag is true or (sps_affine_prof_enabled_flag is true and inter_affine_flag[xSb][ySb] is true), and one or more of the following conditions are true, then in (xInt) L +(xFrac L >>3)-1),yInt L +(yFrac L With >>3)-1) and refPicLX as inputs, the predicted luminance sample value predSamplesLX[x] is derived by calling the luminance integer sample retrieval procedure specified in Clause 8.5.6.3.3. L ][y L ].
[1174] -x L It equals 0.
[1175] -x L It equals sbWidth+1.
[1176] –y L It equals 0.
[1177] –y L It equals sbHeight + 1.
[1178] Otherwise, in the case of (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt) L ,ySbInt L With refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs, the predicted luminance sample values predSamplesLX[xL][yL] are derived by calling the luminance sample 8-tap interpolation filtering procedure specified in Clause 8.5.6.3.2.
[1179] Otherwise (cIdx is not equal to 0), then the following applies:
[1180] – Let (xIntC, yIntC) be the chromaticity position given in units of full samples, and (xFracC, yFracC) be the offset given in units of 1 / 32 samples. These variables are used only in this clause to specify the regular fractional sample positions within the reference sample array refPicLX.
[1181] – Set the top left coordinates (xSbIntC, ySbIntC) of the bounding block used for reference sample filling to equal ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[1182] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) within the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[1183] – Let (refxSb) C ,refySb C ) and (refx C ,refy C) represents the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in units of 1 / 32 samples. The variable refxSb is derived as follows: C ,refySb C refx C and refy C :
[1184] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)
[1185] refx C =((Sign(refxSb) C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)
[1186] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765)
[1187] refy C =((Sign(refySb) C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>> 5(8-766)
[1188] – The variable xInt is derived as follows C yInt C xFrac C and yFrac C :
[1189]
[1190] xFrac C =refy C &31 (8-769)
[1191] yFrac C =refy C &31 (8-770)
[1192] – The predicted sample values predSamplesLX[xC][yC] are derived by invoking the procedure specified in Clause 8.5.6.3.4 with (xIntC,yIntC), (xFracC,yFracC), (xSbIntC,ySbIntC), sbWidth, sbHeight, and refPicLX as inputs.
[1193] 5.7 Examples of using encoding / decoding tools
[1194] 5.7.1 BDOF On / Off Control
[1195] – The variable currPic specifies the current image, and the derivation of the variable bdofFlag is as follows:
[1196] – bdofFlag is set to true if all of the following conditions are true.
[1197] –sps_bdof_enabled_flag equals 1, slice_disable_bdof_dmvr_flag equals 0.
[1198] Both predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx] are equal to 1.
[1199] –DiffPicOrderCnt(currPic, RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1], currPic).
[1200] –RefPicList[0][refIdxL0] is a short-term reference image, and RefPicList[1][refIdxL1] is a short-term reference image.
[1201] –MotionModelIdc[xCb][yCb] equals 0.
[1202] –merge_subblock_flag[xCb][yCb] equals 0.
[1203] –sym_mvd_flag[xCb][yCb] equals 0.
[1204] –ciip_flag[xCb][yCb] equals 0.
[1205] –BcwIdx[xCb][yCb] equals 0.
[1206] Both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are equal to 0.
[1207] –cbWidth is greater than or equal to 8.
[1208] –cbHeight is greater than or equal to 8.
[1209] –cbHeight*cbWidth is greater than or equal to 128.
[1210]
[1211] –cIdx equals 0.
[1212] Otherwise, set bdofFlag to false.
[1213] 5.7.2 DMVR On / Off Control
[1214] – dmvrFlag is set to 1 when all of the following conditions are true:
[1215] –sps_dmvr_enabled_flag is equal to 1, slice_disable_bdof_dmvr_flag is equal to 0
[1216] –general_merge_flag[xCb][yCb] equals 1
[1217] Both –predFlagL0[0][0] and predFlagL1[0][0] are equal to 1.
[1218] –mmvd_merge_flag[xCb][yCb] equals 0
[1219] –ciip_flag[xCb][yCb] equals 0
[1220] –DiffPicOrderCnt(currPic, RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1], currPic)
[1221] –RefPicList[0][refIdxL0] is a short-term reference image, and RefPicList[1][refIdxL1] is a short-term reference image.
[1222] –BcwIdx[xCb][yCb] equals 0
[1223] Both `luma_weight_l0_flag[refIdxL0]` and `luma_weight_l1_flag[refIdxL1]` are equal to 0.
[1224] –cbWidth is greater than or equal to 8
[1225] –cbHeight is greater than or equal to 8
[1226] –cbHeight*cbWidth is greater than or equal to 128
[1227]
[1228] 5.7.3 PROF On / Off Control for Reference Image List X
[1229] The derivation of the variable cbProfFlagLX is as follows:
[1230] - cbProfFlagLX is set to false if one or more of the following conditions are true.
[1231] –sps_affine_prof_enabled_flag equals 0.
[1232] –fallbackModeTriggered equals 1.
[1233] –numCpMv equals 2, cpMvLX[1][0] equals cpMvLX[0][0], cpMvLX[1][1] equals cpMvLX[0][1].
[1234] –numCpMv equals 3, cpMvLX[1][0] equals cpMvLX[0][0], cpMvLX[1][1] equals cpMvLX[0][1], cpMvLX[2][0] equals cpMvLX[0][0], cpMvLX[2][1] equals cpMvLX[0][1].
[1235]
[1236] –[[The pic_width_in_luma_samples of the reference image refPicLX associated with refIdxLX are not equal to the pic_width_in_luma_samples of the current image.
[1237] – The pic_height_in_luma_samples of the reference image refPicLX associated with refIdxLX are not equal to the pic_height_in_luma_samples of the current image.
[1238] Otherwise, set cbProfFlagLX to true.
[1239] 5.7.4 PROF On / Off Control for Reference Image List X (Second Embodiment)
[1240] The derivation of the variable cbProfFlagLX is as follows:
[1241] - cbProfFlagLX is set to false if one or more of the following conditions are true.
[1242] –sps_affine_prof_enabled_flag equals 0.
[1243] –fallbackModeTriggered equals 1.
[1244] –numCpMv equals 2, cpMvLX[1][0] equals cpMvLX[0][0], cpMvLX[1][1] equals cpMvLX[0][1].
[1245] –numCpMv equals 3, cpMvLX[1][0] equals cpMvLX[0][0], cpMvLX[1][1] equals cpMvLX[0][1], cpMvLX[2][0] equals cpMvLX[0][0], cpMvLX[2][1] equals cpMvLX[0][1].
[1246]
[1247] –[[The pic_width_in_luma_samples of the reference image refPicLX associated with refIdxLX are not equal to the pic_width_in_luma_samples of the current image.
[1248] – The pic_height_in_luma_samples of the reference image refPicLX associated with refIdxLX are not equal to the pic_height_in_luma_samples of the current image.
[1249] Otherwise, set cbProfFlagLX to true.
[1250] 5.8 Implementation of Conditional Signaling Notification of Inter-Frame Related Syntax Elements in Picture Header RBSP Syntax
[1251]
[1252]
[1253]
[1254]
[1255]
[1256] 7.4.3.6 Image Header RBSP Semantics
[1257]
[1258] 7.3.7 Header Syntax
[1259] 7.3.7.1 General Strip Header Syntax
[1260]
[1261]
[1262] 5.9 Examples of constrained RPR
[1263] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference image that references the current image of this PPS, respectively. Let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference image, respectively. Bitstream consistency requires that all of the following conditions be met:
[1264] –PicOutputWidthL*2 should be greater than or equal to refPicOutputWidthL.
[1265] –PicOutputHeightL*2 should be greater than or equal to refPicOutputHeightL.
[1266] –PicOutputWidthL should be less than or equal to refPicOutputWidthL*8.
[1267] –PicOutputHeightL should be less than or equal to refPicOutputHeightL*8.
[1268] –(PicOutputWidthL–refPicOutputWidthL)*(PicWidthInLumaSamples–refPicWidthInLumaSamples) should be greater than or equal to 0.
[1269] –(PicOutputHeightL–refPicOutputHeightL)*(PicHeightInLumaSamples–refPicHeightInLumaSamples) should be greater than or equal to 0.
[1270] –135*pic_width_max_in_luma_samples*PicOutputWidthL–(128*refPicOutputWidthL+7*PicOutputWidthL)*PicWidthInLumaSamples should be greater than or equal to 0.
[1271] –135*pic_height_max_in_luma_samples*PicOutputHeightL–(128*refPicOutputHeightL+7*PicOutputHeightL)*PicHeightInLumaSamples should be greater than or equal to 0.
[1272] 5.10 Example of signaling wraparound offset
[1273] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[1274]
[1275] 7.3.2.3 Image Parameter Set RBSP Syntax
[1276]
[1277] A value of 1 for `pps_ref_wraparound_enabled_flag` specifies that horizontal wraparound motion compensation is applied in inter-frame prediction. A value of 0 for `pps_ref_wraparound_enabled_flag` specifies that horizontal wraparound motion compensation is not applied. When the value of `(CtbSizeY / MinCbSizeY+1)` is less than or equal to `(pic_width_in_luma_samples / MinCbSizeY-1)`, the value of `pps_ref_wraparound_enabled_flag` should be 0.
[1278] The increment of 1 in pps_ref_wraparound_offset_minus1 specifies the offset used to calculate the horizontal wraparound position in units of MinCbSizeY luminance samples. The value of pps_ref_wraparound_offset_minus1 should be in the range of (CtbSizeY / MinCbSizeY)+1 to (pic_width_in_luma_samples / MinCbSizeY)-1, inclusive.
[1279] 7.4.4.2 General Constraint Information Semantics
[1280] The value of no_ref_wraparound_constraint_flag equal to 1 specifies that [[sps_ref_wraparound_enabled_flag]] It should be equal to 0. A value of 0 for no_ref_wraparound_constraint_flag specifies that such a constraint should not be imposed.
[1281] 8.5.3.2.2 Bilinear Interpolation Process for Luminance Samples
[1282] The inputs to this process include:
[1283] – Brightness position in units of the entire sample (xInt) L yInt L ),
[1284] – Brightness position in fractional samples (xFrac) L ,yFrac L ),
[1285] –Luminance reference sample array refPicLX L .
[1286] …
[1287] For i = 0..1, the brightness position (xInt) in units of the entire sample point.i ,yInt i The derivation is as follows:
[1288] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1289] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i)(642)
[1290] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i)(643)
[1291] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[1292] xInt i =Clip3(0,picW-1,[[sps_ref_wraparound_enabled_flag]]
[1293] ClipH(([[pps_ref_wraparound_offset_minus1]] +1)*MinCbSizeY,picW,(xInt L +i)): (644)
[1295] xInt L +i)
[1296] yInt i =Clip3(0,picH-1,yInt) L +i) (645)
[1297] …
[1298] 8.5.6.3.2 Brightness Sample Interpolation and Filtering Process
[1299] …
[1300] For i = 0..7, the brightness position (xInt) in units of the entire sample point. i ,yInt i The derivation is as follows:
[1301] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1302] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3) (955)
[1303] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (956)
[1305] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[1306] xInt i =Clip3(0,picW-1,[[sps_ref_wraparound_enabled_flag]]
[1307] ClipH(([[pps_ref_wraparound_offset_minus1]] +1)*MinCbSizeY,picW,xInt L +i-3):(957)
[1308] xInt L +i-3)
[1309] yInt i =Clip3(0,picH-1,yInt) L +i-3) (958)
[1310] …
[1311] 8.5.6.3.3 Luminance Integer Sample Extraction Process
[1312] …
[1313] The brightness positions (xInt, yInt) per unit of the entire sample are derived as follows:
[1314] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1315] xInt=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L ) (966)
[1317] yInt=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L ) (967)
[1319] – Otherwise, the following applies:
[1320] xInt=Clip3(0,picW-1,[[sps_ref_wraparound_enabled_flag]] (968)
[1321] ClipH(([[pps_ref_wraparound_offset_minus1]] +1)*MinCbSizeY,picW,xInt L ):xInt L )
[1322] yInt = Clip3(0, picH-1, yInt) L (969)
[1323] …
[1324] 8.5.6.3.4 Chromaticity Sample Interpolation Process
[1325] …
[1326] The variable xOffset is set to equal to ([[pps_ref_wraparound_offset_minus1]]). +1)*MinCbSizeY) / SubWidthC.
[1327] For i = 0..3, the chromaticity position (xInt) in units of the entire sample point. i ,yInt i The derivation is as follows:
[1328] – If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies:
[1329] xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (971)
[1330] yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (972)
[1331] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies:
[1332] xInt i =Clip3(0,picW C -1,[[sps_ref_wraparound_enabled_flag]] ClipH(xOffset, picW) C ,xInt C +i-1): (973)
[1334] xInt C +i-1)
[1335] yInt i =Clip3(0,picH C -1,yInt C +i-1)
[1336] 5.11 Example of Deblocking Filtering Between Sub-Images
[1337] 8.8.3 Deblocking Filtering Process
[1338] 8.8.3.1 Overview
[1339] The input to this process is the reconstructed image before deblocking, i.e., the array recPicture. L When ChromaArrayType is not equal to 0, the array recPicture Cb and recPictureCr .
[1340] The output of this process is the reconstructed image after removing the blocks, i.e., the array recPicture. L When ChromaArrayType is not equal to 0, the array recPicture Cb and recPicture Cr .
[1341] …
[1342] The deblocking filtering process is applied to all encoded and transformed block edges of the image, except for the following types of edges:
[1343] –The edge at the boundary of the image,
[1344] – [[edges that coincide with the boundary of the subpic, where loop_filter_across_subpic_enabled_flag[SubPicIdx] equals 0,]]
[1345] – When VirtualBoundariesDisabledFlag equals 1, the edge that coincides with the virtual boundary of the image.
[1346] -…
[1347] 8.8.3.2 Deblocking filtering process for one direction
[1348] The inputs to this process include:
[1349] – The variable `treeType` specifies whether the current processing is luminance (DUAL_TREE_LUMA) or chrominance (DUAL_TREE_CHROMA).
[1350] …
[1351] 1. The derivation of the variable filterEdgeFlag is as follows:
[1352] – If edgeType equals EDGE_VER, and one or more of the following conditions are true, then filterEdgeFlag is set to 0:
[1353] – The left boundary of the current encoding / decoding block is the left boundary of the image.
[1354] – [[The left boundary of the current codec block is either the left or right boundary of the subpic, and loop_filter_across_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1355] -…
[1356] Otherwise, if edgeType equals EDGE_HOR, and one or more of the following conditions are true, then the variable filterEdgeFlag is set to 0:
[1357] - The upper boundary of the current luminance codec block is the upper boundary of the image.
[1358] – [[The upper boundary of the current codec block is either the upper or lower boundary of the subpic, and loop_filter_across_subpic_enabled_flag[SubPicIdx] equals 0.]]
[1359] -…
[1360] 8.8.3.6.6 Filtering process of using a short filter on luminance samples
[1361] When nDp is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample p0 is equal to 1, nDp is set to 0.
[1362] When nDq is greater than 0 and the pred_mode_plt_flag of the codec unit containing the codec block with sample q0 is equal to 1, nDq is set to 0.
[1363]
[1364] 8.8.3.6.7 Filtering process of using a long filter for luminance samples
[1365] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where i = 0..maxFilterLengthP-1.
[1366] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The corresponding input sample value q j Instead, where j = 0..maxFilterLengthQ-1.
[1367]
[1368] 8.8.3.6.9 Filtering process for chromaticity samples
[1369] When including sample point p i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value p i 'The corresponding input sample value p i Instead, where 0..maxFilterLengthP-1.
[1370] When including sample point q i When the pred_mode_plt_flag of the codec unit of the codec block is equal to 1, the filtered sample value q i 'The defined input sample value q i Instead, where ii = 0..maxFilterLengthQ-1:
[1371]
[1372] 5.12 Examples of Striped Signaling
[1373] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[1374]
[1375]
[1376]
[1377] A value of 1 for `sps_weighted_bipred_flag` indicates that explicit weighted predictions can be applied to the B-bands of the reference SPS. A value of 0 for `sps_weighted_bipred_flag` indicates that explicit weighted predictions are not applied to the B-bands of the reference SPS.
[1378] A value of 0 for sps_bdof_enabled_flag indicates that bidirectional optical flow inter-frame prediction is disabled. A value of 1 for sps_bdof_enabled_flag indicates that bidirectional optical flow inter-frame prediction is enabled.
[1379] A value of 1 for `sps_smvd_enabled_flag` indicates that symmetric motion vector difference can be used in motion vector decoding. A value of 0 for `sps_smvd_enabled_flag` indicates that symmetric motion vector difference is not used in motion vector encoding and decoding.
[1380] A value of 1 for sps_dmvr_enabled_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement is enabled. A value of 0 for sps_dmvr_enabled_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement is disabled.
[1381] `sps_bcw_enabled_flag` specifies whether bidirectional prediction with CU weights can be used for inter-frame prediction. If `sps_bcw_enabled_flag` equals 0, the syntax should be restricted so that bidirectional prediction with CU weights is not used in CLVS, and `bcw_idx` does not exist in the CLVS codec unit syntax. Otherwise (`sps_bcw_enabled_flag` equals 1), bidirectional prediction with CU weights can be used in CLVS.
[1382] `sps_ciip_enabled_flag` specifies that the `ciip_flag` can exist in the codec unit syntax used for inter-frame codec units. `sps_ciip_enabled_flag` equal to 0 specifies that the `ciip_flag` does not exist in the codec unit syntax used for inter-frame codec units.
[1383] …
[1384] 7.3.2.7 Image Header Structure Syntax
[1385]
[1386]
[1387]
[1388] A ph_inter_slice_allowed_flag value of 0 indicates that all codec slices of the image have a slice_type of 2. A ph_inter_slice_allowed_flag value of 1 indicates that the image may or may not contain one or more codec slices with a slice_type of 0 or 1.
[1389] `ph_intra_slice_allowed_flag` equal to 0 specifies that the slice_type of all codec slices in the image is either 0 or 1. `ph_intra_slice_allowed_flag` equal to 1 specifies that one or more codec slices with a slice_type of 2 may or may not exist in the image. If none exist, the value of `ph_intra_slice_allowed_flag` is inferred to be equal to 1. ...
[1391] A value of 1 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 0. A value of 0 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 1.
[1392] ph_collocated_ref_idx specifies the reference index of the juxtaposed images used for temporal motion vector prediction. When ph_collocated_from_l0_flag equals 1, ph_collocated_ref_idx represents the entry in reference image list 0, and the value of ph_collocated_ref_idx should be in the range of 0 to num_ref_entries[0][RplsIdx[0]]–1, inclusive.
[1393] When ph_collocated_from_l0_flag equals 0, ph_collocated_ref_idx represents the entry in reference image list 1, and the value of ph_collocated_ref_idx should be in the range of 0 to num_ref_entries[1][RplsIdx[1]]–1, inclusive.
[1394] If it does not exist, then it is inferred that the value of ph_collocated_ref_idx is equal to 0. ...
[1396] A `mvd_l1_zero_flag` equal to 1 indicates that the `mvd_coding(x0,y0,1)` syntax structure is not parsed, and for `compIdx = 0..1` and `cpIdx = 0..2`, `MvdL1[x0][y0][compIdx]` and `MvdCpL1[x0][y0][cpIdx][compIdx]` are set to 0. A `mvd_l1_zero_flag` equal to 0 indicates that the `mvd_coding(x0,y0,1)` syntax structure is parsed.
[1397] …
[1398] A value of 1 for ph_disable_bdof_flag indicates that bidirectional inter-frame prediction based on bidirectional optical flow inter-frame prediction is disabled in the stripe associated with PH. A value of 0 for ph_disable_bdof_flag indicates that bidirectional inter-frame prediction based on bidirectional optical flow inter-frame prediction can be enabled or disabled in the stripe associated with PH.
[1399] When ph_disable_bdof_flag does not exist, the following applies:
[1400] - If sps_bdof_enabled_flag equals 1 Therefore, it can be inferred that the value of ph_disable_bdof_flag is equal to 0.
[1401] - Otherwise (sps_bdof_enabled_flag equals 0) This will infer that the value of ph_disable_bdof_flag is equal to 1.
[1402] A value of 1 for ph_disable_dmvr_flag indicates that bidirectional inter-frame prediction based on decoder motion vector refinement is disabled in the slice associated with the PH. A value of 0 for ph_disable_dmvr_flag indicates that bidirectional inter-frame prediction based on decoder motion vector refinement can be enabled or disabled in the slice associated with the PH.
[1403] When ph_disable_dmvr_flag does not exist, the following applies:
[1404] – If sps_dmvr_enabled_flag equals 1 Therefore, it can be inferred that the value of ph_disable_dmvr_flag is equal to 0.
[1405] - Otherwise (sps_dmvr_enabled_flag equals 0) Therefore, it can be inferred that the value of ph_disable_dmvr_flag is equal to 1.
[1406] 7.3.7.1 General Strip Header Syntax
[1407]
[1408] slice_type specifies the encoding / decoding type of the slice according to Table 9.
[1409] Table 9 – Relationship between name and slice_type
[1410] Names of slice_type B (B slice) 0 P (P slice) 1 I (I slice) 2 slice_type
[1411] When it does not exist, the value of slice_type is inferred to be equal to [[2]].
[1412] When ph_intra_slice_allowed_flag equals 0, the value of slice_type should be 0 or 1. When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (inclusive), and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, slice_type should be 2.
[1413] 5.13 Another embodiment of striped signaling
[1414] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[1415]
[1416]
[1417] ...
[1419] The semantic changes for pps_weighted_bipred_flag are as follows:
[1420] A `pps_weighted_bipred_flag` value of 0 specifies that explicit weighted predictions should not be applied to B-strips referencing the PPS. A `pps_weighted_bipred_flag` value of 1 specifies that explicit weighted predictions should be applied to B-strips referencing the PPS. When the value is 0, the value of pps_weighted_bipred_flag should be 0.
[1421] 7.3.2.7 Image Header Structure Syntax
[1422]
[1423]
[1424]
[1425] …
[1426] `ph_inter_slice_allowed_flag` equal to 0 indicates that all codec slices of the image have a slice_type of 2. `ph_inter_slice_allowed_flag` equal to 1 indicates that one or more codec slices with a slice_type of 0 or 1 may or may not exist in the image.
[1427] `ph_intra_slice_allowed_flag` equal to 0 specifies that the slice_type of all codec slices in the image is either 0 or 1. `ph_intra_slice_allowed_flag` equal to 1 specifies that one or more codec slices with a slice_type of 2 may or may not exist in the image. If none exist, the value of `ph_intra_slice_allowed_flag` is inferred to be 1.
[1428]
[1429] …
[1430] `ph_intra_slice_allowed_flag` equal to 0 specifies that the slice_type of all codec slices in the image is equal to 0 or 1. `ph_intra_slice_allowed_flag` equal to 1 specifies that one or more codec slices with slice_type equal to 2 may or may not exist in the image. If they do not exist, it is inferred that the value of `ph_intra_slice_allowed_flag` is equal to [[1]].
[1431] …
[1432] 7.3.7.1 General Strip Header Syntax
[1433]
[1434]
[1435] slice_type specifies the encoding / decoding type of the slice according to Table 9.
[1436] Table 9 – Relationship between name and slice_type
[1437] Names of slice_type B (B slice) 0 P (P slice) 1 I (I slice) 2 Figure 5
[1438] When it does not exist, the value of slice_type is inferred to be equal to [[2]].
[1439] 5.14 Example of Direct Signaling Notification of the Starting Point of the Chromaticity QP Mapping Table
[1440]
[1441] `qp_table_start[[_minus26]][i][[+26]]` specifies the starting luma and chromaticity QP used to describe the i-th chromaticity QP mapping table. The value of `qp_table_start[[_minus26]][i]` should be between `[[-26]]` and `QpBdOffset` and `[
[36] ]`. Within the range, including end values. When qp_table_start[[_minus26]][i] does not exist in the bitstream, it is inferred that the value of qp_table_start[[_minus26]][i] is equal to 0.
[1442] 6. Example Implementations of the Technology of this Disclosure
[1443] Figure 6This is a block diagram of a video processing apparatus 500. Apparatus 500 can be used to implement one or more methods described herein. Apparatus 500 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 500 may include one or more processors 502, one or more memories 504, and video processing hardware 506. The processors(multiple) 502 can be configured to implement one or more methods described in this document. The memories(multiple) 504 can be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 506 can be used to implement some of the techniques described in this document in hardware circuitry and can be partly or entirely part of the processor 502 (e.g., a graphics processing unit (GPU) core or other signal processing circuitry).
[1444] In this document, the term "video processing" or encoding / decoding can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits located at the same position or distributed at different positions within the bitstream. For example, a macroblock can be encoded based on the error residuals from the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream.
[1445] It is understood that the disclosed methods and techniques will be beneficial to video encoder and / or decoder embodiments incorporated in video processing devices such as smartphones, laptops, desktops and similar devices by allowing the use of the techniques disclosed in this document.
[1446] Figure 9 This is a flowchart of an example method 600 for video processing. Method 600 includes: at 610, performing a conversion between the current video block and its codec representation, wherein, during the conversion, if the resolution and / or size of a reference image differs from the resolution and / or size of the current video block, the same interpolation filter is applied to a group of adjacent or non-adjacent samples predicted using the current video block.
[1447] Some embodiments can be described using the following clause-based format.
[1448] 1. A video processing method, comprising:
[1449] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, if the resolution and / or size of the reference image is different from the resolution and / or size of the current video block, the same interpolation filter is applied to the adjacent or non-adjacent sample groups predicted using the current video block.
[1450] 2. According to the method of Clause 1, wherein the same interpolation filter is a vertical interpolation filter.
[1451] 3. According to the method of Clause 1, wherein the same interpolation filter is a horizontal interpolation filter.
[1452] 4. According to the method of Clause 1, wherein the adjacent or non-adjacent sample group includes all samples located in the region of the current video block.
[1453] 5. The method according to Clause 4, wherein the current video block is divided into multiple rectangles, each with a size of MXN.
[1454] 6. The method according to Clause 5, wherein M and / or N are predetermined.
[1455] 7. The method according to Clause 5, wherein M and / or N are derived based on the dimensions of the current video block.
[1456] 8. The method according to Clause 5, wherein signaling notifications M and / or N are given in the codec representation of the current video block.
[1457] 9. The method according to Clause 1, wherein the sample group shares the same motion vector.
[1458] 10. The method according to Clause 9, wherein the sample group shares the same horizontal component and / or the same fractional part of the horizontal component.
[1459] 11. The method according to Clause 9, wherein the sample group shares the same vertical component and / or the same fractional portion of the vertical component.
[1460] 12. The method according to one or more of the clauses 9-11, wherein the same motion vector or its components satisfy at least one or more of the following rules: the resolution of the reference picture, the size of the reference picture, the resolution of the current video block, the size of the current video block, or the precision value.
[1461] 13. The method according to one or more of the clauses 9-11, wherein the same motion vector or its components correspond to the motion information of a sample point located in the current video block.
[1462] 14. The method according to one or more of the clauses 9-11, wherein the same motion vector or its components are set as motion information of virtual samples located within or outside the group.
[1463] 15. A video processing method, comprising:
[1464] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, if the resolution and / or size of the reference image differs from the resolution and / or size of the current video block, only blocks predicted from the current video block are allowed to use integer-valued motion information related to the current block.
[1465] 16. The method according to Clause 15, wherein integer motion information is derived by rounding the original motion information of the current video block.
[1466] 17. The method according to Clause 15, wherein the original motion information of the current video block is in the horizontal and / or vertical direction.
[1467] 18. A video processing method, comprising:
[1468] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, if the resolution and / or size of the reference image is different from the resolution and / or size of the current video block, an interpolation filter is applied to derive a block predicted using the current video block, and wherein the interpolation filter is selected based on rules.
[1469] 19. The method according to Clause 18, wherein the rule relates to the resolution and / or size of a reference image relative to the resolution and / or size of the current video block.
[1470] 20. The method according to Clause 18, wherein the interpolation filter is a vertical interpolation filter.
[1471] 21. The method according to Clause 18, wherein the interpolation filter is a horizontal interpolation filter.
[1472] 22. The method according to Clause 18, wherein the interpolation filter is one of the following: a 1-tap filter, a bilinear filter, a 4-tap filter, or a 6-tap filter.
[1473] 23. The method according to Clause 22, wherein an interpolation filter is used as part of another step in the transformation.
[1474] 24. The method according to Clause 18, wherein the interpolation filter includes the use of filling in sample points.
[1475] 25. The method according to Clause 18, wherein the use of the interpolation filter depends on the color components of the samples of the current video block.
[1476] 26. A video processing method, comprising:
[1477] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, if the resolution and / or size of the reference image differs from the resolution and / or size of the current video block, a deblocking filter is selectively applied, wherein the strength of the deblocking filter is set according to rules relating to the resolution and / or size of the reference image relative to the resolution and / or size of the current video block.
[1478] 27. The method according to Clause 27, wherein the intensity of the deblocking filter changes between one video block and another.
[1479] 28. A video processing method, comprising:
[1480] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, if a sub-picture of the current video block exists, the consistent bitstream satisfies rules relating to the resolution and / or size of a reference picture relative to the resolution and / or size of the current video block.
[1481] 29. The method pursuant to Clause 28 also includes:
[1482] Divide the current video block into one or more sub-images, where the division depends at least on the resolution of the current video block.
[1483] 30. A video processing method, comprising:
[1484] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, the reference image of the current video block is resampled according to a rule based on the dimension of the current video block.
[1485] 31. A video processing method, comprising:
[1486] Perform a conversion between the current video block and its codec representation, wherein, during the conversion, the use of codec tools for the current video block is selectively enabled or disabled depending on the resolution / size of the reference image of the current video block relative to its resolution / size.
[1487] 32. The method according to one or more of the foregoing clauses, wherein the sample group is located in the consistency window.
[1488] 33. The method according to Clause 32, wherein the shape of the consistency window is rectangular.
[1489] 34. The method according to any one or more of the foregoing clauses, wherein the resolution is the resolution of the video block being encoded / decoded or the resolution of the consistency window in the video block being encoded / decoded.
[1490] 35. The method according to any one or more of the foregoing clauses, wherein the size is the size of the video block being encoded / decoded or the size of a consistency window within the video block being encoded / decoded.
[1491] 36. The method according to any one or more of the foregoing clauses, wherein the dimension belongs to the dimension of the video block being encoded / decoded or the dimension of the consistency window in the video block being encoded / decoded.
[1492] 37. The method according to Clause 32, wherein the consistency window is defined by the set of consistency trimming window parameters.
[1493] 38. The method according to Clause 37, wherein at least a portion of the set of conformance pruning window parameters is implicitly or explicitly signaled in the codec representation.
[1494] 39. The method according to any one or more of the foregoing clauses, wherein the signaling notification consistency trimming window parameter set is not permitted in the codec representation.
[1495] 40. The method of any one or more of the foregoing clauses, wherein the position of the reference sample is derived relative to the top-left sample of the current video block in the consistency window.
[1496] 41. A video processing method, comprising:
[1497] Perform a conversion between multiple video blocks and the codec representations of multiple video blocks, wherein, during the conversion, a first consistency window for a first video block and a second consistency window for a second video block are defined, and wherein the ratio of the width and / or height of the first consistency window to the second consistency window is based on a rule at least based on the consistency bitstream.
[1498] 42. A video decoding apparatus, including a processor configured to implement one or more of the methods of clauses 1-41.
[1499] 43. A video encoding apparatus, including a processor configured to implement one or more of the methods of clauses 1-41.
[1500] 44. A computer program product having stored computer code thereon, which, when executed by a processor, causes the processor to implement the method of any one of clauses 1-41.
[1501] 45. A method, apparatus or system described in this document.
[1502] Figure 10This is a block diagram illustrating an example video processing system 900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 900. System 900 may include an input 902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (e.g., Wi-Fi or cellular interfaces).
[1503] System 900 may include codec component 904, which can implement the various codec or encoding methods described in this document. Codec component 904 can reduce the average bit rate of the video from input 902 to the output of codec component 904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 906, the output of codec component 904 can be stored or transmitted via connected communication. The stored or transmitted bitstream (or codec) representation of the video received at input 902 can be used by the component. 908 is used to generate pixel values or displayable video, which are sent to display interface 910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations for reversing the codec results will be performed by the decoder.
[1504] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The techniques described herein can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[1505] Figure 10 This is a block diagram illustrating an example video encoding / decoding system 100 that can utilize the technology disclosed herein.
[1506] like Figure 11 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[1507] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[1508] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems used to generate video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.
[1509] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[1510] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to the user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interact with an external display device.
[1511] The video encoder 114 and the video decoder 124 can operate according to video compression standards (e.g., High Efficiency Video Coding (HEVC) standard, Universal Video Coding (VVC) standard, and other current and / or future standards).
[1512] Figure 10 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 9 The video encoder 114 in the system 100 shown in the figure.
[1513] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[1514] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[1515] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[1516] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for ease of illustration, in Figure 12 The examples are shown respectively.
[1517] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[1518] The mode selection unit 203 can select one of the intra-frame or inter-frame coding / decoding modes, for example, based on the error result, and provide the obtained intra-frame or inter-frame coding / decoding block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[1519] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[1520] For example, depending on whether the current video block is in an I-band, P-band, or B-band, the motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block.
[1521] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in reference images in list 0 or list 1. Then, motion estimation unit 204 can generate a reference index indicating the reference images in list 0 or list 1, which contains the reference video blocks and a motion vector indicating the spatial displacement between the current video block and the reference video blocks. Motion estimation unit 204 can output this reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video blocks indicated by the motion information of the current video block.
[1522] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Then, motion estimation unit 204 can generate reference indices indicating the reference images in lists 0 and 1, which contain reference video blocks and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[1523] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.
[1524] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block relative to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[1525] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[1526] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[1527] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Combined Mode Signaling.
[1528] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on the decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[1529] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the current video block.
[1530] In other examples, for the current video block, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[1531] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[1532] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[1533] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in the buffer 213.
[1534] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[1535] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bit stream including the entropy-encoded data.
[1536] Figure 10 This is a block diagram illustrating an example of a video decoder 300. The video decoder 300 can be... Figure 10 The video decoder 114 in the system 100 shown in the figure.
[1537] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 10 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[1538] exist Figure 9 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations relative to the video encoder 200 (e.g., Figure 13 The encoding process described is largely the opposite of the decoding process.
[1539] The entropy decoding unit 301 can acquire the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list indexes, and other motion information, from the entropy-coded video data. For example, the motion compensation unit 302 can determine such information by performing AMVP and merging modes.
[1540] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter used at sub-pixel precision can be included in the syntax element.
[1541] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[1542] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information used to decode the encoded video sequence.
[1543] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[1544] The reconstruction unit 306 can add the residual block to the corresponding predicted block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation.
[1545] Figure 14 This is a flow representation of a video processing method according to the present disclosure. Method 1300 includes, in operation 1310, performing a conversion between blocks of video in a video unit and a video bitstream. The bitstream conforms to a rule specifying the maximum dimension of the video unit. The dimension of the video unit includes the width, height, or number of samples in the video unit, and the maximum dimension is known prior to the conversion.
[1546] In some embodiments, a video unit includes a slice, wherein the maximum dimension includes the maximum width of the slice, and a rule specifies that the maximum width of the slice is equal to the maximum number of samples in the codec tree block counted in terms of the luminance dimension, and the maximum luminance slice dimension is the maximum luminance slice width.
[1547] In some embodiments, a video unit includes a strip, a sub-picture, or a slice. In some embodiments, the dimension of a video unit is determined based on the number of codec tree blocks in a strip. In some embodiments, the maximum dimension of a video unit is determined based on the maximum number of samples in the codec tree block, counted in terms of luminance.
[1548] In some embodiments, the maximum luminance dimension is indicated in the bitstream. In some embodiments, the bitstream includes syntax elements that indicate the maximum dimension of the video unit allowed in the current sequence, current picture, current strip, or current subpicture. In some embodiments, the maximum luminance dimension is indicated in the bitstream for rectangular strips or rectangular subpictures. In some embodiments, the maximum luminance dimension is indicated in the bitstream for raster scan strips. In some embodiments, the width of the maximum luminance dimension is a fixed value. In some embodiments, the maximum luminance width or height is 1920 or 4096. In some embodiments, the maximum luminance area is 2073600 or 83388608.
[1549] In some embodiments, the rule further specifies the maximum number of sub-units in a video unit. In some embodiments, the maximum number of sub-units is signaled in the bitstream. In some embodiments, the maximum number of sub-units varies depending on the characteristics of the video, including the video profile, level, and hierarchy.
[1550] In some embodiments, the maximum dimension of a video unit varies based on video characteristics. In some embodiments, characteristics include video profiles, levels, or hierarchies.
[1551] Figure 15 This is a flow representation of a video processing method according to the present disclosure. Method 1400 includes, in operation 1410, performing a conversion between a current image of the video and a bitstream of the video. The current image is associated with one or more reference images such that (1) the dimension PW×PH of the current image, (2) the scaling window dimension SW×WH of the current image, (3) the scaling window dimension SW'×SH' of the reference images of the current image, and (4) the maximum permissible dimensions Wmax and Hmax of the image satisfy constraints.
[1552] In some embodiments, the constraint specifies a×Wmax×SW-(b×SW′+c×SW)×PW+offw≥0, where a, b, c, and offw are integers. In some embodiments, the constraint specifies d×Hmax×SH-(e×SH′+f×SH)×PH+offh≥0, where d, e, f, and offw are integers. In some embodiments, the constraint specifies a×Wmax×SW-b×SW′×PW+offw≥0, where a and b are integers. In some embodiments, the constraint specifies d×Hmax×SH-e×SH′×PH+offh≥0, where d and e are integers. For example, d=1, e=1, offh=0.
[1553] In some embodiments, as well as Where the constraint specifies Where Lw, Bw, and offw are integers. In some embodiments, as well as Where the constraint specifies Where Lh, Bh, and offh are integers. In some embodiments, and The constraint specifies that rw ≤ a*Rw + offw, where a and offw are integers. In some embodiments, and The constraint specifies h ≤ b*Rh + offh, where b and offh are integers. In some embodiments, and The constraint specifies that rw ≤ a*Qw + offw, where a and offw are integers. In some embodiments, and The constraint specifies that rh ≤ b * Qh + offh, where b and offh are integers.
[1554] Figure 16 This is a flow representation of a video processing method according to the present disclosure. Method 1500 includes, in operation 1510, performing a conversion between video frames and a video bitstream. The bitstream conforms to rules specifying that quantization parameter information is signaled in the frame header and excluded from the frame parameter set. In some embodiments, the quantization parameter information is signaled for a specific codec tool, which includes Adaptive Color Transform (ACT).
[1555] Figure 17 This is a flow representation of a video processing method according to the present disclosure. Method 1600 includes, in operation 1610, performing a conversion between video and video bitstreams. The bitstreams conform to rules specifying the stripe type in video units associated with multiple stripes, the video units including picture parameter sets or picture headers, and the stripe type including I stripe, P stripe, or B stripe.
[1556] In some embodiments, the bitstream includes a first syntax element indicating whether all plurality of slices associated with a video unit have the same slice type. In some embodiments, the bitstream includes a second syntax element indicating the same slice type shared by all plurality of slices.
[1557] In some embodiments, a second syntax element is conditionally signaled in the bitstream based on a first syntax element. In some embodiments, the second syntax element is signaled in the bitstream only if the first syntax element indicates that all plurality of stripes associated with a video unit have the same stripe type.
[1558] In some embodiments, the second syntax element is binarized or interpreted in the same manner as the third syntax element indicating the stripe type in the stripe header. In some embodiments, the third syntax element is conditionally signaled in the bitstream based at least on the first or second syntax element. In some embodiments, the third syntax element is signaled in the bitstream only when the first syntax element indicates that multiple stripes associated with a video unit have different stripe types. In some embodiments, if the third syntax element is omitted in the bitstream, a default value is inferred from the first and / or second syntax elements.
[1559] In some embodiments, the bitstream includes one or more syntax elements in a sequence parameter set, a picture parameter set, or a picture header that indicate whether a slice type is allowed for a video unit. In some embodiments, semantic constraints are added to the picture parameter set of the bitstream if a B-slice slice type is not allowed in a video unit. In some embodiments, the bitstream conditionally includes one or more syntax elements that constrain the slice type of the video unit based at least on a first syntax element or a second syntax element. In some embodiments, one or more syntax elements are signaled in the bitstream only if the first syntax element indicates that multiple slices associated with a video unit have different slice types. In some embodiments, if one or more syntax elements are omitted in the bitstream, one or more syntax elements are inferred to have default values based at least on the first or second syntax element. In some embodiments, one or more syntax elements include syntax elements indicating whether all multiple slices have an I-slice slice type. In some embodiments, one or more syntax elements include syntax elements indicating whether all multiple slices have a B-slice or P-slice slice type. In some embodiments, one or more syntax elements include syntax elements indicating whether all multiple slices have a P-slice or I-slice slice type. In some embodiments, whether to include signaling notification of one or more inter-frame association syntax elements in the bitstream's picture header is based on the slice type or whether all multiple slices have the same slice type. In some embodiments, one or more inter-frame association syntax elements include at least one of the following: ph_log2_diff_min_qt_min_cb_inter_slice,ph_max_mtt_hierarchy_depth_inter_slice,ph_log2_diff_max_bt_min_qt_inter_slice,ph_log2_diff_max_tt_min_qt_inter_slice,ph_cu_qp_delta_subdiv_inter_slice,ph_cu_chrom a_qp_offset_subdiv_inter_slice,ph_temporal_mvp_enabled_flag,ph_collocated_from_l0_flag,ph_collocated_ref_idx,mvd_l1_z ero_flag, ph_fpel_mmvd_enabled_flag, ph_disable_bdof_flag, ph_disable_dmvr_flag, ph_disable_prof_flag, or pred_weight_table.In some embodiments, whether the strip header includes a syntax element for indicating the strip type depends on whether all the multiple stripes have the same strip type.
[1560] This is a flow representation of a video processing method according to the present disclosure. Method 1700 includes, in operation 1710, performing a conversion between video and video bitstreams. The bitstreams include values indicating the starting point of a colorimetric parameter mapping table.
[1561] In some embodiments, the value is indicated without differential encoding / decoding. In some embodiments, the value is used for signaling notification as an unsigned value instead of a signed value.
[1562] In some embodiments, the conversion includes encoding video into a bitstream. The method also includes storing the bitstream in a computer-readable medium. In some embodiments, the conversion includes decoding the bitstream into video.
[1563] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of the video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to a video block will be performed using the video processing tool or mode enabled based on a decision or determination.
[1564] Some embodiments of the disclosed technology include making a decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in converting video blocks into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that no modifications have been made to the bitstream using a video processing tool or mode that was disabled based on the decision or mode.
[1565] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagated signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[1566] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[1567] The processes and logic flows described in this document can be executed by one or more programmable processors executing one or more computer programs, thereby performing functions by manipulating input data and generating outputs. These processes and logic flows can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[1568] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[1569] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of specific features of particular embodiments of a particular technology. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features from the claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.
[1570] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner to obtain the desired result, or as requiring all illustrated operations to be performed. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[1571] Only a few implementation methods and examples have been described. Other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Perform the conversion between video blocks in the first video unit and the video bitstream. Wherein, the bitstream conforms to a first rule, the first rule specifying the maximum dimension of the first video unit, and wherein the dimension of the first video unit includes the width, height, or number of samples in the first video unit; and Wherein, the maximum dimension is known prior to the transformation; The first video unit includes a slice, wherein the maximum dimension includes the maximum width of the slice, wherein the first rule specifies that the maximum width of the slice is equal to the maximum number of samples counted in the luminance dimension of the codec tree block, and the maximum luminance slice dimension is the maximum luminance slice width.
2. The method according to claim 1, wherein, The first video unit also includes strips or sub-images.
3. The method according to claim 2, wherein, The dimension of the first video unit is determined based on the number of codec tree blocks in the strip.
4. The method according to claim 2, wherein, The maximum dimension of the first video unit is determined based on the maximum number of samples counted in terms of luminance dimension in the codec tree block.
5. The method according to claim 1, wherein, The maximum luminance dimension is indicated in the bitstream.
6. The method according to claim 5, wherein, The bitstream includes syntax elements that indicate the maximum dimension of the video unit allowed in the current sequence, current picture, current stripe, or current sub-picture.
7. The method according to claim 5, wherein, For rectangular strips or rectangular sub-images, the maximum brightness dimension is indicated in the bitstream.
8. The method according to claim 5, wherein, For a raster scan strip, the maximum brightness dimension is indicated in the bit stream.
9. The method according to any one of claims 5 to 8, wherein, The maximum brightness dimension is a fixed value.
10. The method according to claim 9, wherein, The maximum brightness width or height is 1920 or 4096.
11. The method according to claim 9, wherein, The maximum brightness area is 2,073,600 or 83,388,608.
12. The method according to claim 1, wherein, The first rule further specifies the maximum number of sub-units in the first video unit.
13. The method according to claim 12, wherein, The maximum number of signaling notifications for the sub-units in the bit stream.
14. The method according to claim 12, wherein, The maximum number of sub-units varies depending on the characteristics of the video, including its quality, layer, and level.
15. The method according to claim 1, wherein, The maximum dimension of the first video unit changes based on the characteristics of the video.
16. The method according to claim 15, wherein, The characteristics include the video's grade, layer, or level.
17. The method according to claim 1, further comprising: Perform a conversion between the current image of the video and the bitstream of the video, wherein the current image is associated with one or more reference images such that (1) the dimension of the current image is PW×PH, (2) the scaling window dimension of the current image is SW×WH, (3) the scaling window dimension of the reference images of the current image is SW'×SH', and (4) the maximum allowed dimension W of the image. max and H max constraint.
18. The method according to claim 17, wherein, The constraint specifies that a×Wmax×SW- (b×SW'+c×SW)×PW+offw≥0, where a, b, c, and offw are integers.
19. The method of claim 17, wherein, The constraint specifies that d ×Hmax×SH-(e×SH'+f×SH)×PH+offh≥0, where d, e, f, and offh are integers.
20. The method of claim 17, wherein, The constraint specifies that a×Wmax×SW- b×SW'×PW+offw≥0, where a and b are integers.
21. The method according to claim 17, wherein, The constraint specifies that d×Hmax×SH-e×SH'×PH+offh≥0, where d and e are integers.
22. The method according to claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW, Qw=PW' / PW, rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies rw≤Rw+(Lw×(Rw-1)) / Bw+offw, where Lw, Bw and offw are integers.
23. The method according to claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW, Qw=PW' / PW, rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies rh≤Rh+(Lh×(Rh-1)) / Bh+offh, where Lh, Bh and offh are integers.
24. The method of claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW, Qw=PW' / PW, rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies that rw≤a×Rw+offw, where a and offw are integers.
25. The method according to claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW, Qw=PW' / PW, rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies h≤b×Rh+offh, where b and offh are integers.
26. The method according to claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW, Qw=PW' / PW, rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies that rw≤a×Qw+offw, where a and offw are integers.
27. The method according to claim 17, wherein, rw=SW' / SW, Rw=Wmax / PW,Qw=PW' / PW,rh=SH' / SH and Rh=Hmax / PH, Qh=PH' / PH, wherein the constraint specifies that rh≤b×Qh+offh, where b and offh are integers.
28. The method according to claim 1, further comprising: Perform a conversion between images of the video and the bitstream of the video, wherein the bitstream conforms to a second rule, the second rule specifying that quantization parameter information is signaled in the image header and excluded from the image parameter set.
29. The method according to claim 28, wherein, The quantization parameter information is communicated via signaling to a specific codec tool, which includes Adaptive Color Transform (ACT).
30. The method according to claim 1, further comprising: Perform a conversion between the video and the video bitstream, wherein the bitstream conforms to a third rule that specifies the stripe type in a second video unit of the video associated with multiple stripes, the second video unit including a picture parameter set or picture header, the stripe type including I stripe, P stripe, or B stripe.
31. The method according to claim 30, wherein, The bitstream includes a first syntax element that indicates whether all multiple stripes associated with the second video unit have the same stripe type.
32. The method according to claim 31, wherein, The bitstream includes a second syntax element that indicates the same stripe type shared by all multiple stripes.
33. The method according to claim 32, wherein, Based on the first syntax element, the second syntax element is conditionally signaled in the bitstream.
34. The method according to claim 33, wherein, The second syntax element is signaled in the bitstream only when the first syntax element indicates that all multiple stripes associated with the second video unit have the same stripe type.
35. The method according to claim 32, wherein, The second syntax element is binarized or interpreted in the same way as the third syntax element that indicates the strip type in the strip header.
36. The method according to claim 35, wherein, The third syntax element is conditionally signaled in the bitstream based at least on the first syntax element or the second syntax element.
37. The method of claim 36, wherein, The third syntax element is signaled in the bitstream only when the first syntax element indicates that multiple stripes associated with the second video unit have different stripe types.
38. The method according to claim 36, wherein, In the case where the third syntax element is omitted in the bitstream, the third syntax element is inferred to have a default value based on the first and / or second syntax elements.
39. The method according to claim 30, wherein, The bitstream includes one or more syntax elements from a sequence parameter set, a picture parameter set, or a picture header, indicating whether the stripe type is allowed for the second video unit.
40. The method according to claim 39, wherein, In cases where B-striping is not permitted in the second video unit, semantic constraints are added to the image parameter set of the bitstream.
41. The method according to claim 30, wherein, The bitstream conditionally includes one or more syntax elements, which constrain the stripe type of the second video unit based at least on a first syntax element or a second syntax element.
42. The method according to claim 41, wherein, The signaling in the bitstream notifies one or more syntax elements only when the first syntax element indicates that multiple stripes associated with the second video unit have different stripe types.
43. The method according to claim 41, wherein, In the case where one or more syntax elements are omitted in the bitstream, it is inferred that the one or more syntax elements have default values, at least based on the first syntax element or the second syntax element.
44. The method according to claim 42 or 43, wherein, The one or more syntax elements include syntax elements indicating whether all multiple stripes have I-stripes of a stripe type.
45. The method according to claim 42 or 43, wherein, The one or more syntax elements include syntax elements indicating whether all multiple stripes have a stripe type of B-strip or P-strip.
46. The method according to claim 42 or 43, wherein, The one or more syntax elements include syntax elements indicating whether all multiple stripes have a stripe type of P-strip or I-strip.
47. The method according to any one of claims 41-43, wherein, Whether a signaling notification for one or more inter-frame associated syntax elements is included in the image header of the bitstream is based on the stripe type or whether all multiple stripes have the same stripe type.
48. The method according to claim 47, wherein, The syntax elements for the one or more inter-frame associations include at least one of the following: ph_log2_diff_min_qt_min_cb_inter_slice, ph_max_mtt_hierarchy_depth_inter_slice, ph_log2_diff_max_bt_min_qt_inter_slice, ph_log2_diff_max_tt_min_qt_inter_slice, ph_cu_qp_delta_subdiv_inter_slice, ph_cu_chroma_qp_offset_subdiv_inter_slice, ph_temporal_mvp_enabled_flag, ph_collocated_from_l0_flag, ph_collocated_ref_idx, mvd_l1_zero_flag, ph_fpel_mmvd_enabled_flag, ph_disable_bdof_flag, ph_disable_dmvr_flag. ph_disable_prof_flag, or pred_weight_table.
49. The method according to any one of claims 41 to 43, wherein, Whether the stripe header includes a syntax element to indicate the stripe type depends on whether all the multiple stripes have the same stripe type.
50. The method of claim 1, further comprising: Perform a conversion between the video and the video bitstream, wherein the bitstream includes values indicating the starting point of a colorimetric parameter mapping table.
51. The method according to claim 50, wherein, The value is indicated without differential encoding / decoding.
52. The method according to claim 50, wherein, The value is used for signaling notification as an unsigned value instead of a signed value.
53. The method according to any one of claims 1 to 8, 12-43, and 50-52, wherein, The conversion includes encoding the video into the bitstream.
54. The method of claim 53, further comprising: The bitstream is stored in a computer-readable medium.
55. The method according to any one of claims 1 to 8, 12-43, and 50-52, wherein, The conversion includes decoding the bitstream into the video.
56. A video processing apparatus comprising a processor configured to implement the method as claimed in any one of claims 1 to 55.
57. A computer-readable medium having code stored thereon, which, when executed by a processor, causes the processor to perform the method as described in any one of claims 1 to 55.
58. A computer-readable medium storing a bit stream generated according to any one of claims 1 to 55.