Prediction type signaling in video codecs

Through adaptive resolution change and reference image resampling technology, the video resolution and size are dynamically adjusted, solving the problem of low resolution change efficiency in existing technologies, achieving efficient encoding and decoding and optimizing user experience.

CN114556934BActive Publication Date: 2025-09-16DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080071644.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-12
Filing Date
2020-10-12
Publication Date
2025-09-16
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

Existing video codec technologies are inefficient when resolution and size change, resulting in wasted resources and poor user experience. Flexible resolution adjustment is particularly difficult to achieve when network conditions change or during multi-party video conferencing.

Method used

Adaptive Resolution Change (ARC) technology is used to dynamically adjust video resolution and size through reference picture resampling (RPR) and consistency window management, and combined with interpolation filters and deblocking filters to optimize the encoding and decoding process of video blocks.

Benefits of technology

It achieves efficient encoding and decoding when the resolution and size change, reduces resource waste, improves user experience, supports fast startup and seamless switching, and adapts to changes in network conditions and the needs of multi-party video conferencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114556934B_ABST
    Figure CN114556934B_ABST
Patent Text Reader

Abstract

An example method of video processing includes performing conversion between a video comprising a video picture of one or more video units and a bitstream representation of the video, the bitstream representation conforming to a format rule that provides for including a first syntax element in a picture header to indicate allowed prediction types for at least some of the one or more video units in the video picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a Chinese national phase application of International Patent Application No. PCT / CN2020 / 120289, filed on October 12, 2020. This application claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 110902, filed on October 12, 2019. The entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to video encoding and decoding technology, equipment and systems. Background Art

[0004] Currently, efforts are underway to improve the performance of current video codec technologies to provide better compression ratios or to provide video codec schemes that allow for lower complexity or parallel implementations. Several new video codec tools have recently been proposed by industry experts and are currently being tested to determine their effectiveness. Summary of the Invention

[0005] Devices, systems, and methods are described for digital video coding, particularly for managing motion vectors. The methods are applicable to existing video coding standards (e.g., High Efficiency Video Codec (HEVC) or Versatile Video Codec) and future video coding standards or codecs.

[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing conversion between a video and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies indicating in the bitstream representation the applicability of a decoder-side motion vector refinement codec and a bidirectional optical flow codec for a picture of the video.

[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing conversion between a video picture of the video and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies indicating use of a codec in a picture header corresponding to the video picture.

[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing conversion between a video comprising a video picture of one or more video units and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies including a first syntax element in a picture header to indicate allowed prediction types for at least some of the one or more video units in the video picture.

[0009] In another representative aspect, the disclosed technology may be used to provide a method for video processing, the method comprising: performing a conversion between a current video block and a codec representation of the current video block, wherein, during the conversion, if a resolution and / or size of a reference picture differs from the resolution and / or size of the current video block, applying a same interpolation filter to a set of adjacent or non-adjacent samples predicted using the current video block.

[0010] In another representative aspect, the disclosed technology may be used to provide another method for video processing. The method includes performing a conversion between a current video block and a codec representation of the current video block, wherein, during the conversion, if a resolution and / or size of a reference picture differs from the resolution and / or size of the current video block, blocks predicted using the current video block are only allowed to use integer-valued motion information associated with the current block.

[0011] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a codec representation of the current video block, wherein during the conversion, if a resolution and / or size of a reference picture differs from the resolution and / or size of the current video block, applying an interpolation filter to derive a block predicted using the current video block, and wherein the interpolation filter is selected based on a rule.

[0012] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a codec representation of the current video block, wherein during the conversion, if a resolution and / or size of a reference picture differs from the resolution and / or size of the current video block, selectively applying a deblocking filter, wherein the deblocking filter is set according to a rule related to the resolution and / or size of the reference picture relative to the resolution and / or size of the current video block.

[0013] In another representative aspect, the disclosed technology may be used to provide another method for video processing, the method comprising: performing a conversion between a current video block and a codec representation of the current video block, wherein, during the conversion, a reference picture of the current video block is resampled according to a rule based on a dimension of the current video block.

[0014] In another representative aspect, the disclosed technology may be used to provide another method for video processing, the method comprising: performing a conversion between a current video block and a codec representation of the current video block, wherein, during the conversion, use of codec tools for the current video block is selectively enabled or disabled depending on a resolution / size of a reference picture of the current video block relative to a resolution / size of the current video block.

[0015] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing conversion between a plurality of video blocks and codec representations of the plurality of video blocks, wherein during the conversion, a first conformance window is defined for a first video block and a second conformance window is defined for a second video block, and wherein a ratio of a width and / or a height of the first conformance window to the second conformance window complies with a rule based at least on a conforming bitstream.

[0016] In another representative aspect, the disclosed technology may be used to provide another method for video processing, the method comprising: performing conversion between a plurality of video blocks and codec representations of the plurality of video blocks, wherein during the conversion, a first conformance window is defined for a first video block and a second conformance window is defined for a second video block, wherein a ratio of a width and / or a height of the first conformance window to the second conformance window complies with a rule based at least on a conformance bitstream.

[0017] Furthermore, in one representative aspect, an apparatus in a video system is disclosed, the apparatus comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to implement any one or more of the disclosed methods.

[0018] In one representative aspect, a video decoding apparatus is disclosed that includes a processor configured to implement the methods described herein.

[0019] In one representative aspect, a video encoding apparatus is disclosed that includes a processor configured to implement the methods described herein.

[0020] Furthermore, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product comprising program code for performing any one or more of the disclosed methods.

[0021] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Examples of sub-block motion vectors (VSBs) and motion vector differences are shown.

[0023] Figure 2 An example of a 16x16 video block partitioned into 16 4x4 regions is shown.

[0024] Figure 3A Examples of specific locations in the sample points are shown.

[0025] Figure 3B Another example of a specific location in a sample point is shown.

[0026] Figure 3C Yet another example of a specific location in a sample point is shown.

[0027] Figure 4A An example of the position of the current sample point and its reference sample point is shown.

[0028] Figure 4B Another example of the positions of the current sample and its reference sample is shown.

[0029] Figure 5 is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding and decoding techniques described in this document.

[0030] Figure 6 A flow chart illustrating an example method for video encoding and decoding.

[0031] Figure 7 is a block diagram of an exemplary video processing system in which the disclosed technology may be implemented.

[0032] Figure 8 is a block diagram illustrating an exemplary video coding system.

[0033] Figure 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0034] Figure 10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0035] Figure 11 is a flowchart representation of a method for video processing according to the present technology.

[0036] Figure 12 is a flowchart representation of another method for video processing according to the present technology.

[0037] Figure 13 is a flowchart representation of another video processing method according to the present technology. DETAILED DESCRIPTION

[0038] 1. Video Codec in HEVC / H.265

[0039] Video codec standards have primarily evolved through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Video, and the two organizations jointly produced the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of a 50% bitrate reduction compared to HEVC.

[0040] 2. Overview

[0041] 2.1 Adaptive Resolution Change (ARC)

[0042] AVC and HEVC do not have the ability to change resolution without having to introduce IDR or intra-frame random access point (IRAP) pictures; this capability can be called adaptive resolution change (ARC). Some use cases or application scenarios can benefit from the ARC feature, including:

[0043] Rate Adaptation in Video Calling and Conferencing: To adapt the video codec to changing network conditions, the encoder can adapt by encoding smaller resolution pictures when network conditions worsen and available bandwidth becomes lower. Currently, picture resolution changes can only be made after an IRAP picture has been rendered; this presents several challenges. An IRAP picture of reasonable quality would be significantly larger than an inter-codec picture and correspondingly more complex to decode, wasting time and resources. This can also be problematic if the decoder requests a resolution change for load reasons. It can also disrupt low-latency buffer conditions, forcing audio resynchronization and increasing the end-to-end latency of the stream, at least temporarily. This results in a poor user experience.

[0044] -Changes in the active speaker during multi-party video conferencing: In multi-party video conferencing, the active speaker is typically displayed at a larger video size than the other conference participants. When the active speaker changes, the image resolution of each participant may also need to be adjusted. The need for ARC becomes particularly important when such changes occur frequently in the active speaker.

[0045] - Fast start in streaming: For streaming applications, usually the application will buffer up to a certain length of decoded pictures before starting to display. Starting the bitstream at a smaller resolution will allow the application to have enough pictures in the buffer to start displaying faster.

[0046] Adaptive stream switching in streaming: The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature called @mediaStreamStructureId. This enables switching between different representations at open GOP random access points with non-decodable leading pictures (e.g., CRA pictures with associated RASL pictures in HEVC). When two different representations of the same video have different bitrates but the same spatial resolution, and they have the same @mediaStreamStructureId value, switching between the two representations can be performed on CRA pictures with associated RASL pictures, and the RASL pictures associated with the CRA pictures at the switch can be decoded with acceptable quality, resulting in seamless switching. Using ARC, the @mediaStreamStructureId feature can also be used to switch between DASH representations with different spatial resolutions.

[0047] ARC is also known as dynamic resolution switching.

[0048] ARC can also be viewed as a special case of Reference Picture Resampling (RPR), such as H.263 Annex P.

[0049] 2.2 Reference Picture Resampling in H.263 Annex P

[0050] This mode describes an algorithm that warps the reference picture before it is used for prediction. It can be useful for resampling reference pictures that have a different source format than the picture being predicted. By warping the shape, size and position of the reference picture, it can also be used for global motion estimation or rotational motion estimation. The syntax includes the warping parameters to be used as well as the resampling algorithm. The simplest level of operation of the reference picture resampling mode is an implicit factor of 4 resampling, as only FIR filters are needed for the upsampling and downsampling processes. In this case, when the size of the new picture (indicated in the picture header) is different from the size of the previous picture, no additional signaling overhead is required since its usage is understood.

[0051] 2.3.Consistency Window in VVC

[0052] The consistency window in VVC defines a rectangle. Samples within the consistency window belong to the image of interest. Samples outside the consistency window may be discarded during output.

[0053] When a consistency window is applied, the scaling ratio in the RPR is derived based on the consistency window.

[0054] Picture Parameter Set RBSP Syntax

[0055]

[0056] Specifies the width of each decoded picture that references the PPS in units of luma samples. pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples.

[0057] When subpics_present_flag is equal to 1, the value of pic_width_in_luma_samples shall be equal to pic_width_max_in_luma_samples.

[0058] Specifies the height of each decoded picture that references the PPS in units of luma samples. pic_height_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_height_max_in_luma_samples.

[0059] When subpics_present_flag is equal to 1, the value of pic_height_in_luma_samples shall be equal to pic_height_max_in_luma_samples.

[0060] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be the pic_width_in_luma_samples and pic_height_in_luma_samples, respectively, of the reference picture that refers to the current picture of this PPS. It is a requirement for bitstream conformance that all of the following conditions are met:

[0061]

[0062] If conformance_window_flag is equal to 1, the conformance cropping window offset parameter follows immediately in the SPS. If conformance_window_flag is equal to 0, the conformance cropping window offset parameter does not exist.

[0063] and The picture samples in the CVS output from the decoding process are specified in a rectangular area specified in picture coordinates for output. When conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset and conf_win_bottom_offset are inferred to be 0.

[0064] The conforming cropping window contains the luma samples with horizontal picture coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical picture coordinates from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1), inclusive.

[0065] The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) should be less than pic_height_in_luma_samples.

[0066] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0067] PicOutputWidthL=pic_width_in_luma_samples- (7-43)

[0068] SubWidthC*(conf_win_right_offset+conf_win_left_offset)

[0069] PicOutputHeightL=pic_height_in_pic_size_units

[0070] -SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)

[0071] (7-44)

[0072] When ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.

[0073] NOTE – The consistency crop window offset parameter applies only to the output. All internal decoding processes apply to the uncropped The size of the cropped image.

[0074] Assume that ppsA and ppsB are any two PPSs that reference the same SPS. When pic_width_in_luma_samples and pic_height_in_luma_samples of ppsA and ppsB have the same values, ppsA and ppsB shall have the same values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset, respectively, as required for bitstream conformance.

[0075] 2.4. Reference Picture Resampling (RPR)

[0076] In some embodiments, ARC is also referred to as reference picture resampling (RPR). For RPR, TMVP is disabled if the collocated picture has a different resolution than the current picture. Additionally, bidirectional optical flow (BDOF) and decoder-side motion vector refinement (DMVR) are disabled when the reference picture has a different resolution than the current picture.

[0077] To handle normal MC when the reference picture has a different resolution than the current picture, the interpolation part is defined as follows:

[0078] 8.5.6.3 Fractional Sample Interpolation Processing

[0079] 8.5.6.3.1 Overview

[0080] The inputs to this process include:

[0081] – Luma position (xSb, ySb), specifies the upper left sample of the current codec sub-block relative to the upper left luminance sample of the current picture,

[0082] –Variable sbWidth specifies the width of the current codec sub-block.

[0083] –Variable sbHeight specifies the height of the current codec sub-block.

[0084] – motion vector offset mvOffset,

[0085] – refined motion vector refMvLX,

[0086] – the selected reference picture sample array refPicLX,

[0087] – Half-sample interpolation filter index hpelIfIdx,

[0088] – bidirectional optical flow flag bdofFlag,

[0089] –Variable cIdx specifies the color component index of the current block.

[0090] The output of this processing includes:

[0091] – predSamplesLX array of (sbWidth+brdExtSize)x(sbHeight+brdExtSize) prediction sample values.

[0092] The prediction block boundary extension size brdExtSize is derived as follows:

[0093] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&

[0094] sps_affine_prof_enabled_flag))? 2:0 (8-752)

[0095] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0096] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0097] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0098] – If cIdx is equal to 0, the following applies:

[0099] –The scaling factor and its fixed-point representation are defined as

[0100] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0101] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0102] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0103] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0104] – For each luma sample position (x L =0..sbWidth-1+brdExtSize,y L=0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0105] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0106] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)

[0107] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)

[0108] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757)

[0109] refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)

[0110] –Variable xInt L ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0111] xInt L =refx L >>4 (8-759)

[0112] yInt L =refy L >>4 (8-760)

[0113] xFrac L =refx L &15 (8-761)

[0114] yFrac L =refy L &15 (8-762)

[0115] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0116] –x L Equal to 0.

[0117] –x L Equal to sbWidth+1.

[0118] –y L Equal to 0.

[0119] –y L Equal to sbHeight+1.

[0120] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L ),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0121] – Otherwise (cIdx is not equal to 0), the following applies:

[0122] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0123] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0124] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0125] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0126] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0127] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)

[0128] refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)

[0129] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765)

[0130] refy C=((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0131] –Variable xInt C ,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0132] xInt C =refx C >>5 (8-767)

[0133] yInt C =refy C >>5 (8-768)

[0134] xFrac C =refy C &31 (8-769)

[0135] yFrac C =refy C &31 (8-770)

[0136] – The predicted sample values ​​predSamplesLX[xC][yC] are derived by invoking the process specified in clause 8.5.6.3.4, with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC), sbWidth, sbHeight and refPicLX as input.

[0137] 8.5.6.3.2 Luminance Sample Interpolation Filtering

[0138] The input to this process is:

[0139] – Luminance position in full sample units (xInt L ,yInt L ),

[0140] – Luma position in fractional samples (xFrac L ,yFrac L ),

[0141] – Luminance position in units of full samples (xSbInt L ,ySbInt L), which specifies the upper left sample of the boundary block used for reference sample filling relative to the upper left luma sample of the reference picture,

[0142] – Luminance reference sample array refPicLX L ,

[0143] – Half-sample interpolation filter index hpelIfIdx,

[0144] –Variable sbWidth specifies the width of the current sub-block,

[0145] –Variable sbHeight specifies the height of the current sub-block.

[0146] – Luma position (xSb, ySb), specifies the upper left sample point of the current sub-block relative to the upper left luminance sample point of the current picture,

[0147] The output of this process is the predicted luminance sample value predSampleLX L

[0148] The variables shift1, shift2, and shift3 are derived as follows:

[0149] – The variable shift1 is set equal to Min(4, BitDepth Y -8), the variable shift2 is set equal to 6 and the variable shift3 is set equal to Max(2,14-BitDepth Y ).

[0150] – The variable picW is set equal to pic_width_in_luma_samples and the variable picH is set equal to pic_height_in_luma_samples.

[0151] Luma interpolation filter coefficient f for each 1 / 16 fractional sample position p L [p] is equal to xFrac L or yFrac L The derivation is as follows:

[0152] – If MotionModelIdc[xSb][ySb] is greater than 0, and sbWidth and sbHeight are both equal to 4, then the luminance interpolation filter coefficient f L [p] is specified in Table 2.

[0153] – Otherwise, the luminance interpolation filter coefficient f L [p] depends on hpelIfIdx as specified in Table 1.

[0154] Luminance position of all sample points (xInti ,yInt i ) is derived as follows, for i = 0..7:

[0155] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0156] xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3)(8-771)

[0157] yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772)

[0158] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following applies:

[0159] xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?

[0160] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3):

[0161] (8-773)

[0162] xInt L +i-3)

[0163] yInt i =Clip3(0,picH-1,yInt L +i-3) (8-774)

[0164] The brightness position of all sample units is further modified as follows, for i = 0..7:

[0165] xInt i =Clip3(xSbInt L -3,xSbInt L +sbWidth+4,xInt i ) (8-775)

[0166] yInt i =Clip3(ySbInt L -3,ySbInt L +sbHeight+4,yInt i ) (8-776)

[0167] Predicted brightness sample value predSampleLX L The derivation is as follows:

[0168] – If xFrac L and yFrac L are all equal to 0, then predSampleLX L The value of is derived as follows:

[0169] predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3 (8-777)

[0170] – Otherwise, if xFrac L Not equal to 0 and yFrac L If equal to 0, then predSampleLX L The value of is derived as follows:

[0171]

[0172] – Otherwise, if xFrac L Equal to 0 and yFrac L If not equal to 0, then predSampleLX L The value of is derived as follows:

[0173]

[0174] – Otherwise, if xFrac L Not equal to 0 and yFrac L If not equal to 0, then predSampleLX L The value of is derived as follows:

[0175] – Sample array temp[n] where n = 0..7, derived as follows:

[0176]

[0177] –Predicted luminance sample value predSampleLX L The derivation is as follows:

[0178]

[0179] Table 8-11 – Luma interpolation filter coefficients f for each 1 / 16 fractional sample position p L Specification of [p]

[0180]

[0181] Table 8-12 Luma interpolation filter coefficient f for each 1 / 16 fractional sample position p in affine motion mode L Specification of [p]

[0182]

[0183] 8.5.6.3.3 Luminance integer sample acquisition processing

[0184] The input to this process is:

[0185] – Luminance position in full sample units (xInt L ,yInt L ),

[0186] – Luminance reference sample array refPicLX L ,

[0187] The output of this process is the predicted luminance sample value predSampleLX L

[0188] The variable shift is set equal to Max(2,14-BitDepth Y ).

[0189] The variable picW is set equal to pic_width_in_luma_samples and the variable picH is set equal to pic_height_in_luma_samples.

[0190] The brightness position (xInt, yInt) of all sample units is derived as follows:

[0191] xInt=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag? (8-782)

[0192] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L ):xInt L )

[0193] yInt=Clip3(0,picH-1,yInt L ) (8-783)

[0194] Predicted brightness sample value predSampleLX L The derivation is as follows:

[0195] predSampleLX L =refPicLX L [xInt][yInt]< <shift3

[0196] (8-784)

[0197] 8.5.6.3.4 Chroma sample interpolation processing

[0198] The input to this process is:

[0199] – Chroma position of all sample points (xInt C ,yInt C ),

[0200] – Chroma position in 1 / 32 fractional sample units (xFrac C ,yFrac C ),

[0201] – Chroma position in full sample units (xSbIntC, ySbIntC), specifies the upper left sample of the boundary block used for reference sample filling relative to the upper left chroma sample of the reference picture,

[0202] –Variable sbWidth specifies the width of the current sub-block,

[0203] –Variable sbHeight specifies the height of the current sub-block.

[0204] – Chroma reference sample array refPicLX C .

[0205] The output of this process is the predicted chroma sample value predSampleLX C

[0206] The variables shift1, shift2, and shift3 are derived as follows:

[0207] – The variable shift1 is set equal to Min(4, BitDepth C -8), the variable shift2 is set equal to 6 and the variable shift3 is set equal to Max(2,14-BitDepth C ).

[0208] –Variable picWC is set equal to pic_width_in_luma_samples / SubWidthC and the variable picH C Set equal to pic_height_in_luma_samples / SubHeightC.

[0209] Chroma interpolation filter coefficient f for each 1 / 32 fractional sample position p C [p] is equal to xFrac C or yFrac C are specified in Table 3.

[0210] The variable xOffset is set equal to (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC.

[0211] Chroma position of all sample points (xInt i ,yInt i ) is derived as follows, for i = 0..3:

[0212] – If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0213] xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (8-785)

[0214] yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (8-786)

[0215] – Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following applies:

[0216] xInt i =Clip3(0,picW C -1,

[0217] sps_ref_wraparound_enabled_flag? ClipH(xOffset,picWC ,xInt C +i-1): (8-787)

[0218] xInt C +i-1)

[0219] yInt i =Clip3(0,picH C -1,yInt C +i-1) (8-788)

[0220] Chroma position of all sample points (xInt i ,yInt i ) is further modified as follows, for i=0..3:

[0221] xInt i =Clip3(xSbIntC-1,xSbIntC+sbWidth+2,xInt i ) (8-789)

[0222] yInt i =Clip3(ySbIntC-1,ySbIntC+sbHeight+2,yInt i ) (8-790)

[0223] Predicted chroma sample value predSampleLX C The derivation is as follows:

[0224] – If xFrac C and yFrac C are all equal to 0, then predSampleLX C The value of is derived as follows:

[0225] predSampleLX C =refPicLX C [xInt1][yInt1]< <shift3 (8-791)

[0226] – Otherwise, if xFrac C Not equal to 0 and yFrac C If equal to 0, then predSampleLX C The value of is derived as follows:

[0227]

[0228] – Otherwise, if xFrac C Equal to 0 and yFrac CIf not equal to 0, then predSampleLX C The value of is derived as follows:

[0229]

[0230] – Otherwise, if xFrac C Not equal to 0 and yFrac C If not equal to 0, then predSampleLX C The value of is derived as follows:

[0231] – Sample array temp[n] where n = 0..3, derived as follows:

[0232]

[0233] –Predicted chroma sample value predSampleLX C The derivation is as follows:

[0234]

[0235] Table 8-13 – Chroma interpolation filter coefficients f for each 1 / 32 fractional sample position p C Specification of [p]

[0236]

[0237] 2.5. Affine Motion Compensated Prediction Based on Refined Sub-Blocks

[0238] The technology disclosed herein involves using optical flow to refine sub-block affine motion compensation predictions. After performing sub-block affine motion compensation, the prediction samples are refined by adding differences derived from the optical flow equation, a technique called optical flow prediction refinement (PROF). This method enables pixel-level inter-frame prediction without increasing memory access bandwidth.

[0239] To achieve more refined motion compensation, this paper proposes a sub-block-based affine motion compensation prediction method using optical flow refinement. After performing sub-block-based affine motion compensation, the luma prediction samples are refined by adding the differences derived from the optical flow equation. The proposed PROF (Prediction Refinement using Optical Flow) method describes the following four steps.

[0240] Step 1) Perform sub-block based affine motion compensation to generate sub-block prediction I(i, j).

[0241] Step 2) Calculate the spatial gradient g of the sub-block prediction using a 3-tap filter [-1, 0, 1] at each sample point x (i, j) and g y (i, j).

[0242] g x (i,j)=I(i+1,j)-I(i-1,j)

[0243] g y (i,j)=I(i,j+1)-I(i,j-1)

[0244] The sub-block prediction is extended by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, pixels on the extended boundary are copied from the nearest integer pixel position in the reference picture. Thus, additional interpolation in the padded area is avoided.

[0245] Step 3) The brightness prediction is refined (denoted as ΔI), which is calculated using the optical flow equation.

[0246] ΔI(i, j)=g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j)

[0247] where delta MV (denoted as Δv(i, j))) is the difference between the pixel MV (denoted as v(i, j)) calculated for the sample position (i, j) and the sub-block MV of the sub-block to which the pixel (i, j) belongs, as Figure 1 shown.

[0248] Since the affine model parameters and the position of the pixel relative to the sub-block center do not change from sub-block to sub-block, Δv(i, j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the sub-block center, Δv(x, y) can be derived as follows,

[0249]

[0250] For a 4-parameter affine model,

[0251]

[0252] For the 6-parameter affine model,

[0253]

[0254] Where (v 0x ,v 0y ),(v 1x ,v 1y ),(v 2x ,v 2y ) are the control point motion vectors of the upper left corner, upper right corner and lower left corner, w and h are the width and height of the CU.

[0255] Step 4) Finally, the luma prediction refinement is added to the sub-block prediction I(i, j). The final prediction value I' is generated as follows.

[0256] I'(i,j)=I(i,j)+ΔI(i,j)

[0257] Some details are as follows:

[0258] a) How to derive the gradient of PROF

[0259] In some embodiments, the gradient of each sub-block (4×4 sub-blocks in VTM-4.0) is calculated for each reference list. For each sub-block, the nearest integer samples of the reference block are taken to fill the four outer lines of samples.

[0260] Assume that the MV of the current sub-block is (MVx, MVy). The fractional part is calculated as (FracX, FracY) = (MVx & 15, MVy & 15). The integer part is calculated as (IntX, IntY) = (MVx>>4, MVy>>4).

[0261] The offset (OffsetX, OffsetY) is derived as:

[0262] OffsetX=FracX>7?1:0;

[0263] OffsetY=FracY>7?1:0;

[0264] Assuming that the coordinates of the upper left corner of the current sub-block are (xCur, yCur), and the dimension of the current sub-block is W×H, then (xCor0, yCor0), (xCor1, yCor1), (xCor2, yCor2), and (xCor3, yCor3) are calculated as follows:

[0265] (xCor0,yCor0)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY-1);

[0266] (xCor1,yCor1)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY+H);

[0267] (xCor2,yCor2)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY);

[0268] (xCor3,yCor3)=(xCur+IntX+OffsetX+W,yCur+IntY+OffsetY);

[0269] Assume that the pre-samples [x][y], x = 0..W-1, y = 0..H-1 store the predicted samples of the sub-block. Then the filling samples are derived as:

[0270] PredSample[x][-1]=(Ref(xCor0+x,yCor0)< <Shift0)-Rounding,for x=-1..W;

[0271] PredSample[x][H]=(Ref(xCor1+x,yCor1)< <Shift0)-Rounding,for x=-1..W;

[0272] PredSample[-1][y]=(Ref(xCor2,yCor2+y)< <Shift0)-Rounding,for y=0..H-1;

[0273] PredSample[W][y]=(Ref(xCor3,yCor3+y)< <Shift0)-Rounding,for y=0..H-1;

[0274] Where Rec represents the reference picture. Rounding is an integer equal to 2 in the exemplary PROF implementation. 13 Shift0 = Max(2,(14-BitDepth)); PROF attempts to improve the accuracy of the gradient, unlike BIO in VTM-4.0, which outputs the gradient with the same accuracy as the input luminance samples.

[0275] The gradient in PROF is calculated as follows:

[0276] Shift1=Shift0-4.

[0277] gradientH[x][y]=(predSamples[x+1][y]-predSample[x-1][y])>>Shift1

[0278] gradientV[x][y]=(predSample[x][y+1]-predSample[x][y-1])>>Shift1

[0279] It should be noted that predSamples[x][y] maintains the accuracy after interpolation.

[0280] b) How to derive Δv for PROF

[0281] The derivation of Δv (expressed as dMvH[posX][posY] and dMvV[posX][posY], posX=0..W-1, posY=0..H-1) is as follows.

[0282] Assume that the dimension of the current block is cbWidth×cbHeight, the number of control point motion vectors is numCpMv, the control point motion vector is cpMvLX[cpIdx], cpIdx=0..numCpMv-1 and X is 0 or 1 representing two reference lists.

[0283] The variables log2CbW and log2CbH are derived as follows:

[0284] log2CbW=Log2(cbWidth)

[0285] log2CbH=Log2(cbHeight)

[0286] The variables mvScaleHor, mvScaleVer, dHorX, and dVerX are derived as follows:

[0287] mvScaleHor=cpMvLX[0][0]<<7

[0288] mvScaleVer=cpMvLX[0][1]<<7

[0289] dHorX=(cpMvLX[1][0]-cpMvLX[0][0])<<(7-log2CbW)

[0290] dVerX=(cpMvLX[1][1]-cpMvLX[0][1])<<(7-log2CbW)

[0291] The variables dHorY and dVerY are derived as follows:

[0292] - If numCpMv is equal to 3, the following applies:

[0293] dHorY=(cpMvLX[2][0]-cpMvLX[0][0])<<(7-log2CbH)

[0294] dVerY=(cpMvLX[2][1]-cpMvLX[0][1])<<(7-log2CbH)

[0295] - Otherwise (numCpMv is equal to 2), the following applies:

[0296] dHorY=-dVerX

[0297] dVerY=dHorX

[0298] The variables qHorX, qVerX, qHorY, and qVerY are calculated as follows:

[0299] qHorX=dHorX<<2;

[0300] qVerX=dVerX<<2;

[0301] qHorY=dHorY<<2;

[0302] qVerY=dVerY<<2;

[0303] dMvH[0][0] and dMvV[0][0] are calculated as follows:

[0304] dMvH[0][0]=((dHorX+dHorY)<<1)-((qHorX+qHorY)<<1);

[0305] dMvV[0][0]=((dVerX+dVerY)<<1)-((qVerX+qVerY)<<1);

[0306] dMvH[xPos][0] and dMvV[xPos][0] (xPos from 1 to W-1) are derived as follows:

[0307] dMvH[xPos][0]=dMvH[xPos-1][0]+qHorX;

[0308] dMvV[xPos][0]=dMvV[xPos-1][0]+qVerX;

[0309] For yPos from 1 to H-1, the following applies:

[0310] dMvH[xPos][yPos]=dMvH[xPos][yPos-1]+qHorY with xPos=0..W-1

[0311] dMvV[xPos][yPos]=dMvV[xPos][yPos-1]+qVerY with xPos=0..W-1

[0312] Finally, dMvH[xPos][yPos] and dMvV[xPos][yPos] where posX=0..W-1, posY=0..H-1 are right-shifted as follows:

[0313] dMvH[xPos][yPos]=SatShift(dMvH[xPos][yPos],7+2-1);

[0314] dMvV[xPos][yPos]=SatShift(dMvV[xPos][yPos],7+2-1);

[0315] Where SatShift(x,n) and Shift(x,n) are defined as follows:

[0316]

[0317] Shift(x,n)=(x+offset0)>>n

[0318] In one example, offset0 and / or offset1 are set to (1 <<n)> >1.

[0319] c) How to derive ΔI for PROF

[0320] For the position (posX, posY) within the sub-block, its corresponding Δv(i, j) is expressed as (dMvH[posX][posY], dMvV[posX][posY]). Its corresponding gradient is expressed as (gradientH[posX][posY], gradientV[posX][posY]).

[0321] Then ΔI(posX, posY) is derived as follows.

[0322] (dMvH[posX][posY], dMvV[posX][posY]) is clipped to:

[0323] dMvH[posX][posY]=Clip3(-32768,32767,dMvH[posX][posY]);

[0324] dMvV[posX][posY]=Clip3(-32768,32767,dMvV[posX][posY]);

[0325] ΔI(posX,posY)=dMvH[posX][posY]×gradientH[posX][posY]+dMvV[posX][posY]×gradientV[posX][posY];

[0326] ΔI(posX,posY)=Shift(ΔI(posX,posY),1+1+4);

[0327] ΔI(posX,posY)=Clip3(-(2 13 -1),2 13 -1,ΔI(posX,posY));

[0328] d) How to derive I' for PROF

[0329] If the current block is not coded using bidirectional prediction or weighted prediction, then

[0330] I'(posX,posY)=Shift((I(posX,posY)+ΔI(posX,posY)),Shift0),

[0331] I'(posX,posY)=ClipSample(I'(posX,posY)),

[0332] ClipSample clips the sample value to a valid output sample value, and then I'(posX, posY) is output as the inter-frame prediction value.

[0333] Otherwise (the current block is coded using bidirectional prediction or weighted prediction), I'(posX, posY) will be stored and used to generate inter-frame prediction values ​​based on other prediction values ​​and / or weighted values.

[0334] 2.6 Example Strip Header

[0335]

[0336]

[0337]

[0338]

[0339]

[0340]

[0341] 2.7 Example Sequence Parameter Set

[0342]

[0343]

[0344]

[0345]

[0346]

[0347]

[0348] 2.8 Example Image Parameter Set

[0349]

[0350]

[0351]

[0352]

[0353]

[0354] 2.9 Example Adaptive Parameter Set

[0355]

[0356]

[0357]

[0358]

[0359]

[0360]

[0361]

[0362] 2.10 Example Image Header

[0363] In some embodiments, the image header is designed to have the following characteristics:

[0364] 1. The temporal ID and layer ID of the picture header NAL unit are the same as those of the layer access unit containing the picture header.

[0365] 2. The picture header NAL unit should be located before the NAL unit containing the first slice of its associated picture. This establishes the association between the picture header and the picture slice associated with the picture header without the need to signal the picture header ID in the picture header and reference it from the slice header.

[0366] 3. The picture header NAL unit should follow the picture level parameter set or higher level, such as DPS, VPS, SPS, PPS, etc. Therefore, it is required that these parameter sets do not repeat / appear within a picture or access unit.

[0367] 4. The image header contains information about the image type of the image it is associated with. The image type can be used to define the following (this is not an exhaustive list):

[0368] a. The image is an IDR image

[0369] b. The image is a CRA image

[0370] c. The picture is a GDR picture

[0371] d. The image is a non-IRAP, non-GDR image and contains only I stripes

[0372] e. The picture is a non-IRAP, non-GDR picture and may contain only P and I slices

[0373] f. The picture is a non-IRAP, non-GDR picture and contains any B, P and / or I slices

[0374] 5. Move the signaling of picture-level syntax elements in the slice header to the picture header.

[0375] 6. Mark non-picture level syntax elements in the slice header, which are usually the same for all slices of the same picture in the picture header. When these syntax elements are not present in the picture header, they can be signaled in the slice header.

[0376] In some implementations, the concept of a mandatory picture header is used to be sent once per picture as the first VCL NAL unit of a picture. It is also proposed to move syntax elements currently in the slice header to the picture header. For a given picture, syntax elements that functionally only need to be sent once per picture can be moved to the picture header instead of being sent multiple times. For example, syntax elements in the slice header are sent once per slice. The moved slice header syntax elements are constrained to be identical across pictures.

[0377] The syntax elements are already constrained to be the same across all slices of a picture. It can be argued that, without making any changes to the functionality of these syntax elements, moving these fields to the picture header so that they are signaled consistently only once per picture instead of once per slice avoids unnecessary bit redundancy.

[0378] 1. In some implementations, the following semantic constraints exist:

[0379] When present, the values ​​of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and slice_temporal_mvp_enabled_flag shall be the same in all slice headers of a codec picture. Therefore, each of these syntax elements can be moved to the picture header to avoid unnecessary redundant bits.

[0380] In this contribution, recovery_poc_cnt and no_output_of_prior_pics_flag are not moved to the picture header. Their presence in the slice header depends on the conditional check of the slice header nal_unit_type, so if it is desirable to move these syntax elements to the picture header, it is recommended to study them.

[0381] 2. In some implementations, the following semantic constraints exist:

[0382] When present, the value of slice_lmcs_aps_id shall be the same for all slices of a picture.

[0383] When present, the value of slice_scaling_list_aps_id shall be the same for all slices of a picture.Therefore, each of these syntax elements can be moved to the picture header to avoid unnecessary redundant bits.

[0384] In some embodiments, syntax elements are not currently constrained to be the same across all slices of a picture. It is recommended to evaluate the expected usage of these syntax elements to determine which ones can be moved into the picture header to simplify the overall VVC design, as processing a large number of syntax elements in each slice header is reported to have a complexity impact.

[0385] 1. It is proposed to move the following syntax elements to the picture header. There is currently no restriction on them having different values ​​on different slices, but it is claimed that there is no / minimal benefit and codec loss to transmitting them in each slice header, as their intended usage will vary at the picture level:

[0386] a.six_minus_max_num_merge_cand

[0387] b.five_minus_max_num_subblock_merge_cand

[0388] c.slice_fpel_mmvd_enabled_flag

[0389] d.slice_disable_bdof_dmvr_flag

[0390] e.max_num_merge_c and _minus_max_num_triangle_cand

[0391] f.slice_six_minus_max_num_ibc_merge_cand

[0392] 2. It is proposed to move the following syntax elements to the picture header. There is currently no restriction on them having different values ​​on different slices, but it is claimed that there is no / minimal benefit and codec loss to transmitting them in each slice header, as their intended usage will vary at the picture level:

[0393] a.partition_constraints_override_flag

[0394] b.slice_log2_diff_min_qt_min_cb_luma

[0395] c.slice_max_mtt_hierarchy_depth_luma

[0396] d.slice_log2_diff_max_bt_min_qt_luma

[0397] e.slice_log2_diff_max_tt_min_qt_luma

[0398] f.slice_log2_diff_min_qt_min_cb_chroma

[0399] g.slice_max_mtt_hierarchy_depth_chroma

[0400] h.slice_log2_diff_max_bt_min_qt_chroma

[0401] i.slice_log2_diff_max_tt_min_qt_chroma

[0402] The conditional check "slice_type == 1" associated with some of these syntax elements has been removed when moved to the picture header.

[0403] 3. It is proposed to move the following syntax elements to the picture header. There is currently no restriction on them having different values ​​on different slices, but it is claimed that there is no / minimal benefit and codec loss to transmitting them in each slice header, as their intended use will vary at the picture level:

[0404] a.mvd_l1_zero_flag

[0405] The conditional check "slice_type == B" associated with some of these syntax elements has been removed when moved to the picture header.

[0406] 4. It is proposed to move the following syntax elements to the picture header. There is currently no restriction on them having different values ​​on different slices, but it is claimed that there is no / minimal benefit and codec loss to transmitting them in each slice header, as their intended use will vary at the picture level:

[0407] a.dep_quant_enabled_flag

[0408] b.sign_data_hiding_enabled_flag

[0409] 2.10.1 Example Syntax Table

[0410] 7.3.2.8 Picture Header RBSP Syntax

[0411]

[0412]

[0413]

[0414]

[0415]

[0416] 3. Disadvantages of existing implementations

[0417] DMVR and BIO do not involve the original signal during the motion vector refinement process, which may lead to inaccurate motion information for the codec block. In addition, DMVR and BIO sometimes use fractional motion vectors after motion refinement, while screen video usually uses integer motion vectors, which makes the current motion information even more inaccurate and worsens the codec performance.

[0418] When RPR is applied to VVC, RPR (ARC) may have the following problems:

[0419] 1. For RPR, the interpolation filters may be different for adjacent samples in a block, which is undesirable in SIMD (Single Instruction Multiple Data) implementation.

[0420] 2. RPR is not considered in the boundary area.

[0421] 3. It is worth noting that "the consistency crop window offset parameters are applied only to output. All internal decoding processes are applied to the uncropped picture size." However, when RPR is applied, these parameters may be used during the decoding process.

[0422] 4. When deriving the reference sample point position, RPR only considers the ratio between the two consistency windows. However, the difference in the upper left offset between the two consistency windows should also be taken into account.

[0423] 5. In VVC, the ratio of the width / height of the reference picture to the width / height of the current picture is constrained. However, the ratio of the width / height of the consistency window of the reference picture to the width / height of the consistency window of the current picture is not constrained.

[0424] 6. Not all syntax elements in the image header are processed correctly.

[0425] 7. In the current VVC, for TPM and GEO prediction modes, chroma mixing weights are derived regardless of the chroma sample position type of the video sequence. For example, in TPM / GEO, if the chroma weights are derived from the luma weights, the luma weights may need to be downsampled to match the samples of the chroma signal. Chroma downsampling is usually applied assuming the chroma sample position type is 0, which is widely used in ITU-R BT.601 or ITU-R BT.709 containers. However, if different chroma sample position types are used, this may result in misalignment between the chroma samples and the downsampled luma samples, which may reduce codec performance.

[0426] 4. Example Techniques and Embodiments

[0427] The detailed embodiments described below should be considered as examples for explaining the general concept. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any manner.

[0428] The method described below can also be applied to other decoder motion information derivation techniques besides DMVR and BIO mentioned below.

[0429] The motion vector is represented by (mv_x, mv_y), where mv_x is the horizontal component and mv_y is the vertical component.

[0430] In the present disclosure, the resolution (or dimension, width / height or size) of a picture may refer to the resolution (or dimension, width / height or size) of the encoded / decoded picture, or may refer to the resolution (or dimension, width / height or size) of a consistency window in the encoded / decoded picture.

[0431] Motion Compensation in RPR

[0432] 1. When the resolution of the reference picture is different from that of the current picture, or when the width and / or height of the reference picture is larger than the width and / or height of the current picture, the same horizontal and / or vertical interpolation filter can be used to generate prediction values ​​for a set of samples (at least two samples) of the current block.

[0433] a. In one example, the group may include all samples in the region of the block.

[0434] i. For example, the block can be divided into S MxN rectangles that do not overlap each other. Each MxN rectangle is a group. Figure 2 In the example shown, a 16x16 block can be divided into 16 4x4 rectangles, each of which is a group.

[0435] ii. For example, a row of N samples is a group. N is an integer not greater than the block width. In one example, N is 4 or 8 or the block width.

[0436] iii. For example, a column of N samples is a group. N is an integer not greater than the block height. In one example, N is 4 or 8 or the block height.

[0437] iv. M and / or N may be predefined or dynamically derived, for example based on block dimension / codec information or signaled.

[0438] b. In one example, samples in a group may have the same MV (denoted as shared MV).

[0439] c. In one example, samples in a group may have MVs with the same horizontal component (denoted as shared horizontal component).

[0440] d. In one example, samples in a group may have MVs with the same vertical component (denoted as shared vertical component).

[0441] e. In one example, samples in a group may have MVs with the same fractional portion of the horizontal component (denoted as shared fractional horizontal component).

[0442] i. For example, assuming that the MV of the first sample point is (MV1x, MV1y) and the MV of the second sample point is (MV2x, MV2y), then MV1x&(2M -1) is equal to MV2x&(2 M -1), where M represents the MV precision. For example, M=4.

[0443] f. In one example, samples in a group may have MVs with the same fractional portion of the vertical component (denoted as shared fractional vertical component).

[0444] i. For example, assuming that the MV of the first sample point is (MV1x, MV1y) and the MV of the second sample point is (MV2x, MV2y), then MV1y&(2 M -1) is equal to MV2y&(2 M -1), where M represents the MV precision. For example, M=4.

[0445] g. In one example, for the samples in the group to be predicted, the predictions can first be made based on the current picture and the reference picture (e.g., the refx derived in 8.5.6.3.1 of JVET-O2001-v14). L ,refy L )) resolution to export using MV b Then, MV b It can be further modified (eg, rounded / truncated / clipped) to MV' to meet requirements such as those mentioned above, and MV' will be used to derive the predicted samples of the samples.

[0446] i. In one example, MV' has the same b The same integer part, and the fractional part of MV' is set to the shared horizontal and / or vertical fractional part.

[0447] ii. In one example, MV' is set to have a shared fractional horizontal and / or vertical component and is closest to MV b .

[0448] h. A shared motion vector (and / or a shared horizontal component and / or a shared vertical component and / or a shared fractional vertical component and / or a shared fractional vertical component) may be set as the motion vector (and / or a horizontal component and / or a vertical component and / or a fractional vertical component and / or a fractional vertical component) of a specific sample of a group.

[0449] i. For example, a specific sample point can be located at a corner of a rectangular group, such as Figure 3A "A", "B", "C" and "D" shown.

[0450] ii. For example, a specific sample point can be located at the center of a rectangular group, such as Figure 3A "E", "F", "G" and "H" shown.

[0451] iii. For example, a specific sample point may be located at the end of a row or column group, e.g. Figure 3B and 3C "A" and "D" shown in the figure.

[0452] iv. For example, a particular sample point may be located in the middle of a row or column group, e.g. Figure 3B and 3C "B" and "C" shown in the figure.

[0453] v. In one example, the motion vector of a particular sample point can be the MV mentioned in item g. b .

[0454] i. The shared motion vector (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) may be set to the motion vector (and / or horizontal component and / or vertical component and / or fractional vertical component and / or fractional vertical component) of a virtual sample located at a different position than all samples in the group.

[0455] i. In one example, a virtual point is not in the group, but it is in a region covering all points in the group.

[0456] 1) Alternatively, the virtual sample point is located outside the region covering all the sample points in the group, for example, near the lower right corner of the region.

[0457] ii. In one example, the MVs of virtual samples are derived in the same way as real samples but at different locations.

[0458] iii. Figures 3A-3C The “V” in shows an example of three virtual sample points.

[0459] j. The shared MV (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) can be set as a function of the MV (and / or horizontal component and / or vertical component and / or fractional vertical component and / or fractional vertical component) of multiple samples and / or virtual samples.

[0460] i. For example, shared MV (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) can be set to all or part of the samples in the group, or Figure 3A Sample points "E", "F", "G", "H" in, or Figure 3A Sample point "E", "H", or Figure 3A Sample points "A", "B", "C", "D" in the Figure 3A Sample points "A", "D", or Figure 3BSample point "B", "C", or Figure 3B Sample points "A", "D", or Figure 3C Sample point "B", "C", or Figure 3C The average value of the MV (and / or horizontal component and / or vertical component and / or vertical component) of the sample points "A" and "D" in .

[0461] 2. It is proposed that when the resolution of the reference picture is different from that of the current picture, or when the width and / or height of the reference picture is larger than that of the current picture, only integer MVs are allowed to perform motion compensation processing to derive the prediction block of the current block.

[0462] a. In one example, the decoded motion vector for the sample to be predicted is rounded to an integer MV before being used.

[0463] b. In one example, the decoded motion vector for the sample to be predicted is rounded to the integer MV that is closest to the decoded motion vector.

[0464] c. In one example, the decoded motion vector for the sample to be predicted is rounded to the integer MV that is closest to the decoded motion vector in the horizontal direction.

[0465] d. In one example, the decoded motion vector for the sample to be predicted is rounded to the integer MV that is closest to the decoded motion vector in the vertical direction.

[0466] 3. The motion vector used in the motion compensation processing of the samples in the current block (for example, the shared MV / shared horizontal or vertical or fractional component / MV′ mentioned in the above items) can be stored in the decoded picture buffer and used for motion vector prediction of subsequent blocks in the current / different picture.

[0467] a. Alternatively, the motion vector used in the motion compensation process of the samples in the current block (for example, the shared MV / shared horizontal or vertical or fractional component / MV′ mentioned in the above items) may not be allowed to be used for motion vector prediction of subsequent blocks in the current / different picture.

[0468] i. In one example, the decoded motion vector (e.g., MV in the above bullet) b ) can be used for motion vector prediction of subsequent blocks in the current / different picture.

[0469] b. In one example, the motion vector used in the motion compensation process for the samples in the current block may be used in a filtering process (eg, deblocking filter / SAO / ALF).

[0470] i. Alternatively, the decoded motion vectors (e.g., MV in the above item) can be used in the filtering process. b ).

[0471] c. In one example, such MVs may be derived at the sub-block level and may be stored for each sub-block.

[0472] 4. It is proposed that an interpolation filter for deriving a prediction block of a current block in a motion compensation process can be selected according to whether the resolution of the reference picture is different from that of the current picture, or whether the width and / or height of the reference picture is larger than the width and / or height of the current picture.

[0473] a. In one example, when condition A is satisfied, an interpolation filter with fewer taps may be applied, where condition A depends on the dimensions of the current picture and / or the reference picture.

[0474] i. In one example, condition A is that the resolution of the reference picture is different from that of the current picture.

[0475] ii. In one example, condition A is that the width and / or height of the reference picture is greater than the width and / or height of the current picture.

[0476] iii. In one example, condition A is W1>a*W2 and / or H1>b*H2, where (W1, H1) represents the width and height of the reference picture, (W2, H2) represents the width and height of the current picture, and a and b are two factors, for example, a=b=1.5.

[0477] iv. In one example, condition A may also depend on whether bidirectional prediction is used.

[0478] v. In one example, a single-tap filter is applied. In other words, unfiltered integer pixels are output as the interpolation result.

[0479] vi. In one example, when the resolution of the reference picture is different from the current picture, a bilinear filter is applied.

[0480] vii. In one example, when the resolution of the reference picture is different from that of the current picture, or the width and / or height of the reference picture is larger than that of the current picture, a 4-tap filter or a 6-tap filter is applied.

[0481] 1) A 6-tap filter can also be used for affine motion compensation.

[0482] 2) A 4-tap filter can also be used for interpolation of chroma samples.

[0483] b. In one example, when the resolution of the reference picture is different from that of the current picture, or the width and / or height of the reference picture is larger than that of the current picture, interpolation is performed using padding samples.

[0484] c. Whether and / or how to apply the method disclosed in item 4 may depend on the color components.

[0485] i. For example, these methods only apply to the luma component.

[0486] d. Whether and / or how to apply the method disclosed in item 4 may depend on the interpolation filtering direction.

[0487] i. For example, these methods are only applicable to horizontal filtering.

[0488] ii. For example, these methods are only applicable to vertical filtering.

[0489] 5. It is proposed to apply a two-stage process for prediction block generation when the resolution of the reference picture is different from that of the current picture, or when the width and / or height of the reference picture is larger than that of the current picture.

[0490] a. In the first stage, a virtual reference block is generated by upsampling or downsampling an area in the reference picture, depending on the width and / or height of the current picture and the reference picture.

[0491] b. In the second stage, prediction samples are generated from the virtual reference blocks by applying interpolation filtering, independent of the width and / or height of the current picture and the reference picture.

[0492] 6. It is proposed that the upper left corner coordinates (xSbInt) of the boundary block used for reference sample filling in some embodiments can be derived according to the width and / or height of the current picture and the reference picture. L ,ySbInt L ) calculation.

[0493] a. In one example, the brightness position in full sample units is modified as follows:

[0494] xInt i =Clip3(xSbInt L -Dx,xSbInt L +sbWidth+Ux,xInt i ),

[0495] yInt i =Clip3(ySbInt L -Dy,ySbInt L +sbHeight+Uy,yInt i ),

[0496] Dx and / or Dy and / or Ux and / or Uy may depend on the width and / or height of the current picture and the reference picture.

[0497] b. In one example, the chroma position in units of full samples is modified as follows:

[0498] xInti=Clip3(xSbInt C -Dx,xSbInt C +sbWidth+Ux,xInti)

[0499] yInti=Clip3(ySbInt C -Dy,ySbInt C +sbHeight+Uy,yInti)

[0500] Dx and / or Dy and / or Ux and / or Uy may depend on the width and / or height of the current picture and the reference picture.

[0501] 7. Instead of storing / using the motion vector of a block based on the same reference picture resolution as the current picture, it is proposed to use the real motion vector that takes into account the resolution difference.

[0502] a. Alternatively, in addition, when a motion vector is used to generate a prediction block, the motion vector does not need to be further modified according to the resolution of the current picture and the reference picture (eg, (refxL, refyL)).

[0503] Interaction between RPR and other codecs

[0504] 8. Whether / how to apply filtering (eg, a deblocking filter) may depend on the resolution of the reference picture and / or the resolution of the current picture.

[0505] a. In one example, the boundary strength (BS) setting in the deblocking filter can take into account the resolution difference in addition to the motion vector difference.

[0506] i. In one example, the scaled motion vector difference according to the resolution of the current picture and the reference picture can be used to determine the boundary strength.

[0507] b. In one example, if the resolution of at least one reference picture of block A is different from (or smaller than or larger than) the resolution of at least one reference picture of block B, the strength of the deblocking filter at the boundary between block A and block B may be set differently (e.g., increased or decreased) compared to a case where the same resolution is used for both blocks.

[0508] c. In one example, if the resolution of at least one reference picture of block A is different from (or smaller than or larger than) the resolution of at least one reference picture of block B, the boundary between block A and block B is marked to be filtered (e.g., BS is set to 2).

[0509] d. In one example, if the resolution of at least one reference picture for block A and / or block B is different from (or smaller than or larger than) the resolution of the current picture, the strength of the deblocking filter at the boundary between block A and block B may be set differently (e.g., increased / decreased) compared to when the reference picture and the current picture use the same resolution.

[0510] e. In one example, if at least one reference picture of at least one of the two blocks has a different resolution than the resolution of the current picture, the boundary between the two blocks is marked to be filtered (eg, BS is set to 2).

[0511] 9. When sub-pictures exist, the conforming bitstream can satisfy that the reference picture must have the same resolution as the current picture.

[0512] a. Alternatively, when the reference picture has a different resolution than the current picture, there must be no sub-picture in the current picture.

[0513] b. Alternatively, for a sub-picture in the current picture, it is not allowed to use a reference picture with a different resolution than the current picture.

[0514] i. Alternatively, or in addition, reference picture management can be invoked to exclude reference pictures with different resolutions.

[0515] 10. In one example, sub-pictures can be defined for pictures with different resolutions (eg, how to divide a picture into multiple sub-pictures).

[0516] In one example, if the reference picture has a different resolution than the current picture, the corresponding sub-picture in the reference picture may be derived by scaling and / or offsetting a sub-picture of the current picture.

[0517] 11. When the resolution of the reference image is different from the resolution of the current image, PROF (prediction refinement using optical flow) can be enabled.

[0518] a. In one example, a set of MVs (denoted as MVs) can be generated for a set of sample points. g ), and can be used for motion compensation as described in item 1. On the other hand, MV (denoted as MV p ), and MV p and MV gThe difference between (eg, corresponding to Δv used in PROF) and the gradient (eg, the spatial gradient of the motion compensated block) can be used to derive prediction refinement.

[0519] b. In one example, MV p Maybe with MV g With different precision. For example, MV p It can be 1 / N pixel (N>0) accuracy, N=32, 64, etc.

[0520] c. In one example, MV g Can have a different precision than the internal MV precision (e.g., 1 / 16 pel).

[0521] d. In one example, prediction refinements are added to the prediction block to generate a refined prediction block.

[0522] e. In one example, this method can be applied in each prediction direction.

[0523] f. In one example, this method can be applied only to the unidirectional prediction case.

[0524] g. In one example, the method can be applied to unidirectional prediction or / and bidirectional prediction.

[0525] h. In one example, this method can be applied only when the reference picture has a different resolution than the current picture.

[0526] 12. It is proposed that when the resolution of the reference picture is different from that of the current picture, the block / sub-block can only use one MV to perform motion compensation processing to derive the prediction block of the current block.

[0527] a. In one example, a unique MV for a block / sub-block may be defined as a function (eg, an average) of all MVs associated with each sample within the block / sub-block.

[0528] b. In one example, the unique MV of a block / sub-block may be defined as a selected MV associated with a selected sample point (eg, a center sample point) within the block / sub-block.

[0529] c. In one example, a 4x4 block or sub-block (eg, 4x1) may use only one MV.

[0530] d. In one example, BIO can be further used to compensate for the precision loss due to block-based motion vectors.

[0531] 13. When the width and / or height of the reference picture is different from the width and / or height of the current picture, a lazy mode without signaling any block-based motion vectors may be applied.

[0532] a. In one example, motion vectors may not be signaled, and the motion compensation process is a pure resolution change case approximating a still picture.

[0533] b. In one example, when the resolution changes, only the motion vector at the picture / slice / brick / CTU level may be signaled and the block may use the motion vector.

[0534] 14. For blocks coded with affine prediction mode and / or non-affine prediction mode, PROF can be used for approximate motion compensation when the width and / or height of the reference picture is different from the width and / or height of the current picture.

[0535] a. In one example, PROF may be enabled when the width and / or height of the reference picture is different from the width and / or height of the current picture.

[0536] b. In one example, a set of affine motions can be generated by combining the indicated motion and resolution scaling and used by PROF.

[0537] 15. When the width and / or height of the reference picture is different from the width and / or height of the current picture, interlaced prediction (eg, as proposed in JVET-K0102) can be used for approximate motion compensation.

[0538] a. In one example, resolution changes (scaling) are represented as affine motion and interlaced motion prediction can be applied.

[0539] 16. LMCS and / or chroma residual scaling may be disabled when the width and / or height of the current picture is different from the width and / or height of the IRAP picture in the same IRAP period.

[0540] a. In one example, when LMCS is disabled, slice level flags such as slice_lmcs_enabled_flag, slice_lmcs_aps_id, and slice_chroma_residual_scale_flag may not be signaled and inferred to be 0.

[0541] b. In one example, when chroma residual scaling is disabled, slice level flags such as slice_chroma_residual_scale_flag may not be signaled and inferred to be 0.

[0542] Constraints on RPR

[0543] 17. RPR can be applied to codec blocks with block dimension constraints.

[0544] a. In one example, for an M×N encoding / decoding block, where M is the block width and N is the block height, when M×N < T or M×N <= T (e.g., T = 256), RPR may not be used.

[0545] b. In one example, when M < K (or M <= K) (e.g., K = 16) and / or N < L (or N <= L) (e.g., L = 16), RPR may not be used.

[0546] 18. Bitstream consistency can be added to limit the ratio between the width and / or height of an active reference picture (or its consistency window) and the width and / or height of the current picture (or its consistency window). Assume refPicW and refPicH represent the width and height of the reference picture, and curPicW and curPicH represent the width and height of the current picture.

[0547] a. In one example, when (refPicW ÷ curPicW) equals an integer, the reference picture can be marked as an active reference picture.

[0548] i. Alternatively, when (refPicW ÷ curPicW) equals a fraction, the reference picture can be marked as unavailable.

[0549] b. In one example, when (refPicW ÷ curPicW) equals (X * n), where X represents a fraction, e.g., X = 1 / 2, and n represents an integer, e.g., n = 1, 2, 3, 4..., the reference picture can be marked as an active reference picture.

[0550] i. In one example, when (refPicW ÷ curPicW) does not equal (X * n), the reference picture can be marked as unavailable.

[0551] 19. Whether and / or how to enable encoding / decoding tools (e.g., bidirectional prediction / whole triangle prediction mode (TPM) / hybrid processing in TPM) for an M×N block can depend on the resolution of the reference picture (or its consistency window) and / or the current picture (or its consistency window).

[0552] a. In one example, M * N < T or M * N <= T (e.g., T = 64).

[0553] b. In one example, M < K (or M <= K) (e.g., K = 16) and / or N < L (or N <= L) (e.g., L = 16).

[0554] c. In one example, when the width / height of at least one reference picture is different from the current picture, the use of encoding / decoding tools is not allowed.

[0555] i. In one example, when the width / height of at least one reference picture of the block is larger than the width / height of the current picture, the codec tool is not allowed to be used.

[0556] d. In one example, when the width / height of each reference picture of the block is different from the width / height of the current picture, the codec tool is not allowed to be used,

[0557] i. In one example, when the width / height of each reference picture is larger than the width / height of the current picture, the codec tool is not allowed to be used.

[0558] e. Alternatively, and additionally, when codec tools are not enabled, one can use one MV for unidirectional prediction for motion compensation.

[0559] Consistency Window Related

[0560] 20. Consistent cropping window offset parameters (e.g., conf_win_left_offset) are signaled with N pixel precision instead of 1 pixel, where N is a positive integer greater than 1.

[0561] a. In one example, the actual offset can be derived as the signaled offset multiplied by N.

[0562] b. In one example, N is set to 4 or 8.

[0563] 21. It is proposed that the consistency cropping window offset parameter is not only applied to the output. The specific internal decoding process may depend on the cropped picture size (i.e., the resolution of the consistency window in the picture).

[0564] 22. It is proposed that when the width and / or height of the picture represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are the same, the consistency cropping window offset parameters in the first video unit (e.g., PPS) and the second video unit may be different.

[0565] 23. It is proposed that when the width and / or height of the picture represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are different, the consistency cropping window offset parameters in the first video unit (e.g., PPS) and the second video unit should be the same in the consistency bitstream.

[0566] a. It is proposed that the consistency cropping window offset parameters in the first video unit (e.g., PPS) and the second video unit should be the same in the consistency bitstream regardless of whether the width and / or height of the picture represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are the same.

[0567] 24. Assume that the width and height of the consistency window defined in the first video unit (e.g., PPS) are denoted as W1 and H1, respectively. The width and height of the consistency window defined in the second video unit (e.g., PPS) are denoted as W2 and H2, respectively. The upper left position of the consistency window defined in the first video unit (e.g., PPS) is denoted as X1 and Y1. The upper left position of the consistency window defined in the second video unit (e.g., PPS) is denoted as X2 and Y2. The width and height (e.g., pic_width_in_luma_samples and pic_height_in_luma_samples) of the encoded / decoded picture defined in the first video unit (e.g., PPS) are denoted as PW1 and PH1, respectively. The width and height of the encoded / decoded picture defined in the second video unit (e.g., PPS) are denoted as PW2 and PH2.

[0568] a. In one example, in a conforming bitstream, W1 / W2 should be equal to X1 / X2.

[0569] i. Alternatively, in a conforming bitstream, W1 / X1 shall equal W2 / X2.

[0570] ii. Alternatively, in a conforming bitstream, W1*X2 shall be equal to W2*X1.

[0571] b. In one example, in a conforming bitstream, H1 / H2 should be equal to Y1 / Y2.

[0572] i. Alternatively, in a conforming bitstream, H1 / Y1 shall equal H2 / Y2.

[0573] ii. Alternatively, in a conforming bitstream, H1×Y2 shall be equal to H2×Y1.

[0574] c. In one example, in a conforming bitstream, PW1 / PW2 should be equal to X1 / X2.

[0575] i. Alternatively, in a conforming bitstream, PW1 / X1 shall be equal to PW2 / X2.

[0576] ii. Alternatively, in the conforming bitstream, PW1×X2 shall be equal to PW2×X1.

[0577] d. In one example, in a conforming bitstream, PH1 / PH2 should be equal to Y1 / Y2.

[0578] i. Alternatively, in a conforming bitstream, PH1 / Y1 shall be equal to PH2 / Y2.

[0579] ii. Alternatively, in a conforming bitstream, PH1*Y2 shall be equal to PH2*Y1.

[0580] e. In one example, in a conforming bitstream, PW1 / PW2 should be equal to W1 / W2.

[0581] i. Alternatively, in a conforming bitstream, PW1 / W1 shall equal PW2 / W2.

[0582] ii. Alternatively, in a conforming bitstream, PW1*W2 shall be equal to PW2*W1.

[0583] f. In one example, in a conforming bitstream, PH1 / PH2 should be equal to H1 / H2.

[0584] i. Alternatively, in a conforming bitstream, PH1 / H1 shall be equal to PH2 / H2.

[0585] ii. Alternatively, in a conforming bitstream, PH1*H2 shall be equal to PH2*H1.

[0586] g. In a conforming bitstream, if PW1 is greater than PW2, then W1 must be greater than W2.

[0587] h. In a conforming bitstream, if PW1 is less than PW2, then W1 must be less than W2.

[0588] i. In a conforming bitstream, (PW1-PW2)*(W1-W2) shall not be less than 0.

[0589] j. In a conforming bitstream, if PH1 is greater than PH2, then H1 must be greater than H2.

[0590] k. In a conforming bitstream, if PH1 is less than PH2, then H1 must be less than H2.

[0591] l. In a conformant bitstream, (PH1-PH2)*(H1-H2) shall not be less than 0.

[0592] m. In a conforming bitstream, if PW1>=PW2, then W1 / W2 shall not be greater than (or less than) PW1 / PW2.

[0593] n. In a conforming bitstream, if PH1>=PH2, H1 / H2 shall not be greater than (or less than) PH1 / PH2.

[0594] 25. Assume that the width and height of the consistency window of the current picture are denoted as W and H respectively. The width and height of the consistency window of the reference picture are denoted as W' and H' respectively. Then the conforming bitstream shall meet at least one of the following constraints.

[0595] aW*pw>=W'; pw is an integer, for example, 2.

[0596] bW*pw>W'; pw is an integer, such as 2.

[0597] c. W'*pw'>=W; pw' is an integer, for example 8.

[0598] d.W'*pw'>W; pw' is an integer, such as 8.

[0599] eH*ph>=H'; ph is an integer, for example, 2.

[0600] fH*ph>H'; ph is an integer, such as 2.

[0601] g.H'*ph'>=H; ph' is an integer, for example 8.

[0602] h.h'*ph'>h; ph' is an integer, such as 8.

[0603] i. In one example, pw is equal to pw'.

[0604] j. In one example, ph is equal to ph'.

[0605] k. In one example, pw is equal to ph.

[0606] 1. In one example, pw' is equal to ph'.

[0607] In one example, when W and H represent the width and height of the current picture, respectively, the conforming bitstream may be required to satisfy the above sub-items. W' and H' represent the width and height of the reference picture, respectively.

[0608] 26. It is proposed that the consistency window parameter be partially signaled.

[0609] a. In one example, the upper left sample point in the consistency window of the picture is the same as the upper left sample point in the picture.

[0610] b. For example, conf_win_left_offset defined in VVC is not signaled and is inferred to be zero.

[0611] c. For example, conf_win_top_offset defined in VVC is not signaled and is inferred to be zero.

[0612] 27. The derivation of the position of the reference sample points (e.g., (refx L , refy L )) defined in VVC) can depend on the upper - left position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_left_offset, conf_win_top_offset defined in VVC). FIG. 4 shows an example of the sample point positions derived in VVC (a) and the proposed method (b). The dashed rectangle represents the consistency window.

[0613] a. In one example, there is a correlation only when the width and / or height of the current picture is different from the width and / or height of the reference picture.

[0614] b. In one example, the derivation of the horizontal position of the reference sample point (e.g., refx L ) defined in VVC) can depend on the left position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_left_offset defined in VVC).

[0615] i. In one example, calculate the horizontal position of the current sample point relative to the upper - left position of the consistency window in the current picture (denoted as xSb'), and use it to derive the position of the reference sample point.

[0616] 1) For example, calculate xSb' = xSb – (conf_win_left_offset << Prec) and use it to derive the position of the reference sample point, where xSb represents the horizontal position of the current sample point in the current picture. conf_win_left_offset represents the horizontal position of the upper - left sample point in the consistency window of the current picture. Prec represents the precision of xSb and xSb', where (xSb >> Prec) can represent the actual horizontal coordinate of the current sample point relative to the current picture. For example, Prec = 0 or Prec = 4.

[0617] ii. In one example, calculate the horizontal position of the reference sample point relative to the upper - left position of the consistency window in the reference picture (denoted as Rx’).

[0618] 1) The calculation of Rx' can depend on xSb' and / or the motion vector and / or the resampling rate.

[0619] iii. In one example, calculate the horizontal position of the reference sample point relative to the reference picture based on Rx' (denoted as Rx).

[0620] 1) For example, calculate Rx = Rx'+(conf_win_left_offset_ref << Prec), where conf_win_left_offset_ref represents the horizontal position of the top-left sample in the consistency window of the reference picture. Prec represents the precision of Rx and Rx'. For example, Prec = 0 or Prec = 4.

[0621] iv. In one example, Rx can be directly calculated based on xSb′ and / or the motion vector and / or the resampling rate. In other words, the two derivation steps of Rx' and Rx are combined into one calculation step.

[0622] v. Whether and / or how to use the left position of the consistency window of the current picture and / or the reference picture (such as conf_win_left_offset defined in VVC) may depend on the color component and / or the color format.

[0623] 1) For example, conf_win_left_offset can be modified to conf_win_left_offset = conf_win_left_offset * SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example, for the luminance component, SubWidthC is equal to 1. When the color format is 4:2:0 or 4:2:2, SubWidthC of the chrominance component is equal to 2.

[0624] 2) For example, conf_win_left_offset can be modified to conf_win_left_offset = conf_win_left_offset / SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example., for the luminance component, SubWidthC is equal to 1. When the color format is 4:2:0 or 4:2:2, SubWidthC of the chrominance component is equal to 2.

[0625] c. In one example, the derivation of the vertical position of the reference sample (e.g., refy defined in VVC) L can depend on the top position of the consistency window of the current picture and / or the reference picture (e.g., conf_win_top_offset defined in VVC).

[0626] i. In one example, calculate the vertical position of the current sample relative to the top-left position of the consistency window in the current picture (denoted as ySb'), and use it to derive the position of the reference sample.

[0627] 1) For example, calculate ySb' = ySb – (conf_win_top_offset << Prec) and use it to derive the position of the reference sample, where ySb represents the vertical position of the current sample in the current picture. conf_win_top_offset represents the vertical position of the top-left sample in the consistency window of the current picture. Prec represents the precision of ySb and ySb'. For example, Prec = 0 or Prec = 4.

[0628] ii. In one example, calculate the vertical position (denoted as Ry') of the reference sample relative to the top-left position of the consistency window in the reference picture.

[0629] 1) The calculation of Ry' can depend on ySb' and / or the motion vector and / or the resampling rate.

[0630] iii. In one example, calculate the vertical position (denoted as Ry) of the reference sample relative to the reference picture according to Ry'.

[0631] 1) For example, calculate Ry = Ry' + (conf_win_top_offset_ref << Prec), where conf_win_top_offset_ref represents the vertical position of the top-left sample in the consistency window of the reference picture. Prec represents the precision of Ry and Ry'. For example, Prec = 0 or Prec = 4.

[0632] iv. In one example, Ry can be directly calculated according to ySb′ and / or the motion vector and / or the resampling rate. In other words, the two derivation steps of Ry' and Ry are combined into one calculation step.

[0633] v. Whether and / or how to use the top position of the consistency window of the current picture and / or the reference picture (such as conf_win_top_offset defined in VVC) can depend on the color component and / or the color format.

[0634] 1) For example, conf_win_top_offset can be modified to conf_win_top_offset = conf_win_top_offset * SubHeightC, where SubHeightC defines the vertical sampling step of the color component. For example, for the luminance component, SubHeightC is equal to 1. When the color format is 4:2:0, SubHeightC of the chrominance component is equal to 2.

[0635] 2) For example, conf_win_top_offset can be modified to conf_win_top_offset = conf_win_top_offset / SubHeightC, where SubHeightC defines the vertical sample step size of the color component. For example, for the luma component, SubHeightC is equal to 1. When the color format is 4:2:0, SubHeightC for the chroma components is equal to 2.

[0636] 28. It is proposed to clip the integer part of the horizontal coordinate of the reference sample to [minW, maxW]. Assume that the width and height of the consistency window of the reference picture are denoted by W and H respectively. The width and height of the consistency window of the reference picture are denoted by W' and H' respectively. The top left corner position of the consistency window in the reference picture is denoted by (x0, y0).

[0637] a. In one example, minW is equal to 0.

[0638] b. In one example, minW is equal to X0.

[0639] c. In one example, maxW is equal to W-1.

[0640] d. In one example, maxW is equal to W'-1.

[0641] e. In one example, maxW is equal to X0+W'-1.

[0642] f. In one example, minW and / or maxW may be modified based on the color format and / or color components.

[0643] i. For example, minW is modified to minW*SubC.

[0644] ii. For example, change minW to minW / SubC.

[0645] iii. For example, maxW is modified to maxW*SubC.

[0646] iv. For example, maxW is changed to maxW / SubC.

[0647] v. In one example, for the luma component, SubC is equal to 1.

[0648] vi. In one example, when the color format is 4:2:0, SubC is equal to 2 for the chroma components.

[0649] vii. In one example, when the color format is 4:2:2, SubC is equal to 2 for the chroma components.

[0650] viii. In one example, when the color format is 4:4:4, SubC is equal to 1 for the chroma components.

[0651] g. In one example, whether and / or how cropping is performed may depend on the dimensions of the current picture (or the consistency window therein) and the dimensions of the reference picture (or the consistency window therein).

[0652] i. In one example, cropping is performed only when the dimensions of the current picture (or the consistency window therein) are different from the dimensions of the reference picture (or the consistency window therein).

[0653] 29. It is proposed to clip the integer part of the vertical coordinate of the reference sample to [minH, maxH]. Assume that the width and height of the consistency window of the reference picture are denoted by W and H respectively. The width and height of the consistency window of the reference picture are denoted by W' and H' respectively. The upper left corner position of the consistency window in the reference picture is denoted by (X0, Y0).

[0654] a. In one example, minH is equal to 0.

[0655] b. In one example, minH is equal to Y0.

[0656] c. In one example, maxH is equal to H-1.

[0657] d. In one example, maxH is equal to H'-1.

[0658] e. In one example, maxH is equal to Y0+H'-1.

[0659] f. In one example, minH and / or maxH may be modified based on the color format and / or color components.

[0660] i. For example, minH is modified to minH*SubC.

[0661] ii. For example, minH is modified to minH / SubC.

[0662] iii. For example, maxH is modified to maxH*SubC.

[0663] iv. For example, maxH is modified to maxH / SubC.

[0664] v. In one example, for the luma component, SubC is equal to 1.

[0665] vi. In one example, when the color format is 4:2:0, SubC is equal to 2 for the chroma components.

[0666] vii. In one example, when the color format is 4:2:2, SubC is equal to 1 for the chroma components.

[0667] viii. In one example, when the color format is 4:4:4, SubC is equal to 1 for the chroma components.

[0668] g. In one example, whether and / or how to perform cropping may depend on the dimensions of the current picture (or the consistency window therein) and the dimensions of the reference picture (or the consistency window therein).

[0669] i. In one example, cropping is performed only when the dimensions of the current picture (or the consistency window therein) are different from the dimensions of the reference picture (or the consistency window therein).

[0670] In the following discussion, a first syntax element is asserted to "correspond" to a second syntax element if the two syntax elements have equivalent functionality but may be signaled at different video units (e.g., VPS / SPS / PPS / slice header / picture header, etc.).

[0671] 30. It is proposed that a syntax element may be signaled in a first video unit (eg, picture header or PPS) while the corresponding syntax element is not signaled in a second video unit at a higher level (eg, SPS) or a lower level (eg, slice header).

[0672] a. Alternatively, a first syntax element may be signaled in a first video unit (eg, a picture header or PPS), and a corresponding second syntax element may be signaled in a lower-level second video unit (eg, a slice header).

[0673] i. Optionally, an indicator may be signaled in the second video unit to notify whether the second syntax element is signaled thereafter.

[0674] ii. In one example, if a second syntax element is signaled, the slice associated with the second video unit (eg, a slice header) may follow the indication of the second syntax element instead of the first syntax element.

[0675] iii. An indicator associated with the first syntax element may be signaled in the first video unit to inform whether the second syntax element is signaled in any slice (or other video unit) associated with the first video unit.

[0676] b. Optionally, the first syntax element may be signaled in a higher-level first video unit (eg, VPS / SPS / PPS), and the corresponding second syntax element may be signaled in a second video unit (eg, picture header).

[0677] i. Alternatively, an indicator may be signaled to inform whether the second syntax element is signaled thereafter.

[0678] ii. In one example, if a second syntax element is signaled, the pictures associated with the second video unit (which may be divided into slices) may follow the indication of the second syntax element instead of the first syntax element.

[0679] c. The first syntax element in the picture header may have the equivalent functionality to the second syntax element in the slice header, as defined in Section 2.6 (e.g., but not limited to ...) but controls all strips of the picture.

[0680] d. The first syntax element in the SPS as defined in Section 2.6 may have equivalent functionality to the second syntax element in the picture header (e.g., but not limited to ...) but only controls the associated picture (which may be split into strips).

[0681] e. The first syntax element in the PPS as defined in Section 2.7 may have equivalent functionality to the second syntax element in the picture header (e.g., but not limited to ...) but only controls the associated picture (which may be split into strips).

[0682] 31. The syntax elements signaled in the picture header are separated from other syntax elements signaled or derived in the SPS / VPS / DPS.

[0683] 32. The indication of DMVR and BDOF enable / disable can be signaled separately in the picture header instead of by the same flag (e.g. )control.

[0684] 33. In ALF, the enablement / disablement of PROF / cross-component ALF / inter-prediction using geometric partitioning (GEO) can be signaled in the picture header.

[0685] a. Alternatively, the indication of enabling / disabling PROF can be conditionally signaled in the picture header according to the PROF enable flag in the SPS.

[0686] b. Alternatively, the indication of enabling / disabling cross-component ALF (CCALF) may be conditionally signaled in the picture header according to the CCALF enable flag in the SPS.

[0687] c. Alternatively, the indication of enabling / disabling GEO can be conditionally signaled in the picture header according to the GEO enable flag in the SPS.

[0688] d. Alternatively, and in addition, the indication of enabling / disabling PROF / cross-component ALF / inter prediction using geometric partitioning (GEO) may be conditionally signaled in the slice header, depending on which syntax elements are signaled in the picture header instead of the SPS.

[0689] 34. The indication of the prediction type for slices / tiles / slices (or other video units smaller than a picture) in the same picture may be signaled in the picture header.

[0690] a. In one example, an indication of whether all slices / tiles / slices (or other video units smaller than a picture) are intra-coded (eg, all I slices) may be signaled in the picture header.

[0691] i. Alternatively, and additionally, if the indication tells that all slices in the picture are I slices, the slice type may not be signaled in the slice header.

[0692] b. Optionally, an indication of whether at least one of the slices / bricks / slices (or other video units smaller than a picture) is not intra-coded (eg, at least one non-I slice) may be signaled in the picture header.

[0693] c. Optionally, an indication of whether all slices / tiles / slices (or other video units smaller than a picture) have the same prediction type (eg, I / P / B slices) can be signaled in the picture header.

[0694] i. Alternatively, further, the slice type may not be signaled in the slice header.

[0695] ii. Alternatively, in addition, an indication of which tools are allowed for a particular prediction type (eg, DMVR / BDOF / TPM / GEO only for B-slices; dual-tree only for I-slices) may be conditionally signaled based on the indication of the prediction type.

[0696] d. Alternatively, or in addition, the signaling of the indication of enabling / disabling the tool may depend on the indication of the prediction type mentioned in the above sub-item.

[0697] i. Alternatively, or in addition, the indication for enabling / disabling the tool may be derived based on the indication of the prediction type mentioned in the above sub-item.

[0698] 35. In the present disclosure (items 1 to 29), the term "consistency window" may be replaced with other terms, such as "scaling window." The scaling window may be signaled differently from the consistency window and used to derive a scaling ratio and / or an upper left offset for deriving the reference sample position for RPR.

[0699] a. In one example, the scaling window can be constrained by the consistency window. For example, in a conforming bitstream, the scaling window must be contained within the consistency window.

[0700] 36. Whether and / or how the allowed maximum block size of a transform skip codec block is signaled may depend on the maximum block size of a transform codec block.

[0701] a. Alternatively, the maximum block size of a transform skip codec block cannot be larger than the maximum block size of a transform codec block in the conforming bitstream.

[0702] 37. Whether and how to signal the indication of enabling joint Cb Cr residual (JCCR) codec (eg, sps_joint_cbcr_enabled_flag) may depend on the color format (eg, 4:0:0, 4:2:0, etc.).

[0703] a. For example, if the color format is 4:0:0, the indication of enabling the joint Cb Cr residual (JCCR) may not be signaled. The syntax design of the example is as follows:

[0704]

[0705] Downsampling filter type used for chroma blending mask generation in TPM / GEO

[0706] 38. The downsampling filter type used for mixing weight derivation of chroma samples can be signaled at the video unit level (e.g., SPS / VPS / PPS / picture header / sub-picture / slice / slice header / slice / brick / CTU / VPDU level).

[0707] a. In one example, a high-level flag may be signaled to switch between different chroma format types for content.

[0708] i. In one example, a high level flag may be signaled to switch between chroma format type 0 and chroma format type 2.

[0709] ii. In one example, a flag may be signaled to specify whether the left upper and lower sample luma weights in TPM / GEO prediction mode are co-located with the left upper luma weight (ie, chroma sample location type 0).

[0710] iii. In one example, a flag may be signaled to specify whether the luma samples of the left upper and lower samples in TPM / GEO prediction mode are horizontally co-located with the upper left luma sample but vertically offset by 0.5 units relative to the upper left luma sample (i.e., chroma sample position type 2).

[0711] b. In one example, the type of downsampling filter may be signaled for 4:2:0 chroma format and / or 4:2:2 chroma format.

[0712] c. In one example, a flag may be signaled to specify the type of chroma downsampling filter used for TPM / GEO prediction.

[0713] i. In one example, a flag may be signaled to indicate whether downsampling filter A or downsampling filter B is used to derive chroma weights in TPM / GEO prediction mode.

[0714] 39. The downsampling filter type for mixing weight derivation of chroma samples can be derived at the video unit level (e.g., SPS / VPS / PPS / picture header / sub-picture / slice / slice header / slice / brick / CTU / VPDU level).

[0715] a. In one example, a lookup table may be defined to specify the correspondence between chroma subsampling filter types and chroma format types of the content.

[0716] 40. In case of different chroma sample location types, the specified downsampling filter can be used in TPM / GEO prediction mode.

[0717] a. In one example, in the case of a specific chroma sample location type (eg, chroma sample location type 0), the chroma weights of the TPM / GEO may be subsampled from the co-located top-left luma weight.

[0718] b. In one example, under certain chroma sample position types (e.g., chroma sample position type 0 or 2), a specified X-tap filter (X is a constant, e.g., X=6 or 5) may be used for chroma weight subsampling in TPM / GEO prediction mode.

[0719] 41. In a video unit (eg, SPS, PPS, picture header, slice header, etc.), a first syntax element (eg, a flag) may be signaled to indicate whether multiple transform selection (MTS) is disabled for all blocks (slices / pictures).

[0720] a. A second syntax element indicating how MTS is applied to intra-coded blocks (slices / pictures) (e.g., MTS enabled / disabled / implicit / explicit) is conditionally signaled based on the first syntax element. For example, the second syntax element is signaled only when the first syntax element indicates that MTS is not disabled for all blocks (slices / pictures).

[0721] b. A third syntax element indicating how MTS is applied to inter-coded blocks (slices / pictures) (e.g., MTS enabled / disabled / implicit / explicit) is conditionally signaled based on the first syntax element. For example, the third syntax element is signaled only when the first syntax element indicates that MTS is not disabled for all blocks (slices / pictures).

[0722] c. The syntax design of the example is as follows

[0723]

[0724] d. The third syntax element may be conditionally signaled based on whether sub-block transform (SBT) is applied.

[0725] The example grammar is designed as follows

[0726]

[0727]

[0728] e. The example grammar design is as follows

[0729]

[0730] 5. Additional Embodiments

[0731] Below, text changes are shown in underlined bold italic font.

[0732] 5.1 Implementation of Constraints on the Consistency Window

[0733] and The picture samples in the CVS output from the decoding process are specified in a rectangular area specified in picture coordinates for output. When conformance_window_flag is equal to 0, the values ​​of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset and conf_win_bottom_offset are inferred to be 0.

[0734] The conforming cropping window contains the luma samples with horizontal picture coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical picture coordinates from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1), inclusive.

[0735] The value of SubWidthC*(conf_win_left_offset+conf_win_right_offset) should be less than pic_width_in_luma_samples, and the value of SubHeightC*(conf_win_top_offset+conf_win_bottom_offset) should be less than pic_height_in_luma_samples.

[0736] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:

[0737] PicOutputWidthL=pic_width_in_luma_samples- (7-43)

[0738] SubWidthC*(conf_win_right_offset+conf_win_left_offset)

[0739] PicOutputHeightL=pic_height_in_pic_size_units

[0740] -SubHeightC*(conf_win_bottom_offset+conf_win_top_offset)

[0741] (7-44)

[0742] When ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma arrays are samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.

[0743]

[0744] 5.2 Example 1 of Derivation of Reference Sample Point Positions

[0745] 8.5.6.3.1 Overview

[0746]

[0747] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0748] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0749]

[0750] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0751] – If cIdx is equal to 0, the following applies:

[0752] –The scaling factor and its fixed-point representation are defined as

[0753] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0754] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0755] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0756] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0757] –For each luma sample position in the predicted luma sample array predSamplesLX

[0758] (x L=0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0759] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0760]

[0761] refy L =((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)

[0762]

[0763] –Variable xInt L ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0764] xInt L =refx L >>4 (8-759)

[0765] yInt L =refy L >>4 (8-760)

[0766] xFrac L =refx L &15 (8-761)

[0767] yFrac L =refy L &15 (8-762)

[0768] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0769] –x L Equal to 0.

[0770] –x L Equal to sbWidth+1.

[0771] –y L Equal to 0.

[0772] –y L Equal to sbHeight+1.

[0773] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L ),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0774] – Otherwise (cIdx is not equal to 0), the following applies:

[0775] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0776] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0777] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0778] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0779] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0780]

[0781] refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)

[0782]

[0783] refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0784]

[0785] –Variable xInt C ,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0786] xInt C =refx C >>5 (8-767)

[0787] yInt C =refy C >>5 (8-768)

[0788] xFrac C =refy C &31 (8-769)

[0789] yFrac C =refy C &31 (8-770)

[0790] 5.3 Example 2 of Derivation of Reference Sample Point Positions

[0791] 8.5.6.3.1 Overview

[0792]

[0793] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0794] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0795]

[0796] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0797] – If cIdx is equal to 0, the following applies:

[0798] –The scaling factor and its fixed-point representation are defined as

[0799] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0800] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0801] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0802] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0803] – For each luma sample position (x L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0804] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0805]

[0806]

[0807] –Variable xInt L ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0808] xInt L =refx L >>4 (8-759)

[0809] yInt L =refy L >>4 (8-760)

[0810] xFrac L =refx L &15 (8-761)

[0811] yFrac L =refy L &15 (8-762)

[0812] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0813] –x L Equal to 0.

[0814] –x L Equal to sbWidth+1.

[0815] –y L Equal to 0.

[0816] –y L Equal to sbHeight+1.

[0817] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L ),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0818] – Otherwise (cIdx is not equal to 0), the following applies:

[0819] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0820] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0821] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0822] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0823] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0824]

[0825] refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)

[0826]

[0827] refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0828]

[0829] –Variable xInt C ,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0830] xInt C =refx C >>5 (8-767)

[0831] yInt C =refyC>>5 (8-768)

[0832] xFrac C =refy C &31 (8-769)

[0833] yFrac C =refyC&31 (8-770)

[0834] 5.4 Example 3 of Derivation of Reference Sample Point Positions

[0835] 8.5.6.3.1 Overview

[0836]

[0837] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0838] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0839]

[0840] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0841] – If cIdx is equal to 0, the following applies:

[0842] –The scaling factor and its fixed-point representation are defined as

[0843] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0844] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0845] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0846] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0847] – For each luma sample position (x L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0848] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0849]

[0850] –Variable xInt L ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0851] xInt L =refx L >>4 (8-759)

[0852] yInt L =refy L >>4 (8-760)

[0853] xFrac L =refx L &15 (8-761)

[0854] yFrac L =refy L &15 (8-762)

[0855] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0856] –x L Equal to 0.

[0857] –x L Equal to sbWidth+1.

[0858] –y L Equal to 0.

[0859] –y L Equal to sbHeight+1.

[0860] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L ),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0861] – Otherwise (cIdx is not equal to 0), the following applies:

[0862] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0863] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0864] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0865] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0866] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0867]

[0868]

[0869] –Variable xInt C ,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0870] xInt C =refx C >>5 (8-767)

[0871] yInt C =refy C >>5 (8-768)

[0872] xFrac C =refy C &31 (8-769)

[0873] yFrac C =refy C &31 (8-770)

[0874] 5.5 Example 1 of Reference Sample Point Position Cropping

[0875] 8.5.6.3.1 Overview

[0876] The inputs to this process include:

[0877] – Luma position (xSb, ySb), specifies the upper left sample of the current codec sub-block relative to the upper left luminance sample of the current picture,

[0878] –Variable sbWidth specifies the width of the current codec sub-block.

[0879] –Variable sbHeight specifies the height of the current codec sub-block.

[0880] – motion vector offset mvOffset,

[0881] – refined motion vector refMvLX,

[0882] – the selected reference picture sample array refPicLX,

[0883] – Half-sample interpolation filter index hpelIfIdx,

[0884] – bidirectional optical flow flag bdofFlag,

[0885] –Variable cIdx specifies the color component index of the current block.

[0886] The output of this processing includes:

[0887] – predSamplesLX array of (sbWidth+brdExtSize)x(sbHeight+brdExtSize) prediction sample values.

[0888] The prediction block boundary extension size brdExtSize is derived as follows:

[0889] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&sps_affine_prof_enabled_flag))? 2:0 (8-752)

[0890] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0891] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0892] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0893] – If cIdx is equal to 0, the following applies:

[0894] –The scaling factor and its fixed-point representation are defined as

[0895] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0896] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0897] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0898] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0899] – For each luma sample position (x L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0900] –Assume (refxSb L ,refySbL ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0901] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)

[0902] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)

[0903] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757)

[0904] refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)

[0905] –Variable xInt L ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0906]

[0907] xFrac L =refx L &15 (8-761)

[0908] yFrac L =refy L &15 (8-762)

[0909] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0910] –x L Equal to 0.

[0911] –x L Equal to sbWidth+1.

[0912] –y L Equal to 0.

[0913] –y L Equal to sbHeight+1.

[0914] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L ),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0915] – Otherwise (cIdx is not equal to 0), the following applies:

[0916] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0917] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0918] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0919] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0920] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0921] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)

[0922] refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)

[0923] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765)

[0924] refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0925] –Variable xInt C,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0926]

[0927] xFrac C =refy C &31 (8-769)

[0928] yFrac C =refy C &31 (8-770)

[0929] – The predicted sample values ​​predSamplesLX[xC][yC] are derived by invoking the process specified in clause 8.5.6.3.4, with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC), sbWidth, sbHeight and refPicLX as input.

[0930] 5.6 Example 2 of Reference Sample Point Position Cropping

[0931] 8.5.6.3.1 Overview

[0932] The inputs to this process include:

[0933] – Luma position (xSb, ySb), specifies the upper left sample of the current codec sub-block relative to the upper left luminance sample of the current picture,

[0934] –Variable sbWidth specifies the width of the current codec sub-block.

[0935] –Variable sbHeight specifies the height of the current codec sub-block.

[0936] – motion vector offset mvOffset,

[0937] – refined motion vector refMvLX,

[0938] – the selected reference picture sample array refPicLX,

[0939] – Half-sample interpolation filter index hpelIfIdx,

[0940] – bidirectional optical flow flag bdofFlag,

[0941] –Variable cIdx specifies the color component index of the current block.

[0942] The output of this processing includes:

[0943] – predSamplesLX array of (sbWidth+brdExtSize)x(sbHeight+brdExtSize) prediction sample values.

[0944] The prediction block boundary extension size brdExtSize is derived as follows:

[0945] brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb]&&sps_affine_prof_enabled_flag))? 2:0 (8-752)

[0946] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in units of luma samples.

[0947] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in units of luma samples.

[0948] The variable fRefLeftOff is set equal to the conf_win_left_ offset.

[0949] The variable fRefTopOff is set equal to the conf_win_top_ offset.

[0950] The motion vector mvLX is set equal to (refMvLX-mvOffset).

[0951] – If cIdx is equal to 0, the following applies:

[0952] –The scaling factor and its fixed-point representation are defined as

[0953] hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL(8-753)

[0954] vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754)

[0955] – Let (xIntL, yIntL) be the luma position given in full sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are then used in this clause only to specify fractional sample positions within the reference sample array refPicLX.

[0956] – The upper left corner coordinate of the boundary block used for reference sample filling (xSbInt L ,ySbInt L ) is set equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)).

[0957] – For each luma sample position (x L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ]The derivation is as follows:

[0958] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. The variable refxSb L ,refx L ,refySb L , and refy L The derivation is as follows:

[0959] refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755)

[0960] refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756)

[0961] refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757)

[0962] refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)

[0963] –Variable xIntL ,yInt L ,xFrac L and yFrac L The derivation is as follows:

[0964]

[0965] xFrac L =refx L &15 (8-761)

[0966] yFrac L =refy L &15 (8-762)

[0967] – If bdofFlag is equal to true or (sps_affine_prof_enabled_flag is equal to true and inter_affine_flag[xSb][ySb] is equal to true), and one or more of the following conditions are true, the predicted luma sample value predSamplesLX[xSb][ySb] is L ][y L ] is obtained by calling the luminance integer sample acquisition process specified in clause 8.5.6.3.3, with (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX as input.

[0968] –x L Equal to 0.

[0969] –x L Equal to sbWidth+1.

[0970] –y L Equal to 0.

[0971] –y L Equal to sbHeight+1.

[0972] – Otherwise, the predicted luma sample values ​​predSamplesLX[xL][yL] are derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2, with the values ​​(xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)),(xFracL,yFracL),(xSbInt L ,ySbInt L),refPicLX,hpelIfIdx,sbWidth,sbHeight and (xSb,ySb) as input.

[0973] – Otherwise (cIdx is not equal to 0), the following applies:

[0974] – Let (xIntC, yIntC) be the chroma position given in full sample units and (xFracC, yFracC) be the offset given in 1 / 32 sample units.

[0975] These variables are used in this section only to specify general fractional sample locations in the reference sample array refPicLX.

[0976] – The upper left corner coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample filling are set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).

[0977] – For each chroma sample position in the predicted chroma sample array predSamplesLX (xC = 0..sbWidth-1, yC = 0..sbHeight-1), the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:

[0978] –Assume (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variable refxSb C ,refySb C ,refx C and refy C The derivation is as follows:

[0979] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)

[0980] refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)

[0981] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765)

[0982] refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)

[0983] –Variable xInt C ,yInt C ,xFrac C and yFrac C The derivation is as follows:

[0984]

[0985] xFrac C =refy C &31 (8-769)

[0986] yFrac C =refy C &31 (8-770)

[0987] – The predicted sample values ​​predSamplesLX[xC][yC] are derived by invoking the process specified in clause 8.5.6.3.4, with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC), sbWidth, sbHeight and refPicLX as input.

[0988] 6. Example Implementations of Disclosed Technologies

[0989] Figure 5is a block diagram of a video processing device 500. The device 500 can be used to implement one or more methods described herein. The device 500 can be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 500 may include one or more processors 502, one or more memories 504, and video processing hardware 506. The processor 502 can be configured to implement one or more methods described in this document. The memory (multiple memories) 504 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 506 can be used to implement some of the techniques described in this document in hardware circuits and can be partially or completely part of the processor 502 (e.g., a graphics processor core GPU or other signal processing circuitry).

[0990] In this document, the term "video processing" or codec can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block can correspond to bits spread across co-located or different locations within the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on the error residual after transformation and coding, and bits from the header and other fields in the bitstream can also be used.

[0991] It will be appreciated that the disclosed methods and techniques will benefit video encoder and / or decoder embodiments incorporated into video processing devices such as smartphones, laptops, desktops, and similar devices by enabling use of the techniques disclosed in this document.

[0992] Figure 6 6 is a flow chart of an example method 600 for video processing. The method 600 includes, at 610, performing a conversion between a current video block and a codec representation of the current video block, wherein during the conversion, if a resolution and / or size of a reference picture is different from a resolution and / or size of the current video block, applying a same interpolation filter to a set of adjacent or non-adjacent samples predicted using the current video block.

[0993] Some embodiments may be described using the following clause-based format.

[0994] 1. A video processing method, comprising:

[0995] Converting between a current video block and a codec representation of the current video block is performed, wherein during the conversion, if a resolution and / or size of a reference picture is different from the resolution and / or size of the current video block, a same interpolation filter is applied to a set of adjacent or non-adjacent samples predicted using the current video block.

[0996] 2. The method of clause 1, wherein the same interpolation filter is a vertical interpolation filter.

[0997] 3. The method of clause 1, wherein the same interpolation filter is a horizontal interpolation filter.

[0998] 4. The method of clause 1, wherein a set of adjacent or non-adjacent samples includes all samples located in a region of the current video block.

[0999] 5. The method of clause 4, wherein the current video block is divided into a plurality of rectangles, each rectangle having a size of MxN.

[1000] 6. The method of clause 5, wherein M and / or N are predetermined.

[1001] 7. The method of clause 5, wherein M and / or N are derived from the dimensions of the current video block.

[1002] 8. The method of clause 5, wherein M and / or N are signaled in a codec representation of the current video block.

[1003] 9. The method of clause 1, wherein a group of samples share the same motion vector.

[1004] 10. The method of clause 9, wherein a group of samples share the same horizontal component and / or the same fractional portion of the horizontal component.

[1005] 11. The method of clause 9, wherein a group of samples share the same vertical component and / or the same fractional portion of the vertical component.

[1006] 12. The method of one or more of clauses 9-11, wherein the same motion vector or components thereof satisfy one or more rules based on at least one of: a resolution of a reference picture, a size of a reference picture, a resolution of a current video block, a size of a current video block, or a precision value.

[1007] 13. The method of one or more of clauses 9-11, wherein the same motion vector or components thereof correspond to motion information of samples located in the current video block.

[1008] 14. The method according to one or more of clauses 9-11, wherein the same motion vector or its components are set as motion information of virtual samples located inside or outside the group.

[1009] 15. A video processing method, comprising:

[1010] A conversion is performed between the current video block and a codec representation of the current video block, wherein, during the conversion, only integer-valued motion information associated with the current block is permitted for a block predicted using the current video block if a resolution and / or size of a reference picture differs from a resolution and / or size of the current video block.

[1011] 16. The method of clause 15, wherein the integer-valued motion information is derived by rounding the original motion information of the current video block.

[1012] 17. The method of clause 15, wherein the original motion information of the current video block is in a horizontal direction and / or a vertical direction.

[1013] 18. A video processing method, comprising:

[1014] A conversion is performed between the current video block and a codec representation of the current video block, wherein, during the conversion, if a resolution and / or size of a reference picture differs from a resolution and / or size of the current video block, an interpolation filter is applied to derive a block predicted using the current video block, and wherein the interpolation filter is selected based on a rule.

[1015] 19. The method of clause 18, wherein the rule is related to the resolution and / or size of the reference picture relative to the resolution and / or size of the current video block.

[1016] 20. The method of clause 18, wherein the interpolation filter is a vertical interpolation filter.

[1017] 21. The method of clause 18, wherein the interpolation filter is a horizontal interpolation filter.

[1018] 22. The method of clause 18, wherein the interpolation filter is one of: a 1-tap filter, a bilinear filter, a 4-tap filter, or a 6-tap filter.

[1019] 23. The method of clause 22, wherein an interpolation filter is used as part of the other steps of the conversion.

[1020] 24. The method of clause 18, wherein the interpolation filter includes the use of padding samples.

[1021] 25. The method of clause 18, wherein use of the interpolation filter depends on the color components of the samples of the current video block.

[1022] 26. A video processing method, comprising:

[1023] Converting between the current video block and a codec representation of the current video block is performed, wherein during the conversion, if a resolution and / or size of a reference picture differs from a resolution and / or size of the current video block, a deblocking filter is selectively applied, wherein a strength of the deblocking filter is set according to a rule related to the resolution and / or size of the reference picture relative to the resolution and / or size of the current video block.

[1024] 27. The method of clause 27, wherein the strength of the deblocking filter varies from video block to video block.

[1025] 28. A video processing method, comprising:

[1026] A conversion is performed between the current video block and a codec representation of the current video block, wherein during the conversion, if a sub-picture of the current video block exists, the conforming bitstream satisfies rules related to a resolution and / or size of a reference picture relative to a resolution and / or size of the current video block.

[1027] 29. The method of clause 28, further comprising:

[1028] The current video block is partitioned into one or more sub-pictures, where the partitioning depends on at least a resolution of the current video block.

[1029] 30. A video processing method, comprising:

[1030] A conversion is performed between the current video block and a codec representation of the current video block, wherein during the conversion, reference pictures of the current video block are resampled according to rules based on a dimension of the current video block.

[1031] 31. A video processing method, comprising:

[1032] A conversion is performed between a current video block and a codec representation of the current video block, wherein during the conversion, use of codec tools for the current video block is selectively enabled or disabled depending on a resolution and / or size of a reference picture of the current video block relative to a resolution and / or size of the current video block.

[1033] 32. The method of one or more of the above, wherein the set of samples is located within a consistency window.

[1034] 33. The method of clause 32, wherein the consistency window is rectangular.

[1035] 34. The method of one or more of the preceding claims, wherein the resolution is related to the resolution of an encoded / decoded video block or the resolution of a consistency window in an encoded / decoded video block.

[1036] 35. The method as recited in one or more of the preceding claims, wherein the size is related to the size of a coded / decoded video block or the size of a consistency window in a coded / decoded video block.

[1037] 36. The method of one or more of the preceding claims, wherein the dimension is related to a dimension of a coded / decoded video block or a dimension of a consistency window in a coded / decoded video block.

[1038] 37. The method of clause 32, wherein the consistency window is defined by a set of consistency clipping window parameters.

[1039] 38. The method of clause 37, wherein at least a portion of the conforming cropping window parameters are implicitly or explicitly signaled in the codec representation.

[1040] 39. The method as recited in one or more of the preceding claims, wherein signaling of the set of consistent cropping window parameters in a codec representation is not permitted.

[1041] 40. The method as recited in one or more of the preceding claims, wherein the position of the reference sample is derived relative to the top left sample of the current video block in the consistency window.

[1042] 41. A video processing method, comprising:

[1043] Converting between a plurality of video blocks and codec representations of the plurality of video blocks is performed, wherein during the conversion, a first conformance window is defined for a first video block and a second conformance window is defined for a second video block, wherein a ratio of a width and / or a height of the first conformance window to the second conformance window complies with a rule based at least on a conforming bitstream.

[1044] 42. A video decoding device comprising a processor configured to implement one or more of the methods recited in clauses 1 to 41.

[1045] 43. A video encoding apparatus comprising a processor configured to implement one or more of the methods recited in clauses 1 to 41.

[1046] 44. A computer program product storing computer code, which, when executed by a processor, causes the processor to implement the method of any one of clauses 1 to 41.

[1047] 45. A method, apparatus or system as herein described.

[1048] Figure 7is a block diagram illustrating an exemplary video processing system 700 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 700. 700 may include an input 702 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 702 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, passive optical networks (PONs)) and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[1049] System 700 may include a codec component 704 that can implement various codecs or codec methods described in this document. The codec component 704 can reduce the average bit rate of the video from input 702 to the output of the codec component 704 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of the codec component 704 can be stored or sent via a communication connected by a component 706. Component 708 can use the storage or communication bitstream (or codec) representation of the video received at input 702 to generate pixel values ​​or displayable video sent to display interface 710. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the encoding results will be performed by the decoder.

[1050] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[1051] Figure 8 is a block diagram illustrating an exemplary video coding system 100 that may utilize the techniques of this disclosure.

[1052] like Figure 8 As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[1053] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[1054] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[1055] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[1056] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, configured to interface with an external display device.

[1057] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or future standards.

[1058] Figure 9 is shown to be Figure 8 A block diagram of an example of a video encoder 200 is shown for video encoder 114 in system 100 .

[1059] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[1060] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[1061] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.

[1062] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are not shown for the purpose of explanation. Figure 5 In the examples, they are respectively represented.

[1063] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[1064] The mode selection unit 203 can, for example, select a coding mode (intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select a resolution of motion vectors for the block in the case of inter prediction (e.g., sub-pixel or integer pixel precision).

[1065] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 (except the picture associated with the current video block).

[1066] For example, motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice.

[1067] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[1068] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 containing the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[1069] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder.

[1070] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.

[1071] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[1072] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[1073] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.

[1074] Intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 206 performs intra-frame prediction on the current video block, intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.

[1075] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[1076] In other examples, such as in skip mode, for the current video block, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[1077] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[1078] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[1079] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[1080] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[1081] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[1082] Figure 10 is shown to be Figure 8 A block diagram of an example of a video decoder 300 of the video decoder 114 in the system 100 is shown.

[1083] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 10 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[1084] exist Figure 10 In the example of FIG, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those performed with respect to the video encoder 200 (e.g., Figure 9 ) is a decoding process that is roughly the opposite of the encoding process described.

[1085] The entropy decoding unit 301 can retrieve a coded bitstream. The coded bitstream can include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information from the entropy-decoded video data. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge mode.

[1086] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel precision may be included in a syntax element.

[1087] The motion compensation unit 302 may calculate interpolated values ​​of sub-integer pixels of the reference block using the interpolation filter used by the video encoder 20 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.

[1088] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partition information describing how each macroblock of the pictures of the coded video sequence is partitioned, a mode indicating how each partition is to be encoded, one or more reference frames (and reference frame lists) to use for each inter-coded block, and other information for decoding the coded video sequence.

[1089] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[1090] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.

[1091] Figure 11 1 is a flowchart representation of a video processing method according to the present technology. The method 1100 includes, at operation 1110, performing conversion between a video and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies indicating in the bitstream representation the applicability of a decoder-side motion vector refinement codec and a bidirectional optical flow codec for a video picture.

[1092] In some embodiments, the decoder-side motion vector refinement tool includes performing automatic refinement of motion vectors without transmitting additional motion data to perform motion compensation at the decoder side. In some embodiments, the bidirectional optical flow encoding and decoding tool includes enhancing bidirectional prediction prediction samples of a block using higher precision motion vectors derived from two reference blocks. In some embodiments, the applicability is indicated in a picture header field of the bitstream representation.

[1093] Figure 12 1 is a flowchart representation of a video processing method according to the present technology. The method 1200 includes: at operation 1210, performing conversion between a video picture of a video and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies indicating the use of a codec tool in a picture header corresponding to the video picture.

[1094] In some embodiments, the format rule specifies that the indication of the codec tool in the picture header is based on a syntax element associated with the codec tool in a sequence parameter set, a video parameter set, or a decoder parameter set represented by the bitstream. In some embodiments, the codec tool includes at least one of the following: a prediction refinement using optical flow (PROF) codec, a cross-component adaptive loop filtering (CCALF) codec, or an inter-frame prediction codec using geometric partitioning (GEO). In some embodiments, use of the codec tool is indicated by one or more syntax elements.

[1095] In some embodiments, the format rules provide that use of the codec is conditionally indicated based on a syntax element that indicates, in the sequence parameter set, enabling or disabling of the codec. In some embodiments, the codec comprises a prediction refinement using optical flow (PROF) codec. In some embodiments, the codec comprises a cross-component adaptive loop filtering (CCALF) codec. In some embodiments, the codec comprises an inter-frame prediction codec using geometric partitioning (GEO). In some embodiments, use is further conditionally indicated in a slice header based on a picture header.

[1096] Figure 13 13 is a flow chart of a video processing method according to the present technology. The method 1300 includes, at operation 1310, performing conversion between a video including a video picture of one or more video units and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies including a first syntax element in a picture header to indicate allowed prediction types for at least some of the one or more video units in the video picture.

[1097] In some embodiments, the video units of the one or more video units comprise slices, tiles, or slices. In some embodiments, the picture header comprises a first syntax element to indicate whether all video units in the one or more video units have the same prediction type. In some embodiments, the same prediction type is an intra codec type. In some embodiments, if the first syntax element indicates that all slices of the corresponding video picture have the same prediction type, a second syntax element indicating the slice type is omitted from the slice header corresponding to the slice. In some embodiments, the picture header comprises a third syntax element indicating whether at least one of the plurality of video units is not intra coded.

[1098] In some embodiments, based on whether the indicated same prediction type is a specific prediction type, an indication of codec tools allowed for the specific prediction type is conditionally signaled in the bitstream representation. In some embodiments, when the same prediction type is a bidirectional prediction type, the codec tools allowed for the specific prediction type include: a decoder-side motion vector refinement codec tool, a bidirectional optical flow codec tool, a triangular segmentation pattern codec tool, or a geometric segmentation (GEO) codec tool. In some embodiments, when the specific prediction type is an intra-frame codec type, the codec tools allowed for the same prediction type include a dual-tree codec tool. In some embodiments, the format rule provides for indicating the use of the codec tool based on the indication of the same prediction type. In some embodiments, the use of the codec tool is determined based on the indication of the same prediction type.

[1099] In some embodiments, the converting comprises encoding and decoding the video into the bitstream representation. In some embodiments, the converting comprises decoding the bitstream representation into the video.

[1100] Some embodiments of the disclosed technology include making a decision or determining to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of the video block, but will not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when enabled based on the decision or determination, the conversion from the video block to the bitstream representation of the video will use the video processing tool or mode. In another example, when the video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to the video block will be performed using the video processing tool or mode enabled based on the decision or determination.

[1101] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode when converting video blocks to a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was enabled based on the decision or determination.

[1102] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or any combination thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or any combination thereof. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or a plurality of processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[1103] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers, located at one site or distributed across multiple sites and interconnected by a communications network.

[1104] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[1105] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[1106] While this patent document contains many specifics, they should not be construed as limitations on the scope of any subject matter or the claims, but rather as descriptions of features for particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, while the features described above may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features in a claim combination may be removed from the combination, and a claim combination may be directed to a subcombination or variations of a subcombination.

[1107] Likewise, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[1108] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: for converting between a video comprising a video picture including one or more video units and a bitstream of the video, determining a prediction type for at least one of the one or more video units in the video picture based on one or more syntax elements; as well as performing said converting based on said determining, wherein the one or more syntax elements are included in a picture header of the bitstream; wherein the one or more video units are one or more slices; The one or more syntax elements include a first syntax element, the first syntax element indicating whether a prediction type of all slices of the one or more slices is a first prediction type; When the first syntax element indicates that all slices of the video picture are I slices, a syntax element indicating a slice type is not included in a slice header of the bitstream.

2. The method according to claim 1, wherein The first prediction type is intra prediction.

3. The method according to claim 1, wherein The one or more syntax elements conditionally include a second syntax element that indicates whether at least one slice of the one or more slices is a non-I slice.

4. The method according to claim 1, wherein The converting includes encoding the video into the bitstream.

5. The method according to claim 1, wherein The converting includes decoding the video from the bitstream.

6. The method according to claim 1, further comprising: performing conversion between said video and a bitstream representation of said video, The bitstream representation conforms to a format rule that specifies including a third syntax element in a picture header to indicate allowed prediction types for at least some of the one or more video units in the video picture.

7. The method of claim 6, wherein a video unit of the one or more video units comprises a stripe, a brick, or a slice.

8. The method of claim 6, wherein the picture header comprises a fourth syntax element to indicate whether all video units of the one or more video units have the same prediction type.

9. The method of claim 8, wherein the same prediction type is an intra-frame coding type.

10. The method according to claim 8, wherein In case that the fourth syntax element indicates that all slices of the corresponding video picture have the same prediction type, the third syntax element indicating the slice type is omitted in the slice header corresponding to the slice.

11. The method according to claim 8, wherein The picture header includes a fifth syntax element indicating whether at least one of the plurality of video units is not intra-coded.

12. The method according to claim 8, wherein Based on whether the indicated same prediction type is a specific prediction type, an indication of codec tools enabled for the specific prediction type is conditionally signaled in the bitstream representation.

13. The method according to claim 12, wherein: In the case where the same prediction type is a bidirectional prediction type, the codec tools allowed for the specific prediction type include: a decoder-side motion vector refinement codec tool, a bidirectional optical flow codec tool, a triangular partitioning mode codec tool, or a geometric partitioning (GEO) codec tool.

14. The method of claim 12, wherein in case the specific prediction type is an intra codec type, the codec tools allowed for the same prediction type include a dual-tree codec tool.

15. The method according to any one of claims 12 to 14, wherein the format rules provide for indicating the use of the codec tool based on the indication of the same prediction type.

16. The method according to any one of claims 12 to 14, wherein the use of the codec tool is determined based on the indication of the same prediction type.

17. The method of claim 6, wherein the converting comprises encoding the video into the bitstream representation.

18. The method of claim 6, wherein the converting comprises decoding the bitstream representation into the video.

19. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For conversion between a video comprising a video picture comprising one or more video units and a bitstream of the video, determining a prediction type for at least one of the one or more video units in the video picture based on one or more syntax elements; and performing said converting based on said determining, in, The one or more syntax elements are included in a picture header of the bitstream; wherein the one or more video units are one or more slices; The one or more syntax elements include a first syntax element, the first syntax element indicating whether a prediction type of all slices of the one or more slices is a first prediction type; When the first syntax element indicates that all slices of the video picture are I slices, a syntax element indicating a slice type is not included in a slice header of the bitstream.

20. The device according to claim 19, wherein The first prediction type is intra prediction.

21. The apparatus of claim 19, wherein the one or more syntax elements conditionally include a second syntax element that indicates whether at least one slice of the one or more slices is a non-I slice.

22. A non-transitory computer-readable storage medium storing instructions that cause a processor to: For conversion between a video comprising a video picture comprising one or more video units and a bitstream of the video, determining a prediction type for at least one of the one or more video units in the video picture based on one or more syntax elements; and performing said converting based on said determining, in, The one or more syntax elements are included in a picture header of the bitstream; wherein the one or more video units are one or more slices; The one or more syntax elements include a first syntax element, the first syntax element indicating whether a prediction type of all slices of the one or more slices is a first prediction type; When the first syntax element indicates that all slices of the video picture are I slices, a syntax element indicating a slice type is not included in a slice header of the bitstream.

23. The non-transitory computer-readable storage medium of claim 22, in, The first prediction type is intra prediction; and in, The one or more syntax elements conditionally include a second syntax element that indicates whether at least one slice of the one or more slices is a non-I slice.

24. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a video processing device performing a method, wherein the method comprises: determining, for a video comprising a video picture including one or more video units, a prediction type for at least one of the one or more video units in the video picture based on one or more syntax elements; as well as generating the bitstream based on the determination, wherein the one or more syntax elements are included in a picture header of the bitstream; wherein the one or more video units are one or more slices; The one or more syntax elements include a first syntax element, the first syntax element indicating whether a prediction type of all slices of the one or more slices is a first prediction type; When the first syntax element indicates that all slices of the video picture are I slices, a syntax element indicating a slice type is not included in a slice header of the bitstream.

25. The non-transitory computer-readable recording medium according to claim 24, in, The first prediction type is intra prediction; and in, The one or more syntax elements conditionally include a second syntax element that indicates whether at least one slice of the one or more slices is a non-I slice.

26. A method for storing a bitstream of a video, comprising: for a video comprising a video picture comprising one or more video units, determining, based on one or more syntax elements, a prediction type for at least one of the one or more video units in the video picture; generating the bitstream based on the determination; as well as storing the bitstream in a non-transitory computer-readable storage medium, wherein the one or more syntax elements are included in a picture header of the bitstream; wherein the one or more video units are one or more slices; The one or more syntax elements include a first syntax element, the first syntax element indicating whether a prediction type of all slices of the one or more slices is a first prediction type; When the first syntax element indicates that all slices of the video picture are I slices, a syntax element indicating a slice type is not included in a slice header of the bitstream.

27. A video processing device comprising a processor configured to implement the method according to any one of claims 6 to 18.

28. A computer readable medium having stored thereon codes which, when executed by a processor, cause the processor to implement the method of any one of claims 6 to 18.