Signaling of slice types in video picture header
By implementing Adaptive Resolution Change (ARC) technology in the video encoder, the problem of the inability of existing video codecs to flexibly change resolution is solved, achieving efficient resource utilization and seamless video switching when network conditions change, thus improving the user experience.
Patent Information
- Application Number
- CN202080090776.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-27
- Filing Date
- 2020-12-28
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-12-28
AI Technical Summary
Existing video codecs cannot flexibly change resolution without introducing internal random access points, resulting in increased encoding complexity, excessive resource consumption, and a degraded user experience when network conditions change. This is especially true in multi-party video conferencing and streaming, where there are frequent changes in active speakers or rapid startup displays, making seamless switching difficult.
Adaptive Resolution Change (ARC) technology is employed to adjust the resolution of video images to adapt to changes in network conditions by implementing processor configuration in the video encoder, and to omit or infer strip types in the bitstream to achieve flexible resolution conversion.
It reduces encoding complexity and resource consumption without affecting video quality, improves the flexibility of video processing and user experience, and supports rapid startup and seamless switching in multi-party video conferences and streaming.
Smart Images

Figure CN115152230B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] In accordance with the patent law and / or rules applicable under the Paris Convention, this application claims timely protection of the priority and interest of International Patent Application No. PCT / CN2019 / 129069, filed on December 27, 2019. For all purposes required by law, the entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology
[0004] Currently, efforts are underway to improve the performance of existing video codec technologies to provide better compression ratios or to offer video encoding and decoding schemes that allow for lower complexity or parallel implementations. Industry experts have recently proposed several new video codec tools, and they are currently being tested to determine their effectiveness. Summary of the Invention
[0005] This document describes devices, systems, and methods related to digital video encoding and decoding, specifically devices, systems, and methods related to stripe-type signaling in video picture headers. The described methods can be applied to existing video encoding and decoding standards (e.g., High Efficiency Video Codec (HEVC) or Universal Video Codec) and future video encoding and decoding standards or codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: performing a conversion between a video comprising one or more video images and a bitstream of the video, the video images comprising one or more stripes, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that all stripes in the one or more video images are encoded and decoded as I-striped video images, omitting syntax elements related to P-stripes and B-stripes from the image header of the video images.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: performing a conversion between a video comprising one or more video images and a bitstream of the video, the video images comprising one or more stripes, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the image header of each video image includes a syntax element indicating whether all stripes in the video image are encoded and decoded using the exact same codec type.
[0008] In yet another representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a picture type of a picture is signaled in a picture header of the picture.
[0009] In yet another representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a picture type of a picture is signaled in a picture header of the picture.
[0010] In yet another representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a picture type of a picture is signaled in a picture header of the picture.
[0011] In yet another representative aspect, the disclosed technology can be used to provide a method of video processing. The method includes performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a picture type of a picture is signaled in a picture header of the picture.
[0012] In yet another representative aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement a method described above.
[0013] In yet another representative aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement a method described above.
[0014] In yet another representative aspect, a computer readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0015] These and other features will be described in this document. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 An example of sub-block motion vector (VSB) and motion vector difference is shown.
[0017] Figure 2 An example of a 16x16 video block partitioned into 16 4x4 regions is shown.
[0018] Figures 3A-3C An example of a particular position in a sample is shown.
[0019] Figure 4A and Figure 4B An example of a current sample in a current picture and a reference sample in a reference picture is shown.
[0020] Figure 5 An example of decoder-side motion vector refinement is shown.
[0021] Figure 6 An example of the flow of the concatenated DMVR and BDOF process in VTM5.0 is shown. The DMVR SAD operation and the BDOF SAD operation are different and not shared.
[0022] Figure 7 is a block diagram illustrating an example video processing system in which various techniques disclosed herein can be implemented.
[0023] Figure 8 is a block diagram of an example hardware platform for video processing.
[0024] Figure 9 is a block diagram illustrating an example video coding system in which some embodiments of the disclosure can be implemented.
[0025] Figure 10 is a block diagram illustrating an example of an encoder that can implement some embodiments of the disclosure.
[0026] Figure 11 is a block diagram illustrating an example of a decoder that can implement some embodiments of the disclosure.
[0027] Figures 12-17 is a flowchart illustrating an example method of video processing. DETAILED DESCRIPTION
[0028] 1. Video coding in HEVC / H.265
[0029] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure that utilizes temporal prediction plus transform coding in the spatial domain. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). In April 2018, the Joint Video Team (JVT) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard with the goal of 50% bitrate reduction compared to HEVC.
[0030] The latest version of the VVC draft (i.e., Versatile Video Coding (Draft 6)) can be found at the following URL: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 15_Gothenburg / wg11 / JVET-O2001-v14.zip. The latest reference software of VVC (named VTM) can be found at the following URL: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-6.0.
[0031] AVC and HEVC do not have the ability to change the resolution without having to introduce IDR or intra random access point (IRAP) pictures; such an ability can be referred to as adaptive resolution change (ARC). Some use cases or application scenarios can benefit from the ARC feature, including:
[0032] Rate adaptation in video telephony and conferencing: In order to adapt the coded video to changing network conditions, when network conditions become worse so that the available bandwidth becomes lower, the encoder can adapt it by coding smaller resolution pictures. Currently, changing picture resolution can only be done after an IRAP picture; this has several problems. A reasonably good quality IRAP picture will be much larger than an inter-coded picture, and correspondingly more complex to decode: this is time and resource consuming. If the decoder requests a resolution change for loading reasons, this is a problem. It can also break low latency buffering conditions, force audio resynchronization, and the end-to-end delay of the stream will increase, at least temporarily. This can lead to a bad user experience.
[0033] Change of active speaker in multi-party video conferencing: For multi-party video conferencing, the active speaker is usually shown with a video size that is larger than the video of the other conference participants. When the active speaker changes, the picture resolution of each participant can also need to be adjusted. When such change of active speaker happens frequently, the need for an ARC feature becomes more important.
[0034] Fast start in streaming: For streaming applications, the application usually buffers a certain length of decoded pictures before starting display. Starting a bitstream with a smaller resolution will allow the application to have enough pictures in the buffer to start display faster.
[0035] Adaptive stream switching in streaming: The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature called @mediaStreamStructureId. This enables switching between different representations at an open GOP random access point with a non-decodable leading picture (e.g., a CRA picture with an associated RASL picture of HEVC). When two different representations of the same video have different bitrates but the same spatial resolution, while they have the same @mediaStreamStructureId value, switching between the two representations at a CRA picture with an associated RASL picture can be performed, and the RASL pictures associated with the switch at the CRA picture can be decoded with acceptable quality, enabling seamless switching. For ARC, the @mediaStreamStructureId feature can also be used to switch between DASH representations with different spatial resolutions.
[0036] ARC is also known as dynamic resolution conversion.
[0037] ARC can also be considered as a special case of Reference Picture Resampling (RPR), such as H.263 Annex P.
[0038] 2.1. Reference Picture Resampling in H.263 Annex P
[0039] This mode describes an algorithm that warps the reference picture before it is used for prediction. This can be useful for resampling a reference picture that has a source format different from the picture being predicted. By warping the shape, size and position of the reference picture, it can also be used for global motion estimation or rotational motion estimation. The syntax includes the warping parameters to be used and the resampling algorithm. The simplest level of operation for the reference picture resampling mode is an implicit factor-4 resampling, as only FIR filters need to be applied for the up- and down-sampling processes. In this case, no additional signaling overhead is needed, as the use can be understood when the size of the new picture (indicated in the picture header) is different from the size of the previous picture.
[0040] 2.2. ARC contributions to VVC
[0041] Several contributions have been made to address ARC, as follows:
[0042] JVET-M0135, JVET-M0259, JVET-N0048, JVET-N0052, JVET-N0118, JVET-N0279.
[0043] 2.3. Consistency window in VVC
[0044] The consistency window in VVC defines a rectangle. The samples inside the consistency window belong to the picture of interest. When output, the samples outside the consistency window can be discarded.
[0045] When the consistency window is applied, the scaling ratio in RPR is derived based on the consistency window.
[0046] Picture parameter set RBSP syntax
[0047]
[0048]
[0049] pic_width_in_luma_samples specifies the width of each decoded picture referring to the PPS in units of luma samples. pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples.
[0050] When subpics_present_flag is equal to 1, the value of pic_width_in_luma_samples shall be equal to pic_width_max_in_luma_samples.
[0051] pic_height_in_luma_samples specifies the height of each decoded picture referring to the PPS in units of luma samples. pic_height_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_height_max_in_luma_samples.
[0052] When subpics_present_flag is equal to 1, the value of pic_height_in_luma_samples shall be equal to pic_height_max_in_luma_samples.
[0053] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be pic_width_in_luma_samples and pic_height_in_luma_samples, respectively, of the reference picture of the current picture referring to the PPS. The requirement for bitstream conformance is that all of the following conditions are met:
[0054] - pic width in luma samples * 2 shall be greater than or equal to refPicWidthInLumaSamples.
[0055] - pic height in luma samples * 2 shall be greater than or equal to refPicHeightInLumaSamples.
[0056] - pic width in luma samples shall be less than or equal to refPicWidthInLumaSamples * 8.
[0057] - pic height in luma samples shall be less than or equal to refPicHeightInLumaSamples * 8 .
[0058] conformance_window_flag equal to 1 indicates that a conformance clipping window offset parameter follows in the SPS. conformance_window_flag equal to 0 indicates that no conformance clipping window offset parameter is present.
[0059] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset specify, in terms of the rectangular region specified in picture coordinates for output, the samples of the picture in the CVS output from the decoding process. When conformance_window_flag is equal to 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are inferred to be equal to 0.
[0060] The conformance clipping window contains luma samples whose horizontal picture coordinates (inclusive) are from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples - (SubWidthC * conf_win_right_offset + 1), and whose vertical picture coordinates (inclusive) are from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples - (SubHeightC * conf_win_bottom_offset + 1).
[0061] The value of SubWidthC * (conf_win_left_offset + conf_win_right_offset) shall be less than pic_width_in_luma_samples, and the value of SubHeightC * (conf_win_top_offset + conf_win_bottom_offset) shall be less than pic_height_in_luma_samples.
[0062] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
[0063]
[0064]
[0065] When ChromaArrayType is not equal to 0, the corresponding specified sample of the two chroma arrays is the sample with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.
[0066]
[0067] Let ppsA and ppsB be any two PPSs referring to the same SPS. The requirement for bitstream conformance is that ppsA and ppsB shall have the same values of conf win left offset, conf win right offset, conf win top offset and conf win bottom offset, respectively, when ppsA and ppsB have the same values of pic width in luma samples and pic height in luma samples, respectively.
[0068] 2.4. RPR in JVET-O2001-v14
[0069] ARC (also known as RPR (Reference Picture Resampling)) is incorporated in JVET-O2001-v14.
[0070] For RPR in JVET-O2001-v14, TMVP is disabled if the collocated picture has a different resolution than the current picture. In addition to this, BDOF and DMVR are also disabled when the reference picture has a different resolution than the current picture.
[0071] When the reference picture has a different resolution than the current picture, for handling normal MC, the chapter of interpolation is defined as follows:
[0072] 8.5.6.3 Fractional sample interpolation process
[0073] 8.5.6.3.1 Overview
[0074] The inputs of this process are:
[0075] - luma position (xSb, ySb) specifying the top-left sample of the current coding sub-block relative to the top-left luma sample of the current picture,
[0076] - variable sbWidth specifying the width of the current coding sub-block,
[0077] - variable sbHeight specifying the height of the current coding sub-block,
[0078] - motion vector offset mvOffset,
[0079] - refined motion vector refMvLX,
[0080] - selected reference picture sample array refPicLX,
[0081] - half-sample interpolation filter index hpelIfIdx,
[0082] - bidirectional optical flow flag bdofFlag,
[0083] - variable cldx specifying the color component index of the current block.
[0084] The output of the process is:
[0085] - an (sbWidth + brdExtSize) x (sbHeight + brdExtSize) array of prediction sample values, predSamplesLX.
[0086] The prediction block boundary extension size brdExtSize is derived as follows:
[0087] brdExtSize = (bdofFlag || (inter affine flag [xSb] [ySb] && sps affine prof enabled flag))? 2 : 0 (8-752)
[0088] The variable fRefWidth is set equal to PicOutputWidthL in units of luma samples of the reference picture.
[0089] The variable fRefHeight is set equal to PicOutputHeightL in units of luma samples of the reference picture.
[0090] The motion vector mvLX is set to (refMvLX - mvOffset).
[0091] - If cldx is equal to 0, the following applies:
[0092] - The scaling factors and their fixed-point representations are defined as follows
[0093] hori_scale_fp = ((fRefWidth « 14) + (PicOutputWidthL » 1)) / PicOutputWidthL (8-753)
[0094] vert_scale_fp = ((fRefHeight « 14) + (PicOutputHeightL » 1)) / PicOutputHeightL (8-754)
[0095] - Let (xIntL, yIntL) be the luma position given in full-sample units and (xFracL, yFracL) be the offset given in 1 / 16-sample units. These variables are only used in this clause to specify fractional-sample positions within the reference sample array refPicLX.
[0096] - The top-left coordinates (xSbInt L ,ySbInt L ) of the reference sample padded border block are set equal to (xSb + (mvLX[0] » 4), ySb + (mvLX[1] » 4)).
[0097] - For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L = 0..sbHeight - 1 + brdExtSize) within the predicted luma sample array predSamplesLX, the corresponding predicted luma sample value predSamplesLX[x L ][y L ] is derived as follows:
[0098] - Let (refxSb L ,refySb L ) and (refx L ,refy L ) be the luma positions pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16-sample units. The variables refxSb L , refx L , refySb L and refy L are derived as follows:
[0099] refxSb L = ((xSb « 4) + refMvLX[0]) * hori_scale_fp (8-755)
[0100] refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128) » 8) + x L *((hori_scale_fp + 8) » 4)) + 32) » 6 (8-756)
[0101] refySb L = ((ySb « 4) + refMvLX[1]) * vert_scale_fp (8-757)
[0102] refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758)
[0103] – Variable xInt L yInt L xFrac L and yFrac L The derivation is as follows:
[0104] xInt L =refx L >>4 (8-759)
[0105] yInt L =refy L >>4 (8-760)
[0106] xFrac L =refx L &15 (8-761)
[0107] yFrac L =refy L &15 (8-762)
[0108] – If bdofFlag equals TRUE or (sps_affine_prof_enabled_flag equals TRUE and inter_affine_flag[xSb][ySb] equals TRUE), and one or more of the following conditions are true, then the luminance integer sample acquisition procedure specified in Section 8.5.6.3.3 is invoked to obtain (xInt) luminance integer samples. L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) and refPicLX are used as inputs to derive the predicted luminance sample value predSamplesLX[x L ][y L ]
[0109] -x L It equals 0.
[0110] -x L It equals sbWidth+1.
[0111] –y L It equals 0.
[0112] –y L It equals sbHeight + 1.
[0113] – Otherwise, the predicted luma sample value predSamplesLX[ xL ][ yL ] is derived by calling the luma sample 8-tap interpolation filter process as specified in clause 8.5.6.3.2 with ( xIntL - ( brdExtSize > 0? 1 : 0 ), yIntL - ( brdExtSize > 0? 1 : 0 ) ), ( xFracL, yFracL ), ( xSbIntL, ySbIntL ), refPicLX, hpelIdx, sbWidth, sbHeight, and ( xSb, ySb ) as inputs. L ,ySbInt L ).
[0114] – Otherwise ( cIdx is not equal to 0 ), the following applies:
[0115] – Let ( xIntC, yIntC ) be the chroma position in whole sample units and ( xFracC, yFracC ) be the offset in 1 / 32 sample units. These variables are only used in this subclause to specify the general fractional sample positions within the reference sample array refPicLX.
[0116] – The top-left coordinates ( xSbIntC, ySbIntC ) of the boundary block for reference sample padding are set equal to ( ( xSb / SubWidthC ) + ( mvLX[ 0 ] » 5 ), ( ySb / SubHeightC ) + ( mvLX[ 1 ] » 5 ) ).
[0117] – For each chroma sample position ( xC = 0..sbWidth - 1, yC = 0..sbHeight - 1 ) within the predicted chroma sample array predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[ xC ][ yC ] is derived as follows:
[0118] – Let ( refxSb C , refySb C ) and ( refx C , refy C ) be the chroma positions pointed to by the motion vector ( mvLX[ 0 ], mvLX[ 1 ]) in 1 / 32 sample units. The variables refxSb C , refySb C , refx C , and refy C are derived as follows:
[0119] refxSb C= (( xSb / SubWidthC << 5 ) + mvLX[ 0 ] ) * hori_scale_fp (8-763)
[0120] refx C = (( Sign( refxSb C ) * ( ( Abs( refxSb C ) + 256 ) » 9 ) + xC * ( ( hori_scale_fp + 8 ) » 4 ) ) + 16 ) » 5 (8-764)
[0121] refySb C = (( ySb / SubHeightC << 5 ) + mvLX[ 1 ] ) * vert_scale_fp (8-765)
[0122] refy C = (( Sign( refySb C ) * ( ( Abs( refySb C ) + 256 ) » 9 ) + yC * ( ( vert_scale_fp + 8 ) » 4 ) ) + 16 ) » 5 (8-766)
[0123] – variables xInt C , yInt C , xFrac C and yFrac C are derived as follows:
[0124] xInt C = refx C >> 5 (8-767)
[0125] yInt C = refy C >> 5 (8-768)
[0126] xFrac C = refy C & 31 (8-769)
[0127] yFrac C = refy C & 31 (8-770)
[0128] - the prediction sample values predSamplesLX[ xC ][ yC ] are derived by calling the process specified in clause 8.5.6.3.4 with ( xIntC, yIntC ), ( xFracC, yFracC ), ( xSbIntC, ySbIntC ), sbWidth, sbHeight and refPicLX as inputs.
[0129] 8.5.6.3.2 Luma sample interpolation filter process
[0130] The inputs of this process comprise:
[0131] - the luma position in whole sample units ( xInt L , yInt L ),
[0132] - the luma position in fractional sample units ( xFrac L , yFrac L ),
[0133] - the luma position in whole sample units ( xSbInt L , ySbInt L ) specifying the top-left sample of the boundary block relative to the top-left luma sample of the reference picture for reference sample padding,
[0134] - the array of luma reference samples refPicLX L ,
[0135] - the half-sample interpolation filter index hpelIfIdx,
[0136] - the variable sbWidth specifying the width of the current sub-block,
[0137] - the variable sbHeight specifying the height of the current sub-block,
[0138] - the luma position ( xSb, ySb ) specifying the top-left sample of the current sub-block relative to the top-left luma sample of the current picture,
[0139] The output of this process is the prediction luma sample value predSampleLX L
[0140] The variables shift1, shift2 and shift3 are derived as follows:
[0141] - the variable shift1 is set equal to Min( 4, BitDepth Y - 8 ), the variable shift2 is set equal to 6 and the variable shift3 is set equal to Max( 2, 14 - BitDepth Y ).
[0142] - The variable picW is set equal to pic_width_in_luma_samples and the variable picH is set equal to pic_height_in_luma_samples.
[0143] equal to xFrac L or yFrac L , the luma interpolation filter coefficients f L [p] are derived as follows:
[0144] - If MotionModelIdc[xSb][ySb] is greater than 0 and both sbWidth and sbHeight are equal to 4, the luma interpolation filter coefficients f L [p] are specified in Table 2.
[0145] - Otherwise, the luma interpolation filter coefficients f L [p] are specified in Table 1, depending on hpelIfIdx.
[0146] The luma positions (xInt i , yInt i ) in full-sample units are derived as follows, for i = 0..7:
[0147] - If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:
[0148] xInt i = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i - 3 ) (8-771)
[0149] yInt i = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i - 3 ) (8-772)
[0150] - Otherwise (subpic_treated_as_pic_flag[SubPicIdx] is equal to 0), the following applies:
[0151] xInt i=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag?ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3):xInt L +i-3) (8-773)
[0152] yInt i =Clip3(0,picH-1,yInt) L +i-3) (8-774)
[0153] The brightness position, calculated per unit of the full sample point, is further modified as follows for i = 0..7:
[0154] xInt i =Clip3(xSbInt) L -3,xSbInt L +sbWidth+4,xInt i (8-775)
[0155] yInt i =Clip3(ySbInt) L -3,ySbInt L +sbHeight+4,yInt i (8-776)
[0156] Predicted brightness sample value predSampleLX L The derivation is as follows:
[0157] –If xFrac L and yFrac L If both are equal to 0, then predSampleLX L The value of is derived as follows:
[0158] predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3 (8-777)
[0159] – Otherwise, if xFrac L Not equal to 0 and yFrac L If it equals 0, then predSampleLX L The value of is derived as follows:
[0160]
[0161] – Otherwise, if xFrac L L is not equal to 0 and yFrac L is not equal to 0, the value of predSampleLX L is derived as follows:
[0162]
[0163] – Otherwise, if xFrac L is not equal to 0 and yFrac L is not equal to 0, the value of predSampleLX L is derived as follows:
[0164] – The sample array temp[ n ] for n = 0..7 is derived as follows:
[0165]
[0166] – The predicted luma sample value predSampleLX L is derived as follows:
[0167]
[0168] Table 1 - Luma interpolation filter coefficients f for each 1 / 16 fraction sample position p L The specification of f[ p ] is
[0169]
[0170] Table 2 - Luma interpolation filter coefficients f for each 1 / 16 fraction sample position p for affine motion mode L The specification of f[ p ] is
[0171]
[0172]
[0173] 8.5.6.3.3 Luma integer sample fetching process
[0174] The inputs of the process include:
[0175] - the luma position in full sample units (xInt L , yInt L ),
[0176] - the luma reference sample array refPicLX L ,
[0177] The output of the process is the predicted luma sample value predSampleLX L
[0178] The variable shift is set equal to Max(2, 14 - BitDepth Y ).
[0179] The variable picW is set equal to pic_width_in_luma_samples and the variable picH is set equal to pic_height_in_luma_samples.
[0180] The luma position in full-sample units (xInt, yInt) is derived as follows:
[0181] xInt = Clip3(0, picW - 1, sps_ref_wraparound_enabled_flag? ClipH((sps_ref_wraparound_offset_minus1 + 1) * MinCbSizeY, picW, xInt L ): xInt L ) (8-782)
[0182] yInt = Clip3(0, picH - 1, yInt L ) (8-783)
[0183] The predicted luma sample value predSampleLX L is derived as follows:
[0184] predSampleLX L = refPicLX L [xInt][yInt] « shift3 (8-784)
[0185] 8.5.6.3.4 Chroma sample interpolation process
[0186] The inputs to this process are:
[0187] – the chroma position in full-sample units (xInt C , yInt C ),
[0188] – the chroma position in 1 / 32- fraction sample units (xFrac C , yFrac C ),
[0189] – the chroma position in full-sample units (xSbIntC, ySbIntC) specifying the top-left sample of the boundary block relative to the top-left chroma sample of the reference picture used for reference sample padding,
[0190] - variable sbWidth specifying the width of the current sub-block,
[0191] - variable sbHeight specifying the height of the current sub-block,
[0192] - chroma reference sample array refPicLX C .
[0193] The output of the process is the predicted chroma sample value predSampleLX C
[0194] The variables shiftl, shift2 and shift3 are derived as follows:
[0195] - variable shiftl is set equal to Min(4, BitDepth C - 8), variable shift2 is set equal to 6, and variable shift3 is set equal to Max(2, 14 - BitDepth C ).
[0196] - variable picWC is set equal to pic_width_in_luma_samples / SubWidthC and variable picHC is set equal to pic_height_in_luma_samples / SubHeightC.
[0197] The chroma interpolation filter coefficients f C [p] for each 1 / 32 fractional sample position p equal to xFrac C or yFrac C are specified in Table 3.
[0198] - variable xOffset is set equal to (sps_ref_wraparound_offset_minusl + 1) * MinCbSizeY) / SubWidthC.
[0199] The chroma positions in full sample units (xInt i , yInt i ) are derived as follows for i = 0..3:
[0200] - If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:
[0201] xInt i= Clip3( SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xInt L + i) (8-785)
[0202] yInt i = Clip3( SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yInt L + i) (8-786)
[0203] - Otherwise (subpic_treated_as_pic_flag[ SubPicIdx ] is equal to 0), the following applies:
[0204] xInt i = Clip3( 0, picW C - 1, sps_ref_wraparound_enabled_flag? ClipH( xOffset, picW C , xInt C + i - 1 ) : xInt C + i - 1 ) (8-787)
[0205] yInt i = Clip3( 0, picH C - 1, yInt C + i - 1 ) (8-788)
[0206] The chroma positions in full-sample units (xInt i , yInt i ) are further modified as follows for i = 0..3:
[0207] xInt i = Clip3( xSbIntC - 1, xSbIntC + sbWidth + 2, xInt i ) (8-789)
[0208] yInt i = Clip3( ySbIntC - 1, ySbIntC + sbHeight + 2, yInt i ) (8-790)
[0209] The predicted chroma sample value predSampleLX C is derived as follows:
[0210] – If xFrac C and yFrac C are both equal to 0, the value of predSampleLX C is derived as follows:
[0211] predSampleLX C = refPicLX C [ xInt1 ][ yInt1 ] « shift3 (8-791)
[0212] – Otherwise, if xFrac C is not equal to 0 and yFrac C is equal to 0, the value of predSampleLX C is derived as follows:
[0213]
[0214] – Otherwise, if xFrac C is equal to 0 and yFrac C is not equal to 0, the value of predSampleLX C is derived as follows:
[0215]
[0216] – Otherwise, if xFrac C is not equal to 0 and yFrac C is not equal to 0, the value of predSampleLX C is derived as follows:
[0217] – The sample array temp[ n ] for n = 0..3 is derived as follows:
[0218]
[0219] – The predicted chroma sample value predSampleLX C is derived as follows:
[0220] predSampleLX C = ( f C [ yFrac C ] [ 0 ] * temp [ 0 ] + f C [ yFrac C ] [ 1 ] * temp [ 1 ] + f C [ yFrac C ] [ 2 ] * temp [ 2 ] + f C [ yFrac C ] [ 3 ] * temp [ 3 ] ) » shift2 (8-795)
[0221] Table 3 - Chroma interpolation filter coefficients f for each 1 / 32 fraction sample position p C [p] of the specification
[0222]
[0223]
[0224] 2.5. JVET-N0236
[0225] A method of refining subblock-based affine motion compensated prediction with optical flow is presented herein. After performing subblock-based affine motion compensation, the prediction samples are refined by adding the difference derived from the optical flow equation, which is referred to as prediction refinement with optical flow (PROF). The proposed method can achieve pixel-level granularity of inter prediction without increasing memory access bandwidth.
[0226] To achieve finer motion compensation granularity, a method of refining subblock-based affine motion compensated prediction with optical flow is presented herein. After performing subblock-based affine motion compensation, the luma prediction samples are refined by adding the difference derived from the optical flow equation. The proposed PROF (prediction refinement with optical flow) is described as the following four steps.
[0227] Step 1) Perform subblock-based affine motion compensation to generate subblock prediction I(i,j).
[0228] Step 2) Calculate the spatial gradient g x (i,j) of the subblock prediction at each sample position using a 3-tap filter [-1, 0, 1]. y (i,j).
[0229] g x (i,j) = I(i+1,j) - I(i-1,j)
[0230] g y (i,j) = I(i,j+1) - I(i,j-1)
[0231] For gradient calculation, the subblock prediction is extended by one pixel on each side. To reduce memory bandwidth and complexity, the pixels on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Thus, additional interpolation for padding the area is avoided.
[0232] Step 3) Luma prediction refinement (denoted as ΔΙ) calculated from the optical flow equation.
[0233] ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j)
[0234] where delta MV (denoted as Δv(i,j)) is the difference between the pixel MV (denoted as v(i,j)) calculated for sample position (i,j) and the sub-block MV of the sub-block to which the pixel (i,j) belongs, as shown in Figure 1
[0235] Since the affine model parameters and the pixel position relative to the sub-block center are invariant between sub-blocks and sub-blocks, Δv(i,j) can be computed for the first sub-block and reused for other sub-blocks in the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the sub-block center, Δv(x,y) can be derived from the following equation,
[0236]
[0237] For 4-parameter affine model,
[0238]
[0239] For 6-parameter affine model,
[0240]
[0241] where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the top-left control point motion vector, top-right control point motion vector and bottom-left control point motion vector, w and h are the width and height of the CU.
[0242] Step 4) Finally, the luminance prediction refinement is added to the sub-block prediction I(i,j). The final prediction I’ is generated as the following equation.
[0243] I'(i,j) = I(i,j) + ΔI(i,j)
[0244] Some details in JVET-N0236
[0245] a) How to derive the gradient of PROF
[0246] In JVET-N0263, the gradient is computed for each subblock (4x4 subblock in VTM 4.0) of each reference list. For each subblock, the nearest integer samples of the reference block are fetched to fill the four outer rows of samples.
[0247] Assume the MV of the current subblock is (MVx, MVy). Then the fractional part is computed as (FracX, FracY) = (MVx & 15, MVy & 15). The integer part is computed as (IntX, IntY) = (MVx » 4, MVy » 4). The offset (OffsetX, OffsetY) is derived as follows:
[0248] OffsetX = FracX > 7? 1 : 0;
[0249] OffsetY = FracY > 7? 1 : 0;
[0250] Assume the top-left coordinate of the current subblock is (xCur, yCur), and the dimension of the current subblock is W x H.
[0251] Then (xCor0, yCor0), (xCor1, yCor1), (xCor2, yCor2) and (xCor3, yCor3) are computed as
[0252] (xCor0, yCor0) = (xCur + IntX + OffsetX - 1, yCur + IntY + OffsetY - 1);
[0253] (xCor1, yCor1) = (xCur + IntX + OffsetX - 1, yCur + IntY + OffsetY + H);
[0254] (xCor2, yCor2) = (xCur + IntX + OffsetX - 1, yCur + IntY + OffsetY);
[0255] (xCor3, yCor3) = (xCur + IntX + OffsetX + W, yCur + IntY + OffsetY);
[0256] Assume PredSample[x][y] (x = 0..W-1, y = 0..H-1) stores the predicted samples of the subblock. Then the padding samples are derived as
[0257] PredSample[x][-1] = (Ref(xCor0+x, yCor0) « Shift0) - Rounding for x = -1..W;
[0258] PredSample[x][H] = (Ref(xCorl+x, yCorl) « Shift0) - Rounding for x = -1..W;
[0259] PredSample[-1][y] = (Ref(xCor2, yCor2+y) « Shift0) - Rounding for y = 0..H-1;
[0260] PredSample[W][y] = (Ref(xCor3, yCor3+y) « Shift0) - Rounding for y = 0..H-1;
[0261] where Rec denotes the reference picture. Rounding is an integer which equals 2 in the example PROF implementation 13 . Shift0 = Max(2, (14 - BitDepth));
[0262] The PROF tries to improve the precision of the gradients, which is different from BIO in VTM 4.0, where the gradients are outputted with the same precision as the input luma samples.
[0263] The gradient calculation in PROF is as follows:
[0264] Shift1 = Shift0 - 4.
[0265] gradientH[x][y] = (predSamples[x+1][y] - predSample[x-1][y]) » Shift1
[0266] gradientV[x][y] = (predSample[x][y+1] - predSample[x][y-1]) » Shift1
[0267] It should be noted that predSamples[x][y] keeps the precision after the interpolation.
[0268] b) How to derive the Dv of PROF
[0269] The derivation of Dv (denoted as dMvH[posX][posY] and dMvV[posX][posY] with posX = 0..W-1, posY = 0..H-1) can be described as follows
[0270] Let the dimensions of the current block be cbWidth x cbHeight, the number of control point motion vectors be numCpMv, and the control point motion vectors be cpMvLX[ cpIdx ], with cpIdx = 0..numCpMv - 1, and X being 0 or 1, referring to the two reference lists.
[0271] The variables log2CbW and log2CbH are derived as follows:
[0272] log2CbW = Log2( cbWidth )
[0273] log2CbH = Log2( cbHeight )
[0274] The variables mvScaleHor, mvScaleVer, dHorX and dVerX are derived as follows:
[0275] mvScaleHor = cpMvLX[ 0 ][ 0 ] « 7
[0276] mvScaleVer = cpMvLX[ 0 ][ 1 ] « 7
[0277] dHorX = ( cpMvLX[ 1 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) « ( 7 - log2CbW )
[0278] dVerX = ( cpMvLX[ 1 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) « ( 7 - log2CbW )
[0279] The variables dHorY and dVerY are derived as follows:
[0280] - If numCpMv is equal to 3, the following applies:
[0281] dHorY = ( cpMvLX[ 2 ][ 0 ] - cpMvLX[ 0 ][ 0 ] ) « ( 7 - log2CbH )
[0282] dVerY = ( cpMvLX[ 2 ][ 1 ] - cpMvLX[ 0 ][ 1 ] ) « ( 7 - log2CbH )
[0283] - Else ( numCpMv is equal to 2 ), the following applies:
[0284] dHorY = - dVerX
[0285] dVerY = dHorX
[0286] The variables qHorX, qVerX, qHorY and qVerY are derived as follows
[0287] qHorX = dHorX « 2;
[0288] qVerX = dVerX « 2;
[0289] qHorY = dHorY « 2;
[0290] qVerY = dVerY « 2;
[0291] dMvH[0][0] and dMvV[0][0] are calculated as follows
[0292] dMvH[0][0] = ((dHorX + dHorY) « 1) - ((qHorX + qHorY) « 1);
[0293] dMvV[0][0] = ((dVerX + dVerY) « 1) - ((qVerX + qVerY) « 1);
[0294] dMvH[xPos][0] and dMvV[xPos][0] for xPos from 1 to W-1 are derived as follows:
[0295] dMvH[xPos][0] = dMvH[xPos-1][0] + qHorX;
[0296] dMvV[xPos][0] = dMvV[xPos-1][0] + qVerX;
[0297] For yPos from 1 to H-1 the following applies:
[0298] dMvH[xPos][yPos] = dMvH[xPos][yPos-1] + qHorY, xPos = 0..W-1
[0299] dMvV[xPos][yPos] = dMvV[xPos][yPos-1] + qVerY, xPos = 0..W-1
[0300] Finally, dMvH[xPos][yPos] and dMvV[xPos][yPos] (posX = 0..W-1, posY = 0..H-1) are right shifted as follows
[0301] dMvH[xPos][yPos] = SatShift(dMvH[xPos][yPos], 7+2-1);
[0302] dMvV[xPos][yPos] = SatShift(dMvV[xPos][yPos], 7+2-1);
[0303] where SatShift(x, n) and Shift(x, n) are defined as
[0304]
[0305] Shift(x, n) = (x + offset0) » n
[0306] In one example, offset0 and / or offset1 are set to (1 « n) » 1.
[0307] c) How to derive DI of PROF
[0308] For a position (posX, posY) within a sub-block, its corresponding dv(i,j) is denoted as (dMvH[posX][posY], dMvV[posX][posY]). Its corresponding gradient is denoted as (gradientH[posX][posY], gradientV[posX][posY]).
[0309] Then DI(posX, posY) is derived as follows.
[0310] (dMvH[posX][posY], dMvV[posX][posY]) is clipped as
[0311] dMvH[posX][posY] = Clip3(-32768, 32767, dMvH[posX][posY]);
[0312] dMvV[posX][posY] = Clip3(-32768, 32767, dMvV[posX][posY]);
[0313] DI(posX, posY) = dMvH[posX][posY] * gradientH[posX][posY] + dMvV[posX][posY] * gradientV[posX][posY];
[0314] DI(posX, posY) = Shift(DI(posX, posY), 1 + 1 + 4);
[0315] DI(posX, posY) = Clip3(-(2 13 -1), 2 13 -1, DI(posX, posY));
[0316] d) How to derive I' of PROF
[0317] If the current block is not coded as bi-prediction or weighted prediction,
[0318] I'(posX, posY) = Shift((I(posX, posY) + DI(posX, posY)), Shift0),
[0319] I'(posX, posY) = ClipSample(I'(posX, posY)),
[0320] where ClipSample clips the sample value to a valid output sample value.
[0321] I'(posX, posY) is then output as the inter prediction value.
[0322] Else (the current block is coded as bi-prediction or weighted prediction)
[0323] I'(posX, posY) will be stored and used to generate the inter prediction value from the other prediction values and / or weighting values.
[0324] 2.6. Slice header in JVET-O2001-vE
[0325]
[0326]
[0327]
[0328]
[0329]
[0330] 2.7. Sequence parameter set in JVET-O2001-vE
[0331]
[0332]
[0333]
[0334]
[0335]
[0336] 2.8. Picture parameter set in JVET-O2001-vE
[0337]
[0338]
[0339]
[0340]
[0341] 2.9. Adaptation parameter set in JVET-O2001-vE
[0342]
[0343]
[0344]
[0345] 2.10. Picture header proposed in VVC
[0346] Picture header was proposed to VVC in JVET-P0120 and JVET-P0239.
[0347] In JVET-P0120, the picture header is designed to have the following properties:
[0348] 1. The temporal id and layer id of the picture header NAL unit is the same as the temporal id and layer id of the layer access unit that contains the picture header.
[0349] 2. The picture header NAL unit shall precede the NAL unit of the first slice of the picture to which it is associated. This establishes the association between the picture header and the picture slices associated with the picture header without the need to signal the picture header id in the picture header and reference from the slice header.
[0350] 3. The picture header NAL unit shall follow the picture level parameter set or higher level such as DPS, VPS, SPS, PPS, etc. Therefore, this requires these parameter sets not to be duplicated / absent within the picture or access unit.
[0351] 4. The picture header contains information about the picture type of its associated picture. The picture type can be used to define the following (non-exhaustive list)
[0352] a. The picture is an IDR picture
[0353] b. The picture is a CRA picture
[0354] c. The picture is a GDR picture
[0355] d. The picture is a non-IRAP, non-GDR picture and contains only I slices
[0356] e. The picture is a non-IRAP, non-GDR picture and can contain only P and I slices
[0357] f. The picture is a non-IRAP, non-GDR picture and contains any of B, P and / or I slices
[0358] 5. Move the signaling of picture-level syntax elements in slice header to picture header.
[0359] 6. Signal in picture header non-picture-level syntax elements that are typically the same for all slices of the same picture in picture header. When those syntax elements are not present in picture header, they can be signaled in slice header.
[0360] In JVET-P0239, the concept of mandatory picture header was proposed to be signaled once per picture as the first VCL NAL unit of the picture. It was also proposed to move the syntax elements currently in slice header to this picture header. For a given picture, the syntax elements that are functionally only needed to be signaled once per picture can be moved to picture header instead of being signaled multiple times, e.g., syntax elements in slice header are signaled once per slice. The authors claim that there are benefits to moving the syntax elements from slice header because the computation needed for slice header processing can be the limiting factor for the total throughput.
[0361] Moving slice header syntax elements that are constrained to be the same within a picture
[0362] The syntax elements in this section have been constrained to be the same in all slices of a picture. It can be asserted that moving these fields to picture header, so they are signaled only once per picture instead of once per slice, avoids unnecessary redundant bit transmission without any change to the functionality of these syntax elements.
[0363] 1. In section 7.4.7.1 of current draft JVET-O2001-vE, there are the following semantic constraints:
[0364] When present, the value of each of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and slice_temporal_mvp_enabled_flag shall be the same in all slice headers of a coded picture.
[0365] Accordingly, each of these syntax elements can be moved to the picture header to avoid unnecessary redundant bits.
[0366] In this document, recovery_poc_cnt and no_output_of_prior_pics_flag are not moved to the picture header. Their presence in the slice header depends on conditional checks of the slice header nal_unit_type, so if it is desired to move these syntax elements to the picture header, it is suggested to investigate them.
[0367] 2. In section 7.4.7.1 of the current draft JVET-O2001-vE, there are the following semantic constraints:
[0368] When present, the value of slice_lmcs_aps_id shall be the same for all slices of a picture.
[0369] When present, the value of slice_scaling_list_aps_id shall be the same for all slices of a picture.
[0370] Accordingly, each of these syntax elements can be moved to the picture header to avoid unnecessary redundant bits.
[0371] Motion is not constrained to be the same within picture for the same slice header syntax elements
[0372] The syntax elements in this section are currently not constrained to be the same in all slices of a picture. It is suggested to evaluate the intended usage of these syntax elements to determine which can be moved to the picture header to simplify the overall VVC design, as it is claimed that there is a complexity impact to handle a large number of syntax elements in each slice header.
[0373] 1. It is proposed to move the following syntax elements to picture header. Currently there is no restriction for these syntax elements to have different values for different slices, but it is claimed that there is no / good benefit to send them in each slice header and there is coding loss because their intended usage will change at picture level:
[0374] a. six_minus_max_num_merge_cand
[0375] b. five_minus_max_num_subblock_merge_cand
[0376] c. slice_fpel_mmvd_enabled_flag
[0377] d. slice_disable_bdof_dmvr_flag
[0378] e. max_num_merge_cand_minus_max_num_triangle_cand
[0379] f. slice_six_minus_max_num_ibc_merge_cand
[0380] 2. It is proposed to move the following syntax elements to picture header. Currently there is no restriction for these syntax elements to have different values for different slices, but it is claimed that there is no / good benefit to send them in each slice header and there is coding loss because their intended usage will change at picture level:
[0381] a) partition_constraints_override_flag
[0382] b) slice_log2_diff_min_qt_min_cb_luma
[0383] c) slice_max_mtt_hierarchy_depth_luma
[0384] d) slice_log2_diff_max_bt_min_qt_luma
[0385] e) slice_log2_diff_max_tt_min_qt_luma
[0386] f) slice_log2_diff_min_qt_min_cb_chroma
[0387] g) slice_max_mtt_hierarchy_depth_chroma
[0388] h) slice_log2_diff_max_bt_min_qt_chroma
[0389] i) slice_log2_diff_max_tt_min_qt_chroma
[0390] The conditional check "slice_type == I" associated with some of these syntax elements has been removed with the move to picture header.
[0391] 3. It is proposed to move the following syntax elements to the picture header. Currently there is no restriction for these syntax elements to have different values for different slices, but it is claimed that there is no / little benefit to send them in each slice header and there is a coding loss because their intended use will change at picture level:
[0392] a. mvd_l1_zero_flag
[0393] The conditional check "slice_type == B" associated with some of these syntax elements has been removed with the move to picture header.
[0394] 4. It is proposed to move the following syntax elements to the picture header. Currently there is no restriction for these syntax elements to have different values for different slices, but it is claimed that there is no / little benefit to send them in each slice header and there is a coding loss because their intended use will change at picture level:
[0395] a. dep_quant_enabled_flag
[0396] sign_data_hiding_enabled_flag
[0397] 2.10.1. Syntax tables defined in JVET-P1006
[0398] 7.3.2.8 Picture header RBSP syntax
[0399]
[0400]
[0401]
[0402]
[0403]
[0404] 2.11. DMVR in VVC Draft 6
[0405] Decoder-side motion vector refinement (DMVR) derives the motion information of a current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures using bilateral matching (BM). The BM method computes the distortion between two candidate blocks in the reference picture list L0 and list L1. As shown in Figure 1 , the SAD between the red blocks of each MV candidate around the initial MV is computed. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predicted signal. The cost function used in the matching process is the row subsampled SAD (sum of absolute differences). An example of DMVR is shown in Figure 5 .
[0406] In VTM5.0, when a coding unit (CU) is coded using regular merge / skip mode and bi-prediction, DMVR is employed at the decoder to refine the motion vector (MV) of the coding unit (CU), in display order, one reference picture is before the current picture and the other reference picture is after the current picture, the temporal distance between the current picture and one reference picture is equal to the temporal distance between the current picture and the other reference picture, and bi-prediction with CU weights (BCW) selects equal weights. When DMVR is applied, one luma coding block (CB) is divided into several independently processed sub-blocks of size min(cbWidth, 16) x min(cbHeight, 16). DMVR refines the MV of each sub-block by minimizing the SAD between the ½ sub-sampled 10-bit L0 and L1 prediction samples generated by bi-linear interpolation. For each sub-block, an integer AMV search around the initial MV (i.e., the MV of the selected regular merge / skip candidate) is first performed using SAD, and then a fractional AMV derivation is performed to obtain the final MV.
[0407] When a CU is coded using bi-prediction, BDOF refines the luma prediction samples of the CU, one reference picture precedes the current picture in display order and the other reference picture follows the current picture in display order, and BCW selects equal weights. 8-tap interpolation is used to generate initial L0 and L1 prediction samples from the input MVs (e.g. the final MVs of DMVR in case DMVR is enabled). Next, a two-stage early termination process is performed. The first early termination is at sub-block level and the second early termination is at 4x4 block level and is checked when the first early termination does not occur. At each level, first the SAD between the full-sampled 14-bit L0 prediction samples and the L1 prediction samples in each sub-block / 4x4 block is computed. If the SAD is smaller than a threshold, BDOF is not applied to the sub-block / 4x4 block. Otherwise, the BDOF parameters are derived and used to generate the final luma sample prediction values for each 4x4 block. In BDOF, the sub-block size is the same as in DMVR, i.e. min(cbWidth, 16) x min(cbHeight, 16).
[0408] When a CU is coded in regular merge / skip mode, one reference picture precedes the current picture in display order and the other reference picture follows the current picture in display order, the temporal distance between the current picture and one reference picture is equal to the temporal distance between the current picture and the other reference picture, and BCW selects equal weights, both DMVR and BDOF are applied. The flow of the concatenated DMVR process and BDOF process is shown in FIG. 2. Figure 6
[0409] Figure 6 The flow of the concatenated DMVR process and BDOF process in VTM5.0 is shown. The DMVR SAD operation and the BDOF SAD operation are different and not shared.
[0410] To reduce the latency and operations in this critical path, when both DMVR and BDOF are applied, the latest VVC working draft has been revised to reuse the sub-block SAD computed in DMVR for the sub-block early termination in BDOF.
[0411] The SAD computation is defined as follows:
[0412]
[0413] where the two variables nSbW and nSbH specify the width and height of the current sub-block, the two (nSbW+4)x(nSbH+4) arrays pL0 and pL1 contain the prediction samples of L0 and L1, respectively, and (dX, dY) is the integer sample offset in the prediction list L0.
[0414] To reduce the penalty of the uncertainty of DMVR refinement, support of original MV during the DMVR process is proposed. The SAD between the reference blocks referred by the initial (or called original) MV candidate is reduced by 1 / 4 of the SAD value. That is, when both dX and dY in the above equation are equal to 0, the value of sad is modified as follows:
[0415] sad = sad - (sad » 2)
[0416] When the SAD value is smaller than a threshold (2*subblock width*subblock height), there is no need to perform BDOF anymore.
[0417] 3. Drawbacks of existing implementations
[0418] DMVR and BIO do not involve the original signal during the refinement of the motion vector, which can lead to the coded block having inaccurate motion information. In addition, DMVR and BIO sometimes employ fractional motion vectors after motion refinement, while screen video usually has integer motion vectors, which makes the current motion information even less accurate and makes the coding performance worse.
[0419] When RPR is applied in VVC, RPR (ARC) can have the following problems:
[0420] 1. With RPR, the interpolation filter can be different for neighboring samples in a block, which is not desirable in SIMD (Single Instruction Multiple Data) implementations.
[0421] 2. Border region does not consider RPR
[0422] 3. Note that "Consistency cropping window offset parameters are only applied at the output. All internal decoding processes are applied to the uncropped picture size." However, when RPR is applied, these parameters can be used in the decoding process.
[0423] 4. When deriving the reference sample positions, RPR only considers the ratio between two consistency windows. However, the top-left offset difference between two consistency windows should also be considered.
[0424] 5. The ratio between the width / height of the reference picture and the width / height of the current picture is constrained in VVC. However, the ratio between the width / height of the consistency window of the reference picture and the width / height of the consistency window of the current picture is not constrained.
[0425] 6. Not all syntax elements are correctly handled in the picture header.
[0426] 7. In current VVC, for TPM and GEO prediction modes, chroma mixing weights are derived regardless of the chroma sample location type of the video sequence. For example, in TPM / GEO, if the chroma weights are derived from the luma weights, it may be necessary to downsample the luma weights to match the sampling of the chroma signal. Chroma downsampling is typically applied assuming a chroma sample location type of 0, which is widely used in ITU-R BT.601 or ITU-R BT.709 containers. However, if different chroma sample location types are used, this can lead to misalignment between the chroma samples and the downsampled luma samples, potentially degrading encoding / decoding performance.
[0427] 8. Note that the SAD calculation / SAD threshold does not take into account the effect of bit depth. Therefore, for higher bit depths (e.g., 14 or 16-bit input sequences), the threshold for early termination may be too small.
[0428] 9. For non-RPR cases, AMVR (i.e., alternative interpolation filter / switchable interpolation filter) with 1 / 2 pixel MV accuracy is applied together with a 6-tap motion compensation filter, but an 8-tap filter is applied for other cases (e.g., 1 / 16 pixel). However, for RPR cases, the same interpolation filter is applied for all cases, regardless of MV / MVD accuracy. Therefore, signaling for the 1 / 2 pixel case (alternative interpolation filter / switchable interpolation filter) is a waste of bits.
[0429] 10. The decision on whether to allow split tree partitioning depends on the encoding / decoding image resolution, not the output image resolution.
[0430] 11. SMVD / MMVD applications do not consider the RPR case. These methods are based on the assumption of applying symmetric MVD to two reference images. However, this assumption does not hold true when the output images have different resolutions.
[0431] 12. Paired merge candidates are generated by averaging the two MVs of two merge candidates from the same list of reference images. However, averaging is meaningless when the two reference images associated with two merge candidates have different resolutions.
[0432] 13. If all stripes in the current image are I (intra-frame) stripes, it may not be necessary to encode and decode several inter-strip related syntax elements in the image header. Conditionally notifying them with signaling can save syntax overhead, especially for low-resolution sequences with all intra-frame encoding and decoding.
[0433] 14. In current VVC, there is no restriction on the dimension of a slice / tile. Adding proper restriction helps parallel processing of real-time software / hardware decoder, especially for ultra-high resolution sequences that can be larger than 4K / 8K per frame.
[0434] 4. Example techniques and embodiments
[0435] The detailed embodiments described below should be considered as examples to explain the general concepts. These embodiments should not be interpreted in a narrow way. Furthermore, these embodiments can be combined in any way.
[0436] The methods described below can also be applicable to other decoder motion information derivation techniques, in addition to DMVR and BIO mentioned below.
[0437] A motion vector is represented by (mv_x, mv_y), where mv_x is the horizontal component and mv_y is the vertical component.
[0438] In this disclosure, the resolution (or dimension, or width / height, or size) of a picture can refer to the resolution (or dimension, or width / height, or size) of a coded / decoded picture, or can refer to the resolution (or dimension, or width / height, or size) of a consistent window in a coded / decoded picture. In one example, the resolution (or dimension, or width / height, or size) of a picture can refer to parameters related to the RPR (Reference Picture Resampling) process, such as the scaling window / phase offset window. In one example, the resolution (or dimension, or width / height, or size) of a picture is related to the resolution associated with an output picture.
[0439] Motion compensation in RPR
[0440] 1. When the resolution of the reference picture is different from the current picture, or when the width and / or height of the reference picture is larger than the width and / or height of the current picture, the prediction value of a set of samples (at least two samples) of the current block can be generated with the same horizontal and / or vertical interpolation filter.
[0441] a. In one example, the set can include all samples in a region of the block.
[0442] i. For example, the block can be divided into S MxN rectangles that do not overlap with each other. Each MxN rectangle is a set. Figure 2 In the illustrated example, a 16x16 block can be divided into 16 4x4 rectangles, each of which is a set.
[0443] ii. For example, a row with N samples is a set. N is an integer that is not larger than the width of the block. In one example, N is 4 or 8 or the width of the block.
[0444] iii. For example, a column with N samples is a group. N is an integer not greater than the block height. In one example, N is 4 or 8 or the height of the block.
[0445] iv. M and / or N can be predefined or derived on-the-fly, such as based on block dimension / coding information or signaled.
[0446] b. In one example, the samples in a group can have the same MV (denoted as shared MV).
[0447] c. In one example, the samples in a group can have MVs with the same horizontal component (denoted as shared horizontal component).
[0448] d. In one example, the samples in a group can have MVs with the same vertical component (denoted as shared vertical component).
[0449] e. In one example, the samples in a group can have MVs with the same fractional part of the horizontal component (denoted as shared fractional horizontal component).
[0450] i. For example, assume the MV of the first sample is (MV1x, MV1y), and the MV of the second sample is (MV2x, MV2y), it should be satisfied that MV1x & (2 M - 1) is equal to MV2x & (2 M - 1), where M denotes the MV precision. For example, M = 4.
[0451] f. In one example, the samples in a group can have MVs with the same fractional part of the vertical component (denoted as shared fractional vertical component).
[0452] i. For example, assume the MV of the first sample is (MV1x, MV1y), and the MV of the second sample is (MV2x, MV2y), it should be satisfied that MV1y & (2 M - 1) is equal to MV2y & (2 M - 1), where M denotes the MV precision. For example, M = 4.
[0453] g. In one example, for the samples in a group to be predicted, the motion vector can be first derived according to the resolution of the current picture and the reference picture (e.g., derived in 8.5.6.3.1 in JVET-O2001-v14 (refx L , refy L )), denoted as MV b . Then, MV bis further modified (e.g., rounded / truncated / clipped) to MV' to satisfy requirements such as the bullet points above, and MV' will be used to derive the predicted samples for this sample.
[0454] i. In one example, MV' has the same integer part as MV b , and the fractional part of MV' is set to the shared fractional horizontal component and / or fractional vertical component.
[0455] ii. In one example, MV' is set to a value that has the shared fractional horizontal component and / or fractional vertical component, and is closest to MV b .
[0456] h. The shared motion vector (and / or shared horizontal component and / or shared vertical component and / or shared fractional horizontal component and / or shared fractional vertical component) can be set to the motion vector (and / or horizontal component and / or vertical component and / or fractional horizontal component and / or fractional vertical component) of a particular sample in the group.
[0457] i. For example, the particular sample can be located at a corner of a rectangular group, such as "A", "B", "C", and "D" shown in Figure 3A .
[0458] ii. For example, the particular sample can be located at the center of a rectangular group, such as "E", "F", "G", and "H" shown in Figure 3A .
[0459] iii. For example, the particular sample can be located at an end of a row-wise or column-wise group, such as "A" and "D" shown in Figure 3B and Figure 3C .
[0460] iv. For example, the particular sample can be located in the middle of a row-wise or column-wise group, such as "B" and "C" shown in Figure 3B and Figure 3C .
[0461] v. In one example, the motion vector of the particular sample can be MV b .
[0462] i. The shared motion vector (and / or shared horizontal component and / or shared vertical component and / or shared fractional vertical component and / or shared fractional vertical component) can be set to the motion vector (and / or horizontal component and / or vertical component and / or fractional horizontal component and / or fractional vertical component) of a virtual sample that is located at a different position compared to all samples in the group.
[0463] i. In one example, the virtual sample is not in the set, but it is located in an area covering all samples in the set.
[0464] 1) Alternatively, the virtual sample is located outside the area covering all samples in the set, e.g. at the right bottom position next to the area.
[0465] ii. In one example, the MV of the virtual sample is derived in the same way as the real samples, but with a different position.
[0466] iii. Figures 3A-3C The "V" in shows three examples of virtual samples.
[0467] j. The shared MV (and / or shared horizontal component and / or shared vertical component and / or shared fractional horizontal component and / or shared fractional vertical component) can be set as a function of the MVs (and / or horizontal components and / or vertical components and / or fractional horizontal components and / or fractional vertical components) of the multiple samples and / or virtual samples.
[0468] i. For example, the shared MV (and / or shared horizontal component and / or shared vertical component and / or shared horizontal fractional component and / or shared vertical fractional component) can be set as the average of the MVs (and / or horizontal components and / or vertical components and / or fractional horizontal components and / or fractional vertical components) of all or some of the samples in the set, or Figure 3A the average of the MVs of samples "E", "F", "G", "H" in Figure 3A the average of the MVs of samples "E", "H" in Figure 3A the average of the MVs of samples "A", "B", "C", "D" in, or Figure 3A the average of the MVs of samples "A", "D" in, or Figure 3B the average of the MVs of samples "B", "C" in, or Figure 3A the average of the MVs of samples "A", "D" in, or Figure 3C the average of the MVs of samples "B", "C" in, or Figure 3C the average of the MVs of samples "A", "D" in,
[0469] 2. It is proposed to allow only integer MVs to perform the motion compensation process to derive the prediction block of a current block when the resolution of the reference picture is different from the current picture, or when the width and / or height of the reference picture is larger than the width and / or height of the current picture.
[0470] a. In one example, the decoded motion vector of the sample to be predicted is rounded to an integer MV before use.
[0471] b. In one example, the decoded motion vector of the sample to be predicted is rounded to the integer MV closest to the decoded motion vector.
[0472] c. In one example, the decoded motion vector of the sample to be predicted is rounded to the integer MV closest to the decoded motion vector in horizontal direction.
[0473] d. In one example, the decoded motion vector of the sample to be predicted is rounded to the integer MV closest to the decoded motion vector in vertical direction.
[0474] 3. The motion vector used in the motion compensation process for the sample in the current block (e.g. the shared MV / shared horizontal or vertical or fractional component / MV' mentioned in the above bullets) can be stored in the decoded picture buffer and used for motion vector prediction of subsequent blocks in the current / different picture.
[0475] a. Alternatively, the motion vector used in the motion compensation process for the sample in the current block (e.g. the shared MV / shared horizontal or vertical or fractional component / MV' mentioned in the above bullets) can not be allowed to be used for motion vector prediction of subsequent blocks in the current / different picture.
[0476] i. In one example, the decoded motion vector (e.g. the MV in the above bullets b ) can be used for motion vector prediction of subsequent blocks in the current / different picture.
[0477] b. In one example, the motion vector used in the motion compensation process for the sample in the current block can be used in a filtering process (e.g. deblocking filter / SAO / ALF).
[0478] i. Alternatively, the decoded motion vector (e.g. the MV in the above bullets b ) can be used in a filtering process.
[0479] c. In one example, such MVs can be derived at sub-block level and can be stored for each sub-block.
[0480] 4. It is proposed that an interpolation filter can be selected for use in a motion compensation process for deriving a prediction block for a current block depending on whether the resolution of the reference picture is different from the current picture or whether the width and / or height of the reference picture is larger than the width and / or height of the current picture.
[0481] a. In one example, when condition A is fulfilled, an interpolation filter with less taps can be applied, where condition A depends on the dimensions of the current picture and / or the reference picture.
[0482] i. In one example, condition A is that the resolution of the reference picture is different from the current picture.
[0483] ii. In one example, condition A is that the width and / or height of the reference picture is larger than the width and / or height of the current picture.
[0484] iii. In one example, condition A is W1 > a*W2 and / or H1 > b*H2, where (W1, H1 ) denotes the width and height of the reference picture, and (W2, H2) denotes the width and height of the current picture, a and b are two factors, e.g. a = b = 1.5.
[0485] iv. In one example, condition A can also depend on whether bi-prediction is used or not.
[0486] v. In one example, a 1 -tap filter is applied. In other words, the output is the unfiltered integer pixel as the interpolation result.
[0487] vi. In one example, a bi-linear filter is applied when the resolution of the reference picture is different from the current picture.
[0488] vii. In one example, a 4-tap filter or a 6-tap filter is applied when the resolution of the reference picture is different from the current picture, or the width and / or height of the reference picture is larger than the width and / or height of the current picture.
[0489] 1) A 6-tap filter can also be used for affine motion compensation.
[0490] 2) A 4-tap filter can also be used for the interpolation of chroma samples.
[0491] b. In one example, padded samples are used to perform the interpolation when the resolution of the reference picture is different from the current picture, or the width and / or height of the reference picture is larger than the width and / or height of the current picture.
[0492] c. Whether and / or how the methods disclosed in item 4 are applied can depend on the color component.
[0493] i. For example, the methods are only applied to the luma component.
[0494] d. Whether and / or how the methods disclosed in item 4 are applied can depend on the interpolation filter direction.
[0495] i. For example, the methods are only applied to horizontal filtering.
[0496] ii. For example, the methods are only applied to vertical filtering.
[0497] 5. It is proposed to apply a two-stage process for the generation of prediction blocks when the resolution of the reference picture is different from the current picture, or when the width and / or height of the reference picture is greater than the width and / or height of the current picture.
[0498] a. In the first stage, depending on the width and / or height of the current picture and of the reference picture, a virtual reference block is generated by upsampling or downsampling a region in the reference picture.
[0499] b. In the second stage, independently of the width and / or height of the current picture and of the reference picture, prediction samples are generated from the virtual reference block by applying an interpolation filter.
[0500] 6. It is proposed that the computation of the top-left coordinates (xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding defined in 8.5.6.3.1 of JVET-O2001-v14 can be derived depending on the width and / or height of the current picture and of the reference picture.
[0501] a. In one example, the luma position in full-sample units is modified as:
[0502] xInt i = Clip3( xSbInt L - Dx, xSbInt L + sbWidth + Ux, xInt i ),
[0503] yInt i = Clip3( ySbInt L - Dy, ySbInt L + sbHeight + Uy, yInt i ),
[0504] where Dx and / or Dy and / or Ux and / or Uy can depend on the width and / or height of the current picture and of the reference picture.
[0505] b. In one example, the chroma position in full-sample units is modified as:
[0506] xInti = Clip3( xSbInt C - Dx, xSbInt C + sbWidth + Ux, xInti)
[0507] yInti = Clip3( ySbInt C - Dy, ySbInt C + sbHeight + Uy, yInti)
[0508] where Dx and / or Dy and / or Ux and / or Uy can depend on the width and / or height of the current picture and the reference picture.
[0509] 7. Instead of storing / using the motion vector of a block based on the same resolution of the reference picture as the current picture, it is proposed to use the real motion vector that takes into account the resolution difference.
[0510] a. Alternatively, in addition, when using the motion vector to generate the prediction block, there is no need to further change the motion vector according to the resolution of the current picture and the reference picture (e.g., (refx L ,refy L ) derived in 8.5.6.3.1 in JVET-O2001-v14).
[0511] Interaction between RPR and other coding tools
[0512] 8. Whether / how to apply a filtering process (e.g., deblocking filter) can depend on the resolution of the reference picture and / or the resolution of the current picture.
[0513] a. In one example, in addition to the motion vector difference, the boundary strength (BS) setting in the deblocking filter can also take into account the resolution difference.
[0514] i. In one example, the scaled motion vector difference according to the resolution of the current picture and the reference picture can be used to determine the boundary strength.
[0515] b. In one example, if the resolution of at least one reference picture of block A is different from (or smaller than or larger than) the resolution of at least one reference picture of block B, the strength of the deblocking filter for the boundary between block A and block B can be set differently (e.g., increased / decreased) compared to the case where the same resolution is used for both blocks.
[0516] c. In one example, if the resolution of at least one reference picture of block A is different from (or smaller than or larger than) the resolution of at least one reference picture of block B, the boundary between block A and block B is marked to be filtered (e.g., BS is set to 2)
[0517] d. In one example, if the resolution of at least one reference picture of block A and / or block B is different from (or smaller than or larger than) the resolution of the current picture, the strength of the deblocking filter for the boundary between block A and block B can be set differently (e.g., increased / decreased) compared to the case where the same resolution is used for the reference picture and the current picture.
[0518] e.In one example, if at least one reference picture of at least one of the two blocks has a different resolution than the current picture, the boundary between the two blocks is marked to be filtered (e.g., BS is set to 2).
[0519] 9.When sub-pictures are present, a conformant bitstream can satisfy that a reference picture must have the same resolution as the current picture.
[0520] a.Alternatively, when a reference picture has a different resolution than the current picture, there must be no sub-pictures in the current picture.
[0521] b.Alternatively, for a sub-picture in the current picture, it is not allowed to use a reference picture that has a different resolution than the current picture.
[0522] i.Alternatively, in addition, reference picture management can be invoked to exclude those reference pictures that have a different resolution.
[0523] 10.In one example, sub-pictures can be defined separately for pictures that have different resolutions (e.g., how to divide one picture into multiple sub-pictures).
[0524] In one example, if a reference picture has a different resolution than the current picture, a corresponding sub-picture in the reference picture can be derived by scaling and / or offsetting a sub-picture of the current picture.
[0525] 11.When a reference picture has a different resolution than the current picture, PROF (prediction refinement with optical flow) can be enabled.
[0526] a.In one example, a set of MVs (denoted as MV g ) can be generated for a group of samples and can be used for motion compensation as in item 1. On the other hand, MVs (denoted as MV p ) can be derived for each sample and the difference between MV p and MV g (e.g., corresponding to Δv used in PROF) as well as the gradient (e.g., spatial gradient of the motion compensated block) can be used to derive the prediction accuracy.
[0527] b.In one example, MV p can have a different precision than MV g . For example, MV p can have a precision of 1 / N pixels (N>0), N=32, 64, etc.
[0528] c.In one example, MV gIt can have a different precision than the internal MV precision (e.g. 1 / 16 pixel).
[0529] d. In one example, the prediction refinement is added to the prediction block to generate a refined prediction block.
[0530] e. In one example, this approach can be applied for each prediction direction.
[0531] f. In one example, this approach can be applied only for the single prediction case.
[0532] g. In one example, this approach can be applied for single prediction or / and bi-prediction.
[0533] h. In one example, this approach can be applied only when the reference picture has a different resolution than the current picture.
[0534] 12. It is proposed that when the resolution of the reference picture is different than the resolution of the current picture, the motion compensation process can be performed for a block / sub-block with only one MV to derive the prediction block for the current block.
[0535] a. In one example, the only one MV for a block / sub-block can be defined as a function (e.g. average) of all MVs associated with each sample within the block / sub-block.
[0536] b. In one example, the only one MV for a block / sub-block can be defined as a selected MV associated with a selected sample (e.g. center sample) within the block / sub-block.
[0537] c. In one example, only one MV can be utilized for 4x4 block or sub-block (e.g. 4x1).
[0538] d. In one example, BIO can be further applied to compensate the precision loss due to block-based motion vector.
[0539] 13. When the width and / or height of the reference picture is different than the width and / or height of the current picture, a lazy mode without signaling any block-based motion vector can be applied.
[0540] a. In one example, the motion vector can not be signaled and the motion compensation process is approximated to the case of pure resolution change for still image.
[0541] b. In one example, when the resolution changes, only the picture / slice / tile / CTU level motion vector can be signaled and the related blocks can use this motion vector.
[0542] 14. For blocks coded with affine and / or non-affine prediction mode, PROF can be applied to approximate motion compensation when the width and / or height of the reference picture is different from the width and / or height of the current picture.
[0543] a. In one example, PROF can be enabled when the width and / or height of the reference picture is different from the width and / or height of the current picture.
[0544] b. In one example, the affine motion set can be generated by combining the indicated motion and resolution scaling, and used by PROF.
[0545] 15. Interweaving prediction (e.g., proposed in JVET-K0102) can be applied to approximate motion compensation when the width and / or height of the reference picture is different from the width and / or height of the current picture.
[0546] a. In one example, resolution change (zooming) is represented as affine motion, and interweaving motion prediction can be applied.
[0547] 16. LMCS and / or chroma residual scaling can be disabled when the width and / or height of the current picture is different from the width and / or height of the IRAP picture in the same IRAP period.
[0548] a. In one example, when LMCS is disabled, the slice-level flags (such as slice_lmcs_enabled_flag, slice_lmcs_aps_id, and slice_chroma_residual_scale_flag) can not be signaled and inferred to be 0.
[0549] b. In one example, when chroma residual scaling is disabled, the slice-level flag (such as slice_chroma_residual_scale_flag) can not be signaled and inferred to be 0.
[0550] Constraints on RPR
[0551] 17. RPR can be applied to coded blocks with block dimension constraints.
[0552] a. In one example, for MxN coded blocks, where M is the block width and N is the block height, RPR can not be used when M*N < T or M*N <= T (such as T = 256).
[0553] b. In one example, when M < K (or M <= K) (such as K = 16) and / or N < L (or N <= L) (such as L = 16), RPR can not be used.
[0554] 18. Bitstream conformance can be added to limit the ratio between the width and / or height of an active reference picture (or its conformance window) and the width and / or height of the current picture (or its conformance window). Assuming refPicW and refPicH represent the width and height of a reference picture, curPicW and curPicH represent the width and height of the current picture,
[0555] a. In one example, when (refPicW ÷ curPicW) is equal to an integer, the reference picture can be marked as an active reference picture.
[0556] i. Alternatively, when (refPicW ÷ curPicW) is equal to a fraction, the reference picture can be marked as unavailable.
[0557] b. In one example, when (refPicW ÷ curPicW) is equal to (X*n), where X represents a fraction (such as X = 1 / 2) and n represents an integer (such as n = 1, 2, 3, 4,...), the reference picture can be marked as an active reference picture.
[0558] i. In one example, when (refPicW ÷ curPicW) is not equal to (X*n), the reference picture can be marked as unavailable.
[0559] 19. Whether and / or how to enable a coding tool (e.g., bi-prediction / full triangle prediction mode (TPM) / hybrid process in TPM) for an MxN block can depend on the resolution of a reference picture (or its conformance window) and / or the resolution of the current picture (or its conformance window).
[0560] a. In one example, M*N < T or M*N <= T (such as T = 64).
[0561] b. In one example, M < K (or M <= K) (such as K = 16) and / or N < L (or N <= L) (such as L = 16).
[0562] c. In one example, when the width / height of at least one reference picture is different from the current picture, the coding tool is not allowed,
[0563] i. In one example, when the width / height of at least one reference picture of a block is greater than the width / height of the current picture, the coding tool is not allowed.
[0564] d. In one example, the coding tool is not allowed when the width / height of each reference picture is different from the width / height of the current picture.
[0565] i. In one example, the coding tool is not allowed when the width / height of each reference picture is greater than the width / height of the current picture.
[0566] e. Alternatively, in addition, when the coding tool is not allowed to be used, one MV can be utilized as uni-prediction for motion compensation.
[0567] Correlation consistency window
[0568] 20. Signaling the consistent clipping window offset parameter (e.g., conf_win_left_offset) with N pixel precision instead of 1 pixel precision, where N is a positive integer greater than 1.
[0569] a. In one example, the actual offset can be derived as the signaled offset multiplied by N.
[0570] b. In one example, N is set to 4 or 8.
[0571] 21. It is proposed that the consistent clipping window offset parameter is not only applied to output. Some internal decoding processes can depend on the clipped picture size (i.e., the resolution of the consistent window in the picture).
[0572] 22. It is proposed that the consistent clipping window offset parameter in a first video unit (e.g., PPS) and a second video unit can be different when the width and / or height of the picture represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are the same.
[0573] 23. It is proposed that the consistent clipping window offset parameter in a first video unit (e.g., PPS) and a second video unit should be the same in the consistent bitstream when the width and / or height of the picture represented as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and the second video unit are different.
[0574] a. It is proposed that the conformance clipping window offset parameters in a first video unit (e.g., PPS) and in a second video unit in a conformance bitstream should be the same, regardless of whether the width and / or height of the picture denoted as (pic_width_in_luma_samples, pic_height_in_luma_samples) in the first video unit and in the second video unit are the same.
[0575] 24. Assume that the width and height of the conformance window defined in a first video unit (e.g., PPS) are denoted as W1 and H1, respectively. The width and height of the conformance window defined in a second video unit (e.g., PPS) are denoted as W2 and H2, respectively. The top-left position of the conformance window defined in the first video unit (e.g., PPS) is denoted as X1 and Y1. The top-left position of the conformance window defined in the second video unit (e.g., PPS) is denoted as X2 and Y2. The width and height of the coded / decoded picture (e.g., pic_width_in_luma_samples and pic_height_in_luma_samples) defined in the first video unit (e.g., PPS) are denoted as PW1 and PH1, respectively. The width and height of the coded / decoded picture defined in the second video unit (e.g., PPS) are denoted as PW2 and PH2, respectively.
[0576] a. In one example, in a conformance bitstream, W1 / W2 should be equal to X1 / X2.
[0577] i. Alternatively, in a conformance bitstream, W1 / X1 should be equal to W2 / X2.
[0578] ii. Alternatively, in a conformance bitstream, W1*X2 should be equal to W2*X1.
[0579] b. In one example, in a conformance bitstream, H1 / H2 should be equal to Y1 / Y2.
[0580] i. Alternatively, in a conformance bitstream, H1 / Y1 should be equal to H2 / Y2.
[0581] ii. Alternatively, in a conformance bitstream, H1*Y2 should be equal to H2*Y1.
[0582] c. In one example, in a conformance bitstream, PW1 / PW2 should be equal to X1 / X2.
[0583] i. Alternatively, in a conformance bitstream, PW1 / X1 should be equal to PW2 / X2.
[0584] ii. Alternatively, in a consistent bitstream, PW1*X2 should equal PW2*X1.
[0585] d. In one example, in a consistent bitstream, PH1 / PH2 should equal Y1 / Y2.
[0586] i. Alternatively, in a consistent bitstream, PH1 / Y1 should equal PH2 / Y2.
[0587] ii. Alternatively, in a consistent bitstream, PH1*Y2 should equal PH2*Y1.
[0588] e. In one example, in a consistent bitstream, PW1 / PW2 should equal W1 / W2.
[0589] i. Alternatively, in a consistent bitstream, PW1 / W1 should equal PW2 / W2.
[0590] ii. Alternatively, in a consistent bitstream, PW1*W2 should equal PW2*W1.
[0591] f. In one example, in a consistent bitstream, PH1 / PH2 should equal H1 / H2.
[0592] i. Alternatively, in a consistent bitstream, PH1 / H1 should equal PH2 / H2.
[0593] ii. Alternatively, in a consistent bitstream, PH1*H2 should equal PH2*H1.
[0594] g. In a consistent bitstream, if PW1 is greater than PW2, then W1 must be greater than W2.
[0595] h. In a consistent bitstream, if PW1 is less than PW2, then W1 must be less than W2.
[0596] i. In a consistent bitstream, (PW1-PW2)*(W1-W2) must not be less than 0.
[0597] j. In a consistent bitstream, if PH1 is greater than PH2, then H1 must be greater than H2.
[0598] k. In a consistent bitstream, if PH1 is less than PH2, then H1 must be less than H2.
[0599] l. In a consistent bitstream, (PH1-PH2)*(H1-H2) must not be less than 0.
[0600] m. In a consistent bitstream, if PW1 >= PW2, then W1 / W2 must not be greater than (or less than) PW1 / PW2.
[0601] n. In a conformance bitstream, if PH1 >= PH2, then H1 / H2 must not be greater than (or less than) PH1 / PH2.
[0602] 25. Assume that the width and height of the conformance window of the current picture are denoted as W and H, respectively. The width and height of the conformance window of the reference picture are denoted as W' and H', respectively. Then at least one of the following constraints should be followed by the conformance bitstream.
[0603] a. W*pw >= W'; pw is an integer such as 2.
[0604] b. W*pw > W'; pw is an integer such as 2.
[0605] c. W'*pw' >= W; pw' is an integer such as 8.
[0606] d. W'*pw' > W; pw' is an integer such as 8.
[0607] e. H*ph >= H'; ph is an integer such as 2.
[0608] f. H*ph > H'; ph is an integer such as 2.
[0609] g. H'*ph' >= H; ph' is an integer such as 8.
[0610] h. H'*ph' > H; ph' is an integer such as 8.
[0611] i. In one example, pw is equal to pw'.
[0612] j. In one example, ph is equal to ph'.
[0613] k. In one example, pw is equal to ph
[0614] l. In one example, pw' is equal to ph'.
[0615] m. In one example, when W and H denote the width and height of the current picture, respectively, it can be required that the conformance bitstream satisfies the above sub-items modulo. W' and H' denote the width and height of the reference picture.
[0616] 26. It is proposed to partially signal the conformance window parameters.
[0617] a. In one example, the top-left sample in the conformance window of a picture is the same as in the picture.
[0618] b. For example, conf_win_left_offset defined in VVC is not signaled and is inferred to be 0.
[0619] c. For example, conf_win_top_offset defined in VVC is not signaled and is inferred to be 0.
[0620] 27. It is proposed that the derivation of the position of the reference sample (e.g., (refx L , refy L ) as defined in VVC) can depend on the top-left position of the conformance window of the current picture and / or the reference picture (e.g., (conf_win_left_offset, conf_win_top_offset) as defined in VVC). Figure 4A and Figure 4B Examples of the derived sample positions in VVC (a) and the proposed method (b) are shown. The dashed rectangle represents the conformance window.
[0621] a. In one example, the dependency exists only when the width and / or height of the current picture is different from the width and / or height of the reference picture.
[0622] b. In one example, the derivation of the horizontal position of the reference sample (e.g., refx L as defined in VVC) can depend on the left position of the conformance window of the current picture and / or the reference picture (e.g., conf_win_left_offset as defined in VVC).
[0623] i. In one example, the horizontal position of the current sample in the current picture relative to the top-left position of the conformance window is calculated (denoted as xSb’) and is used for deriving the position of the reference sample.
[0624] 1) For example, xSb’ = xSb – (conf_win_left_offset << Prec) is calculated and is used for deriving the position of the reference sample, where xSb represents the horizontal position of the current sample in the current picture. conf_win_left_offset represents the horizontal position of the top-left sample in the conformance window of the current picture. Prec represents the precision of xSb and xSb’, where (xSb >> Prec) can show the actual horizontal coordinate of the current sample relative to the current picture. For example, Prec = 0 or Prec = 4.
[0625] ii. In one example, the horizontal position of the reference sample in the reference picture relative to the top-left position of the conformance window is calculated (denoted as Rx’).
[0626] 1) The computation of Rx’ can depend on xSb’ and / or the motion vector and / or the resampling ratio.
[0627] iii. In one example, the relative horizontal position in the reference picture of the reference sample (denoted as Rx) is computed depending on Rx’.
[0628] 1) For example, Rx = Rx’ + (conf_win_left_offset_ref « Prec) is computed, where conf_win_left_offset_ref denotes the horizontal position of the top-left sample in the conformance window of the reference picture. Prec denotes the precision of Rx and Rx’. For example, Prec = 0 or Prec = 4.
[0629] iv. In one example, Rx can be computed directly depending on xSb’ and / or the motion vector and / or the resampling ratio. In other words, the two steps of the derivation of Rx’ and Rx are combined into one step computation.
[0630] v. Whether and / or how to use the left position of the conformance window of the current picture and / or the reference picture (e.g., conf_win_left_offset as defined in VVC) can depend on the color component and / or the color format.
[0631] 1) For example, conf_win_left_offset can be revised as conf_win_left_offset = conf_win_left_offset * SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example, SubWidthC is equal to 1 for the luma component. SubWidthC is equal to 2 for the chroma component when the color format is 4:2:0 or 4:2:2.
[0632] 2) For example, conf_win_left_offset can be revised as conf_win_left_offset = conf_win_left_offset / SubWidthC, where SubWidthC defines the horizontal sampling step of the color component. For example, SubWidthC is equal to 1 for the luma component. SubWidthC is equal to 2 for the chroma component when the color format is 4:2:0 or 4:2:2.
[0633] c. In one example, the derivation of the vertical position of the reference sample (e.g., refy L ) can depend on the top position of the conformance window of the current picture and / or the reference picture (e.g., conf_win_top_offset as defined in VVC).
[0634] i. In one example, the vertical position of the current sample in the current picture relative to the top-left position of the conformance window is computed (denoted as ySb’) and used to derive the position of the reference sample.
[0635] 1) For example, ySb’ = ySb - (conf_win_top_offset << Prec) is computed and used to derive the position of the reference sample, where ySb denotes the vertical position of the current sample in the current picture. conf_win_top_offset denotes the vertical position of the top-left sample in the conformance window of the current picture. Prec denotes the precision of ySb and ySb’.
[0636] For example, Prec = 0 or Prec = 4.
[0637] ii. In one example, the vertical position of the reference sample in the reference picture relative to the top-left position of the conformance window is computed (denoted as Ry’).
[0638] 1) The computation of Ry’ can depend on ySb’ and / or the motion vector and / or the resampling ratio.
[0639] iii. In one example, the vertical position of the reference sample in the reference picture relative to the top-left position of the conformance window is computed (denoted as Ry) depending on Ry’.
[0640] 1) For example, Ry = Ry’ + (conf_win_top_offset_ref << Prec) is computed, where conf_win_top_offset_ref denotes the vertical position of the top-left sample in the conformance window of the reference picture. Prec denotes the precision of Ry and Ry’. For example, Prec = 0 or Prec = 4.
[0641] iv. In one example, Ry can be computed directly depending on yS’ and / or the motion vector and / or the resampling ratio. In other words, the two steps of derivation for Ry’ and Ry are combined into one step of computation.
[0642] v. Whether and / or how to use the top position of the conformance window of the current picture and / or of the reference picture (e.g. conf_win_top_offset defined in VVC) depends on the color component and / or the color format.
[0643] 1) For example, conf win top offset can be revised as conf win top offset = conf win top offset * SubHeightC, where SubHeightC defines the vertical sampling step size of the color component. For example, for the luma component, SubHeightC is equal to 1. When the color format is 4:2:0, for the chroma components, SubHeightC is equal to 2.
[0644] 2) For example, conf win top offset can be revised as conf win top offset = conf win top offset / SubHeightC, where SubHeightC defines the vertical sampling step size of the color component. For example, for the luma component, SubHeightC is equal to 1. When the color format is 4:2:0, for the chroma components, SubHeightC is equal to 2.
[0645] 28. It is proposed that the integer part of the horizontal coordinate of the reference sample can be clipped to [minW, maxW]. Let the width and height of the consistent window of the reference picture be denoted as W and H, respectively. Let the width and height of the consistent window of the reference picture be denoted as W' and H'. Let the top-left position of the consistent window in the reference picture be denoted as (X0, Y0).
[0646] a. In one example, minW is equal to 0.
[0647] b. In one example, minW is equal to X0.
[0648] c. In one example, maxW is equal to W-1.
[0649] d. In one example, maxW is equal to W'-1.
[0650] e. In one example, maxW is equal to X0+W'-1.
[0651] f. In one example, minW and / or maxW can be modified based on the color format and / or the color component.
[0652] i. For example, minW is modified as minW * SubC.
[0653] ii. For example, minW is modified as minW / SubC.
[0654] iii. For example, maxW is modified as maxW * SubC.
[0655] iv. For example, maxW is modified as maxW / SubC.
[0656] v. In one example, for the luma component, SubC is equal to 1.
[0657] vi. In one example, for the chroma components, SubC is equal to 2 when the color format is 4:2:0.
[0658] vii. In one example, for the chroma components, SubC is equal to 2 when the color format is 4:2:2.
[0659] viii. In one example, for the chroma components, SubC is equal to 1 when the color format is 4:4:4.
[0660] g. In one example, whether and / or how clipping is performed can depend on the size of the current picture (or a consistency window therein) and the size of the reference picture (or a consistency window therein).
[0661] i. In one example, clipping is performed only when the size of the current picture (or a consistency window therein) and the size of the reference picture (or a consistency window therein) are different.
[0662] 29. It is proposed that the integer part of the vertical coordinate of the reference sample can be clipped to [minH, maxH]. Let the width and height of the consistency window of the reference picture be denoted as W and H, respectively. Let the width and height of the consistency window of the reference picture be denoted as W' and H'. Let the top-left position of the consistency window in the reference picture be denoted as (X0, Y0).
[0663] a. In one example, minH is equal to 0.
[0664] b. In one example, minH is equal to Y0.
[0665] c. In one example, maxH is equal to H - 1.
[0666] d. In one example, maxH is equal to H' - 1.
[0667] e. In one example, maxH is equal to Y0 + H' - 1.
[0668] f. In one example, minH and / or maxH can be modified based on the color format and / or the color component.
[0669] i. For example, minH is modified to minH * SubC.
[0670] ii. For example, minH is modified to minH / SubC.
[0671] iii. For example, maxH is modified as maxH * SubC.
[0672] iv. For example, maxH is modified as maxH / SubC.
[0673] v. In one example, for luma components, SubC is equal to 1.
[0674] vi. In one example, for chroma components, SubC is equal to 2 when the color format is 4:2:0.
[0675] vii. In one example, for chroma components, SubC is equal to 1 when the color format is 4:2:2.
[0676] viii. In one example, for chroma components, SubC is equal to 1 when the color format is 4:4:4.
[0677] g. In one example, whether and / or how to perform clipping can depend on the size of the current picture (or a consistency window therein) and the size of the reference picture (or a consistency window therein).
[0678] i. In one example, clipping is performed only when the size of the current picture (or a consistency window therein) and the size of the reference picture (or a consistency window therein) are different.
[0679] In the following discussion, a first syntax element is asserted to "correspond to" a second syntax element if the first syntax element and the second syntax element have equivalent functionality, but can be signaled at different video units (e.g., VPS / SPS / PPS / slice header / picture header, etc.)
[0680] 30. It is proposed that a syntax element can be signaled in a first video unit (e.g., picture header or PPS) without a corresponding syntax element signaled in a second video unit at a higher level (such as SPS) or a lower level (such as slice header).
[0681] a. Alternatively, a first syntax element can be signaled in a first video unit (e.g., picture header or PPS) and a corresponding second syntax element can be signaled in a second video unit at a lower level (e.g., slice header).
[0682] i. Alternatively, an indicator can be signaled in the second video unit to inform whether the second syntax element is signaled thereafter.
[0683] ii. In one example, if the second syntax element is signaled, a slice associated with the second video unit (such as a slice header) can follow the indication of the second syntax element rather than the first syntax element.
[0684] iii. An indicator associated with the first syntax element can be signaled in the first video unit to inform whether the second syntax element is signaled in any slice (or other video unit) associated with the first video unit.
[0685] b. Alternatively, the first syntax element can be signaled in the first video unit at a higher level (such as VPS / SPS / PPS), and the corresponding second syntax element can be signaled in the second video unit (e.g., picture header).
[0686] i. Alternatively, an indicator can be signaled to inform whether the second syntax element is signaled thereafter.
[0687] ii. In one example, if the second syntax element is signaled, a picture associated with the second video unit (which can be partitioned into slices) can follow the indication of the second syntax element rather than the first syntax element.
[0688] c. The first syntax element in the picture header can have the same function as the second syntax element in the slice header specified in Section 2.6 (such as but not limited to slice_temporal_mvp_enabled_flag, cabac_init_flag, six_minus_max_num_merge_cand, five_minus_max_num_subblock_merge_cand, slice_fpel_mmvd_enabled_flag, slice_disable_bdof_dmvr_flag, max_num_merge_cand_minus_max_num_triangle_cand, slice_fpel_mmvd_enabled_flag, slice_six_minus_max_num_ibc_merge_cand, slice_joint_cbcr_sign_flag, slice_qp_delta...), but is controlling all slices of the picture.
[0689] d. The first syntax elements in the SPS specified in Section 2.6 can have the same function as the second syntax elements in the picture header (such as but not limited to sps_bdof_dmvr_slice_present_flag, sps_mmvd_enabled_flag, sps_isp_enabled_flag, sps_mrl_enabled_flag, sps_mip_enabled_flag, sps_cclm_enabled_flag, sps_mts_enabled_flag,...), but only control the associated picture (which can be partitioned into slices).
[0690] e. The first syntax elements in the PPS specified in Section 2.7 can have the same function as the second syntax elements in the picture header (such as but not limited to entropy_coding_sync_enabled_flag, entry_point_offsets_present_flag, cabac_init_present_flag, rpl1_idx_present_flag,...), but only control the associated picture (which can be partitioned into slices).
[0691] 31. The syntax elements signaled in the picture header are decoupled from the other syntax elements signaled or derived from the SPS / VPS / DPS.
[0692] 32. The indication to enable / disable DMVR and BDOF can be signaled separately in the picture header instead of being controlled by the same flag (e.g., pic_disable_bdof_dmvr_flag).
[0693] 33. The indication to enable / disable PROF / cross-component ALF / inter prediction with geometric partitioning (GEO) can be signaled in the picture header.
[0694] a. Alternatively, the indication to enable / disable PROF in the picture header can be conditionally signaled according to the PROF enabling flag in the SPS.
[0695] b. Alternatively, the indication to enable / disable cross-component ALF (CCALF) in the picture header can be conditionally signaled according to the enabling / disabling CCALF enabling flag in the SPS.
[0696] c. Alternatively, the indication to enable / disable GEO in the picture header can be conditionally signaled according to the GEO enabling flag in the SPS.
[0697] d. Alternatively, in addition, the indication of enabling / disabling PROF / cross-component ALF / inter prediction with geometric partitioning (GEO) in slice header can be conditionally signaled according to those syntax elements signaled in picture header but not in SPS.
[0698] 34. The indication of prediction type of slices / bricks / tiles (or other video units smaller than a picture) in the same picture can be signaled in picture header.
[0699] a. In one example, the indication of whether all slices / bricks / tiles (or other video units smaller than a picture) are intra coded (e.g., all I slices) can be signaled in picture header.
[0700] i. Alternatively, in addition, the slice type can not be signaled in slice header if the indication indicates that all slices in the picture are I slices.
[0701] b. Alternatively, the indication of whether at least one of the slices / bricks / tiles (or other video units smaller than a picture) is not intra coded (e.g., at least one non-I slice) can be signaled in picture header.
[0702] c. Alternatively, the indication of whether all slices / bricks / tiles (or other video units smaller than a picture) have the same prediction type (e.g., I / P / B slices) can be signaled in picture header.
[0703] i. Alternatively, in addition, the slice type can not be signaled in slice header.
[0704] ii. Alternatively, in addition, the indication of tools allowed for a certain prediction type (e.g., for B slices, only DMVR / BDOF / TPM / GEO are allowed; for I slices, only dual-hybrid tree is allowed) can be conditionally signaled according to the indication of prediction type.
[0705] d. Alternatively, in addition, the signaling of the indication of enabling / disabling tools can depend on the indication of prediction type mentioned in the above sub-bullets.
[0706] i. Alternatively, in addition, the indication of enabling / disabling tools can be derived depending on the indication of prediction type mentioned in the above sub-bullets.
[0707] 35. In this disclosure (bullet 1 - bullet 29), the term “consistency window” can be replaced by other terms such as “scaling window”. The scaling window can be signaled differently from the consistency window and be used to derive the scaling ratio and / or the top-left offset for deriving the reference sample positions of RPR.
[0708] a. In one example, the scaling window can be constrained by the conformance window. For example, in a conformance bitstream, the scaling window must be contained by the conformance window.
[0709] 36. Whether and / or how to signal the allowed maximum block size of a transform skip coded block can depend on the maximum block size of a transform coded block.
[0710] a. Alternatively, the maximum block size of a transform skip coded block cannot be larger than the maximum block size of a transform coded block in a conformance bitstream.
[0711] 37. Whether and how to signal an indication (such as sps_joint_cbcr_enabled_flag) to enable Joint Cb-Cr Residual (JCCR) coding can depend on the color format (such as 4:0:0, 4:2:0, etc.)
[0712] a. For example, if the color format is 4:0:0, the indication to enable Joint Cb-Cr Residual (JCCR) can not be signaled. An example syntax design is as follows:
[0713] if ( ChromaArrayType!= 0 ) sps_joint_cbcr_enabled_flag u(1)
[0714] Downsampling filter type for chroma blending mask generation in TPM / GEO
[0715] 38. The type of down-sampling filter used for the derivation of the mixing weights for chroma samples can be signaled at the video unit level (such as SPS / VPS / PPS / Picture header / Subpicture / Slice / Slice header / Tile / CTU / VPDU level).
[0716] a. In one example, a high level flag can be signaled to switch between different chroma format types of content.
[0717] i. In one example, a high level flag can be signaled to switch between chroma format type 0 and chroma format type 2.
[0718] ii. In one example, a flag can be signaled to specify whether the top-left down-sampled luma weight in TPM / GEO prediction modes is collocated with the top-left luma weight (i.e., chroma sample position type 0).
[0719] iii. In one example, a flag can be signaled to specify whether the top-left down-sampled luma sample in TPM / GEO prediction modes is co-located horizontally with the top-left luma sample but vertically shifted by 0.5 units relative to the top-left luma sample (i.e., chroma sample position type 2).
[0720] b. In one example, for 4:2:0 chroma format and / or 4:2:2 chroma format, the type of down-sampling filter can be signaled.
[0721] c. In one example, a flag for specifying the type of chroma down-sampling filter for TPM / GEO prediction can be signaled.
[0722] i. In one example, a flag for specifying whether down-sampling filter A or down-sampling filter B is used for chroma weight derivation in TPM / GEO prediction mode can be signaled.
[0723] 39. The type of down-sampling filter for mixed weight derivation of chroma samples can be derived at video unit level (e.g., SPS / VPS / PPS / picture header / sub-picture / slice / slice header / tile / graph tile / CTU / VPDU level).
[0724] a. In one example, a lookup table can be defined to specify the correspondence between the type of chroma sub-sampling filter and the type of chroma format content.
[0725] 40. In the case of different chroma sample position types, the specified down-sampling filter can be used for TPM / GEO prediction mode.
[0726] a. In one example, in the case of certain chroma sample position type (e.g., chroma sample position type 0), the chroma weight of TPM / GEO can be sub-sampled from the collocated top-left luma weight.
[0727] b. In one example, in the case of certain chroma sample position type (e.g., chroma sample position type 0 or 2), a specified X-tap filter (X is a constant, such as X=6 or 5) can be used for chroma weight sub-sampling in TPM / GEO prediction mode.
[0728] 41. In a video unit (e.g., SPS, PPS, picture header, slice header, etc.), a first syntax element (such as a flag) can be signaled to indicate whether MTS is disabled for all blocks (slices / pictures).
[0729] a. A second syntax element indicating how to apply MTS (such as enabling MTS / disabling MTS / implicit MTS / explicit MTS) on intra coded blocks (slices / pictures) is conditionally signaled on the first syntax element. For example, the second syntax element is only signaled when the first syntax element indicates that MTS is not disabled for all blocks (slices / pictures).
[0730] b. A third syntax element indicating how to apply MTS (such as enabling MTS / disabling MTS / implicit MTS / explicit MTS) on inter-coded blocks (slices / pictures) is conditionally signaled on the first syntax element. For example, the third syntax element is signaled only when the first syntax element indicates that MTS is not disabled for all blocks (slices / pictures).
[0731] c. An example syntax design is as follows
[0732]
[0733] d. A third syntax element on whether to apply Sub-Block Transform (SBT) can be conditionally signaled. An example syntax design is as follows
[0734] e. An example syntax design is as follows
[0735]
[0736] Determination of usage of coding tool X
[0737] 42. The determination of whether and / or how to enable coding tool X can depend on the width and / or height of one or more reference pictures and / or a considering picture in the current picture.
[0738] a. The width and / or height of the considering picture in the one or more reference pictures and / or the current picture can be modified for the determination.
[0739] b. The considering picture can be defined by a consistent window or a scaling window, as defined in JVET-P0590.
[0740] i. The considering picture can be the entire picture.
[0741] c. In one example, whether and / or how to enable coding tool X can depend on the width of the picture minus one or more offsets in the horizontal direction and / or the height of the picture minus an offset in the vertical direction.
[0742] i. In one example, the horizontal offset can be defined as scaling_win_left_offset, where scaling_win_left_offset can be as defined in JVET-P0590.
[0743] ii. In one example, the vertical offset can be defined as scaling_win_top_offset, where scaling_win_top_offset can be as defined in JVET-P0590.
[0744] iii. In one example, the horizontal offset can be defined as (scaling_win_right_offset + scaling_win_left_offset), where scaling_win_right_offset and scaling_win_left_offset can be as defined in JVET-P0590.
[0745] iv. In one example, the vertical offset can be defined as (scaling_win_bottom_offset + scaling_win_top_offset), where scaling_win_bottom_offset and scaling_win_top_offset can be as defined in JVET-P0590.
[0746] v. In one example, the horizontal offset can be defined as SubWidthC * (scaling_win_right_offset + scaling_win_left_offset), where SubWidthC, scaling_win_right_offset and scaling_win_left_offset can be as defined in JVET-P0590.
[0747] vi. In one example, the vertical offset can be defined as SubHeightC * (scaling_win_bottom_offset + scaling_win_top_offset), where SubHeightC, scaling_win_bottom_offset and scaling_win_top_offset can be as defined in JVET-P0590.
[0748] d. In one example, if at least one of the two considered reference pictures has a different resolution (width or height) than the current picture, the coding tool X is disabled.
[0749] i. Alternatively, if the dimension (width or height) of at least one of the two output reference pictures is greater than the dimension of the current picture, the coding tool X is disabled.
[0750] e. In one example, the coding tool X is disabled for the reference picture list L if the considered reference picture of the reference picture list L has a different resolution than the current picture.
[0751] i. Alternatively, the coding tool X is disabled for the reference picture list L if the dimension (width or height) of the considered reference picture of the reference picture list L is greater than the dimension of the current picture.
[0752] f. In one example, the coding tool can be disabled if the two considered reference pictures of the two reference picture lists have different resolutions.
[0753] i. Alternatively, the indication of the coding tool can be conditionally signaled according to the resolution.
[0754] ii. Alternatively, the signaling of the indication of the coding tool can be skipped.
[0755] g. In one example, the coding tool can be disabled if the two considered reference pictures of the two merge candidates are used to derive the first pairwise merge candidate of at least one reference picture list, i.e. the first pairwise merge candidate is marked as unavailable.
[0756] i. Alternatively, the coding tool can be disabled if the two considered reference pictures of the two merge candidates are used to derive the first pairwise merge candidate of the two reference picture lists, i.e. the first pairwise merge candidate is marked as unavailable.
[0757] h. In one example, the picture dimension can be considered to modify the decoding process of the coding tool.
[0758] i. In one example, the derivation of the MVD of another reference picture list (e.g. list 1) in SMVD can be based on the resolution difference (e.g. scaling factor) of at least one of the two target SMVD reference pictures.
[0759] ii. In one example, the derivation of the pairwise merge candidate can be based on the resolution difference (e.g. scaling factor) of at least one of the two reference pictures associated with the two reference pictures, e.g. a linear weighted average can be applied instead of equal weights.
[0760] i. In one example, X can be:
[0761] i. DMVR / BDOF / PROF / SMVD / MMVD / other coding tools that refine the motion / prediction at the decoder side
[0762] ii. TMVP / other coding tools that rely on temporal motion information
[0763] iii. MTS or other transform coding tool
[0764] iv. CC-ALF
[0765] v. TPM
[0766] vi. GEO
[0767] vii. Switchable interpolation filter (e.g., alternative interpolation filter for half-pel motion compensation)
[0768] viii. Hybrid process of splitting a block into multiple partitions in TPM / GEO / other coding tools
[0769] ix. Coding tool that depends on information stored in a picture different from the current picture
[0770] x. Paired merge candidate (not generated when certain conditions related to resolution are not met)
[0771] xi. Bi-prediction with CU-level Weight (BCW)
[0772] xii. Weighted prediction
[0773] xiii. Affine prediction
[0774] xiv. Adaptive Motion Vector Resolution (AMVR)
[0775] 43. Whether and / or how to signal the use of a coding tool can depend on the width and / or height of a considered picture in one or more reference pictures and / or the current picture.
[0776] j. The width and / or height of a considered picture in one or more reference pictures and / or the current picture can be modified for the determination.
[0777] k. The considered picture can be defined by a consistent window or a scaled window as defined in JVET-P0590.
[0778] i. The considered picture can be the entire picture.
[0779] l. In one example, X can be Adaptive Motion Vector Resolution (AMVR).
[0780] m. In one example, X can be the merge with MV differences (MMVD) method.
[0781] i. In one example, the construction of the symmetric motion vector difference reference index can depend on the indication of the picture resolution / RPR case of different reference pictures.
[0782] n. In one example, X can be the symmetric MVD (SMVD) method.
[0783] o. In one example, X can be QT / BT / TT or other partition type.
[0784] p. In one example, X can be bi-prediction with CU-level weights (BCW).
[0785] q. In one example, X can be weighted prediction.
[0786] r. In one example, X can be affine prediction.
[0787] s. In one example, the indication of whether to signal the use of the half-pel motion vector precision / switchable interpolation filter can depend on the resolution information / whether RPR is enabled for the current block.
[0788] t. In one example, the signaling of amvr_precision_idx can depend on the resolution information / whether RPR is enabled for the current block.
[0789] u. In one example, the signaling of sym_mvd_flag / mmvd_merge_flag can depend on the resolution information / whether RPR is enabled for the current block.
[0790] v. A conformant bitstream should satisfy that 1 / 2 pel MV and / or MVD precision (e.g., alternative interpolation filter / switchable interpolation filter) is not allowed when the width and / or height of the considered picture(s) of one or more reference pictures is different from the width and / or height of the current output picture.
[0791] 44. It is proposed that for blocks in RPR, AMVR with 1 / 2 pel MV and / or MVD precision (or alternative interpolation filter / switchable interpolation filter) can still be enabled.
[0792] w. Alternatively, in addition, different interpolation filters can be applied to blocks with 1 / 2 pel or other precision.
[0793] 45. The same / different resolution condition check in the above bullet can be replaced by adding a flag of the reference picture and checking the flag associated with the reference picture.
[0794] x.In one example, during the reference picture list construction process, a process can be invoked that sets the flag to true or false (i.e., to indicate whether the reference picture is an RPR case or a non-RPR case).
[0795] i.For example, the following cases can apply:
[0796] fRefWidth is set equal to PicOutputWidthL of the reference picture RefPicList[i][j] in luma sample units, where PicOutputWidthL denotes the width of the considered picture of the reference picture.
[0797] fRefWidth is set equal to PicOutputHeightL of the reference picture RefPicList[i][j] in luma sample units, where PicOutputHeightL denotes the height of the considered picture in the reference picture.
[0798] RefPicScale[i][j][0] = ((fRefWidth « 14) + (PicOutputWidthL » 1)) / PicOutputWidthL, where
[0799] PicOutputWidthL denotes the width of the considered picture of the current picture.
[0800] RefPicScale[i][j][1] = ((fRefHeight « 14) + (PicOutputHeightL » 1)) / PicOutputHeightL, where PicOutputWidthL denotes the height of the considered picture in the current picture.
[0801] RefPicIsScaled[i][j] = (RefPicScale[i][j][0]!= (1 « 14)) ||
[0802] (RefPicScale[i][j][1]!= (1 « 14))
[0803] where RefPicList[i][j] denotes the j-th reference picture in the reference picture list i.
[0804] y. In one example, when RefPicIsScaled[0][refIdxL0] is not equal to 0 or RefPicIsScaled[1][refIdxL1] is not equal to 0, the coding tool X (e.g., DMVR / BDOF / SMVD / MMVD / SMVD / PROF / those mentioned in the above bullets) can be disabled.
[0805] z. In one example, when RefPicIsScaled[0][refIdxL0] and RefPicIsScaled[1][refIdxL1] are not equal to 0, the coding tool X (e.g., DMVR / BDOF / SMVD / MMVD / SMVD / PROF / those mentioned in the above bullets) can be disabled.
[0806] aa. In one example, when RefPicIsScaled[0][refIdxL0] is not equal to 0, the coding tool X (e.g., PROF or those tools mentioned in the above bullets) can be disabled for reference picture list 0.
[0807] bb. In one example, when RefPicIsScaled[1][refIdxL1] is not equal to 0, the coding tool X (e.g., PROF or those tools mentioned in the above bullets) can be disabled for reference picture list 1.
[0808] 46. The SAD and / or the threshold used by BDOF / DMVR can depend on the bit depth.
[0809] a. In one example, the computed SAD value can be first shifted as a function of the bit depth before being used for comparison with the threshold.
[0810] b. In one example, the computed SAD value can be directly compared with a modified threshold, which can depend on a function of the bit depth.
[0811] 42. If the slice_type value of all slices of a picture is equal to I (I slice), the P / B slice related syntax elements can not be signaled in the picture header.
[0812] a. In one example, the syntax element(s) can be added to the picture header to indicate whether the slice_type of all slices included in the picture is equal to I (I slice).
[0813] i. In one example, a first syntax element can be signaled in a picture header. Whether and / or how a second syntax element is signaled / interpreted can depend on the first syntax element, the second syntax element informing slice type information in a slice header of a slice associated with the picture header.
[0814] 1) In one example, the second syntax element can not be signaled and is inferred to be a slice type depending on the first syntax element.
[0815] 2) In one example, the second syntax element can be signaled, but a conformance requirement is that the second syntax element must be one of several given values depending on the first syntax element.
[0816] 3) Alternatively, the first syntax element can be signaled in an AU delimiter RBSP associated with the slice.
[0817] ii. In one example, a new syntax element (e.g., pic_all_X_slices_flag) can be signaled in a picture header to indicate whether only X slices are allowed for the picture, or whether all slices in the picture are X slices. X can be I, or P or B, for example.
[0818] 1) In one example, if all slices are indicated to be I slices in the associated picture header, slice type information is not signaled in the slice header and is inferred to be I slices.
[0819] iii. In one example, a new syntax element (e.g., ph_pic_type) can be signaled in a picture header to indicate a picture type of the picture.
[0820] 1) For example, if ph_pic_type is equal to an I picture (e.g., equal to 0), slice_type of slices in the picture can be allowed to be equal to I only.
[0821] 2) For another example, if ph_pic_type is equal to a non-I picture (such as 1 or 2), slice_type of slices in the picture can be allowed to be equal to I and / or P and / or B.
[0822] b. In one example, a syntax element pic_type in an AU delimiter RBSP can be used to indicate whether all slices of a specified picture are equal to I.
[0823] c. In one example, if all slices in a picture are I slices, a syntax element slice_type in a slice header can not be signaled and is inferred to be I slices (such as 2) for all slices in the picture.
[0824] i. In one example, if a syntax element such as but not limited to pic_all_I_slices_flag, ph_pic_type, pic_type indicates that all slices included in the specified picture are equal to I, a bitstream constraint can be added to specify that slice_type of each slice in the specified picture should be equal to I slice.
[0825] ii. Alternatively, if a syntax element such as but not limited to pic_all_I_slices_flag, ph_pic_type, pic_type indicates that all slices included in the specified picture are equal to I, a bitstream constraint can be added to specify that P or B slices are not allowed in the specified picture.
[0826] d. If the slice / picture is a W slice / W picture, one or more syntax elements (denoted as set X of syntax elements as specified below) allowed for non-W slices in the slice header / picture header can not be signaled. For example, W can be I and non-W can be B or P. In another example, W can be B and non-W can be I or P.
[0827] i. In one example, if all slices in the picture are W slices, set X of syntax elements in the picture header allowed for non-W slices can not be signaled.
[0828] ii. Alternatively, set X of syntax elements (denoted as specified below) in the picture header can be conditionally signaled depending on whether all slices in the picture are W slices.
[0829] iii. Set X of syntax elements can be one or more of the following.
[0830] 1) In one example, X can be reference picture related syntax elements in the picture header such as but not limited to pic_rpl_present_flag, pic_rpl_sps_flag, pic_rpl_idx, pic_poc_lsb_lt, pic_delta_poc_msb_present_flag, pic_delta_poc_msb_cycle_lt...
[0831] If the slice / picture is signaled or inferred to be non-inter slice / picture, X can not be signaled and inferred to be unused.
[0832] 2) In one example, X can be an inter slice related syntax element in picture header, such as but not limited to pic_log2_diff_min_qt_min_cb_inter_slice, pic_max_mtt_hierarchy_depth_inter_slice, pic_log2_diff_max_bt_min_qt_inter_slice, pic_log2_diff_max_tt_min_qt_inter_slice,...
[0833] If the slice / picture is signaled or inferred to be non-inter slice / picture, X can not be signaled and inferred to be not used.
[0834] 3) In one example, X can be an inter prediction related syntax element in picture header, such as but not limited to pic_temporal_mvp_enabled_flag, mvd_l1_zero_flag, pic_six_minus_max_num_merge_cand, pic_five_minus_max_num_subblock_merge_cand, pic_fpel_mmvd_enabled_flag, pic_disable_bdof_flag, pic_disable_dmvr_flag, pic_disable_prof_flag, pic_max_num_merge_cand_minus_max_num_triangle_cand,...
[0835] If the slice / picture is signaled or inferred to be non-inter slice / picture, X can not be signaled and inferred to be not used.
[0836] 4) In one example, X can be a bi-prediction related syntax element in picture header, such as but not limited to pic_disable_bdof_flag, pic_disable_dmvr_flag, mvd_l1_zero_flag,...
[0837] If the slice / picture is signaled or inferred to be non-B slice / picture, X can not be signaled and inferred to be not used.
[0838] Limitation on the dimension of the relevant tile / slice
[0839] 43. The maximum tile width can be specified in the specification.
[0840] a. For example, the maximum slice width can be defined as the maximum luma slice width in CTB.
[0841] b. In one example, in a video unit (e.g., SPS, PPS, picture header, slice header, etc.), new syntax element(s) can be signaled to indicate the maximum slice width allowed in the current sequence / picture / slice / subpicture.
[0842] c. In one example, the maximum luma slice width or the maximum luma slice width in CTB can be signaled.
[0843] d. In one example, the maximum luma slice width can be fixed as N (such as N = 1920 or 4096, etc.).
[0844] e. In one example, different maximum slice width can be specified in different profiles / tiers / layers.
[0845] 44. The maximum slice / subpicture / slice dimension (e.g., width and / or size, and / or length, and / or height) can be specified in the specification.
[0846] a. For example, the maximum slice / subpicture / slice dimension can be defined as the maximum luma dimension in CTB.
[0847] b. For example, the size of slice / subpicture / slice can be defined as the number of CTBs in the slice.
[0848] c. In one example, in a video unit (e.g., SPS, PPS, picture header, slice header, etc.), new syntax element(s) can be signaled to indicate the maximum luma slice / subpicture / slice dimension allowed in the current sequence / picture / slice / subpicture.
[0849] d. In one example, for rectangular slice / subpicture, the maximum luma slice / subpicture width / height or the maximum luma slice / subpicture width / height in CTB can be signaled.
[0850] e. In one example, for raster scan slice, the maximum luma slice size (such as width*height) or the maximum luma slice size in CTB can be signaled.
[0851] f. In one example, for rectangular slice / subpicture and raster scan slice, the maximum luma slice / subpicture size (such as width*height) or the maximum luma slice / subpicture size in CTB can be signaled.
[0852] g. In one example, for raster scan slice, the maximum luma slice length or the maximum luma slice length in CTB can be signaled.
[0853] h. In one example, the maximum luma slice / subpicture width / height can be fixed to N (such as N = 1920 or 4096, etc.)
[0854] i. In one example, the maximum luma slice / subpicture size (such as width*height) can be fixed to N (such as N = 2073600 or 83388608, etc.)
[0855] j. In one example, different maximum slice / subpicture dimensions can be specified in different profiles / tiers / levels.
[0856] k. In one example, the maximum number of slices / subpictures / tiles to be partitioned in a picture / subpicture can be specified in the specification.
[0857] i. The maximum number can be signaled.
[0858] ii. The maximum number can be different in different profiles / tiers / levels.
[0859] 45. Assume the width and height of the current picture are denoted as PW and PH, respectively; the width and height of the scaling window of the current picture are denoted as SW and SH, respectively; the width and height of the scaling window of the reference picture are denoted as SW’ and SH’, respectively. The maximum allowed width and height of the picture are denoted as Wmax and Hmax, respectively. For convenience, let and At least one of the following constraints should be followed by a conformant bitstream. The following constraints should not be interpreted in a narrow way. For example, the constraint (a / b) >= (c / d) (where a, b, c, d are integers greater than 0) can also be interpreted as (a / b) - (c / d) >= 0, or a*d >= c*b, or a*d - c*d >= 0.
[0860] a. a x Wmax x SW - (b x SW’ + c x SW) x PW + offw >= 0, where a, b, c and offw are integers. For example, a = 135, b = 128, c = 7, and offw = 0.
[0861] b. d x Hmax x SH - (e x SH’ + f x SH) x PH + offh >= 0, where d, e, f are integers. For example, d = 135, e = 128, f = 7, and offh = 0.
[0862] c. a x Wmax x SW - b x SW’ x PW + offw >= 0, where a, b are integers. For example, a = 1, b = 1, and offw = 0.
[0863] d.d x Hmax x SH - e x SH' x PH + offh > 0, where d, e are integers. For example, d = 1, e = 1, and offh = 0.
[0864] e. where Lw, Bw, and offw are integers. For example, Lw = 7, Bw = 128, and offw = 0.
[0865] f. where Lh, Bh, and offh are integers. For example, Lh = 7, Bh = 128, and offh = 0.
[0866] g.rw≤ a * Rw + offw, where a and offw are integers. For example, a = 1, and offw = 0.
[0867] h.rh≤ b * Rh + offh, where b and offh are integers. For example, b = 1, and offh = 0.
[0868] i.rw≤ a * Rw + offw, and offw are integers. For example, a = 1, offw = 0.
[0869] j.rh≤ b * Qh + offh, where b and offh are integers. For example, b = 1, and offh = 0.
[0870] 46. QP related information such as delta QP can be signaled in picture header instead of PPS.
[0871] a. In one example, QP related information such as delta QP is specified for a specific coding tool such as adaptive color transform (ACT).
[0872] 5. Additional Embodiments
[0873] The working draft specified in JVET-O2001-vE can be changed in the following embodiments. In the following table, the text changes in the VVC draft are underlined in bold italics and deletions are shown within double bold brackets, e.g., [[a]] indicates that "a" is deleted.
[0874] 5.1. Embodiments of constraints on the conformance window
[0875] conf win left offset, conf win right offset, conf win top offset, and conf win bottom offset specify, in terms of a rectangular region specified in picture coordinates for output, samples of a picture in the CVS output from the decoding process. When conformance window flag is equal to 0, the values of conf win left offset, conf win right offset, conf win top offset, and conf win bottom offset are inferred to be equal to 0.
[0876] The conformance clipping window contains luma samples with horizontal picture coordinates (inclusive) from SubWidthC * conf win left offset to pic width in luma samples - (SubWidthC * conf win right offset + 1), and vertical picture coordinates (inclusive) from SubHeightC * conf win top offset to pic height in luma samples - (SubHeightC * conf win bottom offset + 1).
[0877] The value of SubWidthC * (conf win left offset + conf win right offset) shall be less than pic width in luma samples, and the value of SubHeightC * (conf win top offset + conf win bottom offset) shall be less than pic height in luma samples.
[0878] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
[0879] PicOutputWidthL = pic width in luma samples - SubWidthC * (conf win right offset + conf win left offset) (7-43)
[0880] PicOutputHeightL = pic_height_in_pic_size_units - SubHeightC * ( conf_win_bottom_offset + conf_win_top_offset ) (7-44)
[0881] When ChromaArrayType is not equal to 0, the corresponding specified sample of the two chroma arrays is the sample with picture co-ordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture co-ordinates of the specified luma sample.
[0882]
[0883] 5.2. Embodiment 1 of reference sample position derivation
[0884] 8.5.6.3.1 Overview
[0885] …
[0886] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma sample units.
[0887] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma sample units.
[0888]
[0889] The motion vector mvLX is set to ( refMvLX - mvOffset ).
[0890] - If cldx is equal to 0, the following applies:
[0891] - The scaling factors and their fixed-point representations are defined as follows
[0892] hori_scale_fp = ( ( fRefWidth « 14 ) + ( PicOutputWidthL » 1 ) ) / PicOutputWidthL (8-753)
[0893] vert_scale_fp = ( ( fRefHeight « 14 ) + ( PicOutputHeightL » 1 ) ) / PicOutputHeightL (8-754)
[0894] – Let (xIntL, yIntL) be the luma position given in full-sample units and (xFracL, yFracL) be the offset given in 1 / 16-sample units. These variables are only used in this subclause to specify fractional-sample positions within the reference sample array refPicLX.
[0895] – The top-left coordinates (xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding are set equal to (xSb + (mvLX[0] » 4), ySb + (mvLX[1] » 4)).
[0896] – For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L = 0..sbHeight - 1 + brdExtSize) within the array of predicted luma samples predSamplesLX, the corresponding predicted luma sample value predSamplesLX[x L ][y L ] is derived as follows:
[0897] – Let (refxSb L , refySb L ) and (refx L , refy L ) be the luma positions pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16-sample units. The variables refxSb L , refx L , refySb L and refy L are derived as follows:
[0898]
[0899] refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128) » 8) + x L *((hori_scale_fp + 8) » 4)) + 32) » 6 (8-756)
[0900]
[0901] refy L= ((Sign( refySb ) * ( ( Abs( refySb ) + 128 ) » 8 ) + yL * ( ( vert_scale_fp + 8 ) » 4 )) + 32 ) » 6 (8-758)
[0902] –
[0903] –
[0904] – Variable xInt L , yInt L , xFrac L , and yFrac L are derived as follows:
[0905] xInt L = refx L >> 4 (8-759)
[0906] yInt L = refy L >> 4 (8-760)
[0907] xFrac L = refx L & 15 (8-761)
[0908] yFrac L = refy L & 15 (8-762)
[0909] – If bdofFlag is equal to TRUE (or sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[ xSb ][ ySb ] is equal to TRUE), and one or more of the following conditions are true, the predicted luma sample values predSamplesLX[ xL ][ yL ] are derived by invoking the luma integer sample fetching process specified in clause 8.5.6.3.3 with ( xIntL + ( xFracL » 3 ) - 1 ), yIntL + ( yFracL » 3 ) - 1 ), and refPicLX as inputs.
[0910] – x L is equal to 0.
[0911] – x L is equal to sbWidth + 1.
[0912] – y L is equal to 0.
[0913] – y Lis equal to sbHeight + 1.
[0914] - Otherwise, the predicted luma sample values predSamplesLX[ xL ][ yL ] are derived by calling the luma sample 8-tap interpolation filter process as specified in clause 8.5.6.3.2 with ( xIntL - ( brdExtSize > 0? 1 : 0 ), yIntL - ( brdExtSize > 0? 1 : 0 ) ), ( xFracL, yFracL ), ( xSbIntL, ySbIntL ), refPicLX, hpelIfIdx, sbWidth, sbHeight, and ( xSb, ySb ) as inputs. L ,ySbInt L .
[0915] - Otherwise ( cIdx is not equal to 0 ), the following applies:
[0916] - Let ( xIntC, yIntC ) be the chroma position in whole sample units and ( xFracC, yFracC ) be the offset in 1 / 32 sample units. These variables are only used in this subclause to specify the general fractional sample positions within the reference sample array refPicLX.
[0917] - The top-left coordinates ( xSbIntC, ySbIntC ) of the boundary block used for reference sample padding are set equal to ( ( xSb / SubWidthC ) + ( mvLX[ 0 ] » 5 ), ( ySb / SubHeightC ) + ( mvLX[ 1 ] » 5 ) ).
[0918] - For each chroma sample position ( xC = 0..sbWidth - 1, yC = 0..sbHeight - 1 ) within the predicted chroma sample array predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[ xC ][ yC ] is derived as follows:
[0919] - Let ( refxSb C , refySb C ) and ( refx C , refy C ) be the chroma positions pointed to by the motion vector ( mvLX[ 0 ], mvLX[ 1 ] ) in 1 / 32 sample units. The variables refxSb C , refySb C , refx C , and refy C are derived as follows:
[0920]
[0921] refx C = (( Sign( refxSb C ) * ( ( Abs( refxSb C ) + 256 ) >> 9 ) + xC * ( ( hori_scale_fp + 8 ) >> 4 ) ) + 16 ) >> 5 (8-764)
[0922]
[0923] refy C = (( Sign( refySb C ) * ( ( Abs( refySb C ) + 256 ) >> 9 ) + yC * ( ( vert_scale_fp + 8 ) >> 4 ) ) + 16 ) >> 5 (8-766)
[0924] –
[0925] –
[0926] – The variables xInt C , yInt C , xFrac C , and yFrac C are derived as follows:
[0927] xInt C = refx C >> 5 (8-767)
[0928] yInt C = refy C >> 5 (8-768)
[0929] xFrac C = refy C & 31 (8-769)
[0930] yFrac C = refy C & 31 (8-770)
[0931] 5.3. Reference sample position derivation, embodiment 2
[0932] 8.5.6.3.1 Overview
[0933] …
[0934] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in units of luma samples.
[0935] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma sample units.
[0936]
[0937] The motion vector mvLX is set to (refMvLX - mvOffset).
[0938] - If cldx is equal to 0, the following applies:
[0939] - The scaling factors and their fixed-point representations are defined as follows
[0940] hori_scale_fp = (( fRefWidth « 14 ) + ( PicOutputWidthL » 1 )) / PicOutputWidthL (8-753)
[0941] vert_scale_fp = (( fRefHeight « 14 ) + ( PicOutputHeightL » 1 )) / PicOutputHeightL (8-754)
[0942] - Let (xIntL, yIntL) be the luma position given in full-sample units and (xFracL, yFracL) be the offset given in 1 / 16-sample units. These variables are only used in this subclause to specify fractional-sample positions within the reference sample array refPicLX.
[0943] - The top-left coordinates (xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding are set equal to (xSb + (mvLX[0] » 4), ySb + (mvLX[1] » 4)).
[0944] - For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L = 0..sbHeight - 1 + brdExtSize) within the array of predicted luma samples predSamplesLX, the corresponding predicted luma sample value predSamplesLX[xL][yL] is derived as follows:
[0945] - Let (refxSb L ,refySb L ) and (refx L ,refy L) is the luma position pointed to by the motion vector given in 1 / 16 sample units (refMvLX[0], refMvLX[1]). The variable refxSb L , refx L , refySb L and refy L are derived as follows:
[0946]
[0947]
[0948]
[0949]
[0950]
[0951] – the variables xInt L , yInt L , xFrac L and yFrac L are derived as follows:
[0952] xInt L = refx L >> 4 (8-759)
[0953] yInt L = refy L >> 4 (8-760)
[0954] xFrac L = refx L & 15 (8-761)
[0955] yFrac L = refy L & 15 (8-762)
[0956] – If bdofFlag is equal to TRUE (or sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[ xSb ][ ySb ] is equal to TRUE ), and one or more of the following conditions are true, the luma integer sample fetch process specified in clause 8.5.6.3.3 is invoked with ( xInt L + ( xFrac L >> 3 ) - 1 ), yInt L + ( yFrac L>> 3) - 1) and refPicLX as inputs to derive the predicted luma sample value L ][y L ].
[0957] - x L equals 0.
[0958] - x L equals sbWidth + 1.
[0959] - y L equals 0.
[0960] - y L equals sbHeight + 1.
[0961] - Otherwise, the predicted luma sample value predSamplesLX[xL][yL] is derived by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2 with (xIntL - (brdExtSize > 0? 1 : 0), yIntL - (brdExtSize > 0? 1 : 0)), (xFracL, yFracL), (xSbInt L , ySbInt L ), refPicLX, hpelIfIdx, sbWidth, sbHeight and (xSb, ySb) as inputs.
[0962] - Otherwise (cldx is not equal to 0), the following applies:
[0963] - Let (xIntC, yIntC) be the chroma position in whole sample units and (xFracC, yFracC) be the offset in 1 / 32 sample units. These variables are only used in this subclause to specify the general fractional sample position within the reference sample array refPicLX.
[0964] - The top-left coordinates (xSbIntC, ySbIntC) of the boundary block used for reference sample padding are set equal to ((xSb / SubWidthC) + (mvLX[0] » 5), (ySb / SubHeightC) + (mvLX[1] » 5)).
[0965] - For each chroma sample position (xC = 0..sbWidth - 1, yC = 0..sbHeight - 1) within the predicted chroma sample array predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows:
[0966] - refxSbC and refySbC are the chroma positions in 1 / 32 sample units pointed at by the motion vector (mvLX[0], mvLX[1]). The variables refxSbC, refySbC, refxC and refyC are derived as follows:
[0967]
[0968] refx C = ((Sign(refxSb C ) * ((Abs(refxSb C ) + 256) » 9) + xC * ((hori_scale_fp + 8) » 4)) » 5 (8-764)
[0969]
[0970] refy C = ((Sign(refySb C ) * ((Abs(refySb C ) + 256) » 9) + yC * ((vert_scale_fp + 8) » 4)) » 5 (8-766)
[0971]
[0972]
[0973] - The variables xInt C , yInt C , xFrac C and yFrac C are derived as follows:
[0974] xInt C = refx C >> 5 (8-767)
[0975] yInt C = refy C >> 5 (8-768)
[0976] xFrac C = refy C & 31 (8-769)
[0977] yFrac C = refy C & 31 (8-770)
[0978] 5.4. Embodiment 3 of reference sample position derivation
[0979] 8.5.6.3.1 Overview
[0980] …
[0981] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma sample units.
[0982] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma sample units.
[0983]
[0984] The motion vector mvLX is set to (refMvLX - mvOffset).
[0985] - If cldx is equal to 0, the following applies:
[0986] - The scaling factors and their fixed-point representations are defined as follows
[0987] hori_scale_fp = (( fRefWidth « 14 ) + ( PicOutputWidthL » 1 )) / PicOutputWidthL (8-753)
[0988] vert_scale_fp = (( fRefHeight « 14 ) + ( PicOutputHeightL » 1 )) / PicOutputHeightL (8-754)
[0989] - Let (xIntL, yIntL) be the luma position in full-sample units and (xFracL, yFracL) be the offset in 1 / 16-sample units. These variables are only used in this subclause to specify fractional-sample positions within the reference sample array refPicLX.
[0990] - The top-left coordinates (xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding are set equal to (xSb + (mvLX[0] » 4), ySb + (mvLX[1] » 4)).
[0991] - For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L= 0..sbHeight - 1 + brdExtSize), the corresponding predicted luma sample value predSamplesLX[x L ][y L ] is derived as follows:
[0992] – Let (refxSb L , refySb L ) and (refx L , refy L ) be the luma positions pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16-sample units. The variables refxSb L , refx L , refySb L , and refy L are derived as follows:
[0993]
[0994]
[0995]
[0996]
[0997] – The variables xInt L , yInt L , xFrac L , and yFrac L are derived as follows:
[0998] xInt L = refx L >> 4 (8-759)
[0999] yInt L = refy L >> 4 (8-760)
[1000] xFrac L = refx L & 15 (8-761)
[1001] yFrac L = refy L & 15 (8-762)
[1002] - If bdofFlag is equal to TRUE (or sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[ xSb ][ ySb ] is equal to TRUE), and one or more of the following conditions are true, derive the predicted luma sample values predSamplesLX[ xL ][ yL ] by invoking the luma integer sample fetching process specified in clause 8.5.6.3.3 with ( xIntL L + ( xFracL L >> 3 ) - 1 ), yIntL L + ( yFracL L >> 3 ) - 1 ) and refPicLX as inputs.
[1003] - x L is equal to 0.
[1004] - x L is equal to sbWidth + 1.
[1005] - y L is equal to 0.
[1006] - y L is equal to sbHeight + 1.
[1007] - Otherwise, derive the predicted luma sample values predSamplesLX[ xL ][ yL ] by invoking the luma sample 8-tap interpolation filter process specified in clause 8.5.6.3.2 with ( xIntL - ( brdExtSize > 0? 1 : 0 ), yIntL - ( brdExtSize > 0? 1 : 0 ) ), ( xFracL, yFracL ), ( xSbInt L , ySbInt L ), refPicLX, hpelIfIdx, sbWidth, sbHeight and ( xSb, ySb ) as inputs.
[1008] - Otherwise ( cldx is not equal to 0 ), the following applies:
[1009] - Let ( xIntC, yIntC ) be the chroma position in full sample units and ( xFracC, yFracC ) be the offset in 1 / 32 sample units. These variables are only used in this subclause to specify the general fractional sample positions within the reference sample array refPicLX.
[1010] - The top-left coordinates (xSbIntC, ySbIntC) of the border block used for reference sample padding are set equal to ((xSb / SubWidthC) + (mvLX[0] » 5), (ySb / SubHeightC) + (mvLX[1] » 5)).
[1011] - For each chroma sample position (xC = 0..sbWidth - 1, yC = 0..sbHeight - 1) within the array of predicted chroma samples predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[ xC ][ yC ] is derived as follows:
[1012] - Let (refxSb C , refySb C ) and (refx C , refy C ) be the chroma positions pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variables refxSb C , refySb C , refx C , and refy C are derived as follows:
[1013]
[1014]
[1015]
[1016]
[1017] - The variables xIntC, yIntC, xFracC, and yFracC are derived as follows:
[1018] xInt C = refx C >> 5 (8-767)
[1019] yInt C = refy C >> 5 (8-768)
[1020] xFrac C = refy C & 31 (8-769)
[1021] yFrac C = refy C & 31 (8-770)
[1022] 5.5. Embodiment 1 of reference sample position clipping
[1023] 8.5.6.3.1 General
[1024] The inputs of the process are:
[1025] - luma position (xSb, ySb) specifying the top-left sample of the current coding sub-block relative to the top-left luma sample of the current picture,
[1026] - variable sbWidth specifying the width of the current coding sub-block,
[1027] - variable sbHeight specifying the height of the current coding sub-block,
[1028] - motion vector offset mvOffset,
[1029] - refined motion vector refMvLX,
[1030] - selected reference picture sample array refPicLX,
[1031] - half-sample interpolation filter index hpelIfIdx,
[1032] - bi-directional optical flow flag bdofFlag,
[1033] - variable cldx specifying the color component index of the current block.
[1034] The outputs of the process are:
[1035] - (sbWidth + brdExtSize) x (sbHeight + brdExtSize) array of predicted sample values predSamplesLX.
[1036] The prediction block boundary extension size brdExtSize is derived as follows:
[1037] brdExtSize = (bdofFlag || (inter_affine_flag[ xSb ][ ySb ] && sps_affine_prof_enabled_flag ))? 2 : 0 (8-752)
[1038] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma sample units.
[1039] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma sample units.
[1040] The motion vector mvLX is set to (refMvLX - mvOffset).
[1041] - If cldx is equal to 0, the following applies:
[1042] - The scaling factors and their fixed-point representations are defined as follows
[1043] hori_scale_fp = (( fRefWidth « 14 ) + ( PicOutputWidthL » 1 )) / PicOutputWidthL (8-753)
[1044] vert_scale_fp = (( fRefHeight « 14 ) + ( PicOutputHeightL » 1 )) / PicOutputHeightL (8-754)
[1045] - Let (xIntL, yIntL) be the luma position given in integer sample units and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are only used in this subclause to specify fractional sample positions within the reference sample array refPicLX.
[1046] - The top-left coordinates (xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding are set equal to (xSb + (mvLX[0] » 4), ySb + (mvLX[1] » 4)).
[1047] - For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L = 0..sbHeight - 1 + brdExtSize) within the array of predicted luma samples predSamplesLX, the corresponding predicted luma sample value predSamplesLX[x L ][y L ] is derived as follows:
[1048] - Let (refxSb L ,refySb L ) and (refx L ,refy L ) be the luma positions pointed to by the motion vector (refMvLX[0], refMvLX[1]) in 1 / 16 sample units. The variables refxSb L ,refx L ,
[1049] refySb L and refy L are derived as follows:
[1050] refxSb L = (( xSb « 4 ) + refMvLX[ 0 ] ) * hori_scale_fp (8-755)
[1051] refx L = (( Sign( refxSb ) * ( ( Abs( refxSb ) + 128 ) » 8 ) + x L * ( ( hori_scale_fp + 8 ) » 4 ) ) + 32 ) » 6 (8-756)
[1052] refySb L = (( ySb « 4 ) + refMvLX[ 1 ] ) * vert_scale_fp (8-757)
[1053] refyL = (( Sign( refySb ) * ( ( Abs( refySb ) + 128 ) » 8 ) + yL * ( ( vert_scale_fp + 8 ) » 4 ) ) + 32 ) » 6 (8-758)
[1054] The variables xInt L , yInt L , xFrac L and yFrac L are derived as follows:
[1055]
[1056]
[1057] xFrac L = refx L & 15 (8-761)
[1058] yFrac L = refy L & 15 (8-762)
[1059] If bdofFlag is equal to TRUE (or sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[ xSb ][ ySb ] is equal to TRUE ), and one or more of the following conditions are true, the luma integer sample fetch process specified in clause 8.5.6.3.3 is invoked to obtain ( xInt L + xFrac L>> 3) - 1), yInt L + (yFrac L >> 3) - 1) and refPicLX as inputs to derive the predicted luma sample value predSamplesLX[ x L ][ y L ].
[1060] - x L is equal to 0.
[1061] - x L is equal to sbWidth + 1.
[1062] - y L is equal to 0.
[1063] - y L is equal to sbHeight + 1.
[1064] - Otherwise, the predicted luma sample value predSamplesLX[ xL ][ yL ] is derived by invoking the luma sample 8-tap interpolation filter process as specified in clause 8.5.6.3.2 with ( xIntL - ( brdExtSize > 0? 1 : 0 ), yIntL - ( brdExtSize > 0? 1 : 0 ) ), ( xFracL, yFracL ), ( xSbInt L , ySbInt L ), refPicLX, hpelIfIdx, sbWidth, sbHeight and ( xSb, ySb ) as inputs.
[1065] - Otherwise ( cldx is not equal to 0 ), the following applies:
[1066] - Let ( xIntC, yIntC ) be the chroma position in whole sample units and ( xFracC, yFracC ) be the offset in 1 / 32 sample units. These variables are only used in this subclause to specify the general fractional sample position within the reference sample array refPicLX.
[1067] - The top-left coordinates ( xSbIntC, ySbIntC ) of the boundary block used for reference sample padding are set equal to ( ( xSb / SubWidthC ) + ( mvLX[ 0 ] » 5 ), ( ySb / SubHeightC ) + ( mvLX[ 1 ] » 5 ) ).
[1068] – For each luma sample position (xC = 0..sbWidth - 1, yC = 0..sbHeight - 1) within the array of predicted luma samples predSamplesLX, the corresponding predicted luma sample value predSamplesLX[xC][yC] is derived as follows:
[1069] – Let (refxSb C , refySb C ) and (refx C , refy C ) be the chroma positions pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. The variables refxSb C , refySb C , refx C , and refy C are derived as follows:
[1070] refxSb C = ((xSb / SubWidthC « 5) + mvLX[0]) * hori_scale_fp (8-763)
[1071] refx C = ((Sign(refxSb C ) * ((Abs(refxSb C ) + 256) » 9) + xC * ((hori_scale_fp + 8) » 4)) + 16) » 5 (8-764)
[1072] refySb C = ((ySb / SubHeightC « 5) + mvLX[1]) * vert_scale_fp (8-765)
[1073] refy C = ((Sign(refySb C ) * ((Abs(refySb C ) + 256) » 9) + yC * ((vert_scale_fp + 8) » 4)) + 16) » 5 (8-766)
[1074] – The variables xInt C , yInt C , xFrac C , and yFrac C are derived as follows:
[1075]
[1076]
[1077] xFrac C = refx C & 31 (8-769)
[1078] yFrac C = refy C & 31 (8-770)
[1079] – the process specified in clause 8.5.6.3.4 is invoked with (xIntC, yIntC), (xFracC, yFracC), (xSbIntC, ySbIntC), sbWidth, sbHeight and refPicLX as inputs to derive the prediction sample values predSamplesLX[xC][yC].
[1080] 5.6. Reference sample position clipping, embodiment 2
[1081] 8.5.6.3.1 Overview
[1082] The inputs of the process are:
[1083] – the luma position (xSb, ySb) specifying the top-left sample of the current coding sub-block relative to the top-left luma sample of the current picture,
[1084] – the variable sbWidth specifying the width of the current coding sub-block,
[1085] – the variable sbHeight specifying the height of the current coding sub-block,
[1086] – the motion vector offset mvOffset,
[1087] – the refined motion vector refMvLX,
[1088] – the selected reference picture sample array refPicLX,
[1089] – the half-sample interpolation filter index hpelIfIdx,
[1090] – the bi-directional optical flow flag bdofFlag,
[1091] – the variable cIdx specifying the color component index of the current block.
[1092] The outputs of the process are:
[1093] – an (sbWidth + brdExtSize) x (sbHeight + brdExtSize) array of prediction sample values predSamplesLX.
[1094] The prediction block boundary extension size brdExtSize is derived as follows:
[1095] brdExtSize = ( bdofFlag || ( inter affine flag[ xSb ][ ySb ] && sps affine prof enabled flag ) )? 2 : 0 (8-752)
[1096] The variable fRefWidth is set equal to PicOutputWidthL of the reference picture in luma sample units.
[1097] The variable fRefHeight is set equal to PicOutputHeightL of the reference picture in luma sample units.
[1098]
[1099] The motion vector mvLX is set equal to ( refMvLX - mvOffset ).
[1100] - If cldx is equal to 0, the following applies:
[1101] - The scaling factor and its fixed-point representation are defined as follows
[1102] hori_scale_fp = ( ( fRefWidth « 14 ) + ( PicOutputWidthL » 1 ) ) / PicOutputWidthL (8-753)
[1103] vert_scale_fp = ( ( fRefHeight « 14 ) + ( PicOutputHeightL » 1 ) ) / PicOutputHeightL (8-754)
[1104] - Let ( xIntL, yIntL ) be the luma position given in whole sample units and ( xFracL, yFracL ) be the offset given in 1 / 16 sample units. These variables are only used in this subclause to specify fractional sample positions within the reference sample array refPicLX.
[1105] - The top-left coordinates ( xSbInt L ,ySbInt L ) of the boundary block used for reference sample padding are set equal to ( xSb + ( mvLX[ 0 ] » 4 ), ySb + ( mvLX[ 1 ] » 4 ) ).
[1106] – For each luma sample position (x L = 0..sbWidth - 1 + brdExtSize, y L = 0..sbHeight - 1 + brdExtSize) within the array of predicted luma samples predSamplesLX, the corresponding predicted luma sample value predSamplesLX[x L ][y L ] is derived as follows:
[1107] – Let (refxSb L , refySb L ) and (refx L , refy L ) be the luma positions pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16-sample units. The variables refxSb L , refx L , refySb L and refy L are derived as follows:
[1108] refxSb L = ((xSb « 4) + refMvLX[0]) * hori_scale_fp (8-755)
[1109] refx L = ((Sign(refxSb) * ((Abs(refxSb) + 128) » 8) + x L *((hori_scale_fp + 8) » 4)) + 32) » 6 (8-756)
[1110] refySb L = ((ySb « 4) + refMvLX[1]) * vert_scale_fp (8-757)
[1111] refyL = ((Sign(refySb) * ((Abs(refySb) + 128) » 8) + yL*((vert_scale_fp + 8) » 4)) + 32) » 6 (8-758)
[1112] – The variables xIntL, yIntL, xFracL and yFracL are derived as follows:
[1113]
[1114]
[1115] xFrac L = refx L & 15 (8-761)
[1116] yFrac L = refy L & 15 (8-762)
[1117] - If bdofFlag is equal to TRUE (or sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[ xSb ][ ySb ] is equal to TRUE), and one or more of the following conditions are true, derive the predicted luma sample values predSamplesLX[ xL ][ yL ] by invoking the luma integer sample fetching process as specified in clause 8.5.6.3.3 with ( xIntL + ( xFracL » 3 ) - 1 ), yIntL + ( yFracL » 3 ) - 1 ) and refPicLX as inputs. L L L L L L
[1118] - xL is equal to 0.
[1119] - x L is equal to sbWidth + 1.
[1120] - y L is equal to 0.
[1121] - y L is equal to sbHeight + 1.
[1122] - Otherwise, derive the predicted luma sample values predSamplesLX[ xL ][ yL ] by invoking the luma sample 8-tap interpolation filtering process as specified in clause 8.5.6.3.2 with ( xIntL - ( brdExtSize > 0? 1 : 0 ), yIntL - ( brdExtSize > 0? 1 : 0 ) ), ( xFracL, yFracL ), ( xSbInt L , ySbInt L ), refPicLX, hpelIfIdx, sbWidth, sbHeight and ( xSb, ySb ) as inputs.
[1123] - Otherwise ( cIdx is not equal to 0 ), the following applies:
[1124] – Let (xIntC, yIntC) be the chromaticity position given in full sample units, and (xFracC, yFracC) be the offset given in 1 / 32 sample units. These variables are used only in this sub-clause to specify the general fractional sample positions within the reference sample array refPicLX.
[1125] – The top-left coordinates (xSbIntC, ySbIntC) of the boundary block used for filling the reference sample points are set to equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)).
[1126] – For each chromaticity sample position (xC = 0..sbWidth-1, yC = 0..sbHeight-1) in the predicted chromaticity sample array predSamplesLX, the corresponding predicted chromaticity sample value predSamplesLX[xC][yC] is derived as follows:
[1127] – Let (refxSb) C ,refySb C ) and (refx C ,refy C ) is the chromaticity position pointed to by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. Variable refxSb C ,refySb C refx C and refy C The derivation is as follows:
[1128] refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763)
[1129] refx C =((Sign(refxSb) C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764)
[1130] refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765)
[1131] refy C =((Sign(refySb)C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766)
[1132] – variable xInt C , yInt C , xFrac C and yFrac C are derived as follows:
[1133]
[1134]
[1135] xFrac C = refy C & 31 (8-769)
[1136] yFrac C = refy C & 31 (8-770)
[1137] – The predicted sample values predSamplesLX[ xC ][ yC ] are derived by invoking the process specified in clause 8.5.6.3.4 with ( xIntC, yIntC ), ( xFracC, yFracC ), ( xSbIntC, ySbIntC ), sbWidth, sbHeight and refPicLX as inputs.
[1138] 5.7 Embodiments of usage of coding tools
[1139] 5.7.1 BDOF on / off control
[1140] – The variable currPic specifies the current picture and the variable bdofFlag is derived as follows:
[1141] – bdofFlag is set equal to TRUE if all of the following conditions are true.
[1142] – sps_bdof_enabled_flag is equal to 1 and slice_disable_bdof_dmvr_flag is equal to 0.
[1143] – predFlagL0[ xSbIdx ][ ySbIdx ] and predFlagL1[ xSbIdx ][ ySbIdx ] are both equal to 1.
[1144] - DiffPicOrderCnt( currPic, RefPicList[ 0 ][ refldxLO ] ) is equal to DiffPicOrderCnt( RefPicList[ 1 ][ refldxLl ], currPic ).
[1145] - RefPicList[ 0 ][ refldxLO ] is a short-term reference picture and RefPicList[ 1 ][ refldxLl ] is a short-term reference picture.
[1146] - MotionModelldc[ xCb ][ yCb ] is equal to 0.
[1147] - merge_subblock_flag[ xCb ][ yCb ] is equal to 0.
[1148] - sym_mvd_flag[ xCb ][ yCb ] is equal to 0.
[1149] - ciip_flag[ xCb ][ yCb ] is equal to 0.
[1150] - Bcwldx[ xCb ][ yCb ] is equal to 0.
[1151] - luma_weight_l0_flag[ refldxLO ] and luma_weight_l1_flag[ refldxLl ] are both equal to 0.
[1152] - cbWidth is greater than or equal to 8.
[1153] - cbHeight is greater than or equal to 8.
[1154] - cbHeight * cbWidth is greater than or equal to 128.
[1155] -
[1156] - cldx is equal to 0.
[1157] - Otherwise, bdofFlag is set equal to FALSE.
[1158] 5.7.2 DMVR on / off control
[1159] - When all of the following conditions are true, dmvrFlag is set equal to 1:
[1160] - sps_dmvr_enabled_flag is equal to 1 and slice_disable_bdof_dmvr_flag is equal to 0
[1161] - general_merge_flag[ xCb ][ yCb ] is equal to 1
[1162] - predFlagL0[ 0 ][ 0 ] and predFlagL1[ 0 ][ 0 ] are both equal to 1
[1163] - mmvd_merge_flag[ xCb ][ yCb ] is equal to 0
[1164] - ciip_flag[ xCb ][ yCb ] is equal to 0
[1165] - DiffPicOrderCnt( currPic, RefPicList[ 0 ][ refIdxL0 ] ) is equal to DiffPicOrderCnt( RefPicList[ 1 ][ refIdxL1 ], currPic )
[1166] - RefPicList[ 0 ][ refIdxL0 ] is a short-term reference picture and RefPicList[ 1 ][ refIdxL1 ] is a short-term reference picture.
[1167] - BcwIdx[ xCb ][ yCb ] is equal to 0
[1168] - luma_weight_l0_flag[ refIdxL0 ] and luma_weight_l1_flag[ refIdxL1 ] are both equal to 0
[1169] - cbWidth is greater than or equal to 8
[1170] - cbHeight is greater than or equal to 8
[1171] - cbHeight * cbWidth is greater than or equal to 128
[1172] -
[1173] 5.7.3 PROF on / off control for reference picture list X
[1174] The variable cbProfFlagLX is derived as follows:
[1175] - cbProfFlagLX is set equal to FALSE if one or more of the following conditions are true.
[1176] - sps_affine_prof_enabled_flag is equal to 0.
[1177] - fallbackModeTriggered is equal to 1.
[1178] - numCpMv is equal to 2, cpMvLX[1][0] is equal to cpMvLX[0][0], and cpMvLX[1][1] is equal to cpMvLX[0][1].
[1179] - numCpMv is equal to 3, cpMvLX[1][0] is equal to cpMvLX[0][0], cpMvLX[1][1] is equal to cpMvLX[0][1], cpMvLX[2][0] is equal to cpMvLX[0][0], and cpMvLX[2][1] is equal to cpMvLX[0][1].
[1180] -
[1181] - [[pic_width_in_luma_samples of the reference picture refPicLX associated with refIdxLX is not equal to pic_width_in_luma_samples of the current picture, respectively.
[1182] - pic_height_in_luma_samples of the reference picture refPicLX associated with refIdxLX is not equal to pic_height_in_luma_samples of the current picture, respectively.]]
[1183] - Otherwise, cbProfFlagLX is set equal to TRUE.
[1184] 5.7.4 PROF on / off control for reference picture list X (second embodiment)
[1185] The variable cbProfFlagLX is derived as follows:
[1186] - If one or more of the following conditions are true, cbProfFlagLX is set equal to FALSE.
[1187] - sps_affine_prof_enabled_flag is equal to 0.
[1188] - fallbackModeTriggered is equal to 1.
[1189] - numCpMv is equal to 2, cpMvLX[1][0] is equal to cpMvLX[0][0], and cpMvLX[1][1] is equal to cpMvLX[0][1].
[1190] - numCpMv is equal to 3, cpMvLX[1][0] is equal to cpMvLX[0][0], cpMvLX[1][1] is equal to cpMvLX[0][1], cpMvLX[2][0] is equal to cpMvLX[0][0], and cpMvLX[2][1] is equal to cpMvLX[0][1].
[1191] -
[1192] - pic width in luma samples of the reference picture refPicLX associated with refIdxLX is not equal to pic width in luma samples of the current picture, respectively.
[1193] - pic height in luma samples of the reference picture refPicLX associated with refIdxLX is not equal to pic height in luma samples of the current picture, respectively.
[1194] Otherwise, cbProfFlagLX is set to TRUE.
[1195] 5.8 Embodiment of conditionally signaling inter-related syntax elements in picture header (on top of JVET-P2001-v9)
[1196] 7.3.2.6 Picture header RBSP syntax
[1197]
[1198]
[1199]
[1200]
[1201]
[1202]
[1203] 7.4.3.6 Picture header RBSP semantics
[1204]
[1205] 7.3.7 Slice header syntax
[1206] 7.3.7.1 General slice header syntax
[1207]
[1208] 5.9 Embodiment of constraints on RPR (on top of JVET-P2001-v14)
[1209] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be pic_width_in_luma_samples and pic_height_in_luma_samples of a reference picture of the current picture referring to this PPS, respectively. Let refPicOutputWidthL and refPicOutputHeightL be PicOutputWidthL and PicOutputHeightL of the reference picture, respectively. The requirement of bitstream conformance is that all the following conditions are met:
[1210] – PicOutputWidthL * 2 shall be greater than or equal to refPicOutputWidthL.
[1211] – PicOutputHeightL * 2 shall be greater than or equal to refPicOutputHeightL.
[1212] – PicOutputWidthL shall be less than or equal to refPicOutputWidthL * 8.
[1213] – PicOutputHeightL shall be less than or equal to refPicOutputHeightL * 8.
[1214] – (PicOutputWidthL - refPicOutputWidthL) * (PicWidthInLumaSamples - refPicWidthInLumaSamples) shall be greater than or equal to 0.
[1215] – (PicOutputHeightL - refPicOutputHeightL) * (PicHeightInLumaSamples - refPicHeightInLumaSamples) shall be greater than or equal to 0.
[1216] - 135 * pic_width_max_in_luma_samples * PicOutputWidthL - (128 * refPicOutputWidthL + 7 * PicOutputWidthL) * PicWidthInLumaSamples shall be greater than or equal to 0.
[1217] - 135 * pic_height_max_in_luma_samples * PicOutputHeightL - (128 * refPicOutputHeightL + 7 * PicOutputHeightL) * PicHeightInLumaSamples shall be greater than or equal to 0.
[1218] 5.10 Embodiment of signaling of wraparound offset (on top of JVET-P2001-v14)
[1219] 7.3.2.3 Sequence parameter set RBSP syntax
[1220]
[1221] 7.3.2.3 Picture parameter set RBSP syntax
[1222]
[1223] pps_ref_wraparound_enabled_flag equal to 1 specifies that horizontal wrap-around motion compensation is applied in inter prediction. pps_ref_wraparound_enabled_flag equal to 0 specifies that horizontal wrap-around motion compensation is not applied. When the value of (CtbSizeY / MinCbSizeY + 1) is less than or equal to (pic_width_in_luma_samples / MinCbSizeY - 1), the value of pps_ref_wraparound_enabled_flag shall be equal to 0.
[1224] pps_ref_wraparound_offset_minus1 plus 1 specifies the offset in units of MinCbSizeY luma samples for calculating the horizontal wrap-around position. The value of pps_ref_wraparound_offset_minus1 shall be in the range of (CtbSizeY / MinCbSizeY) + 1 to (pic_width_in_luma_samples / MinCbSizeY) - 1, inclusive.
[1225] 7.4.4.2 General constraint information semantics
[1226] …
[1227] no_ref_wraparound_constraint_flag equal to 1 specifies that the constraint that the reference picture list shall not wrap around shall be applied. no_ref_wraparound_constraint_flag equal to 0 does not impose such a constraint.
[1228] 8.5.3.2.2 Luma sample bilinear interpolation process
[1229] The input to this process are:
[1230] - the luma position in whole sample units (xInt L , yInt L ),
[1231] - the luma position in fractional sample units (xFrac L , yFrac L ),
[1232] - the luma reference sample array refPicLX L .
[1233] …
[1234] The luma position in whole sample units (xInt i , yInt i ) is derived as follows, for i = 0..1:
[1235] - If subpic_treated_as_pic_flag[ SubPicIdx ] is equal to 1, the following applies:
[1236] xInt i = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i ) (642)
[1237] yInt i = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i ) (643)
[1238] - Otherwise (subpic_treated_as_pic_flag[ SubPicIdx ] is equal to 0), the following applies:
[1239]
[1240] yInt i = Clip3( 0, picH - 1, yInt L + i ) (645)
[1241] …
[1242] 8.5.6.3.2 Luma sample interpolation filter process
[1243] …
[1244] The luma position in full-sample units (xInt i , yInt i ) is derived as follows, for i = 0..7:
[1245] - If subpic_treated_as_pic_flag[ SubPicldx ] is equal to 1, the following applies:
[1246] xInt i = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i - 3 ) (955)
[1247] yInt i = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i - 3 ) (956)
[1248] - Otherwise (subpic_treated_as_pic_flag[ SubPicldx ] is equal to 0), the following applies:
[1249]
[1250] yInt i = Clip3( 0, picH - 1, yInt L + i - 3 ) (958)
[1251] …
[1252] 8.5.6.3.3 Luma integer sample obtaining process
[1253] …
[1254] The luma position in full-sample units (xInt, yInt) is derived as follows:
[1255] – If subpic_treated_as_pic_flag[ SubPicldx ] is equal to 1, the following applies:
[1256] xInt = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L ) (966)
[1257] yInt = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L ) (967)
[1258] – Otherwise, the following applies:
[1259]
[1260]
[1261] yInt = Clip3( 0, picH - 1, yInt L ) (969)
[1262] …
[1263] 8.5.6.3.4. Chroma sample interpolation process
[1264] …
[1265] The variable xOffset is set equal to
[1266] The chroma position in full sample units (xInt i , yInt i ) is derived as follows for i = 0..3:
[1267] – If subpic_treated_as_pic_flag[ SubPicldx ] is equal to 1, the following applies:
[1268] xInt i = Clip3( SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xInt L + i) (971)
[1269] yInt i= Clip3( SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yInt L + i) (972)
[1270] – Otherwise (subpic_treated_as_pic_flag[ SubPicIdx ] is equal to 0), the following applies:
[1271]
[1272] yInt i = Clip3( 0, picH C - 1, yInt C + i - 1)
[1273] …
[1274] 5.11 Embodiments of Deblocking Filtering between Sub-Pictures (on top of JVET-P2001-v14)
[1275] 8.8.3 Deblocking filter process
[1276] 8.8.3.1 Overview
[1277] The input of the process is the reconstructed picture before deblocking, i.e., the array recPicture L and the array recPicture Cb and recPicture Cr when ChromaArrayType is not equal to 0.
[1278] The output of the process is the modified reconstructed picture after deblocking, i.e., the array recPicture L and the array recPicture Cb and recPicture Cr when ChromaArrayType is not equal to 0.
[1279] …
[1280] The deblocking filter process is applied to all coded sub-block edges and transform block edges of a picture, except for the following types of edges:
[1281] – Edges on the boundaries of the picture,
[1282] – Edges coinciding with the boundaries of sub-pictures for which loop_filter_across_subpic_enabled_flag[ SubPicIdx ] is equal to 0,
[1283] - an edge coinciding with a virtual boundary of the picture when VirtualBoundariesDisabledFlag is equal to 1,
[1284] - …
[1285] 8.8.3.2 Deblocking filter process for one direction
[1286] The inputs of this process are:
[1287] - The variable treeType specifies whether a luma component (DUAL_TREE_LUMA) or a chroma component (DUAL_TREE_CHROMA) is currently processed,
[1288] …
[1289] 1. The variable filterEdgeFlag is derived as follows:
[1290] - If edgeType is equal to EDGE_VER and one or more of the following conditions are true, filterEdgeFlag is set equal to 0:
[1291] - The left boundary of the current coding block is the left boundary of the picture.
[1292] - [[ The left boundary of the current coding block is the left or right boundary of the subpicture and loop_filter_across_subpic_enabled_flag[ SubPicldx ] is equal to 0. ]]
[1293] - …
[1294] - Otherwise, if edgeType is equal to EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set equal to 0:
[1295] - The top boundary of the current luma coding block is the top boundary of the picture.
[1296] - [[ The top boundary of the current coding block is the top or bottom boundary of the subpicture and loop_filter_across_subpic_enabled_flag[ SubPicldx ] is equal to 0. ]]
[1297] - …
[1298] 8.8.3.6.6 Filtering process for luma samples using short filter
[1299] …
[1300] When nDp is greater than 0 and pred_mode_plt_flag of the coding unit including the coding block containing sample p0 is equal to 1, nDp is set equal to 0
[1301] When nDq is greater than 0 and pred_mode_plt_flag of the coding unit including the coding block containing sample q0 is equal to 1, nDq is set equal to 0:
[1302]
[1303] 8.8.3.6.7 Filtering process for luma samples using long filter
[1304] …
[1305] When pred_mode_plt_flag of the coding unit including the coding block containing sample p i i is equal to 1, the filtered sample value p i ' is replaced by the corresponding input sample value p i , where i = 0..maxFilterLengthP - 1.
[1306] When pred_mode_plt_flag of the coding unit including the coding block containing sample q i j is equal to 1, the filtered sample value q i ' is replaced by the corresponding input sample value q j , where j = 0..maxFilterLengthQ - 1.
[1307]
[1308] 8.8.3.6.9 Filtering process for chroma samples
[1309] …
[1310] When pred_mode_plt_flag of the coding unit including the coding block containing sample p i i is equal to 1, the filtered sample value p i ' is replaced by the corresponding input sample value p i , where i = 0..maxFilterLengthP - 1.
[1311] When pred_mode_plt_flag of the coding unit including the coding block containing sample q i j is equal to 1, the filtered sample value q i ' is replaced by the corresponding input sample value q iInstead, where i = 0..maxFilterLengthQ - 1:
[1312]
[1313] 6. Example implementations of the disclosed technology
[1314] Figure 7 is a block diagram illustrating an example video processing system 7000 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 7000. The system 7000 can include an input 7002 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be in a compressed or encoded format. The input 7002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[1315] The system 7000 can include a codec component 7004 that can implement various coding or encoding methods described in this document. The codec component 7004 can reduce the average bitrate of video from the input 7002 to the output of the codec component 7004 to produce a coded representation of the video. Thus, the codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 7004 can be stored, or transmitted via a connected communication, as represented by component 7006. The component 7008 can use a bitstream (or coded) representation of the video that is stored or transmitted that is received at the input 7002 to generate pixel values or displayable video to be transmitted to a display interface 7010. The process of generating user-viewable video from a bitstream representation (or bitstream) is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “codec” operations or tools, it should be understood that the codec tools or operations are used at an encoder and that corresponding decoding tools or operations, which are inverse to the results of the codec, will be performed by a decoder.
[1316] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices that are capable of performing digital data processing and / or video display.
[1317] Figure 8is a block diagram of a video processing apparatus 8000. The apparatus 8000 can be used to implement one or more methods described herein. The apparatus 8000 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 8000 can include one or more processors 8002, one or more memories 8004, and video processing hardware 8006. The processor(s) 8002 can be configured to implement one or more methods described in the present document (for example, in Figures 12-13
[1318] Figure 9 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As shown in Figure 9 the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which can be referred to as a video decoding device. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[1319] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the I / O interface 116 via the network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.
[1320] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[1321] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120, which is configured to interface with an external display device.
[1322] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[1323] Figure 10 is a block diagram illustrating an example of a video encoder 200 that can be Figure 9 the video encoder 114 in the system 100 shown.
[1324] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 10 examples, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[1325] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202 (which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[1326] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[1327] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately for explanatory purposes. Figure 10 In examples.
[1328] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[1329] Mode selection unit 203 can select one of the coding modes (intra or inter), e.g., based on the error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 can select a combination of intra and inter predication (CIIP) mode, in which prediction is based on both inter prediction signals and intra prediction signals. In the case of inter prediction, mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer-pixel precision).
[1330] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples for pictures from buffer 213 other than the picture associated with the current video block.
[1331] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[1332] In some examples, motion estimation unit 204 can perform single prediction for a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference picture containing the reference video block in list 0 or list 1 and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[1333] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 and also search for another reference video block for the current video block in a reference picture in list 1. The motion estimation unit 204 can then generate a reference index that indicates the reference picture in list 0 and list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[1334] In some examples, the motion estimation unit 204 can output a complete set of motion information for a decoding process of a decoder.
[1335] In some examples, the motion estimation unit 204 can not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[1336] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[1337] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[1338] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictively signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[1339] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[1340] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[1341] In other examples, there can be no residual data for the current video block for the current video block, such as in a skip mode, and residual generation unit 207 can not perform the subtraction operation.
[1342] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[1343] After transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[1344] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202, resulting in a reconstructed video block associated with the current block for storage in buffer 213.
[1345] After reconstruction unit 212 reconstructs the video block, in-loop filtering operations can be performed to reduce video block effect artifacts in the video block.
[1346] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[1347] Figure 11 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 9 the video decoder 114 in the system 100 shown.
[1348] The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 11In examples of the video decoder 300, the video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[1349] In Figure 11 In examples of the video decoder 300, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200 Figure 10
[1350] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy encoded video data and, from the entropy decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[1351] The motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax elements.
[1352] The motion compensation unit 302 can use the interpolation filter used by the video encoder 20 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.
[1353] The motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter coded block, and other information useful for decoding the encoded video sequence.
[1354] The intra prediction unit 303 can use an intra prediction mode received, for example, in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[1355] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for presentation on a display device.
[1356] Figures 12-17 An example method of implementing the above-described techniques is shown. Figures 1-5 An example method of implementing the above-described techniques is shown.
[1357] Figure 12 A flowchart of an example method 1200 of video processing is shown. The method 1200 includes, at operation 1210, performing a conversion between a video comprising one or more video pictures comprising one or more slices and a bitstream of the video, the bitstream conforming to a format rule that specifies, for a video picture for which all slices in the video picture are coded as I slices, omitting, from a picture header of the video picture, syntax elements related to P slices and B slices.
[1358] Figure 13 A flowchart of an example method 1300 of video processing is shown. The method 1300 includes, at operation 1310, performing a conversion between a video comprising one or more video pictures comprising one or more slices and a bitstream of the video, the bitstream conforming to a format rule that specifies that a picture header of each video picture includes a syntax element that indicates whether all slices in the video picture are coded with a same coding type.
[1359] Figure 14 A flowchart of an example method 1400 of video processing is shown. The method 1400 includes, at operation 1410, performing a conversion between a video comprising one or more video pictures and a bitstream of the video, the bitstream conforming to a format rule that specifies that a picture header of each video picture of the one or more video pictures includes a syntax element that indicates a picture type of the video picture.
[1360] Figure 15A flowchart of an example method 1500 of video processing is shown. The method 1500 includes, at operation 1510, performing a conversion between a video comprising one or more video pictures and a bitstream of the video, the bitstream conforming to a format rule that specifies a syntax element indicating a picture type of a picture is signaled in an access unit (AU) delimiter raw byte sequence payload (RBSP), and wherein the syntax element indicates whether all slices in the picture are I slices.
[1361] Figure 16 A flowchart of an example method 1600 of video processing is shown. The method 1600 includes, at operation 1610, performing a conversion between a video comprising a video picture and a bitstream of the video, the video picture comprising one or more video slices, the bitstream conforming to a format rule that specifies, for a picture for which each of a plurality of slices in the picture is an I slice, an indication of a slice type is excluded from slice headers of the plurality of slices in the bitstream during encoding or is inferred to be an I slice during decoding.
[1362] Figure 17 A flowchart of an example method 1700 of video processing is shown. The method 1700 includes, at operation 1710, making a determination for a conversion between a video comprising a W slice or a W picture and a bitstream of the video as to whether one or more non-W-related syntax elements are signaled in a slice header of a W slice or a picture header of a W picture, wherein W is I, B, or P.
[1363] The method 1700 includes, at operation 1720, performing the conversion based on the determination.
[1364] Next, a list of preferred solutions for some embodiments is provided.
[1365] 1. A method of video processing, comprising: performing a conversion between a video comprising one or more video pictures and a bitstream of the video, the video picture comprising one or more slices, wherein the bitstream conforms to a format rule, and wherein the format rule specifies, for a video picture for which all slices in the one or more video pictures are coded as I slices, syntax elements related to P slices and B slices are omitted from a picture header of the video picture.
[1366] 2. The method of solution 1, wherein a first syntax element indicating that all slices of a video unit are I slices is signaled in a picture header.
[1367] 3. The method of solution 2, wherein whether a second syntax element is signaled in the bitstream is based on the first syntax element, and wherein the second syntax element indicates slice type information in a slice header of a slice associated with the picture header.
[1368] 4. The method according to solution 3, wherein the second syntax element is excluded from the bitstream and inferred to be slice type.
[1369] 5. The method according to solution 3, wherein the second syntax element is signaled in the bitstream and is equal to one of a plurality of predetermined values based on conformance requirements.
[1370] 6. The method according to solution 3, wherein the first syntax element indicating that a video unit comprises all I slices is signaled in an access unit (AU) delimiter raw byte sequence payload (RBSP) associated with at least one of the I slices.
[1371] 7. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream of the video, the video picture comprising one or more slices, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a picture header of each video picture comprises a syntax element indicating whether all slices in the video picture are coded with a same coding type.
[1372] 8. The method according to solution 7, wherein all slices are coded as I slices, P slices, or B slices.
[1373] 9. The method according to solution 7, wherein a slice header of a slice does not include slice type information, and the slice is inferred to be an I slice due to the syntax element in the picture header indicating that all slices are I slices.
[1374] 10. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a picture header of each video picture of the one or more video pictures comprises a syntax element indicating a picture type thereof.
[1375] 11. The method according to solution 10, wherein due to the syntax element indicating that the picture is an I picture, a slice type of one or more slices in the picture is only allowed to indicate I slices.
[1376] 12. The method according to solution 10, wherein due to the syntax element indicating that the picture is a non-I picture, a slice type of one or more slices in the picture indicates I slices and / or B slices and / or P slices.
[1377] 13. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a picture type of a picture is signaled in an access unit (AU) delimiter raw byte sequence payload (RBSP), and wherein the syntax element indicates whether all slices in the picture are I slices.
[1378] 14. A method of video processing, comprising performing a conversion between a video comprising a video picture and a bitstream of the video, the video picture comprising one or more video slices, wherein the bitstream conforms to a format rule, and wherein for a picture for which each of a plurality of slices in the picture is an I slice, the format rule specifies that an indication of a slice type is excluded from slice headers of the plurality of slices in the bitstream during encoding or is inferred to be an I slice during decoding.
[1379] 15. The method of solution 14, wherein the bitstream is organized such that each of a plurality of slices in the picture is an I slice.
[1380] 16. The method of solution 14, wherein the bitstream is organized such that no B slice or P slice is included in the picture.
[1381] 17. A method of video processing, comprising, for a conversion between a video comprising W slices or W pictures and a bitstream of the video, making a determination as to whether one or more non-W related syntax elements are signaled in a slice header of a W slice or a picture header of a W picture, wherein W is I, B, or P; and performing the conversion based on the determination.
[1382] 18. The method of solution 17, wherein W is I, and wherein non-W is B or P.
[1383] 19. The method of solution 17, wherein W is B, and wherein non-W is I or P.
[1384] 20. The method of any of solutions 17 to 19, wherein the one or more syntax elements are excluded from the bitstream as a result of all slices in the picture being W slices.
[1385] 21. The method of any of solutions 17 to 19, wherein the one or more syntax elements are conditionally signaled in the bitstream as a result of all slices in the picture being W slices.
[1386] 22. The method of any of solutions 17 to 21, wherein the one or more syntax elements comprise reference picture related syntax elements in a picture header.
[1387] 23. The method of any of solutions 17 to 21, wherein the one or more syntax elements comprise inter slice related syntax elements in a picture header.
[1388] 24. The method of any of solutions 17 to 21, wherein the one or more syntax elements comprise inter prediction related syntax elements in a picture header.
[1389] 25. The method of any of solutions 17 to 21, wherein the one or more syntax elements comprise bi-prediction related syntax elements in a picture header.
[1390] 26. The method of any of solutions 1 to 25, wherein the conversion comprises decoding the video from the bitstream.
[1391] 27. The method of any of solutions 1 to 25, wherein the conversion comprises encoding the video into the bitstream.
[1392] 28. A method of writing a bitstream representing a video to a computer-readable recording medium, comprising generating the bitstream from the video according to the method of any of solutions 1 to 25, and writing the bitstream to the computer-readable recording medium.
[1393] 29. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 28.
[1394] 30. A computer-readable medium having stored thereon instructions that, when executed, cause a processor to implement a method recited in any one or more of solutions 1 to 27.
[1395] 31. A computer-readable medium storing a bitstream generated according to any of solutions 1 to 28.
[1396] 32. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 28.
[1397] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, video compression algorithms can be applied during a conversion from a pixel representation of a video to a corresponding bitstream, and vice versa. A bitstream of a current video block can correspond, for example, to bits that are collocated or scattered at different places within the bitstream as defined by the syntax. For example, a macroblock can be encoded from transformed and coded error residual values and also using bits in headers and other fields in the bitstream.
[1398] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[1399] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[1400] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or by a combination of computer hardware and special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[1401] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[1402] Although the patent document includes many details, these should not be construed as limiting the scope of any subject matter or claimed content to such details, but rather to the features, concepts and the like that are specifically described herein. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments or in any suitable sub-combination. Moreover, although features can be described above as acting in certain combinations and initially claimed as such, one or more features from a claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[1403] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring this particular order or sequential order of operations, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[1404] Only some implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: For the conversion between a first video image of a video comprising one or more video images and the bitstream of the video, the value of a first syntax element included in the image header of the first video image is determined, wherein the first syntax element indicates whether all stripes included in the first video image are I stripes; Based on the value of the first syntax element, determine whether to transmit a second syntax element via signal in the strip header of the strip included in the first video image, wherein the second syntax element indicates the strip type of the strip; The conversion is performed based on the determination. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, the second syntax element is omitted from the stripe headers of all stripes included in the first video image, and the value of the second syntax element is inferred to indicate an I stripe. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, one or more syntax elements are omitted from the image header of the first video image in the bitstream. Wherein, the one or more syntax elements include a third syntax element specifying the use of a first prediction mode for the first video image. In the first prediction mode, the motion vector notified by signaling is refined based on at least one motion vector having an offset from the motion vector notified by signaling. Wherein, when the value of the first syntax element does not indicate that all stripes included in the first video image are I stripes, the one or more syntax elements are conditionally included in the image header.
2. The method according to claim 1, wherein, The one or more syntax elements further include at least one of the following: temporal_mvp_enabled_flag, mvd_l1_zero_flag, a fourth syntax element specifying the use of a second prediction mode for the first video image, or a fifth syntax element specifying the use of a third prediction mode for the first video image. In the second prediction mode, for a video block in the first video image, a bidirectional optical flow tool is used to obtain a motion vector offset based on at least one gradient value corresponding to a sample point in a reference block of the video block, and In the third prediction mode, for a video block in the first video image, initial prediction samples of sub-blocks of the video block encoded and decoded using affine mode are generated, and optical flow operations are applied to generate final prediction samples of the sub-blocks by deriving prediction refinement based on motion vector differences dMvH and / or dMvV, where dMvH and dMvV indicate motion vector differences along the horizontal and vertical directions, respectively.
3. The method according to claim 1, wherein, The one or more syntax elements also include at least one of the following: log2_diff_min_qt_min_cb_inter_slice, max_mtt_hierarchy_depth_inter_slice, log2_diff_max_bt_min_qt_inter_slice, or log2_diff_max_tt_min_qt_inter_slice.
4. The method according to claim 1, wherein, The intra_slice_allowed flag is conditionally included in the image header.
5. The method according to claim 1, wherein, The first video image includes one or more slices, wherein the maximum slice width of the one or more slices is defined as the maximum luminance slice width in codec tree block units, and the maximum slice height of the one or more slices is defined as the maximum luminance slice height in codec tree block units.
6. The method according to claim 1, wherein, The first video image includes one or more stripes, wherein the maximum stripe height of the one or more stripes is defined in codec tree block units.
7. The method according to claim 1, wherein, The conversion includes decoding the video from the bitstream.
8. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
9. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For the conversion between a first video image and the bitstream of a video comprising one or more video images, the value of a first syntax element included in the image header of the first video image is determined, wherein, The first syntax element indicates whether all stripes included in the first video image are I stripes; Based on the value of the first syntax element, determine whether to transmit a second syntax element via signal in the strip header of the strip included in the first video image, wherein the second syntax element indicates the strip type of the strip; The conversion is performed based on the determination. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, the second syntax element is omitted from the stripe headers of all stripes included in the first video image, and the value of the second syntax element is inferred to indicate an I stripe. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, one or more syntax elements are omitted from the image header of the first video image in the bitstream. Wherein, the one or more syntax elements include a third syntax element specifying the use of a first prediction mode for the first video image. In the first prediction mode, the motion vector notified by signaling is refined based on at least one motion vector having an offset from the motion vector notified by signaling. Wherein, when the value of the first syntax element does not indicate that all stripes included in the first video image are I stripes, the one or more syntax elements are conditionally included in the image header.
10. The apparatus according to claim 9, wherein, The one or more syntax elements further include at least one of the following: temporal_mvp_enabled_flag, mvd_l1_zero_flag, a fourth syntax element specifying the use of a second prediction mode for the first video image, or a fifth syntax element specifying the use of a third prediction mode for the first video image. In the second prediction mode, for a video block in the first video image, a bidirectional optical flow tool is used to obtain a motion vector offset based on at least one gradient value corresponding to a sample point in a reference block of the video block, and In the third prediction mode, for a video block in the first video image, initial prediction samples of sub-blocks of the video block encoded and decoded using affine mode are generated, and optical flow operations are applied to generate final prediction samples of the sub-blocks by deriving prediction refinement based on motion vector differences dMvH and / or dMvV, where dMvH and dMvV indicate motion vector differences along the horizontal and vertical directions, respectively.
11. The apparatus according to claim 9, wherein, The one or more syntax elements also include at least one of the following: log2_diff_min_qt_min_cb_inter_slice, max_mtt_hierarchy_depth_inter_slice, log2_diff_max_bt_min_qt_inter_slice, or log2_diff_max_tt_min_qt_inter_slice.
12. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: For the conversion between a first video image and the bitstream of a video comprising one or more video images, the value of a first syntax element included in the image header of the first video image is determined, wherein, The first syntax element indicates whether all stripes included in the first video image are I stripes; Based on the value of the first syntax element, determine whether to transmit a second syntax element via signal in the strip header of the strip included in the first video image, wherein the second syntax element indicates the strip type of the strip; The conversion is performed based on the determination. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, the second syntax element is omitted from the stripe headers of all stripes included in the first video image, and the value of the second syntax element is inferred to indicate an I stripe. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, one or more syntax elements are omitted from the image header of the first video image in the bitstream. Wherein, the one or more syntax elements include a third syntax element specifying the use of a first prediction mode for the first video image. In the first prediction mode, the motion vector notified by signaling is refined based on at least one motion vector having an offset from the motion vector notified by signaling. Wherein, when the value of the first syntax element does not indicate that all stripes included in the first video image are I stripes, the one or more syntax elements are conditionally included in the image header.
13. A non-transitory computer-readable recording medium for storing a bitstream of video, the bitstream being generated by a method performed by a video processing apparatus, wherein the method includes: For a first video image of a video that includes one or more video images, determine the value of a first syntax element included in the image header of the first video image, wherein the first syntax element indicates whether all stripes included in the first video image are I stripes; Based on the value of the first syntax element, determine whether to transmit a second syntax element via signal in the strip header of the strip included in the first video image, wherein the second syntax element indicates the strip type of the strip; The bit stream is generated based on the determination. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, the second syntax element is omitted from the stripe headers of all stripes included in the first video image, and the value of the second syntax element is inferred to indicate an I stripe. Specifically, in response to the value of the first syntax element indicating that all stripes included in the first video image are I stripes, one or more syntax elements are omitted from the image header of the first video image in the bitstream. Wherein, the one or more syntax elements include a third syntax element specifying the use of a first prediction mode for the first video image. In the first prediction mode, the motion vector notified by signaling is refined based on at least one motion vector having an offset from the motion vector notified by signaling. Wherein, when the value of the first syntax element does not indicate that all stripes included in the first video image are I stripes, the one or more syntax elements are conditionally included in the image header.