Method and device for video processing and medium

By employing sub-block boundary-based deblocking or filtering processes and interleaved affine prediction in video encoding and decoding, the problem of insufficient encoding and decoding efficiency in existing technologies is solved. In particular, when dealing with complex motion patterns, more efficient video block motion compensation and quality improvement are achieved.

CN121890089APending Publication Date: 2026-04-17DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-09-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially when dealing with complex motion patterns, where existing technologies struggle to efficiently perform motion compensation for video blocks.

Method used

A deblocking or filtering process based on sub-block boundaries is adopted, combined with interleaved affine prediction and template matching affine, to convert and process video blocks according to information such as strip type, frame type and image order counting distance.

Benefits of technology

It improves the efficiency of video encoding and decoding, especially when dealing with complex motion patterns, enhances the motion compensation effect of video blocks, reduces block artifacts, and improves the quality of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121890089A_ABST
    Figure CN121890089A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. In the method, for a conversion between a current video block of the video and a bitstream of the video, information is determined based on a slice type, the information regarding application of a sub-block boundary based deblocking or filtering process to the current video block. The conversion is performed based on the information. The information indicates at least one of whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to sub-block motion compensation in video encoding and decoding. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: determining information based on a stripe type for a conversion between a current video block and a bitstream of the video, the information relating to the application of a sub-block boundary-based deblocking or filtering process to the current video block; and performing the conversion based on the information, wherein the information indicates at least one of: whether or how a sub-block boundary-based deblocking or filtering process is applied. The method according to the first aspect of this disclosure determines whether and / or how a sub-block boundary-based deblocking or filtering process is applied based on the stripe type.

[0005] In a second aspect, another method for video processing is proposed. This method includes: for a conversion between a current video block and a video bitstream, determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of the following: stripe type, frame type, or picture order count (POC) distance between the stripe or frame and its nearest reference frame; and performing the conversion based on the use of at least one of interleaved affine prediction or template matching affine. The method according to the second aspect of this disclosure enables the use of interleaved affine prediction or template matching affine based on stripe type, frame type, or POC.

[0006] In a third aspect, another method for video processing is proposed. This method includes: a conversion between a current video block and a video bitstream; determining the application of interleaving prediction for affine modes based on sequence resolution or frame resolution; and performing the conversion based on the application of the interleaving prediction. The method according to the third aspect of this disclosure enables the application of interleaving prediction for affine modes based on sequence / frame resolution.

[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.

[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first, second, or third aspect of this disclosure.

[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a device for video processing. The method includes: determining information based on stripe type, the information relating to the application of a sub-block boundary-based deblocking or filtering process to the current video block; and generating a bitstream based on the information, wherein the information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a device for video processing. The method includes: determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of: strip type, frame type, or picture order count (POC) distance between a strip or frame and its nearest reference frame; and generating the bitstream based on the use of at least one of interleaved affine prediction or template matching affine.

[0011] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a device for video processing. The method includes: determining the use of interleaving prediction for affine modes based on sequence resolution or frame resolution; and generating the bitstream based on the use of interleaving prediction.

[0012] In a ninth aspect, a method for storing a bitstream of video is proposed. The method includes: determining information based on stripe type, the information relating to the application of a sub-block boundary-based deblocking or filtering process to the current video block of the video; generating a bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[0013] In a tenth aspect, a method for storing a bitstream of video is proposed. The method includes: determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of the following: strip type, frame type, or picture sequence count (POC) distance between a strip or frame and its nearest reference frame; generating a bitstream based on the use of at least one of interleaved affine prediction or template matching affine; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] In the eleventh aspect, a method for storing video bitstreams is proposed. The method includes: determining the use of interleaving prediction for affine modes based on sequence resolution or frame resolution; generating a bitstream based on the use of interleaving prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0015] This summary aims to present, in a simplified form, the concept choices further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0016] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same parts.

[0017] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 This shows the positions of non-adjacent candidates in the ECM; Figure 6 An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The position of the inherited affine motion predictor is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10The locations of candidate positions for the constructive affine Merge pattern are shown; Figure 11 The spatial nearest neighbor used to derive the affine Merge candidate is shown; Figure 12 This shows candidates for constructive affine merges, ranging from non-nearest neighbors. Figure 13 An example of generating HAPC is shown; Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 16 The template and the corresponding reference template are shown; Figure 17 A template and a reference template are shown for a block with sub-block motion that uses motion information of the current block's sub-blocks; Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; Figure 19 An example of interleaved prediction is shown; Figures 20A-20G An exemplary partitioning pattern for a 16×16 block is shown; Figures 21A-21D An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas; Figures 22A-22C An example is shown of deriving the MV of a partitioning pattern from another partitioning pattern; Figures 23A-23C An example of selecting a partitioning mode based on block dimensions is shown; Figure 24A and Figure 24B An example is shown that derives the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of another partitioning pattern. Figure 25 An example of a CU-level OBMC is shown; Figure 26 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 27 A flowchart of another method for video processing according to an embodiment of the present disclosure is shown; Figure 28 A flowchart of another method for video processing according to embodiments of the present disclosure is shown; and Figure 29 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0018] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation

[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art to understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0020] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0022] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0024] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. The codec images are codec representations of images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.

[0029] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.

[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0035] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0037] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.

[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0041] In one example, motion estimation unit 204 may indicate a value to video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.

[0042] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.

[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0050] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally adjacent blocks.

[0056] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.

[0057] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values ​​for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.

[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while specific embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term video processing includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compression format to another or at different compression bitrates.

[0062] 1. Brief Overview This disclosure relates to video coding and decoding techniques. Specifically, it relates to affine motion prediction methods in video coding and decoding. These ideas can be applied individually or in various combinations to any standard or non-standard video codec.

[0063] 2. Introduction The exponential growth of multimedia data poses a significant challenge to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations have also jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), H.265 / HEVC, and the latest VVC standard. Since H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding.

[0064] 2.1. MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames and is an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded or decoded, thus transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. In H.265 / HEVC, a contention mechanism is involved, where the best MVP is selected from multiple candidates through rate-distortion optimization (RDO). Specifically, Advanced MVP (AMVP) mode and Merge mode are designed using different motion information signaling strategies. With AMVP mode, the reference index, the MVP candidate index referencing the AMVP candidate list, and the motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list, and the details of the building process for these two modes are described below.

[0065] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, the candidate list of motion vectors is constructed as follows: first, the availability of temporally adjacent locations on the left and top is checked, redundant candidates are removed, and zero vectors are added to give the candidate list a fixed length. Figure 4 The diagram illustrates the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list. For the spatial motion vector candidate derivation, the two motion vector candidates are ultimately based on blocks located as follows: Figure 4 Motion vectors for the five blocks at different locations are derived. The five neighboring blocks located at B0, B1, B2, and A0, A1 are classified into two groups: group A includes the three spatially adjacent blocks above, and group B includes the two spatially adjacent blocks to the left. Two MV candidates are derived in a predefined order using the first available candidates from group A and group B, respectively. For temporal motion vector candidate derivation, a motion vector candidate is derived by sequentially checking based on two different co-locations (lower right (C0) and center (C1)), as follows: Figure 4 As shown. To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list.

[0066] Figure 5 The positions of non-adjacent candidates in the ECM are shown.

[0067] Merge mode Similar to the AMVP mode, the MVP candidate list for the Merge mode also consists of spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, a maximum of four candidates are selected in the order A1, B1, B0, A0, and B2. For temporal Merge Candidate (TMVP) derivation, a maximum of one candidate is selected from two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.

[0068] In VVC, the Merge pattern construction process is further improved by introducing a history-based MVP (HMVP), which incorporates motion information from previously encoded / decoded blocks that may be geographically distant from the current block. In VVC, HMVP Merge candidates are appended to the Merge list after the Spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy during the encoding / decoding process. Whenever a non-sub-block inter-frame encoding / decoding CU exists, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0069] During the standardization of VVC, the non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... Figure 5 As shown.

[0070] 2.2. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6 Affine motion models based on control points are shown, such as (a) a 4-parameter affine model and (b) a 6-parameter affine model. Figure 6 As shown, the affine motion field of a block is described by motion information from two control point motion vectors (4 parameters) or three control point motion vectors (6 parameters).

[0071] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1).

[0072] For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2).

[0073] in( mv0x, mv0y ) is the motion vector of the upper left control point, ( mv1x, mv1y ) is the motion vector of the upper right control point, and ( mv2x, mv2y ) is the motion vector of the lower left control point.

[0074] To simplify motion compensation prediction, block-based affine transformation prediction is applied. Figure 7 An example affine MVF for each sub-block is shown. To derive the motion vector for each 4×4 luma sub-block, as follows... Figure 7 As shown, the motion vector of the center sample point of each sub-block is calculated according to the above equation and rounded to 1 / 16 pixel precision. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.

[0075] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.

[0076] 2.2.1. Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and a signal transmission index indicates which CPVM should be used for the current CU. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; – Constructive affine Merge candidate CPMVP derived using translational MV of neighboring CUs.

[0077] – Zero MV.

[0078] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of the neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Figure 8 The positions of inherited affine motion predictors are shown. Candidate blocks are as follows: Figure 8 As shown. For the predictor on the left, the scan order is A0->A1, and for the predictor above, the scan order is B0->B1->B2. Only the first inherited candidate is selected from each side. No pruning check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 The inheritance of control point motion vectors is shown. For example... Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, then the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are... Obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on... Computed. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... Calculated.

[0079] Constructive affine candidates mean building candidates by combining the translational motion information of each control point's neighbors. The motion information of the control points is derived from... Figure 10 The spatial and temporal nearest neighbors shown are derived. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks are checked by B2->B3->A2, and the MV of the first available block is used. For CPMV2, blocks are checked by B1->B0, and for CPMV3, blocks are checked by A1->A0. If available, the TMVP is used as CPMV4.

[0080] After obtaining the motion signatures (MVs) of the four control points, the affine merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0081] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.

[0082] Figure 10 The locations of candidate positions for the constructive affine Merge pattern are shown.

[0083] After the inherited affine merge candidates and the constructed affine merge candidates are checked, if the list is still not full, zero MV is inserted at the end of the list.

[0084] 2.2.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted sub-CPVM is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.

[0085] – A constructive affine AMVP candidate CPMVP derived using the translation MV of neighboring CUs.

[0086] – Translation MV from the neighboring CU.

[0087] – Zero MV.

[0088] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference picture as the current block are considered. No pruning is applied when inserting inherited affine motion predictors into the candidate list.

[0089] Constructive AMVP candidates are from Figure 10 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is also checked. In the checking order, the first block that has been inter-coded and has the same reference picture as the current CU is used. There is a rule that the current CU is encoded and decoded using a 4-parameter affine mode, and... mv0 and mv1 When all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU is encoded and decoded using a 6-parameter affine mode, and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructive AMVP candidates are set to unavailable.

[0090] Figure 11Spatial nearest neighbors are shown for deriving affine Merge candidates: (a) for deriving inherited affine Merge candidates and (b) for deriving constructed affine Merge candidates.

[0091] If, after inserting valid inherited and constructed AMVP candidates, the affine AMVP list still contains fewer than two candidates, then when available, mv0 , mv1 and mv2 They will be added sequentially as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.

[0092] 2.2.3. New Affine Candidate Derivation Method in ECM-8.0 In ECM-6.0, three additional affine Merge and AMVP candidate derivation methods are integrated: non-adjacent spatial domain candidates, historical parameter-based candidates, regression-based affine candidates, and pixel-based affine motion compensation.

[0093] 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, non-adjacent airspace neighbors are studied to provide candidates for both affine Merge and affine AMVP. The pattern for obtaining non-adjacent airspace candidates is... Figure 11 As shown in the diagram, similar to the non-adjacent regular Merge candidate, the distance between the non-adjacent spatial candidate and the current codec block is also defined based on the width and height of the current CU.

[0094] Figure 11 Motion information of non-adjacent spatial neighbors in the current block is used to generate additional inherited and constructed affine merge candidates. Specifically, to generate inherited candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbor encoded in affine mode from each side (e.g., left and top) of the current block is included. Figure 11 As shown in (a), the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For constructive candidates, as... Figure 11 As shown in (b), the positions of a non-adjacent spatial neighbor on the left and a non-adjacent spatial neighbor above are first determined independently; then, the position of the upper-left neighbor can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and above. The motion information of the three non-adjacent neighbors is used to form a CPMV at the upper-left (A), upper-right (B), and lower-left (C) of the virtual block, which is projected onto the current CU to generate corresponding constructivist candidates, such as... Figure 12 As shown.

[0095] 2.2.3.2. Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to be inherited from previously affine-encoded blocks (which may not be adjacent to the current block). A History Parameter Table (HPT) is established. HPT entries store sets of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entries in the HPT are categorized by reference list and reference index. For each reference list in the HPT, five reference indices are supported. The HPT category (denoted as HPTCat) is calculated in a formulaic manner as follows: (3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of seven entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU using the reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.

[0096] Candidates based on historical affine parameters (HAPC) from... Figure 13 The neighboring 4×4 blocks, denoted as A0, A1, B0, B1, or B2, and the affine parameter sets in the corresponding entries stored in the HPT are derived. The MV of the neighboring 4×4 blocks is used as the base MV. In a formulaic manner, the MV of the current block at position (x, y) is calculated as: , (4) Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 block, and (xbase, ybase) represents the center position of the nearest 4×4 block. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.

[0097] Figure 13An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are directly extracted from an entry in the category HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT, along with the center position of A0 (as the base position) and the MV of block A0 (as the base MV), are used to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as regular Merge candidates. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.

[0098] Figure 13 An example of generating HAPC is shown.

[0099] 2.2.3.3. Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.

[0100] Previously encoded affine CUs can be identified by scanning through non-adjacent positions and the affine HMVP table. Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 14 As shown, the information of the adjacent sub-blocks of the current CU is obtained from the 4×4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.

[0101] For each affine CU, at most two affine candidates can be derived. One has neighboring subblock information, and the other does not. All candidates generated by linear regression are pruned and merged into a single candidate subgroup. When ARMC is enabled, an ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.

[0102] 2.2.3.4. Pixel-based Affine Motion Compensation Using pixel-based affine motion compensation, when OBMC is not applied, the minimum affine sub-block size is set to 1x1 for the luma component and is always set to 1x1 for the chroma component.

[0103] 2.3. Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between the template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and the block in the reference image (i.e., the same size as the template). Figure 15 This illustrates template matching execution over the search area surrounding the initial MV. (Example) Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU.

[0104] In AMVP mode, MVP candidates are determined based on template matching error, selecting the candidate that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. Starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), the TM refines the MVP candidate using an iterative diamond search within a search range of [-8, +8] pixels. Depending on the AMVR mode, the AMVP candidate can be further refined: a cross search is performed using full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel searches. This search process ensures that after the TM process, the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode.

[0105] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. Depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., when AMVR is in half-pixel mode), TMMerge can proceed up to 1 / 8-pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy. Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. When both BM and TM are enabled for a CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.

[0106] 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, Adaptive Reordering of Merge Candidates (ARMC) is proposed to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.

[0107] The reordering method is applied to the regular Merge pattern, Template Matching (TM) Merge pattern, and Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.

[0108] After the Merge candidate list is constructed, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.

[0109] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and its corresponding reference template. Figure 16 The template and its corresponding reference template are shown. For example... Figure 16 As shown, the template includes a reconstructed set of samples adjacent to the current block, while the reference template is located using the same motion information as the current block. When the merge candidate utilizes bidirectional prediction, the reference samples of the merge candidate's template are also generated through bidirectional prediction.

[0110] For sub-block size equal to The sub-block-based Merge candidate has a template on top consisting of several sub-templates of size Wsub×K, and a template on the left consisting of several sub-templates of size K×Hsub. Figure 17 This shows a template and a reference template for a block that uses the motion information of its child blocks. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.

[0111] 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. SbTMVP uses the same co-located images as TMVP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables sub-CU-level motion prediction, while TMVP predicts CU-level motion; second, compared to TMVP extracting temporal motion vectors (MVs) from co-located blocks in the co-located image (co-located blocks are the lower right or center blocks relative to the current CU), SbTMVP applies motion displacements before extracting temporal motion information from the co-located image. These motion displacements are obtained by reusing the MV from one of the spatially neighboring blocks of the current CU.

[0112] Figure 18 The derivation of the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first acquired. If any MV in reference list 0 and reference list 1 points to the same frame, the corresponding MV will be identified as a motion displacement. Otherwise, zero MV will be used as the motion displacement.

[0113] Once the motion displacement is determined, a designated region within the same frame is used to derive the sub-block-level motion field. Assuming... Figure 18 As shown, Motion is used as motion displacement. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is extracted to provide motion information, where the MV scaling operation is performed first to align the reference frame of the temporal motion vector with the reference frame of the current CU.

[0114] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion predictions for the current CU. This sub-CU-level MVP candidate list includes the motion field generated by the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. After performing template matching-based reordering, multiple AFFINE candidates are included in the list, with those having lower costs placed earlier.

[0115] 2.6. Block Removal Process in VVC 8.6.2 Deblocking Filter Process 8.6.2.1 Overview The input to this process is the reconstructed image before removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.

[0116] The output of this process is the modified reconstructed image after removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.

[0117] The vertical edges in the image are first filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the image are filtered. Vertical and horizontal edges in the CTB of each CTU are processed individually on a codec unit basis. The vertical edges of the codec blocks within a codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically towards the right-hand side of the codec block. The horizontal edges of the codec blocks within a codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically towards the bottom of the codec block.

[0118] Note – Although the filtering process is specified on an image basis in this specification, the filtering process can be implemented on an encoding / decoding unit basis with equivalent results, provided that the decoder properly considers the processing dependency order in order to produce the same output value.

[0119] The deblocking filter process is applied to all encoded and transformed block edges of the image, except for the following types of edges: – The edge at the boundary of the image; – When loop_filter_across_tiles_enabled_flag equals 0, the edge that coincides with the tile boundary; – When tile_group_loop_filter_across_tile_groups_enabled_flag equals 0, or tile_group_deblocking_filter_disabled_flag equals 1, the edge that coincides with the top or left boundary of the tile group; – Edges within a tile group where tile_group_deblocking_filter_disabled_flag is equal to 1; – Edges that do not correspond to the 8×8 sample grid boundary of the considered component; – Edges on both sides of the edge are predicted using the chromaticity components within the frame; – Edges of chroma transform blocks that are not part of the edges of the associated transform unit.

[0120] [Editor's note: Once the fragments are integrated, the syntax is adjusted.]

[0121] The edge type (vertical or horizontal) is represented by the variable edgeType, as specified in Table 8.17.

[0122] Table 8.17 – Names associated with edgeType

[0123] The following applies when the tile_group_deblocking_filter_disabled_flag of the current tile group is equal to 0: – The variable treeType is deduced as follows: – If tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, then treeType is set to equal DUAL_TREE_LUMA.

[0124] Otherwise, treeType is set to equal SINGLE_TREE.

[0125] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_VER are taken as inputs, and the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.

[0126] – The horizontal edge is filtered by calling the deblocking filter procedure for one direction as specified in Item 8.6.2.2, where the variable treeType, the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_HOR are taken as inputs, and the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.

[0127] – When tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, the following applies: – The variable treeType is set to equal DUAL_TREE_CHROMA.

[0128] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the reconstructed image before deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_VER are taken as inputs, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) are taken as outputs.

[0129] – The horizontal edge is filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_HOR are taken as inputs, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) are taken as outputs.

[0130] 8.6.2.2 Deblocking filter process for one direction The input to this process is: – The variable treeType, which specifies whether to use a single tree (SINGLE_TREE) or a dual tree to partition the CTU, and when using a dual tree, whether the currently processed component is the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA); – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the reconstructed picture before deblocking, i.e., the array recPictureL; – When ChromaArrayType is not equal to 0, and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr; – The variable edgeType, which specifies whether a vertical edge (EDGE_VER) or a horizontal edge (EDGE_HOR) is filtered.

[0131] The output of this process is the modified reconstructed picture after deblocking, i.e.: – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; – When ChromaArrayType is not equal to 0, and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr.

[0132] For each coding unit with a coded block width of log2CbW, a coded block height of log2CbH, and the position of the top-left sample of the coded block being (xCb, yCb), when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered through the following ordered steps: 1. The coded block width nCbW is set to be equal to 1 << log2CbW, and the coded block height nCbH is set to be equal to 1 << log2CbH.

[0133] 2. The variable filterEdgeFlag is derived as follows: – If edgeType is equal to EDGE_VER and one or more of the following conditions are true, then filterEdgeFlag is set to be equal to 0: – The left boundary of the current coded block is the left boundary of the picture.

[0134] – The left boundary of the current codec block is the left boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.

[0135] – The left boundary of the current codec block is the left boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.

[0136] Otherwise, if edgeType equals EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set to 0: – The upper boundary of the current luminance codec block is the upper boundary of the image.

[0137] – The upper boundary of the current codec block is the upper boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.

[0138] – The upper boundary of the current codec block is the upper boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.

[0139] Otherwise, filterEdgeFlag is set to 1.

[0140] [Editor's note: Once the fragments are integrated, the syntax is adjusted.]

[0141] 3. All elements of the two-dimensional (nCbW) x (nCbH) array edgeFlags are initialized to 0.

[0142] 4. The derivation process for the transform block boundary specified in Item 8.6.2.3 is invoked, where the position (xB0, yB0) is set to equal to (0, 0), the block width nTbW is set to equal to nCbW, the block height nTbH is set to equal to nCbH, the variables treeType, filterEdgeFlag, array edgeFlags, and edgeType are taken as inputs, and the modified array edgeFlags is taken as output.

[0143] 5. The derivation process for the codec subblock boundary specified in Item 8.6.2.4 is invoked, with the position (xCb, yCb), codec block width nCbW, codec block height nCbH, array edgeFlags, and variable edgeType as inputs, and the modified array edgeFlags as output.

[0144] 6. The image sample array recPicture is derived as follows: – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, recPicture is set to equal to the array of reconstructed luminance image samples recPictureL before deblocking.

[0145] Otherwise (treeType equals DUAL_TREE_CHROMA), recPicture is set to equal the array of reconstructed chroma image samples recPictureCb before deblocking.

[0146] 7. The derivation of the boundary filter strength specified in Item 8.6.2.5 is invoked, with the image sample array recPicture, the luminance position (xCb, yCb), the codec block width nCbW, the codec block height nCbH, the variable edgeType, and the array edgeFlags as inputs, and the (nCbW) x (nCbH) array verBs as outputs.

[0147] 8. The edge filtering process is invoked as follows: – If edgeType equals EDGE_VER, then the vertical edge filtering procedure for the codec unit, as specified in entry 8.6.2.6.1, is invoked, with the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array verBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.

[0148] Otherwise, if edgeType equals EDGE_HOR, the horizontal edge filtering procedure for the codec unit as specified in entry 8.6.2.6.2 is invoked, with the variable treeType, the modified reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array horBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.

[0149] 8.6.2.3 Derivation of the Transform Block Boundary The input to this process is: – Position (xB0, yB0), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; – The variable nTbW specifies the width of the current block; – The variable nTbH specifies the height of the current block; – The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA). – Variable filterEdgeFlag; – A two-dimensional (nCbW) x (nCbH) array edgeFlags; – The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered.

[0150] The output of this process is a modified two-dimensional (nCbW) x (nCbH) array edgeFlags.

[0151] The maximum transform block size, maxTbSize, is derived as follows: maxTbSize = (treeType= =DUAL_TREE_CHROMA) ? MaxTbSizeY / 2 :MaxTbSizeY (8 862).

[0152] Depending on maxTbSize, the following applies: - If nTbW is greater than maxTbSize or nTbH is greater than maxTbSize, then the following ordered steps apply.

[0153] 1. The variables newTbW and newTbH are derived as follows: newTbW = ( nTbW>maxTbSize ) ? ( nTbW / 2 ) : nTbW (8 863) newTbH = ( nTbH>maxTbSize ) ? ( nTbH / 2 ) :nTbH (8 864).

[0154] 2. The derivation process for the transform block boundary specified in this entry is invoked, with the position (xB0, yB0), the variable nTbW set to be equal to newTbW, the variable nTbH set to be equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType as inputs, and the output being a modified version of the array edgeFlags.

[0155] 3. If nTbW is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0 + newTbW, yTb0), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.

[0156] 4. If nTbH is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0, yTb0 + newTbH), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.

[0157] 5. If nTbW is greater than maxTbSize and nTbH is greater than maxTbSize, then the derivation procedure for the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal (xTb0 + newTbW, yTb0 + newTbH), the variable nTbW is set to equal newTbW, the variable nTbH is set to equal newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.

[0158] - Otherwise, the following applies: - If edgeType equals EDGE_VER, then edgeFlags[xB0][yB0 + k] (k = 0..nTbH) The value of 1) is derived as follows: - If xB0 equals 0, then edgeFlags[xB0][yB0 + k] is set to equal filterEdgeFlag.

[0159] Otherwise, edgeFlags[xB0][yB0 + k] is set to 1.

[0160] - Otherwise (edgeType equals EDGE_HOR), edgeFlags[xB0 + k][yB0] (k = 0..nTbW) The value of 1) is derived as follows: - If yB0 equals 0, then edgeFlags[xB0 + k][yB0] is set to equal filterEdgeFlag.

[0161] Otherwise, edgeFlags[xB0 + k][yB0] is set to 1.

[0162] 8.6.2.4 Derivation of Encoding / Decoding Sub-Block Boundaries The input to this process is: - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block; - A two-dimensional (nCbW) x (nCbH) array edgeFlags; - The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered.

[0163] The output of this process is a modified two-dimensional (nCbW) x (nCbH) array edgeFlags.

[0164] The number of horizontal codec sub-blocks, numSbX, and the number of vertical codec sub-blocks, numSbY, are derived as follows: - If CupredMode[xCb][yCb] == MODE_INTRA, then numSbX and numSbY are both set to 1.

[0165] Otherwise, numSbX and numSbY are set to equal NumSbX[xCb][yCb] and NumSbY[xCb][yCb], respectively.

[0166] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER and numSbX is greater than 1, then the following applies to i = 1..min( (nCbW / 8 ) 1, numSbX 1), k = 0..nCbH 1: .

[0167] Otherwise, if edgeType equals EDGE_HOR and numSbY is greater than 1, the following applies to j = 1..min( ( nCbH / 8 ) 1, numSbY 1), k = 0..nCbW 1: .

[0168] 8.6.2.5 Derivation of Boundary Filter Strength The input to this process is: - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block; - The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered; - A two-dimensional (nCbW) x (nCbH) array edgeFlags.

[0169] The output of this process is a two-dimensional (nCbW) x (nCbH) array bS that specifies the boundary filter strength.

[0170] The variables xDi, yDj, xN, and yN are derived as follows: - If edgeType equals EDGE_VER, then xDi is set to equal to (i<<3), yDj is set to equal to (j<<2), and xN is set to equal to Max(0, (nCbW / 8). 1), and yN is set to equal to (nCbH / 4). 1.

[0171] Otherwise (edgeType equals EDGE_HOR), xDi is set to equal to (i<<2), yDj is set to equal to (j<<3), and xN is set to equal to (nCbW / 4). 1, and yN is set to equal Max( 0, ( nCbH / 8 ) ). 1).

[0172] For xDi where i = 0..xN and yDj where j = 0..yN, the following applies: - If edgeFlags[xDi][yDj] equals 0, then the variable bS[xDi][yDj] is set to equal 0.

[0173] - Otherwise, the following applies: - The sample values ​​p0 and q0 are derived as follows: - If edgeType equals EDGE_VER, then p0 is set to equal recPicture[xCb + xDi] 1 ][ yCb + yDj ], and q0 is set to equal recPicture [ xCb + xDi ][ yCb + yDj ].

[0174] Otherwise (edgeType equals EDGE_HOR), p0 is set to equal recPicture[xCb + xDi][yCb + yDj] 1], and q0 is set to equal recPicture [ xCb + xDi ][ yCb + yDj ].

[0175] - The variable bS[ xDi ][ yDj ] is derived as follows: - If sample p0 or q0 is in the codec block of a codec unit that uses intra-prediction mode coding and decoding, then bS[xDi][yDj] is set to equal 2.

[0176] Otherwise, if the block edge is also the transform block edge, and sample p0 or q0 is in a transform block containing one or more non-zero transform coefficient levels, then bS[xDi][yDj] is set to equal 1.

[0177] Otherwise, bS[xDi][yDj] is set to 1 if one or more of the following conditions are true: - For prediction of a codec subblock containing sample p0, use a different reference picture or a different number of motion vectors than for prediction of a codec subblock containing sample q0.

[0178] Note 1 – Determining whether the reference pictures used for two codec subblocks are the same or different is based solely on which pictures are referenced, without considering whether an index in reference picture list 0 or reference picture list 1 is used to form the prediction, and also without considering whether the index positions within the reference picture lists are different.

[0179] Note 2 – The number of motion vectors used to predict the top left sample coverage (xSb, ySb) of the codec subblock is equal to PredFlagL0[xSb][ySb] + PredFlagL1[xSb][ySb].

[0180] - A motion vector is used to predict the codec subblock containing sample p0, and a motion vector is used to predict the codec subblock containing sample q0, and the absolute difference between the horizontal or vertical components of the motion vectors used is greater than or equal to 4 (in quarter-luminance samples).

[0181] - Two motion vectors and two different reference images are used to predict the codec subblock containing sample p0, and two motion vectors and the same two reference images are used to predict the codec subblock containing sample q0. When predicting the two codec subblocks, the absolute difference between the horizontal or vertical components of the two motion vectors used for the same reference image is greater than or equal to 4 (in quarter-luminance samples).

[0182] - Two motion vectors for the same reference image are used to predict a codec subblock containing sample p0, and two motion vectors for the same reference image are used to predict a codec subblock containing sample q0, and both of the following conditions are true: - The absolute difference between the horizontal or vertical components of the motion vector in list 0 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in list 1 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples).

[0183] - The absolute difference between the horizontal or vertical components of the motion vector in List 0 used when predicting the codec subblock containing sample p0 and the motion vector in List 1 used when predicting the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in List 1 used when predicting the codec subblock containing sample p0 and the motion vector in List 0 used when predicting the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples).

[0184] Otherwise, the variable bS[ xDi ][ yDj ] is set to 0.

[0185] 8.6.2.6 Edge Filtering Process 8.6.2.6.1 Vertical Edge Filtering Process The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, reconstruct the image before the block, i.e., the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block.

[0186] The output of this process is the modified reconstructed image after removing the blocks, i.e.: - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.

[0187] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the filtering process for the edges in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8)). 1), and yN is set to equal to (nCbH / 4). 1.

[0188] 2. For xDk of k = 0..nN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.

[0189] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.

[0190] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the filtering process for the edges in the chroma codec block of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8)). 1), and yN is set to equal Max(0, (nCbH / 8) 1).

[0191] 2. The variable edgeSpacing is set to equal to 8 / SubWidthC.

[0192] 3. The variable edgeSections is set to equal yN. (2 / SubHeightC).

[0193] 4. For k = 0..xN, the expression equal to k For edgeSpacing, xDk and m = 0..edgeSections, yDm is equal to m<<2, the following applies: - When bS[ xDk SubWidthC ][ yDm When SubHeightC equals 2 and (((xCb / SubWidthC + xDk)>>3)<<3) equals xCb / SubWidthC + xDk, the following ordered steps apply: a. The filtering procedure for the edges of the chroma block, as specified in entry 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.

[0194] b. The filtering procedure for the edges of the chroma block, as specified in Item 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.

[0195] When treeType equals DUAL_TREE_CHROMA, the filtering process for the edges in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8)). 1), and yN is set to equal to (nCbH / 4). 1.

[0196] 2. For xDk of k = 0..xN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.

[0197] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCb is taken as output.

[0198] c. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType, set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCr is taken as output.

[0199] 8.6.2.6.2 Horizontal Edge Filtering Process The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, reconstruct the image before the block, i.e., the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block.

[0200] The output of this process is the modified reconstructed image after removing the blocks, i.e.: - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.

[0201] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the filtering process for the edges in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8)). 1), and xN is set to equal (nCbW / 4). 1.

[0202] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.

[0203] b. The filtering procedure for block edges as specified in Item 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decision dEp, dEp and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.

[0204] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the filtering process for the edges in the chroma codec block of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8)). 1), and yN is set to equal Max(0, (nCbH / 8) 1).

[0205] 2. The variable edgeSpacing is set to equal to 8 / SubHeightC.

[0206] 3. The variable edgeSections is set to equal xN. (2 / SubWidthC).

[0207] 4. For m = 0..yN, the equality is... The following applies to yDm and xDk where k = 0..edgeSections equals k<<2: - when When yCb / SubHeightC + yDm equals 2 and (((yCb / SubHeightC + yDm)>>3)<<3) equals yCb / SubHeightC + yDm, the following ordered steps apply: a. The filtering procedure for the edges of the chroma block, as specified in entry 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.

[0208] b. The filtering procedure for the chroma block edges as specified in Item 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.

[0209] When treeType equals DUAL_TREE_CHROMA, the filtering process for the edges in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8)). 1), and xN is set to equal (nCbW / 4). 1.

[0210] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.

[0211] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCb is taken as output.

[0212] c. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCr is taken as output.

[0213] 8.6.2.6.3 Decision-making process for block edges The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - Position (xBl, yBl), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - Variable bS, which specifies the boundary filter strength; - Variable bD, which specifies the bit depth of the current component.

[0214] The output of this process is: – Includes decision variables dE, dEp, and dEq; – Variable tC.

[0215] If edgeType equals EDGE_VER, then the sample values ​​pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ] (8 867) pi,k = recPictureL[ xCb + xBl i 1 ][ yCb + yBl + k ] (8 868).

[0216] Otherwise (edgeType equals EDGE_HOR), the sample values ​​pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPicture[ xCb + xBl + k ][ yCb + yBl + i ] (8 869) pi,k = recPicture[ xCb + xBl + k ][ yCb + yBl i 1] (8870).

[0217] The variable qpOffset is derived as follows: - If sps_ladf_enabled_flag equals 1 and treeType equals SINGLE_TREE or DUAL_TREE_LUMA, then the following applies: - The variable lumaLevel for reconstructing the brightness level is derived as follows: lumaLevel = ( ( p0,0 + p0,3 + q0,0 + q0,3 )>>2 ) (8 871).

[0218] - The variable qpOffset is set to equal to sps_ladf_lowest_interval_qp_offset and modified as follows: for( i = 0; i <sps_num_ladf_intervals_minus2 + 1; i++ ) { if( lumaLevel>SpsLadfIntervalLowerBound[ i + 1 ] ) qpOffset = sps_ladf_qp_offset[ i ] (8 872) else break } - Otherwise (treeType equals DUAL_TREE_CHROMA), qpOffset is set to 0.

[0219] The variables QpQ and QpP are derived as follows: - If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, then QpQ and QpP are set to the QpY value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.

[0220] - Otherwise (treeType equals DUAL_TREE_CHROMA), QpQ and QpP are set to the QpC value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.

[0221] The variable qP is derived as follows: qP = ( ( QpQ + QpP + 1 )>>1 ) + qpOffset (8 873).

[0222] The value of variable β′ is determined based on the quantization parameter Q derived as follows, as specified in Table 8.18: Q = Clip3( 0, 63, qP + ( tile_group_beta_offset_div2<<1 ) ) (8 874) Where tile_group_beta_offset_div2 is the value of the syntax element tile_group_beta_offset_div2 that contains the patch group of sample points q0,0.

[0223] The variable β is derived as follows: .

[0224] The value of variable tC′ is determined based on the quantization parameter Q derived as follows, as specified in Table 8.18: Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that contains the patch group of sample points q0,0.

[0225] The variable tC is derived as follows: .

[0226] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:

[0227] dpq0 = dp0 + dq0 (8 882) dpq3 = dp3 + dq3 (8 883) dp = dp0 + dp3 (8 884) dq = dq0 + dq3 (8 885) d = dpq0 + dpq3 (8 886).

[0228] 2. The variables dE, dEp, and dEq are set to 0.

[0229] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.

[0230] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values ​​p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.

[0231] c. The variable dpq is set to equal 2. dpq3.

[0232] d. For the sample location (xCb + xBl, yCb + yBl + 3), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values ​​p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.

[0233] e. The variable dE is set to equal 1.

[0234] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.

[0235] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.

[0236] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.

[0237] - Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:

[0238] dpq0 = dp0 + dq0 (8 891) dpq3 = dp3 + dq3 (8 892) dp = dp0 + dp3 (8 893) dq = dq0 + dq3 (8 894) d = dpq0 + dpq3 (8 895).

[0239] 2. The variables dE, dEp, and dEq are set to 0.

[0240] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.

[0241] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values ​​p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.

[0242] c. The variable dpq is set to equal 2. dpq3.

[0243] d. For the sample location (xCb + xBl + 3, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values ​​p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.

[0244] e. The variable dE is set to equal 1.

[0245] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.

[0246] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.

[0247] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.

[0248] Table 8.18 - Deriving threshold variables β′ and tC′ from input Q

[0249] 8.6.2.6.4 Filtering process for block edges The input to this process is: - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - Position (xBl, yBl), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - Variables including dE, dEp, and dEq for decision-making; - Variable tC.

[0250] The output of this process is a modified image sample array recPicture.

[0251] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The sample values ​​pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ] (8 896) pi,k = recPictureL[ xCb + xBl i 1 ][ yCb + yBl + k ] (8 897).

[0252] 2. When dE is not equal to 0, for each sample point location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: a. The filtering procedure for the sample points as specified in item 8.6.2.6.7 is invoked, where the sample point values ​​pi,k,qi,k (i = 0..3) and positions (xPi, yPi) are set to equal (xCb + xBl). i 1, yCb + yBl + k) and (xQi, yQi) are set to be equal to (xCb + xBl + i, yCb + yBl + k) (i = 0..2), decision dE, variables dEp and dEq and variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample values ​​pi' and qj' are taken as outputs.

[0253] b. When nDp is greater than 0, i = 0..nDp The filtered sample value pi' of 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl i 1 ][ yCb + yBl + k ] = pi' (8 898).

[0254] c. When nDq is greater than 0, j = 0..nDq The filtered sample value qj' of 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + j ][ yCb + yBl + k ] = qj' (8 899).

[0255] - Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The sample values ​​pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl + i ] (8 900) pi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl i 1] (8 901).

[0256] 2. When dE is not equal to 0, for each sample point location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: a. The filtering procedure for the sample points as specified in item 8.6.2.6.7 is invoked, where the sample point values ​​pi,k,qi,k (i = 0..3) and positions (xPi, yPi) are set to equal (xCb + xBl + k, yCb + yBl) i 1) and (xQi, yQi) are set to be equal to (xCb + xBl + k, yCb + yBl + i) (i = 0..2), decision dE, variables dEp and dEq and variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample values ​​pi' and qj' are taken as outputs.

[0257] b. When nDp is greater than 0, i = 0..nDp The filtered sample value pi' of 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl i 1 ] = pi' (8 902).

[0258] c. When nDq is greater than 0, j = 0..nDq The filtered sample value qj' of 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl + j ] = qj' (8 903).

[0259] 8.6.2.6.5 Filtering process for chroma block edges This procedure is invoked only if ChromaArrayType is not equal to 0.

[0260] The input to this process is: - Chroma image sample array s′; - Chroma position (xCb, yCb), which specifies the position of the top-left sample of the current chroma codec block relative to the top-left chroma sample of the current image; - Chroma position (xBl, yBl), which specifies the position of the top-left sample of the current chroma block relative to the top-left sample of the current chroma codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - The variable cQpPicOffset specifies the offset of the image-level color quantization parameter.

[0261] The output of this process is the modified chroma image sample array s′.

[0262] If edgeType equals EDGE_VER, then the values ​​of pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + i ][ yCb + yBl + k ] (8 904) pi,k = s′[ xCb + xBl i 1 ][ yCb + yBl + k ] (8 905).

[0263] Otherwise (edgeType equals EDGE_HOR), the sample values ​​pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + k ][ yCb + yBl + i ] (8 906) pi,k = s′[ xCb + xBl + k ][ yCb + yBl i 1] (8907).

[0264] The variables QpQ and QpP are set to be equal to the QpY value of the codec unit, which includes the codec block containing samples q0,0 and p0,0, respectively.

[0265] If ChromaArrayType equals 1, then the variable QpC is determined based on the index qPi derived as follows, as specified in Table 8.15: qPi = ( ( QpQ + QpP + 1 )>>1 ) + cQpPicOffset (8 908).

[0266] Otherwise (ChromaArrayType is greater than 1), the variable QpC is set to equal Min(qPi, 63).

[0267] Note – The variable cQpPicOffset provides adjustments to the values ​​of pps_cb_qp_offset or pps_cr_qp_offset based on whether the filtered chrominance component is a Cb or Cr component. However, to avoid the need to change the adjustment amount within the image, the filtering process does not include adjustments to the values ​​of tile_group_cb_qp_offset or tile_group_cr_qp_offset.

[0268] The value of variable tC′ is determined based on the colorimetric parameter Q derived as follows, as specified in Table 8.18: Q = Clip3( 0, 65, QpC + 2 + ( tile_group_tc_offset_div2<<1 ) ) (8909) Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that contains the patch group of sample points q0,0.

[0269] The variable tC is derived as follows: .

[0270] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then for each sample location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: 1. The filtering procedure for chromaticity samples as specified in item 8.6.2.6.8 is invoked, where the sample values ​​pi,k,qi,k (i = 0..1), and the position (xCb + xBl) are specified. 1, yCb + yBl + k) and (xCb + xBl, yCb + yBl + k) and variable tC are used as inputs, and the filtered sample values ​​p0′ and q0′ are used as outputs.

[0271] 2. Replace the corresponding samples in the sample array s' with the filtered sample values ​​p0′ and q0′, as shown below: s′[ xCb + xBl ][ yCb + yBl + k ]= q0′ (8 911) s′[ xCb + xBl 1 ][ yCb + yBl + k ] = p0′ (8 912).

[0272] - Otherwise (edgeType equals EDGE_HOR), for each sample location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: 1. The filtering procedure for chromaticity samples as specified in item 8.6.2.6.8 is invoked, where the sample values ​​pi,k,qi,k (i = 0..1) and positions (xCb + xBl + k, yCb + yBl) are specified. 1) and (xCb + xBl + k, yCb + yBl) and variable tC are used as inputs, and the filtered sample values ​​p0′ and q0′ are used as outputs.

[0273] 2. Replace the corresponding samples in the sample array s' with the filtered sample values ​​p0′ and q0′, as shown below: s′[ xCb + xBl + k ][ yCb + yBl ]= q0′ (8 913) s′[ xCb + xBl + k ][ yCb + yBl 1 ] = p0′ (8 914).

[0274] 8.6.2.6.6 Decision-making process for sample points The input to this process is: - Sample values ​​p0, p3, q0, and q3; - Variables dpq, β, and tC.

[0275] The output of this process is the variable dSam, which contains the decision.

[0276] The variable dSam is defined as follows: - If dpq is less than (β >> 2), Abs(p3) p0 ) + Abs( q0 q3) is less than (β>>3), and Abs(p0) q0) is less than (5) If tC + 1 )>>1, then dSam is set to equal to 1.

[0277] Otherwise, dSam is set to 0.

[0278] 8.6.2.6.7 Filtering process for sample points The input to this process is: - Sample values ​​pi and qi for i = 0..3; - The positions of pi and qi are (xPi, yPi) and (xQi, yQi), where i = 0..2; - Variable dE; - dEp and dEq, respectively, contain the decision variables for filtering samples p1 and q1; - Variable tC.

[0279] The output of this process is: - The number of filtered samples, nDp and nDq; - i = 0..nDp 1. j = 0..nDq The filtered sample values ​​pi′ and qj′ of 1.

[0280] Depending on the value of dE, the following applies: - If variable dE equals 2, then nDp and nDq are both set to equal 3, and the following strong filtering applies: p0′ = Clip3( p0 2 tC, p0 + 2 tC, (p2 + 2) p1 + 2 p0 + 2 q0 + q1 + 4 )>>3 ) (8 915) p1′ = Clip3( p1 2 tC, p1 + 2 tC, ( p2 + p1 + p0 + q0 + 2 )>>2 ) (8916) p2′ = Clip3( p2 2 tC, p2 + 2 tC, ( 2 p3 + 3 p2 + p1 + p0 + q0 + 4)>>3 ) (8 917) q0′ = Clip3( q0 2 tC, q0 + 2 tC, (p1 + 2) p0 + 2 q0 + 2 q1 + q2 + 4)>>3) (8 918) q1′ = Clip3( q1 2 tC, q1 + 2 tC, ( p0 + q0 + q1 + q2 + 2 )>>2 ) (8919) q2′= Clip3( q2 2 tC, q² + 2 tC, (p0 + q0 + q1 + 3) q² + 2 q3 + 4)>>3) (8 920).

[0281] - Otherwise, nDp and nDq are both set to 0, and the following weak filtering applies: - The following applies:

[0282] - when Less than tC At 10:00, the following sequential steps apply: - The filtered sample values ​​p0′ and q0′ are defined as follows: = Clip3( tC, tC, (8 922) p0′ = Clip1Y( p0 + (8 923) q0′ = Clip1Y( q0 ) (8 924).

[0283] - When dEp equals 1, the filtered sample value p1′ is defined as follows: p = Clip3( ( tC>>1 ), tC>>1, ( ( ( p2 + p0 + 1 )>>1 ) p1 + )>>1) (8 925) p1′ = Clip1Y( p1 + p ) (8 926).

[0284] - When dEq equals 1, the filtered sample value q1′ is defined as follows: q = Clip3( ( tC>>1 ), tC>>1, ( ( ( q2 + q0 + 1 )>>1 ) q1 )>>1 )(8 927) q1′ = Clip1Y( q1 + q) (8 928).

[0285] - nDp is set to equal dEp + 1, and nDq is set to equal dEq + 1.

[0286] nDp is set to 0 when nDp is greater than 0 and one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 ][ yP0 ] equals 1.

[0287] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.

[0288] nDq is set to 0 when nDq is greater than 0 and one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xQ0 ][ yQ0 ] equals 1.

[0289] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.

[0290] 8.6.2.6.8 Filtering process for chromaticity samples This procedure is invoked only if ChromaArrayType is not equal to 0.

[0291] The input to this process is: - Colorimetric sample values ​​pi and qi for i = 0..1; - The chromaticity positions of p0 and q0 are (xP0, yP0) and (xQ0, yQ0); - Variable tC.

[0292] The output of this process is the filtered sample values ​​p0′ and q0′.

[0293] The filtered sample values ​​p0′ and q0′ are derived as follows: = Clip3( tC, tC, ( ( ( ( q0 p0 )<<2 ) + p1 q1 + 4)>>3)) (8 929) p0′ = Clip1C( p0 + (8 930) q0′ = Clip1C( q0 ) (8 931).

[0294] The filtered sample value p0′ is replaced by the corresponding input sample value p0 when one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 SubWidthC ][yP0 SubHeightC equals 1.

[0295] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.

[0296] The filtered sample value q0′ is replaced by the corresponding input sample value q0 when one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[xQ0] SubWidthC ][yQ0 SubHeightC equals 1.

[0297] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.

[0298] 8.6.3 Sample point adaptive compensation process 8.6.3.1 Overview The input to this process is the array recPictureL of reconstructed image samples before adaptive compensation, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.

[0299] The output of this process is a modified reconstructed image sample array saoPictureL after sample adaptive compensation, and arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0.

[0300] This process is performed on top of CTB after the deblocking filter process for the decoded image is completed.

[0301] The sample values ​​in the modified reconstructed image sample array saoPictureL, as well as the sample values ​​in the arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0, were initially set to be equal to the sample values ​​in the reconstructed image sample array recPictureL, as well as the sample values ​​in the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0, respectively.

[0302] For each CTU with CTB location (rx, ry), where rx = 0..PicWidthInCtbsY 1 and ry = 0..PicHeightInCtbsY 1. The following applies: – When the tile_group_sao_luma_flag of the current tile group is equal to 1, the CTB modification procedure as specified in entry 8.6.3.2 is invoked, where recPicture is set to equal recPictureL, cIdx is set to equal 0, (rx, ry), and nCtbSw and nCtbSh are both set to equal CtbSizeY as input, and the modified luminance image sample array saoPictureL is output.

[0303] – When ChromaArrayType is not equal to 0 and the tile_group_sao_chroma_flag of the current slice group is equal to 1, the CTB modification process as specified in item 8.6.3.2 is called, where recPicture is set to be equal to recPictureCb, cIdx is set to be equal to 1, (rx, ry), nCtbSw is set to be equal to (1 << CtbLog2SizeY) / SubWidthC, and nCtbSh is set to be equal to (1 << CtbLog2SizeY) / SubHeightC as inputs, and the modified chroma picture sample array saoPictureCb is used as the output.

[0304] – When ChromaArrayType is not equal to 0 and the tile_group_sao_chroma_flag of the current slice group is equal to 1, the CTB modification process as specified in item 8.6.3.2 is called, where recPicture is set to be equal to recPictureCr, cIdx is set to be equal to 2, (rx, ry), nCtbSw is set to be equal to (1 << CtbLog2SizeY) / SubWidthC, and nCtbSh is set to be equal to (1 << CtbLog2SizeY) / SubHeightC as inputs, and the modified chroma picture sample array saoPictureCr is used as the output.

[0305] 8.6.3.2 CTB Modification Process The inputs to this process are: – The picture sample array recPicture for color component cIdx; – The variable cIdx that specifies the color component index; – A pair of variables (rx, ry) that specify the CTB position; – The CTB width nCtbSw and height nCtbSh.

[0306] The output of this process is the modified picture sample array saoPicture for color component cIdx.

[0307] The variable bitDepth is derived as follows: – If cIdx is equal to 0, then bitDepth is set to be equal to BitDepthY.

[0308] – Otherwise, bitDepth is set to be equal to BitDepthC.

[0309] The position (xCtb, yCtb) specifies the position of the top-left sample of the current CTB for the color component cIdx relative to the top-left sample of the current image component cIdx, and it is derived as follows: (xCtb, yCtb) = (rx) nCtbSw,ry nCtbSh)(8 932)。

[0310] The current sample point locations within the CTB are derived as follows: (xSi, ySj) = (xCtb + i, yCtb + j) (8 933) (xYi, yYj) = (cIdx = = 0) ? (xSi, ySj) : (xSi SubWidthC, ySj SubHeightC) (8 934).

[0311] For i = 0..nCtbSw 1 and j = 0..nCtbSh All sample locations (xSi, ySj) and (xYi, yYj) of 1 depend on the values ​​of pcm_loop_filter_disabled_flag, pcm_flag[xYi][yYj], and cu_transquant_bypass_flag of the codec unit that covers the codec block that includes recPicture[xSi][ySj], as follows: – If one or more of the following conditions are true, then saoPicture[xSi][ySj] will not be modified: – pcm_loop_filter_disabled_flag and pcm_flag[ xYi ][ yYj ] are both equal to 1.

[0312] – cu_transquant_bypass_flag equals 1.

[0313] – SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 0.

[0314] [Editor's Note: The highlighted section will be modified based on future decision changes / quantification bypasses.] Otherwise, if SaoTypeIdx[cIdx][rx][ry] equals 2, then the following ordered steps apply: 1. The values ​​of hPos[k] and vPos[k] for k = 0..1 are based on SaoEoClass[cIdx][rx][ry] as specified in Table 8.19.

[0315] 2. The variable edgeIdx is derived as follows: – The modified sample point locations (xSik′, ySjk′) and (xYik′, yYjk′) are derived as follows: (xSik′, ySjk′) = (xSi + hPos[ k ], ySj + vPos[ k ]) (8 935) (xYik′, yYjk′) = (cIdx = = 0) ? (xSik′, ySjk′) : (xSik′ SubWidthC,ySjk′ SubHeightC) (8 936).

[0316] – If one or more of the following conditions are true for all sample locations (xSik′, ySjk′) and (xYik′, yYjk′) at k = 0..1, then edgeIdx is set to equal to 0: – The sample point at location (xSik′, ySjk′) is outside the image boundary.

[0317] – The sample points at location (xSik′, ySjk′) belong to different patch groups, and one of the following two conditions is true: – MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSi ][ ySj ] belongs is equal to 0.

[0318] – MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSik′ ][ ySjk′ ] belongs is equal to 0.

[0319] – loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at position (xSik′, ySjk′) belongs to a different tile.

[0320] [Editor's Note: When merging slices that do not contain slice groups, modify the highlighted portion.] Otherwise, edgeIdx is derived as follows: – The following applies: edgeIdx = 2 + Sign( recPicture[ xSi ][ ySj ] recPicture[ xSi + hPos[0 ] ][ ySj + vPos[ 0 ] ]) + Sign( recPicture[ xSi ][ ySj ] recPicture[ xSi + hPos[ 1 ] ][ ySj +vPos[ 1 ] ]) (8 937).

[0321] – When edgeIdx is equal to 0, 1, or 2, edgeIdx is modified as follows: edgeIdx = ( edgeIdx == 2 ) ? 0 : ( edgeIdx + 1 ) (8 938).

[0322] 3. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) 1, recPicture[xSi][ySj]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ edgeIdx ]) (8 939).

[0323] – Otherwise (SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 1), the following ordered steps apply: 1. The variable bandShift is set to equal bitDepth. 5.

[0324] 2. The variable saoLeftClass is set to equal sao_band_position[ cIdx ][ rx ][ ry ].

[0325] 3. The list `bandTable` is defined with 32 elements, and all elements are initially set to 0. Then, its four elements (indicating the starting position of the band for the explicit offset) are modified as follows: for( k = 0; k<4; k++ ) bandTable[ ( k + saoLeftClass )&31 ] = k + 1 (8 940).

[0326] 4. The variable bandIdx is set to equal bandTable[recPicture[xSi][ySj]>>bandShift].

[0327] 5. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) 1, recPicture[xSi][ySj]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ bandIdx ]) (8 941).

[0328] Table 8.19 – Specifications for hPos and vPos based on the sample adaptive compensation category

[0329] 2.7. OBMC in ECM When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information from neighboring blocks with weighted prediction.

[0330] The following conditions should not be used for OBMC: • When OBMC is disabled at the SPS level.

[0331] • When the current block has intra-frame mode or IBC mode.

[0332] • When applying LIC to the current block.

[0333] • When the current luminance block area is less than or equal to 32.

[0334] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using motion information from neighboring sub-blocks. It is enabled for sub-block-based codec tools. • Affine AMVP mode; • Affine Merge pattern and sub-block-based temporal motion vector prediction (SbTMVP). • Bilateral matching based on sub-blocks.

[0335] When OBMC mode is used in CIIP mode with LMCS, inter-frame blending is performed before the LMCS mapping of the inter-frame samples. LMCS is applied to the blended inter-frame samples, which are combined with the intra-frame samples to which LMCS is applied in CIIP mode.

[0336] ,in This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by the motion of neighboring blocks in the original domain, and and It's the weight.

[0337] 2.8. Intertwined Prediction To overcome the problems in sub-block-based prediction, interleaving prediction in video encoding and decoding is proposed.

[0338] Using interleaved prediction, a block is divided into sub-blocks with more than one partitioning pattern. A partitioning pattern is defined as the way a block is divided into sub-blocks, including the size and position of the sub-blocks. For each partitioning pattern, a corresponding prediction block can be generated by deriving the motion information of each sub-block based on the partitioning pattern. Therefore, even for a single prediction direction, multiple prediction blocks can be generated from multiple partitioning patterns. Alternatively, for each prediction direction, only the partitioning pattern can be applied.

[0339] Suppose there are X partitioning patterns, and the X predicted blocks (denoted as P0, P1, ..., PX-1) of the current block are generated through sub-block-based predictions with X partitioning patterns. The final prediction (denoted as P) of the current block can be generated as follows: (15) Where (x, y) are the coordinates of the pixels in the block, and These are the weighted values ​​of Pi. Without losing generality, assume... , where N is a non-negative value. Figure 19 An example of interleaved prediction with two partitioning modes is shown.

[0340] The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0341] Use of interleaving prediction for different codec tools 1. Interleaving prediction can be applied to one, some, or all of the codecs that have sub-block-based prediction. In one example, affine prediction applies interleaving prediction, while other codecs with sub-block-based prediction (such as ATMVP, STMVP, FRUC, and BIO) do not. In another example, affine, ATMVP, and STMVP apply interleaving prediction.

[0342] Definition of partitioning mode 2. The partitioning pattern can have different shapes, sizes, or positions of the sub-blocks. In one example, the partitioning pattern could result in irregular sub-block sizes. Figures 20A-20G Several exemplary partitioning patterns for 16×16 blocks are shown. Figure 20A In this context, blocks are divided into 4x4 sub-blocks, as in JEM. Figure 20B In the middle, the block is divided into 8×8 sub-blocks. Figure 20C and Figure 20D In the middle, the block is divided into 8×4 sub-blocks and 4×8 sub-blocks respectively. Figure 20E and Figure 20F In the block, the block is also divided into 4×4 sub-blocks, but with different positions. Pixels at the block boundary that cannot be divided into complete 4×4 sub-blocks can be divided into smaller sub-blocks of size 2×4, 4×2, or 2×2, such as... Figure 20E As shown, they can also be merged into adjacent 4×4 sub-blocks to form larger sub-blocks of size 6×4, 4×6, or 6×6, such as... Figure 20F As shown; in Figure 20GIn this model, blocks are also divided into 8×8 sub-blocks, but in different positions. Pixels at the block boundaries that cannot be divided into complete 8×8 sub-blocks can be divided into smaller sub-blocks of size 8×4, 4×8, or 4×4.

[0343] 3. The shape and size of the sub-blocks in the sub-block prediction can depend on the shape and / or size of the encoded block and / or the encoded block information (e.g., whether it is an affine or ATMVP mode).

[0344] a. In one example, when the current block has a size of M×N, the child block has a size of 4×N (or 8×N, etc.).

[0345] b. In one example, when the current block has a size of M×N, the child block has a size of M×4 (or M×8, etc.).

[0346] c. In one example, when the current block has a size of M×N and M>N, the child block has a size of A×B, such as 8×4, which is greater than B; otherwise, the child block has a size of B×A, such as 4×8.

[0347] d. In one example, assuming the current block has a size of M×N, the sub-block has a size of A×B when M×N<=T (or Min(M,N)<=T, or Max(M,N)<=T, etc.), and a size of C×D when M×N>T (or Min(M,N)>T, or Max(M,N)>T, etc.), where A<=C and B<=D. For example, if M×N<=256, the sub-block has a size of 4×4; otherwise, the sub-block has a size of 8×8.

[0348] Enabling / disabling interleaving prediction and the encoding / decoding process of interleaving prediction 4. Whether to apply interleaved prediction depends on the inter-frame prediction direction.

[0349] a. In one example, interleaved forecasts can be applied to bidirectional forecasts but not to unidirectional forecasts.

[0350] b. In one example, when multiple hypotheses are applied, interleaved predictions can be applied to a prediction direction when there is more than one reference block.

[0351] 5. How to apply interleaved prediction depends on the inter-frame prediction direction.

[0352] a. In one example, a bidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two different partitioning patterns for two different reference lists. In one example, when predicting from reference list 0 (L0), the block is divided into 4×8 sub-blocks, such as... Figure 20DAs shown, however, when predicting from reference list 1 (L1), the block is divided into 8×4 sub-blocks, as follows: Figure 20C As shown. And the final prediction P is calculated as (16) Where P0 and P1 are predictions from L0 and L1, respectively. w0 and w1 are weighted values ​​for L0 and L1, respectively. Without loss of generality, assume... (Where N is a non-negative integer value).

[0353] b. In one example, a unidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two or more distinct partitioning patterns. For example, the prediction PL for list L (L=0 or 1) is calculated as follows: (17) Where XL is the number of partitioning patterns for list L; It is a prediction generated using the i-th partitioning pattern, and yes The weighted value. For example, XL is 2. Using the 0th partitioning pattern, the block is divided into 4×8 sub-blocks, such as... Figure 20D As shown. Using the first partitioning pattern, the block is divided into 8×4 sub-blocks, as follows. Figure 20C As shown.

[0354] c. In one example, a bidirectional prediction block with sub-block-based predictions is considered as a combination of two unidirectional prediction blocks from L0 and L1, respectively. Predictions from each list can be derived as described in the example above. The final prediction P can be computed as... (18) The parameters a and b are two additional weights applied to the two internal prediction blocks. In one example, both a and b are equal to 1.

[0355] d. In one example, for a block encoded using multiple hypotheses, there may be more than one prediction block generated by different partitioning patterns for each prediction direction (or reference image list). Multiple prediction blocks can be used to generate a final version with additional weights applied. In one example, the additional weights can be set to 1 / M, where M is the total number of prediction blocks generated.

[0356] 6. Whether and how interleaving prediction can be applied can be transmitted from the encoder to the decoder at the sequence level, picture level, view level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, PU level, or TU level, or slice level, or slice group level, or region level (which may include multiple CU / PU / TU / LCU)). This information can be transmitted via signaling in the first block of the sequence parameter set (SPS), view parameter set (VPS), picture parameter set (PPS), stripe header (SH), picture header, sequence header, or slice level or slice group level, CTU (also known as LCU), CU, PU, ​​TU, or region.

[0357] a. In one example, interleaving prediction implicitly applies to existing sub-block methods such as ATMVP, STMVP, FRUC, BIO, or affine. In this example, no additional signaling cost is required.

[0358] b. In another example, new sub-block Merge candidates generated by interleaving prediction are inserted into the Merge list, such as interleaving prediction + ATMVP, interleaving prediction + STMVP, interleaving prediction + FRUC, etc.

[0359] c. In one example, a flag can be transmitted via signaling to indicate whether interleaving prediction is used. In one example, if the current block is affine inter-frame encoded, the flag is transmitted via signaling to indicate whether interleaving prediction is used.

[0360] d. In one example, if the current block is encoded and decoded using affine Merge and one-way prediction is applied, a flag can be signaled to indicate whether interleaved prediction is used.

[0361] e. In one example, if the current block is encoded / decoded using affine Merge, a flag can be transmitted via signaling to indicate whether interleaved prediction is used.

[0362] f. In one example, if the current block is encoded and decoded using an affine Merge codec and one-way prediction is applied, interleaved prediction can always be used.

[0363] g. In one example, if the current block is encoded or decoded using affine Merge, interleaving prediction can always be used.

[0364] h. In one example, a flag indicating whether to use interleaving prediction can be inherited without being transmitted via signaling.

[0365] i. In one example, inheritance can be used if the current block is encoded or decoded using affine Merge.

[0366] ii. In one example, the flag can be inherited from the flag of a neighboring block that inherits the affine model.

[0367] iii. In one example, the flag is inherited from a predefined neighboring block (such as the neighboring block to the left or above).

[0368] iv. In one example, the flag can be inherited from the first encountered affine-coded neighboring block.

[0369] v. In one example, if no neighboring blocks are affine encoded, the flag can be presumed to be zero.

[0370] vi. In one example, the flag can only be inherited if one-way prediction is applied to the current block.

[0371] vii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to be inherited are in the same CTU.

[0372] viii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to be inherited are in the same CTU line.

[0373] ix. In one example, when the affine model is derived from a temporal neighbor block, the flag cannot be inherited from the flag of the neighbor block.

[0374] x. In one example, the flag cannot be inherited from the flag of a neighboring block located in the same LCU or LCU line or video data processing unit (such as 64×64 or 128×128).

[0375] xi. In one example, how the flag is transmitted and / or deduced via signaling may depend on the block dimension of the current block and / or encoded / decoded information.

[0376] i. In one example, if the reference image is the current image, then interleaved predictions are not applied.

[0377] i. In one example, if the reference image is the current image, a flag indicating whether interleaving prediction is used is not transmitted through the signal.

[0378] Weighted values 7. The weighting value w is fixed. For example, in equations (15) and (16) .

[0379] 8. The weighting value can depend on the position and the partitioning pattern, that is, for different (x, y). It may differ. Alternatively, the weighting may further depend on the codec tool based on sub-block prediction (e.g., affine or ATMVP) and / or other codec information (e.g., skip or non-skip modes and / or MV information, etc.).

[0380] 9. Weighting values ​​can be transmitted from the encoder to the decoder at the sequence level, picture level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, or PU level, or region level, which may include multiple CUs / PUs / TUs / LCUs)). They can be transmitted via signaling in the first block of the sequence parameter set (SPS), picture parameter set (PPS), stripe header (SH), CTU (also known as LCU), CU or PU, or region.

[0381] a. In an alternative solution, in addition, for some blocks, weights can be inherited from spatial and / or temporal neighboring blocks.

[0382] Partially intertwined predictions 10. In one embodiment, interleaved predictions are applied to a portion of the current block. Prediction samples at some locations are calculated as a weighted sum of two or more sub-block-based predictions. Prediction samples at other locations are not. For example, these prediction samples are copied from sub-block-based predictions with a specific partitioning pattern. Figures 21A-21D An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas.

[0383] a. In one example, the current block is predicted using sub-block-based predictions P1 and P2 with partitioning patterns D0 and D1. The final prediction is calculated as P = w0 × P0 + w1 × P1. At some locations, w0 ≠ 0 and w1 ≠ 0. But at some other locations, w0 = 1 and w1 = 0, meaning that interleaved predictions are not applied to those locations.

[0384] b. In one example, such as Figure 21A As shown, interleaving prediction is not applied to the four corner sub-blocks.

[0385] c. In one example, such as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns.

[0386] d. In one example, such as Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.

[0387] e. In one example, such as Figure 21DAs shown, interleaving prediction is not applied to the topmost sub-row, the bottommost sub-row, the leftmost sub-column, and the rightmost sub-column.

[0388] f. In one example, whether and how partial interleaving prediction is applied can depend on the size / shape of the current block.

[0389] i. For example, if the size of the current block meets certain conditions, the interleaving prediction is applied to the entire block; otherwise, the interleaving prediction is applied to a portion (or parts) of the block. Conditions include, but are not limited to: (assuming the width and height of the current block are W and H respectively, and T, T1, and T2 are integer values): 1. W>=T1 and H>=T2; 2. W <= T1 and H <= T2; 3. W>=T1 or H>=T2; 4. W <= T1 or H <= T2; 5. W + H >= T; 6. W + H <= T; 7. W×H>=T; 8. W×H<=T.

[0390] ii. For example, if W>=H, then as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.

[0391] iii. For example, if W>H, then as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.

[0392] g. It is proposed that whether and how interleaved predictions are applied can differ for different regions within a block.

[0393] i. For example, suppose the current block is predicted using sub-block-based predictions P1 and P2 with partitioning patterns D0 and D1. The final prediction is computed as P(x, y) = w0 × P0(x, y) + w1 × P1(x, y). If the location (x, y) belongs to a sub-block of dimension S0 × H0 with partitioning pattern D0; and belongs to a sub-block of dimension S1 × H1 with partitioning pattern D1, then w0 = 1 and w1 = 0 if one or more of the following conditions are met. (That is, interleaved predictions are not applied to this location).

[0394] 1. S1 < T1; 2. H1 < T2; 3. S1 < T1 and H1 < T2; 4. S1 < T1 or H1 < T2; T1 and T2 are integers. For example, T1 = T2 = 4.

[0395] Encoder problem 11. In one embodiment, interlaced prediction is not applied during the motion estimation (ME) process.

[0396] a. For example, interlaced prediction is not applied during the ME process for 6-parameter affine prediction.

[0397] b. For example, if the size of the current block satisfies certain conditions, such as (assuming the width and height of the current block are W and H respectively, and T, T1, T2 are integer values): i. W >= T1 and H >= T2; ii. W <= T1 and H <= T2; iii. W >= T1 or H >= T2; iv. W <= T1 or H <= T2; v. W + H >= T; vi. W + H <= T; vii. W × H >= T; viii. W × H <= T.

[0398] c. For example, if the current block is partitioned from a parent block and the parent block does not select the affine mode at the encoder, then interlaced prediction is not applied during the ME process.

[0399] i. Alternatively, if the current block is partitioned from a parent block and the parent block does not select the affine mode at the encoder, then the affine mode is not checked at the encoder.

[0400] MV derivation In the following discussion, SatShift(x, n) is defined as .

[0401] Shift(x, n) is defined as Shift(x, n) = (x + offset0) >> n.

[0402] In one example, offset0 and / or offset1 are set to (1 << n) >> 1 or (1 << (n - 1)). In another example, offset0 and / or offset1 are set to 0.

[0403] 12. The MV of each sub-block within a partitioning pattern can be derived directly from the affine model (such as by using equation (1)), or it can be derived from the MV of a sub-block within another partitioning pattern.

[0404] a. In one example, the MV of sub-block B with partitioning pattern 0 can be derived from the MVs of all or some sub-blocks within partitioning pattern 1 that overlaps with sub-block B.

[0405] b. Figures 22A-22C An example is shown. Figure 22A In this process, the MV1(x,y) of a specific sub-block within partitioning mode 1 will be derived. Figure 22B The diagram shows partitioning pattern 0 (solid) and partitioning pattern 1 (dashed), indicating that there are four sub-blocks in partitioning pattern 0 that overlap with specific sub-blocks in partitioning pattern 1. Figure 22C The diagram shows four MVs for four sub-blocks within partitioning pattern 0 that overlap with a specific sub-block within partitioning pattern 1: MV0(x-2, y-2), MV0(x+2, y-2), MV0(x-2, y+2), and MV0(x+2, y+2). MV1(x, y) is then derived from MV0(x-2, y-2), MV0(x+2, y-2), MV0(x-2, y+2), and MV0(x+2, y+2).

[0406] c. Suppose that the MV' of a sub-block within partitioning pattern 1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks within partitioning pattern 0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.

[0407] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.

[0408] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.

[0409] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.

[0410] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.

[0411] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.

[0412] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.

[0413] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.

[0414] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.

[0415] x. with Figures 22A-22C For example, MV1(x, y) can be derived as: 1. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 2. MV1(x,y) = Shift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 3. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 4. MV1(x,y) = Shift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 5. MV1(x,y) = SatShift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 6. MV1(x,y) = Shift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 7. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y+2), 1); 8. MV1(x,y) = Shift( MV0(x-2,y-2)+ MV0(x+2,y+2), 1); 9. MV1(x,y) = SatShift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 10. MV1(x,y) = Shift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 11. MV1(x,y) = SatShift(MV0(x+2,y-2) +MV0(x+2,y+2), 1); 12. MV1(x,y) = Shift( MV0(x+2,y-2) +MV0(x+2,y+2), 1); 13. MV1(x,y) = SatShift(MV0(x+2,y-2)+MV0(x-2,y+2), 1); 14. MV1(x,y) = Shift( MV0(x+2,y-2)+MV0(x-2,y+2), 1); 15. MV1(x,y) = MV0(x-2,y-2); 16. MV1(x,y) = MV0(x+2,y-2); 17. MV1(x,y) = MV0(x-2,y+2); 18. MV1(x,y) = MV0(x+2,y+2).

[0416] 13. The choice of partitioning mode can depend on the width and height of the current block. Figures 23A-23C An example of selecting a partitioning mode based on block dimensions is shown.

[0417] a. For example, if width > T1 and height > T2 (e.g., T1 = T2 = 4), then both partitioning modes are selected. Figure 23A Examples of two partitioning patterns are shown.

[0418] b. For example, if the height is less than or equal to T2 (e.g., T2 = 4), then the other two partitioning modes are selected. Figure 23B Examples of two partitioning patterns are shown.

[0419] c. For example, if the width is less than or equal to T1 (e.g., T1 = 4), then two more partitioning patterns are selected. Figure 23C Examples of two partitioning patterns are shown.

[0420] 14. The MV of each sub-block within a partitioning mode of a color component C1 can be derived from the MV of the sub-block within another partitioning mode of another color component C0.

[0421] a. For example, C1 refers to a color component that is encoded / decoded after another color component, such as Cb or Cr or U or V or R or B.

[0422] b. For example, C0 refers to a color component that is encoded / decoded before another color component, such as Y or G.

[0423] c. In one example, how to derive the MV of a sub-block within a partitioning mode of a color component from the MV of a sub-block within another partitioning mode of another color component can depend on the color format, such as 4:2:0, or 4:2:2, or 4:4:4.

[0424] d. In one example, the MV of a sub-block B in a color component C1 with partitioning mode C1Pt (t=0 or 1) can be derived from the MV of all or some sub-blocks in a color component C0 with partitioning mode C0Pr (r=0 or 1) that overlaps with sub-block B after scaling or scaling coordinates according to the color format.

[0425] i. In one example, C0Pr is always equal to C0P0.

[0426] e. Figure 24A and Figure 24B An example is shown. Figure 24A and Figure 24B This example illustrates the derivation of the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of a partitioning pattern. The color format is 4:2:0. The MV of a sub-block in the Cb component is derived from the MV of a sub-block in the Y component.

[0427] i. at Figure 24A On the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 0 will be derived. Figure 24A The right side shows four Y-blocks within partition mode 0, which overlap with sub-block B (Cb) when scaled at a 2:1 ratio. Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').

[0428] ii. In Figure 24B On the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 1 will be derived. Figure 24B The right side shows four Y-blocks within partition mode 0, which overlap with sub-block B (Cb) when scaled at a 2:1 ratio. Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').

[0429] f. Suppose that the MV' of a sub-block of color component C1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks of color component C0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.

[0430] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.

[0431] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.

[0432] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.

[0433] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.

[0434] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.

[0435] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.

[0436] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.

[0437] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.

[0438] x. with Figure 24A and Figure 24B For example, MVCbt(x',y') t = 0 or 1 can be derived as: 1. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 2. MVCbt(x',y') = Shift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 3. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 4. MVCbt(x',y') = Shift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 5. MVCbt(x',y') = SatShift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 6. MVCbt(x',y') = Shift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 7. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y+2), 1); 8. MVCbt(x',y') = Shift( MV0(x-2,y-2)+ MV0(x+2,y+2), 1); 9. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 10. MVCbt(x',y') = Shift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 11. MVCbt(x',y') = SatShift(MV0(x+2,y-2) +MV0(x+2,y+2), 1); 12. MVCbt(x',y') = Shift( MV0(x+2,y-2) +MV0(x+2,y+2), 1); 13. MVCbt(x',y') = SatShift(MV0(x+2,y-2)+MV0(x-2,y+2), 1); 14. MV1(x,y) = Shift( MV0(x+2,y-2)+MV0(x-2,y+2), 1); 15. MVCbt(x',y') = MV0(x-2,y-2); 16. MVCbt(x',y') = MV0(x+2,y-2); 17. MVCbt(x',y') = MV0(x-2,y+2); 18. MVCbt(x',y') = MV0(x+2,y+2).

[0439] Interleaved forecasting for bidirectional forecasting 15. When interleaved prediction is applied to bidirectional prediction, the following methods can be applied to preserve the increased internal bit depth due to the different weights: a. For list X (X=0 or 1), PX(x, y) = Shift( W0(x, y) PX0(x,y) + W1(x,y) PX1(x,y), SW), where PX(x,y) is the prediction for list X, and PX0(x,y) and PX1(x,y) are the predictions for list X with partitioning mode 0 and partitioning mode 1. W0 and W1 are integers representing the weighted values ​​of the interleaved predictions, and SW represents the precision of the weighted values.

[0440] b. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB), where Wb0 and Wb1 are integers used in weighted bidirectional forecasting, and SWB is the precision. When there is no weighted bidirectional forecasting, Wb0=Wb1=SWB=1.

[0441] c. In some embodiments, PX0(x,y) and PX1(x,y) can maintain the accuracy of the interpolation filter. For example, they can be 16-bit unsigned integers. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y)+ Wb1(x,y) P1(x,y), SWB+PB), where PB is the additional precision from the interpolation filter, for example, PB = 6. In this case, W0(x,y) PX0(x,y) or W1(x,y) PX1(x,y) can exceed 16 bits. It is proposed that PX0(x,y) and PX1(x,y) are first right-shifted to a lower precision to avoid exceeding 16 bits.

[0442] i. For example, for a list X (X = 0 or 1), PX(x, y) = Shift(W0(x, y)). PLX0(x,y) + W1(x,y) PLX1(x,y), SW), where PLX0(x,y) = Shift( PX0(x,y), M), PLX1(x,y) = Shift( PX1(x,y), M). The final prediction is derived as P(x,y) = Shift( Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB+PB-M). For example, M is set to 2 or 3.

[0443] d. The above method can also be applied to other bidirectional forecasting methods with different weighting factors for two reference forecast blocks, such as generalized bidirectional forecasting (GBi, where the weights can be, for example, 3 / 8, 5 / 8) and weighted forecasting (where the weights can be very large values).

[0444] e. The above method can also be applied to other multi-hypothesis one-way or two-way forecasting methods that have different weighting factors for different reference forecast blocks.

[0445] Block size correlation 16. Whether and / or how to apply interleaving prediction can depend on the block width W and height H.

[0446] a. In one example, whether and / or how to apply interleaved prediction can depend on the size of the VPDU (Video Processing Data Unit, which typically represents the maximum permissible block size to be processed in a hardware design).

[0447] b. In one example, the original prediction method can be utilized when interleaving prediction is disabled for a specific block dimension (or a block with specific encoded / decoded information).

[0448] i. Alternatively, affine mode can be directly disabled for this type of block.

[0449] c. In one example, interleaved prediction cannot be used when W > T1 and H > T2. For example, T1 = T2 = 64; d. In one example, interleaved prediction cannot be used when W > T1 or H > T2. For example, T1 = T2 = 64; e. In one example, when W H > T, the interleaved prediction cannot be used. For example, T = 64 64; f. In one example, when W < T1 and H < T2, the interleaved prediction cannot be used. For example, T1 = T2 = 16; g. In one example, when W < T1 or H > T2, the interleaved prediction cannot be used. For example, T1 = T2 = 16; h. In one example, when W H < T, the interleaved prediction cannot be used. For example, T = 16 16.

[0450] i. In one example, for sub - blocks not located at block boundaries (e.g., codec units), the interleaved affine can be disabled for that sub - block. Alternatively, in addition, the prediction result using the original affine prediction method can be directly used as the final prediction for that sub - block.

[0451] j. In one example, when W > T1 and H > T2, the interleaved prediction is used in a different way. For example, T1 = T2 = 64; k. In one example, when W > T1 or H > T2, the interleaved prediction is used in a different way. For example, T1 = T2 = 64; l. In one example, when W H > T, the interleaved prediction is used in a different way. For example, T = 64 64; m. In one example, when W < T1 and H < T2, the interleaved prediction is used in a different way. For example, T1 = T2 = 16; n. In one example, when W < T1 or H > T2, the interleaved prediction is used in a different way. For example, T1 = T2 = 16; o. In one example, when W H < T, the interleaved prediction is used in a different way. For example, T = 16 16.

[0452] p. In one example, when H > X (e.g., H equals 128, X = 64), the interleaved prediction is not applied to the samples of the sub - blocks belonging to the upper W (H / 2) partition and the lower W (H / 2) partition of the current block.

[0453] q. In one example, when W > X (e.g., W equals 128, X = 64), the interleaved prediction is not applied to the samples of the sub - blocks belonging to the left (W / 2) H-segmentation and right (W / 2) Sample points of sub-blocks divided by H.

[0454] r. In one example, when W>X and H>Y (e.g., W=H=128, X=Y=64), i. Interleaving predictions are not applied to left (W / 2) segments that span the current block. H-segmentation and right (W / 2) Sample points of sub-blocks divided by H.

[0455] ii. Interleaving predictions are not applied to W-levels that span the current block. (H / 2) Segmentation and Lower W Sample points of sub-blocks divided by (H / 2).

[0456] s. In one example, interleaving prediction is enabled only for blocks with a specific set of widths and / or heights.

[0457] In one example, interleaving prediction is disabled only for blocks with a specific set of widths and / or heights.

[0458] u. In one example, interleaving prediction is only used for specific types of picture / strip / group / piece / or other kinds of video data units.

[0459] i. For example, interleaving prediction is only used for P-images or B-images.

[0460] ii. For example, a flag is transmitted via signal in the header of a picture / strip / piece / group of pictures to indicate whether interleaving prediction can be used.

[0461] 1. For example, the flag is transmitted via signal only when affine prediction is permitted.

[0462] 17. A message is proposed to be transmitted via signaling to indicate whether / how interleaving prediction is applied and its dependency on width and height. The message can be transmitted via signaling in SPS / VPS / PPS / strip header / picture header / slice / slice group header / CTU / CTU line / multiple CTUs / or other types of video processing units.

[0463] 18. In one example, bidirectional forecasting is not allowed when using interleaved forecasting.

[0464] a. For example, when using interleaved prediction, the index indicating whether bidirectional prediction is used is not transmitted via signaling.

[0465] b. Alternatively, an indication of whether bidirectional prediction is not allowed can be transmitted via signaling in SPS / VPS / PPS / strip header / picture header / film / film group header / CTU / CTU line / multiple CTUs.

[0466] 19. A method was proposed to further refine the motion information of sub-blocks based on motion information derived from two or more modes.

[0467] a. In one example, refined motion information can be used to predict subsequent blocks to be encoded or decoded.

[0468] b. In one example, refined motion information can be used in filtering processes such as deblocking, SAO, and ALF.

[0469] c. Whether to store refined information can be based on the position of the sub-block relative to the whole block / CTU / CTU line / slice / strip / slice group / image.

[0470] d. Whether to store refined information may be based on the encoded / decoded patterns of the current block and / or neighboring blocks.

[0471] e. Whether to store refined information can be based on the dimension of the current block.

[0472] f. Whether to store refined information can be based on image / strip type / reference image list, etc.

[0473] 20. It is proposed that whether and / or how to apply the deblocking process or other kinds of filtering processes (such as SAO, adaptive loop filter) can depend on whether interleaved prediction is applied.

[0474] a. In one example, if an edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then that edge is not deblocked.

[0475] b. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then the deblocking of that edge is weakened.

[0476] i. In one example, for such edges, bS[xDi][yDj] described in the VVC deblocking process decreases.

[0477] ii. In one example, for such edges, the β reduction described in the VVC deblocking process.

[0478] iii. In one example, for such edges, the Δ described in the VVC deblocking process decreases.

[0479] iv. In one example, for such edges, tC as described in the VVC deblocking process decreases.

[0480] c. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then the deblocking of that edge is strengthened.

[0481] i. In one example, for such an edge, bS[xDi][yDj] described in the VVC deblocking process increases.

[0482] ii. In one example, for such edges, the β increase described in the VVC deblocking process.

[0483] iii. In one example, for such edges, the Δ described in the VVC deblocking process increases.

[0484] iv. In one example, for such edges, the tC described in the VVC deblocking process is increased.

[0485] 21. It is proposed that whether and / or how local illumination compensation or weighted prediction is applied to blocks / subblocks may depend on whether interleaved prediction is applied.

[0486] a. In one example, when a block is encoded or decoded in interleaved prediction mode, local illumination compensation or weighted prediction is not allowed.

[0487] b. Alternatively, if interleaving prediction is applied to blocks / subblocks, there is no need to indicate the need to enable local lighting compensation via signal transmission.

[0488] 22. It was proposed that bidirectional optical flow (BIO) can be skipped when weighted prediction is applied to a block or sub-block.

[0489] a. In one example, BIO can be applied to blocks with weighted predictions.

[0490] b. In one example, BIO can be applied to blocks with weighted predictions; however, certain conditions must be met.

[0491] i. In one example, at least one parameter is required to be within a range or equal to a specific value.

[0492] ii. In one example, specific reference image constraints may be applied.

[0493] 2.9. Affine Control Point Refinement Based on Template Matching (TM-Affine) In this disclosure, template matching is proposed to refine the affine CPMV. For a given affine candidate in the affine candidate list, the CPMV can be further refined using template matching, and the refined affine candidate is then used to derive sub-block or pixel-level affine motion information for the current block.

[0494] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0495] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, ​​TU, PB, TB.

[0496] The term "affine block" can refer to a block encoded using affine Merge, affine AMVP, or any other affine variant mode (i.e., affine MMVD, etc.), which can be described by motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). The term "CPMV" can refer to the motion information of the affine block at the top left, top right, and / or bottom left corners.

[0497] The term "template" can refer to a reconstructed region that can be used to refine a CPMV, and it can refer to either a "single template" or a "uniform template." Here, a "single template" can refer to a reconstructed region that can be used to refine a single CPMV (i.e., one of the top-left, top-right, and / or bottom-left corners), while a "uniform template" can refer to a reconstructed region that can be used to refine all or any (multiple) CPMVs for a block. The term "template matching cost" or "TM cost" can refer to the matching cost of a single template or the matching cost of a uniform template.

[0498] In this disclosure, "blocks encoded and decoded in mode N" can refer to a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, merging, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).

[0499] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.

[0500] 1. In one example, affine motion compensation can be refined by using previously decoded samples.

[0501] a) In one example, at least one CPMV can be refined.

[0502] b) In one example, at least one MV of a sub-block of affine motion compensation can be refined.

[0503] c) In one example, at least one affine parameter (such as a, b, c, d, e, f (This can be further refined.)

[0504] d) In one example, previously decoded samples can be the template of the current block.

[0505] e) In one example, the previously decoded sample can be a template of the reference block.

[0506] f) In one example, the template represents the reconstructed region that can be used to refine the CPMV.

[0507] g) In one example, for a block that has been affinely encoded or decoded, different individual templates can be used for different control points.

[0508] i. In one example, for a control point, the corresponding individual template may include samples from adjacent and / or non-adjacent locations in the already reconstructed area.

[0509] 1) In one example, individual templates for all control points are collected from adjacent reconstruction regions.

[0510] 2) In one example, individual template samples for some control points are collected from adjacent reconstructed regions of the current block, while for the remaining control points, template samples are collected from non-adjacent reconstructed regions.

[0511] a) In one example, specifically, template samples for the top left corner are collected from non-adjacent regions, while template samples for the top right and / or bottom left corners are collected from adjacent regions.

[0512] 3) In one example, both adjacent and non-adjacent samples are used for a specific control point.

[0513] ii. In one example, the shape of the individual template can be different for different control points.

[0514] 1) In one example, for a specific control point, a separate template for an L-shape (e.g., including both the upper neighboring sample and the left neighboring sample) is used.

[0515] 2) In one example, for a specific control point, an I-shaped or "-" shaped template (e.g., including the left neighboring sample point or the upper neighboring sample point (but not both)) can be used.

[0516] iii. In one example, what shape of template is used for CPMV refinement can be based on the position / location of control points.

[0517] 1) In one example, the CPMV at the top left corner of the current video unit can use an L-shaped template (e.g., including both the top neighbor sample and the left neighbor sample).

[0518] 2) In one example, the CPMV at the top right corner of the current video unit can use a "-" shaped template (e.g., only including the upper neighboring sample).

[0519] 3) In one example, the CPMV at the top left corner of the current video unit can use an I-shaped template (e.g., only including the left neighboring sample).

[0520] 4) In one example, for a specific CPMV, the templates in the current image and the reference image have the same shape.

[0521] a) For example, such as Figure 16 As shown, the template for CPMV can refer to a set of neighboring samples in the current image (e.g., the template in the current image) and a second set of neighboring samples in the reference image (e.g., the template in the reference image).

[0522] iv. In one example, the number of sample points used in the template can be different for different control points.

[0523] 1) Alternative sites: For affine blocks, the number of samples used is the same for different control points.

[0524] 2) For different control points, the lines (or rows or columns) of the sample points used in the template can be different.

[0525] h) In one example, a uniform template is used during CPMV refinement.

[0526] i. In one example, the TM cost associated with the uniform template is used to determine the MV shift value.

[0527] ii. In one example, the TM cost associated with the uniform template is used to determine the CPMV combination.

[0528] iii. In one example, a uniform template can include all or some of the adjacent samples of the entire block, i.e., as shown below. Figure 16 As shown.

[0529] i) The template may include samples from only one component (such as luminance) or samples from multiple components (such as luminance and chrominance).

[0530] j) In one example, for any template, a reference template region with the same shape can be positioned using MV, such as Figure 16 As shown.

[0531] k) In one example, the template may not necessarily contain all the pixels in a specific area; it may contain only a portion of the pixels in the specified area.

[0532] 2. When constructing the affine candidate list, CPMV refinement can be performed first on the potential affine candidates, and then the refined candidates are inserted into the affine candidate list.

[0533] a) In one example, alternatively, CPMV refinement is performed after the affine candidate list has been constructed.

[0534] i. In one example, only the affine candidates with (multiple) specific indices need to be refined using CPMV.

[0535] ii. In one example, all or some affine candidates need to be refined using CPMV.

[0536] iii. In one example, a similarity check is performed first to determine whether the candidate needs to undergo CPMV refinement.

[0537] 1) In one example, the candidates in the list are traversed in a specific order. Only when the i-th candidate... C_i and{ C_ 1…,C_i-1 If any candidate in} is sufficiently similar to another, refinement may not be necessary.

[0538] a) In one example, the traversal order can be derived based on the TM cost.

[0539] b) In one example, specifically, a pair of candidates can be considered sufficiently similar if all CPMV differences are less than a threshold and / or the same prediction direction and / or reference frame are used.

[0540] c) In one example, { C_1…,C_i-1 The candidates in} can be refined using TM.

[0541] 3. In one example, the first affine candidate list is constructed first, followed by the construction of the second affine candidate list.

[0542] a) For example, the input for generating the second affine candidate list can be based on the output generated by the first affine candidate list.

[0543] b) For example, the first affine candidate list can be constructed without CPMV refinement.

[0544] c) For example, a second affine candidate list can be generated by applying CPMV refinement to the CPMV candidates in the first affine candidate list.

[0545] i. For example, at least one CPMV candidate in the first affine candidate list can be refined.

[0546] ii. Alternatively, more than one CPMV candidate in the affine candidate list can be refined.

[0547] iii. For example, CPMV refinement can be based on TM.

[0548] d) For example, the first affine candidate list can be constructed using a candidate reordering process.

[0549] i. For example, the reordering process can be based on TM.

[0550] e) For example, a second affine candidate list can be constructed without any candidate reordering process.

[0551] f) For example, different deduplication rules can be used in the first deduplication and the second deduplication.

[0552] i. For example, the generation of the first affine candidate list can be applied in conjunction with the first deduplication method.

[0553] ii. For example, the generation of the second affine candidate list can be applied in conjunction with the second deduplication method.

[0554] iii. For example, the thresholds for motion similarity checks in the first deduplication method and the second deduplication method can be different.

[0555] iv. For example, a threshold based on block dimensions (e.g., block width and / or height) can be used in a second deduplication method.

[0556] v. For example, alternatively, a fixed threshold can be used in a second deduplication method.

[0557] 4. For a given affine candidate, some or all CPMVs can be refined based on TM, and the refined CPMVs are then used to derive affine motion information for the current block and / or (multiple) sub-blocks.

[0558] a) In one example, both integer precision and fractional precision can be used to refine control points.

[0559] i. In one example, only integer precision was used to refine the control points, and fractional precision searches were skipped.

[0560] 1) In one example, whether a fractional precision search is needed depends on the result of an integer precision search.

[0561] ii. In one example, a specific interpolation filter is proposed to generate a reference template for the motion vector pointing to the fractional position.

[0562] 1) In one example, a simplified interpolation filter can be applied.

[0563] 2) In one example, the simplified interpolation filter can be a 2-tap bilinear filter, or alternatively, a 4-tap, 6-tap, or 8-tap filter belonging to DCT, DST, Lanczos, or any other interpolation type.

[0564] 3) In one example, a more complex interpolation filter (e.g., one with longer filter taps) can be applied.

[0565] iii. In one example, whether and / or how the above methods (e.g., integer precision, different interpolation filters) can be determined in the bitstream via signal transmission (such as in SPS, PPS, picture header, strip header, CTU, CU, etc.) or on the fly based on the decoded information.

[0566] 1) In one example, which method is applied can depend on the encoding / decoding tool.

[0567] 2) In one example, which method to be applied can depend on the block dimension.

[0568] b) In one example, different control points are refined separately, which means that the MV shift value (i.e., the difference between the initial CPMV and the corresponding refined CPMV) can be different for different control points.

[0569] i. In one example, all or some control points can first be refined by TM, and then a combination of control points is determined by iterating through all or some combinations of CPMV before and after refinement (i.e., M (e.g., M=4) combinations for a 4-parameter model and -N (e.g., N=8) combinations for a 6-parameter model), and a set of CPMVs that minimizes the TM cost of the current block is derived.

[0570] 1) In one example, under the above circumstances, all or some control points can be refined first through the corresponding individual templates.

[0571] 2) In one example, for each combination of CPMV, sub-block level motion information is calculated for the boundary sub-blocks, and then according to Section 2.4 and Figure 17 The method described in [the document] computes a uniform TM cost. The optimal combination that produces the minimum TM cost is selected as a refined affine candidate.

[0572] a) In one example, only some of the boundary sub-blocks need to have their TM costs calculated.

[0573] 3) In one example, alternatively, it is not necessary to traverse all combinations, and all control points are directly used as refined affine candidates through combinations refined by TM.

[0574] 4) In one example, a second pass of control point refinement can be performed to further refine each control point as the refined affine candidate is derived.

[0575] a) In one example, each CPMV is further iteratively refined to minimize the TM cost of the current block. In each iteration, one CPMV is refined while the others are fixed.

[0576] c) In one example, alternatively, multiple control points are refined simultaneously, where the same MV shift value is shared for all or more control points.

[0577] i. In one example, all or some of the MV shift values ​​in a given MV shift set are traversed one by one. The traversed MV shift values ​​are assigned to all or more CPMVs, and then the motion information of the boundary sub-blocks associated with the refined CPMVs is calculated, and the TM cost is expressed accordingly. In this process, the one that produces the minimum TM cost is determined as the optimal motion displacement value, which can ultimately be used to refine the CPMVs.

[0578] 1) In one example, specifically, multiple integer MV shift values ​​are traversed separately, and the one that produces the minimum TM cost is determined as the initial search point for the fractional shift value.

[0579] d) In one example, the refined affine candidate can replace the original affine candidate.

[0580] i. In one example, the refined affine candidate will always replace the original affine candidate.

[0581] ii. In one example, alternatively, the refined affine candidate will conditionally replace the original affine candidate.

[0582] 1) In one example, specifically, compared to the original CPMV (referred to as C_beforeTM ) and refined CPMV ( C_ afterTM The associated TM costs are calculated separately, and only when C_afterTM and C_beforeTM The ratio is less than (or greater than) a constant or an adaptively determined value. TH Only when the refined affine candidate is used will the original affine candidate be replaced.

[0583] a) In one example, under the above conditions, different codec modes (e.g., affine Merge / affine AMVP / affine MMVD) can have different... TH Value settings.

[0584] iii. Alternatively, refined affine candidates can be used as new candidates.

[0585] 1) In one example, the refined affine candidate can be placed adjacent to the original affine candidate in the affine candidate list (i.e., immediately before or after the original affine candidate).

[0586] 2) In one example, alternatively, the refined affine candidate can be placed anywhere in the affine candidate list.

[0587] 3) In one example, a refined affine candidate can be compared with at least one candidate already in the candidate list. If they are the same or similar, the candidate is not added to the list.

[0588] 5. CPMV refinement can be used in conjunction with regression-based affine candidate derivation methods.

[0589] a) In one example, after all or some of the CPMVs utilize TM, (resulting in...) Affine_model_ TM ),and Affine_model_TM The motion information of the associated boundary sub-blocks is derived, and then fed into the regression model to output a new affine model (called...). Affine_model_R Then calculate and utilize respectively. Affine_model_TM and Affine_model_R The TM costs of the boundary sub-blocks are compared. The one with the smaller TM cost is determined as the final refined affine candidate.

[0590] i. In one example, all or some of the CPMVs can first perform integer precision TM refinement (producing...) Affine_ model_TM_I Then perform fractional precision™ refinement (to produce...) Affine_model_TM_F And with Affine_model_ TM_I The motion information of the associated boundary sub-blocks is derived, and then fed into the regression model to output a new affine model. Affine_model_R Finally, calculate and compare the utilization. Affine_model_TM_F and Affine_model_R The TM cost of the boundary sub-blocks, and the one with the smaller TM cost is determined as the final refined affine candidate.

[0591] ii. In one example, only some sub-blocks may require computation of the TM cost to generate. Affine_model_TM , Affine_model_TM_I and / or Affine_model_TM_F .

[0592] 6. In one example, TM-based refinement can be applied to affine Merge or affine AMVP (affine inter-frame).

[0593] a) In one example, the MVP(s) that affine AMVP can be refined based on TM.

[0594] i. Alternatively, multiple MVPs that affine AMVP can be refined based on DMVR.

[0595] b) In one example, for affine AMVP (inter-frame) mode, whether TM-based refinement is performed may depend on the precision of MV or MVD.

[0596] i. In one example, specifically, TM-based refinement is only performed when a specific MV / MVD precision is used for the block.

[0597] c) In one example, different MV shift sets or search procedures can be used for affine Merge and affine AMVP (affine inter-frame).

[0598] i. In one example, specifically, different numbers of MV shift values ​​can be used for affine Merge and affine AMVP (affine inter-frame).

[0599] ii. In one example, for affine AMVP (affine inter-frame), if the initial CPMV and the CPMV with MV shift produce the same CPMV after rounding to a certain precision, then it is not necessary to calculate the TM cost for the specific MV shift.

[0600] 7. In one example, TM-based refinement can be applied together with DMVR-based refinement to affine-encoded blocks.

[0601] a) In one example, TM-based refinement can be applied before DMVR.

[0602] b) In one example, TM-based refinement can be applied after DMVR.

[0603] c) Alternatively, TM-based refinement can be applied exclusively to affine-coded blocks, just like DMVR-based refinement.

[0604] 8. In one example, the derivation of the TM cost can depend on whether the block is predicted bidirectionally or unidirectionally.

[0605] a) If the block is bidirectionally predicted, then the TM cost can be derived based on the bidirectional prediction on the TM.

[0606] i. In one example, make TM ref0 and TM ref1 If the references TM are associated with list 0 and list 1 respectively, then the final reference TM (TM) bi This can be derived as: TM bi = a * TM ref0 + (1 – a ) * TM ref1 .

[0607] 1) In one example, a It equals 0.5.

[0608] 2) In one example, a The BCW index was used to determine this.

[0609] 3) In one example, TM ref0 The CPMV is generated based on list 0, and / or TM. ref1 The CPMV is generated based on List 1.

[0610] b) Alternatively, if the block is bidirectionally predicted, the TM cost can be computed separately for list 0 and list 1.

[0611] 9. In one example, the refinement of CPMV can be done iteratively.

[0612] a) For example, in one step of refinement, one CPMV is refined while other CPMVs are fixed.

[0613] b) In one example, when a subsequent CPMV needs to be refined, (multiple) already refined CPMVs can be used.

[0614] i. In one example, alternatively, when a subsequent CPMV needs to be refined, the CPMV before refinement is used.

[0615] c) In one example, for bidirectional predictive blocks, the refinement of CPMV can be done iteratively.

[0616] i. In one example, the CPMV associated with list K (K = 0 or 1) can be refined first, and then the CPMV associated with list (1-K) can be refined.

[0617] 1) Whether and / or how the CPMV in the later list (1-K) can be refined can be determined based on the refined CPMV in the earlier list K.

[0618] ii. In one example, the CPMV associated with list 0 and list 1 can be refined separately.

[0619] 1) In one example, specifically, as the CPMV in list K (K = 0 or 1) is being refined, for each search step, a one-way reference TM in list K is generated based on the corresponding CPMV, and the TM cost is thereby calculated to determine the optimal MV shift value.

[0620] iii. In one example, alternatively, the CPMV associated with list 0 and list 1 can be jointly refined.

[0621] 1) In one example, specifically, as the CPMV in list K (K = 0 or 1) is being refined, for each search step, a bidirectional reference TM is generated based on the CPMV information of the two lists (as described in Item 8). The one that produces the minimum TM cost is determined as the optimal MV shift value.

[0622] 10. Multi-round refinement can be performed on CPMV.

[0623] a) In one example, all or part of the CPMV can be refined in each round of refinement.

[0624] b) In one example, all or part of the CPMV may have been refined in the front wheel refinement and then further refined in the rear wheel refinement.

[0625] 11. Whether and / or how CPMV refinement based on TM can be determined based on the prediction direction of the current block.

[0626] a) In one example, CPMV may only need to be refined by TM if the current block is unidirectionally predictable.

[0627] b) In one example, CPMV may only need to be refined by TM if the current block is bidirectionally predictable.

[0628] c) In one example, CPMV may always need to be refined via TM, regardless of whether the current block is bidirectionally predictable.

[0629] 12. For blocks that have been affine encoded or decoded, more than one CPMV refinement process can be cascaded.

[0630] a) For example, the affine DMVR process can be applied based on a refined CPMV.

[0631] i. For example, the first step in CPMV refinement for affine patterns (e.g., affine Merge and / or affine AMVP) can be an explicit (MMVD-based) CPMV refinement process and / or an implicit (e.g., TM-based) CPMV refinement process.

[0632] 1) For example, for explicit refinement, block-level syntax elements (such as MMVD indexes and / or MMVD steps and / or MMVD distances) can be transmitted via signals in the bitstream.

[0633] 2) For example, for implicit refinement, CPMV can be refined on the decoder side (e.g., based on TM) without block-level signaling.

[0634] ii. For example, the second step of CPMV refinement for affine patterns (e.g., affine Merge and / or affine AMVP) can be a DMVR-based refinement process (e.g., regression-based affine DMVR and / or affine DMVR with translational offset added to the CPMV), and the affine Merge DMVR process can be based on the refined CPMV obtained in the first step.

[0635] 1) In addition, for example, two different affine DMVR processes can be applied, such as one based on regression and the other not based on regression.

[0636] iii. For example, bilateral costs can be calculated and used to determine DMVR-based offsets.

[0637] b) Alternatively, the CPMV can first be refined based on DMVR, and then the DMVR-refined CPMV can be further refined through an explicit (MMVD-based) CPMV refinement process and / or an implicit (e.g., TM-based) CPMV refinement process.

[0638] 13. For example, pixel / sample-based affine prediction can be applied based on a refined CPMV.

[0639] a) For example, the refined CPMV can be derived based on explicit methods (e.g., based on MMVD) and / or TM-based refinement and / or DMVR-based refinement.

[0640] b) For example, the refined CPMV can be applied to one-way affine prediction and / or two-way affine prediction.

[0641] c) For example, pixel / sample affine prediction based on refined CPMV can be applied to affine Merge and / or affine AMVP modes.

[0642] 14. If affine prediction is used as an assumption, the disclosed method can be applied to blocks encoded and decoded by MHP (Multiple Hypothesis Prediction).

[0643] 15. An affine candidate list can contain two alternative versions of an affine candidate at the same time (e.g., a refined version and a non-refined version).

[0644] a) In one example, the original candidate and the TM-refined (or bilaterally matched) version of the same candidate can appear in the affine list at the same time.

[0645] i. In one example, for the same candidate, whether two alternative versions appear in the affine list can depend on the prediction direction of the candidate (e.g., bidirectional or unidirectional prediction).

[0646] ii. In one example, after the initial affine list is constructed, candidates in the list can be examined in a specific order (e.g., TM cost). If the current candidate satisfies a specific condition (e.g., one-way prediction), an alternative version of the candidate is used to append to the list.

[0647] 1) In one example, if the current candidate has been refined, the alternative version of the candidate (i.e., the non-refined version) can be used to append to the list, and vice versa.

[0648] 2) In one example, the alternative version can be used to replace a specific candidate or any candidate.

[0649] a) In one example, an alternative version could be used to replace a version with a specific index. A Candidates or in a specific index A The subsequent candidates.

[0650] i. In one example, A It can be a constant value.

[0651] ii. In one example, the first candidate index with all-zero CPMV was determined as A .

[0652] iii. In one example, A The number of candidates can depend on the specific type of affine candidate.

[0653] iii. In one example, the final affine candidate index is based on the initial affine index. I and variables B It has been confirmed.

[0654] 1) In one example, the initial affine index I It can be parsed in a bitstream. . 2) In one example, B It can be any value ranging from 0 to the maximum allowed number of affine candidates.

[0655] 3) In one example, the index of the first candidate in the affine list with all zero CPMV can be determined as B .

[0656] 4) In one example, if I If it is less than B, then I Used to specify affine candidates.

[0657] a) Alternative locations, if I Greater than or equal to B The adjusted index is then deduced and used to specify the affine candidate.

[0658] i. In one example, the adjusted index could be ( I – B ).

[0659] ii. In one example, the adjusted index may depend on I , B And / or the candidate prediction direction.

[0660] iii. In one example, the alternative version of the candidate specified by the adjusted index can be used as the final affine candidate.

[0661] 1. In one example, if the candidate specified by the adjusted index has been refined, the unrefined version will be used to generate affine predictions, and vice versa.

[0662] 16. Whether and / or how the methods disclosed above are applied can be determined based on (multiple) syntax elements.

[0663] a) For example, at least one syntax element is transmitted via signal in the bit stream.

[0664] b) For example, whether and / or how the disclosed methods can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0665] c) For example, whether and / or how the disclosed methods can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.

[0666] d) For example, whether and / or how to apply the methods disclosed above may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.

[0667] e) For example, whether a syntax element (i.e., an indicator of whether TM refinement is applied to CPMV) is transmitted via signaling can be determined based on another syntax element.

[0668] 2.10. An extension to the construction of a motion vector prediction list based on template matching cost ranking This disclosure proposes an optimized MVP list derivation method based on template matching cost ranking. Instead of constructing the MVP list based on a predefined traversal order, an optimized MVP selection method is investigated by utilizing the matching cost in the reconstructed template region, allowing more suitable candidates to be included in the list.

[0669] This disclosure also proposes to improve the inter-frame encoding / decoding process by introducing more TMVPs, where the TMVPs are generated based on motion information from multiple co-located frames. Furthermore, we improve the inter-frame encoding / decoding tool by introducing template matching. It should be noted that the proposed strategy can be applied to any encoding / decoding tool that requires temporal motion information, including but not limited to regular Merge and AMVP list construction, Merge with Motion Vector Difference (MMVD), affine motion compensation, sub-block-based temporal motion vector prediction (SbTMVP), adaptive DMVR, etc. It should also be noted that the proposed strategy for TMVP derivation can be utilized in any encoding / decoding tool that requires an MVP list construction process, including but not limited to regular Merge and AMVP list construction, Merge with Motion Vector Difference (MMVD), affine motion compensation, sub-block-based temporal motion vector prediction (SbTMVP), adaptive DMVR, etc.

[0670] It should be noted that the proposed strategy for MVP list construction can be used in the regular Merge and AMVP list construction process, and can also be easily extended to other modules that require MVP derivation, such as Merge with motion vector difference (MMVD), affine motion compensation, and sub-block-based temporal motion vector prediction (SbTMVP).

[0671] In the following discussion, a category indicates the affiliation of MVP candidates; for example, non-adjacent MVP candidates belong to one category, and HMVP candidates belong to another. A group represents a set of MVP candidates, which contains one or more MVP candidates. In one example, a single group represents a set of MVP candidates where all candidates in the set belong to one category, such as adjacent MVP, non-adjacent MVP, HMVP, etc. In another example, a joint group represents a set of MVP candidates, which contains candidates from multiple categories. A list can be an MVP candidate list, a TMVP candidate list, a motion displacement candidate list, or a sub-CU level MVP candidate list, where an MVP candidate list represents a set of MVP candidates that can be selected as MVPs during video encoding and decoding. A TMVP candidate list represents a set of TMVPs, where each candidate within the group has the potential to be selected as a candidate in the MVP candidate list. A motion displacement candidate list represents a set of MV candidates pointing to co-located frames during video encoding and decoding. A sub-CU level MVP candidate list represents a set of motion candidates that provide sub-CU level motion fields, including SbTMVP candidates, AFFINE candidates, etc.

[0672] The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way. Combinations between this patent application and other patent applications are also applicable.

[0673] 1. Multiple thresholds for determining whether a candidate can be added to the candidate list can be used in candidate deduplication.

[0674] a) Thresholds can be used to determine whether a potential candidate can be added to the candidate list.

[0675] i. For example, if the absolute difference between at least one component of a potential candidate's MV and the corresponding component of a candidate's MV already in the candidate list is less than a threshold, then the potential candidate is not added to the list.

[0676] ii. For example, if the absolute difference between all components of a potential candidate's MV and the corresponding components of a candidate's MV already in the candidate list is less than a threshold, then the potential candidate is not added to the list.

[0677] b) In one example, the candidate is an MVP candidate, the candidate deduplication process is an MVP candidate deduplication process, and the candidate list is a motion candidate list.

[0678] i. In one example, the motion candidate list is the Merge candidate list.

[0679] ii. In one example, the motion candidate list is the AMVP candidate list.

[0680] iii. In one example, the motion candidate list is an extended Merge or AMVP list, such as a sub-block Merge candidate list, an affine Merge candidate list, an MMVD list, a GPM list, a template matching Merge list, a bilateral matching Merge list, etc.

[0681] c) In one example, the deduplication threshold can be different for the two groups, where the group can be a single group (containing only one category of candidates) or a joint group (containing at least two categories of candidates).

[0682] d) Alternatively, for all potential MVP candidates, use only one threshold, regardless of category and / or group.

[0683] e) In one example, N (e.g., N=2) thresholds are used in the deduplication process.

[0684] i. Assumption A It is an MVP set, which contains all available MVP candidates regardless of category. In one example, a first threshold is used for the set.A The first subset of candidates in the set, and the second threshold is used for the set. A The second subset of candidates (e.g., the remaining candidates other than those in the first subset).

[0685] ii. In one example, the first threshold is used for a single group represented by A, and the second threshold is used for another group (single or combined) / multiple other groups / the remaining candidates that do not have the same category as those candidates in A.

[0686] 1) In one example, the first threshold is used for a single group of adjacent candidates, and the second threshold is used for the remaining candidates, including but not limited to non-adjacent MVPs, HMVPs, paired MVPs, and zero MVPs.

[0687] iii. The first threshold can be greater than or less than the second threshold.

[0688] f) Alternatively, additionally, the threshold for an MVP category or group may depend on the decoded information, such as block dimension / encoding / decoding method (e.g., CIIP / MMVD) and / or the variance of motion information within the category or group.

[0689] 2. Multiple reorderings can be performed to build the MVP list.

[0690] a) In one example, multiple iterations can involve different reordering criteria.

[0691] b) In one example, multiple reorderings can be performed on multiple single / joint groups, where at least two single / joint groups may have overlapping MVP candidates or non-overlapping MVP candidates.

[0692] c) In one example, K-times (e.g., K=2) reordering is used to build the MVP list.

[0693] i. In one example, in the first pass, single / joint groups A First, the order is reordered based on the first cost (e.g., template matching cost), and A Candidates with the highest cost (CL) are identified and then moved to another single / joint group. B (For example, B may include the remaining candidates that do not belong to the same category as those candidates in A). Subsequently, the group... B The second through Kth iterations of reordering are performed based on the first cost (or other cost metric) sort. Finally, the group... A (Except for CL) and groups B Candidates (including CL) are included in the MVP list according to their sorting order.

[0694] ii. In one example, group A in the above case is a single group of adjacent candidates, and group B is a combined group of non-adjacent candidates and HMVP.

[0695] iii. Alternatively, Group A and Group B can be any other single candidate group or joint candidate group.

[0696] iv. In one example, in the first pass, one or more individual / joint groups are first sorted and reordered based on a first cost (e.g., template matching cost). Then, an initial MVP list is constructed by inserting some candidates from each group into a list in sorted order. Subsequently, the initial MVP list undergoes a second pass of reordering to select a subset of candidates for the final MVP list.

[0697] 1) In one example, different single / joint groups may have overlapping candidates or non-overlapping candidates.

[0698] 2) In one example, all candidates in the initial MVP list are selected from sorted single / joint groups.

[0699] 3) Alternatively, some candidates in the initial MVP list are selected from the sorted groups, and the remaining candidates are included in the list according to other rules.

[0700] 4) In one example, in the second pass, all candidates in the initial list, regardless of their corresponding category, are sorted based on cost (e.g., template matching cost), and based on the sorting order, only a limited number of candidates are included in the final MVP list.

[0701] a) All candidates in the preliminary MVP list, including alternatives, additional candidates, are included in the final MVP list according to their sorting order.

[0702] 5) Costs calculated in previous iterations (e.g., template matching costs) can be reused in later iterations.

[0703] a) In one example, when the cost for a particular candidate is computed in a previous iteration, it will be stored in a variable or any other data structure in case the same cost is needed in a later iteration.

[0704] b) In one example, in a later iteration, if the cost for a particular candidate is needed, it will first check if that cost has already been computed. If the cost has already been computed and / or saved, and / or is accessible in the current iteration, it will be retrieved in the current iteration instead of being computed again.

[0705] 3. At least one virtual candidate (e.g., pairwise MVP and zero MVP) may be involved in at least one group.

[0706] a) In one example, all virtual candidates are treated as a joint group.

[0707] i. Alternatively, each category of the virtual candidate is treated as a single group.

[0708] ii. In one example, paired MVPs and / or zero MVPs are included in a single / joint group.

[0709] iii. Alternatively, additionally, groups containing dummy candidates are reordered and then placed into the candidate list.

[0710] b) Alternatively, virtual candidates (e.g., paired MVPs and / or zero MVPs) shall not be included in any single / joint group.

[0711] i. Alternative, additional, no reordering process is applied to the virtual candidate.

[0712] 1) Alternative sites, additional sites, which can be further added to the candidate list.

[0713] ii. In one example, one or more single / joint groups are constructed, where some or all of the groups are reordered. In this case, at least one position in the MVP list is reserved for a dummy candidate (e.g., a pair of MVPs and / or a zero MVP), which is appended to the MVP list as the last entry or any other entry.

[0714] iii. In one example, additionally, individual groups of adjacent candidates are first included in the MVP list, and then combined groups of non-adjacent and HMVP candidates are reordered and subsequently appended to the MVP list. In this case, at least one position is reserved for dummy candidates (e.g., paired MVPs and / or zero MVPs), which are appended to the MVP list as the last entry or any other entry.

[0715] iv. In one example, additionally, the combined groups of adjacent candidates, non-adjacent candidates, and HMVPs are reordered and subsequently appended to the MVP list, and dummy candidates (e.g., paired MVPs and / or zero MVPs) are appended to the MVP list as the last entry or any other entry.

[0716] c) Alternatively, virtual candidates of one category (e.g., paired MVPs) may be included in a single / joint group, while virtual candidates of another category may not be included.

[0717] d) In one example, when a reordering operation is performed on the MVP list build, no dummy candidates (e.g., paired MVPs and / or zero MVPs) appear in the final MVP list.

[0718] 4. The number of candidates in a single / joint group may not exceed the maximum number of candidates.

[0719] a) In one example, single / joint groups utilize the maximum number of N i A finite number of candidates constrained by the conditions are constructed, where i ∈[0,1,…,K] is the index of the corresponding group. For different... i , N i They can be the same or they can be different.

[0720] b) In one example, the number of partial candidates in a single / joint group is limited to a maximum. N i .

[0721] i. In one example, candidates for one or more categories within a group utilize a limited number of... N i It is constructed, and other categories in the same group can be included in any number.

[0722] 1) In one example, the categories include, but are not limited to, adjacent candidates, non-adjacent candidates, HMVP, paired candidates, etc.

[0723] c) Alternative sites, the first single / joint group can be in maximum order. N i One MVP candidate is constructed, while the second single / joint group may not have such a constraint.

[0724] d) In one example, N i It is a fixed value shared by the encoder and decoder.

[0725] i. Alternative locations N i Determined by the encoder and transmitted as a signal in the bitstream. And decoded by the decoder. N i Value, and then construct the corresponding first i A single / joint group, which has the most N i One candidate.

[0726] ii. Alternative sites, N iThe same operation is derived in both the encoder and decoder, eliminating the need for signal transmission. N i value.

[0727] 1) In one example, the encoder and decoder can be based on the first... i The variance of all available motion information of the group is used to derive N i value.

[0728] 2) Alternatively, the encoder and decoder can be based on the first... i The number of all available candidates in the group is used to derive the result. N i value.

[0729] 3) In one example, the encoder and decoder can be derived based on the number of available neighboring candidates. N i value.

[0730] a) In one example, N i Set as N – NADJ ,in N It is a constant. N ADJ It is the number of available adjacent candidates.

[0731] 4) Alternate sites, additional sites, encoders, and decoders can be derived based on any information that the encoder / decoder can access when building the MVP list. N i value.

[0732] e) In one example, all or part of a single / joint group can share the same maximum number of candidates. N .

[0733] 5. The construction of single / joint groups can depend on the maximum number of constraints. N i .

[0734] a) In one example, for the first i All available MVP candidates for a group are included in the group in a specific order. Once the number of candidates in the current group reaches [a certain threshold], [the process continues]. N i Then for the group i The construction was terminated.

[0735] b) In one example, under the above conditions, the order of group construction can be deduced based on the distance between the CU to be encoded / decoded and the MVP candidate, where the closer MVP candidate is assigned a higher priority.

[0736] c) Alternatively, the order can be derived based on cost (such as template matching), where the MVP with the lower cost has a higher priority.

[0737] d) In one example, the construction of a single / joint group utilizes at least one deduplication operation performed within or between at least one group.

[0738] e) In one example, the constructed single / joint groups are further reordered based on at least one cost method (e.g., template matching cost), and then some or all of the candidates in the group can be included in the MVP list.

[0739] i. Alternatively, the candidates in the constructed single / joint groups will not be further reordered, and some or all of the candidates in the group will be included in the MVP list in the same order as they were included in the group.

[0740] 6. Regarding how to deduplicate MVP candidates.

[0741] a) In one example, K iterations (e.g., K=2) of deduplication are performed to build the MVP list.

[0742] 1) In one example, the first deduplication can be performed within at least one single / joint group, and the second deduplication can be performed between at least two candidates belonging to different groups.

[0743] a) In one example, in the first deduplication pass, the deduplication thresholds for the two single / joint groups can be the same or different.

[0744] b) In one example, additionally, in the first pass of deduplication, some single / joint groups can share the same threshold, while other single / joint groups can use different thresholds.

[0745] 2) In one example, additionally, the threshold for a particular pass or group is determined by decoding information, including but not limited to block size and the encoding / decoding tools used (e.g., TM, DMVR, Adaptive DMVR, CIIP, AFFINE, AMVP-Merge).

[0746] a) Alternatively, the threshold can be determined by at least one syntax element transmitted to the decoder via a signal.

[0747] 7. A proposal was made to introduce [the following] during the video encoding and decoding process. K(For example, K =2) co-position frames.

[0748] a) In one example, motion vectors stored in at least one of the K co-frames can be used to encode / decode the current frame.

[0749] b) In one example, these co-located frames can be any reconstructed frame in the Decoded Picture Buffer (DPB).

[0750] c) In one example, these co-located frames can be any reconstructed frame from any reference list.

[0751] i. In one example, if the frame to be encoded or decoded currently has one or more reference lists, then the co-frame can be selected from one or more lists.

[0752] 1) In one example, if each list has an index N (For example, N If the reference frames (=0) are not the same reference frame (e.g., have different POC values), then these reference frames are selected as co-position frames.

[0753] a) In another example, after performing a redundancy check, each list has a previous... N The reference frame for each index is selected as the co-position frame.

[0754] b) Alternatively, any reference frame from any list may be selected as a co-position frame.

[0755] 2) In one example, if the frame to be encoded or decoded has one or more reference lists, the selected co-position frame may come from only one reference list.

[0756] a) In one example, if the contents of each reference list are the same (e.g., in a low-latency case), the selected co-frames may come from only one reference list.

[0757] 8. Regarding how to select co-located frames. S This represents a set containing all available reconstructed frames from the DPB or reference list, and SK yes S Any candidate in the given list. Then: a) SK Whether a frame can be selected as a co-location frame depends on the frame to be encoded and decoded. SK The distance between POCs.

[0758] i. In one example, S All candidates are sorted based on the Proof-of-Concept (POC) distance between the frame to be encoded / decoded and each candidate, and then the candidates with the smallest distance are ranked.N ( N >0) candidates were selected as the same frame.

[0759] ii. In one example, only when SK The distance between the frame to be encoded or decoded is less than or greater than a threshold. T ( T When >0), SK Only then can it be selected as a co-position frame.

[0760] b) SK Whether a frame can be selected as a co-position frame depends on the quantization parameter (QP) value. 。

[0761] i. In one example, S All candidates are sorted based on their QP values, and then the top candidates with the smallest or largest QP are ranked. N ( N >0) candidates were selected as the same frame.

[0762] ii. In one example, S All candidates are sorted based on the absolute QP distance between the frame to be encoded / decoded and each candidate, and then the candidates with the smallest distance are ranked. N ( N >0) candidates were selected as the same frame.

[0763] iii. In one example, only when SK The absolute QP difference between the frame to be encoded and decoded is less than or greater than a threshold. T ( T When >0), SK Only then can it be selected as a co-position frame.

[0764] c) SK Whether a frame can be selected as a co-located frame depends on the frame type.

[0765] i. In one example, if SK If it is an I-frame or a P-frame, it cannot be selected as a co-frame.

[0766] 1) Alternative locations, even SK It can be an I-frame or a P-frame, or it can be selected as a co-located frame.

[0767] d) SK Whether a frame can be selected as a co-location frame can depend on the temporal layer or Tid of the frame to be encoded or decoded.

[0768] i. In one example, if the Tid or temporal layer of the frame to be encoded or decoded is less than (or greater than, or equal to) a threshold. T ( T If >0), then SK It cannot be selected as a co-frame.

[0769] ii. In one example, the maximum number of co-location frames that can be used may depend on the Tid or temporal layer of the frame to be encoded or decoded.

[0770] 1) In one example, if the Tid or temporal layer of the frame to be encoded or decoded is less than (or equal to) the threshold. T ( T If >0), then at most N ( N >0) co-position frames can be used.

[0771] 2) In one example, if the Tid or temporal layer of the frame to be encoded or decoded is greater than (or equal to) the threshold. T ( T If >0), then at most M ( M >0) co-position frames can be used.

[0772] 3) In one example, the above example M and N They can be the same value or different values.

[0773] e) In one example, multiple metrics are combined to determine which one is used as the co-frame.

[0774] i. In one example, if the reference list or DPB is... N ( N >1) If two frames have an equal POC distance relative to the frame to be encoded or decoded, then those frames with a larger (or smaller) QP (or an absolute QP distance relative to the frame to be encoded or decoded) have a higher priority and are selected as co-position frames.

[0775] ii. In one example, if the reference list or DPB contains N ( N >1) If two frames have the same QP, then those frames with a larger (or smaller) POC (or absolute QP distance relative to the frame to be encoded or decoded) have a higher priority and are selected as co-position frames.

[0776] iii. In one example, alternatively, if in the reference list or DPB N ( N >1) If two frames have an equal absolute QP distance relative to the frame to be encoded or decoded, then those frames with a smaller (or larger) QP (or POC distance relative to the frame to be encoded or decoded) have a higher priority and are selected as co-position frames.

[0777] 9. (Multiple) selected co-occurring frames may be transmitted in the bitstream via signaling, including but not limited to strip headers or SPS or PPS or picture parameter headers.

[0778] a) Alternatively, both the encoder and decoder derive co-location frames based on predefined rules, so that no additional information needs to be transmitted.

[0779] b) In one example, some co-occurring frames need to be transmitted via signaling by syntax elements, while other co-occurring frames are deduced based on predefined rules, so that no additional information needs to be transmitted.

[0780] c) In one example, this information indicates which list the co-frame comes from (e.g., whether it comes from list 0), and the corresponding reference index is transmitted via signal in the syntax element.

[0781] i. In one example, specifically, if the number of reference frames in any reference list is zero, then the information indicating the reference list does not need to be transmitted via signaling.

[0782] ii. In one example, specifically, if only one frame exists in the corresponding list, the reference index may not need to be transmitted via signaling.

[0783] d) In one example, the number of (multiple) co-frames (denoted as N) can be encoded and decoded in the bitstream.

[0784] e) In one example, after a number N is transmitted via signaling, an indication of N co-located frames can be transmitted via signaling.

[0785] f) In one example, co-location frames may be indicated by a reference list and / or a reference index.

[0786] g) In one example, the signaling of the first co-frame may depend on the second co-frame that was previously transmitted via signaling.

[0787] h) More than one co-located frame can be co-encoded and decoded.

[0788] i) Syntax elements used for transmitting (multiple) co-occurring frames via signaling can be binary-coded using fixed-length encoding / decoding, unary encoding / decoding, rounded unary encoding / decoding, exponential Columbus encoding / decoding, or any other encoding / decoding method.

[0789] j) In one example, information associated with (multiple) co-frames can only be transmitted via signaling when TMVP is enabled.

[0790] 10. Multiple syntax elements can be transmitted via signals in a bitstream to identify multiple pariframes, where each syntax element can specify a different pariframe.

[0791] a) In one example, when a new candidate is being examined, it first needs to check if the same frame (e.g., one with the same POC number) has already been selected previously. If no such frame has been selected previously, it can be selected as a peer frame if it meets certain conditions.

[0792] b) Alternatively, two or more syntax elements may identify the same pariframe.

[0793] i. In one example, if at least K ( k >=0) co-position frames need to be selected, and the number of available distinct frames is less than K In this case, redundant frames can be used as co-position frames.

[0794] ii. In one example, the same co-position frame can be identified by different syntax elements, and these different syntax elements can be in different lists of reference frames.

[0795] iii. In one example, the same co-frame can be identified by different syntax elements, and these different syntax elements can also provide different temporal information if the co-frame is used more than once.

[0796] 1) In one example, when the same co-located frame is used more than twice, MVs from different reference frame lists can be used.

[0797] a) In one example, the MV in list 0 or list 1 is used.

[0798] b) In one example, the MVs in the two lists are used in a combined manner.

[0799] 11. The number of co-location frames used can depend on the codec configuration.

[0800] a) In one example, the number of co-frames used in different codec configurations can be different.

[0801] i. In one example, the codec configuration may include random access (RA), low latency B (LDB), or low latency P (LDP) or any other configuration.

[0802] ii. In one example, the maximum allowed number of co-op frames used in the RA configuration may be greater than (or less than, or equal to) the maximum allowed number of co-op frames in the LDB or LDP configuration.

[0803] 1) In one example, at most K (For example K=2) co-located frames are used in the RA configuration, and at most M (For example M =1) co-location frames are used in the LDB / LDP configuration.

[0804] 2) In one example, the above example M and K They can be the same value or different values.

[0805] 12. The number of co-location frames used may depend on the state of the reference frame list.

[0806] a) In one example, the number of co-frames used during encoding and decoding can depend on whether the POC values ​​of all reference frames are less than (or greater than) the POC value of the frame to be encoded or decoded.

[0807] i. In one example, if all reference frames have a POC value that is smaller (or larger) than the POC value of the frame to be encoded / decoded, then at most M (For example M =1) co-position frames can be used.

[0808] ii. In one example, if some reference frames have a smaller POC value compared to the POC value of the frame to be encoded / decoded, and some other reference frames have a larger POC value compared to the POC value of the frame to be encoded / decoded, then at most K (For example K =2) co-position frames can be used.

[0809] iii. In the example above, M and K They can be the same value or different values.

[0810] b) In one example, whether a reference frame can be used as a co-location frame may depend on its point of origin (POC) distance to the frame to be encoded or decoded.

[0811] i. In one example, if the POC distance between the reference frame and the frame to be encoded / decoded is greater than (or less than, or equal to) a threshold. T If so, the reference frame may not be used as a co-position frame.

[0812] 1) In one example, in the example above, T It can be a constant or an adaptively determined value.

[0813] ii. In one example, for each reference frame in the reference frame list, the POC distance between it and the current frame is calculated, and if the minimum POC distance value is less than a threshold... T If no corresponding frame is used for the current frame, then no corresponding frame is used for the current frame.

[0814] 1) In one example, in the example above, T It can be a constant or an adaptively determined value.

[0815] 2) In one example, specifically, in this case, there is no TMVP (or SbTMVP / time-domain AFFINE control point) in the corresponding candidate list.

[0816] 13. In one example, the determination of (multiple) co-op frames, such as the number of co-op frames and whether a reference frame is a co-op frame, may depend on the encoding / decoding information of at least one reference frame.

[0817] a) In one example, if the reference frame is an I-frame, it cannot be used as a co-frame.

[0818] b) In one example, if the number of intra-frame codec blocks in the reference frame is greater than a threshold, the reference frame cannot be used as a co-frame.

[0819] 14. The number of co-op frames and / or which (which) reference frames are used (as co-op frames) can depend on the specific characteristics of the codec block.

[0820] a) In one example, for any two codec blocks belonging to the same frame / strip / slice, the co-frames used may be the same or different.

[0821] i. In one example, for any two codec blocks belonging to the same frame / strip / slice, the maximum allowed number of co-occurring frames that can be used can be the same or different.

[0822] ii. In one example, for any two codec blocks belonging to the same frame / strip / slice, the number of co-op frames and / or which (which) reference frames are used (as co-op frames) can be the same or different.

[0823] b) In one example, for a given codec block, the number of co-op frames and / or which (which) reference frames are used (as co-op frames) can depend on the characteristics of the codec block.

[0824] i. In one example, the feature could be: 1) Block size.

[0825] 2) Qp.

[0826] 3) Forecast information (e.g., one-way or two-way forecasts).

[0827] c) In one example, for a particular codec block, how many co-op frames and / or which (which) reference frames are used (as co-op frames) can depend on the codec tools used.

[0828] i. In one example, the encoding / decoding tools may include, but are not limited to: 1) Regular / TM / CIIP / MMVD / GPM / TPM / Subblock Merge mode.

[0829] 2) AMVP.

[0830] 3) AMVP-Merge.

[0831] 4) AFFINE.

[0832] 5) BDOF.

[0833] 6) LIC.

[0834] 15. A proposal was made to introduce [something] during the video encoding and decoding process. K ( k There are >=1) TMVPs, which can be located in one or more co-located frames.

[0835] a) In one example, at most C ( C >=0) TMVPs are inserted into the MVP / TMVP candidate list, regardless of whether they come from one or more co-located frames, where C It is a constant or an adaptively determined number.

[0836] b) In one example, the maximum allowed number of TMVPs in a given co-frame in the MVP / TMVP candidate list is constrained by a constant or adaptively determined number.

[0837] c) In one example, the encoder traverses all co-located frames in a predefined or adaptively determined order to obtain the total C ( C >=0) TMVPs, and for each co-position frame, at most obtain D ( D >=0) TMVPs, where from one co-frame to another co-frame D It can change. When the total number of TMVPs reaches C The traversal process terminates when all corresponding frames have been traversed.

[0838] d) In one example, the number of TMVPs to be used in the list can be transmitted via signals in the bitstream, such as in SPS / PPS / Picture Header / Strip Header / etc.

[0839] 16. Different co-frames can be assigned different priorities.

[0840] a) In one example, the priority of co-op frames is determined based on their corresponding QP values, with those co-op frames having larger QP values ​​being assigned higher priority.

[0841] i. Alternatively, co-frames with smaller QP are assigned higher priority.

[0842] b) In one example, the priority of co-located frames is determined based on their temporal distance relative to the current frame, with those co-located frames having a smaller distance being assigned a higher priority.

[0843] i. Alternatively, co-frames with greater distances are assigned higher priority.

[0844] c) In one example, the priority of co-position frames is determined based on their index in the corresponding reference list.

[0845] i. In one example, a reference frame with a smaller index has a higher priority.

[0846] d) In one example, this priority is associated with the TMVP construction process and any other process in video encoding and decoding.

[0847] i. In one example, all co-located frames are traversed in descending priority order. When any co-located frame... Fi When selected, K ( K (0) positions will be traversed to include the TMVP in the TMVP / MVP candidate list, where K The same can be applied to each co-occurring frame, or the same can be applied between co-occurring frames. This applies once the maximum allowed number of TMVPs is reached. M ( M If the frame priority is >=0, the iteration terminates, and the corresponding frame with lower priority is skipped. Otherwise, all corresponding frames will be traversed to construct the TMVP / MVP candidate list.

[0848] ii. In another example, K ( K The positions >=0 are checked individually in a specific order to obtain TMVP candidates. For the positions checked... i Each position is traversed in descending order of priority within all corresponding frames, and available TMVPs are included in the list. Once the number of existing TMVP candidates reaches the maximum allowed number... M ( M If the value is greater than or equal to 0, the iteration terminates, and the remaining positions and corresponding frames with lower priority are skipped. Otherwise, all positions and corresponding frames will be traversed to construct a TMVP / MVP candidate list.

[0849] 1) Alternative locations: Under the above circumstances, for a specific TMVP position, at most... N ( N (>=0) TMVPs are included in the list.

[0850] a) In one example, for a specific TMVP location, if a higher priority co-frame... N ( N If (>=0) TMVPs have already been included in the list, then the same position in a low-priority co-frame will not be checked.

[0851] b) Alternative locations, for a specific TMVP position, any N ( N (0) TMVPs can be included in the list, regardless of priority.

[0852] e) In one example, a lower-priority co-frame is used as a backup, which is activated only if the TMVP or any other information in the higher-priority co-frame is absent.

[0853] i. In one example, for a specific position in different co-frames, if the required information (e.g., TMVP) is available in a higher-priority co-frame, that information is used during encoding / decoding, and the checking process for subsequent co-frames is skipped. Otherwise, the same position in a lower-priority co-frame is checked, and if corresponding information exists, that corresponding information is used.

[0854] ii. Alternatively, in the above case, even if information from a higher-priority co-frame exists, information from a lower-priority co-frame is also used.

[0855] f) Alternatively, different co-occurring frames are assigned equal priority.

[0856] i. In one example, the TMVP candidate set is constructed to include all or part of the potential TMVPs that can be located in any co-located frame, and then it is sorted in a specific order (e.g., template matching cost), and the top of the sorted list... N ( N (0 or more) candidates will be selected as the final TMVP.

[0857] g) In one example, different co-position frames can come from different lists of reference frames.

[0858] i. In one example, high-priority co-frames can only be selected from reference list 0 (or list 1).

[0859] 1) In one example, alternatively, higher priority co-frames can only be selected from reference list 1 (or list 0).

[0860] ii. In one example, whether a low-priority sibling frame is selected from list 0 or list 1 can depend on the high-priority sibling frame.

[0861] 1) In one example, if a higher priority co-frame comes from list 0, a lower priority co-frame can come from either list 0 or list 1.

[0862] 2) In one example, if a higher priority co-frame comes from list 1, then a lower priority co-frame can only come from list 1.

[0863] 3) In one example, alternatively, if a higher priority co-frame comes from list 1, a lower priority co-frame may come from list 0 or list 1.

[0864] 17. The proposed co-location frames can be used in any encoding / decoding tool during the video encoding / decoding process, including but not limited to regular / CIIP / MMVD / GPM / TPM / subblock Merge, AMVP, AFFINE, adaptive DMVR, etc.

[0865] a) In one example, M ( M >=0) TMVPs from N ( N The TMVP is selected from 0 (>=0) co-occurring frames, where each co-occurring frame selects an equal number of TMVPs or selects an unequal number of TMVPs.

[0866] b) In one example, the TMVP candidate list is first constructed, and then all or part of the candidates in the TMVP list are included in the final MVP list.

[0867] i. In one example, S ( S A list of TMVP candidates (>=0) is first constructed, and then they are sorted separately according to some specific metric (e.g., template matching cost), and the top candidates in each TMVP list... M ( M >=0) candidates were included in the final Merge candidate list, where K The constant can be the same for all co-op frames, or it can be different between one co-op frame and another.

[0868] ii. In one example, a TMVP candidate list is constructed for each co-located frame, each list including all or part of the available TMVP candidates in the corresponding co-located frame, and then they are sorted according to a specific metric (e.g., template matching cost), and the top... M ( M >=0) candidates were included in the final Merge candidate list, where K The constant can be the same for all co-op frames, or it can be different between one co-op frame and another.

[0869] iii. In one example, alternatively, only one TMVP candidate list is constructed to include all or a constant number of available TMVP candidates across all co-located frames, and then it is sorted according to a specific metric (e.g., template matching cost), and the top... M ( M (0) candidates were included in the final Merge candidate list.

[0870] iv. In one example, the sorting metric described above could also be the distance between a particular candidate and the current block.

[0871] c) In one example, alternatively, no TMVP candidate list needs to be constructed, and the TMVP associated with a specific location in one or more co-located frames is directly included in the MVP list.

[0872] d) In one example, alternatively, H ( H >=0) TMVP candidates are first included in a joint candidate group containing multiple types of MVP candidates, then the first M ( M (0) candidates were included in the final Merge candidate list based on some specific metrics.

[0873] i. The types of MVP candidates in the joint group include, but are not limited to, adjacent candidates, non-adjacent candidates, HMVP, zero candidates, constructed candidates, etc.

[0874] ii. H ( H (>=0) TMVP candidates can be collected from some or all of the same frame.

[0875] iii. In one example, the ranking metric could be template matching cost or bilateral matching cost.

[0876] 18. In one example, at least two TMVPs from different co-located frames can be used together to generate the final prediction.

[0877] a) In one example, the average or weighted average of two or more TMVPs can be used as the MV or MVP of the current block.

[0878] b) In one example, predictions generated by two or more TMVPs can be averaged or weighted to generate the prediction for the current block.

[0879] 19. A method is proposed to construct at least one list of motion displacements to derive motion displacements for SbTMVP / AFFINE control points / TMVP.

[0880] a) In one example, each candidate in the motion displacement list is an MV pointing to the corresponding co-frame.

[0881] b) In one example, each candidate in the motion displacement list is the MV used to locate the SbTMVP / AFFINE control point / TMVP in the corresponding co-position frame.

[0882] c) In one example, motion displacement candidates can be included in the motion displacement list before being rounded to a specific precision.

[0883] i. In one example, specifically, motion displacement candidates can be included in the motion displacement list before being rounded to integer precision.

[0884] d) In one example, whether to construct a list of motion displacements may depend on, for example... Figure 4 The availability of the template shown.

[0885] i. In one example, if the corresponding template for the current CU does not exist or is unavailable, then the motion displacement list does not need to be constructed.

[0886] e) In one example, the motion displacement candidates in the list can be obtained from a specific block (such as a CU that has already been encoded and decoded).

[0887] i. In one example, the motion information of the encoded and decoded CU is first obtained, and if the corresponding MV points to a specific co-frame, the MV is inserted into the motion displacement list after performing a redundancy check.

[0888] ii. In one example, specifically, the motion displacement candidate can be a neighboring candidate, which is obtained from the neighboring CU.

[0889] 1) In one example, only a few fixed positions can be used to obtain adjacent candidates.

[0890] 2) Alternate locations: Any adjacent location can be used to obtain adjacent candidates.

[0891] iii. In one example, specifically, motion displacement candidates can be non-adjacent candidates, which are obtained from non-neighboring CUs.

[0892] 1) In one example, only a few fixed positions can be used to obtain non-adjacent candidates.

[0893] 2) Alternate locations: Any non-adjacent location can be used to obtain non-adjacent candidates.

[0894] iv. In one example, specifically, motion displacement candidates can be obtained from a list of MVs that maintains the MVs of CUs in history.

[0895] 1) In one example, motion displacement candidates are obtained from a history-based list of MVPs (HMVPs).

[0896] v. In one example, specifically, the motion displacement candidate can be a dummy candidate.

[0897] 1) In one example, a motion candidate can be a zero candidate or a constructed candidate.

[0898] vi. Alternatively, the motion displacement candidate can be any MV pointing to the same frame.

[0899] f) In one example, the MVP candidate list built for some specific codec mode (e.g., regular / CIIP / MMVD / GPM / TPM) can be reused to obtain motion displacement candidates.

[0900] 20. The list of motion displacements can be constructed together with deduplication.

[0901] a) In one example, deduplication is used to avoid duplicate or redundant movement shifts in the list, which can be achieved by using an appropriate threshold. TH To achieve this.

[0902] b) In one example, if two motion displacement candidates point to the same co-position frame, then only if one or both of the absolute differences between the corresponding X and Y components are greater than (or not less than) the specified values. TH Only then can they all be included in the list of motion displacements.

[0903] c) The deduplication threshold can be transmitted via signal in the bitstream.

[0904] i. In one example, the deduplication threshold can be transmitted via signaling at the PU, CU, CTU, or strip level.

[0905] d) The deduplication threshold can depend on the characteristics of the current block.

[0906] i. In one example, the threshold can be derived by analyzing the diversity among the candidates.

[0907] ii. In one example, the optimal threshold can be derived using RDO.

[0908] iii. In one example, the threshold could be the λ value used in the RDO process, or some other value derived based on the λ value.

[0909] twenty one. K ( K A list of >=1 motion displacements can be constructed to derive at least one SbTMVP / AFFINE control point / TMVP candidate.

[0910] a) In one example, the number of candidates in each list may not exceed a certain constant.

[0911] b) In one example, the list of motion displacements is constructed by iterating through motion candidates in a predefined order.

[0912] i. In one example, motion candidates could be: 1) Neighboring MVPs; 2) (Multiple) adjacent neighboring MVPs at specific locations; 3) TMVP MVP; 4) HMVP MVP; 5) Non-adjacent MVPs; 6) Constructed MVPs (such as pairwise MVPs); 7) Inherited affine MV candidates; 8) Constructing affine MV candidates; 9) SbTMVP candidate.

[0913] c) In one example, if the number of candidates in the list reaches the maximum allowed number. K If this happens, the construction of the motion displacement list will terminate.

[0914] d) In one example, M ( M A list of motion displacements >= 0 was constructed to derive SbTMVP / AFFINE control points / TMVP candidates.

[0915] e) In one example D ( D >=0) motion displacements are selected from each list, where D can be the same value for any list of motion candidates, or can be different from one to another.

[0916] f) In one example, the number of motion displacement lists to be constructed may depend on the number of co-frames.

[0917] i. In one example, only one list of motion displacements is constructed for all co-located frames.

[0918] 1) In one example, specifically, the motion information of potential motion displacement candidates is first obtained if the corresponding MV in any list of reference images points to... M ( M If any of the (>=1) co-located frames is selected, the MV is inserted into the motion displacement list after performing a redundancy check.

[0919] ii. In one example, the number of lists can be equal to the number of co-frames, where a list of motion displacements is constructed for each co-frame.

[0920] 1) In one example, for each of the co-located frames, a corresponding motion displacement list is constructed by including motion candidates, where each motion candidate has a motion displacement value (MV) in any reference list pointing to the current co-located frame. Specifically, motion information of potential motion displacement candidates is first obtained, and if the corresponding MV points to the current co-located frame, the MV is inserted into the motion displacement list constructed for the current co-located frame after performing a redundancy check.

[0921] 22. The constructed list of motion displacements can be sorted based on at least one specific metric.

[0922] a) In one example, the metric could be template matching cost or bilateral matching cost.

[0923] b) In one example, the reference template is located in the same frame and the current template is located in the frame to be encoded or decoded, and then the template matching cost is calculated for some or all of the motion displacement candidates.

[0924] c) In one example, the list of motion displacements can be sorted based on template matching of the motion displacement candidates in the list.

[0925] d) In one example, the motion displacement list can be executed. Q ( Q >1) Secondary sorting process.

[0926] i. The Q-ordering process can be executed in a cascaded or parallel manner.

[0927] ii. In one example, for a specific type of candidate, the candidate group is first constructed, and then... K (0< K < QA final reordering pass is performed to select a subset of candidates from the motion displacement list. The last reordering pass is then performed to reorder all candidates from the motion displacement list.

[0928] e) In one example, alternatively, the list of motion displacements may not need to be sorted.

[0929] i. In one example, the list of motion displacements is constructed by iterating through motion candidates in a predefined order.

[0930] ii. In one example, if the number of candidates in the list reaches the maximum allowed number. K If this happens, the construction of the motion displacement list will terminate.

[0931] twenty three. H ( H >=1) SbTMVP / AFFINE control point / TMVP candidates can be included in the sub-CU / CU level MVP candidate list.

[0932] a) In one example, some or all of the candidates in the motion displacement list are used to derive SbTMVP / AFFINE control points / TMVP candidates.

[0933] b) In one example, at most Q ( Q >0) motion displacements are used to derive SbTMVP / AFFINE control point / TMVP candidates in a list of motion displacements.

[0934] c) In one example, no motion displacements in a particular list of motion displacements were used to derive SbTMVP / AFFINE control points / TMVP candidates.

[0935] 24. The list of motion displacements with the lowest metric cost. M ( M >0) candidates can be used to derive SbTMVP / AFFINE control points / TMVP candidates.

[0936] a) In one example, M The values ​​can be the same for each list, or they can be different from one to another.

[0937] b) In one example, temporal motion information (e.g., MV, reference index, etc.) at the location specified by the motion displacement is used by the current block.

[0938] i. In one example, the MV at the position specified by the motion displacement can first undergo a scaling process, and then the scaled MV is used by the current block.

[0939] ii. In one example, alternatively, the MV at the position specified by the motion displacement can be used directly by the current block without performing a scaling process.

[0940] iii. In one example, the MV scaling process can be performed based on the POC distance between the frame to be encoded / decoded and the co-frame, as well as the distance between the co-frame and the corresponding reference frame.

[0941] iv. In one example, specifically, temporal motion information can be used to derive TMVP, SbTMVP, or any other motion information.

[0942] c) In one example, the SbTMVP / AFFINE control point / TMVP candidates constructed based on motion displacement can be further reordered.

[0943] i. In one example, the top of the sorted SbTMVP / AFFINE control points / TMVP set K ( K >0) candidates can be included in the MVP candidate list at the sub-CU / CU level.

[0944] d) In one example, the costs of motion displacement candidates belonging to different motion displacement lists can be compared to determine which one(s) can be used to derive SbTMVP / AFFINE control point / TMVP candidates.

[0945] i. In one example, a list A The first in i The cost of each motion displacement candidate can be compared with the list. B The first in j The costs of the candidates are compared, and the one with the smaller cost will be used to derive the SbTMVP / AFFINE control point / TMVP candidate.

[0946] 25. Before motion displacement is used to derive SbTMVP / AFFINE control points / TMVP candidates, it can be refined or not refined first through the template matching process.

[0947] a) In one example, during template matching refinement, only integer positions are searched to obtain the refined motion displacement.

[0948] b) In one example, the search step size (i.e., the closest distance between two search positions) can be... K ( K >0) pixels.

[0949] 26. New motion displacements can be constructed from existing motion displacements in the list.

[0950] a) In one example, the motion displacement can be constructed by averaging any K (K>1) displacements in the list.

[0951] b) In one example, the constructed motion displacements are reordered together with the motion displacements in the motion displacement list.

[0952] i. In one example, if the constructed motion displacement satisfies a specific condition for being another normal displacement candidate, it can be used to derive an SbTMVP candidate.

[0953] 27. Regarding the reordering of SbTMVP candidates. Once SbTMVPs are derived based on motion displacements selected from one or more lists of motion displacements, they will be included in the sub-CU level MVP candidate list and can then be reordered according to a specific metric.

[0954] a) In one example, the metric could be template matching cost or bilateral matching cost.

[0955] b) In one example, the metrics (e.g., template matching cost) for some or all of the SbTMVP and AFFINE candidates in the list are computed, and then some or all of the candidates are reordered in descending (or ascending) order of the metrics.

[0956] i. In one example, the metric may be further adjusted before reordering, for all or part of the candidates.

[0957] 1) In one example, such that C ori The original metric value representing a specific candidate, and the adjusted metric value. C adj Calculated as: C adj = a C ori + b, in a b can be a fixed constant or an adaptively determined value.

[0958] 2) In one example, whether the metric is adjusted for a particular candidate may depend on the candidate category.

[0959] a) In one example, only some (or all) of the AFFINE (or sbTMVP) candidates need to have their metrics adjusted.

[0960] 3) In one example, whether the metric is adjusted for a particular candidate may depend on whether it is a one-way or two-way prediction.

[0961] ii. In one example, all SbTMVP candidates are reordered based on a specific metric, and then all sorted SbTMVP candidates can be placed before (or after) all or part of the AFFINE candidates.

[0962] iii. In one example, some SbTMVP candidates are reordered along with all AFFINE candidates, while other SbTMVP candidates will always be placed in a fixed position in the list.

[0963] 1) In one example, specifically, which SbTMVP candidates are reordered or not can depend on the co-frame.

[0964] a) In one example, which SbTMVP candidate(s) is reordered or not may depend on the priority of the co-position frame(s) in which the SbTMVP candidate(s) are located.

[0965] i. In one example, co-frames that are closer to the frame to be encoded or decoded are assigned higher priority.

[0966] ii. In one example, located in the area with the previous W ( W >0) All or some of the SbTMVP candidates in the highest priority co-frame will not be reordered.

[0967] 1. In one example, these SbTMVP candidates can always be placed at the top of the list.

[0968] 2. In one example, alternatively, these SbTMVP candidates can be placed in any fixed position in the list.

[0969] 3. Alternatively, these SbTMVP candidates can be placed anywhere in the list.

[0970] 2) In one example, specifically, which SbTMVP candidates are reordered or not may depend on the ranking of the corresponding motion displacement in the motion displacement list.

[0971] a) In one example, SbTMVP candidates with higher (or lower) motion displacements may not be reordered, but instead placed at the top of the list.

[0972] 3) In one example, specifically, which SbTMVP candidates are reordered or not may depend on both the corresponding rank of the motion displacements in the motion displacement list and the co-frames in which they are located.

[0973] a) In one example, if the SbTMVP is in the top M (M>0) highest priority frames and the associated motion displacement is in the top N (N>0) of the corresponding motion displacement list, then the SbTMVP can be placed in the first (or any other) position in the sorted sub-CU level MVP candidate list.

[0974] 28. In one example, the motion displacements obtained from the list of motion displacements can be refined before they are used to locate the position in at least one co-located frame for TMVP or sbTMVP.

[0975] a) In one example, motion displacement can be refined through template matching.

[0976] b) In one example, motion displacement can be refined using bilateral matching.

[0977] c) In one example, the motion displacement can be refined by adding an increment MV.

[0978] d) In one example, the motion displacement can be refined by limiting.

[0979] e) In one example, motion displacement can be refined by shifting.

[0980] 29. In one example, multiple motion displacements (denoted as SM0, SM1, ..., SMn) from a list of motion displacements can be combined to derive a final motion displacement (denoted as SMf) to locate the position in at least one co-located frame for a TMVP or sbTMVP. For example, SMf = F(SM0, SM1, ..., SMn).

[0981] a) In one example, SMf = (SM0 + SM1 + … + SMn) / n.

[0982] b) In one example, SMf = max(SM0, SM1, ..., SMn).

[0983] c) In one example, SMf = min(SM0, SM1, ..., SMn).

[0984] d) In one example, SMf = middle(SM0, SM1, ..., SMn).

[0985] e) In one example, SMf = (W1*SM0 +W*SM1 + … + W3*SMn) / (W1+W2+..+Wn).

[0986] 30. After deriving the SbTMVP candidate, the sub-block motion can be further refined through template matching.

[0987] 3. Problem Existing sub-block-based prediction techniques have the following problems: 1) They face a dilemma. If the sub-block size is smaller, the motion information of each sub-block can be more accurate. However, smaller sub-blocks impose higher bandwidth requirements in the MC.

[0988] 2) Motion information derived for smaller sub-blocks can be dangerous, especially when there is noise in the block. Fixing the sub-block size within a single block may be suboptimal.

[0989] 4. Detailed Solution This disclosure proposes further improvements to sub-block-based motion compensation. Specifically, interleaving prediction is improved to be compatible with pixel-based affine motion compensation. Furthermore, regression affine is also improved.

[0990] The specific embodiments described below should be considered as examples for interpreting general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0991] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, ​​TU, PB, TB.

[0992] In this disclosure, "blocks encoded and decoded in mode N" can refer to a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, merging, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).

[0993] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.

[0994] Regarding interlaced affine 1. The use of interleaved prediction in affine mode can depend on another encoding / decoding method.

[0995] a) In one example, if pixel-based affine motion compensation is applied to the codec block, then interleaving prediction may not be applied.

[0996] b) In one example, interleaved affine prediction and pixel-based affine prediction cannot be used simultaneously for encoding / decoding blocks.

[0997] i. In one example, alternatively, interleaved affine prediction and pixel-based affine prediction can be used simultaneously for encoding and decoding blocks.

[0998] 1. In one example, for a codec block, some pixels can be predicted by interleaved affine mapping, while (some) other pixels can be predicted by pixel-based affine mapping.

[0999] 2. In one example, interleaved affine predictions and pixel-based affine predictions can be mixed to generate new predictions.

[1000] 2. In one example, the use of interleaving prediction on affine mode can depend on the sequence / frame resolution.

[1001] 3. Interleaved affine prediction and BDOF can be used simultaneously for encoding and decoding blocks.

[1002] i. In one example, BDOF can be used to obtain predictions if the affine-coded block is bidirectionally predictable and / or some specific conditions are met.

[1003] 1. In one example, specifically, BDOF and interleaved affine prediction can be used in a cascaded manner.

[1004] a) In one example, BDOF can be used before or after interleaved affine prediction.

[1005] i. In one example, after the first version of the affine motion (or prediction) is generated, BDOF can be applied to obtain the second version of the affine motion (or prediction), and then interleaved affine can be used to obtain the third version of the affine motion (or prediction) based on the second version.

[1006] 1. In one example, the third version prediction is used as the final affine prediction.

[1007] 2. In one example, specifically, multiple versions of the prediction can be combined using a weighted average to obtain a fourth version of the prediction, which is used as the final affine prediction.

[1008] 2. In one example, specifically, BDOF and interleaved affine prediction can be used in parallel.

[1009] a) In one example, specifically, after the first version of the affine motion (or prediction) is generated, the third affine motion (or prediction) and the third affine motion (or prediction) are generated in parallel via BDOF and interleaved affine prediction, respectively.

[1010] 1. In one example, specifically, multiple versions of the prediction can be combined using a weighted average to obtain a fourth version of the prediction, which is used as the final affine prediction.

[1011] 3. In one example, alternatively, when BDOF is used for affine codec blocks, interleaved affine prediction is not used.

[1012] 4. The use of interleaved affine prediction for codec blocks can depend on whether it is bidirectional or unidirectional prediction.

[1013] a) In one example, interleaved affine prediction can be used only for blocks of bidirectional prediction.

[1014] b) In one example, alternatively, interleaved affine prediction can be used only for blocks that are predicted in one direction.

[1015] c) In one example, interleaved affine prediction can be used for blocks that are predicted unidirectionally and blocks that are predicted bidirectionally.

[1016] 5. The use of interleaved affine prediction and / or TM-affine can depend on the strip / frame type.

[1017] a) In one example, interleaved affine prediction and / or TM-affine may not be applied to GPB (Generalized P / B Frame) stripes / frames.

[1018] i. In one example, if all reference images / frames / strips of the image / frame / strip to be encoded / decoded are collected from the forward direction according to the POC, then the current strip / frame is a GPB strip / frame. ii. In one example, specifically, if all reference pictures / frames / strips of the picture / frame / strip to be encoded / decoded are collected from the forward direction about the POC, then interleaved affine prediction and / or TM-affine prediction may not be applied.

[1019] b) In one example, interleaved affine prediction and / or TM-affine can be applied to a portion of GPB (Generalized P / B Frame) strips / frames.

[1020] i. In one example, for a GPB stripe / frame, if the POC distance between it and the nearest reference frame is greater than 1 (i.e., a GPB frame in the random access configuration), interleaved affine prediction and / or TM-affine prediction can be applied to the GPB stripe / frame.

[1021] 1. In one example, otherwise, interleaved affine prediction and / or TM-affine may not be applied to the GPB strip / frame.

[1022] ii. In one example, for a GPB stripe / frame, interleaved affine prediction and / or TM-affine can be applied to the GPB stripe / frame if the POC distance between it and the nearest reference frame is greater than (or less than) a certain value.

[1023] 6. The use of interleaved affine prediction and / or pixel-based affine motion compensation can depend on the block dimension.

[1024] a) In one example, interleaved affine prediction is used if the block dimension meets certain conditions.

[1025] i. In one example, specifically, interleaved affine prediction can be used if both (or either) the height and width are greater than or less than a threshold, or if the ratio of height to width is greater than or less than a threshold.

[1026] ii. In one example, specifically, interleaved affine prediction can be used if the block area is greater than or equal to a threshold.

[1027] b) In one example, alternatively, if the block dimensions do not meet the same conditions, pixel-based affine mapping can be used instead.

[1028] c) In one example, dimension conditions can be applied together with other conditions.

[1029] i. In one example, the OBMC flag can be used together with dimension conditions.

[1030] 1. In one example, specifically, if OBMC is used for blocks and the block dimensions meet certain conditions, then pixel affine mapping is used to generate predictions.

[1031] d) In one example, whether to apply interleaved affines can depend on the sub-block size.

[1032] i. In one example, if the sub-block size is smaller than a threshold, such as 4×4, then interleaved affine is not applied.

[1033] 7. In one example, the size of the largest sub-block to be used for affine prediction is set to a constant when using interleaved affine or not using pixel-based affine.

[1034] 8. In one example, PROF (Affine Prediction Refinement with Optical Flow) can be performed for interleaved affines.

[1035] a) In one example, the predictions generated by the two partitioning patterns are both refined by PROF.

[1036] i. In one example, only predictions generated by a specific partitioning pattern are refined by PROF.

[1037] ii. In one example, if interleaved affines are used for the block, PROF is not applied.

[1038] b) In one example, for a specific partitioning pattern, only a subset of predicted samples are refined by PROF.

[1039] i. In one example, if the width or height of a sub-block is less than a threshold (e.g., 4), PROF will not be applied to that sub-block.

[1040] 9. It is proposed that whether and / or how to apply sub-block boundary deblocking procedures or other types of sub-block boundary-based filtering procedures (such as SAO, adaptive loop filter, OBMC) can depend on whether affine mode / pixel-based affine / interleaved prediction is applied.

[1041] a) In one example, if affine mode / pixel-based affine / interleaved prediction is applied, some or all kinds of sub-block boundary-based deblocking / filtering processes are not applied.

[1042] i. In one example, specifically, if affine mode / pixel-based affine / interlacing prediction is applied, then deblocking and / or OBMC facing the sub-block boundary is not applied.

[1043] 1. In one example, specifically, whether / how a particular type of sub-block boundary-based deblocking / filtering process is applied can also depend on the stripe type.

[1044] a) In one example, specifically, whether / how a particular type of sub-block boundary-based deblocking / filtering process is applied may also depend on whether the particular frame is a GPB.

[1045] b) In one example, specifically, whether / how a particular type of sub-block boundary-based deblocking / filtering process is applied may also depend on the POC distance between a particular frame and the nearest reference frame.

[1046] About OBMC 10. It was proposed that whether and / or how to apply CU-level OBMC can depend on the encoding / decoding mode.

[1047] a) In one example, if a sub-block level inter-frame mode (such as affine, SbTMVP, multi-pass DMVR) is used for the current block, each boundary sub-block will be filtered independently at the CU level OBMC level.

[1048] i. In one example, specifically, boundary sub-blocks are iterated one by one. Suppose A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block used for OBMC filtering. If A and A1 have different motions, then A will be filtered. After A has been filtered, subsequent sub-blocks will undergo OBMC in a similar manner.

[1049] b) In one example, alternatively, if a non-sub-block level inter-frame mode is used for the current block, multiple boundary sub-blocks can be filtered together using OBMC.

[1050] i. Figure 25 An example of a CU-level OBMC is shown. In one example, assume A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block. If A1 and subsequent... N ( N >=0) consecutive sub-blocks (e.g. Figure 25 The two sub-blocks B1 and C1 in the middle have the same motion. M_nei And at the same time M_nei Unlike the motion of A, then ( N +1) sub-blocks (including A) will be filtered together. Otherwise, if M_nei The motion equal to A, then ( N +1) consecutive sub-blocks (including A) will skip CU-level OBMC filtering.

[1051] Regarding regressive affine 11. Whether a previously affine-encoded block can be used to generate regression affine candidates may depend on the block size.

[1052] a) In one example, if the block width / height / width-to-height ratio / area is less than or greater than a predefined threshold or an adaptively derived threshold, the previously affine-encoded block may not be used to generate regression affine candidates.

[1053] b) In one example, if the ratio of the area of ​​a previously affine-encoded block to the area of ​​the current block is less than or greater than a predefined threshold or an adaptively derived threshold, then the previously affine-encoded block may not be used to generate regression affine candidates.

[1054] 12. To generate regression affine candidates, a method using... N ( N >1) The motion field of the previously encoded and decoded CU is used as the input to the regression process.

[1055] a) In one example, for any CU used to generate a regression affine candidate, at least one sub-block is used to provide the motion field.

[1056] b) In one example, for any CU used to generate a regression affine candidate, all sub-blocks of that CU are used to provide the motion field.

[1057] c) In one example N A CU can be collected from adjacent locations, non-adjacent locations, or historical CU / parameter tables.

[1058] d) In one example, at least one previously encoded / decoded CU can be collected from the following locations: i. (Multiple) adjacent or neighboring locations.

[1059] ii. (Multiple) adjacent locations at a specific location.

[1060] iii. (Multiple) co-located or adjacent time-domain locations.

[1061] iv. Historical table.

[1062] v. (Multiple) non-adjacent spatial / temporal locations.

[1063] e) In one example N Each CU is encoded and decoded via inter-frame encoding / decoding.

[1064] i. In one example, N Each CU is encoded and decoded using affine mode.

[1065] ii. In one example, alternatively, N At least one of them K ( K (0) are affine encoded / decoded.

[1066] iii. In one example, the motion fields of at least one affine-coded CU and at least one non-affine-coded CU can be used to generate regressive affine candidates.

[1067] f) In one example, all N Each CU may need to share the same prediction direction (i.e., bidirectional or unidirectional prediction, and / or the reference list used) and / or the same reference index / frame.

[1068] i. In one example, alternatively, they may have different prediction directions or reference frames.

[1069] g) In one example, only when N A regression affine candidate can only be generated when the number of sub-blocks of a CU is greater than a constant threshold.

[1070] h) In one example, sports fields in adjacent or non-adjacent locations can also be used as additional inputs.

[1071] i. In one example, M Row / column adjacent sub-blocks and / or F Non-adjacent sub-blocks in rows / columns can be used as input, where M, F >=0.

[1072] 1. In one example, the location of the non-adjacent positions used can depend on the block dimension.

[1073] i) In one example, the proposed regression candidate can be used when the number of existing regression candidates has not reached the maximum allowed number.

[1074] j) In one example, the proposed regression candidates can be reordered based on a specific metric (such as ARMC or template matching).

[1075] k) In one example, the number of proposed regression candidates may not exceed a constant or an adaptively determined value.

[1076] l) In one example, the proposed regression affine candidate can be used to generate affine Merge / affine AMVP / affine MMVD / affine Adaptive DMVR / affine TM / affine DMVR and / or any other affine-related method that requires the construction of an affine candidate list.

[1077] 13. A motion field of at least K (K>1, e.g. K=2) codec blocks can be used to generate regression affine candidates for affine AMVP patterns.

[1078] a) In one example, a codec block can only be used to generate regression affine candidates if the reference index or reference frame used by the codec block is the same as the reference index of the current block.

[1079] 14. The affine candidate list can contain regressive affine candidates generated from different numbers of previously encoded / decoded blocks.

[1080] a) In one example, at least one regressive affine candidate in the list is generated using M (M>0) previously encoded / decoded blocks, and / or at least one regressive affine candidate in the list is generated using N (N>0) previously encoded / decoded blocks, where M and N are not the same.

[1081] b) In one example, regression affine candidates generated using more previously encoded blocks have a higher priority for being included in the affine candidate list compared to candidates generated using fewer previously encoded blocks.

[1082] Regarding SbTMVP 15. The use of SbTMVP can depend on the temporal layer or temporal layer index (Tid) of the frame / strip / slice associated with the codec block.

[1083] a) In one example, SbTMVP is enabled for blocks when the time layer or Tid is less than or greater than a constant (or determined on the fly).

[1084] i. In one example, specifically, if the time-domain layer or Tid is less than or greater than a constant (or instantaneously determined) value, no SbTMVP candidate is allowed to be inserted into the sub-block-based candidate list.

[1085] b) In one example, the maximum allowed number of SbTMVP candidates that can be included in the sub-block-based candidate list may depend on the time-domain layer or Tid.

[1086] i. In one example, specifically, when Tid is less than or greater than the threshold, at most K One SbTMVP candidate can be included in the sub-block-based candidate list, otherwise at most M Each SbTMVP candidate can be included in a sub-block-based candidate list. Here K and M These can be different values.

[1087] 1. In one example, specifically, the threshold can be a constant (or an on-the-fly determined) value.

[1088] 2. In one example, specifically, multiple thresholds can be used to determine whether and / or how many SbTMVP candidates can be inserted into the sub-block-based candidate list.

[1089] 16. Once one or more SbTMVPs are derived, they can be included in a sub-block-based candidate list and then reordered according to a specific metric.

[1090] a) In one example, the metric could be template matching cost or bilateral matching cost.

[1091] b) In one example, a metric (e.g., template matching cost) is computed for some or all of the SbTMVP and affine candidates in the list, and then some or all of the candidates are reordered in descending (or ascending) order of the metric.

[1092] i. In one example, all SbTMVP candidates can be reordered based on a specific metric, and then all sorted SbTMVP candidates can be placed before (or after) all or part of the affine candidates.

[1093] ii. In one example, some or all of the SbTMVP candidates may always occupy some fixed positions in the reordered list, while the order of other SbTMVP candidates may be determined based on a specific metric (i.e., template matching or bilateral matching cost).

[1094] 1. In one example, specifically, for some or all of the SbTMVP candidates, a particular metric or cost is calculated, which can then be sorted together with the same cost for all or some of the affine candidates, and the order of these SbTMVP and / or affine candidates in the candidate list based on the sub-block can be determined based on the cost sorting result.

[1095] a) In one example, the order of the candidates in the list could be based on the cost order.

[1096] 2. In one example, specifically, whether a particular SbTMVP candidate is reordered can be determined by the temporal layer or Tid of the codec block / image.

[1097] a) In one example, if the temporal layer or Tid is less than (or greater than) a threshold, all SbTMVP candidates are reordered together with affine candidates based on a specific metric.

[1098] i. In one example, the order of candidates in the reordered list could be based on the sorting order of the metric.

[1099] b) In one example, alternatively, if the temporal layer or Tid is greater than (or less than) a threshold, a particular(multiple) SbTMVP candidates may directly occupy one or more fixed positions in the reordered list.

[1100] i. In one example, specifically, if the temporal layer or Tid is greater than (or less than) a threshold, then a particular(multiple) SbTMVP candidates can directly occupy the first or top(multiple) positions in the reordered list without reordering.

[1101] ii. In one example, specifically, for the remaining SbTMVP candidates, a specific metric or cost is computed, which can then be sorted together with the same cost for all or some of the affine candidates, and the order of these SbTMVP and / or affine candidates in the candidate list based on the sub-block can be determined based on the cost sorting result.

[1102] 3. In one example, specifically, whether a particular SbTMVP candidate is reordered may depend on the co-frame.

[1103] a) In one example, which SbTMVP candidates are reordered or not may depend on the priority of the co-frame in which they are located.

[1104] i. In one example, co-frames that are closer to the frame to be encoded or decoded are assigned higher priority.

[1105] ii. In one example, a co-frame with a larger (or lower) Qp can be assigned a higher priority.

[1106] iii. In one example, located in the area with the previous W ( W >0) All or some of the SbTMVP candidates in the highest priority co-frame will not be reordered.

[1107] 1. In one example, these SbTMVP candidates can always be placed at the top of the list.

[1108] 2. In one example, alternatively, these SbTMVP candidates can be placed in any fixed position in the list.

[1109] 3. Alternatively, these SbTMVP candidates can be placed anywhere in the list.

[1110] 4. In one example, specifically, whether a particular SbTMVP candidate is reordered may depend on the ranking of the corresponding motion displacement in the motion displacement list.

[1111] a) In one example, SbTMVP candidates with higher (or lower) motion displacements may not be reordered, but instead placed at the top of the list.

[1112] 5. In one example, specifically, whether a particular SbTMVP candidate is reordered may depend on a combination of the following conditions: (1) the corresponding rank of the motion displacement in the motion displacement list; (2) the co-position frame in which the SbTMVP is located; and (3) the temporal layer or Tid of the codec block / picture.

[1113] a) In one example, if the SbTMVP is in the top M (M>0) highest priority frames, and / or the associated motion displacement is in the top N (N>0) of the corresponding motion displacement list, and / or the temporal layer or Tid of the codec block / frame is greater than (or less than) the threshold, then the SbTMVP can be placed in the first (or any other) position in the sorted sub-CU level MVP candidate list without reordering.

[1114] General aspects 17. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.

[1115] 18. Whether and / or how the methods disclosed above can be applied can be transmitted via signal at the sequence level / picture group level / picture level / strip level / film group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / film group header.

[1116] 19. Whether and / or how to apply the above methods may depend on the following information: a) Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Piece Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video Codec Unit.

[1117] b) Location of CU / PU / TU / block / video codec unit.

[1118] c) The block dimension of the current block and / or its neighboring blocks.

[1119] d) The block shape of the current block and / or its neighboring blocks.

[1120] e) The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode.

[1121] f) Indication of color format (such as 4:2:0, 4:4:4).

[1122] g) Encoding / decoding tree structure.

[1123] h) Strip / panel type and / or image type.

[1124] i) Color components (e.g., can be applied only to the chromaticity component or the luminance component).

[1125] j) Temporal layer ID.

[1126] k) Standard grade / level / layer.

[1127] Further embodiments will be described below. Figure 26 A flowchart of a method 2600 for video processing according to an embodiment of the present disclosure is shown. Method 2600 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.

[1128] At box 2610, the conversion between the current video block and the video bitstream is based on stripe type determination information about applying a deblocking or filtering process based on sub-block boundaries to the current video block. As used herein, the term "current video block" refers to the video block or video unit to be processed, which may also be referred to as the "target video block" or "current video unit".

[1129] At box 2620, the conversion is performed based on information. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how such a process is applied. In some embodiments, the conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the conversion includes decoding the current video block from the bitstream.

[1130] Method 2600 enables the determination of whether and / or how to apply sub-block boundary-based deblocking or filtering processes based on the stripe type. Therefore, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.

[1131] In some embodiments, the information is based on whether the frame is a generalized P or B (GPB) frame.

[1132] In some embodiments, the information is also based on the picture order count (POC) distance between a frame and its nearest reference frame.

[1133] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a video processing apparatus. In this method, information regarding the application of a sub-block boundary-based deblocking or filtering process to a current video block is determined based on the stripe type. The bitstream is generated based on this information. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[1134] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, information regarding the application of a sub-block boundary-based deblocking or filtering process to a current video block is determined based on the stripe type. A bitstream is generated based on the information. The bitstream is stored in a non-transitory computer-readable recording medium. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[1135] Figure 27 A flowchart of a method 2700 for video processing according to an embodiment of the present disclosure is shown. Method 2700 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.

[1136] At box 2710, for the conversion between the current video block and the video bitstream, the use of at least one of interleaved affine prediction or template matching affine is determined based on at least one of the following: strip type, frame type, or picture order count (POC) distance between the strip or frame and its nearest reference frame.

[1137] At box 2720, the conversion is performed based on the use of at least one of interleaved affine prediction or template matching affine. In some embodiments, the conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the conversion includes decoding the current video block from the bitstream.

[1138] Method 2700 enables the determination of the use of interleaved affine prediction or template matching affine based on stripe type, frame type, and / or POC. In this way, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.

[1139] In some embodiments, a reference image, reference frame, or reference stripe of the image or frame or stripe to be encoded or decoded is collected from the forward direction according to the POC, and the stripe type of the current stripe or the frame type of the current frame is a generalized P or B (GPB) stripe or frame.

[1140] In some embodiments, for a generalized P or B (GPB) stripe or frame, the POC distance between the GPB stripe or frame and its nearest reference frame is greater than 1, or the GPB stripe or frame is a GPB frame in the random access configuration, and at least one of interleaved affine prediction or template matching affine is not applied to the GPB stripe or frame.

[1141] In some embodiments, for a generalized P or B (GPB) strip or frame, the POC distance between the GPB strip or frame and its nearest reference frame is greater than or less than a predefined value, and at least one of interleaved affine prediction or template matching affine is applied to the GPB strip or frame.

[1142] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a video processing apparatus. In this method, the use of at least one of interleaved affine prediction or template matching affine is determined based on at least one of: stripe type, frame type, or picture order count (POC) distance between a stripe or frame and its nearest reference frame. The bitstream is generated based on the use of at least one of interleaved affine prediction or template matching affine.

[1143] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, the use of at least one of interleaved affine prediction or template matching affine is determined based on at least one of: stripe type, frame type, or picture order count (POC) distance between a stripe or frame and its nearest reference frame. The bitstream is generated based on the use of at least one of interleaved affine prediction or template matching affine. The bitstream is stored in a non-transitory computer-readable recording medium.

[1144] Figure 28 A flowchart of a method 2800 for video processing according to an embodiment of the present disclosure is shown. Method 2800 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.

[1145] At box 2810, for the conversion between the current video block and the video bitstream, the interleaving prediction used for the affine mode is determined based on the sequence resolution or frame resolution.

[1146] At box 2820, the conversion is performed based on the use of interleaving prediction. In some embodiments, the conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the conversion includes decoding the current video block from the bitstream.

[1147] Method 2800 enables the determination of interleaving prediction for affine modes based on sequence / frame resolution. Therefore, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.

[1148] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video, the bitstream of which is generated by a method performed by means of a video processing apparatus. In this method, interleaving prediction is used for affine modes based on sequence resolution or frame resolution. The bitstream is generated based on the use of interleaving prediction.

[1149] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, an affine mode is determined based on sequence resolution or frame resolution using interleaving prediction. The bitstream is generated based on the use of interleaving prediction. The bitstream is stored in a non-transitory computer-readable recording medium.

[1150] As used herein, a video block or video unit may refer to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of a block, sub-region within a block, or region containing more than one sample or pixel.

[1151] In some embodiments, whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlapped block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP) is based on syntax elements in the bitstream. For example, whether and / or how method 2600, method 2700, and / or method 2800 are applied may be based on syntax elements in the bitstream.

[1152] In some embodiments, the syntax element is located at at least one of the following: sequence level, picture group level, picture level, strip level, or slice group level, or wherein the syntax element is included in at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[1153] In some embodiments, whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlapping block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP) is based on at least one of the following: messages in the bitstream, the position of video units, the block dimension of the current video block, the block dimension of the neighboring blocks of the current video block, the block shape of the current video block, the block shape of the neighboring blocks of the current video block, the codec mode of the block, the color format indication, the codec tree structure, the stripe or slice group type, the picture type, the color components, the temporal layer identifier (ID), or the grade or level or layer of the standard. For example, whether and / or how to apply method 2600, method 2700 and / or method 2800 may be based on at least one of the following: messages in the bitstream, the position of the video unit, the block dimension of the current video block, the block dimension of the neighboring blocks of the current video block, the block shape of the current video block, the block shape of the neighboring blocks of the current video block, the encoding / decoding mode of the block, the color format indication, the encoding / decoding tree structure, the stripe or slice group type, the picture type, the color components, the temporal layer ID, or the standard grade or level or layer.

[1154] In some embodiments, the message is included in at least one of the following: Decoding Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Slice Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Row, LCU Group, Transform Unit (TU), Prediction Unit (PU) Block, or Video Codec Unit.

[1155] In some embodiments, the video unit includes one of the following: a codec unit (CU), a prediction unit (PU), a transform unit (TU), a block or video codec unit.

[1156] In some embodiments, the encoding / decoding mode includes at least one of the following: intra-block copy (IBC) mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

[1157] In some embodiments, the color format includes one of the following: 4:2:0 format or 4:4:4 format.

[1158] In some embodiments, schemes such as method 2600, method 2700 and / or method 2800 are applied to at least one of the chromaticity component or the luminance component.

[1159] It should be understood that methods 2600 and / or 2700 and / or 2800 can be applied individually or in any combination. Using these methods, encoding / decoding efficiency and effectiveness can be improved.

[1160] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[1161] Item 1. A method for video processing, comprising: a conversion between a current video block and a bitstream of video; determining information based on a stripe type, the information relating to the application of a sub-block boundary-based deblocking or filtering process to the current video block; and performing the conversion based on the information, wherein the information indicates at least one of: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[1162] Item 2. According to the method of Item 1, where the information is based on whether the frame is a generalized P or B (GPB) frame.

[1163] Item 3. According to the method of Item 1 or 2, the information is also based on the picture order count (POC) distance between the frame and its nearest reference frame.

[1164] Item 4. A method for video processing, comprising: a conversion between a current video block and a bitstream of the video, determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of: strip type, frame type, or picture order count (POC) distance between a strip or frame and its nearest reference frame; and performing the conversion based on the use of at least one of interleaved affine prediction or template matching affine.

[1165] Item 5. According to the method of Item 4, wherein the reference picture, reference frame or reference strip of the picture or strip to be encoded or decoded is collected from the forward direction according to the POC, and the strip type of the current strip or the frame type of the current frame is a generalized P or B (GPB) strip or frame.

[1166] Item 6. According to the method of Item 4 or 5, wherein for a generalized P or B (GPB) stripe or frame, the POC distance between the GPB stripe or frame and its nearest reference frame is greater than 1, or the GPB stripe or frame is a GPB frame in the case of random access configuration, and at least one of interleaved affine prediction or template matching affine is not applied to the GPB stripe or frame.

[1167] Item 7. According to the method of Item 4 or 5, wherein for a generalized P or B (GPB) strip or frame, the POC distance between the GPB strip or frame and its nearest reference frame is greater than or less than a predefined value, and at least one of interleaved affine prediction or template matching affine is applied to the GPB strip or frame.

[1168] Item 8. A method for video processing, comprising: a conversion between a current video block and a bitstream of the video; determining the use of interleaving prediction for an affine mode based on sequence resolution or frame resolution; and performing the conversion based on the use of the interleaving prediction.

[1169] Item 9. The method according to any one of Items 1 to 8, wherein a video block or video unit includes one of the following: color component, subpicture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.

[1170] Item 10. The method according to any one of Items 1 to 9, wherein whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlapping block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP) is based on syntax elements in the bitstream.

[1171] Item 11. According to the method of Item 10, wherein the syntax element is located at at least one of the following: sequence level, picture group level, picture level, strip level, or slice group level, or wherein the syntax element is included in at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[1172] Item 12. The method according to any one of Items 1 to 9, wherein whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlapping block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP), is based on at least one of the following: messages in the bitstream, the position of the video unit, the block dimension of the current video block, the block dimension of the neighboring blocks of the current video block, the block shape of the current video block, the block shape of the neighboring blocks of the current video block, the codec mode of the block, the color format indication, the codec tree structure, the strip or slice type, the picture type, the color components, the temporal layer identifier (ID), or the grade or level or layer of the standard.

[1173] Item 13. According to the method of Item 12, the message is included in at least one of the following: Decoding Parameter Set (DPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Strip Header, Slice Header, Maximum Codec Unit (LCU), Codec Unit (CU), LCU Line, LCU Group, Transform Unit (TU), Prediction Unit (PU) Block, or Video Codec Unit.

[1174] Item 14. According to the method of Item 12, the video unit includes one of the following: a codec unit (CU), a prediction unit (PU), a transform unit (TU), a block or video codec unit.

[1175] Item 15. The method according to Item 12, wherein the encoding / decoding mode includes at least one of the following: intra-block copy (IBC) mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

[1176] Item 16. According to the method of Item 12, the color format includes one of the following: 4:2:0 format or 4:4:4 format.

[1177] Item 17. The method according to Item 12, wherein the scheme is applied to at least one of the chromaticity component or the luminance component.

[1178] Item 18. The method according to any one of items 1 to 17, wherein the transformation includes encoding the current video block into a bitstream.

[1179] Item 19. The method according to any one of items 1 to 17, wherein the conversion includes decoding the current video block from the bitstream.

[1180] Item 20. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 19.

[1181] Item 21. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of items 1 to 19.

[1182] Item 22. A non-transitory computer-readable recording medium storing a bitstream of video, the bitstream of video being generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining information based on a stripe type, the information relating to the application of a sub-block boundary-based deblocking or filtering process to a current video block of the video; and generating the bitstream based on the information, wherein the information indicates at least one of: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[1183] Item 23. A method for storing a bitstream of video, comprising: determining information based on a stripe type, the information relating to the application of a sub-block boundary-based deblocking or filtering process to a current video block of the video; generating a bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how a sub-block boundary-based deblocking or filtering process is applied.

[1184] Item 24. A non-transitory computer-readable recording medium storing a bitstream of video, the bitstream of video being generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of: strip type, frame type, or picture sequence count (POC) distance between a strip or frame and its nearest reference frame; and generating the bitstream based on the use of at least one of interleaved affine prediction or template matching affine.

[1185] Item 25. A method for storing a bitstream of video, comprising: determining the use of at least one of interleaved affine prediction or template matching affine based on at least one of: strip type, frame type, or picture sequence count (POC) distance between a strip or frame and its nearest reference frame; generating a bitstream based on the use of at least one of interleaved affine prediction or template matching affine; and storing the bitstream in a non-transitory computer-readable recording medium.

[1186] Item 26. A non-transitory computer-readable recording medium storing a bitstream of video, the bitstream of video being generated by a method performed by means of means for video processing, wherein the method includes: determining the use of interleaving prediction for affine modes based on sequence resolution or frame resolution; and generating the bitstream based on the use of interleaving prediction.

[1187] Item 27. A method for storing a bitstream of video, comprising: determining the use of interleaving prediction for an affine mode based on sequence resolution or frame resolution; generating a bitstream based on the use of interleaving prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[1188] Example device Figure 29A block diagram of a computing device 2900 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2900 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[1189] It should be understood that, Figure 29 The computing device 2900 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[1190] like Figure 29 As shown, computing device 2900 includes general-purpose computing device 2900. Computing device 2900 may include at least one or more processors or processing units 2910, memory 2920, storage unit 2930, one or more communication units 2940, one or more input devices 2950, ​​and one or more output devices 2960.

[1191] In some embodiments, the computing device 2900 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2900 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[1192] Processing unit 2910 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2900. Processing unit 2910 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[1193] Computing device 2900 ty...

Claims

1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, information is determined based on stripe type, which relates to applying a deblocking or filtering process based on sub-block boundaries to the current video block; and The conversion is performed based on the information. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how the sub-block boundary-based deblocking or filtering process is applied.

2. The method of claim 1, wherein the information is based on whether the frame is a generalized P or B (GPB) frame.

3. The method according to claim 1 or 2, wherein the information is further based on the picture sequence count (POC) distance between the frame and its nearest reference frame.

4. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, the use of at least one of interleaved affine prediction or template matching affine is determined based on at least one of the following: strip type, frame type, or picture order count (POC) distance between the strip or frame and its nearest reference frame; and The transformation is performed based on the use of at least one of the interleaved affine prediction or the template matching affine.

5. The method according to claim 4, wherein a reference picture, reference frame, or reference stripe of the picture, frame, or stripe to be encoded or decoded is collected from the forward direction according to the POC, and the stripe type of the current stripe or the frame type of the current frame is a generalized P or B (GPB) stripe or frame.

6. The method according to claim 4 or 5, wherein for a generalized P or B (GPB) stripe or frame, the POC distance between the GPB stripe or frame and its nearest reference frame is greater than 1, or the GPB stripe or frame is a GPB frame in a random access configuration, and at least one of the interleaved affine prediction or the template matching affine is not applied to the GPB stripe or frame.

7. The method according to claim 4 or 5, wherein for a generalized P or B (GPB) stripe or frame, the POC distance between the GPB stripe or frame and its nearest reference frame is greater than or less than a predefined value, and at least one of the interleaved affine prediction or the template matching affine is applied to the GPB stripe or frame.

8. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, interleaving prediction is used for the affine mode based on the sequence resolution or frame resolution. as well as The transformation is performed based on the interleaving prediction.

9. The method according to any one of claims 1 to 8, wherein the video block or video unit comprises one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.

10. The method according to any one of claims 1 to 9, wherein whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlap block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP) is based on the syntax elements in the bitstream.

11. The method of claim 10, wherein the syntax element is located at at least one of the following: sequence level, picture group level, picture level, strip level, or slice group level, or The syntax elements mentioned therein include at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

12. The method according to any one of claims 1 to 9, wherein whether and / or how a scheme is applied regarding at least one of interleaved affine mode, codec unit (CU) level overlap block motion compensation (OBMC), regressive affine mode, or sub-block-based temporal motion vector prediction (SbTMVP), is based on at least one of the following: The message in the bitstream, the position of the video unit, the block dimension of the current video block, the block dimension of the neighboring blocks of the current video block, the block shape of the current video block, the block shape of the neighboring blocks of the current video block, the block's encoding / decoding mode, the color format indicator, the encoding / decoding tree structure, the stripe or slice group type, the picture type, the color components, the temporal layer identifier (ID), or the standard's grade or level or layer.

13. The method of claim 12, wherein the message is included in at least one of the following: a Decoding Parameter Set (DPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), a Picture Parameter Set (PPS), an Adaptive Parameter Set (APS), a Picture Header, a Strip Header, a Slice Header, a Maximum Codec Unit (LCU), a Codec Unit (CU), an LCU Line, an LCU Group, a Transform Unit (TU), a Prediction Unit (PU) Block, or a Video Codec Unit.

14. The method of claim 12, wherein the video unit comprises one of the following: a codec unit (CU), a prediction unit (PU), a transform unit (TU), a block or video codec unit.

15. The method of claim 12, wherein the encoding / decoding mode includes at least one of the following: intra-block copy (IBC) mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

16. The method of claim 12, wherein the color format includes one of the following: 4:2:0 format or 4:4:4 format.

17. The method of claim 12, wherein the scheme is applied to at least one of the chromaticity component or the luminance component.

18. The method of any one of claims 1 to 17, wherein the conversion comprises encoding the current video block into the bitstream.

19. The method according to any one of claims 1 to 17, wherein the conversion comprises decoding the current video block from the bitstream.

20. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 19.

21. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 19.

22. A non-transitory computer-readable recording medium storing a bitstream of video, the bitstream of video being generated by a method performed by means of a video processing apparatus, wherein the method comprises: Based on strip type determination information, the information relates to the application of a deblocking or filtering process based on sub-block boundaries to the current video block of the video; as well as The bit stream is generated based on the information. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how the sub-block boundary-based deblocking or filtering process is applied.

23. A method for storing a bitstream of video, comprising: Based on strip type determination information, the information relates to the application of a deblocking or filtering process based on sub-block boundaries to the current video block of the video; The bit stream is generated based on the information; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The information indicates at least one of the following: whether a sub-block boundary-based deblocking or filtering process is applied, or how the sub-block boundary-based deblocking or filtering process is applied.

24. A non-transitory computer-readable recording medium for storing a bitstream of video, the bitstream of video being generated by a method performed by means of a video processing apparatus, wherein the method includes: The use of at least one of the following in interleaved affine prediction or template matching affine is determined based on at least one of the following: strip type, frame type, or picture order count (POC) distance between a strip or frame and its nearest reference frame; and The bitstream is generated based on the use of at least one of the interleaved affine prediction or the template matching affine.

25. A method for storing a bitstream of video, comprising: The use of at least one of the following is determined based on at least one of the following: strip type, frame type, or picture order count (POC) distance between a strip or frame and its nearest reference frame; The bitstream is generated based on the use of at least one of the interleaved affine prediction or the template matching affine; as well as The bitstream is stored in a non-transitory computer-readable recording medium.

26. A non-transitory computer-readable recording medium for storing a bitstream of video, the bitstream of video being generated by a method performed by means of a video processing apparatus, wherein the method comprises: The use of interleaved prediction for affine modes is determined based on sequence resolution or frame resolution. as well as The bitstream is generated based on the interleaving prediction.

27. A method for storing a bitstream of video, comprising: The use of interleaved prediction for affine modes is determined based on sequence resolution or frame resolution. The bitstream is generated based on the use of the interleaving prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.