Method and device for video processing and medium
By using interleaved affine prediction and OBMC technology, the problem of low efficiency in existing video encoding and decoding is solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202480025173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-13
- Filing Date
- 2024-04-12
- Publication Date
- 2025-11-11
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved.
By employing interleaved affine prediction and codec unit-level overlapping block motion compensation (OBMC) techniques, encoding and decoding efficiency is improved by determining the conversion between video blocks and bitstreams.
It improves the efficiency of video encoding and decoding.
Smart Images

Figure CN120937358A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to interlaced affine prediction. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between a target video block and a video bitstream, determining the use of interleaved affine prediction based on at least one of a first encoding / decoding tool or encoding / decoding information of the target video block; and performing the conversion based on the use of interleaved affine prediction. The method according to the first aspect of this disclosure improves encoding / decoding efficiency.
[0005] In a second aspect, another method for video processing is proposed. This method includes: a conversion between a target video block and a video bitstream for the purpose of video processing; determining the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level based on the encoding / decoding mode of the target video block; and performing the conversion based on the use of the CU-level OBMC. The method according to the second aspect of this disclosure improves encoding / decoding efficiency.
[0006] In a third aspect, another method for video processing is proposed. This method includes: for a conversion between a target video block and a video bitstream, determining the motion fields of at least two codec units encoded and decoded prior to the target video block; determining at least one regressive affine candidate for the target video block based on the motion fields of the at least two codec units; and performing a conversion based on the at least one regressive affine candidate. The method according to the third aspect of this disclosure improves encoding and decoding efficiency.
[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.
[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first, second, or third aspect of this disclosure.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the use of interleaved affine prediction based on at least one of the encoding / decoding information of a first encoding / decoding tool or a target video block of the video; and generating a bitstream based on the use of interleaved affine prediction.
[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the use of interleaved affine prediction based on at least one of the encoding / decoding information of a first encoding / decoding tool or a target video block of the video; generating a bitstream based on the use of interleaved affine prediction; and storing the bitstream in the non-transitory computer-readable recording medium.
[0011] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level based on the encoding / decoding mode of a target video block; and generating a bitstream based on the use of the CU-level OBMC.
[0012] In the ninth aspect, a method for storing video bitstreams is proposed. The method includes: determining the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level based on the encoding / decoding mode of the target video block; generating a bitstream based on the use of CU-level OBMC; and storing the bitstream in a non-transitory computer-readable recording medium.
[0013] In a tenth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the motion fields of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate for the target video block based on the motion fields of the at least two codec units; and generating a bitstream based on the at least one regressive affine candidate.
[0014] In the eleventh aspect, a method for storing a bitstream of video is proposed. The method includes: determining the motion fields of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate for the target video block based on the motion fields of the at least two codec units; generating a bitstream based on the at least one regressive affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0015] This summary is provided to present, in a simplified form, the selected concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0016] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0017] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; Figure 5 This shows the positions of non-adjacent candidates in the ECM; Figure 6 An affine motion model based on control points is shown; Figure 7 An example affine MVF for each sub-block is shown; Figure 8 The position of the inherited affine motion predictor is shown; Figure 9 This demonstrates the inheritance of control point motion vectors; Figure 10 The locations of candidate positions for constructive affine Merge patterns are shown; Figure 11 The spatial nearest neighbor used to derive the affine Merge candidate is shown; Figure 12 This shows candidates for constructive affine merges, ranging from non-nearest neighbors. Figure 13An example of generating HAPC is shown; Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 15 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 16 The template and the corresponding reference template are shown; Figure 17 The template and reference template of a block with sub-block motion information using the motion information of the current block are shown; Figure 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; Figure 19 An example of interleaved prediction is shown; Figures 20A-20G An exemplary partitioning pattern for 16×16 blocks is shown; Figures 21A-21D An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas; Figures 22A-22C An example is shown of deriving the MV of a partitioning pattern from another partitioning pattern; Figures 23A-23C An example of selecting a partitioning mode based on block dimensions is shown; Figure 24A and Figure 24B An example is shown that derives the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of another partitioning pattern. Figure 25 An example of a CU-level OBMC is shown; Figure 26 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 27 A flowchart of another method for video processing according to an embodiment of the present disclosure is shown; Figure 28 A flowchart of another method for video processing according to embodiments of the present disclosure is shown; and Figure 29 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0018] In all the accompanying drawings, the same or similar reference numerals usually indicate the same or similar elements. Detailed Implementation
[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0020] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0022] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0024] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. The codec images are codec representations of images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0029] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0035] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based on error results, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-frame joint prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0037] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0041] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0042] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0056] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0057] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of a video block, to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the block size of the frames(multiple) and / or stripes(multiple) used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term "video processing" includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0062] 1. Brief Overview This disclosure relates to video encoding and decoding technologies. Specifically, it pertains to affine motion prediction methods in video encoding and decoding. This idea can be applied alone or in various combinations to any standard or non-standard video codec.
[0063] 2. Introduction The exponential growth of multimedia data has posed significant challenges to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations have jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), H.265 / HEVC, and the latest VVC standard. Since H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding.
[0064] 2.1. MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames and is an indispensable component of hybrid video coding and decoding frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded or decoded, thereby transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the intermediate MV of a neighboring block as the MVP. In H.265 / HEVC, a contention mechanism is involved, where rate-distortion optimization (RDO) selects the best MVP from multiple candidates. Specifically, Advanced MVP (AMVP) mode and Merge mode are designed using different motion information signaling strategies. With AMVP mode, the reference index, the MVP candidate index referencing the AMVP candidate list, and the motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list. The details of the construction process for these two modes are described below.
[0065] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporal neighbors to the left and top, removing redundant candidates, and adding zero vectors to make the candidate list a fixed length. Figure 4 This illustrates the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list. For the spatial motion vector candidate derivation, the final result is based on blocks located as follows: Figure 1 The motion vectors of five blocks at different locations are shown to derive two motion vector candidates. The five neighboring blocks located at B0, B1, B2 and A0, A1 are classified into two groups: group A includes the three spatially adjacent blocks above, and group B includes the two spatially adjacent blocks to the left. The two motion vector candidates are derived separately in a predefined order using the first available candidates from group A and group B. For temporal motion vector candidate derivation, a motion vector candidate is derived by sequentially checking based on two different co-locations (lower right (C0) and center (C1)), as follows: Figure 1 As shown. To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list.
[0066] Figure 5 The positions of non-adjacent candidates in the ECM are shown.
[0067] Merge Mode: Similar to AMVP mode, the MVP candidate list for Merge mode also consists of spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, up to four candidates are selected in the order A1, B1, B0, A0, and B2. For temporal Merge Candidate (TMVP) derivation, up to one candidate is selected from two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.
[0068] In VVC, the Merge pattern construction process is further improved by introducing a history-based MVP (HMVP), where the HMVP incorporates motion information from previously encoded / decoded blocks that may be geographically distant from the current block. In VVC, HMVP Merge candidates are appended to the Merge list after the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, the table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy. Whenever a non-sub-block inter-frame encoding / decoding CU exists, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0069] During the standardization of VVC, a non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... Figure 2 As shown.
[0070] 2.2. Affine Motion Compensation Prediction In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 6 Affine motion models based on control points are shown, such as (a) a 4-parameter affine model and (b) a 6-parameter affine model. Figure 6 As shown, the affine motion field of a block is described by motion information from two control point motion vectors (4 parameters) or three control point motion vectors (6 parameters).
[0071] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1).
[0072] For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2).
[0073] Where (mv0x, mv0y) is the motion vector of the top left control point, (mv1x, mv1y) is the motion vector of the top right control point, and (mv2x, mv2y) is the motion vector of the bottom left control point.
[0074] To simplify motion compensation prediction, a block-based affine transformation prediction is applied. Figure 7 An example affine MVF for each sub-block is shown. To derive the motion vector for each 4×4 luma sub-block, as follows... Figure 7 As shown, the motion vector of the center sample point of each sub-block is calculated according to the above formula and rounded to 1 / 16 pixel precision. Then, a motion-compensated interpolation filter is applied, and the prediction for each sub-block is generated using the derived motion vector. The sub-block size of the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper left luminance sub-block and the lower right luminance sub-block in the corresponding 8×8 luminance region.
[0075] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0076] 2.2.1. Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and a signal transmission index indicates which CPVM candidate should be used for the current CU. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs.
[0077] – Constructive affine Merge candidate CPMVP derived using translational MV of neighboring CUs.
[0078] – Zero MV.
[0079] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of the neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Figure 8 The positions of inherited affine motion predictors are shown. Candidate blocks are as follows: Figure 8 As shown. For the predictor on the left, the scan order is A0->A1, and for the predictor above, the scan order is B0->B1->B2. Only the first inherited candidate is selected from each side. No pruning check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. Figure 9 The inheritance of control point motion vectors is shown. For example... Figure 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, according to Calculate the two CPMVs of the current CU. When block A is encoded and decoded using a 6-parameter affine model, according to... Calculate the three CPMVs of the current CU.
[0080] Constructive affine candidates mean building candidates by combining the translational motion information of each control point's neighbors. The motion information of the control points is derived from... Figure 10 The spatial and temporal nearest neighbors shown are derived. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, the block is checked by B2->B3->A2, and the MV of the first available block is used. For CPMV2, the block is checked by B1->B0, and for CPMV3, the block is checked by A1->A0. If available, the TMVP is used as CPMV4.
[0081] After obtaining the motion signatures (MVs) of the four control points, affine merge candidates are constructed based on this motion information. The following combinations of control point MVs are used in sequence for construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.
[0082] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0083] Figure 10 The locations of candidate positions for constructive affine Merge patterns are shown.
[0084] After the inherited affine merge candidates and the constructed affine merge candidates have been checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0085] 2.2.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted sub-CPVM is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.
[0086] – A constructive affine AMVP candidate CPMVP derived using the translation MV of neighboring CUs.
[0087] – Translation MV from the neighboring CU.
[0088] – Zero MV.
[0089] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference picture as the current block are considered. No pruning is applied when inserting inherited affine motion predictors into the candidate list.
[0090] Constructive AMVP candidates are from Figure 10 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. In the checking order, the first block that is inter-coded and has the same reference picture as the current CU is used. If the current CU is encoded in a 4-parameter affine mode and both mv0 and mv1 are available, they are added as candidates in the affine AMVP list. If the current CU is encoded in a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable.
[0091] Figure 11 Spatial nearest neighbors for deriving affine Merge candidates are shown: (a) for deriving inherited affine Merge candidates, and (b) for deriving constructed affine Merge candidates.
[0092] If the affine AMVP list still has fewer than two candidates after inserting valid inherited and constructed affine AMVP candidates, then mv0, mv1, and mv2 will be added sequentially as translational MVs to predict all control point MVs of the current CU, when available. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.
[0093] 2.2.3. New Affine Candidate Derivation Method in ECM-8.0 ECM-6.0 integrates three additional affine Merge and AMVP candidate derivation methods: non-adjacent spatial domain candidates, historical parameter-based candidates, regression-based affine candidates, and pixel-based affine motion compensation.
[0094] 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, the concept of non-adjacent airspace neighbors was studied to provide candidates for affine Merge and affine AMVP. The pattern for obtaining non-adjacent airspace candidates is... Figure 11 As shown in the diagram, similar to non-adjacent regular merge candidates, the distance between non-adjacent spatial candidates and the current codec block is also defined based on the width and height of the current CU.
[0095] Figure 11 Motion information of non-adjacent spatial neighbors in the current block is used to generate additional inherited and constructed affine merge candidates. Specifically, to generate inherited candidates, non-adjacent spatial neighbors are examined based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbor encoded in affine mode from each side (e.g., left and top) of the current block is included. Figure 11 As shown in (a), checks on the left and top neighbors are performed from bottom to top and from right to left, respectively. For constructive candidates, as... Figure 11 As shown in (b), the positions of a non-adjacent spatial neighbor on the left and a non-adjacent spatial neighbor above are first determined independently; then, the position of the upper left neighbor can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and above. The motion information of the three non-adjacent neighbors is used to form a CPMV at the upper left (A), upper right (B), and lower left (C) of the virtual block, which is projected onto the current CU to generate corresponding constructivist candidates, such as... Figure 12 As shown.
[0096] 2.2.3.2. Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to inherit from previously affine-encoded blocks (which may not be adjacent to the current block). A History Parameter Table (HPT) is established. HPT entries store sets of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entries in the HPT are categorized by reference lists and reference indices. Each reference list in the HPT supports 5 reference indices. The HPT category (denoted as HPTCat) is calculated in a formulaic manner as follows: (3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of 7 entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU using the reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0097] Candidates based on historical affine parameters (HAPC) are derived from... Figure 13 The MV is derived from the set of affine parameters in the corresponding entries stored in the HPT, represented as the neighboring 4×4 blocks of A0, A1, B0, B1, or B2. The MV of the neighboring 4×4 blocks is used as the base MV. The MV of the current block at position (x, y) is calculated in a formulaic manner as follows: (4) Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 blocks, and (xbase, ybase) represents the center position of the nearest 4×4 blocks. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0098] Figure 13An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are directly extracted from an entry in the class HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT, along with the center position of A0 (as the base position) and the MV of block A0 (as the base MV), are used to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as regular Merge candidates. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate newly added regular Merge candidates.
[0099] Figure 13 An example of generating HAPC is shown.
[0100] 2.2.3.3. Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0101] Previously encoded and decoded affine CUs can be identified by scanning non-adjacent positions and the affine HMVP table. Figure 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. Figure 14 As shown, information about the neighboring sub-blocks of the current CU is obtained from the 4x4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.
[0102] For each affine CU, a maximum of two affine candidates can be derived: one with neighboring subblock information and one without. All candidates generated by linear regression are pruned and merged into a single candidate subgroup. When ARMC is enabled, an ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.
[0103] 2.2.3.4. Pixel-based Affine Motion Compensation Using pixel-based affine motion compensation, when OBMC is not applied, the minimum affine sub-block size for the luma component is set to 1x1, and the minimum sub-block size for the chroma component is always set to 1x1.
[0104] 2.3. Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between the template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and the block in the reference image (i.e., the same size as the template). Figure 15 This illustrates template matching execution over the search area surrounding the initial MV. (Example) Figure 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU.
[0105] In AMVP mode, MVP candidates are determined based on template matching error. The candidate that minimizes the difference between the current block and the reference block template is selected, and then TM performs MV refinement only on that specific MVP candidate. TM starts with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) and refines the MVP candidate using an iterative diamond search within a search range of [-8, +8] pixels. Depending on the AMVR mode, the AMVP candidate can be further refined: a cross search is performed using full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel precision. This search process ensures that after the TM process, the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode.
[0106] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. Depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode), TMMerge can proceed up to 1 / 8-pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy. Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. When both BM and TM are enabled for a CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0107] 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, we propose Adaptive Reordering of Merge Candidates (ARMC) to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0108] The reordering method is applied to the regular Merge pattern, the Template Matching (TM) Merge pattern, and the Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.
[0109] After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0110] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and its corresponding reference template. Figure 16 The template and its corresponding reference template are shown. For example... Figure 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information of the current block. When the merge candidate utilizes bidirectional prediction, the reference samples of the merge candidate's template are also generated through bidirectional prediction.
[0111] For sub-block size equal to The sub-block-based Merge candidate has an upper template that includes several sub-templates of size Wsub×K, and a left template that includes several sub-templates of size K×Hsub. Figure 17 This shows a template and a reference template for a block that uses the motion information of its child blocks. For example... Figure 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0112] 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. SbTMVP uses the same co-located images as TMVP. SbTMVP differs from TMVP primarily in two aspects. First, SbTMVP enables motion prediction at the sub-CU level, while TMVP predicts CU-level motion. Second, compared to TMVP extracting temporal motion vectors (MVs) from co-located blocks in the co-located image (co-located blocks are the lower right or center blocks relative to the current CU), SbTMVP applies motion shifting before extracting temporal motion information from the co-located image. This motion shifting is achieved by reusing the MV from a spatially neighboring block of the current CU.
[0113] Figure 18 The derivation process for the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first obtained. If any MV in reference list 0 and reference list 1 points to the same frame, the corresponding MV will be marked as a motion shift. Otherwise, zero MV will be used as a motion shift.
[0114] Once the motion displacement is determined, a specified region within the same frame is used to derive the sub-block level motion field. Assuming... Figure 18 As shown, the motion of A1 is used for motion shifting. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-located image is extracted to provide motion information, where first an MV scaling operation is performed to align the reference frame of the temporal motion vector with the reference frame of the current CU.
[0115] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion predictions for the current CU. This sub-CU-level MVP candidate list includes the motion field generated by the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. After performing template matching-based reordering, multiple AFFINE candidates are included in the list, with those AFFINE candidates with lower costs placed earlier.
[0116] 2.6. Block Removal Process in VVC 8.6.2 Deblocking Filter Process 8.6.2.1 Overview The input to this process is the reconstructed image before removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0117] The output of this process is the modified reconstructed image after removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0118] The vertical edges in the image are first filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the image are filtered. Vertical and horizontal edges in the CTB of each CTU are processed individually on a codec unit basis. The vertical edges of the codec blocks within a codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically towards the right-hand side of the codec block. The horizontal edges of the codec blocks within a codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically towards the bottom of the codec block.
[0119] Note – Although the filtering process is specified in terms of images in this specification, the filtering process can be implemented on a per-encoder / decoder basis with equivalent results, provided that the decoder properly considers the processing dependency order in order to produce the same output value.
[0120] The deblocking filtering process is applied to all encoded and transformed block edges of the image, except for the following types of edges: – The edge at the boundary of the image; – When loop_filter_across_tiles_enabled_flag equals 0, the edge that coincides with the tile boundary; – When tile_group_loop_filter_across_tile_groups_enabled_flag equals 0 or tile_group_deblocking_filter_disabled_flag equals 1, the edge that coincides with the top or left boundary of the tile group; – Edges within a tile group where tile_group_deblocking_filter_disabled_flag is equal to 1; - Edges that do not correspond to the 8x8 sample grid boundary of the considered component; – Edges on both sides of the edge are predicted using the chromaticity components within the frame; – Edges of chroma transform blocks that are not part of the edges of the associated transform unit.
[0121] [Editor's Note: Once the chips are integrated, the syntax will be adjusted.] The edge type (vertical or horizontal) is represented by the variable edgeType, as specified in Table 8.17.
[0122] Table 8.17 – Names associated with edgeType
[0123] The following applies when the tile_group_deblocking_filter_disabled_flag of the current tile group is equal to 0: – The variable treeType is deduced as follows: – If tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, then treeType is set to equal DUAL_TREE_LUMA.
[0124] Otherwise, treeType is set to equal SINGLE_TREE.
[0125] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_VER are taken as inputs, and the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.
[0126] – Horizontal edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the modified deblocked reconstructed image (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_HOR are taken as inputs, and the modified deblocked reconstructed image (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.
[0127] – When tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, the following applies: – The variable treeType is set to equal DUAL_TREE_CHROMA.
[0128] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, with the variable treeType, the reconstructed image before deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_VER as input, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) as output.
[0129] – Horizontal edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, with the variable treeType, the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_HOR as input, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) as output.
[0130] 8.6.2.2 Deblocking filter process for one direction The input to this process is: – A variable treeType that specifies whether to use a single tree (SINGLE_TREE) or a dual tree to split the CTU, and when using a dual tree, whether to process the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) currently; – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the reconstructed picture before deblocking, i.e., the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr; – A variable edgeType that specifies whether to filter vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR).
[0131] The output of this process is the modified reconstructed picture after deblocking, i.e.: – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr.
[0132] For each coding unit with a coded block width of log2CbW, a coded block height of log2CbH, and the position of the top-left sample of the coded block being (xCb, yCb), when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edges are filtered in the following ordered steps: 1. The coded block width nCbW is set to be equal to 1<<log2CbW, and the coded block height nCbH is set to be equal to 1<<log2CbH.
[0133] 2. The variable filterEdgeFlag is derived as follows: – If edgeType is equal to EDGE_VER and one or more of the following conditions are true, then filterEdgeFlag is set to be equal to 0: – The left boundary of the current coded block is the left boundary of the picture.
[0134] – The left boundary of the current codec block is the left boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.
[0135] – The left boundary of the current codec block is the left boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0136] Otherwise, if edgeType equals EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set to 0: – The upper boundary of the current luminance codec block is the upper boundary of the image.
[0137] – The upper boundary of the current codec block is the upper boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.
[0138] – The upper boundary of the current codec block is the upper boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0139] Otherwise, filterEdgeFlag is set to 1.
[0140] [Editor's Note: Once the chips are integrated, the syntax will be adjusted.] 3. All elements of the two-dimensional (nCbW)x(nCbH) array edgeFlags are initialized to 0.
[0141] 4. The derivation process for the transform block boundary specified in Item 8.6.2.3 is invoked, where the position (xB0, yB0) is set to equal to (0, 0), the block width nTbW is set to equal to nCbW, the block height nTbH is set to equal to nCbH, the variables treeType, filterEdgeFlag, array edgeFlags, and edgeType are taken as inputs, and the modified array edgeFlags is taken as output.
[0142] 5. The derivation process for the codec subblock boundary specified in Item 8.6.2.4 is invoked, with the position (xCb, yCb), codec block width nCbW, codec block height nCbH, array edgeFlags, and variable edgeType as inputs, and the modified array edgeFlags as output.
[0143] 6. The image sample array recPicture is derived as follows: – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, then recPicture is set to equal to the array of reconstructed luminance image samples recPictureL before deblocking.
[0144] – Otherwise (treeType equals DUAL_TREE_CHROMA), recPicture is set to equal the reconstructed chroma image sample array recPictureCb before deblocking.
[0145] 7. The derivation process for the boundary filter intensity specified in item 8.6.2.5 is invoked, with the image sample array recPicture, the luminance position (xCb, yCb), the codec block width nCbW, the codec block height nCbH, the variable edgeType, and the array edgeFlags as inputs, and the (nCbW)x(nCbH) array verBs as outputs.
[0146] 8. The edge filtering process is invoked as follows: – If edgeType equals EDGE_VER, the vertical edge filtering procedure of the codec unit specified in entry 8.6.2.6.1 is invoked, with the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array verBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.
[0147] Otherwise, if edgeType equals EDGE_HOR, the horizontal edge filtering procedure of the codec unit specified in entry 8.6.2.6.2 is invoked, with the variable treeType, the modified reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array horBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.
[0148] 8.6.2.3 Derivation of the Transform Block Boundary The input to this process is: – Specifies the position (xB0, yB0) of the top-left sample of the current block relative to the top-left sample of the current codec block; – The variable nTbW specifies the width of the current block; – A variable nTbH that specifies the height of the current block; – Specifies the treeType variable to split the CTU, whether to use a single tree (SINGLE_TREE) or a dual tree, and when using a dual tree, whether to process the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA). – Variable filterEdgeFlag; – Two-dimensional (nCbW)x(nCbH) array edgeFlags; – Specifies the variable edgeType for filtering either the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR).
[0149] The output of this process is a modified two-dimensional (nCbW)x(nCbH) array edgeFlags.
[0150] The maximum transform block size, maxTbSize, is derived as follows: maxTbSize = (treeType= =DUAL_TREE_CHROMA) ? MaxTbSizeY / 2 :MaxTbSizeY(8 862) Based on maxTbSize, the following applies: – If nTbW is greater than maxTbSize or nTbH is greater than maxTbSize, then the following ordered steps apply.
[0151] 1. The variables newTbW and newTbH are derived as follows: newTbW = ( nTbW>maxTbSize ) ? ( nTbW / 2 ) : nTbW(8 863) newTbH = ( nTbH>maxTbSize ) ? ( nTbH / 2 ) : nTbH(8 864).
[0152] 2. The derivation procedure for the transform block boundary specified in this entry is invoked, with the position (xB0, yB0), the variable nTbW set to be equal to newTbW and the variable nTbH set to be equal to newTbH, the variable filterEdgeFlag, the array edgeFlags and the variable edgeType as inputs, and the output being a modified version of the array edgeFlags.
[0153] 3. If nTbW is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal (xTb0 + newTbW, yTb0), the variable nTbW is set to equal newTbW, the variable nTbH is set to equal newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0154] 4. If nTbH is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0, yTb0 + newTbH), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0155] 5. If nTbW is greater than maxTbSize and nTbH is greater than maxTbSize, then the derivation procedure for the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0 + newTbW, yTb0 + newTbH), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0156] – Otherwise, the following applies: – If edgeType equals EDGE_VER, then the value of edgeFlags[xB0][yB0 + k] (k = 0..nTbH-1) is derived as follows: – If xB0 equals 0, then edgeFlags[ xB0 ][ yB0 + k ] is set to equal filterEdgeFlag.
[0157] Otherwise, edgeFlags[xB0][yB0 + k] is set to 1.
[0158] Otherwise (edgeType equals EDGE_HOR), the value of edgeFlags[xB0 + k][yB0] (k = 0..nTbW-1) is derived as follows: – If yB0 equals 0, then edgeFlags[ xB0 + k ][ yB0 ] is set to equal filterEdgeFlag.
[0159] Otherwise, edgeFlags[xB0 + k][yB0] is set to 1.
[0160] 8.6.2.4 Derivation of Encoding / Decoding Sub-Block Boundaries The input to this process is: – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – The variable nCbW specifies the width of the current codec block; – The variable nCbH specifies the height of the current codec block; – Two-dimensional (nCbW)x(nCbH) array edgeFlags; – Specifies the variable edgeType for filtering either the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR).
[0161] The output of this process is a modified two-dimensional (nCbW)x(nCbH) array edgeFlags.
[0162] The number of horizontal codec sub-blocks, numSbX, and the number of vertical codec sub-blocks, numSbY, are derived as follows: – If CupredMode[xCb][yCb] == MODE_INTRA, then numSbX and numSbY are both set to 1.
[0163] Otherwise, numSbX and numSbY are set to equal NumSbX[xCb][yCb] and NumSbY[xCb][yCb], respectively.
[0164] The following applies depending on the value of edgeType: – If edgeType equals EDGE_VER and numSbX is greater than 1, then the following applies to i = 1..min( (nCbW / 8 ) - 1, numSbX - 1), k = 0..nCbH - 1: .
[0165] Otherwise, if edgeType equals EDGE_HOR and numSbY is greater than 1, the following applies to j = 1..min( ( nCbH / 8 ) - 1, numSbY - 1 ), k = 0..nCbW - 1: .
[0166] 8.6.2.5 Derivation of Boundary Filter Strength The input to this process is: – Image sample array recPicture; – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – The variable nCbW specifies the width of the current codec block; – The variable nCbH specifies the height of the current codec block; – Specifies the variable edgeType to indicate whether filtering is applied to the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR); – A two-dimensional (nCbW)x(nCbH) array edgeFlags.
[0167] The output of this process is a two-dimensional (nCbW)x(nCbH) array bS that specifies the boundary filter strength.
[0168] The variables xDi, yDj, xN, and yN are derived as follows: – If edgeType equals EDGE_VER, then xDi is set to equal to (i<<3), yDj is set to equal to (j<<2), xN is set to equal to Max(0, (nCbW / 8) - 1), and yN is set to equal to (nCbH / 4) - 1.
[0169] Otherwise (edgeType equals EDGE_HOR), xDi is set to equal to (i<<2), yDj is set to equal to (j<<3), xN is set to equal to (nCbW / 4) - 1, and yN is set to equal to Max( 0, ( nCbH / 8 ) - 1 ).
[0170] For xDi where i = 0..xN and yDj where j = 0..yN, the following applies: – If edgeFlags[xDi][yDj] equals 0, then the variable bS[xDi][yDj] is set to equal 0.
[0171] – Otherwise, the following applies: – The sample values p0 and q0 are derived as follows: – If edgeType equals EDGE_VER, then p0 is set to equal recPicture [ xCb + xDi - 1 ][ yCb + yDj ], and q0 is set to equal recPicture [ xCb + xDi ][ yCb + yDj ].
[0172] Otherwise (edgeType equals EDGE_HOR), p0 is set to equal recPicture [ xCb + xDi ][yCb + yDj - 1 ], and q0 is set to equal recPicture [ xCb + xDi ][ yCb + yDj ].
[0173] – The variable bS[ xDi ][ yDj ] is derived as follows: – If sample p0 or q0 is in the codec block of a codec unit encoded in intra-prediction mode, then bS[xDi][yDj] is set to equal 2.
[0174] Otherwise, if the block edge is also the transform block edge, and sample p0 or q0 is in a transform block containing one or more non-zero transform coefficient levels, then bS[xDi][yDj] is set to equal to 1.
[0175] Otherwise, bS[xDi][yDj] is set to 1 if one or more of the following conditions are true: – For prediction of a codec subblock containing sample p0, use a different reference picture or a different number of motion vectors than for prediction of a codec subblock containing sample q0.
[0176] Note 1 – Determining whether the reference pictures used for two codec subblocks are the same or different is based solely on which pictures are referenced, without considering whether the indices in reference picture list 0 or reference picture list 1 are used to form the prediction, and also without considering whether the index positions within the reference picture lists are different.
[0177] Note 2 – The number of motion vectors for the codec sub-block used to predict the top-left sample coverage (xSb, ySb) is equal to PredFlagL0[xSb][ySb] + PredFlagL1[xSb][ySb].
[0178] – A motion vector is used to predict the codec subblock containing sample p0, and a motion vector is used to predict the codec subblock containing sample q0, and the absolute difference between the horizontal or vertical components of the motion vectors used is greater than or equal to 4 (in quarter-luminance samples).
[0179] – Two motion vectors and two different reference images are used to predict the codec subblock containing sample p0, and two motion vectors and the same two reference images are used to predict the codec subblock containing sample q0. When predicting the two codec subblocks, the absolute difference between the horizontal or vertical components of the two motion vectors used for the same reference image is greater than or equal to 4 (in quarter-luminance samples).
[0180] – Two motion vectors for the same reference image are used to predict a codec subblock containing sample p0, and two motion vectors for the same reference image are used to predict a codec subblock containing sample q0, provided that both of the following conditions are true: – The absolute difference between the horizontal or vertical components of the motion vector in list 0 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in list 1 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples).
[0181] – The absolute difference between the horizontal or vertical components of the motion vector in List 0 used to predict the codec subblock containing sample p0 and the motion vector in List 1 used to predict the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in List 1 used to predict the codec subblock containing sample p0 and the motion vector in List 0 used to predict the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples).
[0182] Otherwise, the variable bS[ xDi ][ yDj ] is set to 0.
[0183] 8.6.2.6 Edge Filtering Process 8.6.2.6.1 Vertical Edge Filtering Process The input to this process is: – Specifies the treeType variable to split the CTU, whether to use a single tree (SINGLE_TREE) or a dual tree, and when using a dual tree, whether to process the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA). – When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, reconstruct the image before the block, i.e., the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – The variable nCbW specifies the width of the current codec block; – The variable nCbH specifies the height of the current codec block.
[0184] The output of this process is the modified reconstructed image after removing the blocks, i.e.: – When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.
[0185] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the edge filtering process in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal (nCbH / 4) - 1.
[0186] 2. For xDk of k = 0..xN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: – When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for the block edge as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0187] b. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, where the image sample array recPicture is set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.
[0188] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the edge filtering process in the chroma codec block of the current codec unit consists of the following ordered steps: 1. Variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal Max(0, (nCbH / 8) - 1).
[0189] 2. The variable edgeSpacing is set to equal to 8 / SubWidthC.
[0190] 3. The variable edgeSections is set to equal to .
[0191] 4. For k = 0..xN, the equality is... The following applies to xDk and yDm where m = 0..edgeSections are equal to m<<2: - when When xCb / SubWidthC + xDk equals 2 and (((xCb / SubWidthC + xDk)>>3)<<3) equals xCb / SubWidthC + xDk, the following ordered steps apply: a. The filtering process for the chroma block edges as specified in entry 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.
[0192] b. The filtering process for the chroma block edges as specified in entry 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.
[0193] When treeType equals DUAL_TREE_CHROMA, the edge filtering process in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal (nCbH / 4) - 1.
[0194] 2. For xDk of k = 0..xN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: – When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0195] b. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, where the image sample array recPicture is set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as inputs, and the modified chroma image sample array recPictureCb is taken as output.
[0196] c. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, where the image sample array recPicture is set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as inputs, and the modified chroma image sample array recPictureCr is taken as output.
[0197] 8.6.2.6.2 Horizontal Edge Filtering Process The input to this process is: – Specifies the treeType variable to split the CTU, whether to use a single tree (SINGLE_TREE) or a dual tree, and when using a dual tree, whether to process the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA). – When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the reconstructed image before removing the blocks, i.e., the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – The variable nCbW specifies the width of the current codec block; – The variable nCbH specifies the height of the current codec block.
[0198] The output of this process is the modified reconstructed image after removing the blocks, i.e.: – When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; – When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.
[0199] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the edge filtering process in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8) - 1), and xN is set to equal (nCbW / 4) - 1.
[0200] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: – When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for the block edge as specified in Item 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0201] b. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, where the image sample array recPicture is set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dEp, dEp and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.
[0202] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the edge filtering process in the chroma codec block of the current codec unit consists of the following ordered steps: 1. Variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal Max(0, (nCbH / 8) - 1).
[0203] 2. The variable edgeSpacing is set to equal to 8 / SubHeightC.
[0204] 3. The variable edgeSections is set to equal to .
[0205] 4. For m = 0..yN, the equality is... The following applies to yDm and xDk where k = 0..edgeSections equals k<<2: - when When yCb / SubHeightC + yDm equals 2 and (((yCb / SubHeightC + yDm)>>3)<<3) equals yCb / SubHeightC + yDm, the following ordered steps apply: a. The filtering process for the chroma block edges as specified in entry 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.
[0206] b. The filtering process for the chroma block edges as specified in entry 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.
[0207] When treeType equals DUAL_TREE_CHROMA, the edge filtering process in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8) - 1), and xN is set to equal (nCbW / 4) - 1.
[0208] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: – When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0209] b. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, wherein the image sample array recPicture is set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCb is taken as output.
[0210] c. The block edge filtering process specified in Item 8.6.2.6.4 is invoked, where the image sample array recPicture is set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCr is taken as output.
[0211] 8.6.2.6.3 Decision-making process at block edges The input to this process is: – Specifies the treeType variable to split the CTU, whether to use a single tree (SINGLE_TREE) or a dual tree, and when using a dual tree, whether to process the luminance component (DUAL_TREE_LUMA) or the chrominance component (DUAL_TREE_CHROMA). – Image sample array recPicture; – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – Specifies the position (xBl, yBl) of the top-left sample of the current block relative to the top-left sample of the current codec block; – Specifies the variable edgeType to indicate whether filtering is applied to the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR); – The variable bS specifies the boundary filter strength; – The variable bD specifies the bit depth of the current component.
[0212] The output of this process is: – Includes decision variables dE, dEp, and dEq; – Variable tC.
[0213] If edgeType equals EDGE_VER, then the sample values pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ](8 867) pi,k = recPictureL[ xCb + xBl - i - 1 ][ yCb + yBl + k ](8 868) Otherwise (edgeType equals EDGE_HOR), the sample values pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPicture[ xCb + xBl + k ][ yCb + yBl + i ](8 869) pi,k = recPicture[ xCb + xBl + k ][ yCb + yBl - i - 1 ](8 870) The variable qpOffset is derived as follows: – If sps_ladf_enabled_flag equals 1 and treeType equals SINGLE_TREE or DUAL_TREE_LUMA, then the following applies: – The variable lumaLevel for reconstructing the brightness level is derived as follows: lumaLevel = ( ( p0,0 + p0,3 + q0,0 + q0,3 )>>2 ),(8 871) – The variable qpOffset is set to equal to sps_ladf_lowest_interval_qp_offset and modified as follows: For (i = 0; i <sps_num_ladf_intervals_minus2 + 1; i++ ) { if( lumaLevel>SpsLadfIntervalLowerBound[ i + 1 ] ) qpOffset = sps_ladf_qp_offset[i](8 872) else break } Otherwise (treeType equals DUAL_TREE_CHROMA), qpOffset is set to 0.
[0214] The variables QpQ and QpP are derived as follows: – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, then QpQ and QpP are set to the QpY value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.
[0215] Otherwise (treeType equals DUAL_TREE_CHROMA), QpQ and QpP are set to the QpC value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.
[0216] The variable qP is derived as follows: qP = ( ( QpQ + QpP + 1 )>>1 ) + qpOffset(8 873) The value of variable β′ is determined based on the quantization parameter Q, as specified in Table 8.18, derived as follows: Q = Clip3( 0, 63, qP + ( tile_group_beta_offset_div2<<1 ) )(8 874) Where tile_group_beta_offset_div2 is the value of the syntax element tile_group_beta_offset_div2 that includes the patch group of sample points q0,0.
[0217] The variable β is derived as follows:
[0218] The value of variable tC′ is determined based on the quantization parameter Q, as specified in Table 8.18, derived as follows:
[0219] Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that includes the sample points q0,0.
[0220] The variable tC is derived as follows:
[0221] The following applies depending on the value of edgeType: – If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:
[0222] dpq0 = dp0 + dq0(8 882) dpq3 = dp3 + dq3(8 883) dp = dp0 + dp3(8 884) dq = dq0 + dq3(8 885) d = dpq0 + dpq3(8 886) 2. The variables dE, dEp, and dEq are set to 0.
[0223] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.
[0224] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in entry 8.6.2.6.6 is invoked, where the sample values p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.
[0225] c. The variable dpq is set to equal 2. dpq3.
[0226] d. For the sample location (xCb + xBl, yCb + yBl + 3), the decision process for the sample as specified in entry 8.6.2.6.6 is invoked, where the sample values p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.
[0227] e. The variable dE is set to equal 1.
[0228] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.
[0229] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.
[0230] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.
[0231] – Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:
[0232] 2. The variables dE, dEp, and dEq are set to 0.
[0233] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.
[0234] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in entry 8.6.2.6.6 is invoked, where the sample values p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.
[0235] c. The variable dpq is set to equal 2. dpq3.
[0236] d. For sample locations (xCb + xBl + 3, yCb + yBl), the decision process for the sample as specified in entry 8.6.2.6.6 is invoked, where the sample values p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.
[0237] e. The variable dE is set to equal 1.
[0238] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.
[0239] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.
[0240] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.
[0241] Table 8.18 – Deriving threshold variables β′ and tC′ from input Q
[0242] 8.6.2.6.4 Filtering process for block edges The input to this process is: – Image sample array recPicture; – Specifies the position (xCb, yCb) of the top-left corner sample of the current codec block relative to the top-left corner sample of the current image; – Specifies the position (xBl, yBl) of the top-left sample of the current block relative to the top-left sample of the current codec block; – Specifies the variable edgeType to indicate whether filtering is applied to the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR); – Includes decision variables dE, dEp, and dEq; – Variable tC.
[0243] The output of this process is a modified image sample array recPicture.
[0244] The following applies depending on the value of edgeType: – If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The sample values pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ](8 896) pi,k = recPictureL[ xCb + xBl - i - 1 ][ yCb + yBl + k ](8 897) 2. When dE is not equal to 0, for each sample point location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: a. The sample filtering procedure specified in Item 8.6.2.6.7 is invoked, where the sample values pi,k,qi,k at i = 0..3, the positions (xPi,yPi) at i = 0..2 are set to equal (xCb + xBl - i - 1, yCb + yBl + k) and (xQi,yQi) are set to equal (xCb + xBl + i, yCb + yBl + k), the decision dE, the variables dEp and dEq, and the variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample values pi' and qj' are taken as outputs.
[0245] b. When nDp is greater than 0, the filtered sample value pi' of i = 0..nDp - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl - i - 1 ][ yCb + yBl + k ]= pi'(8 898) c. When nDq is greater than 0, the filtered sample value qj' of j = 0..nDq - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + j ][ yCb + yBl + k ]= qj'(8 899) – Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The sample values pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl + i ](8 900) pi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl - i - 1 ](8 901) 2. When dE is not equal to 0, for each sample point location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: a. The sample filtering procedure specified in Item 8.6.2.6.7 is invoked, where the sample values pi,k,qi,k at i = 0..3, the positions (xPi, yPi) at i = 0..2 are set to equal (xCb + xBl + k, yCb + yBl - i - 1) and (xQi, yQi) are set to equal (xCb + xBl + k, yCb + yBl + i), the decision dE, the variables dEp and dEq, and the variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample values pi' and qj' are taken as outputs.
[0246] b. When nDp is greater than 0, the filtered sample value pi' of i = 0..nDp - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl - i - 1 ]= pi'(8 902) c. When nDq is greater than 0, the filtered sample value qj' of j = 0..nDq - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl + j ]= qj'(8 903) 8.6.2.6.5 Filtering process for chroma block edges This procedure is invoked only when ChromaArrayType is not equal to 0.
[0247] The input to this process is: – Chroma image sample array s′; – Specifies the chroma position (xCb, yCb) of the top-left corner sample of the current chroma codec block relative to the top-left corner chroma sample of the current image; – Specifies the chroma position (xBl, yBl) of the top-left sample point of the current chroma block relative to the top-left sample point of the current chroma codec block; – Specifies the variable edgeType to indicate whether filtering is applied to the vertical edge (EDGE_VER) or the horizontal edge (EDGE_HOR); – Specifies the offset of the image-level color metric parameter cQpPicOffset.
[0248] The output of this process is the modified chroma image sample array s′.
[0249] If edgeType equals EDGE_VER, then the values of pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + i ][ yCb + yBl + k ](8 904) pi,k = s′[ xCb + xBl - i - 1 ][ yCb + yBl + k ](8 905) Otherwise (edgeType equals EDGE_HOR), the sample values pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + k ][ yCb + yBl + i ](8 906) pi,k = s′[ xCb + xBl + k ][ yCb + yBl - i - 1 ](8 907) The variables QpQ and QpP are set to be equal to the QpY value of the codec unit that includes codec blocks containing samples q0,0 and p0,0, respectively.
[0250] If ChromaArrayType equals 1, then the variable QpC is determined according to the following derivation of the index qPi as specified in Table 8.15: qPi = ( ( QpQ + QpP + 1 )>>1 ) + cQpPicOffset(8 908) Otherwise (ChromaArrayType is greater than 1), the variable QpC is set to equal Min(qPi, 63).
[0251] Note – The variable cQpPicOffset provides adjustments to the values of pps_cb_qp_offset or pps_cr_qp_offset based on whether the filtered chroma components are Cb or Cr components. However, to avoid the need for adjustments within the image, the filtering process does not include adjustments to the values of tile_group_cb_qp_offset or tile_group_cr_qp_offset.
[0252] The value of variable tC′ is determined based on the colorimetric quantization parameter Q, as specified in Table 8.18, derived as follows: Q = Clip3( 0, 65, QpC + 2 + ( tile_group_tc_offset_div2<<1 ) )(8 909) Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that includes the sample points q0,0.
[0253] The variable tC is derived as follows: tC = tC′ ( 1<<( BitDepthC - 8 ) )(8 910) The following applies depending on the value of edgeType: – If edgeType equals EDGE_VER, then for each sample location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: 1. The chromaticity sample filtering process specified in Item 8.6.2.6.8 is invoked, where the sample values pi,k and qi,k with i = 0..1, the positions (xCb + xBl - 1, yCb + yBl + k) and (xCb + xBl, yCb + yBl + k), and the variable tC are taken as inputs, and the filtered sample values p0′ and q0′ are taken as outputs.
[0254] 2. Replace the corresponding samples in the sample array s' with the filtered sample values p0′ and q0′, as shown below: s′[ xCb + xBl ][ yCb + yBl + k ]= q0′(8 911) s′[ xCb + xBl - 1 ][ yCb + yBl + k ]= p0′(8 912) Otherwise (edgeType equals EDGE_HOR), for each sample location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: 1. The chromaticity sample filtering process specified in Item 8.6.2.6.8 is invoked, where the sample values pi,k and qi,k with i = 0..1, the positions (xCb + xBl + k, yCb + yBl - 1) and (xCb + xBl + k, yCb + yBl), and the variable tC are taken as inputs, and the filtered sample values p0′ and q0′ are taken as outputs.
[0255] 2. Replace the corresponding samples in the sample array s' with the filtered sample values p0′ and q0′, as shown below: s′[ xCb + xBl + k ][ yCb + yBl ]= q0′(8 913) s′[ xCb + xBl + k ][ yCb + yBl - 1 ] = p0′(8 914) 8.6.2.6.6 Decision-making process for sample points The input to this process is: – Sample values p0, p3, q0, and q3; – Variables dpq, β, and tC.
[0256] The output of this process is the variable dSam, which contains the decision.
[0257] The variable dSam is defined as follows: – If dpq is less than (β>>2), Abs(p3 - p0) + Abs(q0 - q3) is less than (β>>3), and Abs(p0 - q0) is less than (5). If tC + 1 )>>1, then dSam is set to equal to 1.
[0258] Otherwise, dSam is set to 0.
[0259] 8.6.2.6.7 Sample Filtering Process The input to this process is: – Sample values pi and qi for i = 0..3; – The positions (xPi, yPi) and (xQi, yQi) of pi and qi for i = 0..2; – Variable dE; – dEp and dEq are variables that respectively contain the decision of filtering samples p1 and q1; – Variable tC.
[0260] The output of this process is: – The number of filtered samples, nDp and nDq; – Filtered sample values pi′ and qj′ for i = 0..nDp - 1 and j = 0..nDq - 1.
[0261] Based on the value of dE, the following applies: – If variable dE equals 2, then nDp and nDq are both set to equal 3, and the following strong filtering applies:
[0262] Otherwise, nDp and nDq are both set to 0, and the following weak filtering applies: – The following applies:
[0263] - when Less than tC At 10:00, the following sequential steps apply: – The filtered sample values p0′ and q0′ are defined as follows:
[0264] – When dEp equals 1, the filtered sample value p1′ is defined as follows:
[0265]
[0266] – When dEq equals 1, the filtered sample value q1′ is defined as follows:
[0267]
[0268] – nDp is set to equal dEp + 1, and nDq is set to equal dEq + 1.
[0269] nDp is set to 0 when nDp is greater than 0 and one or more of the following conditions are true: – pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 ][ yP0 ] equals 1.
[0270] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.
[0271] nDq is set to 0 when nDq is greater than 0 and one or more of the following conditions are true: – pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xQ0 ][ yQ0 ] equals 1.
[0272] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.
[0273] 8.6.2.6.8 Filtering process for chromaticity samples This procedure is invoked only when ChromaArrayType is not equal to 0.
[0274] The input to this process is: – Colorimetric sample values pi and qi for i = 0..1; – The chromaticity positions of p0 and q0 are (xP0, yP0) and (xQ0, yQ0); – Variable tC.
[0275] The output of this process is the filtered sample values p0′ and q0′.
[0276] The filtered sample values p0′ and q0′ are derived as follows:
[0277] The filtered sample value p0′ is replaced by the corresponding input sample value p0 when one or more of the following conditions are true: – pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 SubWidthC ][yP0 SubHeightC equals 1.
[0278] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.
[0279] The filtered sample value q0′ is replaced by the corresponding input sample value q0 when one or more of the following conditions are true: – pcm_loop_filter_disabled_flag equals 1, and pcm_flag[xQ0] SubWidthC ][yQ0 SubHeightC equals 1.
[0280] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.
[0281] 8.6.3 Sample point adaptive compensation process 8.6.3.1 Overview The input to this process is the array recPictureL of reconstructed image samples before adaptive compensation, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0282] The output of this process is a modified array of reconstructed image samples, saoPictureL, after adaptive compensation, and arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0.
[0283] This process is performed on top of CTB after the deblocking filter process for the decoded image is completed.
[0284] The sample values in the modified reconstructed picture sample array saoPictureL, and the sample values in the arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0, are initially set to be equal to the sample values in the reconstructed picture sample array recPictureL, and the sample values in the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0, respectively.
[0285] For each CTU with CTB position (rx, ry), where rx = 0..PicWidthInCtbsY - 1 and ry = 0..PicHeightInCtbsY - 1, the following applies: – When the tile_group_sao_luma_flag of the current slice group is equal to 1, the CTB modification process specified in Clause 8.6.3.2 is called, where recPicture is set to be equal to recPictureL, cIdx is set to be equal to 0, (rx, ry), and both nCtbSw and nCtbSh are set to be equal to CtbSizeY as input, and the modified luma picture sample array saoPictureL is used as output.
[0286] – When ChromaArrayType is not equal to 0 and the tile_group_sao_chroma_flag of the current slice group is equal to 1, the CTB modification process specified in Clause 8.6.3.2 is called, where recPicture is set to be equal to recPictureCb, cIdx is set to be equal to 1, (rx, ry), nCtbSw is set to be equal to (1 << CtbLog2SizeY) / SubWidthC, and nCtbSh is set to be equal to (1 << CtbLog2SizeY) / SubHeightC as input, and the modified chroma picture sample array saoPictureCb is used as output.
[0287] – When ChromaArrayType is not equal to 0 and the tile_group_sao_chroma_flag of the current slice group is equal to 1, the CTB modification process specified in item 8.6.3.2 is called, where recPicture is set to be equal to recPictureCr, cIdx is set to be equal to 2, (rx, ry), nCtbSw is set to be equal to (1 << CtbLog2SizeY) / SubWidthC, and nCtbSh is set to be equal to (1 << CtbLog2SizeY) / SubHeightC as inputs, and the modified chroma picture sample array saoPictureCr is used as the output.
[0288] 8.6.3.2 CTB Modification Process The inputs to this process are: – The picture sample array recPicture of color component cIdx; – The variable cIdx specifying the color component index; – A pair of variables (rx, ry) specifying the CTB position; – The CTB width nCtbSw and height nCtbSh.
[0289] The output of this process is the modified picture sample array saoPicture for color component cIdx.
[0290] The variable bitDepth is derived as follows: – If cIdx is equal to 0, bitDepth is set to be equal to BitDepthY.
[0291] – Otherwise, bitDepth is set to be equal to BitDepthC.
[0292] The position (xCtb, yCtb) of the top - left sample of the current CTB for color component cIdx relative to the top - left sample of the current picture component cIdx is derived as follows: ( xCtb, yCtb ) = ( rx nCtbSw, ry nCtbSh )(8 932) The sample positions within the current CTB are derived as follows: ( xSi, ySj ) = ( xCtb + i, yCtb + j )(8 933) ( xYi, yYj ) = ( cIdx == 0 )? ( xSi, ySj ) : ( xSi SubWidthC, ySj SubHeightC )(8 934) For all sample locations (xSi, ySj) and (xYi, yYj) of i = 0..nCtbSw - 1 and j = 0..nCtbSh - 1, the following applies based on the values of pcm_loop_filter_disabled_flag, pcm_flag[xYi][yYj], and cu_transquant_bypass_flag of the codec unit that includes the codec block covering recPicture[xSi][ySj]. – If one or more of the following conditions are true, then saoPicture[ xSi ][ ySj ] will not be modified: – pcm_loop_filter_disabled_flag and pcm_flag[ xYi ][ yYj ] are both equal to 1.
[0293] – cu_transquant_bypass_flag equals 1.
[0294] – SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 0.
[0295] [Editor's Note: The highlighted parts have been modified based on future decisions, including transformations / quantization bypasses.] Otherwise, if SaoTypeIdx[cIdx][rx][ry] equals 2, then the following ordered steps apply: 1. The values of hPos[k] and vPos[k] for k = 0..1 are based on the values of SaoEoClass[cIdx][rx][ry] as specified in Table 8.19.
[0296] 2. The variable edgeIdx is derived as follows: – The modified sample point locations (xSik′, ySjk′) and (xYik′, yYjk′) are derived as follows: ( xSik′, ySjk′ ) = ( xSi + hPos[ k ], ySj + vPos[ k ])(8 935) ( xYik′, yYjk′ ) = ( cIdx = = 0 ) ? ( xSik′, ySjk′ ) : ( xSik′ SubWidthC, ySjk′ SubHeightC )(8 936) – edgeIdx is set to 0 if one or more of the following conditions are true for all sample locations (xSik′, ySjk′) and (xYik′, yYjk′) at k = 0..1: – The sample point at location (xSik′, ySjk′) is outside the image boundary.
[0297] – The sample points at location (xSik′, ySjk′) belong to different patch groups, and one of the following two conditions is true: – MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSi ][ ySj ] belongs is equal to 0.
[0298] – MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSik′ ][ ySjk′ ] belongs is equal to 0.
[0299] – loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at position (xSik′, ySjk′) belongs to a different tile.
[0300] [Editor's Note: When merging slices that do not contain slice groups, modify the highlighted portion.] Otherwise, edgeIdx is derived as follows: – The following applies: edgeIdx = 2 + Sign( recPicture[ xSi ][ ySj ]- recPicture[ xSi + hPos[ 0 ] ][ ySj + vPos[ 0 ] ]) + Sign( recPicture[ xSi ][ ySj ]- recPicture[ xSi + hPos[ 1 ] ][ ySj +vPos[ 1 ] ])(8 937) – When edgeIdx is equal to 0, 1, or 2, edgeIdx is modified as follows: edgeIdx = ( edgeIdx == 2 ) ? 0 : ( edgeIdx + 1 )(8 938) 3. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) - 1, recPicture[xSi ][ ySj ]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ edgeIdx ])(8 939) – Otherwise (SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 1), the following ordered steps apply: 1. The variable bandShift is set to equal bitDepth - 5.
[0301] 2. The variable saoLeftClass is set to equal sao_band_position[ cIdx ][ rx ][ ry ].
[0302] 3. The list `bandTable` is defined with 32 elements, and all elements are initially set to 0. Then, its four elements (indicating the starting position of the band for the explicit offset) are modified as follows: for( k = 0; k<4; k++ ) bandTable[ ( k + saoLeftClass )&31 ] = k + 1(8 940) 4. The variable bandIdx is set to equal bandTable[recPicture[xSi][ySj]>>bandShift].
[0303] 5. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) - 1, recPicture[xSi ][ ySj ]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ bandIdx ])(8 941) Table 8.19 – Specifications for hPos and vPos based on the sample adaptive compensation category
[0304] 2.7. OBMC in ECM When applying OBMC, motion information from neighboring blocks with weighted predictions is used to refine the top and left boundary samples of the CU.
[0305] The following conditions should not be used for OBMC: • When OBMC is disabled at the SPS level.
[0306] • When the current block has intra-frame mode or IBC mode.
[0307] • When applying LIC to the current block.
[0308] • When the current luminance block area is less than or equal to 32.
[0309] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary samples using motion information from neighboring sub-blocks. Enabled for sub-block-based codec tools: • Affine AMVP mode; • Affine Merge pattern and sub-block-based temporal motion vector prediction (SbTMVP). • Bilateral matching based on sub-blocks.
[0310] When using OBMC mode in CIIP mode with LMCS, inter-frame blending is performed before LMCS mapping of inter-frame samples. LMCS is applied to the blended inter-frame samples, which are combined with intra-frame samples to which LMCS is applied in CIIP mode.
[0311] ,in This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by the motion of neighboring blocks in the original domain, and and It's the weight.
[0312] 2.8. Interleaved Prediction To overcome the problems in sub-block-based prediction, interleaving prediction in video encoding and decoding is proposed.
[0313] Using interleaved prediction, a block is divided into sub-blocks with more than one partitioning pattern. A partitioning pattern is defined as the way the block is divided into sub-blocks, including the size and position of the sub-blocks. For each partitioning pattern, the motion information of each sub-block can be derived based on the partitioning pattern to generate the corresponding predicted block. Therefore, even for a single prediction direction, multiple predicted blocks can be generated using multiple partitioning patterns. Alternatively, only one partitioning pattern can be applied for each prediction direction.
[0314] Suppose there are X partitioning patterns, and the X predicted blocks (denoted as P0, P1, ..., PX-1) of the current block are generated through sub-block-based predictions with the X partitioning patterns. The final prediction (denoted as P) of the current block can be generated as follows: (15) Where (x, y) are the coordinates of the pixels in the block, and These are the weighted values of Pi. Without loss of generality, assume... , where N is a non-negative value. Figure 19 An example of interleaved prediction with two partitioning modes is shown.
[0315] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0316] Use of interleaving prediction for different codec tools 1. Interleaving prediction can be applied to one, some, or all of the codecs that have sub-block-based prediction. In one example, affine prediction applies interleaving prediction, while other codecs with sub-block-based prediction (such as ATMVP, STMVP, FRUC, and BIO) do not. In another example, affine, ATMVP, and STMVP apply interleaving prediction.
[0317] Definition of partitioning mode 2. The partitioning pattern can have different shapes, sizes, or positions of the sub-blocks. In one example, the partitioning pattern could result in irregular sub-block sizes. Figures 20A-20GSeveral exemplary partitioning patterns for 16×16 blocks are shown. Figure 20A In this context, blocks are divided into 4×4 sub-blocks, as in JEM; Figure 20B In the middle, the block is divided into 8×8 sub-blocks; Figure 20C and Figure 20D In the middle, the block is divided into 8x4 sub-blocks and 4x8 sub-blocks; in Figure 20E and Figure 20F In the block, the block is also divided into 4×4 sub-blocks, but in different positions. Pixels at the block boundary that cannot be divided into complete 4×4 sub-blocks can be divided into smaller sub-blocks of size 2×4, 4×2, or 2×2, such as... Figure 20E As shown, they can also be merged into adjacent 4×4 sub-blocks to form larger sub-blocks of size 6×4, 4×6, or 6×6, such as... Figure 20F As shown; in Figure 20G In this model, blocks are also divided into 8×8 sub-blocks, but in different locations. Pixels at the block boundaries that cannot be divided into complete 8×8 sub-blocks can be divided into smaller sub-blocks of size 8×4, 4×8, or 4×4.
[0318] 3. The shape and size of sub-blocks in sub-block-based predictions may depend on the shape and / or size of the encoded / decoded blocks and / or the encoded / decoded block information (e.g., affine or ATMVP modes).
[0319] a. In one example, when the size of the current block is M×N, the size of the sub-block is 4×N (or 8×N, etc.).
[0320] b. In one example, when the size of the current block is M×N, the size of the child block is M×4 (or M×8, etc.).
[0321] c. In one example, when the size of the current block is M×N and M>N, the size of the sub-block is A×B (where A>B), such as 8×4; otherwise, the size of the sub-block is B×A, such as 4×8.
[0322] d. In one example, assuming the current block size is M×N, the sub-block size is A×B when M×N<=T (or Min(M, N)<=T, or Max(M, N)<=T, etc.), and the sub-block size is C×D when M×N>T (or Min(M, N)>T, or Max(M, N)>T, etc.), where A<=C and B<=D. For example, if M×N<=256, the sub-block size is 4×4; otherwise, the sub-block size is 8×8.
[0323] Enabling / disabling interleaving prediction and the encoding / decoding process of interleaving prediction 4. Whether to apply interleaved prediction depends on the inter-frame prediction direction.
[0324] a. In one example, interleaved forecasts can be applied to bidirectional forecasts but not to unidirectional forecasts.
[0325] b. In one example, when multiple hypotheses are applied, interleaved predictions can be applied to a prediction direction when there is more than one reference block.
[0326] 5. How to apply interleaved prediction depends on the inter-frame prediction direction.
[0327] a. In one example, a bidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two different partitioning patterns for two different reference lists. In one example, when predicting from reference list 0 (L0), the block is divided into 4×8 sub-blocks, such as... Figure 20D As shown, however, when predicting from reference list 1 (L1), the block is divided into 8×4 sub-blocks, as follows: Figure 20C As shown. And the final prediction P is calculated as ZSA. (16) Where P0 and P1 are predictions from L0 and LFR1, respectively. w0 and w1 are weighted values for L0 and L1, respectively. Without loss of generality, assume (Where N is a non-negative integer value).
[0328] b. In one example, a unidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two or more distinct partitioning patterns. For example, the prediction PL for list L (L = 0 or 1) is calculated as follows: (17) Where XL is the number of partitioning patterns for list L; It is a prediction generated using the i-th partitioning pattern, and yes The weighted value. For example, XL is 2. Using the 0th partitioning pattern, the block is divided into 4×8 sub-blocks, such as... Figure 20D As shown. Using the first partitioning pattern, the block is divided into 8×4 sub-blocks, as follows. Figure 20C As shown.
[0329] c. In one example, a bidirectional prediction block with sub-block-based predictions is considered as a combination of two unidirectional prediction blocks from L0 and L1, respectively. As described in the example above, predictions from each list can be derived. The final prediction P can be calculated as... (18) The parameters a and b are two additional weights applied to the two internal prediction blocks. In one example, both a and b are equal to 1.
[0330] d. In one example, for a block encoded using multiple hypotheses, there may be more than one prediction block generated with different partitioning patterns for each prediction direction (or reference image list). Multiple prediction blocks can be used to generate a final version with additional weights applied. In one example, the additional weights can be set to 1 / M, where M is the total number of prediction blocks generated.
[0331] 6. Whether and how interleaving prediction can be applied from the encoder to the decoder at the sequence level, picture level, view level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, PU level, or TU level, or slice level, or slice group level, or region level (which may contain multiple CU / PU / TU / LCU)). This information can be transmitted via signaling in the first block of the sequence parameter set (SPS), view parameter set (VPS), picture parameter set (PPS), stripe header (SH), picture header, sequence header, or slice level or slice group level, CTU (also known as LCU), CU, PU, TU, or region.
[0332] a. In one example, interleaving prediction is implicitly applied to existing sub-block methods such as ATMVP, STMVP, FRUC, BIO, or affine. In this example, no additional signaling cost is required.
[0333] b. In another example, new sub-block Merge candidates generated by interleaving prediction are inserted into the Merge list, such as interleaving prediction + ATMVP, interleaving prediction + STMVP, interleaving prediction + FRUC, etc.
[0334] c. In one example, a signaling flag can be used to indicate whether interleaving prediction is used. In one example, if the current block is affine inter-frame encoded, a signaling flag is used to indicate whether interleaving prediction is used.
[0335] d. In one example, if the current block is encoded and decoded using affine Merge and one-way prediction is applied, a signaling flag can be used to indicate whether interleaved prediction is used.
[0336] e. In one example, if the current block is encoded and decoded using an affine Merge codec, a signaling flag can be used to indicate whether interleaving prediction is used.
[0337] f. In one example, if the current block is encoded / decoded using affine Merge and one-way prediction is applied, interleaved prediction can always be used.
[0338] g. In one example, if the current block is encoded / decoded using affine Merge, interleaved prediction can always be used.
[0339] h. In one example, a flag indicating whether interleaving prediction is used can be inherited without signal transmission.
[0340] i. In one example, inheritance can be used if the current block is encoded or decoded using affine Merge.
[0341] ii. In one example, the flag can be inherited from the flag of a neighboring block that inherits the affine model.
[0342] iii. In one example, the flag is inherited from a predefined neighboring block (such as the left or top neighboring block).
[0343] iv. In one example, the flag can be inherited from the first encountered affine-coded neighboring block.
[0344] v. In one example, if no neighboring blocks are affine encoded, the flag can be presumed to be zero.
[0345] vi. In one example, the flag can only be inherited if one-way prediction is applied to the current block.
[0346] vii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to inherit are in the same CTU.
[0347] viii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to inherit are in the same CTU line.
[0348] ix. In one example, when the affine model is derived from a temporal neighbor block, the flag may not be inherited from the flags of the neighbor block.
[0349] x. In one example, the flag may not be inherited from the flags of neighboring blocks located in the same LCU or LCU line or video data processing unit (such as 64x64 or 128x128).
[0350] xi. In one example, how flags are transmitted and / or derived via signaling may depend on the block dimension of the current block and / or encoded / decoded information.
[0351] i. In one example, if the reference image is the current image, then interleaved prediction is not applied.
[0352] i. In one example, if the reference image is the current image, a flag indicating whether interleaving prediction is used is not transmitted through the signal.
[0353] Weighted values 7. The weighting value w is fixed. For example, in equations (15) and (16) .
[0354] 8. The weighting value can depend on the position and the partitioning pattern, that is, for different (x, y). It may differ. Alternatively, the weighting may further depend on the codec tool based on sub-block prediction (e.g., affine or ATMVP) and / or other codec information (e.g., skip or non-skip modes and / or MV information, etc.).
[0355] 9. Weights can be transferred from the encoder to the decoder at the sequence level, picture level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, PU level, or region level (which can contain multiple CUs / PUs / TUs / LCUs)). They can be transmitted via signaling in the first block of the sequence parameter set (SPS), picture parameter set (PPS), stripe header (SH), CTU (also known as LCU), CU or PU, or region.
[0356] a. In an alternative solution, in addition, for some blocks, weights can be inherited from spatial and / or temporal neighboring blocks.
[0357] Partial Interleaving Prediction 10. In one embodiment, interleaved predictions are applied to a portion of the current block. Prediction samples at some locations are calculated as a weighted sum of two or more sub-block-based predictions. Prediction samples at other locations are not. For example, these prediction samples are copied from sub-block-based predictions with a specific partitioning pattern. Figures 21A-21D An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas.
[0358] a. In one example, the current block is predicted using sub-block-based predictions P1 and P2, respectively, with partitioning patterns D0 and D1. The final prediction is calculated as P = w0 × P0 + w1 × P1. At some locations, w0 ≠ 0 and w1 ≠ 0. But at other locations, w0 = 1 and w1 = 0, meaning that interleaved predictions are not applied at these locations.
[0359] b. In one example, such as Figure 21A As shown, interleaving prediction is not applied to the four corner blocks.
[0360] c. In one example, such as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns.
[0361] d. In one example, such as Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0362] e. In one example, such as Figure 21D As shown, interleaving prediction is not applied to the topmost sub-row, the bottommost sub-row, the leftmost sub-column, and the rightmost sub-column.
[0363] f. In one example, whether and how partial interleaving prediction is applied can depend on the size / shape of the current block.
[0364] i. For example, if the size of the current block meets certain conditions, the interleaving prediction is applied to the entire block; otherwise, the interleaving prediction is applied to a portion (or parts) of the block. Conditions include, but are not limited to: (assuming the width and height of the current block are W and H respectively, and T, T1, and T2 are integer values): 1. W>=T1 and H>=T2; 2. W <= T1 and H <= T2; 3. W>=T1 or H>=T2; 4. W <= T1 or H <= T2; 5. W + H >= T; 6. W + H <= T; 7. W×H>=T; 8. W×H<=T.
[0365] ii. For example, if W>=H, then as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0366] iii. For example, if W>H, then as Figure 21B As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... Figure 21C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0367] g. It is proposed that whether and how interleaved predictions are applied can differ for different regions within a block.
[0368] i. For example, assume that the current block is predicted by sub-block based prediction P1 and P2 with partition patterns D0 and D1 respectively. The final prediction is calculated as P(x, y) = w0 × P0(x, y) + w1 × P1(x, y). If the position (x, y) belongs to a sub-block of dimension S0 × H0 with partition pattern D0; and belongs to a sub-block S1 × H1 with partition pattern D1. If one or more of the following conditions are met, then set w0 = 1 and w1 = 0. (That is, the interleaved prediction is not applied to this position).
[0369] 1. S1 < T1; 2. H1 < T2; 3. S1 < T1 and H1 < T2; 4. S1 < T1 or H1 < T2; T1 and T2 are integers. For example, T1 = T2 = 4.
[0370] Encoder issues 11. In one embodiment, the interleaved prediction is not applied during the motion estimation (ME) process.
[0371] a. For example, the interleaved prediction is not applied during the ME process for 6-parameter affine prediction.
[0372] b. For example, if the size of the current block meets specific conditions, such as (assuming the width and height of the current block are W and H respectively, and T, T1, T2 are integer values), then the interleaved prediction is not applied during the ME process: i. W >= T1 and H >= T2; ii. W <= T1 and H <= T2; iii. W >= T1 or H >= T2; iv. W <= T1 or H <= T2; v. W + H >= T; vi. W + H <= T; vii. W × H >= T; viii. W × H <= T.
[0373] c. For example, if the current block is partitioned from a parent block and the parent block does not select the affine mode at the encoder, then the interleaved prediction is not applied during the ME process.
[0374] i. Alternatively, if the current block is partitioned from a parent block and the parent block does not select the affine mode at the encoder, then the affine mode is not checked at the encoder.
[0375] MV derivation In the following discussion, SatShift(x, n) is defined as
[0376] Shift(x, n) is defined as Shift(x, n) = (x + offset0) >> n.
[0377] In one example, offset0 and / or offset1 are set to (1 < 0.05).<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.
[0378] 12. The MV of each sub-block within a partitioning pattern can be derived directly from an affine model (such as using equation (1)), or it can be derived from the MV of a sub-block within another partitioning pattern.
[0379] a. In one example, the MV of sub-block B with partition pattern 0 can be derived from the MV of all or some sub-blocks within partition pattern 1 that overlap with sub-block B.
[0380] b. Figures 22A-22C An example is shown. Figure 22A In this process, the MV1(x,y) of a specific sub-block within partitioning mode 1 will be derived. Figure 22B The diagram shows partitioning pattern 0 (solid line) and partitioning pattern 1 (dashed line) within the block, indicating that there are four sub-blocks within partitioning pattern 0 that overlap with specific sub-blocks within partitioning pattern 1. Figure 22C The diagram shows four MVs for four sub-blocks that overlap with specific sub-blocks in partitioning pattern 1 within partitioning pattern 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2), and MV0(x+2,y+2). MV1(x,y) is then derived from MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2), and MV0(x+2,y+2).
[0381] c. Suppose that the MV' of a sub-block within partitioning pattern 1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks within partitioning pattern 0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.
[0382] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.
[0383] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.
[0384] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.
[0385] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.
[0386] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.
[0387] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.
[0388] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.
[0389] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.
[0390] x. with Figures 22A-22C For example, MV1(x,y) can be derived as:
[0391]
[0392] 13. The choice of partitioning mode can depend on the width and height of the current block. Figures 23A-23C An example of selecting a partitioning mode based on block dimensions is shown.
[0393] a. For example, if width > T1 and height > T2 (e.g., T1 = T2 = 4), then both partitioning modes are selected. Figure 23A Examples of two partitioning patterns are shown.
[0394] b. For example, if the height is less than or equal to T2 (e.g., T2 = 4), then the other two partitioning modes are selected. Figure 23B Examples of two partitioning patterns are shown.
[0395] c. For example, if the width is less than or equal to T1 (e.g., T1 = 4), then two more partitioning patterns are selected. Figure 23C Examples of two partitioning patterns are shown.
[0396] 14. The MV of each sub-block within a partitioning mode of a color component C1 can be derived from the MV of the sub-block within another partitioning mode of another color component C0.
[0397] a. For example, C1 refers to a color component that is encoded / decoded after another color component, such as Cb or Cr or U or V or R or B.
[0398] b. For example, C0 refers to a color component that is encoded / decoded before another color component, such as Y or G.
[0399] c. In one example, how to derive the MV of a sub-block within a partitioning mode of a color component from the MV of a sub-block within another partitioning mode of another color component can depend on the color format, such as 4:2:0, or 4:2:2, or 4:4:4.
[0400] d. In one example, the MV of a sub-block B in a color component C1 with partitioning pattern C1Pt (t=0 or 1) can be derived from the MV of all or some sub-blocks in a partitioning pattern C0Pr (r=0 or 1) that overlaps with sub-block B, after scaling or scaling coordinates according to the color format.
[0401] i. In one example, C0Pr is always equal to C0P0.
[0402] e. Figure 24A and Figure 24B An example is shown. Figure 24A and Figure 24B This example illustrates the derivation of the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of a different partitioning pattern. The color format is 4:2:0. The MV of a sub-block in the Cb component is derived from the MV of a sub-block in the Y component.
[0403] i. at Figure 24A On the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 0 will be derived. Figure 24A The right side shows four Y-blocks within partition mode 0. When scaled at a 2:1 ratio, these four Y-blocks overlap with sub-block B (Cb). Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').
[0404] ii. In Figure 24BOn the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 1 will be derived. Figure 24B The right side shows four Y-blocks within partition mode 0. When scaled at a 2:1 ratio, these four Y-blocks overlap with sub-block B (Cb). Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').
[0405] f. Suppose that the MV' of a sub-block of color component C1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks of color component C0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.
[0406] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.
[0407] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.
[0408] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.
[0409] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.
[0410] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.
[0411] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.
[0412] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.
[0413] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.
[0414] x. with Figure 24A and Figure 24B For example, MVCbt(x',y') t= 0 or 1 can be derived as:
[0415]
[0416] Interleaved forecasting for bidirectional forecasting 15. When interleaved prediction is applied to bidirectional prediction, the following methods can be used to save the increased internal bit depth due to different weights: a. For list X (X = 0 or 1), PX(x, y) = Shift(W0(x,y)). PX0(x,y) + W1(x,y) PX1(x,y), SW), where PX(x,y) is the prediction for list X, and PX0(x,y) and PX1(x,y) are the predictions for list X with partitioning mode 0 and partitioning mode 1. W0 and W1 are integers representing the weighted values of the interleaved predictions, and SW represents the precision of the weighted values.
[0417] b. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB), where Wb0 and Wb1 are integers used in weighted bidirectional forecasting, and SWB is the precision. When there is no weighted bidirectional forecasting, Wb0=Wb1=SWB=1.
[0418] c. In some embodiments, PX0(x,y) and PX1(x,y) can maintain the accuracy of the interpolation filter. For example, they can be 16-bit unsigned integers. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y)+ Wb1(x,y) P1(x,y), SWB+PB), where PB is the additional precision from the interpolation filter, for example, PB=6. In this case, W0(x,y) PX0(x,y) or W1(x,y) PX1(x,y) can exceed 16 bits. It is proposed that PX0(x,y) and PX1(x,y) are first right-shifted to a lower precision to avoid exceeding 16 bits.
[0419] i. For example, for a list X (X = 0 or 1), PX(x, y) = Shift(W0(x,y)). PLX0(x,y) + W1(x,y) PLX1(x,y), SW), where PLX0(x,y) = Shift( PX0(x,y), M), PLX1(x,y) = Shift(PX1(x,y), M). The final prediction is derived as P(x,y) = Shift( Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB+PB-M). For example, M is set to 2 or 3.
[0420] d. The above method can also be applied to other bidirectional forecasting methods with different weighting factors for two reference forecast blocks, such as generalized bidirectional forecasting (GBi, where the weights can be, for example, 3 / 8, 5 / 8) and weighted forecasting (where the weights can be very large values).
[0421] e. The above method can also be applied to other multi-hypothesis one-way or two-way forecasting methods with different weighting factors for different reference forecast blocks.
[0422] Block size dependence 16. Whether and / or how to apply interleaving prediction can depend on the block width W and height H.
[0423] a. In one example, whether and / or how to apply interleaved prediction can depend on the size of the VPDU (Video Processing Data Unit, which typically represents the maximum permissible block size to be processed in a hardware design).
[0424] b. In one example, the original prediction method can be utilized when interleaving prediction is disabled for a specific block dimension (or a block with specific encoded / decoded information).
[0425] i. Alternatively, affine mode can be disabled directly for this type of block.
[0426] c. In one example, interleaved prediction cannot be used when W > T1 and H > T2. For example, T1 = T2 = 64; d. In one example, interleaved prediction cannot be used when W > T1 or H > T2. For example, T1 = T2 = 64; e. In one example, when W Interleaving prediction cannot be used when H > T. For example, when T = 64 64; f. In one example, when W < T1 and H < T2, interlaced prediction cannot be used. For example, T1 = T2 = 16; g. In one example, when W < T1 or H > T2, interlaced prediction cannot be used. For example, T1 = T2 = 16; h. In one example, when W H < T, interlaced prediction cannot be used. For example, T = 16 16.
[0427] i. In one example, for sub - blocks not located at block boundaries (e.g., codec units), the interlaced affine for that sub - block can be disabled. Alternatively, in addition, the prediction result using the original affine prediction method can be directly used as the final prediction for that sub - block.
[0428] j. In one example, when W > T1 and H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 64; k. In one example, when W > T1 or H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 64; l. In one example, when W H > T, interlaced prediction is used in a different way. For example, T = 64 64; m. In one example, when W < T1 and H < T2, interlaced prediction is used in a different way. For example, T1 = T2 = 16; n. In one example, when W < T1 or H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 16; o. In one example, when W H < T, interlaced prediction is used in a different way. For example, T = 16 16.
[0429] p. In one example, when H > X (e.g., H equals 128, X = 64), interlaced prediction is not applied to the samples of sub - blocks belonging to the upper W (H / 2) partition and the lower W (H / 2) partition of the current block.
[0430] q. In one example, when W > X (e.g., W equals 128, X = 64), interlaced prediction is not applied to the samples of sub - blocks belonging to the left (W / 2) H partition and the right (W / 2) H partition of the current block.
[0431] r. In one example, when W>X and H>Y (e.g., W=H=128, X=Y=64), i. Interleaving predictions are not applied to left (W / 2) segments that span the current block. H-segmentation and right (W / 2) Sample points of sub-blocks divided by H.
[0432] ii. Interleaving predictions are not applied to W-levels that span the current block. (H / 2) Segmentation and Lower W Sample points of sub-blocks divided by (H / 2).
[0433] s. In one example, interleaving prediction is enabled only for blocks with a specific set of widths and / or heights.
[0434] t. In one example, interleaving prediction is disabled only for blocks with a specific set of widths and / or heights.
[0435] u. In one example, interleaving prediction is only used for specific types of picture / strip / group / piece / or other kinds of video data units.
[0436] i. For example, interleaving prediction is only used for P-images or B-images.
[0437] ii. For example, a signal transmission flag in the header of the image / strip / piece group / piece indicates whether interleaving prediction can be used.
[0438] 1. For example, the flag is transmitted via signal only if affine prediction is permitted.
[0439] 17. A method is proposed to transmit messages via signaling to indicate whether / how interleaving prediction is applied and its dependence on width and height. Messages can be transmitted via signaling in SPS / VPS / PPS / strip header / picture header / slice / slice group header / CTU / CTU line / multiple CTUs / or other types of video processing units.
[0440] 18. In one example, bidirectional forecasting is not allowed when using interleaved forecasting.
[0441] a. For example, when using interleaved prediction, the index indicating whether bidirectional prediction is used is not transmitted via signaling.
[0442] b. Alternatively, an indication of whether bidirectional prediction is not allowed can be transmitted via signal transmission in SPS / VPS / PPS / strip header / image header / film / film group header / CTU / CTU line / multiple CTUs.
[0443] 19. A method was proposed to further refine the motion information of sub-blocks based on motion information derived from two or more modes.
[0444] a. In one example, refined motion information can be used to predict subsequent blocks to be encoded or decoded.
[0445] b. In one example, refined motion information can be used in filtering processes such as deblocking, SAO, and ALF.
[0446] c. Whether to store refined information can be based on the position of the sub-block relative to the whole block / CTU / CTU line / slice / strip / slice group / image.
[0447] d. Whether to store refined information may be based on the encoded / decoded patterns of the current block and / or neighboring blocks.
[0448] e. Whether to store refined information can be based on the dimension of the current block.
[0449] f. Whether to store refined information can be based on image / strip type / reference image list, etc.
[0450] 20. It is proposed that whether and / or how to apply the deblocking process or other kinds of filtering processes (such as SAO, adaptive loop filter) can depend on whether interleaved prediction is applied.
[0451] a. In one example, if an edge between two sub-blocks in one partitioning pattern of a block is within a sub-block in another partitioning pattern of the block, then that edge is not deblocked.
[0452] b. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then the deblocking of that edge is weakened.
[0453] i. In one example, for such edges, bS[xDi][yDj] described in the VVC deblocking process decreases.
[0454] ii. In one example, for such edges, the β reduction described in the VVC deblocking process.
[0455] iii. In one example, for such edges, the VVC deblocking process is described Decrease.
[0456] iv. In one example, for such edges, tC as described in the VVC deblocking process decreases.
[0457] c. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is within a sub-block in another partitioning pattern of the block, then the deblocking of that edge is strengthened.
[0458] i. In one example, for such an edge, bS[xDi][yDj] described in the VVC deblocking process increases.
[0459] ii. In one example, for such edges, the β increase described in the VVC deblocking process.
[0460] iii. In one example, for such edges, the VVC deblocking process is described Increase.
[0461] iv. In one example, for such edges, the tC described in the VVC deblocking process is increased.
[0462] 21. It is proposed that whether and / or how local illumination compensation or weighted prediction is applied to blocks / subblocks may depend on whether interleaved prediction is applied.
[0463] a. In one example, when a block is encoded and decoded in interleaved prediction mode, local illumination compensation or weighted prediction is not allowed.
[0464] b. Alternatively, if interleaving prediction is applied to blocks / subblocks, there is no need to indicate the need to enable local lighting compensation via signal transmission.
[0465] 22. It was proposed that bidirectional optical flow (BIO) can be skipped when weighted prediction is applied to a block or sub-block.
[0466] a. In one example, BIO can be applied to blocks with weighted predictions.
[0467] b. In one example, BIO can be applied to blocks with weighted predictions, however, certain conditions must be met.
[0468] i. In one example, at least one parameter is required to be within a range or equal to a specific value.
[0469] ii. In one example, specific reference image restrictions may be applied.
[0470] 3. Problem Existing sub-block-based prediction techniques have the following problems: 1) They face a dilemma. If the sub-block size is smaller, the motion information for each sub-block can be more accurate. However, smaller sub-blocks impose higher bandwidth requirements in the MC.
[0471] 2) Motion information derived for smaller sub-blocks can be dangerous, especially when there is noise in the block. Fixing the size of sub-blocks within a block can be suboptimal.
[0472] 4. Detailed Solution This disclosure proposes further improvements to sub-block-based motion compensation. Specifically, interleaving prediction is improved to be compatible with pixel-based affine motion compensation. Furthermore, regression affine is also improved.
[0473] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0474] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0475] In this disclosure, with respect to "blocks encoded and decoded in mode N", "mode N" can be a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, merging, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).
[0476] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0477] Regarding interlaced affine 1. The use of interleaved prediction in affine mode can depend on another encoding / decoding method.
[0478] a) In one example, if pixel-based affine motion compensation is applied to the codec block, then interleaving prediction may not be applied.
[0479] b) In one example, interleaved affine prediction and pixel-based affine prediction cannot be used simultaneously for encoding / decoding blocks.
[0480] i. In one example, alternatively, interleaved affine prediction and pixel-based affine prediction can be used simultaneously for encoding and decoding blocks.
[0481] 1) In one example, for a codec block, some pixels can be predicted by interleaved affine mapping, while (some) other pixels can be predicted by pixel-based affine mapping.
[0482] 2) In one example, interleaved affine predictions and pixel-based affine predictions can be mixed to generate new predictions.
[0483] 2. The use of interleaved affine prediction and / or pixel-based affine motion compensation can depend on the block dimension.
[0484] a) In one example, interleaved affine prediction is used if the block dimension meets certain conditions.
[0485] i. In one example, specifically, interleaved affine prediction can be used if both (or either) the height and width are greater than or less than a threshold, or if the ratio of the height to the width is greater than or less than a threshold.
[0486] b) In one example, alternatively, if the block dimensions do not meet the same conditions, pixel-based affine mapping can be used instead.
[0487] c) In one example, dimension conditions can be applied together with other conditions.
[0488] i. In one example, the OBMC flag can be used together with dimension conditions.
[0489] 1) In one example, specifically, if OBMC is used for blocks and the block dimensions meet certain conditions, then pixel affine mapping is used to generate predictions.
[0490] d) In one example, whether to apply interlaced affines can depend on the sub-block size.
[0491] i. In one example, if the sub-block size is smaller than a threshold, such as 4x4, then interleaved affine is not applied.
[0492] 3. In one example, the size of the largest sub-block to be used for affine prediction is set to a constant when using interleaved affine or not using pixel-based affine.
[0493] 4. In one example, PROF (Affine Prediction Refinement with Optical Flow) can be performed on interleaved affines.
[0494] a) In one example, the predictions generated by the two partitioning patterns are both refined by PROF.
[0495] i. In one example, predictions generated only by a specific partitioning pattern are refined using PROF.
[0496] ii. In one example, if interleaving affine is used for a block, PROF is not applied.
[0497] b) In one example, for a specific partitioning pattern, only a portion of the predicted samples are refined by PROF.
[0498] i. In one example, if the width or height of a sub-block is less than a threshold, such as 4, then PROF will not be applied to that sub-block.
[0499] 5. It is proposed that whether and / or how to apply sub-block boundary deblocking procedures or other types of sub-block boundary-based filtering procedures (such as SAO, adaptive loop filter, OBMC) can depend on whether affine mode / pixel-based affine / interleaved prediction is applied.
[0500] a) In one example, if affine mode / pixel-based affine / interleaved prediction is applied, some or all kinds of sub-block boundary-based deblocking / filtering processes are not applied.
[0501] i. In one example, specifically, if affine mode / pixel-based affine / interlacing prediction is applied, then deblocking and / or OBMC facing the sub-block boundary is not applied.
[0502] About OBMC Figure 25 An example of a CU-level OBMC is shown.
[0503] 6. It was proposed that whether and / or how to apply CU-level OBMC can depend on the encoding / decoding mode.
[0504] a) In one example, if a sub-block level inter-frame mode (such as affine, SbTMVP, multi-pass DMVR) is used for the current block, each boundary sub-block will be filtered independently at the CU level OBMC level.
[0505] i. In one example, specifically, boundary sub-blocks are iterated one by one. Suppose A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block used for OBMC filtering. If A and A1 have different motions, then A will be filtered. After A has been filtered, subsequent sub-blocks will undergo OBMC in a similar manner.
[0506] b) In one example, alternatively, if a non-sub-block level inter-frame mode is used for the current block, multiple boundary sub-blocks can be filtered together using OBMC.
[0507] i. In one example, suppose A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block. If A1 and the subsequent N (N>=0) consecutive sub-blocks (e.g., Figure 25 If two sub-blocks B1 and C1 have the same motion M_nei, and M_nei is different from the motion of A, then (N + 1) sub-blocks (including A) will be filtered together. Otherwise, if M_nei is equal to the motion of A, then (N + 1) consecutive sub-blocks (including A) will skip CU-level OBMC filtering.
[0508] Regarding regressive affine 7. To generate regression affine candidates, it is proposed to use N (N>1) previously encoded and decoded CU motion fields as inputs to the regression process.
[0509] a) In one example, N CUs can be collected from adjacent locations, non-adjacent locations, or historical CU / parameter tables.
[0510] b) In one example, all N CUs are inter-frame encoded and decoded.
[0511] i. In one example, all N CUs are encoded and decoded using affine mode.
[0512] ii. In one example, alternatively, at least K (K>=0) of the N are affine encoded or decoded.
[0513] c) In one example, all N CUs may need to share the same prediction direction (i.e., bidirectional or unidirectional prediction, and / or the reference list used) and / or the same reference index / frame.
[0514] i. In one example, alternatively, they may have different prediction directions or reference frames.
[0515] d) In one example, regression affine candidates can only be generated if the number of sub-blocks of N CUs is greater than a constant threshold.
[0516] e) In one example, sports fields in adjacent or non-adjacent locations can also be used as additional inputs.
[0517] i. In one example, M rows / columns of adjacent sub-blocks and / or F rows / columns of non-adjacent sub-blocks can be used as input, where M, F>= 0.
[0518] 1) In one example, the location of the non-adjacent positions used can depend on the block dimension.
[0519] f) In one example, the proposed regression candidate can be used when the number of existing regression candidates has not reached the maximum allowed number.
[0520] g) In one example, the proposed regression candidates can be reordered based on a specific metric, such as ARMC or template matching.
[0521] h) In one example, the number of proposed regression candidates may not exceed a constant or an adaptively determined value.
[0522] Figure 26 A flowchart of a method 2600 for video processing according to an embodiment of the present disclosure is shown. Method 2600 implements the conversion between video blocks or video units of a video and a video bitstream.
[0523] At box 2610, for the conversion between the target video block and the video bitstream, the use of interleaved affine prediction is determined based on at least one of the first codec tool or the codec information of the target video block. As used herein, the term "target video unit" refers to the video block or video unit to be processed, which may also be referred to as the "current video block". As used herein, the term "interleaved prediction" refers to a prediction generated by dividing a video block into sub-blocks using at least two partitioning modes, as described in Section 2.8. The term "interleaved affine prediction" may refer to interleaved prediction under affine modes.
[0524] At box 2620, a transformation is performed based on the use of interleaved affine prediction. In some embodiments, the transformation includes encoding the target video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the target video block from the bitstream.
[0525] Method 2600 enables the determination of whether to apply interleaved affine prediction based on encoding / decoding information (such as dimensionality information) from another encoding / decoding tool or the target video block. In this way, encoding / decoding efficiency can be improved.
[0526] In some embodiments, the first encoding / decoding tool includes pixel-based affine motion compensation.
[0527] In some embodiments, the codec information of the target video block includes the dimensionality information of the target video block. Determining the use of interleaved affine prediction includes determining whether the dimensionality information satisfies at least one condition. At least one condition includes at least one of the following: a first condition, at least one of the height or width of the target video block is greater than or less than a dimensionality threshold; a second condition, the ratio of the height to the width of the target video block is greater than or less than a ratio threshold; or a third condition, the sub-block size of the target video block is greater than or equal to a threshold size. If the dimensionality information of the target video block satisfies at least one condition, interleaved affine prediction is enabled. As used herein, enabling a codec tool or codec mode means that the codec tool or codec mode is applied to the video block. Similarly, disabling a codec tool or codec mode means that the codec tool or codec mode is not applied to the video block.
[0528] In some embodiments, determining the use of interleaved affine prediction further includes disabling interleaved affine prediction based on the determination that the dimensional information of the target video block does not meet at least one condition.
[0529] In some embodiments, method 2600 further includes: determining the use of a first encoding / decoding tool based on dimensional information.
[0530] In some embodiments, determining the use of the first codec tool includes: enabling the first codec tool based on the determination that the dimensional information of the target video block does not meet at least one condition.
[0531] In some embodiments, the use of at least one of the interleaved affine prediction or the first codec tool is also determined based on the flag of Overlap Block Motion Compensation (OBMC).
[0532] In some embodiments, the first codec tool is enabled if the OBMC flag indicates that OBMC is enabled for the target video block and the dimension information meets at least one condition.
[0533] In some embodiments, the use of interleaved affine prediction is based on information associated with the first encoding / decoding tool.
[0534] In some embodiments, if the first codec tool is enabled for the target video block, interleaved affine prediction is disabled.
[0535] In some embodiments, interleaved affine prediction and the first encoding / decoding tool are not used simultaneously for the target video block, or interleaved affine prediction and the first encoding / decoding tool are used simultaneously for the target video block.
[0536] In some embodiments, some pixels of the target video block are predicted by interleaved affine prediction, and other pixels of the target video block are predicted by a first encoding / decoding tool.
[0537] In some embodiments, a first prediction from interleaved affine prediction and a second prediction from a first encoding / decoding tool are mixed to generate a third prediction.
[0538] In some embodiments, if interleaved affine prediction is enabled or the first codec tool is disabled, the size of the largest sub-block of the target video block is constant.
[0539] In some embodiments, interleaved affine prediction is enabled for the target video block, and the method further includes performing prediction refinement (PROF) using optical flow for the interleaved affine prediction.
[0540] In some embodiments, performing PROF includes: obtaining a first prediction and a second prediction of a target video block by applying interleaved affine prediction, the first prediction and the second prediction corresponding to a first partitioning pattern and a second partitioning pattern; and obtaining a first refinement prediction and a second refinement prediction by performing PROF on the first prediction and the second prediction.
[0541] In some embodiments, performing PROF includes: obtaining a first prediction and a second prediction of a target video block by applying interleaved affine prediction, the first prediction and the second prediction corresponding to a first partitioning pattern and a second partitioning pattern; and obtaining a first refined prediction by performing PROF on the first prediction without refining the second prediction.
[0542] In some embodiments, for a first partitioning mode, a portion of the predicted samples of the first prediction are refined using PROF.
[0543] In some embodiments, if the width or height of a sub-block of a target video block is less than a threshold, PROF is disabled for the predicted samples of the sub-block.
[0544] In some embodiments, interleaved affine prediction is enabled for the target video block, and prediction refinement using optical flow (PROF) is disabled for interleaved affine prediction.
[0545] In some embodiments, method 2600 further includes determining whether and / or how to apply at least one sub-block boundary-based process to the target video block based on whether at least one of the following is enabled: affine mode, pixel-based affine or interleaved prediction.
[0546] In some embodiments, at least one sub-block boundary-based process includes at least one of the following: sub-block boundary deblocking process, sub-block boundary-based filtering process, sample adaptive compensation (SAO), adaptive loop filter, or overlapping block motion compensation (OBMC).
[0547] In some embodiments, if at least one of affine mode, pixel-based affine, or interleaving prediction is enabled, at least a portion of at least one sub-block boundary-based process is disabled.
[0548] In some embodiments, if at least one of affine mode, pixel-based affine, or interleaving prediction is enabled, at least one of the following is disabled: deblocking or overlapping block motion compensation based on sub-block boundaries (OBMC).
[0549] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the use of interleaved affine prediction based on at least one of the encoding / decoding information of a first encoding / decoding tool or a target video block of the video; and generating a bitstream based on the use of interleaved affine prediction.
[0550] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining the use of interleaved affine prediction based on at least one of the encoding / decoding information of a first encoding / decoding tool or a target video block of the video; generating a bitstream based on the use of interleaved affine prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0551] Figure 27 A flowchart of a method 2700 for video processing according to an embodiment of the present disclosure is shown. Method 2700 implements the conversion between video units or video blocks of a video and a bitstream of the video.
[0552] At box 2710, for the conversion between the target video block and the video bitstream, the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level is determined based on the codec mode of the target video block.
[0553] At box 2720, a conversion is performed based on the use of the CU-level OBMC. In some embodiments, the conversion includes encoding the target video block into a bitstream. Alternatively or additionally, in some embodiments, the conversion includes decoding the target video block from the bitstream.
[0554] Method 2700 enables the determination of whether to apply CU-level OBMC based on the encoding / decoding mode of the video block. In this way, encoding / decoding efficiency can be improved.
[0555] In some embodiments, the use of CU-level OBMC includes at least one of the following: whether to apply CU-level OBMC, or how to apply CU-level OBMC.
[0556] In some embodiments, if a sub-block level inter-frame coding / decoding mode is enabled for a target video block, each boundary sub-block of the target video block is independently filtered at the CU level using OBMC.
[0557] In some embodiments, the sub-block level inter-frame coding and decoding mode includes at least one of the following: affine mode, sub-block-based temporal motion vector prediction (SbTMVP), or multi-pass decoder-side motion vector refinement (DMVR).
[0558] In some embodiments, applying CU-level OBMC filtering includes traversing a plurality of boundary sub-blocks of a target video block. During the traversal of a first boundary sub-block among the plurality of boundary sub-blocks of the target video block, if the motion of the first boundary sub-block differs from the motion of its corresponding neighboring sub-blocks for OBMC, then OBMC filtering is applied to the first boundary sub-block. After the first boundary sub-block has been traversed, subsequent boundary sub-blocks can be traversed in a similar manner.
[0559] In some embodiments, if a non-sub-block level inter-frame coding / decoding mode is enabled for a target video block, multiple boundary sub-blocks of the target video block are filtered together at the CU level using OBMC.
[0560] In some embodiments, applying CU-level OBMC filtering includes traversing a plurality of boundary sub-blocks. During the traversal of a first boundary sub-block among the plurality of boundary sub-blocks, if the motion of the first boundary sub-block differs from the motion set of a set of neighboring sub-blocks, a group of boundary sub-blocks among the plurality of sub-blocks is filtered together. This motion set is identical, and the sub-block set includes the corresponding neighboring sub-blocks of the first boundary sub-block and a first number of consecutive sub-blocks of the neighboring sub-blocks, and the group of boundary sub-blocks includes the first boundary sub-block and the first number of consecutive boundary sub-blocks. After the first boundary sub-block has been traversed, subsequent boundary sub-blocks can be traversed in a similar manner.
[0561] In some embodiments, if the motion of the first boundary sub-block is the same as the motion set, then CU-level OBMC filtering is skipped for the first boundary sub-block.
[0562] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the use of overlay block motion compensation (OBMC) at the codec unit (CU) level based on the encoding / decoding mode of a target video block of the video; and generating a bitstream based on the use of the CU-level OBMC.
[0563] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining the use of overlay block motion compensation (OBMC) at the codec unit (CU) level based on the encoding / decoding mode of a target video block of the video; generating a bitstream based on the use of the CU-level OBMC; and storing the bitstream in a non-transitory computer-readable recording medium.
[0564] Figure 28 A flowchart of a method 2800 for video processing according to an embodiment of the present disclosure is shown. Method 2800 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.
[0565] At box 2810, for the conversion between the target video block and the video bitstream, determine the motion field of at least two codec units that were encoded and decoded before the target video block.
[0566] At box 2820, at least one regressive affine candidate for the target video block is determined based on the motion field of at least two codec units.
[0567] At block 2830, a transformation is performed based on at least one regressive affine candidate. In some embodiments, the transformation includes encoding the target video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the target video block from the bitstream.
[0568] Method 2800 enables the generation of regressive affine candidates based on the motion fields of at least two previously encoded / decoded units. In this way, encoding / decoding efficiency and encoding / decoding effectiveness can be improved.
[0569] In some embodiments, at least two codec units are collected from at least one adjacent location, at least one non-adjacent location, or a historical codec unit table or parameter table.
[0570] In some embodiments, at least two encoding / decoding units are inter-frame encoded / decoded.
[0571] In some embodiments, at least two codec units are encoded and decoded using an affine mode, or some of the codec units in at least two codec units are encoded and decoded using an affine mode.
[0572] In some embodiments, none of the at least two codec units are encoded or decoded using an affine mode.
[0573] In some embodiments, at least two codec units share at least one of the following: the same prediction direction, the same reference list, the same reference index, or the same frame.
[0574] In some embodiments, at least two codec units have at least one of the following: different prediction directions, different reference lists, different reference indices, or different frames.
[0575] In some embodiments, the predicted direction may be bidirectional or unidirectional.
[0576] In some embodiments, if the number of sub-blocks of at least two codec units is greater than a threshold, then at least one regression affine candidate is determined.
[0577] In some embodiments, at least one regression affine candidate is also determined based on the motion field at nearby or non-adjacent locations.
[0578] In some embodiments, the sports field at adjacent or non-adjacent locations includes at least one of the following: a first number of adjacent sub-block rows, a first number of adjacent sub-block columns, a second number of non-adjacent sub-block rows, or a second number of non-adjacent sub-block columns.
[0579] In some embodiments, the positions of non-adjacent locations are based on the block dimension of the target video block.
[0580] In some embodiments, if the number of existing regression candidates for the target video block is less than the maximum allowed number, at least one regression candidate is used for the target video block.
[0581] In some embodiments, method 2800 further includes: reordering at least one regression candidate based on a metric.
[0582] In some embodiments, the metric includes at least one of the following: adaptive reordering of Merge candidates (AMRC) or template matching.
[0583] In some embodiments, the number of at least one regression candidate is less than or equal to a value, wherein the value is a constant or adaptively determined.
[0584] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the motion fields of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate for the target video block based on the motion fields of the at least two codec units; and generating a bitstream based on the at least one regressive affine candidate.
[0585] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining the motion fields of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate for the target video block based on the motion fields of the at least two codec units; generating a bitstream based on the at least one regressive affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0586] It should be understood that methods 2600, 2700, and / or 2800 can be applied individually or in any combination. Using these methods, encoding / decoding efficiency and effectiveness can be improved.
[0587] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0588] Item 1. A method for video processing, comprising: for a conversion between a target video block and a bitstream of the video, determining the use of interleaved affine prediction based on at least one of a first encoding / decoding tool or encoding / decoding information of the target video block; and performing the conversion based on the use of the interleaved affine prediction.
[0589] Item 2. The method according to Item 1, wherein the first encoding / decoding tool includes pixel-based affine motion compensation.
[0590] Item 3. The method according to any one of items 1-2, wherein the encoding / decoding information of the target video block includes dimensional information of the target video block, and wherein determining the use of the interleaved affine prediction comprises: determining whether the dimensional information satisfies at least one condition, the at least one condition including at least one of the following: a first condition, at least one of the height or width of the target video block is greater than or less than a dimensional threshold; a second condition, the ratio of the height to the width of the target video block is greater than or less than a ratio threshold; or a third condition, the sub-block size of the target video block is greater than or equal to a threshold size; and enabling the interleaved affine prediction based on determining that the dimensional information of the target video block satisfies the at least one condition.
[0591] Item 4. The method according to Item 3, wherein determining the use of the interleaved affine prediction further comprises: disabling the interleaved affine prediction based on determining that the dimensional information of the target video block does not satisfy at least one condition.
[0592] Item 5. The method according to Item 3 or 4 further includes: determining the use of the first encoding / decoding tool based on the dimensional information.
[0593] Item 6. The method according to Item 5, wherein determining the use of the first codec tool comprises: enabling the first codec tool based on determining that the dimensional information of the target video block does not satisfy at least one condition.
[0594] Item 7. The method according to any one of Items 5-6, wherein the use of at least one of the interleaved affine prediction or the first codec tool is further determined based on the flag of Overlap Block Motion Compensation (OBMC).
[0595] Item 8. The method according to Item 7, wherein the first codec tool is enabled if the flag of the OBMC indicates that the OBMC is enabled for the target video block and the dimension information satisfies the at least one condition.
[0596] Item 9. The method according to Item 1 or 2, wherein the use of the interleaved affine prediction is based on information associated with the first encoding / decoding tool.
[0597] Item 10. The method according to Item 9, wherein if the first codec tool is enabled for the target video block, the interleaved affine prediction is disabled.
[0598] Item 11. The method according to Item 1 or 2, wherein the interleaved affine prediction and the first codec tool are not used simultaneously for the target video block, or wherein the interleaved affine prediction and the first codec tool are used simultaneously for the target video block.
[0599] Item 12. The method according to any one of Item 1 or 2, wherein a portion of the pixels of the target video block are predicted by the interleaved affine prediction, and the remaining pixels of the target video block are predicted by the first encoding / decoding tool.
[0600] Item 13. The method according to Item 10 or 11, wherein a first prediction from the interleaved affine prediction and a second prediction from the first encoding / decoding tool are mixed to generate a third prediction.
[0601] Item 14. The method according to any one of items 1-13, wherein if the interleaved affine prediction is enabled or the first codec tool is disabled, the size of the largest sub-block of the target video block is constant.
[0602] Item 15. The method according to any one of items 1-14, wherein the interleaved affine prediction is enabled for the target video block, and the method further comprises performing prediction refinement (PROF) using optical flow for the interleaved affine prediction.
[0603] Item 16. The method according to Item 15, wherein performing the PROF comprises: obtaining a first prediction and a second prediction of the target video block by applying the interleaved affine prediction, the first prediction and the second prediction corresponding to a first partitioning pattern and a second partitioning pattern; and obtaining a first refinement prediction and a second refinement prediction by performing the PROF on the first prediction and the second prediction.
[0604] Item 17. The method according to Item 16, wherein performing the PROF includes: obtaining a first prediction and a second prediction of the target video block by applying the interleaved affine prediction, the first prediction and the second prediction corresponding to a first partitioning pattern and a second partitioning pattern; and obtaining a first refined prediction by performing the PROF on the first prediction without refining the second prediction.
[0605] Item 18. The method according to Item 16 or 17, wherein, for the first partitioning pattern, a portion of the predicted samples of the first prediction are refined by the PROF.
[0606] Item 19. The method according to Item 18, wherein if the width or height of a sub-block of the target video block is less than a threshold, the PROF is disabled for the predicted samples of the sub-block.
[0607] Item 20. The method according to any one of items 1-14, wherein the interleaved affine prediction is enabled for the target video block, and prediction refinement (PROF) utilizing optical flow is disabled for the interleaved affine prediction.
[0608] Item 21. The method according to any one of items 1-20 further comprises: determining whether and / or how to apply at least one sub-block boundary-based process to the target video block based on whether at least one of the following is enabled: affine mode, pixel-based affine or interleaved prediction.
[0609] Item 22. The method according to Item 21, wherein the at least one sub-block boundary-based process includes at least one of the following: a sub-block boundary deblocking process, a sub-block boundary-based filtering process, sample adaptive compensation (SAO), an adaptive loop filter, or overlapping block motion compensation (OBMC).
[0610] Item 23. The method according to Item 21 or 22, wherein if at least one of the affine mode, the pixel-based affine, or the interleaving prediction is enabled, at least a portion of the at least one sub-block boundary-based process is disabled.
[0611] Item 24. The method according to Item 21 or 22, wherein if at least one of the affine mode, the pixel-based affine mode, or the interleaving prediction is enabled, at least one of the following is disabled: deblocking based on sub-block boundaries or overlapping block motion compensation (OBMC).
[0612] Item 25. A method for video processing, comprising: a conversion between a target video block and a bitstream of the video; determining the use of codec unit (CU) level overlapping block motion compensation (OBMC) based on the codec mode of the target video block; and performing the conversion based on the use of the CU level OBMC.
[0613] Item 26. The method according to Item 25, wherein the use of the CU-level OBMC includes at least one of the following: whether to apply the CU-level OBMC, or how to apply the CU-level OBMC.
[0614] Item 27. The method according to Item 25 or 26, wherein if a sub-block level inter-frame coding / decoding mode is enabled for the target video block, each boundary sub-block of the target video block is independently subjected to the CU-level OBMC filtering.
[0615] Item 28. The method according to Item 27, wherein the sub-block level inter-frame coding and decoding mode includes at least one of the following: affine mode, sub-block-based temporal motion vector prediction (SbTMVP), or multi-pass decoder-side motion vector refinement (DMVR).
[0616] Item 29. The method according to Item 28, wherein applying the CU-level OBMC filtering comprises: traversing a plurality of boundary sub-blocks of the target video block, wherein during the traversal of a first boundary sub-block among the plurality of boundary sub-blocks of the target video block, the OBMC filtering is applied to the first boundary sub-block based on determining that the motion of the first boundary sub-block is different from the motion of a corresponding neighboring sub-block for OBMC.
[0617] Item 30. The method according to Item 25 or 26, wherein if a non-sub-block level inter-frame coding / decoding mode is enabled for the target video block, then multiple boundary sub-blocks of the target video block are subjected to the CU-level OBMC filtering together.
[0618] Item 31. The method according to Item 30, wherein applying the CU-level OBMC filtering comprises: traversing the plurality of boundary sub-blocks, wherein during the traversal of a first boundary sub-block among the plurality of boundary sub-blocks, filtering a group of boundary sub-blocks among the plurality of sub-blocks together based on determining that the motion of the first boundary sub-block is different from the motion set of a set of neighboring sub-blocks, wherein the motion set is the same, the sub-block set comprising the corresponding neighboring sub-block of the first boundary sub-block and a first number of consecutive sub-blocks of the neighboring sub-blocks, and the group of boundary sub-blocks comprising the first boundary sub-block and the first number of consecutive boundary sub-blocks.
[0619] Item 32. The method according to Item 30, wherein if the motion of the first boundary sub-block is the same as the motion set, the CU-level OBMC filtering is skipped for the first boundary sub-block.
[0620] Item 33. A method for video processing, comprising: for a conversion between a target video block and a bitstream of the video, determining a motion field of at least two codec units encoded and decoded prior to the target video block; determining at least one regressive affine candidate of the target video block based on the motion field of the at least two codec units; and performing the conversion based on the at least one regressive affine candidate.
[0621] Item 34. The method according to Item 33, wherein the at least two codec units are collected from at least one adjacent position, at least one non-adjacent position, or a historical codec unit table or parameter table.
[0622] Item 35. The method according to Item 33 or 34, wherein the at least two encoding / decoding units are inter-frame encoded / decoded.
[0623] Item 36. The method according to any one of items 33-35, wherein the at least two codec units are encoded and decoded using an affine mode, or wherein some of the at least two codec units are encoded and decoded using an affine mode.
[0624] Item 37. The method according to any one of items 33-35, wherein none of the at least two encoding / decoding units is encoded / decoded using an affine mode.
[0625] Item 38. The method according to any one of items 33-37, wherein the at least two codec units share at least one of the following: the same prediction direction, the same reference list, the same reference index, or the same frame.
[0626] Item 39. The method according to any one of items 33-37, wherein the at least two encoding / decoding units have at least one of the following: different prediction directions, different reference lists, different reference indices, or different frames.
[0627] Item 40. The method according to Item 38 or 39, wherein the predicted direction includes bidirectional or unidirectional directions.
[0628] Item 41. The method according to any one of items 33-40, wherein the at least one regression affine candidate is determined if the number of sub-blocks of the at least two encoding / decoding units is greater than a threshold.
[0629] Item 42. The method according to any one of items 33-41, wherein the at least one regression affine candidate is further determined based on the motion field at a neighboring or non-adjacent location.
[0630] Item 43. The method according to Item 42, wherein the sports field of the adjacent or non-adjacent locations includes at least one of the following: a first number of adjacent sub-block rows, a first number of adjacent sub-block columns, a second number of non-adjacent sub-block rows, or a second number of non-adjacent sub-block columns.
[0631] Item 44. The method according to Item 42 or 43, wherein the position of the non-adjacent position is based on the block dimension of the target video block.
[0632] Item 45. The method according to any one of items 33-44, wherein the at least one regression candidate is used for the target video block if the number of existing regression candidates for the target video block is less than the maximum allowed number.
[0633] Item 46. The method according to any one of items 33-45 further includes: reordering the at least one regression candidate based on a metric.
[0634] Item 47. The method according to Item 46, wherein the metric includes at least one of the following: adaptive reordering of Merge candidates (AMRC) or template matching.
[0635] Item 48. The method according to any one of items 33-47, wherein the number of the at least one regression candidate is less than or equal to a value, wherein the value is a constant or adaptively determined.
[0636] Item 49. The method according to any one of items 1-48, wherein the conversion includes encoding the target video block into the bitstream.
[0637] Item 50. The method according to any one of items 1-48, wherein the conversion includes decoding the target video block from the bitstream.
[0638] Item 51. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-50.
[0639] Item 52. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-50.
[0640] Item 53. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining the use of interleaved affine prediction based on at least one of a first encoding / decoding tool or encoding / decoding information of a target video block of the video; and generating the bitstream based on the use of the interleaved affine prediction.
[0641] Item 54. A method for storing a bitstream of video, comprising: determining the use of interleaved affine prediction based on at least one of a first codec tool or codec information of a target video block of the video; generating the bitstream based on the use of the interleaved affine prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0642] Item 55. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining the use of overlapping block motion compensation (OBMC) at the codec unit (CU) level based on the codec mode of a target video block of the video; and generating the bitstream based on the use of the CU-level OBMC.
[0643] Item 56. A method for storing a bitstream of video, comprising: determining the use of codec unit (CU) level overlapping block motion compensation (OBMC) based on the codec mode of a target video block of the video; generating the bitstream based on the use of the CU level OBMC; and storing the bitstream in a non-transitory computer-readable recording medium.
[0644] Item 57. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method comprises: determining a motion field of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate of the target video block based on the motion field of the at least two codec units; and generating the bitstream based on the at least one regressive affine candidate.
[0645] Item 58. A method for storing a bitstream of video, comprising: determining a motion field of at least two codec units encoded and decoded prior to a target video block of the video; determining at least one regressive affine candidate of the target video block based on the motion field of the at least two codec units; generating the bitstream based on the at least one regressive affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0646] Example device Figure 29 A block diagram of a computing device 2900 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2900 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0647] It should be understood that, Figure 29 The computing device 2900 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0648] like Figure 29As shown, computing device 2900 includes general-purpose computing device 2900. Computing device 2900 may include at least one or more processors or processing units 2910, memory 2920, storage unit 2930, one or more communication units 2940, one or more input devices 2950, and one or more output devices 2960.
[0649] In some embodiments, the computing device 2900 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2900 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0650] Processing unit 2910 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2900. Processing unit 2910 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0651] Computing device 2900 typically includes various computer storage media. Such media can be any media accessible by computing device 2900, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2930 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2900.
[0652] The computing device 2900 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 29 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0653] Communication unit 2940 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 2900 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0654] Input device 2950 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2960 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2940, computing device 2900 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2900 can also communicate with one or more devices that enable a user to interact with computing device 2900, or any device that enables computing device 2900 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0655] In some embodiments, some or all components of computing device 2900 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to perform the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may be provided through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0656] In embodiments of this disclosure, computing device 2900 can be used to implement video encoding / decoding. Memory 2920 may include one or more video codec modules 2925 having one or more program instructions. These modules can be accessed and executed by processing unit 2910 to perform the functions of the various embodiments described herein.
[0657] In an example embodiment of performing video encoding, input device 2950 may receive video data as input 2970 to be encoded. The video data may be processed, for example, by video codec module 2925 to generate an encoded bitstream. The encoded bitstream may be provided as output 2980 via output device 2960.
[0658] In an example embodiment of performing video decoding, input device 2950 may receive an encoded bitstream as input 2970. The encoded bitstream may be processed, for example, by a video codec module 2925 to generate decoded video data. The decoded video data may be provided as output 2980 via output device 2960.
[0659] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between a target video block and the bitstream of the video, the use of interleaved affine prediction is determined based on at least one of a first encoding / decoding tool or the encoding / decoding information of the target video block; and The transformation is performed based on the interleaved affine prediction.
2. The method of claim 1, wherein the first encoding / decoding tool includes pixel-based affine motion compensation.
3. The method according to any one of claims 1-2, wherein the codec information of the target video block includes dimensional information of the target video block, and wherein determining the use of the interleaved affine prediction includes: Determine whether the dimensional information satisfies at least one condition, wherein the at least one condition includes at least one of the following: The first condition is that at least one of the height or width of the target video block is greater than or less than the dimension threshold. The second condition is that the ratio of the height to the width of the target video block is greater than or less than a ratio threshold; or The third condition is that the size of the sub-blocks of the target video block is greater than or equal to the threshold size; as well as The interleaved affine prediction is enabled based on the determination that the dimensional information of the target video block satisfies at least one of the conditions.
4. The method of claim 3, wherein determining the use of the interlaced affine prediction further comprises: The interleaved affine prediction is disabled if the dimensional information of the target video block does not satisfy at least one of the conditions.
5. The method according to claim 3 or 4, further comprising: The use of the first encoding / decoding tool is determined based on the dimensional information.
6. The method of claim 5, wherein determining the use of the first codec tool comprises: The first encoding / decoding tool is activated if the dimensional information of the target video block does not meet at least one condition.
7. The method according to any one of claims 5-6, wherein the use of at least one of the interleaved affine prediction or the first codec tool is further determined based on the flag of Overlap Block Motion Compensation (OBMC).
8. The method of claim 7, wherein the first codec tool is enabled if the flag of the OBMC indicates that the OBMC is enabled for the target video block and the dimension information satisfies the at least one condition.
9. The method of claim 1 or 2, wherein the use of the interleaved affine prediction is based on information associated with the first encoding / decoding tool.
10. The method of claim 9, wherein if the first codec tool is enabled for the target video block, the interleaved affine prediction is disabled.
11. The method according to claim 1 or 2, wherein the interleaved affine prediction and the first encoding / decoding tool are not used simultaneously for the target video block, or The interleaved affine prediction and the first encoding / decoding tool are used simultaneously for the target video block.
12. The method of any one of claims 1 or 2, wherein a portion of the pixels of the target video block are predicted by the interleaved affine prediction, and the remaining pixels of the target video block are predicted by the first encoding / decoding tool.
13. The method of claim 10 or 11, wherein the first prediction from the interleaved affine prediction and the second prediction from the first encoding / decoding tool are mixed to generate a third prediction.
14. The method according to any one of claims 1-13, wherein if the interleaved affine prediction is enabled or the first encoding / decoding tool is disabled, the size of the largest sub-block of the target video block is constant.
15. The method of any one of claims 1-14, wherein the interleaved affine prediction is enabled for the target video block, and the method further comprises performing prediction refinement (PROF) using optical flow for the interleaved affine prediction.
16. The method of claim 15, wherein performing the PROF comprises: The first and second predictions of the target video block are obtained by applying the interleaved affine prediction, wherein the first and second predictions correspond to a first partitioning pattern and a second partitioning pattern, respectively; and The first refined prediction and the second refined prediction are obtained by performing the PROF on the first prediction and the second prediction.
17. The method of claim 16, wherein performing the PROF comprises: The first and second predictions of the target video block are obtained by applying the interleaved affine prediction, wherein the first and second predictions correspond to a first partitioning pattern and a second partitioning pattern, respectively; and A first refined prediction is obtained by performing the PROF on the first prediction, without refining the second prediction.
18. The method of claim 16 or 17, wherein, for the first partitioning mode, a portion of the predicted samples of the first prediction are refined by the PROF.
19. The method of claim 18, wherein if the width or height of a sub-block of the target video block is less than a threshold, the PROF is disabled for the predicted samples of the sub-block.
20. The method of any one of claims 1-14, wherein the interleaved affine prediction is enabled for the target video block, and prediction refinement using optical flow (PROF) is disabled for the interleaved affine prediction.
21. The method according to any one of claims 1-20, further comprising: Based on whether at least one of the following is enabled: affine mode, pixel-based affine or interleaved prediction, determine whether and / or how to apply at least one sub-block boundary-based process to the target video block.
22. The method of claim 21, wherein the at least one sub-block boundary-based process comprises at least one of the following: a sub-block boundary deblocking process, a sub-block boundary-based filtering process, sample adaptive compensation (SAO), an adaptive loop filter, or overlapping block motion compensation (OBMC).
23. The method of claim 21 or 22, wherein if at least one of the affine mode, the pixel-based affine, or the interleaving prediction is enabled, at least a portion of the at least one sub-block boundary-based process is disabled.
24. The method of claim 21 or 22, wherein if at least one of the affine mode, the pixel-based affine, or the interleaving prediction is enabled, at least one of the following is disabled: deblocking or overlapping block motion compensation (OBMC) based on sub-block boundaries.
25. A method for video processing, comprising: For the conversion between the target video block and the bitstream of the video, based on the encoding / decoding mode of the target video block, the use of Overlapping Block Motion Compensation (OBMC) at the Codec Unit (CU) level is determined; and The conversion is performed based on the use of the CU-level OBMC.
26. The method of claim 25, wherein the use of the CU-level OBMC includes at least one of: whether to apply the CU-level OBMC, or how to apply the CU-level OBMC.
27. The method of claim 25 or 26, wherein if a sub-block level inter-frame coding / decoding mode is enabled for the target video block, each boundary sub-block of the target video block undergoes the CU-level OBMC filtering independently.
28. The method of claim 27, wherein the sub-block level inter-frame coding and decoding mode includes at least one of the following: affine mode, sub-block-based temporal motion vector prediction (SbTMVP), or multi-pass decoder-side motion vector refinement (DMVR).
29. The method of claim 28, wherein applying the CU-level OBMC filtering comprises: Traverse multiple boundary sub-blocks of the target video block, wherein during the traversal of a first boundary sub-block among the multiple boundary sub-blocks of the target video block, Based on the determination that the motion of the first boundary sub-block is different from the motion of the corresponding neighboring sub-block for OBMC, the OBMC filter is applied to the first boundary sub-block.
30. The method of claim 25 or 26, wherein if a non-sub-block level inter-frame coding / decoding mode is enabled for the target video block, then the multiple boundary sub-blocks of the target video block are subjected to the CU-level OBMC filtering together.
31. The method of claim 30, wherein applying the CU-level OBMC filtering comprises: Traverse the plurality of boundary sub-blocks, wherein during the traversal of the first boundary sub-block among the plurality of boundary sub-blocks, Based on the fact that the motion of the first boundary sub-block is different from the motion set of its neighboring sub-blocks, a group of boundary sub-blocks from among the plurality of sub-blocks are filtered together. The motion set is the same, the sub-block set includes the corresponding neighboring sub-blocks of the first boundary sub-block and a first number of consecutive sub-blocks of the neighboring sub-blocks, and the set of boundary sub-blocks includes the first boundary sub-block and the first number of consecutive boundary sub-blocks.
32. The method of claim 30, wherein if the motion of the first boundary sub-block is the same as the motion set, then the CU-level OBMC filtering is skipped for the first boundary sub-block.
33. A method for video processing, comprising: For the conversion between the target video block and the bitstream of the video, determine the motion field of at least two codec units that were encoded and decoded before the target video block; Based on the motion field of the at least two encoding and decoding units, at least one regression affine candidate of the target video block is determined; as well as The transformation is performed based on the at least one regression affine candidate.
34. The method of claim 33, wherein the at least two codec units are collected from at least one adjacent position, at least one non-adjacent position, or a historical codec unit table or parameter table.
35. The method of claim 33 or 34, wherein the at least two encoding / decoding units are inter-frame encoded / decoded.
36. The method according to any one of claims 33-35, wherein the at least two encoding / decoding units are encoded / decoded using affine patterns, or Some of the at least two codec units are encoded and decoded using affine modes.
37. The method according to any one of claims 33-35, wherein none of the at least two encoding / decoding units is encoded / decoded using an affine mode.
38. The method according to any one of claims 33-37, wherein the at least two encoding / decoding units share at least one of the following: the same prediction direction, the same reference list, the same reference index, or the same frame.
39. The method according to any one of claims 33-37, wherein the at least two encoding / decoding units have at least one of the following: different prediction directions, different reference lists, different reference indices, or different frames.
40. The method of claim 38 or 39, wherein the predicted direction includes bidirectional or unidirectional directions.
41. The method according to any one of claims 33-40, wherein the at least one regression affine candidate is determined if the number of sub-blocks of the at least two encoding / decoding units is greater than a threshold.
42. The method according to any one of claims 33-41, wherein the at least one regression affine candidate is further determined based on the motion field of a neighboring or non-adjacent location.
43. The method of claim 42, wherein the sports field at the adjacent or non-adjacent location comprises at least one of the following: The first number of neighboring sub-block lines, The first number of neighboring sub-blocks, The second number of non-adjacent sub-blocks, or The second number of non-adjacent sub-blocks.
44. The method of claim 42 or 43, wherein the position of the non-adjacent position is based on the block dimension of the target video block.
45. The method according to any one of claims 33-44, wherein if the number of existing regression candidates for the target video block is less than the maximum allowed number, then the at least one regression candidate is used for the target video block.
46. The method according to any one of claims 33-45, further comprising: The at least one regression candidate is reordered based on a metric.
47. The method of claim 46, wherein the metric comprises at least one of: adaptive reordering of Merge candidates (AMRC) or template matching.
48. The method according to any one of claims 33-47, wherein the number of the at least one regression candidate is less than or equal to a value, wherein the value is a constant or adaptively determined.
49. The method according to any one of claims 1-48, wherein the conversion comprises encoding the target video block into the bitstream.
50. The method according to any one of claims 1-48, wherein the conversion comprises decoding the target video block from the bitstream.
51. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-50.
52. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-50.
53. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The use of interleaved affine prediction is determined based on at least one of the encoding and decoding information of the first encoding and decoding tool or the target video block of the video; as well as The bitstream is generated based on the interleaved affine prediction.
54. A method for storing a bitstream of video, comprising: The use of interleaved affine prediction is determined based on at least one of the encoding and decoding information of the first encoding and decoding tool or the target video block of the video; The bitstream is generated based on the use of the interleaved affine prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
55. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Based on the encoding and decoding mode of the target video block of the video, determine the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level; as well as The bitstream is generated based on the use of the CU-level OBMC.
56. A method for storing a bitstream of video, comprising: Based on the encoding and decoding mode of the target video block of the video, determine the use of Overlapping Block Motion Compensation (OBMC) at the codec unit (CU) level; The bitstream is generated based on the use of the CU-level OBMC; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
57. A non-transitory computer-readable recording medium storing instructions, the instructions being a bitstream generated by a method executed by means of a video processing apparatus, wherein the method comprises: Determine the motion field of at least two codec units that are encoded and decoded before the target video block of the video; Based on the motion field of the at least two encoding and decoding units, at least one regression affine candidate of the target video block is determined; as well as The bit stream is generated based on the at least one regression affine candidate.
58. A method for storing a bitstream of video, comprising: Determine the motion field of at least two codec units that are encoded and decoded before the target video block of the video; Based on the motion field of the at least two encoding and decoding units, at least one regression affine candidate of the target video block is determined; The bitstream is generated based on the at least one regression affine candidate; as well as The bitstream is stored in a non-transitory computer-readable recording medium.