Method and device for video processing and medium
By determining the motion field of the current video block and multiple codec units, and utilizing regression affine candidates and affine advanced motion vector prediction modes, the problem of low encoding and decoding efficiency in existing technologies is solved, achieving more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202480027909.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-23
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-25
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved, especially when dealing with complex motion patterns, as existing methods struggle to effectively utilize motion information from neighboring blocks for efficient encoding and decoding.
By determining the motion field of the current video block and multiple codec units, and utilizing regression affine candidates and affine advanced motion vector prediction modes, encoding and decoding efficiency is improved.
It improves the efficiency of video encoding and decoding, especially when dealing with complex motion patterns, and can more effectively utilize the motion information of neighboring blocks for encoding and decoding.
Smart Images

Figure CN121014206A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to regression affine candidate. BACKGROUND
[0002] Nowadays, digital video capability is being applied to various aspects of people's life. Various types of video compression techniques have been proposed for video coding / decoding, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, the coding efficiency of video coding techniques is generally expected to be further improved. SUMMARY
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method comprises: determining, for a conversion between a current video block of a video and a bitstream of the video, motion fields of a plurality of coded units (CUs) that are coded before the current video block, wherein at least one coded unit of the plurality of coded units is collected from at least one of: adjacent neighboring positions, adjacent neighboring positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block; determining, based on the motion fields of the plurality of coded units, a regression affine candidate for the current video block; and performing the conversion based on the regression affine candidate. The method according to the first aspect of the present disclosure determines the regression affine candidate based on the motion fields of a plurality of previously coded CUs, thereby improving the coding efficiency.
[0005] In a second aspect, another method for video processing is proposed. The method comprises: determining, for a conversion between a current video block of a video and a bitstream of the video, motion fields of a plurality of coded blocks; determining, based on the motion fields of the plurality of coded blocks, a regression affine candidate for the current video block, the current video block being in an affine advanced motion vector prediction (AMVP) mode; and performing the conversion based on the regression affine candidate. The method according to the second aspect of the present disclosure determines the regression affine candidate for the block in the affine AMVP mode based on the motion fields of a plurality of coded blocks or coded units, thereby improving the coding efficiency.
[0006] In a third aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect or the second aspect of the present disclosure.
[0007] In a fourth aspect, a non-transitory computer-readable storage medium is presented. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first aspect or the second aspect of the present disclosure.
[0008] In a fifth aspect, another non-transitory computer-readable recording medium is presented. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method comprises: determining motion fields of a plurality of coding units coded before a current video block of the video, wherein at least one of the plurality of coding units is collected from at least one of: neighboring spatial locations, neighboring spatial locations at a location, collocated temporal locations, neighboring temporal locations, non-neighboring spatial locations, non-neighboring temporal locations, or a history table of the current video block; determining an affine candidate for the current video block based on the motion fields of the plurality of coding units; and generating the bitstream based on the affine candidate.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is presented. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method comprises: determining motion fields of a plurality of coding units coded before a current video block of the video, wherein at least one of the plurality of coding units is collected from at least one of: neighboring spatial locations, neighboring spatial locations at a location, collocated temporal locations, neighboring temporal locations, non-neighboring spatial locations, non-neighboring temporal locations, or a history table of the current video block; determining a regression affine candidate for the current video block based on the motion fields of the plurality of coding units; generating the bitstream based on the regression affine candidate; and storing the bitstream in the non-transitory computer-readable recording medium.
[0010] In a seventh aspect, another non-transitory computer-readable recording medium is presented. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method comprises: determining motion fields of a plurality of coding blocks; determining a regression affine candidate for a current video block of the video based on the motion fields of the plurality of coding blocks, the current video block being in an affine advanced motion vector prediction (AMVP) mode; and generating the bitstream based on the regression affine candidate.
[0011] In an eighth aspect, a method for storing a bitstream of a video is presented. The method comprises: determining motion fields of a plurality of coding blocks; determining a regression affine candidate for a current video block of the video based on the motion fields of the plurality of coding blocks, the current video block being in an affine advanced motion vector prediction (AMVP) mode; generating the bitstream based on the regression affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0012] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0013] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same parts.
[0014] FIG. 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; FIG. 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; FIG. 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; FIG. 4 This shows the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list; FIG. 5 This shows the positions of non-adjacent candidates in the ECM; FIG. 6 An affine motion model based on control points is shown; FIG. 7 An example affine MVF for each sub-block is shown; FIG. 8 The position of the inherited affine motion predictor is shown; FIG. 9 This demonstrates the inheritance of control point motion vectors; FIG. 10 The locations of candidate positions for the constructive affine Merge pattern are shown; FIG. 11 The spatial nearest neighbor used to derive the affine Merge candidate is shown; FIG. 12 This shows candidates for constructive affine merges, ranging from non-nearest neighbors. FIG. 13 An example of generating HAPC is shown; FIG. 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. FIG. 15 This demonstrates template matching execution over the search area surrounding the initial MV; FIG. 16 The template and the corresponding reference template are shown; FIG. 17 A template and a reference template are shown for a block with sub-block motion that uses motion information of the current block's sub-blocks; FIG. 18 The derivation of the sub-CU motion field obtained by applying motion displacement based on neighbor motion information is shown; FIG. 19 An example of interleaved prediction is shown; FIGS. 20A-20G An exemplary partitioning pattern for a 16×16 block is shown; FIGS. 21A-21D An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas; FIGS. 22A-22C An example is shown of deriving the MV of a partitioning pattern from another partitioning pattern; FIGS. 23A-23C An example of selecting a partitioning mode based on block dimensions is shown; FIG. 24A and FIG. 24B An example is shown that derives the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of another partitioning pattern. FIG. 25 An example of a CU-level OBMC is shown; FIG. 26 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; FIG. 27 A flowchart of another method for video processing according to embodiments of the present disclosure is shown; and FIG. 28 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0015] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0016] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0017] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0018] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0019] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0021] Example Environment FIG. 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0022] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0023] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. The codec images are codec representations of images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0024] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0025] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0026] FIG. 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... FIG. 1 An example of a video encoder 114 in system 100 is shown.
[0027] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. FIG. 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0028] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0029] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0030] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... FIG. 2 The examples are shown separately.
[0031] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0032] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0033] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0034] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0035] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0036] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0037] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0038] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0039] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0040] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0041] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0042] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0043] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0044] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0045] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0046] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0047] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0048] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0049] FIG. 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... FIG. 1 An example of video decoder 124 in system 100 is shown.
[0050] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. FIG. 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0051] exist FIG. 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0052] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0053] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0054] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0055] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0056] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0057] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0058] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while specific embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term video processing includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compression format to another or at different compression bitrates.
[0059] 1. Brief Overview This disclosure relates to video coding and decoding techniques. Specifically, it relates to affine motion prediction methods in video coding and decoding. These ideas can be applied individually or in various combinations to any standard or non-standard video codec.
[0060] 2. Introduction The exponential growth of multimedia data poses a significant challenge to video encoding and decoding. To meet the ever-increasing demand for more efficient compression technologies, the ITU-T and ISO / IEC have developed a series of video encoding and decoding standards over the past few decades. Specifically, the ITU-T developed the H.261 and H.263 standards, and ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations have also jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), H.265 / HEVC, and the latest VVC standard. Since H.262 / MPEG-2, a hybrid video encoding and decoding framework has been adopted, utilizing intra / inter-frame prediction plus transform encoding and decoding.
[0061] 2.1. MVP in Video Encoding and Decoding Inter-frame prediction aims to eliminate temporal redundancy between adjacent frames and is an indispensable component in hybrid video codec frameworks. Specifically, inter-frame prediction utilizes the content specified by motion vectors (MVs) as the predicted version of the current block to be encoded or decoded, thus transmitting only residual signals and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an efficient mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. In H.265 / HEVC, a contention mechanism is involved, where the best MVP is selected from multiple candidates through rate-distortion optimization (RDO). Specifically, Advanced MVP (AMVP) mode and Merge mode are designed using different motion information signaling strategies. With AMVP mode, the reference index, the MVP candidate index referencing the AMVP candidate list, and the motion vector difference (MVD) are transmitted via signaling. Regarding Merge mode, only the Merge index referencing the Merge candidate list is transmitted via signaling, and all motion information associated with the Merge candidate is inherited. Both the AMVP and Merge modes require building an MVP candidate list, and the details of the building process for these two modes are described below.
[0062] AMVP mode: AMVP utilizes the spatial-temporal correlation of motion vectors with neighboring blocks for explicit transfer of motion parameters. For each list of reference images, the candidate list of motion vectors is constructed as follows: first, the availability of temporally adjacent locations on the left and top is checked, redundant candidates are removed, and zero vectors are added to give the candidate list a fixed length. FIG. 4 The diagram illustrates the locations of spatial and temporal neighbor blocks used in the construction of the AMVP / Merge candidate list. For the spatial motion vector candidate derivation, the two motion vector candidates are ultimately based on blocks located as follows: FIG. 4 Motion vectors for the five blocks at different locations are derived. The five neighboring blocks located at B0, B1, B2, and A0, A1 are classified into two groups: group A includes the three spatially adjacent blocks above, and group B includes the two spatially adjacent blocks to the left. Two MV candidates are derived in a predefined order using the first available candidates from group A and group B, respectively. For temporal motion vector candidate derivation, a motion vector candidate is derived by sequentially checking based on two different co-locations (lower right (C0) and center (C1)), as follows: FIG. 4 As shown. To avoid redundant MV candidates, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than 2, additional zero motion vector candidates are added to the list.
[0063] FIG. 5 The positions of non-adjacent candidates in the ECM are shown.
[0064] Merge mode Similar to the AMVP mode, the MVP candidate list for the Merge mode also consists of spatial and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, a maximum of four candidates are selected in the order A1, B1, B0, A0, and B2. For temporal Merge Candidate (TMVP) derivation, a maximum of one candidate is selected from two temporally neighboring blocks (C0 and C1). When there are not enough Merge candidates using both spatial and temporal candidates, combined bidirectional prediction Merge candidates and zero MV candidates are added to the MVP candidate list. The Merge candidate list construction process terminates once the number of available Merge candidates reaches the maximum allowed number for signal transmission.
[0065] In VVC, the Merge pattern construction process is further improved by introducing a history-based MVP (HMVP), which incorporates motion information from previously encoded / decoded blocks that may be geographically distant from the current block. In VVC, HMVP Merge candidates are appended to the Merge list after the Spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained using a first-in, first-out (FIFO) strategy during the encoding / decoding process. Whenever a non-sub-block inter-frame encoding / decoding CU exists, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0066] During the standardization of VVC, the non-adjacent MVP was proposed to facilitate better motion information derivation by utilizing non-adjacent regions. In ECM software, the non-adjacent MVP is inserted between the TMVP and HMVP, where the distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block, such as... FIG. 5 As shown.
[0067] 2.2. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. FIG. 6 Affine motion models based on control points are shown, such as (a) a 4-parameter affine model and (b) a 6-parameter affine model. FIG. 6 As shown, the affine motion field of a block is described by motion information from two control point motion vectors (4 parameters) or three control point motion vectors (6 parameters).
[0068] For the 4-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (1).
[0069] For the 6-parameter affine motion model, the motion vector at the sample point position (x, y) in the block is derived as: (2).
[0070] in( mv0x, mv0y ) is the motion vector of the upper left control point, ( mv1x, mv1y ) is the motion vector of the upper right control point, and ( mv2x, mv2y ) is the motion vector of the lower left control point.
[0071] To simplify motion compensation prediction, block-based affine transformation prediction is applied. FIG. 7 An example affine MVF for each sub-block is shown. To derive the motion vector for each 4×4 luma sub-block, as follows... FIG. 7 As shown, the motion vector of the center sample point of each sub-block is calculated according to the above equation and rounded to 1 / 16 pixel precision. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8×8 luminance region.
[0072] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0073] 2.2.1. Affine Merge Prediction The Affine Merge pattern can be applied to CUs with a width and height greater than or equal to 8. In this pattern, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and a signal transmission index indicates which CPVM should be used for the current CU. In VVC, the following three types of CPVM candidates are used to form the Affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; – Constructive affine Merge candidate CPMVP derived using translational MV of neighboring CUs.
[0074] – Zero MV.
[0075] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of the neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU.FIG. 8 The positions of inherited affine motion predictors are shown. Candidate blocks are as follows: FIG. 8 As shown. For the predictor on the left, the scan order is A0->A1, and for the predictor above, the scan order is B0->B1->B2. Only the first inherited candidate is selected from each side. No pruning check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. FIG. 9 The inheritance of control point motion vectors is shown. For example... FIG. 9 As shown, if the adjacent lower-left block A is encoded and decoded in affine mode, then the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are... Obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on... Computed. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... Calculated.
[0076] Constructive affine candidates mean building candidates by combining the translational motion information of each control point's neighbors. The motion information of the control points is derived from... FIG. 10 The spatial and temporal nearest neighbors shown are derived. CPMVk (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks are checked by B2->B3->A2, and the MV of the first available block is used. For CPMV2, blocks are checked by B1->B0, and for CPMV3, blocks are checked by A1->A0. If available, the TMVP is used as CPMV4.
[0077] After obtaining the motion signatures (MVs) of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.
[0078] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0079] FIG. 10 The locations of candidate positions for the constructive affine Merge pattern are shown.
[0080] After the inherited affine Merge candidates and the constructed affine Merge candidates have been checked, if the list is still not full, zero MV is inserted at the end of the list.
[0081] 2.2.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted sub-CPVM is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.
[0082] – A constructive affine AMVP candidate CPMVP derived using the translation MV of neighboring CUs.
[0083] – Translation MV from the neighboring CU.
[0084] – Zero MV.
[0085] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference picture as the current block are considered. No pruning is applied when inserting inherited affine motion predictors into the candidate list.
[0086] Constructive AMVP candidates are from FIG. 10 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is also checked. In the checking order, the first block that has been inter-coded and has the same reference picture as the current CU is used. There is a rule that the current CU is encoded and decoded using a 4-parameter affine mode, and... mv0 and mv1 When all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU is encoded and decoded using a 6-parameter affine mode, and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructive AMVP candidates are set to unavailable.
[0087] FIG. 11Spatial nearest neighbors are shown for deriving affine Merge candidates: (a) for deriving inherited affine Merge candidates and (b) for deriving constructed affine Merge candidates.
[0088] If, after inserting valid inherited and constructed AMVP candidates, the affine AMVP list still contains fewer than two candidates, then when available, mv0 , mv1 and mv2 They will be added sequentially as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.
[0089] 2.2.3. New Affine Candidate Derivation Method in ECM-8.0 In ECM-6.0, three additional affine Merge and AMVP candidate derivation methods are integrated: non-adjacent spatial domain candidates, historical parameter-based candidates, regression-based affine candidates, and pixel-based affine motion compensation.
[0090] 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, non-adjacent airspace neighbors are studied to provide candidates for both affine Merge and affine AMVP. The pattern for obtaining non-adjacent airspace candidates is... FIG. 11 As shown in the diagram, similar to the non-adjacent regular Merge candidate, the distance between the non-adjacent spatial candidate and the current codec block is also defined based on the width and height of the current CU.
[0091] FIG. 11 Motion information of non-adjacent spatial neighbors in the current block is used to generate additional inherited and constructed affine merge candidates. Specifically, to generate inherited candidates, non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, only the first available neighbor encoded in affine mode from each side (e.g., left and top) of the current block is included. FIG. 11 As shown in (a), the checks of the left and top nearest neighbors are performed from bottom to top and from right to left, respectively. For constructive candidates, as... FIG. 11 As shown in (b), the positions of a non-adjacent spatial neighbor on the left and a non-adjacent spatial neighbor above are first determined independently; then, the position of the upper-left neighbor can be determined accordingly to form a rectangular virtual block together with the non-adjacent neighbors on the left and above. The motion information of the three non-adjacent neighbors is used to form a CPMV at the upper-left (A), upper-right (B), and lower-left (C) of the virtual block, which is projected onto the current CU to generate corresponding constructivist candidates, such as... FIG. 12 As shown.
[0092] 2.2.3.2. Affine Candidates Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to be inherited from previously affine encoded / decoded blocks (which may not be adjacent to the current block). A History Parameter Table (HPT) is established. Each entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entry points in the HPT are categorized by reference list and reference index. For each reference list in the HPT, five reference indices are supported. The HPT category (denoted as HPTCat) is calculated in a formulaic manner as follows: (3) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of seven entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding the affine-encoded CU using the reference lists RefListcur and RefIdxcur, the affine parameters are used to update the entries in the category HPTCat(RefListcur, RefIdxcur) in a manner similar to HMVP table updates.
[0093] Candidates based on historical affine parameters (HAPC) from... FIG. 13 The set of affine parameters in the neighboring 4×4 blocks, denoted as A0, A1, B0, B1, or B2, and the corresponding entries stored in the HPT, is derived. The MV of the neighboring 4×4 blocks is used as the base MV. In a formulaic manner, the MV of the current block at position (x, y) is calculated as: , (4) Where (mvhbase, mvvbase) represents the MV of the nearest 4×4 block, and (xbase, ybase) represents the center position of the nearest 4×4 block. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0094] FIG. 13An example of how to derive the HAPC from block A0 is shown. The affine parameters {a0, b0, c0, d0} are directly extracted from an entry in the category HPTIdx(RefListA0, refIdx0A0) in the HPT. The affine parameters from the HPT, along with the center position of A0 (as the base position) and the MV of block A0 (as the base MV), are used to derive the CPMV for either the affine MergeHAPC or the affine AMVP HAPC. They can also be used to derive the MV located at the center of the current block as regular Merge candidates. The HAPC can be placed into the sub-block-based Merge candidate list, the affine AMVP candidate list, or the regular Merge candidate list. In response to the introduction of new HAPCs, the size of the sub-block-based Merge candidate list is increased from 5 to 10 and 12 for random access and low-latency B configurations, respectively. Furthermore, for the random access configuration, the size of the regular Merge candidate list is increased from 10 to 11 to accommodate the newly added regular Merge candidates.
[0095] FIG. 13 An example of generating HAPC is shown.
[0096] 2.2.3.3. Regression-based Affine Candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0097] Previously encoded affine CUs can be identified by scanning through non-adjacent positions and the affine HMVP table. FIG. 14 A diagram illustrating the regression-based affine Merge candidate derivation is shown. FIG. 14 As shown, the information of the adjacent sub-blocks of the current CU is obtained from the 4×4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used.
[0098] For each affine CU, at most two affine candidates can be derived. One has neighboring subblock information, and the other does not. All candidates generated by linear regression are pruned and merged into a single candidate subgroup. When ARMC is enabled, an ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.
[0099] 2.2.3.4. Pixel-based Affine Motion Compensation Using pixel-based affine motion compensation, when OBMC is not applied, the minimum affine sub-block size is set to 1x1 for the luma component and is always set to 1x1 for the chroma component.
[0100] 2.3. Template Matching Merge / AMVP Pattern in ECM Template Matching (TM) Merge / AMVP mode is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between the template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and the block in the reference image (i.e., the same size as the template). FIG. 15 This illustrates template matching execution over the search area surrounding the initial MV. (Example) FIG. 15 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU.
[0101] In AMVP mode, MVP candidates are determined based on template matching error, selecting the candidate that minimizes the difference between the current block and the reference block template. Then, the TM process performs MV refinement only on that specific MVP candidate. Starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), the TM refines the MVP candidate using an iterative diamond search within a search range of [-8, +8] pixels. Depending on the AMVR mode, the AMVP candidate can be further refined: a cross search is performed using full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel searches. This search process ensures that after the TM process, the MVP candidate maintains the same MV precision as indicated by the Adaptive Motion Vector Resolution (AMVR) mode.
[0102] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. Depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., when AMVR is in half-pixel mode), TMMerge can proceed up to 1 / 8-pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy. Furthermore, when TM mode is enabled, template matching can operate as a standalone process, or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. When both BM and TM are enabled for a CU, the TM search process stops at half-pixel MVD accuracy, and the resulting MV is further refined using the same model-based MVD derivation method as in DMVR.
[0103] 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, Adaptive Reordering of Merge Candidates (ARMC) is proposed to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected through the RDO process and should therefore be placed earlier in the list to reduce signaling costs.
[0104] The reordering method is applied to the regular Merge pattern, the Template Matching (TM) Merge pattern, and the Affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, the Merge candidates are reordered before the refinement process.
[0105] After the Merge candidate list is constructed, the Merge candidates are divided into several subgroups. The subgroup size is set to 5. The Merge candidates in each subgroup are reordered in ascending order based on the cost value of template matching. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered.
[0106] Template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and its corresponding reference template. FIG. 16 The template and its corresponding reference template are shown. For example... FIG. 16 As shown, the template includes a set of reconstructed samples adjacent to the current block, while the reference template is located using the same motion information as the current block. When the merge candidate utilizes bidirectional prediction, the reference samples of the merge candidate's template are also generated through bidirectional prediction.
[0107] For sub-block size equal to The sub-block-based Merge candidate has a template on top consisting of several sub-templates of size Wsub×K, and a template on the left consisting of several sub-templates of size K×Hsub. FIG. 17 This shows a template and a reference template for a block that uses the motion information of its child blocks. For example... FIG. 17 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0108] 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to TMVP, SbTMVP leverages motion fields in co-located images to facilitate more accurate MVP derivation. SbTMVP uses the same co-located images as TMVP. SbTMVP differs from TMVP primarily in two ways. First, SbTMVP enables sub-CU-level motion prediction, while TMVP predicts CU-level motion; second, compared to TMVP extracting temporal motion vectors (MVs) from co-located blocks in the co-located image (co-located blocks are the lower right or center blocks relative to the current CU), SbTMVP applies motion shifting before extracting temporal motion information from the co-located image. This motion shifting is achieved by reusing the MV from one of the spatially neighboring blocks of the current CU.
[0109] FIG. 18 The derivation of the sub-block level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first acquired. If any MV in reference list 0 and reference list 1 points to the same frame, the corresponding MV will be identified as a motion shift. Otherwise, zero MV will be used as a motion shift.
[0110] Once the motion shift is determined, a designated region within the same frame is used to derive the sub-block level motion field. Assuming... FIG. 19 As shown, Motion is used for motion shifting. So for each sub-CU, motion information of its corresponding block (the smallest motion grid covering the center sample point) in the co-location image is extracted to provide motion information, where the MV scaling operation is performed first to align the reference frame of the temporal motion vector with the reference frame of the current CU.
[0111] In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion predictions for the current CU. This sub-CU-level MVP candidate list includes the motion field generated by the SbTMVP and AFFINE methods. Specifically, only one SbTMVP candidate is included, and this SbTMVP candidate is always placed as the first entry in the constructed sub-CU-level MVP candidate list. After performing template matching-based reordering, multiple AFFINE candidates are included in the list, with those having lower costs placed earlier.
[0112] 2.6. Block Removal Process in VVC 8.6.2 Deblocking Filter Process 8.6.2.1 Overview The input to this process is the reconstructed image before removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0113] The output of this process is the modified reconstructed image after removing the blocks, i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0114] The vertical edges in the image are first filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the image are filtered. Vertical and horizontal edges in the CTB of each CTU are processed individually on a codec unit basis. The vertical edges of the codec blocks within a codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically towards the right-hand side of the codec block. The horizontal edges of the codec blocks within a codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically towards the bottom of the codec block.
[0115] Note – Although the filtering process is specified on an image basis in this specification, the filtering process can be implemented on an encoding / decoding unit basis with equivalent results, provided that the decoder properly considers the processing dependency order in order to produce the same output value.
[0116] The deblocking filter process is applied to all encoded and transformed block edges of the image, except for the following types of edges: – The edge at the boundary of the image; – When loop_filter_across_tiles_enabled_flag equals 0, the edge that coincides with the tile boundary; – When tile_group_loop_filter_across_tile_groups_enabled_flag equals 0, or tile_group_deblocking_filter_disabled_flag equals 1, the edge that coincides with the top or left boundary of the tile group; – Edges within a tile group where tile_group_deblocking_filter_disabled_flag is equal to 1; – Edges that do not correspond to the 8×8 sample grid boundary of the considered component; – Edges on both sides of the edge are predicted using the chromaticity components within the frame; – Edges of chroma transform blocks that are not part of the edges of the associated transform unit.
[0117] [Editor's note: Once the fragments are integrated, the syntax is adjusted.]
[0118] The edge type (vertical or horizontal) is represented by the variable edgeType, as specified in Table 8.17.
[0119] Table 8.17 – Names associated with edgeType
[0120] The following applies when the tile_group_deblocking_filter_disabled_flag of the current tile group is equal to 0: – The variable treeType is deduced as follows: – If tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, then treeType is set to equal DUAL_TREE_LUMA.
[0121] Otherwise, treeType is set to equal SINGLE_TREE.
[0122] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_VER are taken as inputs, and the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.
[0123] – The horizontal edge is filtered by calling the deblocking filter procedure for one direction as specified in Item 8.6.2.2, where the variable treeType, the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE), and the variable edgeType set to equal EDGE_HOR are taken as inputs, and the modified reconstructed image after deblocking (i.e., the array recPictureL, and arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 or treeType is equal to SINGLE_TREE) are taken as outputs.
[0124] – When tile_group_type equals 1 and qtbtt_dual_tree_intra_flag equals 1, the following applies: – The variable treeType is set to equal DUAL_TREE_CHROMA.
[0125] – Vertical edges are filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the reconstructed image before deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_VER are taken as inputs, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) are taken as outputs.
[0126] – The horizontal edge is filtered by calling the deblocking filter procedure for one direction as specified in entry 8.6.2.2, where the variable treeType, the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr), and the variable edgeType set to equal EDGE_HOR are taken as inputs, and the modified reconstructed image after deblocking (i.e., arrays recPictureCb and recPictureCr) are taken as outputs.
[0127] 8.6.2.2 Deblocking filter process for one direction The input to this process is: – A variable treeType that specifies whether a single tree (SINGLE_TREE) or a dual tree is used to partition the CTU, and when using a dual tree, whether the currently processed component is the luma component (DUAL_TREE_LUMA) or the chroma component (DUAL_TREE_CHROMA); – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the reconstructed picture before deblocking, i.e., the array recPictureL; – When ChromaArrayType is not equal to 0, and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr; – A variable edgeType that specifies whether a vertical edge (EDGE_VER) or a horizontal edge (EDGE_HOR) is filtered.
[0128] The output of this process is the modified reconstructed picture after deblocking, i.e.: – When treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; – When ChromaArrayType is not equal to 0, and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the arrays recPictureCb and recPictureCr.
[0129] For each coding unit with a coded block width of log2CbW, a coded block height of log2CbH, and the position of the top-left sample of the coded block being (xCb, yCb), when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered through the following ordered steps: 1. The coded block width nCbW is set to be equal to 1 << log2CbW, and the coded block height nCbH is set to be equal to 1 << log2CbH.
[0130] 2. The variable filterEdgeFlag is derived as follows: <000– The left boundary of the current codec block is the left boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.
[0132] – The left boundary of the current codec block is the left boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0133] Otherwise, if edgeType equals EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set to 0: – The upper boundary of the current luminance codec block is the upper boundary of the image.
[0134] – The upper boundary of the current codec block is the upper boundary of the tile, and loop_filter_across_tiles_enabled_flag is equal to 0.
[0135] – The upper boundary of the current codec block is the upper boundary of the tile group, and tile_group_loop_filter_across_tile_groups_enabled_flag is equal to 0.
[0136] Otherwise, filterEdgeFlag is set to 1.
[0137] [Editor's note: Once the fragments are integrated, the syntax is adjusted.]
[0138] 3. All elements of the two-dimensional (nCbW) x (nCbH) array edgeFlags are initialized to 0.
[0139] 4. The derivation process for the transform block boundary specified in Item 8.6.2.3 is invoked, where the position (xB0, yB0) is set to equal to (0, 0), the block width nTbW is set to equal to nCbW, the block height nTbH is set to equal to nCbH, the variables treeType, filterEdgeFlag, array edgeFlags, and edgeType are taken as input, and the modified array edgeFlags is taken as output.
[0140] 5. The derivation process for the codec subblock boundary specified in Item 8.6.2.4 is invoked, with the position (xCb, yCb), codec block width nCbW, codec block height nCbH, array edgeFlags, and variable edgeType as input, and the modified array edgeFlags as output.
[0141] 6. The image sample array recPicture is derived as follows: – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, recPicture is set to equal to the array of reconstructed luminance image samples recPictureL before deblocking.
[0142] Otherwise (treeType equals DUAL_TREE_CHROMA), recPicture is set to equal the array of reconstructed chroma image samples recPictureCb before deblocking.
[0143] 7. The derivation process for the boundary filter strength specified in item 8.6.2.5 is invoked, with the image sample array recPicture, the luminance position (xCb, yCb), the codec block width nCbW, the codec block height nCbH, the variable edgeType, and the array edgeFlags as inputs, and the (nCbW) x (nCbH) array verBs as outputs.
[0144] 8. The edge filtering process is invoked as follows: – If edgeType equals EDGE_VER, then the vertical edge filtering procedure for the codec unit, as specified in entry 8.6.2.6.1, is invoked, with the variable treeType, the reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array verBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.
[0145] Otherwise, if edgeType equals EDGE_HOR, the horizontal edge filtering procedure for the codec unit as specified in entry 8.6.2.6.2 is invoked, with the variable treeType, the modified reconstructed image before deblocking (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA), position (xCb, yCb), codec block width nCbW, codec block height nCbH, and array horBs as input, and the modified reconstructed image (i.e., the array recPictureL when treeType equals SINGLE_TREE or DUAL_TREE_LUMA, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0 and treeType equals SINGLE_TREE or DUAL_TREE_CHROMA) as output.
[0146] 8.6.2.3 Derivation of the Transform Block Boundary The input to this process is: – Position (xB0, yB0), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; – The variable nTbW specifies the width of the current block; – The variable nTbH specifies the height of the current block; – The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). – Variable filterEdgeFlag; – A two-dimensional (nCbW) x (nCbH) array edgeFlags; – The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered.
[0147] The output of this process is a modified two-dimensional (nCbW) x (nCbH) array edgeFlags.
[0148] The maximum transform block size, maxTbSize, is derived as follows: maxTbSize = (treeType= =DUAL_TREE_CHROMA) ? MaxTbSizeY / 2 :MaxTbSizeY (8 862).
[0149] Depending on maxTbSize, the following applies: - If nTbW is greater than maxTbSize or nTbH is greater than maxTbSize, then the following ordered steps apply.
[0150] 1. The variables newTbW and newTbH are derived as follows: newTbW = ( nTbW>maxTbSize ) ? ( nTbW / 2 ) : nTbW (8 863) newTbH = ( nTbH>maxTbSize ) ? ( nTbH / 2 ) :nTbH (8 864).
[0151] 2. The derivation process for the transform block boundary specified in this entry is invoked, with the position (xB0, yB0), the variable nTbW set to be equal to newTbW, the variable nTbH set to be equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType as inputs, and the output being a modified version of the array edgeFlags.
[0152] 3. If nTbW is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0 + newTbW, yTb0), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0153] 4. If nTbH is greater than maxTbSize, the derivation process of the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal to (xTb0, yTb0 + newTbH), the variable nTbW is set to equal to newTbW, the variable nTbH is set to equal to newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0154] 5. If nTbW is greater than maxTbSize and nTbH is greater than maxTbSize, then the derivation procedure for the transform block boundary specified in this entry is invoked, where the luminance position (xTb0, yTb0) is set to equal (xTb0 + newTbW, yTb0 + newTbH), the variable nTbW is set to equal newTbW, the variable nTbH is set to equal newTbH, the variable filterEdgeFlag, the array edgeFlags, and the variable edgeType are taken as inputs, and the output is a modified version of the array edgeFlags.
[0155] - Otherwise, the following applies: - If edgeType equals EDGE_VER, then the value of edgeFlags[xB0][yB0 + k] (k = 0..nTbH-1) is derived as follows: - If xB0 equals 0, then edgeFlags[xB0][yB0 + k] is set to equal filterEdgeFlag.
[0156] Otherwise, edgeFlags[xB0][yB0 + k] is set to 1.
[0157] - Otherwise (edgeType equals EDGE_HOR), the value of edgeFlags[xB0 + k][yB0] (k = 0..nTbW-1) is derived as follows: - If yB0 equals 0, then edgeFlags[xB0 + k][yB0] is set to equal filterEdgeFlag.
[0158] Otherwise, edgeFlags[xB0 + k][yB0] is set to 1.
[0159] 8.6.2.4 Derivation of Encoding / Decoding Sub-Block Boundaries The input to this process is: - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block; - A two-dimensional (nCbW) x (nCbH) array edgeFlags; - The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered.
[0160] The output of this process is a modified two-dimensional (nCbW) x (nCbH) array edgeFlags.
[0161] The number of horizontal codec sub-blocks, numSbX, and the number of vertical codec sub-blocks, numSbY, are derived as follows: - If CupredMode[xCb][yCb] == MODE_INTRA, then numSbX and numSbY are both set to 1.
[0162] Otherwise, numSbX and numSbY are set to equal NumSbX[xCb][yCb] and NumSbY[xCb][yCb], respectively.
[0163] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER and numSbX is greater than 1, then the following applies to i = 1..min( (nCbW / 8 ) - 1, numSbX - 1), k = 0..nCbH - 1: .
[0164] Otherwise, if edgeType equals EDGE_HOR and numSbY is greater than 1, the following applies to j = 1..min( ( nCbH / 8 ) - 1, numSbY - 1 ), k = 0..nCbW - 1: .
[0165] 8.6.2.5 Derivation of Boundary Filter Strength The input to this process is: - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block; - The variable edgeType specifies whether vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR) are filtered; - A two-dimensional (nCbW) x (nCbH) array edgeFlags.
[0166] The output of this process is a two-dimensional (nCbW) x (nCbH) array bS that specifies the boundary filter strength.
[0167] The variables xDi, yDj, xN, and yN are derived as follows: - If edgeType equals EDGE_VER, then xDi is set to equal to (i<<3), yDj is set to equal to (j<<2), xN is set to equal to Max(0, (nCbW / 8) - 1), and yN is set to equal to (nCbH / 4) - 1.
[0168] - Otherwise (edgeType equals EDGE_HOR), xDi is set to equal to (i<<2), yDj is set to equal to (j<<3), xN is set to equal to (nCbW / 4) - 1, and yN is set to equal to Max(0, (nCbH / 8) - 1).
[0169] For xDi where i = 0..xN and yDj where j = 0..yN, the following applies: - If edgeFlags[xDi][yDj] equals 0, then the variable bS[xDi][yDj] is set to equal 0.
[0170] - Otherwise, the following applies: - The sample values p0 and q0 are derived as follows: - If edgeType equals EDGE_VER, then p0 is set to equal recPicture[xCb + xDi - 1][yCb + yDj], and q0 is set to equal recPicture[xCb + xDi][yCb + yDj].
[0171] - Otherwise (edgeType equals EDGE_HOR), p0 is set to equal recPicture [ xCb + xDi ][yCb + yDj - 1 ], and q0 is set to equal recPicture [ xCb + xDi ][ yCb + yDj ].
[0172] - The variable bS[ xDi ][ yDj ] is derived as follows: - If sample p0 or q0 is in the codec block of a codec unit that uses intra-prediction mode coding and decoding, then bS[xDi][yDj] is set to equal 2.
[0173] - Otherwise, if the block edge is also the transform block edge, and sample p0 or q0 is in a transform block containing one or more non-zero transform coefficient levels, then bS[ xDi ][ yDj ] is set to equal to 1.
[0174] Otherwise, bS[xDi][yDj] is set to 1 if one or more of the following conditions are true: - For prediction of a codec subblock containing sample p0, use a different reference picture or a different number of motion vectors than for prediction of a codec subblock containing sample q0.
[0175] Note 1 – Determining whether the reference pictures used for two codec subblocks are the same or different is based solely on which pictures are referenced, without considering whether indices in reference picture list 0 or reference picture list 1 are used to form the prediction, and also without considering whether the index positions within the reference picture lists are different.
[0176] Note 2 – The number of motion vectors used to predict the top left sample coverage (xSb, ySb) of the codec subblock is equal to PredFlagL0[xSb][ySb] + PredFlagL1[xSb][ySb].
[0177] - A motion vector is used to predict the codec subblock containing sample p0, and a motion vector is used to predict the codec subblock containing sample q0, and the absolute difference between the horizontal or vertical components of the motion vectors used is greater than or equal to 4 (in quarter-luminance samples).
[0178] - Two motion vectors and two different reference images are used to predict the codec subblock containing sample p0, and two motion vectors and the same two reference images are used to predict the codec subblock containing sample q0. When predicting the two codec subblocks, the absolute difference between the horizontal or vertical components of the two motion vectors used for the same reference image is greater than or equal to 4 (in quarter-luminance samples).
[0179] - Two motion vectors for the same reference image are used to predict a codec subblock containing sample p0, and two motion vectors for the same reference image are used to predict a codec subblock containing sample q0, and both of the following conditions are true: - The absolute difference between the horizontal or vertical components of the motion vector in list 0 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in list 1 used when predicting two codec subblocks is greater than or equal to 4 (in quarter-luminance samples).
[0180] - The absolute difference between the horizontal or vertical components of the motion vector in List 0 used when predicting the codec subblock containing sample p0 and the motion vector in List 1 used when predicting the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples), or the absolute difference between the horizontal or vertical components of the motion vector in List 1 used when predicting the codec subblock containing sample p0 and the motion vector in List 0 used when predicting the codec subblock containing sample q0 is greater than or equal to 4 (in quarter-luminance samples).
[0181] Otherwise, the variable bS[ xDi ][ yDj ] is set to 0.
[0182] 8.6.2.6 Edge Filtering Process 8.6.2.6.1 Vertical Edge Filtering Process The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, reconstruct the image before the block, i.e., the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block.
[0183] The output of this process is the modified reconstructed image after removing the blocks, i.e.: - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.
[0184] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the filtering process for the edges in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal (nCbH / 4) - 1.
[0185] 2. For xDk of k = 0..nN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter intensity bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0186] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.
[0187] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the filtering process for the edges in the chroma codec block of the current codec unit consists of the following ordered steps: 1. Variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal Max(0, (nCbH / 8) - 1).
[0188] 2. The variable edgeSpacing is set to equal to 8 / SubWidthC.
[0189] 3. The variable edgeSections is set to equal yN. (2 / SubHeightC).
[0190] 4. For k = 0..xN, the expression equal to k For edgeSpacing, xDk and m = 0..edgeSections, yDm is equal to m<<2, the following applies: - When bS[ xDk SubWidthC ][ yDm When SubHeightC equals 2 and (((xCb / SubWidthC + xDk)>>3)<<3) equals xCb / SubWidthC + xDk, the following ordered steps apply: a. The filtering procedure for the edges of the chroma block, as specified in entry 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.
[0191] b. The filtering procedure for the edges of the chroma block, as specified in Item 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_VER, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.
[0192] When treeType equals DUAL_TREE_CHROMA, the filtering process for the edges in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal (nCbH / 4) - 1.
[0193] 2. For xDk of k = 0..xN and equal to k<<3, and yDm of m = 0..yN and equal to m<<2, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0194] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCb is taken as output.
[0195] c. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType, set to be equal to EDGE_VER, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCr is taken as output.
[0196] 8.6.2.6.2 Horizontal Edge Filtering Process The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, reconstruct the image before the block, i.e., the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - The variable nCbW specifies the width of the current codec block; - The variable nCbH specifies the height of the current codec block.
[0197] The output of this process is the modified reconstructed image after removing the blocks, i.e.: - When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the array recPictureL; - When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, arrays recPictureCb and recPictureCr.
[0198] When treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the filtering process for the edges in the luma codec block of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8) - 1), and xN is set to equal (nCbW / 4) - 1.
[0199] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthY are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0200] b. The filtering procedure for block edges as specified in Item 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the luminance image sample array recPictureL, the position of the luminance codec block (xCb, yCb), the luminance position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decision dEp, dEp and dEq, and the variable tC are taken as inputs, and the modified luminance image sample array recPictureL is taken as output.
[0201] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE, the filtering process for the edges in the chroma codec block of the current codec unit consists of the following ordered steps: 1. Variable xN is set to equal Max(0, (nCbW / 8) - 1), and yN is set to equal Max(0, (nCbH / 8) - 1).
[0202] 2. The variable edgeSpacing is set to equal to 8 / SubHeightC.
[0203] 3. The variable edgeSections is set to equal xN. (2 / SubWidthC).
[0204] 4. For m = 0..yN, the equality is... The following applies to yDm and xDk where k = 0..edgeSections equals k<<2: - when When yCb / SubHeightC + yDm equals 2 and (((yCb / SubHeightC + yDm)>>3)<<3) equals yCb / SubHeightC + yDm, the following ordered steps apply: a. The filtering procedure for the edges of the chroma block, as specified in entry 8.6.2.6.5, is invoked, with the chroma picture sample array recPictureCb, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cb_qp_offset as input, and the modified chroma picture sample array recPictureCb as output.
[0205] b. The filtering procedure for the chroma block edges as specified in Item 8.6.2.6.5 is invoked, with the chroma picture sample array recPictureCr, the position of the chroma codec block (xCb / SubWidthC, yCb / SubHeightC), the chroma position of the block (xDk, yDm), the variable edgeType set to equal EDGE_HOR, and the variable cQpPicOffset set to equal pps_cr_qp_offset as input, and the modified chroma picture sample array recPictureCr as output.
[0206] When treeType equals DUAL_TREE_CHROMA, the filtering process for the edges in the two chroma codec blocks of the current codec unit consists of the following ordered steps: 1. The variable yN is set to equal Max(0, (nCbH / 8) - 1), and xN is set to equal (nCbW / 4) - 1.
[0207] 2. For yDm equal to m<<3 and xDk equal to k<<2 for m = 0..yN, the following applies: - When bS[xDk][yDm] is greater than 0, the following ordered steps apply: a. The decision process for block edges as specified in entry 8.6.2.6.3 is invoked, where treeType, the image sample array recPicture set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the position of the chroma block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the boundary filter strength bS[xDk][yDm], and the bit depth bD set to be equal to BitDepthC are taken as inputs, and the decisions dE, dEp, and dEq and the variable tC are taken as outputs.
[0208] b. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCb, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCb is taken as output.
[0209] c. The filtering procedure for block edges as specified in entry 8.6.2.6.4 is invoked, wherein the image sample array recPicture, set to be equal to the chroma image sample array recPictureCr, the position of the chroma codec block (xCb, yCb), the chroma position of the block (xDk, yDm), the variable edgeType set to be equal to EDGE_HOR, the decisions dE, dEp, and dEq, and the variable tC are taken as input, and the modified chroma image sample array recPictureCr is taken as output.
[0210] 8.6.2.6.3 Decision-making process for block edges The input to this process is: - The variable treeType specifies whether to use a single tree (SINGLE_TREE) or a dual tree to segment the CTU, and when using a dual tree, whether the current processing is luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA). - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - Position (xBl, yBl), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - Variable bS, which specifies the boundary filter strength; - Variable bD, which specifies the bit depth of the current component.
[0211] The output of this process is: – Includes decision variables dE, dEp, and dEq; – Variable tC.
[0212] If edgeType equals EDGE_VER, then the sample values pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ] (8 867) pi,k = recPictureL[ xCb + xBl - i - 1 ][ yCb + yBl + k ] (8 868).
[0213] Otherwise (edgeType equals EDGE_HOR), the sample values pi,k and qi,k for i = 0..3 and k = 0 and 3 are derived as follows: qi,k = recPicture[ xCb + xBl + k ][ yCb + yBl + i ] (8 869) pi,k = recPicture[ xCb + xBl + k ][ yCb + yBl - i - 1 ] (8 870).
[0214] The variable qpOffset is derived as follows: - If sps_ladf_enabled_flag equals 1 and treeType equals SINGLE_TREE or DUAL_TREE_LUMA, then the following applies: - The variable lumaLevel for reconstructing the brightness level is derived as follows: lumaLevel = ( ( p0,0 + p0,3 + q0,0 + q0,3 )>>2 ) (8 871).
[0215] - The variable qpOffset is set to equal to sps_ladf_lowest_interval_qp_offset and modified as follows: for( i = 0; i <sps_num_ladf_intervals_minus2 + 1; i++ ) { if( lumaLevel>SpsLadfIntervalLowerBound[ i + 1 ] ) qpOffset = sps_ladf_qp_offset[ i ] (8 872) else break } - Otherwise (treeType equals DUAL_TREE_CHROMA), qpOffset is set to 0.
[0216] The variables QpQ and QpP are derived as follows: - If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, then QpQ and QpP are set to the QpY value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.
[0217] - Otherwise (treeType equals DUAL_TREE_CHROMA), QpQ and QpP are set to the QpC value of the codec unit that includes the codec block containing samples q0,0 and p0,0, respectively.
[0218] The variable qP is derived as follows: qP = ( ( QpQ + QpP + 1 )>>1 ) + qpOffset (8 873).
[0219] The value of variable β′ is determined based on the quantization parameter Q derived as follows, as specified in Table 8.18: Q = Clip3( 0, 63, qP + ( tile_group_beta_offset_div2<<1 ) ) (8 874) Where tile_group_beta_offset_div2 is the value of the syntax element tile_group_beta_offset_div2 that contains the patch group of sample points q0,0.
[0220] The variable β is derived as follows: .
[0221] The value of variable tC′ is determined based on the quantization parameter Q derived as follows, as specified in Table 8.18: Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that contains the patch group of sample points q0,0.
[0222] The variable tC is derived as follows: .
[0223] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:
[0224] dpq0 = dp0 + dq0 (8 882) dpq3 = dp3 + dq3 (8 883) dp = dp0 + dp3 (8 884) dq = dq0 + dq3 (8 885) d = dpq0 + dpq3 (8886).
[0225] 2. The variables dE, dEp, and dEq are set to 0.
[0226] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.
[0227] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.
[0228] c. The variable dpq is set to equal 2. dpq3.
[0229] d. For the sample location (xCb + xBl, yCb + yBl + 3), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.
[0230] e. The variable dE is set to equal 1.
[0231] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.
[0232] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.
[0233] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.
[0234] - Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The variables dpq0, dpq3, dp, dq, and d are derived as follows:
[0235] dpq0 = dp0 + dq0 (8 891) dpq3 = dp3 + dq3 (8 892) dp = dp0 + dp3 (8 893) dq = dq0 + dq3 (8 894) d = dpq0 + dpq3 (8 895).
[0236] 2. The variables dE, dEp, and dEq are set to 0.
[0237] 3. When d is less than β, the following ordered steps apply: a. The variable dpq is set to equal 2. dpq0.
[0238] b. For the sample location (xCb + xBl, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values p0,0, p3,0, q0,0 and q3,0, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam0.
[0239] c. The variable dpq is set to equal 2. dpq3.
[0240] d. For the sample location (xCb + xBl + 3, yCb + yBl), the decision process for the sample as specified in Item 8.6.2.6.6 is invoked, where the sample values p0,3, p3,3, q0,3 and q3,3, the variables dpq, β and tC are taken as inputs, and the output is assigned to the decision dSam3.
[0241] e. The variable dE is set to equal 1.
[0242] f. When dSam0 equals 1 and dSam3 equals 1, the variable dE is set to equal 2.
[0243] g. When dp is less than (β + (β>>1))>>3, the variable dEp is set to equal to 1.
[0244] h. When dq is less than (β + (β>>1))>>3, the variable dEq is set to equal to 1.
[0245] Table 8.18 - Deriving threshold variables β′ and tC′ from input Q
[0246] 8.6.2.6.4 Filtering process for block edges The input to this process is: - Image sample array recPicture; - Position (xCb, yCb), which specifies the position of the top-left sample of the current codec block relative to the top-left sample of the current image; - Position (xBl, yBl), which specifies the position of the top-left sample of the current block relative to the top-left sample of the current codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - Variables including dE, dEp, and dEq for decision-making; - Variable tC.
[0247] The output of this process is a modified image sample array recPicture.
[0248] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then the following ordered steps apply: 1. The sample values pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + i ][ yCb + yBl + k ] (8 896) pi,k = recPictureL[ xCb + xBl - i - 1 ][ yCb + yBl + k ] (8 897).
[0249] 2. When dE is not equal to 0, for each sample point location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: a. The filtering procedure for the sample points as specified in item 8.6.2.6.7 is invoked, wherein the sample point values pi,k,qi,k (i = 0..3), the positions (xPi, yPi) are set to equal (xCb + xBl - i - 1, yCb + yBl + k) and (xQi, yQi) are set to equal (xCb + xBl + i, yCb + yBl + k) (i = 0..2), the decision dE, the variables dEp and dEq, and the variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample point values pi' and qj' are taken as outputs.
[0250] b. When nDp is greater than 0, the filtered sample value pi' of i = 0..nDp - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl - i - 1 ][ yCb + yBl + k ] = pi' (8 898).
[0251] c. When nDq is greater than 0, the filtered sample value qj' of j = 0..nDq - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + j ][ yCb + yBl + k ] = qj' (8 899).
[0252] - Otherwise (edgeType equals EDGE_HOR), the following ordered steps apply: 1. The sample values pi,k and qi,k for i = 0..3 and k = 0..3 are derived as follows: qi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl + i ] (8 900) pi,k = recPictureL[ xCb + xBl + k ][ yCb + yBl - i - 1 ] (8 901).
[0253] 2. When dE is not equal to 0, for each sample point location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: a. The filtering procedure for the sample points as specified in Item 8.6.2.6.7 is invoked, wherein the sample point values pi,k,qi,k (i = 0..3), the positions (xPi, yPi) are set to equal (xCb + xBl + k, yCb + yBl - i - 1) and (xQi, yQi) are set to equal (xCb + xBl + k, yCb + yBl + i) (i = 0..2), the decision dE, the variables dEp and dEq, and the variable tC are taken as inputs, and the number of filtered samples nDp and nDq from each side of the block boundary and the filtered sample point values pi' and qj' are taken as outputs.
[0254] b. When nDp is greater than 0, the filtered sample value pi' of i = 0..nDp - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl - i - 1 ] = pi' (8 902).
[0255] c. When nDq is greater than 0, the filtered sample value qj' of j = 0..nDq - 1 replaces the corresponding sample in the sample array recPicture, as shown below: recPicture[ xCb + xBl + k ][ yCb + yBl + j ] = qj' (8 903).
[0256] 8.6.2.6.5 Filtering process for chroma block edges This procedure is invoked only if ChromaArrayType is not equal to 0.
[0257] The input to this process is: - Chroma image sample array s′; - Chroma position (xCb, yCb), which specifies the position of the top-left sample of the current chroma codec block relative to the top-left chroma sample of the current image; - Chroma position (xBl, yBl), which specifies the position of the top-left sample of the current chroma block relative to the top-left sample of the current chroma codec block; - The variable edgeType specifies whether filtering is applied to vertical edges (EDGE_VER) or horizontal edges (EDGE_HOR); - The variable cQpPicOffset specifies the offset of the image-level color quantization parameter.
[0258] The output of this process is the modified chroma image sample array s′.
[0259] If edgeType equals EDGE_VER, then the values of pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + i ][ yCb + yBl + k ] (8 904) pi,k = s′[ xCb + xBl - i - 1 ][ yCb + yBl + k ] (8 905).
[0260] Otherwise (edgeType equals EDGE_HOR), the sample values pi and qi for i = 0..1 and k = 0..3 are derived as follows: qi,k = s′[ xCb + xBl + k ][ yCb + yBl + i ] (8 906) pi,k = s′[ xCb + xBl + k ][ yCb + yBl - i - 1 ] (8 907).
[0261] The variables QpQ and QpP are set to be equal to the QpY value of the codec unit, which includes the codec block containing samples q0,0 and p0,0, respectively.
[0262] If ChromaArrayType equals 1, then the variable QpC is determined based on the index qPi derived as follows, as specified in Table 8.15: qPi = ( ( QpQ + QpP + 1 )>>1 ) + cQpPicOffset (8 908).
[0263] Otherwise (ChromaArrayType is greater than 1), the variable QpC is set to equal Min(qPi, 63).
[0264] Note – The variable cQpPicOffset provides adjustments to the values of pps_cb_qp_offset or pps_cr_qp_offset based on whether the filtered chrominance component is a Cb or Cr component. However, to avoid the need to change the adjustment amount within the image, the filtering process does not include adjustments to the values of tile_group_cb_qp_offset or tile_group_cr_qp_offset.
[0265] The value of variable tC′ is determined based on the colorimetric parameter Q derived as follows, as specified in Table 8.18: Q = Clip3( 0, 65, QpC + 2 + ( tile_group_tc_offset_div2<<1 ) ) (8909) Where tile_group_tc_offset_div2 is the value of the syntax element tile_group_tc_offset_div2 that contains the patch group of sample points q0,0.
[0266] The variable tC is derived as follows: .
[0267] Depending on the value of edgeType, the following applies: - If edgeType equals EDGE_VER, then for each sample location (xCb + xBl, yCb + yBl + k), k = 0..3, the following ordered steps apply: 1. The filtering procedure for chromaticity samples as specified in Item 8.6.2.6.8 is invoked, wherein the sample values pi,k,qi,k (i = 0..1), positions (xCb + xBl - 1, yCb + yBl + k) and (xCb + xBl, yCb + yBl + k) and variable tC are taken as inputs, and the filtered sample values p0′ and q0′ are taken as outputs.
[0268] 2. Replace the corresponding samples in the sample array s' with the filtered sample values p0′ and q0′, as shown below: s′[ xCb + xBl ][ yCb + yBl + k ]= q0′ (8 911) s′[ xCb + xBl - 1 ][ yCb + yBl + k ] = p0′ (8 912).
[0269] - Otherwise (edgeType equals EDGE_HOR), for each sample location (xCb + xBl + k, yCb + yBl), k = 0..3, the following ordered steps apply: 1. The filtering procedure for chromaticity samples as specified in item 8.6.2.6.8 is invoked, wherein the sample values pi,k,qi,k (i = 0..1), positions (xCb + xBl + k, yCb + yBl - 1) and (xCb + xBl + k, yCb + yBl) and variable tC are taken as inputs, and the filtered sample values p0′ and q0′ are taken as outputs.
[0270] 2. Replace the corresponding samples in the sample array s' with the filtered sample values p0′ and q0′, as shown below: s′[ xCb + xBl + k ][ yCb + yBl ]= q0′ (8 913) s′[ xCb + xBl + k ][ yCb + yBl - 1 ] = p0′ (8 914).
[0271] 8.6.2.6.6 Decision-making process for sample points The input to this process is: - Sample values p0, p3, q0, and q3; - Variables dpq, β, and tC.
[0272] The output of this process is the variable dSam, which contains the decision.
[0273] The variable dSam is defined as follows: - If dpq is less than (β>>2), Abs(p3 - p0) + Abs(q0 - q3) is less than (β>>3), and Abs(p0 - q0) is less than (5). If tC + 1 )>>1, then dSam is set to equal to 1.
[0274] Otherwise, dSam is set to 0.
[0275] 8.6.2.6.7 Filtering process for sample points The input to this process is: - Sample values pi and qi for i = 0..3; - The positions of pi and qi are (xPi, yPi) and (xQi, yQi), where i = 0..2; - Variable dE; - dEp and dEq, respectively, contain the decision variables for filtering samples p1 and q1; - Variable tC.
[0276] The output of this process is: - The number of filtered samples, nDp and nDq; - Filtered sample values pi′ and qj′ of i = 0..nDp - 1 and j = 0..nDq - 1.
[0277] Depending on the value of dE, the following applies: - If variable dE equals 2, then nDp and nDq are both set to equal 3, and the following strong filtering applies: p0′ = Clip3( p0 - 2 tC, p0 + 2 tC, (p2 + 2) p1 + 2 p0 + 2 q0 + q1 + 4)>>3) (8 915) p1′ = Clip3( p1 - 2 tC, p1 + 2 tC, ( p2 + p1 + p0 + q0 + 2 )>>2 ) (8 916) p2′ = Clip3( p2 - 2 tC, p2 + 2 tC, ( 2 p3 + 3 p2 + p1 + p0 + q0 +4 )>>3 ) (8 917) q0′ = Clip3( q0 - 2 tC, q0 + 2 tC, (p1 + 2) p0 + 2 q0 + 2 q1 + q2 + 4)>>3) (8 918) q1′ = Clip3( q1 - 2 tC, q1 + 2 tC, ( p0 + q0 + q1 + q2 + 2 )>>2 ) (8 919) q2′ = Clip3(q2 - 2) tC, q² + 2 tC, (p0 + q0 + q1 + 3) q² + 2 q3 + 4)>>3) (8 920).
[0278] - Otherwise, nDp and nDq are both set to 0, and the following weak filtering applies: - The following applies:
[0279] - when Less than tC At 10:00, the following sequential steps apply: - The filtered sample values p0′ and q0′ are defined as follows: = Clip3( -tC, tC, (8 922) p0′ = Clip1Y( p0 + (8 923) q0′ = Clip1Y( q0 - ) (8 924).
[0280] - When dEp equals 1, the filtered sample value p1′ is defined as follows: p = Clip3( -( tC>>1 ), tC>>1, ( ( ( p2 + p0 + 1 )>>1 ) - p1 + )>>1 ) (8 925 ) p1′ = Clip1Y( p1 + p ) (8 926).
[0281] - When dEq equals 1, the filtered sample value q1′ is defined as follows: q = Clip3( -( tC>>1 ), tC>>1, ( ( ( q2 + q0 + 1 )>>1 ) - q1 - )>>1 ) (8 927) q1′ = Clip1Y( q1 + q) (8 928).
[0282] - nDp is set to equal dEp + 1, and nDq is set to equal dEq + 1.
[0283] nDp is set to 0 when nDp is greater than 0 and one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 ][ yP0 ] equals 1.
[0284] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.
[0285] nDq is set to 0 when nDq is greater than 0 and one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xQ0 ][ yQ0 ] equals 1.
[0286] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.
[0287] 8.6.2.6.8 Filtering process for chromaticity samples This procedure is invoked only if ChromaArrayType is not equal to 0.
[0288] The input to this process is: - Colorimetric sample values pi and qi for i = 0..1; - The chromaticity positions of p0 and q0 are (xP0, yP0) and (xQ0, yQ0); - Variable tC.
[0289] The output of this process is the filtered sample values p0′ and q0′.
[0290] The filtered sample values p0′ and q0′ are derived as follows: = Clip3( -tC, tC, ( ( ( ( q0 - p0 )<<2 ) + p1 - q1 + 4 )>>3 ) ) (8929) p0′ = Clip1C( p0 + (8 930) q0′ = Clip1C( q0 - ) (8 931).
[0291] The filtered sample value p0′ is replaced by the corresponding input sample value p0 when one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[ xP0 SubWidthC ][yP0 SubHeightC equals 1.
[0292] - The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample p0, is equal to 1.
[0293] The filtered sample value q0′ is replaced by the corresponding input sample value q0 when one or more of the following conditions are true: - pcm_loop_filter_disabled_flag equals 1, and pcm_flag[xQ0] SubWidthC ][yQ0 SubHeightC equals 1.
[0294] – The cu_transquant_bypass_flag of the codec unit, including the codec block containing sample q0, is equal to 1.
[0295] 8.6.3 Sample point adaptive compensation process 8.6.3.1 Overview The input to this process is the array recPictureL of reconstructed image samples before adaptive compensation, and the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0.
[0296] The output of this process is a modified reconstructed image sample array saoPictureL after sample adaptive compensation, and arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0.
[0297] This process is performed on top of CTB after the deblocking filter process for the decoded image is completed.
[0298] The sample values in the modified reconstructed picture sample array saoPictureL, and the sample values in the arrays saoPictureCb and saoPictureCr when ChromaArrayType is not equal to 0, are initially set to be equal to the sample values in the reconstructed picture sample array recPictureL, and the sample values in the arrays recPictureCb and recPictureCr when ChromaArrayType is not equal to 0, respectively.
[0299] For each CTU with CTB position (rx, ry), where rx = 0..PicWidthInCtbsY - 1 and ry = 0..PicHeightInCtbsY - 1, the following applies: – When the tile_group_sao_luma_flag of the current slice group is equal to 1, the CTB modification process as specified in entry 8.6.3.2 is called, where recPicture is set to be equal to recPictureL, cIdx is set to be equal to 0, (rx, ry), and both nCtbSw and nCtbSh are set to be equal to CtbSizeY as input, and the modified luma picture sample array saoPictureL is used as output. <00– When ChromaArrayType is not equal to 0 and the tile_group_sao_chroma_flag of the current slice group is equal to 1, the CTB modification process as specified in item 8.6.3.2 is called, where recPicture is set to be equal to recPictureCr, cIdx is set to be equal to 2, (rx, ry), nCtbSw is set to be equal to (1 << CtbLog2SizeY) / SubWidthC, and nCtbSh is set to be equal to (1 << CtbLog2SizeY) / SubHeightC as inputs, and the modified chroma picture sample array saoPictureCr is used as the output.
[0302] 8.6.3.2 CTB Modification Process The inputs to this process are: – The picture sample array recPicture for color component cIdx; – The variable cIdx that specifies the color component index; – A pair of variables (rx, ry) that specify the CTB position; [[ID= SubWidthC, ySj SubHeightC) (8 934).
[0308] For all sample locations (xSi, ySj) and (xYi, yYj) of i = 0..nCtbSw - 1 and j = 0..nCtbSh - 1, depending on the values of pcm_loop_filter_disabled_flag, pcm_flag[xYi][yYj], and cu_transquant_bypass_flag of the codec unit that covers the codec block covering recPicture[xSi][ySj], the following applies: – If one or more of the following conditions are true, then saoPicture[xSi][ySj] will not be modified: – pcm_loop_filter_disabled_flag and pcm_flag[ xYi ][ yYj ] are both equal to 1.
[0309] – cu_transquant_bypass_flag equals 1.
[0310] – SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 0.
[0311] [Editor's Note: The highlighted section will be modified based on future decision changes / quantification bypasses.] Otherwise, if SaoTypeIdx[cIdx][rx][ry] equals 2, then the following ordered steps apply: 1. The values of hPos[k] and vPos[k] for k = 0..1 are based on SaoEoClass[cIdx][rx][ry] as specified in Table 8.19.
[0312] 2. The variable edgeIdx is derived as follows: – The modified sample point locations (xSik′, ySjk′) and (xYik′, yYjk′) are derived as follows: (xSik′, ySjk′) = (xSi + hPos[ k ], ySj + vPos[ k ]) (8 935) (xYik′, yYjk′) = (cIdx = = 0) ? (xSik′, ySjk′) : (xSik′ SubWidthC,ySjk′ SubHeightC) (8 936).
[0313] – If one or more of the following conditions are true for all sample locations (xSik′, ySjk′) and (xYik′, yYjk′) at k = 0..1, then edgeIdx is set to equal to 0: – The sample point at location (xSik′, ySjk′) is outside the image boundary.
[0314] – The sample points at location (xSik′, ySjk′) belong to different patch groups, and one of the following two conditions is true: – MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSi ][ ySj ] belongs is equal to 0.
[0315] – MinTbAddrZs[ xYi>>MinTbLog2SizeY ][ yYj>>MinTbLog2SizeY ] is less than MinTbAddrZs[ xYik′>>MinTbLog2SizeY ][ yYjk′>>MinTbLog2SizeY ], and the tile_group_loop_filter_across_tile_groups_enabled_flag in the tile group to which the sample recPicture[ xSik′ ][ ySjk′ ] belongs is equal to 0.
[0316] – loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at position (xSik′, ySjk′) belongs to a different tile.
[0317] [Editor's Note: When merging slices that do not contain slice groups, modify the highlighted portion.] Otherwise, edgeIdx is derived as follows: – The following applies: edgeIdx = 2 + Sign( recPicture[ xSi ][ ySj ]- recPicture[ xSi + hPos[ 0 ] ][ ySj + vPos[ 0 ] ]) + Sign( recPicture[ xSi ][ ySj ]- recPicture[ xSi + hPos[ 1 ] ][ ySj +vPos[ 1 ] ]) (8 937).
[0318] – When edgeIdx is equal to 0, 1, or 2, edgeIdx is modified as follows: edgeIdx = ( edgeIdx == 2 ) ? 0 : ( edgeIdx + 1 ) (8 938).
[0319] 3. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) - 1, recPicture[xSi ][ ySj ]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ edgeIdx ]) (8 939).
[0320] – Otherwise (SaoTypeIdx[ cIdx ][ rx ][ ry ] equals 1), the following ordered steps apply: 1. The variable bandShift is set to equal bitDepth - 5.
[0321] 2. The variable saoLeftClass is set to equal sao_band_position[ cIdx ][ rx ][ ry ].
[0322] 3. The list `bandTable` is defined with 32 elements, and all elements are initially set to 0. Then, its four elements (indicating the starting position of the band for the explicit offset) are modified as follows: for( k = 0; k<4; k++ ) bandTable[ ( k + saoLeftClass )&31 ] = k + 1 (8 940).
[0323] 4. The variable bandIdx is set to equal bandTable[recPicture[xSi][ySj]>>bandShift].
[0324] 5. The modified image sample array saoPicture[xSi][ySj] is derived as follows: saoPicture[ xSi ][ ySj ]= Clip3( 0, ( 1< <bitDepth ) - 1, recPicture[xSi ][ ySj ]+ SaoOffsetVal[ cIdx ][ rx ][ ry ][ bandIdx ]) (8 941).
[0325] Table 8.19 – Specifications for hPos and vPos based on the sample adaptive compensation category
[0326] 2.7. OBMC in ECM When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information from neighboring blocks with weighted prediction.
[0327] The following conditions should not be used for OBMC: • When OBMC is disabled at the SPS level.
[0328] • When the current block has intra-frame mode or IBC mode.
[0329] • When applying LIC to the current block.
[0330] • When the current luminance block area is less than or equal to 32.
[0331] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using motion information from neighboring sub-blocks. It is enabled for sub-block-based codec tools. • Affine AMVP mode; • Affine Merge pattern and sub-block-based temporal motion vector prediction (SbTMVP). • Bilateral matching based on sub-blocks.
[0332] When OBMC mode is used in CIIP mode with LMCS, inter-frame blending is performed before the LMCS mapping of the inter-frame samples. LMCS is applied to the blended inter-frame samples, which are combined with the intra-frame samples to which LMCS is applied in CIIP mode.
[0333] ,in This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by the motion of neighboring blocks in the original domain, and and It's the weight.
[0334] 2.8. Intertwined Prediction To overcome the problems in sub-block-based prediction, interleaving prediction in video encoding and decoding is proposed.
[0335] Using interleaved prediction, a block is divided into sub-blocks with more than one partitioning pattern. A partitioning pattern is defined as the way a block is divided into sub-blocks, including the size and position of the sub-blocks. For each partitioning pattern, a corresponding prediction block can be generated by deriving the motion information of each sub-block based on the partitioning pattern. Therefore, even for a single prediction direction, multiple prediction blocks can be generated from multiple partitioning patterns. Alternatively, for each prediction direction, only the partitioning pattern can be applied.
[0336] Suppose there are X partitioning patterns, and the X predicted blocks (denoted as P0, P1, ..., PX-1) of the current block are generated through sub-block-based predictions with the X partitioning patterns. The final prediction (denoted as P) of the current block can be generated as follows: (15) Where (x, y) are the coordinates of the pixels in the block, and These are the weighted values of Pi. Without losing generality, assume... , where N is a non-negative value. Use of Interweaved Prediction for Different Coding Tools An example of interleaved prediction with two partitioning modes is shown.
[0337] The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0338] Definition of Partition Modes 1. Interleaving prediction can be applied to one, some, or all of the codecs that have sub-block-based prediction. In one example, affine prediction applies interleaving prediction, while other codecs with sub-block-based prediction (such as ATMVP, STMVP, FRUC, and BIO) do not. In another example, affine, ATMVP, and STMVP apply interleaving prediction.
[0339] FIGS. 20A-20G 2. The partitioning pattern can have different shapes, sizes, or positions of the sub-blocks. In one example, the partitioning pattern could result in irregular sub-block sizes. FIG. 20A Several exemplary partitioning patterns for 16×16 blocks are shown. FIG. 20B In this context, blocks are divided into 4x4 sub-blocks, as in JEM. FIG. 20C In the middle, the block is divided into 8×8 sub-blocks. FIG. 20D and FIG. 20E In the middle, the block is divided into 8×4 sub-blocks and 4×8 sub-blocks respectively. FIG. 20F and FIG. 20E In the block, the block is also divided into 4×4 sub-blocks, but with different positions. Pixels at the block boundary that cannot be divided into complete 4×4 sub-blocks can be divided into smaller sub-blocks of size 2×4, 4×2, or 2×2, such as... FIG. 20F As shown, they can also be merged into adjacent 4×4 sub-blocks to form larger sub-blocks of size 6×4, 4×6, or 6×6, such as... FIG. 20G As shown; in Enabling / Disabling Interweaved Prediction and Coding Process for Interweaved Prediction In this model, blocks are also divided into 8×8 sub-blocks, but in different positions. Pixels at the block boundaries that cannot be divided into complete 8×8 sub-blocks can be divided into smaller sub-blocks of size 8×4, 4×8, or 4×4.
[0340] 3. The shape and size of the sub-blocks in the sub-block prediction can depend on the shape and / or size of the encoded block and / or the encoded block information (e.g., whether it is an affine or ATMVP mode).
[0341] a. In one example, when the current block has a size of M×N, the child block has a size of 4×N (or 8×N, etc.).
[0342] b. In one example, when the current block has a size of M×N, the child block has a size of M×4 (or M×8, etc.).
[0343] c. In one example, when the current block has a size of M×N and M>N, the child block has a size of A×B, such as 8×4, which is greater than B; otherwise, the child block has a size of B×A, such as 4×8.
[0344] d. In one example, assuming the current block has a size of M×N, the sub-block has a size of A×B when M×N<=T (or Min(M,N)<=T, or Max(M,N)<=T, etc.), and a size of C×D when M×N>T (or Min(M,N)>T, or Max(M,N)>T, etc.), where A<=C and B<=D. For example, if M×N<=256, the sub-block has a size of 4×4; otherwise, the sub-block has a size of 8×8.
[0345] FIG. 20D 4. Whether to apply interleaved prediction depends on the inter-frame prediction direction.
[0346] a. In one example, interleaved forecasts can be applied to bidirectional forecasts but not to unidirectional forecasts.
[0347] b. In one example, when multiple hypotheses are applied, interleaved predictions can be applied to a prediction direction when there is more than one reference block.
[0348] 5. How to apply interleaved prediction depends on the inter-frame prediction direction.
[0349] a. In one example, a bidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two different partitioning patterns for two different reference lists. In one example, when predicting from reference list 0 (L0), the block is divided into 4×8 sub-blocks, such as... FIG. 20C As shown, however, when predicting from reference list 1 (L1), the block is divided into 8×4 sub-blocks, as follows: FIG. 20D As shown. And the final prediction P is calculated as (16) Where P0 and P1 are predictions from L0 and L1, respectively. w0 and w1 are weighted values for L0 and L1, respectively. Without loss of generality, assume... (where N is a non-negative integer value).
[0350] b. In one example, a unidirectional prediction block with sub-block-based predictions is divided into sub-blocks with two or more distinct partitioning patterns. For example, the prediction PL for list L (L=0 or 1) is calculated as follows: (17) Where XL is the number of partitioning patterns for list L; It is a prediction generated using the i-th partitioning pattern, and yes The weighted value. For example, XL is 2. Using the 0th partitioning pattern, the block is divided into 4×8 sub-blocks, such as... FIG. 20C As shown. Using the first partitioning pattern, the block is divided into 8×4 sub-blocks, as follows. Weighting Values As shown.
[0351] c. In one example, a bidirectional prediction block with sub-block-based predictions is considered as a combination of two unidirectional prediction blocks from L0 and L1, respectively. Predictions from each list can be derived as described in the example above. The final prediction P can be computed as... (18) The parameters a and b are two additional weights applied to the two internal prediction blocks. In one example, both a and b are equal to 1.
[0352] d. In one example, for a block encoded using multiple hypotheses, there may be more than one prediction block generated by different partitioning patterns for each prediction direction (or reference image list). Multiple prediction blocks can be used to generate a final version with additional weights applied. In one example, the additional weights can be set to 1 / M, where M is the total number of prediction blocks generated.
[0353] 6. Whether and how interleaving prediction can be applied can be transmitted from the encoder to the decoder at the sequence level, picture level, view level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, PU level, or TU level, or slice level, or slice group level, or region level (which may include multiple CU / PU / TU / LCU)). This information can be transmitted via signaling in the first block of the sequence parameter set (SPS), view parameter set (VPS), picture parameter set (PPS), stripe header (SH), picture header, sequence header, or slice level or slice group level, CTU (also known as LCU), CU, PU, TU, or region.
[0354] a. In one example, interleaving prediction implicitly applies to existing sub-block methods such as ATMVP, STMVP, FRUC, BIO, or affine. In this example, no additional signaling cost is required.
[0355] b. In another example, new sub-block Merge candidates generated by interleaving prediction are inserted into the Merge list, such as interleaving prediction + ATMVP, interleaving prediction + STMVP, interleaving prediction + FRUC, etc.
[0356] c. In one example, a flag may be transmitted via signaling to indicate whether interleaving prediction is used. In one example, if the current block is affine inter-frame encoded, the flag is transmitted via signaling to indicate whether interleaving prediction is used.
[0357] d. In one example, if the current block is encoded and decoded using affine Merge and one-way prediction is applied, a flag can be signaled to indicate whether interleaved prediction is used.
[0358] e. In one example, if the current block is encoded / decoded using affine Merge, a flag can be transmitted via signaling to indicate whether interleaved prediction is used.
[0359] f. In one example, if the current block is encoded and decoded using an affine Merge codec and one-way prediction is applied, interleaved prediction can always be used.
[0360] g. In one example, if the current block is encoded or decoded using affine Merge, interleaving prediction can always be used.
[0361] h. In one example, a flag indicating whether to use interleaving prediction can be inherited without being transmitted via signaling.
[0362] i. In one example, inheritance can be used if the current block is encoded or decoded using affine Merge.
[0363] ii. In one example, the flag can be inherited from the flag of a neighboring block that inherits the affine model.
[0364] iii. In one example, the flag is inherited from a predefined neighboring block (such as the neighboring block to the left or above).
[0365] iv. In one example, the flag can be inherited from the first encountered affine-coded neighboring block.
[0366] v. In one example, if no neighboring blocks are affine encoded, the flag can be presumed to be zero.
[0367] vi. In one example, the flag can only be inherited if one-way prediction is applied to the current block.
[0368] vii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to be inherited are in the same CTU.
[0369] viii. In one example, the flag can only be inherited if the current block and the neighboring block from which it is to be inherited are in the same CTU line.
[0370] ix. In one example, when the affine model is derived from a temporal neighbor block, the flag cannot be inherited from the flag of the neighbor block.
[0371] x. In one example, the flag cannot be inherited from the flag of a neighboring block located in the same LCU or LCU line or video data processing unit (such as 64×64 or 128×128).
[0372] xi. In one example, how the flag is transmitted and / or deduced via signaling may depend on the block dimension of the current block and / or encoded / decoded information.
[0373] i. In one example, if the reference image is the current image, then interleaved predictions are not applied.
[0374] i. In one example, if the reference image is the current image, a flag indicating whether interleaving prediction is used is not transmitted through the signal.
[0375] FIGS. 21A-21D 7. The weighting value w is fixed. For example, in equations (15) and (16) .
[0376] 8. The weighting value can depend on the position and the partitioning pattern, that is, for different (x, y). It may differ. Alternatively, the weighting may further depend on the codec tool based on sub-block prediction (e.g., affine or ATMVP) and / or other codec information (e.g., skip or non-skip modes and / or MV information, etc.).
[0377] 9. Weighting values can be transmitted from the encoder to the decoder at the sequence level, picture level, stripe level, codec tree unit (CTU) (also known as maximum codec unit (LCU) level, CU level, or PU level, or region level, which may include multiple CUs / PUs / TUs / LCUs)). They can be transmitted via signaling in the first block of the sequence parameter set (SPS), picture parameter set (PPS), stripe header (SH), CTU (also known as LCU), CU or PU, or region.
[0378] a. In an alternative solution, in addition, for some blocks, weights can be inherited from spatial and / or temporal neighboring blocks.
[0379] Partial Interleaving Prediction 10. In one embodiment, interleaved predictions are applied to a portion of the current block. Prediction samples at some locations are calculated as a weighted sum of two or more sub-block-based predictions. Prediction samples at other locations are not. For example, these prediction samples are copied from sub-block-based predictions with a specific partitioning pattern. FIG. 21A An example of partial interleaving prediction is shown. Interleaving prediction is not applied to shaded areas.
[0380] a. In one example, the current block is predicted using sub-block-based predictions P1 and P2 with partitioning patterns D0 and D1. The final prediction is calculated as P = w0 × P0 + w1 × P1. At some locations, w0 ≠ 0 and w1 ≠ 0. But at some other locations, w0 = 1 and w1 = 0, meaning that interleaved predictions are not applied to those locations.
[0381] b. In one example, such asFIG. 21B As shown, interleaving prediction is not applied to the four corner sub-blocks.
[0382] c. In one example, such as FIG. 21C As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns.
[0383] d. In one example, such as FIG. 21D As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0384] e. In one example, such as FIG. 21B As shown, interleaving prediction is not applied to the topmost sub-row, the bottommost sub-row, the leftmost sub-column, and the rightmost sub-column.
[0385] f. In one example, whether and how partial interleaving prediction is applied can depend on the size / shape of the current block.
[0386] i. For example, if the size of the current block meets certain conditions, the interleaving prediction is applied to the entire block; otherwise, the interleaving prediction is applied to a portion (or parts) of the block. Conditions include, but are not limited to: (assuming the width and height of the current block are W and H respectively, and T, T1, and T2 are integer values): 1. W>=T1 and H>=T2; 2. W <= T1 and H <= T2; 3. W>=T1 or H>=T2; 4. W <= T1 or H <= T2; 5. W + H >= T; 6. W + H <= T; 7. W×H>=T; 8. W×H<=T.
[0387] ii. For example, if W>=H, then as FIG. 21C As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... FIG. 21B As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0388] iii. For example, if W>H, then as FIG. 21C As shown, interleaving prediction is not applied to the leftmost and rightmost sub-block columns; otherwise, as... FIGS. 22A-22C As shown, interleaving prediction is not applied to the topmost and bottommost sub-rows.
[0389] g. It is proposed that whether and how interleaved predictions are applied can differ for different regions within a block.
[0390] i. For example, assume that the current block is predicted by sub-block based prediction P1 and P2 with partitioning patterns D0 and D1. The final prediction is calculated as P(x, y) = w0 × P0(x, y) + w1 × P1(x, y). If the position (x, y) belongs to a sub-block with partitioning pattern D0 of dimension S0 × H0; and belongs to a sub-block with partitioning pattern D1 of dimension S1 × H1. If one or more of the following conditions are met, then set w0 = 1 and w1 = 0. (That is, the interleaved prediction is not applied to this position).
[0391] 1. S1 < T1; 2. H1 < T2; 3. S1 < T1 and H1 < T2; 4. S1 < T1 or H1 < T2; T1 and T2 are integers. For example, T1 = T2 = 4.
[0392] Encoder issues 11. In one embodiment, the interleaved prediction is not applied during the motion estimation (ME) process.
[0393] a. For example, the interleaved prediction is not applied during the ME process for 6-parameter affine prediction.
[0394] b. For example, if the size of the current block satisfies specific conditions, such as (assuming the width and height of the current block are W and H respectively, and T, T1, T2 are integer values): i. W >= T1 and H >= T2; ii. W <= T1 and H <= T2; iii. W >= T1 or H >= T2; iv. W <= T1 or H <= T2; v. W + H >= T; vi. W + H <= T; vii. W × H >= T; viii. W × H <= T. In the following discussion, SatShift(x, n) is defined as .
[0398] Shift(x, n) is defined as Shift(x, n) = (x + offset0) >> n.
[0399] In one example, offset0 and / or offset1 are set to (1 < 0.05).<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.
[0400] 12. The MV of each sub-block within a partitioning pattern can be derived directly from the affine model (such as by using equation (1)), or it can be derived from the MV of a sub-block within another partitioning pattern.
[0401] a. In one example, the MV of sub-block B with partitioning pattern 0 can be derived from the MVs of all or some sub-blocks within partitioning pattern 1 that overlaps with sub-block B.
[0402] b. FIG. 22A An example is shown. FIG. 22B In this process, the MV1(x,y) of a specific sub-block within partitioning mode 1 will be derived. FIG. 22C The diagram shows partitioning pattern 0 (solid) and partitioning pattern 1 (dashed), indicating that there are four sub-blocks in partitioning pattern 0 that overlap with specific sub-blocks in partitioning pattern 1. FIGS. 22A-22C The diagram shows four MVs for four sub-blocks within partitioning pattern 0 that overlap with a specific sub-block within partitioning pattern 1: MV0(x-2, y-2), MV0(x+2, y-2), MV0(x-2, y+2), and MV0(x+2, y+2). MV1(x, y) is then derived from MV0(x-2, y-2), MV0(x+2, y-2), MV0(x-2, y+2), and MV0(x+2, y+2).
[0403] c. Suppose that the MV' of a sub-block within partitioning pattern 1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks within partitioning pattern 0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.
[0404] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.
[0405] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.
[0406] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.
[0407] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.
[0408] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.
[0409] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.
[0410] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.
[0411] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.
[0412] x. with FIGS. 23A-23C For example, MV1(x, y) can be derived as: 1. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 2. MV1(x,y) = Shift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 3. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 4. MV1(x,y) = Shift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 5. MV1(x,y) = SatShift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 6. MV1(x,y) = Shift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 7. MV1(x,y) = SatShift( MV0(x-2,y-2)+MV0(x+2,y+2), 1); 8. MV1(x,y) = Shift( MV0(x-2,y-2)+ MV0(x+2,y+2), 1); 9. MV1(x,y) = SatShift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 10. MV1(x,y) = Shift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 11. MV1(x,y) = SatShift(MV0(x+2,y-2) +MV0(x+2,y+2), 1); 12. MV1(x,y) = Shift( MV0(x+2,y-2) +MV0(x+2,y+2), 1); 13. MV1(x,y) = SatShift(MV0(x+2,y-2)+MV0(x-2,y+2), 1); 14. MV1(x,y) = Shift( MV0(x+2,y-2)+MV0(x-2,y+2), 1); 15. MV1(x,y) = MV0(x-2,y-2); 16. MV1(x,y) = MV0(x+2,y-2); 17. MV1(x,y) = MV0(x-2,y+2); 18. MV1(x,y) = MV0(x+2,y+2).
[0413] 13. The choice of partitioning mode can depend on the width and height of the current block. FIG. 23A An example of selecting a partitioning mode based on block dimensions is shown.
[0414] a. For example, if width > T1 and height > T2 (e.g., T1 = T2 = 4), then both partitioning modes are selected. FIG. 23B Examples of two partitioning patterns are shown.
[0415] b. For example, if the height is less than or equal to T2 (e.g., T2 = 4), then the other two partitioning modes are selected. FIG. 23C Examples of two partitioning patterns are shown.
[0416] c. For example, if the width is less than or equal to T1 (e.g., T1 = 4), then two more partitioning patterns are selected. FIG. 24A Examples of two partitioning patterns are shown.
[0417] 14. The MV of each sub-block within a partitioning mode of a color component C1 can be derived from the MV of the sub-block within another partitioning mode of another color component C0.
[0418] a. For example, C1 refers to a color component that is encoded / decoded after another color component, such as Cb or Cr or U or V or R or B.
[0419] b. For example, C0 refers to a color component that is encoded / decoded before another color component, such as Y or G.
[0420] c. In one example, how to derive the MV of a sub-block within a partitioning mode of a color component from the MV of a sub-block within another partitioning mode of another color component can depend on the color format, such as 4:2:0, or 4:2:2, or 4:4:4.
[0421] d. In one example, the MV of a sub-block B in a color component C1 with partitioning mode C1Pt (t=0 or 1) can be derived from the MV of all or some sub-blocks in a color component C0 with partitioning mode C0Pr (r=0 or 1) that overlaps with sub-block B after scaling or scaling coordinates according to the color format.
[0422] i. In one example, C0Pr is always equal to C0P0.
[0423] e. FIG. 24B and FIG. 24A An example is shown. FIG. 24B and FIG. 24A This example illustrates the derivation of the MV of a sub-block within a component of a partitioning pattern from the MV of a sub-block within another component of a partitioning pattern. The color format is 4:2:0. The MV of a sub-block in the Cb component is derived from the MV of a sub-block in the Y component.
[0424] i. at FIG. 24A On the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 0 will be derived. FIG. 24B The right side shows four Y-blocks within partition mode 0, which overlap with sub-block B (Cb) when scaled at a 2:1 ratio. Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').
[0425] ii. In FIG. 24B On the left, the MVCb0(x',y') of a specific Cb subblock B within partitioning mode 1 will be derived. FIG. 24A The right side shows four Y-blocks within partition mode 0, which overlap with sub-block B (Cb) when scaled at a 2:1 ratio. Assume x = 2. x' and y=2 y', the four MVs of the four Y sub-blocks in partition mode 0: MV0(x-2,y-2), MV0(x+2,y-2), MV0(x-2,y+2) and MV0(x+2,y+2) are used to derive MVCb0(x',y').
[0426] f. Suppose that the MV' of a sub-block of color component C1 is derived from the MV0, MV1, MV2, ..., MVk of the (k-1) sub-blocks of color component C0. MV' can be derived as: i. MV' = MVn, where n is any one of 0…k.
[0427] ii. MV' = f( MV0, MV1, MV2, …, MVk). f is a linear function.
[0428] iii. MV' = f( MV0, MV1, MV2, …, MVk). f is a nonlinear function.
[0429] iv. MV' = Average(MV0, MV1, MV2, …, MVk). Average is the averaging operation.
[0430] v. MV' = Median(MV0, MV1, MV2, …, MVk). Median is the operation to get the median.
[0431] vi. MV' = Max(MV0, MV1, MV2, …, MVk). Max is the operation to get the maximum value.
[0432] vii. MV' = Min(MV0, MV1, MV2, …, MVk). Min is the operation to find the minimum value.
[0433] viii. MV' = MaxAbs(MV0, MV1, MV2, …, MVk). MaxAbs is the operation that retrieves the value with the largest absolute value.
[0434] ix. MV' = MinAbs(MV0, MV1, MV2, …, MVk). MinAbs is the operation that retrieves the value with the smallest absolute value.
[0435] x. with FIG. 24B and Interweaved Prediction for Bi-prediction For example, MVCbt(x',y') t = 0 or 1 can be derived as: 1. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 2. MVCbt(x',y') = Shift( MV0(x-2,y-2)+MV0(x+2,y-2)+MV0(x-2,y+2)+MV0(x+2,y+2), 2); 3. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 4. MVCbt(x',y') = Shift( MV0(x-2,y-2)+MV0(x+2,y-2), 1); 5. MVCbt(x',y') = SatShift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 6. MVCbt(x',y') = Shift(MV0(x-2,y+2)+MV0(x+2,y+2), 1); 7. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+MV0(x+2,y+2), 1); 8. MVCbt(x',y') = Shift( MV0(x-2,y-2)+ MV0(x+2,y+2), 1); 9. MVCbt(x',y') = SatShift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 10. MVCbt(x',y') = Shift( MV0(x-2,y-2)+ MV0(x-2,y+2), 1); 11. MVCbt(x',y') = SatShift(MV0(x+2,y-2) +MV0(x+2,y+2), 1); 12. MVCbt(x',y') = Shift( MV0(x+2,y-2) +MV0(x+2,y+2), 1); 13. MVCbt(x',y') = SatShift(MV0(x+2,y-2)+MV0(x-2,y+2), 1); 14. MV1(x,y) = Shift( MV0(x+2,y-2)+MV0(x-2,y+2), 1); 15. MVCbt(x',y') = MV0(x-2,y-2); 16. MVCbt(x',y') = MV0(x+2,y-2); 17. MVCbt(x',y') = MV0(x-2,y+2); 18. MVCbt(x',y') = MV0(x+2,y+2).
[0436] Block size dependency 15. When interleaved prediction is applied to bidirectional prediction, the following methods can be applied to preserve the increased internal bit depth due to the different weights: a. For a list X (X = 0 or 1), PX(x, y) = Shift(W0(x, y)). PX0(x,y) + W1(x,y) PX1(x,y), SW), where PX(x,y) is the prediction for list X, and PX0(x,y) and PX1(x,y) are the predictions for list X with partitioning mode 0 and partitioning mode 1. W0 and W1 are integers representing the weighted values of the interleaved predictions, and SW represents the precision of the weighted values.
[0437] b. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB), where Wb0 and Wb1 are integers used in weighted bidirectional forecasting, and SWB is the precision. When there is no weighted bidirectional forecasting, Wb0=Wb1=SWB=1.
[0438] c. In some embodiments, PX0(x,y) and PX1(x,y) can maintain the accuracy of the interpolation filter. For example, they can be 16-bit unsigned integers. The final predicted value is derived as P(x,y) = Shift(Wb0(x,y)). P0(x,y)+ Wb1(x,y) P1(x,y), SWB+PB), where PB is the additional precision from the interpolation filter, for example, PB = 6. In this case, W0(x,y) PX0(x,y) or W1(x,y) PX1(x,y) can exceed 16 bits. It is proposed that PX0(x,y) and PX1(x,y) are first right-shifted to a lower precision to avoid exceeding 16 bits.
[0439] i. For example, for a list X (X = 0 or 1), PX(x, y) = Shift(W0(x, y)). PLX0(x,y) + W1(x,y) PLX1(x,y), SW), where PLX0(x,y) = Shift( PX0(x,y), M), PLX1(x,y) = Shift( PX1(x,y), M). The final prediction is derived as P(x,y) = Shift( Wb0(x,y)). P0(x,y) + Wb1(x,y) P1(x,y), SWB+PB-M). For example, M is set to 2 or 3.
[0440] d. The above method can also be applied to other bidirectional forecasting methods with different weighting factors for two reference forecast blocks, such as generalized bidirectional forecasting (GBi, where the weights can be, for example, 3 / 8, 5 / 8) and weighted forecasting (where the weights can be very large values).
[0441] e. The above method can also be applied to other multi-hypothesis one-way or two-way forecasting methods that have different weighting factors for different reference forecast blocks.
[0442] On Interweaved Affine 16. Whether and / or how to apply interleaving prediction can depend on the block width W and height H.
[0443] a. In one example, whether and / or how to apply interleaved prediction can depend on the size of the VPDU (Video Processing Data Unit, which typically represents the maximum permissible block size to be processed in a hardware design).
[0444] b. In one example, when interlaced prediction is disabled for a specific block dimension (or a block with information having a specific warp decoding), the original prediction method can be utilized.
[0445] i. Alternatively, the affine mode can be directly disabled for such a block.
[0446] c. In one example, when W > T1 and H > T2, interlaced prediction cannot be used. For example, T1 = T2 = 64; d. In one example, when W > T1 or H > T2, interlaced prediction cannot be used. For example, T1 = T2 = 64; e. In one example, when W H > T, interlaced prediction cannot be used. For example, T = 64 64; f. In one example, when W < T1 and H < T2, interlaced prediction cannot be used. For example, T1 = T2 = 16; g. In one example, when W < T1 or H > T2, interlaced prediction cannot be used. For example, T1 = T2 = 16; h. In one example, when W H < T, interlaced prediction cannot be used. For example, T = 16 16.
[0447] i. In one example, for a sub - block not located at the block boundary (e.g., the coding / decoding unit), interlaced affine can be disabled for this sub - block. Alternatively, in addition, the prediction result using the original affine prediction method can be directly used as the final prediction for this sub - block.
[0448] j. In one example, when W > T1 and H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 64; k. In one example, when W > T1 or H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 64; l. In one example, when W H > T, interlaced prediction is used in a different way. For example, T = 64 <00者01563>64; m. In one example, when W < T1 and H < T2, interlaced prediction is used in a different way. For example, T1 = T2 = 16; n. In one example, when W < T1 or H > T2, interlaced prediction is used in a different way. For example, T1 = T2 = 16; o. In one example, when W When H < T, the interleaved prediction is used in a different way. For example, T = 16 16
[0449] p. In one example, when H > X (e.g., H equals 128, X = 64), the interleaved prediction is not applied to the samples of the sub - blocks belonging to the upper W (H / 2) split and the lower W (H / 2) split
[0450] q. In one example, when W > X (e.g., W equals 128, X = 64), the interleaved prediction is not applied to the samples of the sub - blocks belonging to the left (W / 2) H split and the right (W / 2) H split
[0451] r. In one example, when W > X and H > Y (e.g., W = H = 128, X = Y = 64), i. The interleaved prediction is not applied to the samples of the sub - blocks belonging to the left (W / 2) H split and the right (W / 2) H split
[0452] ii. The interleaved prediction is not applied to the samples of the sub - blocks belonging to the upper W (H / 2) split and the lower W (H / 2) split
[0453] s. In one example, the interleaved prediction is enabled only for blocks with a specific set of widths and / or heights
[0454] t. In one example, the interleaved prediction is disabled only for blocks with a specific set of widths and / or heights
[0455] u. In one example, the interleaved prediction is used only for specific types of pictures / strips / groups of pictures / slices / or other kinds of video data units
[0456] i. For example, the interleaved prediction is used only for P pictures or B pictures
[0457] ii. For example, a flag is signaled in the header of the picture / strip / group of pictures / slice to indicate whether the interleaved prediction can be used
[0458] 1. For example, the flag is signaled only when affine prediction is allowed
[0459] 17. A message is proposed to be transmitted via signaling to indicate whether / how interleaving prediction is applied, and the dependency between width and height. The message can be transmitted via signaling in SPS / VPS / PPS / strip header / picture header / slice / slice group header / CTU / CTU line / multiple CTUs / or other types of video processing units.
[0460] 18. In one example, bidirectional forecasting is not allowed when using interleaved forecasting.
[0461] a. For example, when using interleaved prediction, the index indicating whether bidirectional prediction is used is not transmitted via signaling.
[0462] b. Alternatively, an indication of whether bidirectional prediction is not allowed can be transmitted via signaling in SPS / VPS / PPS / strip header / picture header / film / film group header / CTU / CTU line / multiple CTUs.
[0463] 19. A method was proposed to further refine the motion information of sub-blocks based on motion information derived from two or more modes.
[0464] a. In one example, refined motion information can be used to predict subsequent blocks to be encoded or decoded.
[0465] b. In one example, refined motion information can be used in filtering processes such as deblocking, SAO, and ALF.
[0466] c. Whether to store refined information can be based on the position of the sub-block relative to the whole block / CTU / CTU line / slice / strip / slice group / image.
[0467] d. Whether to store refined information may be based on the encoded / decoded patterns of the current block and / or neighboring blocks.
[0468] e. Whether to store refined information can be based on the dimension of the current block.
[0469] f. Whether to store refined information can be based on image / strip type / reference image list, etc.
[0470] 20. It is proposed that whether and / or how to apply the deblocking process or other kinds of filtering processes (such as SAO, adaptive loop filter) can depend on whether interleaved prediction is applied.
[0471] a. In one example, if an edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then that edge is not deblocked.
[0472] b. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then the deblocking of that edge is weakened.
[0473] i. In one example, for such edges, bS[xDi][yDj] described in the VVC deblocking process decreases.
[0474] ii. In one example, for such edges, the β reduction described in the VVC deblocking process.
[0475] iii. In one example, for such edges, the Δ described in the VVC deblocking process decreases.
[0476] iv. In one example, for such edges, tC as described in the VVC deblocking process decreases.
[0477] c. In one example, if the edge between two sub-blocks in one partitioning pattern of a block is inside a sub-block in another partitioning pattern of the block, then the deblocking of that edge is strengthened.
[0478] i. In one example, for such an edge, bS[xDi][yDj] described in the VVC deblocking process increases.
[0479] ii. In one example, for such edges, the β increase described in the VVC deblocking process.
[0480] iii. In one example, for such edges, the Δ described in the VVC deblocking process increases.
[0481] iv. In one example, for such edges, the tC described in the VVC deblocking process is increased.
[0482] 21. It is proposed that whether and / or how local illumination compensation or weighted prediction is applied to blocks / subblocks may depend on whether interleaved prediction is applied.
[0483] a. In one example, when a block is encoded or decoded in interleaved prediction mode, local illumination compensation or weighted prediction is not allowed.
[0484] b. Alternatively, if interleaving prediction is applied to blocks / subblocks, there is no need to indicate the need to enable local lighting compensation via signal transmission.
[0485] 22. It was proposed that bidirectional optical flow (BIO) can be skipped when weighted prediction is applied to a block or sub-block.
[0486] a. In one example, BIO can be applied to blocks with weighted predictions.
[0487] b. In one example, BIO can be applied to blocks with weighted predictions; however, certain conditions must be met.
[0488] i. In one example, at least one parameter is required to be within a range or equal to a specific value.
[0489] ii. In one example, specific reference image constraints may be applied.
[0490] 3. Problem Existing sub-block-based prediction techniques have the following problems: 1) They face a dilemma. If the sub-block size is smaller, the motion information of each sub-block can be more accurate. However, smaller sub-blocks impose higher bandwidth requirements in the MC.
[0491] 2) Motion information derived for smaller sub-blocks can be dangerous, especially when there is noise in the block. Fixing the sub-block size within a single block may be suboptimal.
[0492] 4. Detailed Solution This disclosure proposes further improvements to sub-block-based motion compensation. Specifically, interleaving prediction is improved to be compatible with pixel-based affine motion compensation. Furthermore, regression affine is also improved.
[0493] The specific embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0494] The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0495] In this disclosure, "blocks encoded and decoded in mode N" can refer to a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, AMVP, SMVD, merging, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.).
[0496] It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0497] On OBMC 1. The use of interleaved prediction in affine mode can depend on another encoding / decoding method.
[0498] a) In one example, if pixel-based affine motion compensation is applied to the codec block, then interleaving prediction may not be applied.
[0499] b) In one example, interleaved affine prediction and pixel-based affine prediction cannot be used simultaneously for encoding / decoding blocks.
[0500] i. In one example, alternatively, interleaved affine prediction and pixel-based affine prediction can be used simultaneously for encoding and decoding blocks.
[0501] 1) In one example, for a codec block, some pixels can be predicted by interleaved affine mapping, while (some) other pixels can be predicted by pixel-based affine mapping.
[0502] 2) In one example, interleaved affine predictions and pixel-based affine predictions can be mixed to generate new predictions.
[0503] 2. The use of interleaved affine prediction and / or pixel-based affine motion compensation can depend on the block dimension.
[0504] a) In one example, interleaved affine prediction is used if the block dimension meets certain conditions.
[0505] i. In one example, specifically, interleaved affine prediction can be used if both (or either) the height and width are greater than or less than a threshold, or if the ratio of height to width is greater than or less than a threshold.
[0506] b) In one example, alternatively, if the block dimensions do not meet the same conditions, pixel-based affine mapping can be used instead.
[0507] c) In one example, dimension conditions can be applied together with other conditions.
[0508] i. In one example, the OBMC flag can be used together with dimension conditions.
[0509] 1) In one example, specifically, if OBMC is used for blocks and the block dimensions meet certain conditions, then pixel affine mapping is used to generate predictions.
[0510] d) In one example, whether to apply interlaced affines can depend on the sub-block size.
[0511] i. In one example, if the sub-block size is smaller than a threshold, such as 4×4, then interleaved affine is not applied.
[0512] 3. In one example, the size of the largest sub-block to be used for affine prediction is set to a constant when using interleaved affine or not using pixel-based affine.
[0513] 4. In one example, PROF (Affine Prediction Refinement with Optical Flow) can be performed for interleaved affines.
[0514] a) In one example, the predictions generated by the two partitioning patterns are both refined by PROF.
[0515] i. In one example, only predictions generated by a specific partitioning pattern are refined by PROF.
[0516] ii. In one example, if interleaved affines are used for the block, PROF is not applied.
[0517] b) In one example, for a specific partitioning pattern, only a subset of predicted samples are refined by PROF.
[0518] i. In one example, if the width or height of a sub-block is less than a threshold (e.g., 4), PROF will not be applied to that sub-block.
[0519] 5. It is proposed that whether and / or how to apply sub-block boundary deblocking procedures or other types of sub-block boundary-based filtering procedures (such as SAO, adaptive loop filter, OBMC) can depend on whether affine mode / pixel-based affine / interleaved prediction is applied.
[0520] a) In one example, if affine mode / pixel-based affine / interleaved prediction is applied, some or all kinds of sub-block boundary-based deblocking / filtering processes are not applied.
[0521] i. In one example, specifically, if affine mode / pixel-based affine / interlacing prediction is applied, then deblocking and / or OBMC facing the sub-block boundary is not applied.
[0522] FIG. 25 FIG. 25 An example of a CU-level OBMC is shown.
[0523] 6. It was proposed that whether and / or how to apply CU-level OBMC can depend on the encoding / decoding mode.
[0524] a) In one example, if a sub-block level inter-frame mode (such as affine, SbTMVP, multi-pass DMVR) is used for the current block, each boundary sub-block will be filtered independently at the CU level OBMC level.
[0525] i. In one example, specifically, boundary sub-blocks are iterated one by one. Suppose A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block used for OBMC filtering. If A and A1 have different motions, then A will be filtered. After A has been filtered, subsequent sub-blocks will undergo OBMC in a similar manner.
[0526] b) In one example, alternatively, if a non-sub-block level inter-frame mode is used for the current block, multiple boundary sub-blocks can be filtered together using OBMC.
[0527] i. In one example, suppose A is an arbitrary boundary sub-block being traversed, and A1 is the corresponding neighboring sub-block. If A1 and subsequent... N ( N >=0) consecutive sub-blocks (e.g. nei The two sub-blocks B1 and C1 in the middle have the same motion. M_ M_nei And at the same time nei Unlike the motion of A, then ( N +1) sub-blocks (including A) will be filtered together. Otherwise, if M_ On Regressed Affine The motion equal to A, then ( N +1) consecutive sub-blocks (including A) will skip CU-level OBMC filtering.
[0528] M, F 7. To generate regression affine candidates, a method using... N (N >1) The motion field of the previously encoded and decoded CU is used as the input to the regression process.
[0529] a) In one example, for any CU used to generate a regression affine candidate, at least one sub-block is used to provide the motion field.
[0530] b) In one example, for any CU used to generate a regression affine candidate, all sub-blocks of that CU are used to provide the motion field.
[0531] c) In one example N A CU can be collected from adjacent locations, non-adjacent locations, or historical CU / parameter tables.
[0532] d) In one example, at least one previously encoded / decoded CU can be collected from the following locations: i. (Multiple) adjacent and neighboring locations; ii. (Multiple) adjacent locations at a specific location.
[0533] iii. (Multiple) co-located or adjacent time-domain locations iv. Historical table.
[0534] v. (Multiple) non-adjacent spatial / temporal locations.
[0535] e) In one example N Each CU is encoded and decoded via inter-frame encoding / decoding.
[0536] i. In one example, N Each CU is encoded and decoded using affine mode.
[0537] ii. In one example, alternatively, N At least one of them K ( K (0) are affine encoded / decoded.
[0538] iii. In one example, the motion fields of at least one affine-coded CU and at least one non-affine-coded CU can be used to generate regressive affine candidates.
[0539] f) In one example, all N Each CU may need to share the same prediction direction (i.e., bidirectional or unidirectional prediction, and / or the reference list used) and / or the same reference index / frame.
[0540] i. In one example, alternatively, they may have different prediction directions or reference frames.
[0541] g) In one example, only when N A regression affine candidate can only be generated when the number of sub-blocks of a CU is greater than a constant threshold.
[0542] h) In one example, sports fields in adjacent or non-adjacent locations can also be used as additional inputs.
[0543] i. In one example, M Row / column adjacent sub-blocks and / or F Non-adjacent sub-blocks in rows / columns can be used as input, where FIG. 26 >=0.
[0544] 1) In one example, the location of the non-adjacent positions used can depend on the block dimension.
[0545] i) In one example, the proposed regression candidate can be used when the number of existing regression candidates has not reached the maximum allowed number.
[0546] j) In one example, the proposed regression candidates can be reordered based on a specific metric (such as ARMC or template matching).
[0547] k) In one example, the number of proposed regression candidates may not exceed a constant or an adaptively determined value.
[0548] l) In one example, the proposed regression affine candidate can be used to generate affine Merge / affine AMVP / affine MMVD / affine Adaptive DMVR / affine TM / affine DMVR and / or any other affine-related method that requires the construction of an affine candidate list.
[0549] 8. A motion field of at least K (K>1, e.g. K=2) codec blocks can be used to generate regression affine candidates for affine AMVP patterns.
[0550] a) In one example, a codec block can only be used to generate regression affine candidates if the reference index or reference frame used by the codec block is the same as the reference index of the current block.
[0551] 9. The affine candidate list can contain regressive affine candidates generated from different numbers of previously encoded / decoded blocks.
[0552] a) In one example, at least one regressive affine candidate in the list is generated using M (M>0) previously encoded / decoded blocks, and / or at least one regressive affine candidate in the list is generated using N (N>0) previously encoded / decoded blocks, where M and N are not the same.
[0553] b) In one example, regression affine candidates generated using more previously encoded blocks have a higher priority for being included in the affine candidate list compared to candidates generated using fewer previously encoded blocks.
[0554] FIG. 27 A flowchart of a method 2600 for video processing according to an embodiment of the present disclosure is shown. Method 2600 is implemented during the conversion between video block video units and the video bitstream.
[0555] At box 2610, for the conversion between the current video block and the video bitstream, the motion field of a plurality of codec units encoded and decoded prior to the current video block is determined. At least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a given location, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block. As used herein, the term “current video block” refers to the video block or video unit to be processed, which may also be referred to as “target video block” or “current video unit”. It should be understood that these example locations or tables are for illustrative purposes only and do not imply any limitation. Any suitable location or table may be applied. The scope of this disclosure is not limited herein.
[0556] At box 2620, regression affine candidates for the current video block are determined based on the motion field of multiple codec units.
[0557] At box 2630, a transformation is performed based on regressive affine candidates. In some embodiments, the transformation includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the current video block from the bitstream.
[0558] Method 2600 enables the generation of regressive affine candidates using the motion fields of multiple previously encoded CUs. In this way, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.
[0559] In some embodiments, for a codec unit among a plurality of codec units, the motion field includes at least one motion field provided by at least one sub-block of the codec unit. For example, for any CU used to generate a regressive affine candidate, at least one sub-block is used to provide the motion field. Alternatively, in some embodiments, the motion field includes at least one motion field provided by all sub-blocks of the codec unit. For example, for any CU used to generate a regressive affine candidate, all sub-blocks of that CU are used to provide the motion field.
[0560] In some embodiments, the plurality of codec units include at least one affine codec unit and at least one non-affine codec unit, and the motion fields of at least one affine codec unit and at least one non-affine codec unit are used to determine regressive affine candidates.
[0561] In some embodiments, regression affine candidates are used to determine at least one of the following: affine merge, affine advanced motion vector prediction (AMVP), affine MMVD, adaptive DMVR for affine, affine template matching (TM), affine DMVR, or other affine-related information that requires the construction of an affine candidate list.
[0562] In some embodiments, the affine candidate list for the current video block includes multiple regressive affine candidates based on different numbers of previously encoded / decoded codec units.
[0563] In some embodiments, a first regressive affine candidate in the affine candidate list is determined based on a first number of previously encoded / decoded units, and a second regressive affine candidate in the affine candidate list is determined based on a second number of previously encoded / decoded units, the second number being different from the first number. For example, at least one regressive affine candidate in the list is generated using M (M>0) previously encoded / decoded blocks, and / or at least one regressive affine candidate in the list is generated using N (N>0) previously encoded / decoded blocks. M and N may be different.
[0564] In some embodiments, if the first number is less than the second number, the second regressive affine candidate has a higher priority in being included in the affine candidate list compared to the first regressive affine candidate. That is, the priority of the regressive affine candidate can be based on the number of previously encoded / decoded units used to determine the regressive affine candidate.
[0565] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, the motion field of a plurality of codec units encoded and decoded prior to the current video block is determined. At least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a given location, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block. The affine candidate of the current video block is determined based on the motion field of the plurality of codec units. The bitstream is generated based on the regressive affine candidate.
[0566] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, the motion field of a plurality of codec units encoded and decoded prior to the current video block is determined. At least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a given location, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block. Affine candidates for the current video block are determined based on the motion field of the plurality of codec units. The bitstream is generated based on regressive affine candidates. The bitstream is stored in a non-transitory computer-readable recording medium.
[0567] FIG. 28 A flowchart of another method 2700 for video processing according to an embodiment of the present disclosure is shown. Method 2700 is implemented during the conversion between video blocks or video units of a video and a bitstream of a video.
[0568] At box 2710, the motion field of multiple codec blocks is determined for the conversion between the current video block and the video bitstream.
[0569] At box 2720, regression affine candidates are determined for the current video block based on the motion fields of multiple codec blocks. The current video block is in Affine Advanced Motion Vector Prediction (AMVP) mode.
[0570] At box 2730, a transformation is performed based on regressive affine candidates. In some embodiments, the transformation includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, the transformation includes decoding the current video block from the bitstream.
[0571] Method 2700 enables the use of motion fields from multiple encoded / decoded blocks to determine regression affine candidates for blocks encoded / decoded via affine AMVP. In this way, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.
[0572] In some embodiments, a codec block is used to determine regression affine candidates if the reference index or reference frame for the codec block is the same as another reference index or another reference frame for the current video block. For example, a codec block is used to generate regression affine candidates only if the reference index or reference frame used by the codec block is the same as the reference index of the current block.
[0573] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, motion fields of a plurality of codec blocks are determined. Based on the motion fields of the plurality of codec blocks, regression affine candidates for a current video block are determined. The current video block is in affine high-level motion vector prediction (AMVP) mode. A bitstream is generated based on the regression affine candidates.
[0574] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, motion fields of a plurality of codec blocks are determined. Based on the motion fields of the plurality of codec blocks, regression affine candidates for a current video block are determined. The current video block is in Affine Advanced Motion Vector Prediction (AMVP) mode. A bitstream is generated based on the regression affine candidates. The bitstream is stored in a non-transitory computer-readable recording medium.
[0575] It should be understood that methods 2600 and / or 2700 can be applied individually or in any combination. Using these methods, encoding / decoding efficiency and effectiveness can be improved.
[0576] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0577] Item 1. A method for video processing, comprising: for a conversion between a current video block of a video and a bitstream of the video; determining a motion field of a plurality of codec units previously encoded and decoded in the current video block, wherein at least one of the plurality of codec units is collected from at least one of: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block; determining a regression affine candidate for the current video block based on the motion field of the plurality of codec units; and performing the conversion based on the regression affine candidate.
[0578] Item 2. The method according to Item 1, wherein for each of the plurality of encoding / decoding units, the motion field includes at least one motion field provided by at least one sub-block of the encoding / decoding unit.
[0579] Item 3. The method according to Item 1, wherein for each of the plurality of encoding / decoding units, the motion field includes at least one motion field provided by all sub-blocks of the encoding / decoding unit.
[0580] Item 4. The method according to any one of items 1-3, wherein the plurality of encoding and decoding units comprises at least one affine encoding and decoding unit and at least one non-affine encoding and decoding unit, and the motion field of the at least one affine encoding and decoding unit and the at least one non-affine encoding and decoding unit is used to determine the regressive affine candidate.
[0581] Item 5. The method according to any one of Items 1-34, wherein the regression affine candidate is used to determine at least one of the following: affine merge, affine advanced motion vector prediction (AMVP), affine (MMVD), adaptive for affine (DMVR), affine template matching (TM), affine DMVR, or additional affine-related information required for the construction of the affine candidate list.
[0582] Item 6. The method according to any one of items 1-54, wherein the affine candidate list of the current video block includes multiple regressive affine candidates based on a different number of previously encoded / decoded codec units.
[0583] Item 7. The method according to Item 6, wherein the first regressive affine candidate in the affine candidate list is determined based on a first number of previously encoded / decoded units, and the second regressive affine candidate in the affine candidate list is determined based on a second number of previously encoded / decoded units, the second number being different from the first number.
[0584] Item 8. The method according to Item 7, wherein if the first number is less than the second number, the second regression affine candidate has a higher priority to be included in the affine candidate list compared to the first regression affine candidate.
[0585] Item 9. A method for video processing, comprising: determining motion fields of a plurality of codec blocks for a conversion between a current video block and a bitstream of the video; determining regression affine candidates for the current video block, the current video block being in an affine advanced motion vector prediction (AMVP) mode, based on the motion fields of the plurality of codec blocks; and performing the conversion based on the regression affine candidates.
[0586] Item 10. The method according to Item 9, wherein the codec block is used to determine the regression affine candidate if the reference index or reference frame for the codec block is the same as another reference index or another reference frame for the current video block.
[0587] Item 11. The method according to any one of items 1-10, wherein the conversion includes encoding the current video block into the bitstream.
[0588] Item 12. The method according to any one of items 1-10, wherein the conversion includes decoding the current video block from the bitstream.
[0589] Item 13. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1-12.
[0590] Item 14. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-12.
[0591] Item 15. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method comprises: determining a motion field of a plurality of codec units encoded and decoded prior to a current video block of the video, wherein at least one of the plurality of codec units is collected from at least one of: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block; determining an affine candidate of the current video block based on the motion field of the plurality of codec units; and generating the bitstream based on the regressive affine candidate.
[0592] Item 16. A method for storing a bitstream of video, comprising: determining a motion field of a plurality of codec units encoded and decoded prior to a current video block of the video, wherein at least one of the plurality of codec units is collected from at least one of: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or a history table of the current video block; determining a regression affine candidate of the current video block based on the motion field of the plurality of codec units; generating the bitstream based on the regression affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0593] Item 17. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: determining a motion field of a plurality of codec blocks; determining, based on the motion field of the plurality of codec blocks, a regression affine candidate for a current video block of the video, the current video block being in an affine advanced motion vector prediction (AMVP) mode; and generating the bitstream based on the regression affine candidate.
[0594] Item 18. A method for storing a bitstream of video, comprising: determining a motion field of a plurality of codec blocks; determining a regression affine candidate for a current video block of the video based on the motion field of the plurality of codec blocks, the current video block being in an affine high-level motion vector prediction (AMVP) mode; generating the bitstream based on the regression affine candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0595] Example device FIG. 28 A block diagram of a computing device 2800 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2800 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0596] It should be understood that, FIG. 28 The computing device 2800 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0597] like FIG. 28 As shown, computing device 2800 includes general-purpose computing device 2800. Computing device 2800 may include at least one or more processors or processing units 2810, memory 2820, storage unit 2830, one or more communication units 2840, one or more input devices 2850, and one or more output devices 2860.
[0598] In some embodiments, the computing device 2800 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2800 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0599] Processing unit 2810 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2800. Processing unit 2810 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0600] Computing device 2800 typically includes various computer storage media. Such media can be any media accessible by computing device 2800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2830 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2800.
[0601] The computing device 2800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0602] Communication unit 2840 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2800 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0603] Input device 2850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2860 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2840, computing device 2800 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2800 can also communicate with one or more devices that enable a user to interact with computing device 2800, or any device that enables computing device 2800 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0604] In some embodiments, some or all components of computing device 2800 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to achieve the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider offers applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0605] In embodiments of this disclosure, computing device 2800 can be used to implement video encoding / decoding. Memory 2820 may include one or more video codec modules 2825 having one or more program instructions. These modules can be accessed and executed by processing unit 2810 to perform the functions of the various embodiments described herein.
[0606] In an example embodiment of performing video encoding, input device 2850 may receive video data as input 2870 to be encoded. The video data may be processed, for example, by video codec module 2825 to generate an encoded bitstream. The encoded bitstream may be provided as output 2880 via output device 2860.
[0607] In an example embodiment of performing video decoding, input device 2850 may receive an encoded bitstream as input 2870. The encoded bitstream may be processed, for example, by a video codec module 2825 to generate decoded video data. The decoded video data may be provided as output 2880 via output device 2860.
[0608] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine the motion field of a plurality of codec units that were encoded and decoded before the current video block, wherein at least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or the history table of the current video block. Based on the motion field of the plurality of encoding and decoding units, the regression affine candidate of the current video block is determined; as well as The transformation is performed based on the regression affine candidate.
2. The method of claim 1, wherein for each of the plurality of encoding / decoding units, the motion field includes at least one motion field provided by at least one sub-block of the encoding / decoding unit.
3. The method of claim 1, wherein for each of the plurality of encoding / decoding units, the motion field comprises at least one motion field provided by all sub-blocks of the encoding / decoding unit.
4. The method according to any one of claims 1-3, wherein the plurality of encoding and decoding units comprises at least one affine encoding and decoding unit and at least one non-affine encoding and decoding unit, and the motion field of the at least one affine encoding and decoding unit and the at least one non-affine encoding and decoding unit is used to determine the regressive affine candidate.
5. The method according to any one of claims 1-34, wherein the regression affine candidate is used to determine at least one of the following: affine merge, affine advanced motion vector prediction (AMVP), affine (MMVD), adaptive affine for affine (DMVR), affine template matching (TM), affine DMVR, or other affine-related information required for the construction of the affine candidate list.
6. The method according to any one of claims 1-54, wherein the affine candidate list of the current video block includes a plurality of regressive affine candidates based on a different number of previously encoded / decoded codec units.
7. The method of claim 6, wherein the first regressive affine candidate in the affine candidate list is determined based on a first number of previously encoded / decoded units, and the second regressive affine candidate in the affine candidate list is determined based on a second number of previously encoded / decoded units, the second number being different from the first number.
8. The method of claim 7, wherein if the first number is less than the second number, the second regression affine candidate has a higher priority to be included in the affine candidate list compared to the first regression affine candidate.
9. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine the motion field of multiple codec blocks; Based on the motion field of the plurality of codec blocks, regression affine candidates are determined for the current video block, which is in affine advanced motion vector prediction (AMVP) mode; as well as The transformation is performed based on the regression affine candidate.
10. The method of claim 9, wherein the codec block is used to determine the regression affine candidate if the reference index or reference frame for the codec block is the same as another reference index or another reference frame for the current video block.
11. The method according to any one of claims 1-10, wherein the conversion comprises encoding the current video block into the bitstream.
12. The method according to any one of claims 1-10, wherein the conversion comprises decoding the current video block from the bitstream.
13. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-12.
14. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-12.
15. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine the motion field of a plurality of codec units that were encoded and decoded before the current video block of the video, wherein at least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or the history table of the current video block; Based on the motion field of the plurality of encoding and decoding units, the regression affine candidate of the current video block is determined; as well as The bitstream is generated based on the regression affine candidate.
16. A method for storing a bitstream of video, comprising: Determine the motion field of a plurality of codec units that were encoded and decoded before the current video block of the video, wherein at least one of the plurality of codec units is collected from at least one of the following: adjacent neighbor positions, adjacent neighbor positions at a position, co-located temporal positions, adjacent temporal positions, non-adjacent spatial positions, non-adjacent temporal positions, or the history table of the current video block; Based on the motion field of the plurality of encoding and decoding units, the regression affine candidate of the current video block is determined; The bitstream is generated based on the regression affine candidate; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
17. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine the motion field of multiple codec blocks; Based on the motion field of the plurality of codec blocks, a regression affine candidate for the current video block of the video is determined, the current video block being in affine advanced motion vector prediction (AMVP) mode; as well as The bitstream is generated based on the regression affine candidate.
18. A method for storing a bitstream of video, comprising: Determine the motion field of multiple codec blocks; Based on the motion field of the plurality of codec blocks, a regression affine candidate for the current video block of the video is determined, the current video block being in affine advanced motion vector prediction (AMVP) mode; The bitstream is generated based on the regression affine candidate; as well as The bitstream is stored in a non-transitory computer-readable recording medium.