Method and apparatus for adaptive motion compensation filtering

By using adaptive motion compensation filtering technology, template matching and filter optimization of inter-frame prediction blocks are employed, which solves the problem of insufficient encoding/decoding efficiency of inter-frame coding blocks and improves the performance and compression efficiency of video encoding and decoding.

CN120958798APending Publication Date: 2025-11-14BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480023010.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-12
Filing Date
2024-04-12
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies still have room for improvement in terms of encoding/decoding efficiency of inter-frame coding blocks, especially based on the VVC standard, where further optimization of encoding and decoding tools is needed to improve encoding and decoding efficiency.

Method used

An adaptive motion compensation filtering technique is adopted. Motion vector candidates of neighboring reconstructed samples are obtained through the decoder and encoder. Inter-frame prediction blocks are filtered using template matching cost and template filter, and the final prediction block is generated by combining the intra-frame prediction blocks.

Benefits of technology

It improves the encoding/decoding efficiency of inter-frame coding blocks, enhances the performance of video encoding and decoding, especially the encoding quality and compression efficiency in complex video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958798A_ABST
    Figure CN120958798A_ABST
Patent Text Reader

Abstract

Methods, apparatus, and non-transitory computer-readable storage media for video decoding and encoding are provided. In a video decoding method, a decoder may obtain, by the decoder, a template matching cost of a plurality of motion vector candidates of a plurality of first reconstruction samples adjacent to a current inter-coded block in response to determining that adaptive motion compensation filtering is applied to the inter-coded block; obtaining, by the decoder, a target motion vector from the plurality of motion vector candidates based on the template matching cost of the plurality of first reconstruction sample points; and obtaining, by a decoder, a plurality of prediction samples based on the target motion vector and the current inter-coded block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is filed and claims priority to U.S. Provisional Application No. 63 / 458,913, filed April 12, 2023, entitled “Adaptive Motion Compensated Filtering for Bi-Prediction,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to video encoding / decoding and compression, and specifically, but not limited to, methods and apparatus for improving the encoding / decoding efficiency of inter-frame coded blocks. Background Technology

[0004] Various video codec techniques can be used to compress video data. Video codecs are performed according to one or more video codec standards. For example, video codec standards include Universal Video Codec (VVC), High Efficiency Video Codec (H.265 / HEVC), High-Advanced Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically use predictive methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilize redundancy present in video images or sequences. A key goal of video codec techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing video quality degradation.

[0005] The first version of the VVC standard was completed in July 2020, offering approximately 50% bitrate savings or equivalent perceived quality compared to its predecessor, HEVC. While the VVC standard provides significant codec improvements over its predecessor, there is evidence that superior codec efficiency can be achieved using additional codec tools. Recently, the Joint Video Exploration Group (JVET), in collaboration with ITU-T VECG and ISO / IEC MPEG, began exploring advanced technologies that can significantly improve codec efficiency compared to VVC. In April 2021, a software codebase called the Enhanced Compression Model (ECM) was established for future video codec exploration work. The ECM reference software is based on the VVC Test Model (VTM) developed by JVET for VVC and further extends and / or improves several existing modules (e.g., intra / inter-frame prediction, transform, loop filters, etc.). In the future, any new codec tools beyond the VVC standard will need to be integrated into the ECM platform and tested using the JVET Common Test Conditions (CTC). Summary of the Invention

[0006] This disclosure provides examples of techniques related to improving the encoding / decoding efficiency of inter-frame coded blocks.

[0007] According to a first aspect of this disclosure, a video decoding method for an inter-frame coded block is provided. The method includes: in response to determining that an adaptive motion compensation filter is applied to a current inter-frame coded block, obtaining by a decoder template matching costs for a plurality of motion vector candidates from a plurality of first reconstructed samples adjacent to the current inter-frame coded block; obtaining by the decoder a target motion vector from the plurality of motion vector candidates based on the template matching costs of the plurality of first reconstructed samples; and obtaining by the decoder a plurality of predicted samples based on the target motion vector and the current inter-frame coded block.

[0008] According to a second aspect of this disclosure, a video decoding method for inter-frame coded blocks is provided. The method includes: obtaining intra-prediction blocks of a current inter-frame coded block by a decoder; obtaining a plurality of inter-frame prediction blocks of the current inter-frame coded block by the decoder in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coded block; obtaining a filtered inter-frame prediction block by the decoder based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on a current template of the current inter-frame coded block, wherein the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coded block; and obtaining a final prediction block by the decoder by combining the intra-frame prediction blocks and the filtered inter-frame prediction blocks.

[0009] According to a third aspect of this disclosure, a video coding method for inter-frame coding blocks is provided. The method includes: in response to determining that an adaptive motion compensation filter is applied to a current inter-frame coding block, an encoder obtains template matching costs for a plurality of motion vector candidates from a plurality of first reconstructed samples adjacent to the current inter-frame coding block; the encoder obtains a target motion vector from the plurality of motion vector candidates based on the template matching costs of the plurality of first reconstructed samples; and the encoder obtains a plurality of predicted samples based on the target motion vector and the current inter-frame coding block.

[0010] According to a fourth aspect of this disclosure, a video coding method for inter-frame coded blocks is provided. The method includes: obtaining intra-prediction blocks of a current inter-frame coded block by an encoder; in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coded block, obtaining a plurality of inter-frame prediction blocks of the current inter-frame coded block by the encoder; obtaining a filtered inter-frame prediction block by the encoder based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on a current template of the current inter-frame coded block, wherein the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coded block; and obtaining a final prediction block by the encoder by combining the intra-frame prediction blocks and the filtered inter-frame prediction blocks.

[0011] According to a fifth aspect of this disclosure, an apparatus for video decoding is provided. The apparatus may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, when executing the instructions, are configured to perform a method according to the first or second aspect.

[0012] According to a sixth aspect of this disclosure, an apparatus for video encoding is provided. The apparatus may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, when executing the instructions, are configured to perform a method according to a third or fourth aspect.

[0013] According to a seventh aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect or the second aspect.

[0014] According to an eighth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the third or fourth aspect.

[0015] According to a ninth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by a method according to the first or second aspect.

[0016] According to a tenth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by a method according to a third or fourth aspect. Attached Figure Description

[0017] A more specific description of the examples disclosed herein will be presented by reference to the specific examples illustrated in the accompanying drawings. Given that these drawings depict only a few examples and should therefore not be considered as limiting the scope, these examples will be described and explained in more specific and detailed manner using the accompanying drawings.

[0018] Figure 1A This is a block diagram illustrating a system for encoding and decoding video blocks according to some examples of this disclosure.

[0019] Figure 1B This is a block diagram of an encoder based on some examples of this disclosure.

[0020] Figures 1C to 1F This is a block diagram illustrating how, according to some examples of this disclosure, a frame is recursively partitioned into multiple video blocks of different sizes and shapes.

[0021] Figure 1G This is a block diagram illustrating an exemplary video encoder according to some examples of this disclosure.

[0022] Figure 2A This is a block diagram of a decoder based on some examples of this disclosure.

[0023] Figure 2B This is a block diagram illustrating an exemplary video decoder according to some examples of this disclosure.

[0024] Figure 3A This is a simplified diagram illustrating block partitioning in a multi-type tree structure according to some examples of this disclosure.

[0025] Figure 3B This is a simplified diagram illustrating block partitioning in a multi-type tree structure according to some examples of this disclosure.

[0026] Figure 3C This is a simplified diagram illustrating block partitioning in a multi-type tree structure according to some examples of this disclosure.

[0027] Figure 3D This is a simplified diagram illustrating block partitioning in a multi-type tree structure according to some examples of this disclosure.

[0028] Figure 3E This is a simplified diagram illustrating block partitioning in a multi-type tree structure according to some examples of this disclosure.

[0029] Figure 4 Some examples of d are shown according to this disclosure. x and d y This is an example of the horizontal and vertical values ​​of MV.

[0030] Figure 5 An example is shown where one of the MVs according to this disclosure has a small value and an interpolation filter is applied to generate the corresponding predicted sample at the location of the small sample.

[0031] Figure 6 shows examples of two diamond filter shapes according to some examples of this disclosure.

[0032] Figure 7 Subsampled 1-D Laplace calculations for gradient computation in all directions are shown as examples of some of the methods described in this disclosure.

[0033] Figure 8 This is a simplified diagram illustrating local illumination compensation (LIC) for unidirectional prediction according to some examples of this disclosure.

[0034] Figure 9A and Figure 9B This is a simplified diagram illustrating the generation of LIC template prediction samples for affine patterns according to some examples of this disclosure.

[0035] Figure 10 This is a block diagram illustrating video coding via adaptive filtering for bidirectional prediction, based on some examples of this disclosure.

[0036] Figure 11 This is a block diagram illustrating video decoding using adaptive filtering for bidirectional prediction, based on some examples of this disclosure.

[0037] Figure 12 This is a simplified diagram illustrating an adaptive motion compensation filter for template-based bidirectional predictive samples according to some examples of this disclosure.

[0038] Figure 13 This is a simplified diagram illustrating an adaptive motion compensation filter for template-based unidirectional predictive samples according to some examples of this disclosure.

[0039] Figure 14 This is a simplified diagram illustrating an OBMC process for encoding and decoding a CU without sub-block motion compensation, according to some examples of this disclosure.

[0040] Figure 15 This is a simplified diagram illustrating the OBMC process for a CU that is encoded and decoded in sub-block mode according to some examples of this disclosure.

[0041] Figure 16 This is a simplified diagram illustrating a template-based OBMC according to some examples of this disclosure.

[0042] Figure 17A and Figure 17B This is a simplified diagram illustrating non-adjacent neighboring blocks of different sizes according to some examples of this disclosure.

[0043] Figure 18 This is a simplified diagram illustrating the template used for cost calculation in non-subblock merging modes in ARMC according to some examples of this disclosure, and its corresponding reference samples.

[0044] Figure 19 This is a simplified diagram illustrating the template used for cost calculation of sub-block merging patterns in ARMC according to some examples of this disclosure, and its corresponding reference samples.

[0045] Figure 20 It is a simplified diagram illustrating the refined position of a basic candidate (reflected by a central asterisk) along a k×π / 8 diagonal angle (reflected by dots with three different patterns) according to some examples of this disclosure.

[0046] Figure 21 This is a simplified diagram illustrating how an adaptive MC filter, according to some examples of this disclosure, is applied when generating prediction samples for the current block but is bypassed when calculating the template cost of different merge / MMVD candidates.

[0047] Figure 22 This is a simplified diagram illustrating the application of adaptive MC filtering according to some examples of this disclosure in generating prediction samples for the current block and in calculating the template cost of different merge / MMVD candidates.

[0048] Figure 23 This is a simplified diagram illustrating the selection of merge candidates for an AMVP-merging mode when at least one merge candidate is associated with an adaptive motion compensation filter, according to some examples of this disclosure.

[0049] Figure 24 This is a simplified diagram illustrating a computing environment coupled to a user interface according to some examples of this disclosure.

[0050] Figure 25 This is a flowchart illustrating some examples of video decoding methods according to this disclosure.

[0051] Figure 26 This is a flowchart illustrating some examples of video decoding methods according to this disclosure.

[0052] Figure 27 This is a flowchart illustrating some examples of video encoding methods according to this disclosure.

[0053] Figure 28 This is a flowchart illustrating some examples of video encoding methods according to this disclosure. Detailed Implementation

[0054] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.

[0055] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular forms “a / an,” “the,” and “the” in this disclosure and the appended claims are also intended to include the plural forms unless otherwise expressly indicated throughout the disclosure. It should also be understood that the term “and / or” as used in this disclosure refers to and includes one or any one or all possible combinations of the listed related items.

[0056] Throughout this specification, references to "an embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise expressly stated, the features, structures, elements, or characteristics described in connection with one or more embodiments also apply to other embodiments.

[0057] Throughout this disclosure, the terms “first,” “second,” “third,” etc., are used as a nomenclature and are used only to refer to relevant elements, such as equipment, components, parts, steps, etc., and do not imply any spatial or temporal order unless otherwise expressly stated. For example, “first equipment” and “second equipment” can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be named arbitrarily.

[0058] The terms "module," "submodule," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "subunit" can include memory (shared, dedicated, or grouped) that stores code or instructions executable by one or more processors. A module can include one or more circuits, with or without stored code or instructions. A module or circuit can include one or more components that are directly or indirectly connected. These components may or may not be physically attached to each other or adjacent to each other.

[0059] As used herein, depending on the context, the terms "if" or "when" can be understood to mean "at the time of" or "in response to". If these terms appear in the claims, they may not indicate that the relevant limitation or feature is conditional or optional. For example, a method may include the steps of: i) performing a function or action X' when or if condition X exists, and ii) performing a function or action Y' when or if condition Y exists. The method may simultaneously have the capability to perform both function or action X' and function or action Y'. Therefore, functions X' and Y' may be performed at different times in multiple executions of the method.

[0060] Units or modules can be implemented purely in software, purely in hardware, or a combination of both. For example, in a purely software implementation, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.

[0061] Figure 1A This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1A As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.

[0062] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, switch, base station, or any other means that may facilitate communication from source device 12 to target device 14.

[0063] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.

[0064] like Figure 1A As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.

[0065] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be sent directly to the target device 14 via the output interface 22 of the source device 12. Alternatively, the encoded video data can be stored on the storage device 32 for later access by the target device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or transmitter.

[0066] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0067] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0068] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to a specific video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.

[0069] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0070] In some embodiments, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20 or components included in video encoder 20 as described below with reference to FIG. 2, and output interface 22) and / or at least a portion of the components of target device 14 (e.g., input interface 28, video decoder 30 or components included in video decoder 30 as described below with reference to FIG. 3, and display device 34) can operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), where the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of source device 12 and / or target device 14 not included in the cloud computing service network can be located in one or more client devices, and these one or more client devices can communicate with a server computer in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a Global Navigation Satellite System (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In embodiments, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein can be implemented by one or more client devices. In some embodiments, the cloud computing service network can be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” are used interchangeably as appropriate. It should be understood that this disclosure is not limited to implementation in the aforementioned cloud computing service network. Rather, this disclosure can be implemented in any other type of computing environment currently known or developed in the future.

[0071] Like HEVC, VVC is built on a block-based hybrid video codec framework. Figure 1B This is a block diagram illustrating a block-based video encoder according to some embodiments of the present disclosure. In encoder 100, the input video signal is processed block by block (called encoding unit (CU)). Encoder 100 may be as follows: Figure 1AThe video encoder 20 is shown. In VTM-1.0, the CU can be up to 128×128 pixels. However, unlike HEVC, which partitions blocks solely based on quadtrees, in VVC, a coding tree unit (CTU) is split into multiple CUs to accommodate different local characteristics based on quadtrees, binary trees, and ternary trees. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, there is no longer a division of CUs, prediction units (PUs), and transform units (TUs) in VVC. Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In a multi-type tree structure, a CTU is first partitioned according to a quadtree structure. Then, each quadtree leaf node can be further partitioned according to binary tree and ternary tree structures.

[0072] Figures 3A to 3E This is a schematic diagram illustrating a multi-type tree partitioning pattern according to some embodiments of the present disclosure. Figures 3A to 3E Five partition types are shown, including quad partitioning ( Figure 3A ), vertical binary partitioning ( Figure 3B ), horizontal binary partitioning ( Figure 3C Vertical triangular partitioning ( Figure 3D ) and horizontal triangular partitioning ( Figure 3E ).

[0073] For each given video block, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Similarly, if multiple reference pictures are supported, an additional reference picture index is sent to identify which reference picture in the reference picture storage the temporal prediction signal originates from.

[0074] Following spatial and / or temporal prediction, the intra / inter-frame mode decision circuit 121 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Then, the block prediction value 120 is subtracted from the current video block; and the resulting prediction residual is decorrelated using transform circuit 102 and quantization circuit 104. The resulting quantization residual coefficients are dequantized by inverse quantization circuit 116 and inversely transformed by inverse transform circuit 118 to form a reconstruction residual, which is then added back to the prediction block to form the reconstructed CU signal. Further, before placing the reconstructed CU in the reference image storage of image buffer 117 and using it for encoding and decoding future video blocks, loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), can be applied to the reconstructed CU. In order to form the output video bitstream 114, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are sent to the entropy coding unit 106 for further compression and packaging to form the bitstream.

[0075] For example, deblocking filters are available in the current version of VVC, as well as in AVC and HEVC. In HEVC, an additional loop filter called SAO is defined to further improve encoding and decoding efficiency. In the current version of the VVC standard, another loop filter called ALF is under active research and is very likely to be included in the final standard.

[0076] These loop filter operations are optional. Performing these operations helps improve encoding / decoding efficiency and visual quality. Encoder 100 can also decide to disable these operations to save computational complexity.

[0077] It should be noted that if these filter options are enabled by encoder 100, intra-frame prediction is typically based on unfiltered reconstructed pixels, while inter-frame prediction is based on filtered reconstructed pixels.

[0078] Figure 2A This is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video codec standards. The decoder 200 is located in... Figure 1B The reconstruction-related parts in encoder 100 are similar. Block-based video decoder 200 can be as follows: Figure 1AThe video decoder 30 is shown. In decoder 200, the incoming video bitstream 201 is first decoded by entropy decoding 202 to obtain quantization coefficient levels and prediction-related information. Then, the quantization coefficient levels are processed by inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residual. The block prediction mechanism implemented in the intra / inter-frame mode selector 212 is configured to perform intra-frame prediction 208 or motion compensation 210 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residual from inverse transform 206 with the prediction output generated by the block prediction mechanism using adder 214.

[0079] The reconstructed blocks can be further passed through loop filter 209 and then stored in image buffer 213, which serves as a reference image storage. The reconstructed video in image buffer 213 can be sent to drive the display device and used to predict future video blocks. With loop filter 209 enabled, filtering operations are performed on these reconstructed pixels to obtain the final reconstructed video output 222.

[0080] Figure 1G This is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy of video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy of video data within adjacent video frames or pictures in a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".

[0081] like Figure 1GAs shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copying (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (such as a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that, regarding CCSAO technology, this application is not limited to the embodiments described herein, but can be applied to situations where an offset is selected for any other component among the luminance, Cb, and Cr chrominance components based on any one of the luminance, Cb, and Cr chrominance components to modify that component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein can be any one of the luminance, Cb, and Cr chrominance components; the second component mentioned herein can be any other one of the luminance, Cb, and Cr chrominance components; and the third component mentioned herein can be the remaining component among the luminance, Cb, and Cr chrominance components. In some examples, the loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or can be divided among one or more of the fixed or programmable hardware units illustrated.

[0082] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1A The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding modes) when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.

[0083] like Figure 1GAs shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of samples with sample values. Samples in the array may also be referred to as pixels or image elements (pel). The number of samples in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of samples with sample values, but its dimension is smaller than that of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. By iteratively using, for example, QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof, a video block can be further segmented into one or more block partitions or sub-blocks (which can again form blocks). It should be noted that the term "block" or "video block" as used herein can refer to a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring, for example, to HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.

[0084] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.

[0085] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.

[0086] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.

[0087] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values ​​for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional-pixel accuracy.

[0088] The motion estimation unit 42 calculates the motion vector for a video block in an inter-frame predictive coding frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1), where each reference frame list identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.

[0089] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation unit 42. Upon receiving motion vectors for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values ​​of the prediction block provided by motion compensation unit 44 from the pixel values ​​of the currently encoded video block. The pixel differences forming the residual video block may include luminance component differences or chrominance component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.

[0090] In some implementations, the intra-BC unit 48 can generate vectors and acquire prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-BC unit 48 can, for example, use various intra-prediction modes to encode the current block during individual encoding passes and test their performance through rate-distortion analysis. Next, the intra-BC unit 48 can select a suitable intra-prediction mode from the various tested intra-prediction modes to use and generate an intra-mode indicator accordingly. For example, the intra-BC unit 48 can use rate-distortion analysis to calculate rate-distortion values ​​for the various tested intra-prediction modes and select the intra-prediction mode with the best rate-distortion characteristics from the tested modes as the suitable intra-prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-frame prediction mode exhibits the optimal rate-distortion value for the block.

[0091] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction according to the embodiments described herein. In any case, for intra-frame block copying, in terms of pixel differences, the predicted block may be a block considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values ​​for sub-integer pixel positions.

[0092] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values ​​of the predicted block from the pixel values ​​of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.

[0093] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra-block copy prediction performed by the intra-BC unit 48, the intra-prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. The intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.

[0094] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.

[0095] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.

[0096] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1A The video decoder 30 shown, or archived in, for example Figure 1A The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.

[0097] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values ​​for use in motion estimation.

[0098] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.

[0099] Figure 2B This is a block diagram illustrating another exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 1G The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.

[0100] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).

[0101] Video data memory 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding modes) when decoding video data. Video data memory 79 and DPB 92 can be formed of any memory device from a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown in... Figure 2B The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip along with other components of the video decoder 30, or off-chip relative to those components.

[0102] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0103] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-prediction unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-prediction mode transmitted by the signal and reference data from the previous decoded block of the current frame.

[0104] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.

[0105] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.

[0106] Motion compensation unit 82 and / or intra-frame prediction (BC) unit 85 determine prediction information for video blocks in the current video frame by parsing motion vectors and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for encoding video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the reference frame list for the frame, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.

[0107] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.

[0108] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values ​​for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.

[0109] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0110] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1A On the display device 34).

[0111] In current VVC and AVS3 standards, motion information for the current coded block is either copied from spatially or temporally neighboring blocks specified by the merge candidate index, or obtained through explicit signals of motion estimation. This disclosure focuses on improving the accuracy of motion vectors in affine merging patterns by refining the derivation method of affine merge candidates. For ease of description, existing affine merging pattern designs in the VVC standard are used as examples to illustrate the proposed ideas. Note that while existing affine pattern designs in the VVC standard are used as examples throughout this disclosure, the proposed techniques can be applied to different designs of affine motion prediction patterns or other codec tools with the same or similar design principles by those skilled in the art of modern video coding and decoding.

[0112] In a typical video encoding and decoding process, a video sequence typically consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of chroma samples (Cb). SCr is a two-dimensional array of chroma samples (Cr). In other instances, a frame may be monochromatic and therefore consist of only a two-dimensional array of luma samples.

[0113] like Figure 1C As shown, the video encoder 20 (or more specifically, the partitioning unit in the predictive processing unit of the video encoder 20) generates an encoded representation of the frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size, namely one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 1D As shown, each CTU may include a luma sample CTB, two corresponding chroma sample coded tree blocks, and syntax elements for encoding and decoding the samples of the coded tree blocks. The syntax elements describe the properties of different types of units within the pixel coded block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In monochrome images or images with three separate color planes, the CTU may include a single coded tree block and syntax elements for encoding and decoding the samples of the coded tree block. The coded tree block can be an N×N sample block.

[0114] To achieve better performance, the video encoder 20 can recursively perform tree partitioning (such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof) on the coding tree blocks of the CTU, and divide the CTU into smaller CUs. Figure 1E The diagram depicts the 64×64 CTU 400 first divided into four smaller CUs, each with a block size of 32×32. Within these four smaller CUs, CU 410 and CU 420 are each divided into four 16×16 CUs. These two 16×16 CUs, 430 and 440, are further divided into four 8×8 CUs. Figure 1F The illustration depicts, as shown Figure 1E The final result of the partitioning process of the CTU 400, as depicted in the diagram, is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU with a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 1DThe CTU depicted in the image may include a luma sample CB and two corresponding chroma sample coded blocks (frames of the same size), as well as syntax elements for encoding and decoding the samples of the coded blocks. In monochrome images or images with three separate color planes, the CU may include a single coded block and syntax structures for encoding and decoding the samples of the coded block. It should be noted that... Figures 1E to 1F The quadtree partitioning depicted is for illustrative purposes only, and a CTU can be divided into multiple CUs to accommodate different local characteristics based on quadtree / tritree / binary tree partitioning. In a multi-type tree structure, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to a binary tree or ternary tree structure. Figures 3A to 3E As shown, a coding block with width W and height H has five possible partition types: quad partition, horizontal binary partition, vertical binary partition, horizontal triad partition, and vertical triad partition.

[0115] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB for luma samples, two corresponding PBs for chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.

[0116] Video encoder 20 can generate prediction blocks for a PU using intra-frame prediction or inter-frame prediction. If video encoder 20 uses intra-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0117] After the video encoder 20 generates predicted luminance blocks, predicted Cb blocks, and predicted Cr blocks for one or more PUs of the CU, the video encoder 20 can generate luminance residual blocks for the CU by subtracting the predicted luminance blocks of the CU from the original luminance coding blocks of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0118] In addition, such as Figure 1E As shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of the luminance samples, two corresponding transform blocks of the chrominance samples, and syntax elements for transforming the transform block samples. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.

[0119] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0120] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.

[0121] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding samples of the predicted blocks of the PU for the current CU to corresponding samples of the transformed blocks of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0122] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.

[0123] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "Motion Vector Prediction" (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0124] Instead of the above combination Figure 1BThe method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit, into the video bitstream, and subtracting the predicted motion vector value of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0125] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to encode and decode the current CU using the same motion vector prediction value from the motion vector candidate list.

[0126] Some embodiments of this disclosure aim to further improve inter-frame encoding / decoding efficiency by applying adaptive enhancement filters to the motion-compensated prediction signal of the bidirectional prediction block. Some embodiments of this disclosure aim to further improve the chroma encoding / decoding efficiency of the motion compensation module applied in ECM. A brief review of some relevant encoding / decoding tools used in the transform and entropy coding processes in ECM is provided below. Subsequently, some shortcomings in existing motion compensation designs are discussed. Finally, solutions to improve existing designs are proposed.

[0127] Motion Compensation Prediction (MCP)

[0128] Motion-compensated prediction (MCP) (also known as motion compensation) is one of the most widely used video coding and decoding techniques in modern video coding and decoding standards. In MCP, a video frame is divided into multiple blocks (called prediction units (PUs)). Each PU is predicted from a block of the same size in a time reference picture, which significantly reduces the overhead required to transmit the block via signaling. In all existing video coding and decoding standards, each inter-frame PU is associated with a set of motion parameters, consisting of one or two MVs and a reference picture index. Inter-frame PUs in P-strips have only one list of reference pictures, while PUs in B-strips can use up to two lists of reference pictures. In MCP, the corresponding inter-frame prediction samples are generated based on the corresponding regions in the reference pictures identified by the MV and the reference picture index. The MV specifies the horizontal and vertical displacement between the current block and its reference block in the reference picture. Figure 4 It shows where d x and d y This is an example of the horizontal and vertical values ​​of an MV. In reality, an MV value may have decimal precision. When an MV has a decimal value, an interpolation filter is applied to generate the corresponding predicted sample at the decimal sample location, such as... Figure 5 As shown. VVC supports MV, which is 1 / 16 of the distance between two adjacent luminance samples in luminance MC, and MV, which is 1 / 32 of the distance between two adjacent chrominance samples in chrominance MC.

[0129] Adaptive Loop Filtering

[0130] In VVC and ECM, Adaptive Loop Filtering (ALF) selects one of 25 filters for each 4×4 block based on the direction and activity of the local gradient.

[0131] Filter shapes: Two diamond filter shapes were used (e.g. Figures 6A to 6B (As shown). A 7×7 rhombus is used for the luminance component, and a 5×5 rhombus is used for the chrominance component.

[0132] Block Classification: For the luminance component, each 4×4 block is classified into one of 25 classes. The class index C is based on its directionality D and the quantization value of the activity. The results are as follows:

[0133]

[0134] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using the 1-D Laplacian operator:

[0135]

[0136] Here, indices i and j refer to the coordinates of the top-left sample point within the 4×4 block, and R(i,j) represents the reconstructed sample point at coordinates (i,j). To reduce the complexity of block classification, such as... Figure 7 As shown, subsampled 1-D Laplace calculation is applied to gradient calculation in all directions.

[0137] Then, the maximum and minimum values ​​of the gradient of D in the horizontal and vertical directions are set as follows:

[0138]

[0139] The maximum and minimum values ​​of the gradients in the two diagonal directions are set as follows:

[0140]

[0141] To obtain the value of the directionality D, these values ​​are compared with each other and with two thresholds t1 and t2:

[0142] Step 1. If If all are true, then D is set to 0.

[0143] Step 2. If If so, continue from step 3; otherwise, continue from step 4.

[0144] Step 3. If If so, set it to 2; otherwise, set D to 1.

[0145] Step 4. If If so, set D to 4; otherwise, set D to 3.

[0146] Activity value A is calculated as:

[0147]

[0148] A is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as No classification method should be applied to the chromaticity components in the image.

[0149] Geometric transformation of filter coefficients and cutoff values

[0150] Before filtering each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flips are applied to the filter coefficients f(k,l) and the corresponding filter cutoff values ​​c(k,l), depending on the gradient values ​​calculated for the block. This is equivalent to applying these transformations to samples in the filter support region. The idea is to make the blocks more similar by aligning the orientations of the different blocks to which the ALF is applied.

[0151] Three geometric transformations are provided: diagonal, vertical flip, and rotation.

[0152]

[0153] Here, K is the size of the filter, and 0 ≤ k, l ≤ K⁻¹ are the coefficient coordinates, such that position (0, 0) is at the top left corner and position (K⁻¹, K⁻¹) is at the bottom right corner. These transformations are applied to the filter coefficients f(k, l) and the cutoff value c(k, l), depending on the gradient value calculated for that block. The relationship between the transformations and the four gradients in these four directions is summarized in Table 1 below.

[0154] gradient value Transformation gd2 < gd1 and gh < gv No transformation gd2 < gd1 and gv < gh diagonal gd1 < gd2 and gh < gv Vertical flip gd1 < gd2 and gv < gh Rotation

[0155] Table 1

[0156] Filtering process

[0157] When ALF is enabled for CTB, each sample point R(i,j) within the CU is filtered to obtain the sample value R′(i,j), as shown below.

[0158]

[0159] Where f(k,l) represents the decoded filter coefficients, K(x,y) is the truncation function, and c(k,l) represents the decoded truncation parameters. Variables k and l are located in... to The range is defined as L, where L represents the filter length. Clip3(-y,y,x) is the clipping function that cuts the input value x to the range [-y,y]. Clipping introduces nonlinearity, making ALF more efficient by reducing the influence of neighboring sample values ​​that differ too much from the current sample value.

[0160] Local illumination compensation

[0161] Local Illumination Compensation (LIC) is an encoding / decoding tool researched during the development of VVC, designed to address local illumination variations in temporally adjacent images. LIC is based on a linear model, from which scaling factors and offsets are derived to enhance the predicted samples of the current block. Specifically, LIC can be mathematically modeled using the following equation:

[0162] P(x,y)=α·P r (x+v x ,y+v y )+β (8)

[0163] Where P(x,y) is the prediction signal for the current block at coordinate (x,x); P r (x+v x ,y+v y) is based on motion vector (v x ,v y The generated prediction block; α and β are the corresponding scaling factors and offsets. Figure 8 The LIC process is illustrated. (For example...) Figure 8 As shown, when LIC is applied to a video block, it makes the neighboring samples of the current block (i.e., Figure 8 The template in the template and its corresponding predicted sample points (i.e., Figure 8 Minimizing the difference between the template predictions in the model yields a linear model (i.e., scaling factor α and offset β).

[0164] Since the scaling factor and offset are derived based on the current block and template and their corresponding prediction signals, there is no overhead in signaling LIC parameters. Additionally, for an inter-block without merging, a LIC flag is signaled to indicate whether LIC mode is enabled. For merged inter-blocks, the LIC flag is considered part of the motion information. Specifically, when constructing the merge list, in addition to the MV and reference index, the LIC flag inherits the flags of its corresponding neighboring blocks. LIC mode is also applied to affine inter-blocks. When affine mode is applied, an inter-block is divided into multiple sub-blocks, and a specific MV is derived for each sub-block based on the affine model. Based on this design, when LIC is applied to an affine block, the corresponding LIC parameters are derived based on the motion information of the sub-blocks on the top and left boundaries of the block; then, the derived LIC model is applied to the prediction samples of the entire block, as shown in Figure 9. Since the MVs of each boundary sub-block may be different, the template prediction signal is also generated based on the sub-blocks, and the prediction samples of each template sub-block are generated using the MVs of the corresponding sub-blocks on the coding block boundaries.

[0165] Finally, it should be noted that in the current LIC design, LIC is only applicable to inter-frame blocks with unidirectional prediction.

[0166] Bidirectional prediction with CU-level weights

[0167] In HEVC, a bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In VVC, the bidirectional prediction mode goes beyond simple averaging, allowing a weighted average of the two prediction signals, i.e.,

[0168] P bi-pred =((8-w)*P0+w*P1+4)>>3 (9)

[0169] Weighted average bidirectional prediction allows five weights, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) for non-merging CUs, using the signaling weight index; 2) for merged CUs, inheriting the weight index from one of the neighboring blocks based on the merge candidate index. Additionally, in VVC, all five weights are used for low-latency images (i.e., all reference images are displayed before the current image). For non-low-latency images (where at least one reference image is displayed after the current image), only three weights (w∈{3,4,5}) are used.

[0170] Overlapping block motion compensation

[0171] OBMC is a codec technique for removing block artifacts during the MC stage. The basic idea of ​​OBMC is to perform motion compensation on the current block using the MVs from neighboring blocks and combine multiple prediction signals using neighboring MVs to generate the final prediction signal for the CU. For each inter-frame CU, OBMC is performed on the top and left boundaries of the block. Additionally, when a video block is encoded and decoded in a sub-block mode (e.g., affine, ATMVP, or DMVR), OBMC is also performed on all internal boundaries of each sub-block (i.e., top, left, bottom, and right boundaries). Figure 15 The diagram illustrates the OBMC process applied to a CU that does not perform sub-block level motion compensation. When OBMC is applied to a sub-block (e.g., Figure 15 When predicting a sub-block (A), in addition to the left and top neighboring sub-blocks of the current sub-block, the MV of the right and bottom neighboring sub-blocks of the current sub-block are also used to derive the prediction signal; then the average of these four prediction blocks is taken to generate the final prediction signal of the current sub-block.

[0172] In the current ECM software, a template-based OBMC scheme is applied. Specifically, the method for deriving the predicted values ​​of CU boundary samples does not use fixed weights to combine multiple motion compensation assumptions, but rather determines them based on the template matching cost, including using only the motion information of the current block, or also using the motion information of neighboring blocks and adopting one of these hybrid modes.

[0173] In this scheme, for each 4×4 block on the top CU boundary, the template size is equal to 4×1. If N adjacent blocks have the same motion information, the template size is increased to 4N×1 because the MC operation can be processed simultaneously. For each 4×4 left block at the left CU boundary, the left template size is equal to 1×4 or 1×4N (e.g., ...). Figure 16 (As shown).

[0174] For each 4×4 top block (or N groups of 4×4 blocks), follow these steps to obtain the predicted values ​​of the boundary samples.

[0175] Taking the current block A and its adjacent block AboveNeighbor_A as an example, the operations on the left-hand blocks are performed in the same way.

[0176] First, based on the following three types of motion information, the three template matching costs (Cost1, Cost2, Cost3) are measured by the SAD between the template reconstruction samples obtained by the MC process and their corresponding reference samples:

[0177] Cost1 is calculated based on the motion information of A.

[0178] Cost2 is calculated based on the motion information of AboveNeighbor_A.

[0179] Cost3 is calculated based on a weighted prediction of the motion information of A and AboveNeighbor_A, with weighting factors of 3 / 4 and 1 / 4, respectively.

[0180] Secondly, a method is selected to calculate the final prediction results of the boundary samples by comparing Cost1, Cost2, and Cost3.

[0181] The original MC result using the motion information of the current block is represented as Pixel1, and the MC result using the motion information of neighboring blocks is represented as Pixel2. The final prediction result is represented as NewPixel.

[0182] If Cost1 is the smallest, then NewPixel(i,j) = Pixel1(i,j).

[0183] If (Cost2+(Cost2>>2)+(Cost2>>3))<=Cost1, then use mixed mode 1.

[0184] For a luminance block, the number of mixed pixel rows is 4.

[0185] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5

[0186] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3

[0187] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4

[0188] NewPixel(i,3)=(31×NewPixel1(i,3)+Pixel2(i,3)+16)>>5

[0189] For chroma blocks, the number of mixed pixel rows is 1.

[0190] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5

[0191] If Cost1 <= Cost2, then use Mixed Mode 2.

[0192] For a luminance block, the number of mixed pixel rows is 2.

[0193] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4

[0194] NewPixel(i,1)=(31×Pixel1(i,1)+Pixel2(i,1)+16)>>5

[0195] For chroma blocks, the number of mixed pixel rows / columns is 1.

[0196] NewPixel(i,0)=(15×NewPixel1(i,0)+Pixel2(i,0)+8)>>4

[0197] Otherwise, use mixed mode 3.

[0198] For a luminance block, the number of mixed pixel rows is 4.

[0199] NewPixel(i,1)=(7×Pixel1(i,1)+NewPixel2(i,1)+4)>>3

[0200] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4

[0201] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5

[0202] For chroma blocks, the number of mixed pixel rows is 1.

[0203] NewPixel(i,0)=(7×Pixel1(i,0)+Pixel2(i,0)+4)>>3

[0204] Adaptive reordering of merged candidates using template matching

[0205] In ECM, a reordering tool called adaptive reordering of merge candidates with template matching (ARMC) is applied to the merge mode of inter-frame encoding and decoding. When this method is applied, the merge candidates are adaptively ordered according to the template matching (TM) cost. This method is applicable to both regular merge mode and affine merge mode.

[0206] Specifically, in the ARMC design, an initial merge candidate list is first constructed, which includes multiple merge candidates, such as spatial, TMVP, non-adjacent, HMVP, and pairwise merge candidates. The candidates in the initial list are then divided into one or more subgroups. The merge candidates in each subgroup are reordered according to the cost based on template matching to generate a reordered merge candidate list. Finally, the indices of the selected merge candidates in the reordered merge candidate list are signaled from the encoder to the decoder.

[0207] During the reordering process, the template matching cost of the merging candidates is measured by the SAD (Search Aspect Ratio) between the template samples of the current block and their corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. Reference samples of the template are located using the motion information of the merging candidates. When the merging candidates utilize bidirectional prediction, reference samples of the merging candidate templates are also generated through bidirectional prediction, such as... Figure 18 As shown.

[0208] For affine patterns, since different sub-blocks may represent different motion vectors, the predicted samples of the template are generated based on sub-block-based motion compensation. Specifically, as... Figure 19 As shown, assuming the size of the candidate sub-blocks for affine merging based on sub-blocks is equal to Wsub × Hsub, the upper template includes several sub-templates of size Wsub × 1, and the left template includes several sub-templates of size 1 × Hsub. The motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.

[0209] Merging patterns based on template matching and MVD

[0210] In addition to the merging mode that directly uses implicitly derived motion information to generate predicted samples for the current CU, a merging mode utilizing motion vector difference (MMVD) will be applied to both the regular merging mode and the affine merging mode. For signal transmission, an MMVD flag is transmitted immediately after the regular merging flag to indicate whether the MMVD mode is used for the CU. In the ECM, 16 refined positions along the k×π / 8 diagonal angle are defined for the MMVD mode, such as... Figure 20 As shown. Additionally, the top N motion candidates in the candidate list before reordering are used as the base candidates for MMVD and affine MMVD. For MMVD, N equals 3; while for affine MMVD, N equals 1 or 3 depending on the affine flags of neighboring blocks. When a base candidate is predicted bidirectionally, two ways of adding MMVD offsets are allowed: 'bilateral' and 'single-sided'. In 'bilateral' MMVD mode, depending on the POC relationship between the current image and its reference images in L0 and L1, the same selected MMVD offset (or its opposite) is applied to the candidate's L0 and L1 MVs. In 'single-sided' MMVD mode, the selected MMVD offset is applied only to the MVs in one reference image list (L0 and L1), while the MVs in the other reference list remain unchanged. Accordingly, based on this design, the MMVD mode has a total of 16 × 6 × 3 = 288 refinement positions. To save on signal transmission overhead, all 288 possible refinement positions are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position, and only the first 36 refined positions after reordering are allowed to be selected, indicated by an MMVD index in the bitstream.

[0211] AMVP - Merge Mode

[0212] In ECM, a novel bidirectional coding mode is introduced, consisting of Advanced Motion Vector Prediction (AMVP) predictions in one direction and merged predictions in the other. This mode is enabled for a coding block when the selected merged and AMVP predictions meet the following conditions: one reference image is from the past and one is from the future relative to the current image, and the distances from both reference images to the current image are the same. If bilateral matching is enabled, bilateral matching MV refinement is applied starting with the merged MV candidate and the AMVP MVP. Otherwise, if template matching is enabled, template matching MV refinement is applied to either the merged prediction or the AMVP prediction, depending on which has a higher template matching cost.

[0213] The AMVP portion of this mode is semaphored as a regular one-way AMVP, that is, the reference index and MVD are semaphored, and if template matching is used, it has a derived MVP index, or the MVP index is semaphored when template matching is disabled.

[0214] For the AMVP direction LX, where X can be 0 or 1, the merging portion in the other direction (1-LX) is implicitly derived by minimizing the bilateral matching cost between the AMVP prediction and the merging prediction (i.e., for a pair of AMVPs and the merging motion MV). For each merging candidate in the merging candidate list with a motion vector in the other direction (1-LX), the bilateral matching cost is computed using the merging candidate MV and the AMVP MV. The merging candidate with the minimum cost is selected. Bilateral matching refinement is applied to the coded block starting from the selected merging candidate MV and AMVP MV.

[0215] The new bidirectional encoding mode is indicated by a flag, and if the mode is enabled, the AMVP direction LX is further indicated by another flag.

[0216] When the current block uses the bilateral matching (BM) AMVP-merge mode and template matching is enabled, no MVD semaphore is used. An additional pair of AMVP-merge MVPs is introduced. The merge candidate list is sorted in ascending order based on the BM cost. A semaphore index (0 or 1) indicates which merge candidate in the sorted merge candidate list to use. When there is only one candidate in the merge candidate list, a pair of AMVP MVPs and merge MVPs that have not undergone bilateral matching MV refinement will be filled.

[0217] Combined intra-frame and inter-frame prediction

[0218] In VVC, when encoding and decoding a CU in merged mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64) and if both the CU width and height are less than 128 luma samples, an additional flag indicating whether to apply Combined Intra-Inter-Frame Prediction (CIIP) mode to the current CU is sent. In CIIP mode, the prediction signal is obtained by combining the inter-frame prediction signal and the intra-frame prediction signal.

[0219] The inter-frame prediction signal in CIIP mode is derived using the same inter-frame prediction process applied in the regular merging mode; and the intra-frame prediction signal in CIIP mode is derived after utilizing the regular intra-frame prediction process in the planar mode. Then, a weighted average is used to combine the intra-frame prediction signal and the inter-frame prediction signal, where the weight values ​​are calculated based on the coding modes of the top neighbor block and the left neighbor block of the current CU, as follows:

[0220] If the top neighboring block is available and intra-frame encoding / decoding has been performed, set isIntraTop to 1; otherwise, set isIntraTop to 0. If the left neighboring block is available and intra-frame encoding / decoding has been performed, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. If (isIntraLeft + isIntraTop) equals 2, set the weight value to 3; otherwise, if (isIntraLeft + isIntraTop) equals 1, set the weight value to 2; otherwise, set the weight value to 1. The prediction signal is derived as follows:

[0221] Pred ciip =((4-wt)*P inter +wt*P intra +2)>>2

[0222] Here, Pinter is the inter-prediction signal in CIIP mode, Pintra is the intra-prediction signal in CIIP mode, wt is the weight value, and >> indicates a right shift operation. Furthermore, when LIC is enabled, the generation of inter-prediction samples in CIIP mode always bypasses the LIC process; that is, scaling and offset are not applied to adjust the inter-prediction samples before they are blended with the intra-prediction samples.

[0223] Among all existing video codec standards, MCP plays a crucial role in ensuring the efficiency of inter-frame coding and decoding. MCP allows the prediction of the video signal to be encoded and decoded based on temporally adjacent signals, transmitting only the prediction error, MV, and reference picture index. As analyzed earlier, ALF can effectively improve the quality of reconstructed video, thereby improving the performance of inter-frame coding and decoding by providing high-quality reference pictures. LIC can be considered an enhancement of conventional motion compensation prediction. While both tools can improve the efficiency of inter-frame coding and decoding, the quality of temporal prediction may still be insufficient for the following reasons.

[0224] Video signals can be encoded and decoded using coarse quantization (i.e., high quantization parameter (QP) values). When coarse quantization is applied, the reconstructed image may contain severe encoding and decoding artifacts, such as block artifacts and ringing artifacts. Given that the reconstructed signal of the current image will be used as a reference for time prediction, this distortion may reduce the effectiveness of MCP, thereby reducing the inter-frame encoding and decoding efficiency of subsequent images.

[0225] While LIC can efficiently compensate for lighting variations between different images, it can only be applied to unidirectional prediction blocks. It is well known that combining multiple prediction blocks can effectively suppress encoding / decoding noise present in motion-compensated signals (this noise is caused by the quantization / dequantization process). Therefore, bidirectional prediction is generally more efficient than unidirectional prediction, meaning that more bidirectional prediction blocks are needed. This implies that unidirectional LIC cannot fully utilize the encoding / decoding gains that LIC tools can achieve.

[0226] According to the existing OBMC design in ECM, OBMC is always disabled for inter-frame CUs using LIC encoding / decoding. This design is not optimal in terms of encoding / decoding efficiency because block artifacts exist between inter-frame blocks encoded / decoded using LIC and those not. Furthermore, even when LIC is applied to two adjacent blocks, potential block artifacts may still exist at the block boundaries because the LIC parameters applied to the two blocks may differ.

[0227] This disclosure presents a method and apparatus for improving the efficiency of motion compensation, thereby improving the quality of time prediction. Specifically, it proposes applying adaptive filtering at the prediction samples of a bidirectional prediction block. To reduce signaling overhead, filter coefficients are derived from the neighboring reconstructed samples (i.e., templates) of the current block and their corresponding prediction samples. In this way, the energy of the prediction residual is mitigated, thereby reducing the overhead of residual signal transmission.

[0228] Figure 10A block diagram of the video encoder applying the proposed adaptive bidirectional predictive filter is presented. First, similar to a conventional video encoder, the motion estimation and compensation module generates a motion-compensated signal by matching the current block with one block (unidirectional prediction) or two blocks (bidirectional prediction) of the reference image using the optimal MV. Then, for a bidirectional prediction block, motion-compensated samples (luminance and chrominance) are provided to the proposed adaptive filter to generate filtered motion-compensated prediction samples for the current block. The original signal is then subtracted from the prediction signal to eliminate temporal redundancy and generate the corresponding residual signal. The residual signal is transformed and quantized, then entropy-coded and output to the bitstream. To obtain the reconstructed signal, the residual signal is reconstructed through inverse quantization and inverse transform. The reconstructed residual is then added to the motion-compensated prediction. Further, loop filtering processes (e.g., deblocking, ALF, and SAO) are applied to the reconstructed video signal for output. As will be discussed later, the filter coefficients of the proposed adaptive bidirectional predictive filter are derived directly from neighboring reconstructed luminance and chrominance samples at the decoder. In addition, to maximize the encoding and decoding gain of the proposed method, additional syntax can be signaled at a given block level (e.g., CTU, CU, or PU level) to indicate whether the proposed filter is applied to the current block for motion compensation.

[0229] Figure 11 A block diagram of the proposed decoder is shown, which receives signals from... Figure 10 The encoder generates a bitstream. At the decoder, the bitstream is first parsed by an entropy decoder. The residual coefficients are then dequantized and inversely transformed to obtain the reconstructed residuals. For time prediction, a prediction signal is first generated by obtaining motion-compensated blocks using prediction information transmitted in the signal (i.e., MV and reference index). Then, for bidirectional prediction blocks, parsing is performed from the bitstream to determine whether adaptive filtering is enabled for that block. If adaptive filtering is enabled, the motion-compensated luma and chroma signals are further processed by the proposed adaptive filtering; otherwise, the motion-compensated chroma signal is not filtered. The motion-compensated signal (filtered or unfiltered) and the reconstructed residuals are then added to obtain the reconstructed video. The reconstructed video may also need to undergo loop filtering before it can be stored in the reference image storage area for display and / or for decoding future video signals.

[0230] Adaptive bidirectional predictive filtering based on template bidirectional predictive samples

[0231] This section proposes an adaptive filtering scheme for bidirectional prediction, where the filter coefficients are derived from bidirectional prediction samples of a template of a bidirectional prediction block. Specifically, in the proposed scheme, bidirectional prediction samples of the template are first generated based on the motion vector of the current block; then, the least squares mean error (LMSE) algorithm is applied to minimize the difference between the template prediction samples and the template samples to obtain the filter parameters. Figure 12 The proposed adaptive filtering method based on template bidirectional prediction samples is illustrated. Figure 12 As shown, T represents the template of the current bidirectional prediction block; T0 and T1 are the L0 and L1 prediction samples of the template, which are bidirectional motion vectors of the current block. and The generated data is based on these representations. In the proposed scheme, bidirectional prediction samples of the template are first generated by averaging the two unidirectional predictions of the templates in L0 and L1.

[0232] T bi =w0*T0+w1*T1 (10)

[0233] Here, w0 and w1 are the weights applied in the L0 and L1 directions when generating bidirectional prediction samples for the current block. When BCW is not applied, the weights are equal to 0.5, while when BCW is applied, the weights may be equal to -0.125, 0.375, 0.625, and 1.125. Based on the obtained template bidirectional prediction samples, using LMSE derivation, the coefficients of the adaptive filter are calculated by minimizing the difference between the template samples and their bidirectional prediction samples, i.e.,

[0234]

[0235] Among them, f * This indicates that it is applied to a template prediction sample T. bi The coefficients of the filter corresponding to the H×L neighborhood region of (x,y), where, In practice, various filters of different sizes and shapes can be applied to provide different trade-offs between encoding / decoding performance and complexity. Larger filters can make the template prediction samples closer to the template samples, but at the cost of increased computational complexity. Finally, the derived filter coefficients are applied to modify the original bidirectional prediction signal of the current block, as shown below.

[0236]

[0237] Among them, P bi (x,y) and P′ bi(x,y) are the bidirectional predicted samples before and after applying the proposed adaptive filtering. Furthermore, to further improve the encoding / decoding gain, the proposed method introduces an offset and some nonlinear terms when deriving the filter coefficients, which can further reduce the distortion between the template samples and their predicted samples. Specifically, after this modification, the derivation of the filter coefficients in equation (11) becomes...

[0238]

[0239] Furthermore, the filter in (12) is applied as

[0240]

[0241] Where o is the offset, nl k It is a nonlinear term, which is represented as a template prediction sample T. bi The sum of a series of powers of (2x, 2y) (i.e., k = 2, ..., K-1).

[0242] In one or more examples, a linear model (i.e., scaling factor and offset) is proposed to derive a two-tap filter to enhance the predicted samples of a bidirectional prediction block. Specifically, a bidirectional prediction LIC is proposed, which operates as follows: 1) generating bidirectional prediction samples of a template, as shown in (10); 2) using the template samples and their corresponding bidirectional prediction samples to derive the scaling factor and offset, as shown below.

[0243]

[0244] Where α and β are the scaling factor and offset of the LIC linear model; N is the number of template samples involved in the derivation process. Then, the final bidirectional prediction for the current block is generated, as shown below.

[0245] P′ bi (x,y)=α·P bi (x,y)+β (16)

[0246] Adaptive bidirectional predictive filtering based on template unidirectional samples

[0247] This section proposes an adaptive bidirectional prediction filtering scheme for unidirectional prediction samples using a template of a bidirectional prediction block. For example, in this method, the template's prediction samples are applied twice in a unilateral manner: two sets of filter coefficients are obtained and applied to the prediction samples in L0 and L1 respectively; then, the weighted average of these two filtered unidirectional prediction samples forms the final prediction sample of the current block. Figure 13 The proposed solution is illustrated. For example... Figure 13As shown, based on L0 and L1 MV, two unidirectional predictions T0 and T1 are generated for the template. Then, based on the individual minimization of the distortion between T0 and T and between T1 and T, two sets of filter parameters f0 and f1 can be obtained for the L0 and L1 directions respectively, as described below:

[0248]

[0249] Where N represents the number of template samples involved; T is the number of template samples in the current block; This indicates that a one-way prediction of the template sample points is performed based on the MV (L0 or L1) of the current block. Then, these two filters are applied to the two one-way predictions of the current block respectively, and then combined to generate the final two-way prediction for the current block, as shown below.

[0250] P′ bi (x,y)=w0*P′0(x,y)+w1*P′1(x,y) (18)

[0251] in,

[0252]

[0253] Where P0(x,y) and P1(x,y) are the two unidirectional prediction samples of the current block before the proposed adaptive filtering is applied. Similar to (13) and (14), in addition, to further improve the encoding and decoding gain, offset and nonlinear terms can be introduced when deriving the filter coefficients. With such modifications, the filter coefficients can be derived as follows.

[0254]

[0255] Furthermore, the filtered one-way prediction samples of the current block are calculated as follows:

[0256]

[0257] In one or more examples, a linear model (i.e., scaling factors and offsets) is proposed to derive a two-tap filter to enhance the two unidirectional predictions of a bidirectional prediction block. Specifically, a bidirectional prediction LIC is proposed, which operates as follows: 1) generating two unidirectional predictions of a template; 2) using template samples and their corresponding unidirectional prediction samples to derive two sets of scaling factors and offsets, as shown below.

[0258]

[0259] Where α0 and β0 are the scaling factors and offsets of the L0 unidirectional LIC linear model, and α1 and β1 are the scaling factors and offsets of the L1 unidirectional LIC linear model; N is the number of template samples involved in the derivation process. Then, the final bidirectional prediction for the current block is generated, as shown below.

[0260] P′ bi (x,y)=w0*(α0·P0(x,y)+β0)+w1*(α1·P0(x,y)+β1) (23)

[0261] Here, w0 and w1 are the BCW weights applied to the current block.

[0262] Recursive unidirectional filtering based on adaptive bidirectional predictive filtering

[0263] exist Figure 13 In this context, since the filter coefficients for the two unidirectional prediction signals applied to the template are derived separately, the resulting bidirectional prediction signal (i.e., a weighted combination of the two filtered unidirectional prediction signals) may not be optimal when minimizing the distortion between the template sample and its corresponding prediction sample. To address this issue, an iterative scheme is proposed to derive the optimal filter coefficients for the two unidirectional prediction signals applied to the template of a bidirectional prediction block. The proposed scheme operates iteratively, alternately optimizing the prediction filter for one prediction direction while keeping the prediction filter for the other prediction direction fixed. Specifically, the derivation process of the two unidirectional prediction filter coefficients is summarized as follows:

[0264] Step 1: Given the initial prediction direction L (0) By enabling one-way prediction Minimizing the distortion between the filter and template T yields the initial filter coefficients for the initial prediction direction. Right now,

[0265]

[0266] Step 2: Based on filter coefficients Calculate the filtered one-way prediction And set k=1.

[0267]

[0268] Step 3: Select the target prediction direction L (k) =1-L (k-1) And calculate the target template sample points for the current block, as shown below.

[0269]

[0270] Step 4: By enabling one-way prediction With template T (k) Minimizing the distortion between them yields the initial prediction direction L. (k) filter coefficients Right now,

[0271]

[0272] Step 5: Based on filter coefficients Calculate the filtered one-way prediction As shown below

[0273]

[0274] Step 6: Set k = k + 1 and go to step 3.

[0275] The obtained filter is used as the corresponding filter for the two unidirectional predictions applied to the current block, and then the filtered prediction samples are combined to generate the final bidirectional prediction for the current block, as shown in (18) and (19). Similarly, the offset and nonlinear terms shown in (20) and (21) can also be applied to the proposed iterative bidirectional prediction filter derivation scheme. In addition, in one or more examples, it is proposed to use a linear model (i.e., scaling factor and offset) to derive a two-tap filter through the proposed iterative filter derivation scheme: 1) generate two unidirectional predictions of the template; 2) derive two sets of scaling factors and offsets based on the iterative algorithm shown in steps 1 to 6; 3) calculate the final bidirectional prediction samples for the current block, as shown in (23).

[0276] In practical applications, different numbers of iterations can be applied to the above iterative filter derivation scheme. Generally, the more iterations, the smaller the distortion between the template and its predicted signal (i.e., the better the encoding / decoding gain), but this comes at the cost of increased computational complexity. Different methods are proposed below to determine the number of iterations applied in the proposed algorithm. One method proposes using a fixed number of iterations (i.e., 3) at both the encoder and decoder. A second method proposes allowing the encoder to freely choose a specific number of iterations and transmitting the corresponding value to the decoder as a signal. When applying this method, new syntax elements (multiple) can be added at the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), picture header, strip header, or even the code block level to indicate the applied iteration value. A third method proposes adaptively determining the iteration value applied to a block based on block statistics (e.g., sample changes, motion vector differences, and additional information). In one or more examples, the difference between the original L0 and L1 predicted samples of a bidirectional prediction block is proposed as the criterion for selecting the number of iterations applied. For example, when the difference between two predicted samples (i.e., the sum of absolute differences (SAD), the sum of squared differences (SSD), and other matrices) is greater than a threshold, a larger number of iterations are applied to the block; otherwise (i.e., the difference is less than the threshold), a smaller number of iterations are applied.

[0277] Last but not least, different initial prediction directions can be applied in the proposed scheme. One approach proposes always using L0 as the initial prediction direction. Another approach proposes using L1 as the initial prediction direction. A third approach proposes selecting the initial prediction direction based on the stripe type, prediction structure, and QP of the current block. For example, L0 can be used as the initial prediction direction for non-low-latency images, while L1 can be used for low-latency images.

[0278] Adaptive motion compensation filtering for signal transmission

[0279] In practice, various signal transmission schemes can be applied to indicate the use of adaptive motion compensation filtering for bidirectional prediction inter-frame blocks. In one embodiment of this disclosure, for explicit inter-frame mode (i.e., AMVP mode), it is proposed to signal a control flag to explicitly indicate whether adaptive motion compensation filtering is applied to the current block. When the flag is one, it indicates that adaptive filtering is applied to motion compensation prediction samples, and the corresponding filter coefficients are derived from template samples using one of the methods discussed above. Otherwise, when the flag is zero, it indicates that adaptive filtering is not applied to the current block. On the other hand, for merging mode, it is proposed to inherit its control flag from its corresponding selected merging candidate (indicated by the merging index), in addition to other motion information (e.g., MV, reference index, and additional information).

[0280] Adaptive motion compensation filtering based on non-adjacent spatial neighbor blocks

[0281] In some embodiments, the blocks surrounding the current block are defined as the neighboring blocks of the current block. For example... Figure 17A and Figure 17B As shown, the blank neighboring blocks without shades are considered adjacent neighboring blocks, while those with shades are considered non-adjacent neighboring blocks. In the above method, the coefficients of the proposed motion compensation filter are always derived from the reconstructed samples adjacent to the current coding block (i.e., the neighboring samples immediately above and to the left). This scheme can be efficient when the current block is highly correlated with its adjacent spatial neighboring blocks. However, in real-world encoding / decoding scenarios, due to encoding / decoding noise (e.g., noise caused by quantization / dequantization and block artifacts introduced during the motion compensation stage), the current block may be more correlated with samples in the reconstructed regions that are not adjacent to the current block. Based on this consideration, this section proposes an adaptive motion compensation filtering scheme based on non-adjacent neighboring blocks. Using this scheme, the coefficients of the adaptive motion compensation filter can be derived using samples in non-adjacent regions. Different methods can be applied to locate the non-adjacent reconstructed samples used to derive the filter coefficients. In one or more embodiments, non-adjacent neighboring blocks can be scanned from the left and top regions of the current block. The scan distance can be defined as the number of scan block sizes from the left or top of the current block.

[0282] As shown in Figure 17, multiple rows (columns) of non-adjacent neighboring blocks can be scanned at the top or left of the current block. The distances shown in Figure 17 represent the number of scan block sizes from each candidate location to the current block, with each scan block size representing one distance unit. For example, a region to the left of the current block with a "distance 2" indicates that candidate neighboring blocks located in that region are 2 scan block sizes away from the current block. Based on this pattern, different scan block sizes can be applied:

[0283] In one method, such as Figure 17A As shown, non-adjacent neighboring blocks at each distance can have the same block size as the current block. Note that when this method is applied, the granularity of the block scan is adaptively adjusted according to the partition granularity of the current block; that is, larger coded blocks have a greater chance of utilizing more distant non-adjacent reconstructed samples to compute the coefficients of the adaptive filter.

[0284] In another approach, non-neighboring blocks that can be used for filter coefficient derivation can be defined based on fixed blocks (e.g., 4×4, 8×8).

[0285] In the third method, a combination approach can be used to define the scan pattern. For example, for small blocks, a fixed scan block size (Ws × Hs) can be applied, where Ws and Hs are the width and height of the fixed scan block size; while for large blocks, the scan block size is defined as the current block size. Specifically, let xStep and yStep represent the width and height of the scan block size, with corresponding values ​​of xStep = max(Ws, width) and yStep = max(Hs, height), where the width and height are the width and height of the current block, respectively.

[0286] To indicate the use of non-adjacent neighbor blocks for filter derivation, the spatial candidate list can be formed to include adjacent neighbor blocks (i.e., spatially adjacent reconstructed samples immediately above and to the left) and non-adjacent neighbor blocks. In some embodiments, an index can be signaled from the encoder to the decoder to specify which spatial candidate to select for deriving the filter coefficients.

[0287] Alternatively, in some examples, it is proposed to apply the proposed non-adjacent spatial neighbor blocks to an existing LIC design, where the proposed adaptive motion compensation filter degenerates into a 2-tap filter (i.e., one scaling and one offset). Specifically, based on the motion information of the current block (one-way or two-way prediction), the method uses the motion information to generate corresponding prediction signals for selected non-adjacent blocks, which are then used to derive the corresponding LIC parameters by minimizing the difference between the reconstructed samples of the non-adjacent blocks and their corresponding predictions.

[0288] Adaptive motion compensation filtering based on historical filter coefficients

[0289] In the aforementioned scheme based on non-adjacent neighbor blocks, the filter coefficients are derived from reconstructed regions far from the current block, requiring additional on-chip memory to store these non-adjacent reconstructed samples. This is relatively costly for practical hardware codec implementations. Therefore, to reduce implementation costs, a history-based adaptive motion compensation filtering method is proposed. In this method, the filter coefficients of a previously encoded block are stored in a table and can be used for filtering motion compensation samples of future blocks. In some embodiments, this table can be a list of candidate filters. The table with multiple sets of filter coefficients can be maintained and synchronized during the encoding and decoding processes. Whenever an inter-frame block is encoded or decoded, a set of filter coefficients can be derived based on its reconstructed samples and their predicted samples, and then this set of filter coefficients is added as a new candidate to the last entry of the table. To maintain the size of the table, a first-in-first-out (FIFO) rule can be used, where redundancy checks can be applied to check if there is a candidate in the table that is the same as the new candidate. If so, the same candidate is removed from the table, and all other candidates are moved forward, with the new candidate added at the last entry. When the table is full and there are no duplicate candidates, the first candidate is removed from the table and a new candidate is added at the end. The candidate set of filter coefficients can then be selected for filtering motion-compensated samples in future coding blocks. For signal transmission, when selecting historically based filter coefficient derivations, an index can be used to indicate which candidate set in the table will be used to derive the filter coefficients for the current block. In another embodiment, to reduce the number of filter coefficient derivations, it is proposed to include only the filter coefficients of coding blocks that have selected adaptive motion-compensated filtering in the table.

[0290] Alternatively, in some examples, it is proposed to apply the proposed history-based filter derivation scheme to existing LIC designs, where the proposed adaptive motion compensation filter degenerates into a 2-tap filter. Specifically, in this case, each candidate in the table consists of two parameters, namely a scaling and an offset, which can be selected by an inter-frame coding block to adjust its prediction samples.

[0291] Combination of Adaptive Motion Compensation Filtering and OBMC

[0292] This section presents a method for applying the proposed adaptive motion compensation filtering method to the OBMC process. Specifically, in some example methods, in addition to the motion vectors of neighboring blocks, the influence of the LIC parameters of each neighboring block on its corresponding motion compensation prediction samples is considered when performing the OBMC process on the current block. For ease of description, the proposed method is illustrated below using conventional inter-frame prediction without sub-block partitioning as an example. For example, let P... obmc(x,y) represents the mixed prediction sample at coordinates (x,y) after combining the prediction signal of the current CU with multiple prediction signals of the MV based on its spatial neighboring blocks. cur (x,y) represents the predicted sample point at the current CU coordinates (x,y); P top (x,y) and P left (x,y) represents the predicted sample points located at the same position as the current CU but using the MV of the left and right neighboring blocks of the CU, respectively. In some embodiments, as shown in equation (29), P obmc (x, y) can be P cur (x,y), P top (x,y) and P left The weighted average of (x,y).

[0293] P obmc (x,y)=w cur *P cur (x,y)+w top *P top (x,y)+w left *P left (x,y) (29)

[0294] Furthermore, for ease of explanation, assume that adaptive motion compensation filtering is applied to the current block and its top and left spatially neighboring blocks, and that the applied filter is a tapped filter (i.e., a scaling factor and an offset), with filter coefficients as follows: For the current block, α... cur and β cur The top neighboring block is α top and β top And the left neighbor is α left and β left The proposed scheme first generates predicted samples for the current block, as shown below.

[0295] P cur (x,y)=α cur ·P org cur (x,y)+β cur (30)

[0296] Among them, P org cur (x,y) are the raw prediction samples of the current block using the current block's motion vector without applying filtering. Then, the boundary prediction samples of the current CU are updated using the MVs of the top and left causal neighboring blocks of the current CU. First, the top neighboring block of the current block is checked. If the block is an inter-frame block, its MV and filter coefficients (i.e., α) are... top and β topThe signal will be assigned to the current block to generate the prediction signal P at the corresponding position of the current block. top (x,y), as shown below.

[0297] P top (x,y)=α top ·P org top (x,y)+β top (31)

[0298] Among them, P org top (x, y) are the original predicted samples of the current block using the motion vectors of the top neighboring block without applying filtering. Then, following the same procedure, corresponding predicted samples are generated based on the motion vectors of the left neighboring block and the LIC parameters, as shown below.

[0299] P left (x,y)=α left ·P org left (x,y)+β left (32)

[0300] Among them, P org left (x,y) are the original predicted samples of the current block using the motion vectors of its left neighboring blocks without any filtering applied. Finally, these three predicted signals are combined according to a template-based OBMC mixing process (as shown in the "Overlapping Block Motion Compensation" section) to generate the final predicted samples of the current block.

[0301] When the current block is encoded and decoded in a sub-block mode (e.g., affine, ATMVP, and DMVR), the proposed motion-compensated filtering-based OBMC can also be applied to the internal OBMC of the sub-blocks within the current CU. Specifically, when applying this scheme, the same filtering process as shown in equations (29) to (31) can be applied to generate the corresponding prediction samples of each sub-block using the top, left, bottom, and right neighboring sub-blocks. However, in the derivation of the prediction samples in the internal OBMC process, the filter coefficients of the current CU, rather than the LIC parameters of the spatially neighboring blocks, will always be applied.

[0302] To achieve different complexity / performance tradeoffs, this paper proposes two methods for applying the proposed motion-compensated filter OBMC. In one method, the filter-based OBMC is proposed to be applied only to prediction samples on the CU boundary, and not to prediction samples of sub-blocks within the CU (i.e., internal OBMC). In this case, for internal OBMC, only the neighboring motion vectors of neighboring blocks of each sub-block are considered to generate its OBMC prediction samples. In the other method, the filter-based OBMC is proposed to be applied to prediction samples on the CU boundary as well as to prediction samples along the sub-block boundaries of the internal sub-blocks of the CU.

[0303] Furthermore, in equations (30) and (31), adaptive filter parameters of neighboring blocks are applied to generate corresponding predicted samples for the OBMC process of the current block. Such a design can increase the complexity of the hardware / software implementation due to the derivation of the neighboring block's filter. To reduce complexity, in one embodiment of this disclosure, it is proposed to use the filter parameters of the current block instead of the filter parameters of the neighboring blocks when generating OBMC predicted samples from the neighboring blocks. Specifically, when encoding and decoding the current block with adaptive motion compensation filtering enabled, the filter coefficients of the current block are used to modify the OBMC predicted samples generated from the motion information of each neighboring block. Otherwise, if adaptive motion compensation filtering is not applied to the current block, it is not used in the OBMC process to generate any predicted samples from the neighboring blocks, even if the neighboring blocks themselves have applied adaptive motion compensation filtering.

[0304] Furthermore, in a specific example, it is proposed to apply the above methods to existing LIC designs. Specifically, for all the methods discussed above, the adaptive motion compensation filtering process degenerates into a 2-tap filter, i.e., a scaling factor plus an offset.

[0305] Combination of adaptive motion compensation filtering and template matching-based inter-frame tools

[0306] As discussed in the "Introduction" section, several template-matching-based techniques are introduced in ECM to reduce the overhead of saving merged patterns. For example, in ARMC, candidates in the initial merge candidate list are divided into subgroups, and candidates in each subgroup are reordered based on the cost between the template sample and its corresponding predicted sample (i.e., the reference sample). In this way, candidates with better MV (i.e., lower template cost) are associated with smaller merge indices. Similarly, in MMVD patterns, all possible MMVD refinement positions are reordered using template costs, and only the encoder / decoder is allowed to select the first few positions after reordering. In this disclosure, a method is proposed to apply the proposed adaptive motion compensation filtering to the cost calculation of the template-matching-based scheme.

[0307] In the first approach, when applying adaptive motion compensation filtering to a merging candidate, it is proposed to always bypass the adaptive motion compensation filtering when calculating its template cost. However, if the candidate is selected (e.g., as shown in the merging index), the adaptive motion compensation filtering is still applied to generate the predicted samples for the block. To illustrate the above approach, see... Figure 21 As shown, assume there are L merger candidates, namely, M0, M1, ..., M L-1 Furthermore, this approach lacks generality; it assumes that candidate M will be merged. i and M j An adaptive motion-compensated filter is applied, while other merging candidates are not. This method bypasses the adaptive motion-compensated filter when calculating the difference between the template sample and the corresponding template prediction sample using the motion of L merging candidates during the template-based reordering process. However, this depends on whether M is ultimately selected. i Or M j Adaptive motion compensation filtering can still be applied to generate the final predicted samples for the block.

[0308] In the second method, adaptive motion compensation filtering is proposed to be applied to the calculation of template cost and the generation of predicted samples for the block. For example... Figure 22 As shown, unlike the first method, an adaptive motion compensation filter is applied to adjust M before calculating its corresponding cost. i and M jThe template prediction samples are obtained. Furthermore, during the reordering process, when a merge candidate is bidirectionally predicted, different adaptive filtering methods can be applied to adjust the template prediction samples. In one method (Method #1), the method discussed in the "Adaptive Bidirectional Prediction Filtering Based on Template Bidirectional Prediction Samples" section is proposed to generate the template prediction samples for each bidirectional prediction merge candidate. Specifically, in the proposed scheme, bidirectional prediction samples of the template samples are first generated based on the L0 and L1 MV of the merge candidate; then, an adaptive filter is derived and applied to the template prediction samples of the bidirectional prediction, as shown in (14). In a second method (Method #2), the method discussed in the "Adaptive Bidirectional Prediction Filtering Based on Template Unidirectional Samples" section is proposed to generate the template prediction samples for each bidirectional prediction merge candidate. Specifically, in this method, template prediction samples in L0 and L1 are first generated using the MV in L0 and L1 respectively; then, two adaptive filters are derived and applied to the L0 and L1 prediction samples of the template in a one-sided manner, and then they are combined to generate the final prediction samples of the template, as shown in (18) and (19). In the third method (method #3), it is proposed to use the method discussed in the section on "Adaptive Bidirectional Prediction Filtering Based on Recursive One-Way Filtering" to generate template prediction samples for each bidirectional prediction merging candidate. Specifically, with this scheme, template prediction samples in L0 and L1 are first generated using the MV in L0 and L1 respectively; then, two filters are iteratively derived and applied to the two one-way prediction samples of the template, and then they are combined to generate the final prediction samples of the template. In fact, different adaptive filtering schemes can be applied to different template matching schemes, which may lead to different encoding / decoding efficiency / complexity trade-offs. In a specific example, it is proposed to apply method #3 to ARMC mode and regular MMVD mode, and method #2 to affine MMVD mode.

[0309] Furthermore, in a specific example, it is proposed to apply the above methods to existing LIC designs. Specifically, for all the methods discussed above, the adaptive motion compensation filtering process degenerates into a 2-tap filter, i.e., a scaling factor plus an offset.

[0310] Combination of adaptive motion compensation filtering and AMVP-merging mode

[0311] As discussed previously, in the AMVP-merge model, the merge candidate corresponding to a given AMVP part is implicitly determined by minimizing the bilateral matching cost between the AMVP part and the merged part. For each merge candidate in the merge candidate list, the bilateral matching cost is calculated using the merge candidate MV and the AMVP MV. The merge candidate with the minimum cost for the AMVP MV is selected. Furthermore, when bilateral matching is enabled, the bilateral matching is refined and applied to the coding block starting from the selected merge candidate MV and the AMVP MV; otherwise, if template matching is enabled, the template matching is refined and applied to the coding block starting from the selected merge candidate MV and the AMVP MV. As analyzed previously, the proposed adaptive motion compensation filter can compensate for illumination variations between the predicted block and the current block. Therefore, when applying the adaptive motion compensation filter to a merge candidate, bilateral matching, which aims to measure the average illumination difference between the two blocks, may not efficiently evaluate the effectiveness of the candidate. Based on this consideration, in one embodiment of this disclosure, for a given AMVP MV, when there is one or more merge candidates in its corresponding merge candidate list, it is proposed to use template matching cost to select the merge candidate associated with the AMVP MV. Specifically, in this case, for each merge candidate in the merge candidate list, the template matching cost is calculated using the merge candidate MV and the AMVP MV. Furthermore, for merge candidates associated with adaptive motion compensation filtering, a filtering process is applied when calculating the corresponding template cost. Figure 23 An example is given to illustrate this approach. In another embodiment, it is proposed that a bilateral matching cost is always applied to select a merge candidate for each AMVP MV, regardless of whether any merge candidates are associated with the adaptive motion compensation filter.

[0312] Furthermore, in a specific example, it is proposed to apply the above methods to existing LIC designs. Specifically, for all the methods discussed above, the adaptive motion compensation filtering process degenerates into a 2-tap filter, i.e., a scaling factor plus an offset.

[0313] Combination of Adaptive Motion Compensation Filtering and CIIP

[0314] As discussed previously, in existing CIIP designs, inter-frame prediction samples are always generated solely based on the corresponding motion before mixing inter-frame prediction samples and intra-frame prediction samples. In one embodiment of this disclosure, to further improve encoding and decoding performance, it is proposed to also apply the proposed adaptive motion compensation filtering process to generate the corresponding CIIP inter-frame prediction samples. Specifically, if the merging candidate for the inter-frame portion is associated with adaptive motion compensation filtering (e.g., through merging inheritance), adaptive motion filtering is applied to modify the motion compensation prediction samples generated based on their MV. In the above method, when adaptive motion compensation filtering is enabled for a current block, it is always invoked to generate the CIIP inter-frame prediction samples. Such a design may introduce non-negligible encoding / decoding complexity. To control computational complexity, in another embodiment of this disclosure, it is proposed to enable adaptive motion compensation filtering for a CIIP block only when the current image is a low-latency image and the POC of all reference images is no greater than the POC of the current image. In another embodiment of this disclosure, it is proposed that adaptive motion compensation filtering be enabled for the CIIP block only when the current image is a low-latency image and the POC distance between the current image and the first reference image in list L0 is equal to 1.

[0315] Furthermore, in a specific example, it is proposed to apply the above methods to existing LIC designs. Specifically, for all the methods discussed above, the adaptive motion compensation filtering process degenerates into a 2-tap filter, i.e., a scaling factor plus an offset.

[0316] Figure 24 A computing environment (or computing device) 2410 coupled to a user interface 2460 is shown. The computing environment 2410 may be part of a data processing server. In some embodiments, the computing device 2410 may perform any of the various methods or processes (such as encoding / decoding methods or processes) described above according to various examples of this disclosure. The computing environment 2410 may include a processor 2420, a memory 2440, and an I / O interface 2450.

[0317] Processor 2420 typically controls the overall operation of computing environment 2410, such as operations associated with display, data acquisition, data communication, and image processing. Processor 2420 may include one or more processors for executing instructions to perform all or some of the steps described above. Furthermore, processor 2420 may include one or more modules that facilitate interaction between processor 2420 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, GPU, etc.

[0318] Memory 2440 is configured to store various types of data to support the operation of computing environment 2410. Memory 2440 may include predefined software 2442. Examples of such data include instructions for any application or method operating on computing environment 2410, video datasets, image data, etc. Memory 2440 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0319] I / O interface 2450 provides an interface between processor 2420 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, home button, start scan button, and stop scan button. I / O interface 2450 can be coupled to encoder and decoder.

[0320] In an embodiment, a non-transitory computer-readable storage medium is also provided, which, for example, includes in memory 2430 a plurality of programs executable by processor 2420 in computing environment 2410 to perform the methods described above, and / or stores a bitstream generated by the encoding method described above or a bitstream to be decoded by the decoding method described above. In one example, the plurality of programs may be executed by processor 2420 in computing environment 2410 to receive (e.g., from...) Figure 1G The video encoder 20 in the computing environment 2410 includes a bitstream or data stream of encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and can also be executed by the processor 2420 in the computing environment 2410 to perform the above-described decoding method based on the received bitstream or data stream. In another example, multiple programs can be executed by the processor 2420 in the computing environment 2410 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 2420 in the computing environment 2410 to transmit the bitstream or data stream (e.g., transmit to...). Figure 2B The video decoder 30 in the middle). Alternatively, a non-transitory computer-readable storage medium may store a bit stream or data stream therein, the bit stream or data stream comprising data generated by an encoder (e.g., Figure 1G The video encoder 20 in the code uses, for example, the encoding method described above to generate encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.) for the decoder (e.g., Figure 2BThe video decoder 30) decodes the video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0321] In one embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In another embodiment, a bitstream is provided, including encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.

[0322] In one embodiment, a computing device is also provided, the computing device comprising: one or more processors (e.g., processor 2420); and a non-transitory computer-readable storage medium or memory 2430 storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the above-described method when executing the plurality of programs.

[0323] In one embodiment, a computer program product is also provided, having instructions for storing or transmitting a bitstream, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product is also provided, which includes, for example, multiple programs in a memory 2430, executable by a processor 2420 in a computing environment 2410, for performing the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0324] In an embodiment, the computing environment 2410 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0325] In one embodiment, a method for storing a bitstream is also provided, comprising storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.

[0326] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.

[0327] Figure 25This is a flowchart illustrating an example of a video decoding method according to the present disclosure. The method can be implemented by a decoder for decoding inter-frame coded blocks. In step 2501, the method includes, in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coded block, obtaining by the decoder a template matching cost of a plurality of motion vector candidates from a plurality of first reconstructed samples adjacent to the current inter-frame coded block. In step 2502, the method includes obtaining a target motion vector from the plurality of motion vector candidates by the decoder based on the template matching cost of the plurality of first reconstructed samples. In step 2503, the method includes obtaining a plurality of predicted samples by the decoder based on the target motion vector and the current inter-frame coded block.

[0328] In some examples, obtaining the target motion vector from multiple motion vector candidates based on the template matching cost of multiple first reconstructed samples includes: reordering multiple motion vector candidates in the motion vector candidate list by the decoder based on the template matching cost of multiple first reconstructed samples; and obtaining the target motion vector by the decoder based on the motion vector candidate list.

[0329] In some examples, the template matching cost for obtaining multiple motion vector candidates for multiple first reconstructed samples includes: obtaining candidate predictions for the first reconstructed samples by the decoder; obtaining filtered candidate predictions by the decoder based on at least one candidate filter and the candidate predictions; and obtaining the template matching cost for the filtered candidate predictions, wherein at least one candidate filter is obtained based on candidate templates for the first reconstructed samples, wherein the candidate templates include multiple second reconstructed samples adjacent to the first reconstructed samples.

[0330] In some examples, at least one candidate filter is obtained by the following steps: obtaining a candidate template of a first reconstructed sample by a decoder; predicting at least one candidate template of the first reconstructed sample by the decoder based on the motion vector candidate of the first reconstructed sample; and obtaining at least one candidate filter by the decoder based on the prediction of at least one candidate template and the candidate template.

[0331] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and candidate template includes: obtaining a combined candidate template prediction based on multiple candidate template predictions; and obtaining the coefficients of a candidate filter by minimizing the difference between the combined candidate template prediction and the candidate template.

[0332] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and candidate prediction includes: obtaining a combined candidate prediction based on multiple candidate predictions; and obtaining a filtered candidate prediction by applying a candidate filter to the combined candidate prediction.

[0333] In some examples, obtaining candidate predictions for the first reconstructed sample includes obtaining a first candidate prediction and a second candidate prediction.

[0334] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and the candidate template includes: obtaining a first coefficient of the first candidate filter by minimizing the difference between the first candidate template prediction and the candidate template; and obtaining a second coefficient of the second candidate filter by minimizing the difference between the second candidate template prediction and the candidate template.

[0335] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and a candidate prediction includes: obtaining a first filtered candidate prediction by applying a first candidate filter to a first candidate prediction; obtaining a second filtered candidate prediction by applying a second candidate filter to a second candidate prediction; and obtaining a filtered candidate prediction by combining the first filtered candidate prediction and the second filtered candidate prediction.

[0336] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and the candidate template includes: calculating a target template based on the candidate template and a previously filtered candidate template prediction; obtaining the coefficients of the current candidate filter by minimizing the difference between the current candidate template prediction and the target template; and calculating the current filtered candidate template prediction by applying the current candidate filter to the current candidate template prediction.

[0337] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and candidate template further includes: obtaining the coefficients of the first candidate filter by minimizing the difference between the first candidate template prediction and the candidate template; and calculating the previously filtered candidate template prediction by applying the first candidate filter to the first candidate template prediction.

[0338] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and a candidate prediction includes: obtaining a first filtered candidate prediction by applying a first candidate filter to a first candidate prediction; obtaining a second filtered candidate prediction by applying a second candidate filter to a second candidate prediction; and obtaining a filtered candidate prediction by combining the first filtered candidate prediction and the second filtered candidate prediction.

[0339] In some examples, the list of motion vector candidates is applied to the Advanced Motion Vector Prediction (AMVP) - merge mode.

[0340] In some examples, the method further includes: in response to determining that the adaptive motion compensation filter has not been applied to any first reconstructed sample, obtaining by the decoder a bilateral matching cost for each of the plurality of first reconstructed samples; and obtaining by the decoder a target motion vector based on the bilateral matching costs of the plurality of first reconstructed samples.

[0341] In some examples, the bilateral matching cost of obtaining each of the multiple first reconstructed samples includes: obtaining the AMVP block of the current inter-coding block by the decoder by applying the AMVP mode; obtaining the merged prediction block of the current inter-coding block by the decoder by applying the merge mode; and obtaining the bilateral matching cost by the decoder based on the AMVP block and the merged prediction block.

[0342] Figure 26 This is a flowchart illustrating an example of a video decoding method according to the present disclosure. The method can be implemented by a decoder for decoding inter-frame coded blocks. In step 2601, the method includes obtaining an intra-prediction block of the current inter-frame coded block by the decoder. In step 2602, the method includes obtaining a plurality of inter-frame prediction blocks of the current inter-frame coded block by the decoder in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coded block. In step 2603, the method includes obtaining a filtered inter-frame prediction block by the decoder based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on a current template of the current inter-frame coded block, wherein the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coded block. In step 2604, the method includes obtaining a final prediction block by the decoder by combining the intra-prediction block and the filtered inter-frame prediction block.

[0343] In some examples, at least one template filter is obtained by the following steps: obtaining a current template of the current inter-frame coding block by the decoder; obtaining multiple template predictions of the current template corresponding to the multiple prediction blocks of the current inter-frame coding block by the decoder; and obtaining at least one template filter by the decoder based on the multiple template predictions and the current template.

[0344] In some examples, the method further includes, in response to determining that the picture order count (POC) of all reference pictures to be used in inter-frame coding is less than the POC of the current picture of the current inter-frame coding block, the decoder determines to apply adaptive motion compensation filtering to the current inter-frame coding block.

[0345] In some examples, the POC of all reference pictures to be used in inter-frame coding and decoding is less than the POC of the current picture in the current inter-frame coding block, including: the POC distance between the current picture and the first reference picture is equal to 1, and the POC distance between the first reference picture and the current picture is the smallest among all reference pictures.

[0346] Figure 27 This is a flowchart illustrating an example of a video coding method according to the present disclosure. The method can be implemented by an encoder for encoding inter-frame coding blocks. In step 2701, the method includes, in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coding block, obtaining by the encoder a template matching cost of a plurality of motion vector candidates from a plurality of first reconstructed samples adjacent to the current inter-frame coding block. In step 2702, the method includes obtaining a target motion vector from the plurality of motion vector candidates by the encoder based on the template matching cost of the plurality of first reconstructed samples. In step 2703, the method includes obtaining a plurality of predicted samples by the encoder based on the target motion vector and the current inter-frame coding block.

[0347] In some examples, obtaining the target motion vector from multiple motion vector candidates based on the template matching cost of multiple first reconstructed samples includes: reordering multiple motion vector candidates in the motion vector candidate list by the encoder based on the template matching cost of multiple first reconstructed samples; and obtaining the target motion vector by the encoder based on the motion vector candidate list.

[0348] In some examples, the template matching cost for obtaining multiple motion vector candidates for multiple first reconstructed samples includes: obtaining candidate predictions for the first reconstructed samples by an encoder; obtaining filtered candidate predictions by an encoder based on at least one candidate filter and the candidate predictions; and obtaining the template matching cost for the filtered candidate predictions, wherein at least one candidate filter is obtained based on candidate templates for the first reconstructed samples, wherein the candidate templates include multiple second reconstructed samples adjacent to the first reconstructed samples.

[0349] In some examples, at least one candidate filter is obtained by the following steps: obtaining a candidate template of a first reconstructed sample by an encoder; predicting at least one candidate template of the first reconstructed sample by the encoder based on the motion vector candidate of the first reconstructed sample; and obtaining at least one candidate filter by the encoder based on the prediction of at least one candidate template and the candidate template.

[0350] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and candidate template includes: obtaining a combined candidate template prediction based on multiple candidate template predictions; and obtaining the coefficients of a candidate filter by minimizing the difference between the combined candidate template prediction and the candidate template.

[0351] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and candidate prediction includes: obtaining a combined candidate prediction based on multiple candidate predictions; and obtaining a filtered candidate prediction by applying a candidate filter to the combined candidate prediction.

[0352] In some examples, obtaining candidate predictions for the first reconstructed sample includes obtaining a first candidate prediction and a second candidate prediction.

[0353] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and the candidate template includes: obtaining a first coefficient of the first candidate filter by minimizing the difference between the first candidate template prediction and the candidate template; and obtaining a second coefficient of the second candidate filter by minimizing the difference between the second candidate template prediction and the candidate template.

[0354] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and a candidate prediction includes: obtaining a first filtered candidate prediction by applying a first candidate filter to a first candidate prediction; obtaining a second filtered candidate prediction by applying a second candidate filter to a second candidate prediction; and obtaining a filtered candidate prediction by combining the first filtered candidate prediction and the second filtered candidate prediction.

[0355] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and the candidate template includes: calculating a target template based on the candidate template and a previously filtered candidate template prediction; obtaining the coefficients of the current candidate filter by minimizing the difference between the current candidate template prediction and the target template; and calculating the current filtered candidate template prediction by applying the current candidate filter to the current candidate template prediction.

[0356] In some examples, obtaining at least one candidate filter based on at least one candidate template prediction and candidate template further includes: obtaining the coefficients of the first candidate filter by minimizing the difference between the first candidate template prediction and the candidate template; and calculating the previously filtered candidate template prediction by applying the first candidate filter to the first candidate template prediction.

[0357] In some examples, obtaining a filtered candidate prediction based on at least one candidate filter and a candidate prediction includes: obtaining a first filtered candidate prediction by applying a first candidate filter to a first candidate prediction; obtaining a second filtered candidate prediction by applying a second candidate filter to a second candidate prediction; and obtaining a filtered candidate prediction by combining the first filtered candidate prediction and the second filtered candidate prediction.

[0358] In some examples, the list of motion vector candidates is applied to the Advanced Motion Vector Prediction (AMVP) - merge mode.

[0359] In some examples, the method further includes: in response to determining that the adaptive motion compensation filter has not been applied to any first reconstructed sample, obtaining by the encoder a bilateral matching cost for each of a plurality of first reconstructed samples; and obtaining by the encoder a target motion vector based on the bilateral matching costs of the plurality of first reconstructed samples.

[0360] In some examples, the bilateral matching cost of obtaining each of the multiple first reconstructed samples includes: obtaining the AMVP block of the current inter-coding block by the encoder by applying the AMVP mode; obtaining the merged prediction block of the current inter-coding block by the encoder by applying the merge mode; and obtaining the bilateral matching cost by the encoder based on the AMVP block and the merged prediction block.

[0361] Figure 28 This is a flowchart illustrating an example of a video coding method according to the present disclosure. The method can be implemented by an encoder for encoding inter-frame coded blocks. In step 2801, the method includes obtaining an intra-prediction block of the current inter-frame coded block by the encoder. In step 2802, the method includes obtaining a plurality of inter-frame prediction blocks of the current inter-frame coded block by the encoder in response to determining that an adaptive motion compensation filter is applied to the current inter-frame coded block. In step 2803, the method includes obtaining a filtered inter-frame prediction block by the encoder based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on a current template of the current inter-frame coded block, wherein the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coded block. In step 2804, the method includes obtaining a final prediction block by the encoder by combining the intra-prediction block and the filtered inter-frame prediction block.

[0362] In some examples, at least one template filter is obtained by the following steps: obtaining a current template of the current inter-coding block by the encoder; obtaining multiple template predictions of the current template corresponding to the multiple prediction blocks of the current inter-coding block by the encoder; and obtaining at least one template filter by the encoder based on the multiple template predictions and the current template.

[0363] In some examples, the method further includes, in response to determining that the picture order count (POC) of all reference pictures to be used in inter-frame coding is less than the POC of the current picture of the current inter-frame coding block, the encoder determines to apply adaptive motion compensation filtering to the current inter-frame coding block.

[0364] In some examples, the POC of all reference pictures to be used in inter-frame coding and decoding is less than the POC of the current picture in the current inter-frame coding block, including: the POC distance between the current picture and the first reference picture is equal to 1, and the POC distance between the first reference picture and the current picture is the smallest among all reference pictures.

[0365] In some examples, an apparatus for video encoding and decoding is provided. The apparatus includes a processor 2420 and a memory 2440, the memory being configured to store instructions executable by the processor; wherein the processor, when executing the instructions, is configured to perform actions such as... Figures 25 to 28 Any of the methods shown.

[0366] In some other examples, a non-transitory computer-readable storage medium is provided in which instructions are stored. When the instructions are executed by processor 2420, the instructions cause the processor to perform actions such as Figures 19 to 21 Any of the methods shown. In one example, multiple programs can be executed by processor 2420 in computing environment 2410 to receive (e.g., from...) Figure 1G The video encoder 20 in the computing environment 2410 includes a bitstream or data stream of encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and can also be executed by the processor 2420 in the computing environment 2410 to perform the above-described decoding method based on the received bitstream or data stream. In another example, multiple programs can be executed by the processor 2420 in the computing environment 2410 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 2420 in the computing environment 2410 to transmit the bitstream or data stream (e.g., transmit to...). Figure 2B The video decoder 30 in the middle). Alternatively, a non-transitory computer-readable storage medium may store a bit stream or data stream therein, the bit stream or data stream comprising data generated by an encoder (e.g., Figure 1G The video encoder 20 in the code uses encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.) generated by, for example, the encoding method described above, for the decoder (e.g., Figure 2B The video decoder 30) decodes the video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0367] The description in this disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.

[0368] Unless otherwise specified, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.

[0369] The examples were chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.

[0370] The methods described above can be implemented using an apparatus comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus can be used in conjunction with other hardware or software components for performing the methods described above. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using these one or more circuits.

[0371] Considering the specification and practice of the disclosure herein, other examples of this disclosure will be apparent to those skilled in the art. This application is intended to cover any changes, uses, or adaptations of the disclosure made in accordance with its general principles, including deviations from this disclosure within the scope of known or customary practice in the art. The specification and examples are intended to be illustrative only.

[0372] It should be understood that this disclosure is not limited to the exact examples shown in the description and figures above, and various modifications and changes can be made without departing from its scope.

Claims

1. A video decoding method, comprising: In response to determining that adaptive motion compensation filtering is applied to the current inter-frame coding block, the decoder obtains template matching costs for multiple motion vector candidates of multiple first reconstructed samples adjacent to the current inter-frame coding block; The decoder obtains the target motion vector from the plurality of motion vector candidates based on the template matching cost of the plurality of first reconstructed samples; as well as The decoder obtains multiple prediction samples based on the target motion vector and the current inter-frame coding block.

2. The method as described in claim 1, wherein, Obtaining the target motion vector from the plurality of motion vector candidates based on the template matching cost of the plurality of first reconstructed samples includes: The decoder reorders the plurality of motion vector candidates in the motion vector candidate list based on the template matching cost of the plurality of first reconstructed samples; and The decoder obtains the target motion vector based on the motion vector candidate list.

3. The method as described in claim 1, wherein, The template matching cost for obtaining the plurality of motion vector candidates for the plurality of first reconstructed samples includes: The decoder obtains candidate predictions for the first reconstructed sample. The decoder obtains filtered candidate predictions based on at least one candidate filter and the candidate predictions; and Obtain the template matching cost of the filtered candidate prediction. The at least one candidate filter is obtained based on a candidate template of the first reconstructed sample point, wherein the candidate template includes a plurality of second reconstructed sample points adjacent to the first reconstructed sample point.

4. The method of claim 3, wherein, The at least one candidate filter is obtained through the following steps: The decoder obtains the candidate template for the first reconstructed sample. The decoder obtains at least one candidate template prediction for the first reconstructed sample based on the motion vector candidates of the first reconstructed sample; as well as The decoder obtains the at least one candidate filter based on the at least one candidate template prediction and the candidate template.

5. The method of claim 4, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: A combined candidate template prediction is obtained based on predictions of multiple candidate templates; and The coefficients of a candidate filter are obtained by minimizing the difference between the combined candidate template prediction and the candidate template.

6. The method of claim 5, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: Obtain a combined candidate prediction based on multiple candidate predictions; and The filtered candidate prediction is obtained by applying the candidate filter to the combined candidate prediction.

7. The method of claim 4, wherein, Obtaining the candidate predictions for the first reconstructed sample points includes: Obtain the first candidate prediction and the second candidate prediction.

8. The method of claim 7, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: The first coefficients of the first candidate filter are obtained by minimizing the difference between the prediction of the first candidate template and the candidate template; and The second coefficients of the second candidate filter are obtained by minimizing the difference between the prediction of the second candidate template and the candidate template.

9. The method of claim 8, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: A first filtered candidate prediction is obtained by applying the first candidate filter to the first candidate prediction; A second filtered candidate prediction is obtained by applying the second candidate filter to the second candidate prediction; and The filtered candidate prediction is obtained by combining the first filtered candidate prediction and the second filtered candidate prediction.

10. The method of claim 7, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: The target template is calculated based on the candidate template and the previously filtered candidate template prediction; The coefficients of the current candidate filter are obtained by minimizing the difference between the current candidate template prediction and the target template; and The current filtered candidate template prediction is calculated by applying the current candidate filter to the current candidate template prediction.

11. The method of claim 10, wherein, Predicting the at least one candidate template and obtaining the at least one candidate filter based on the at least one candidate template further includes: The coefficients of the first candidate filter are obtained by minimizing the difference between the prediction of the first candidate template and the candidate template; and The previously filtered candidate template prediction is computed by applying the first candidate filter to the first candidate template prediction.

12. The method of claim 10, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: A first filtered candidate prediction is obtained by applying a first candidate filter to the first candidate prediction; A second filtered candidate prediction is obtained by applying a second candidate filter to the second candidate prediction; and The filtered candidate prediction is obtained by combining the first filtered candidate prediction and the second filtered candidate prediction.

13. The method of claim 2, wherein, The candidate list of motion vectors is applied to the Advanced Motion Vector Prediction (AMVP) - Merge mode.

14. The method of claim 13, further comprising: In response to determining that the adaptive motion compensation filter was not applied to any first reconstructed sample, the decoder obtains the bilateral matching cost for each of the plurality of first reconstructed samples; as well as The target motion vector is obtained by the decoder based on the bilateral matching cost of the plurality of first reconstructed samples.

15. The method of claim 14, wherein, The bilateral matching cost for obtaining each of the plurality of first reconstructed samples includes: The decoder obtains the AMVP block of the current inter-frame coded block by applying the AMVP mode; The decoder obtains the merged prediction block of the current inter-frame coded block by applying a merging mode; and The decoder obtains the bilateral matching cost based on the AMVP block and the merged prediction block.

16. A video decoding method, comprising: The decoder obtains the intra-predicted block of the current inter-frame coded block; In response to determining that adaptive motion compensation filtering is applied to the current inter-frame coded block, the decoder obtains multiple inter-frame prediction blocks of the current inter-frame coded block; The decoder obtains a filtered inter-frame prediction block based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on the current template of the current inter-frame coding block, and the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coding block; and The decoder obtains the final prediction block by combining the intra-frame prediction block with the filtered inter-frame prediction block.

17. The method of claim 16, wherein, The at least one template filter is obtained through the following steps: The current template of the current inter-frame coded block is obtained by the decoder; The decoder obtains multiple template predictions of the current template, each corresponding to one of the multiple prediction blocks of the current inter-frame coded block; as well as The decoder obtains the at least one template filter based on the multiple template predictions and the current template.

18. The method of claim 16, further comprising: In response to determining that the Picture Order Count (POC) of all reference pictures to be used in inter-frame coding and decoding is less than the POC of the current picture of the current inter-frame coding block, the decoder determines to apply the adaptive motion compensation filter to the current inter-frame coding block.

19. The method of claim 18, wherein, The Proof of Concept (POC) of all reference images to be used in inter-frame encoding and decoding is less than the POC of the current image in the current inter-frame coding block, including: The POC distance between the current image and the first reference image is equal to 1, and among all reference images, the POC distance between the first reference image and the current image is the smallest.

20. A video coding method, comprising: In response to determining that adaptive motion compensation filtering is applied to the current inter-frame coding block, the encoder obtains template matching costs for multiple motion vector candidates of multiple first reconstructed samples adjacent to the current inter-frame coding block; The encoder obtains the target motion vector from the plurality of motion vector candidates based on the template matching cost of the plurality of first reconstructed samples; as well as The encoder obtains multiple prediction samples based on the target motion vector and the current inter-frame coding block.

21. The method of claim 20, wherein, Obtaining the target motion vector from the plurality of motion vector candidates based on the template matching cost of the plurality of first reconstructed samples includes: The encoder reorders the plurality of motion vector candidates in the motion vector candidate list based on the template matching cost of the plurality of first reconstructed samples; and The encoder obtains the target motion vector based on the motion vector candidate list.

22. The method of claim 20, wherein obtaining the template matching cost of the plurality of motion vector candidates for the plurality of first reconstructed samples comprises: The encoder obtains candidate predictions for the first reconstructed sample points; The encoder obtains filtered candidate predictions based on at least one candidate filter and the candidate predictions; as well as Obtain the template matching cost of the filtered candidate prediction. The at least one candidate filter is obtained based on a candidate template of the first reconstructed sample point, wherein the candidate template includes a plurality of second reconstructed sample points adjacent to the first reconstructed sample point.

23. The method of claim 22, wherein, The at least one candidate filter is obtained through the following steps: The encoder obtains the candidate template for the first reconstructed sample point; The encoder obtains at least one candidate template prediction for the first reconstructed sample based on the motion vector candidates of the first reconstructed sample; as well as The encoder obtains the at least one candidate filter based on the at least one candidate template prediction and the candidate template.

24. The method of claim 23, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: A combined candidate template prediction is obtained based on predictions of multiple candidate templates; and The coefficients of a candidate filter are obtained by minimizing the difference between the combined candidate template prediction and the candidate template.

25. The method of claim 24, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: Obtain a combined candidate prediction based on multiple candidate predictions; and The filtered candidate prediction is obtained by applying the candidate filter to the combined candidate prediction.

26. The method of claim 23, wherein, Obtaining the candidate predictions for the first reconstructed sample points includes: Obtain the first candidate prediction and the second candidate prediction.

27. The method of claim 26, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: The first coefficients of the first candidate filter are obtained by minimizing the difference between the prediction of the first candidate template and the candidate template; and The second coefficients of the second candidate filter are obtained by minimizing the difference between the prediction of the second candidate template and the candidate template.

28. The method of claim 27, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: A first filtered candidate prediction is obtained by applying the first candidate filter to the first candidate prediction; A second filtered candidate prediction is obtained by applying the second candidate filter to the second candidate prediction; and The filtered candidate prediction is obtained by combining the first filtered candidate prediction and the second filtered candidate prediction.

29. The method of claim 26, wherein, Obtaining the at least one candidate filter based on the at least one candidate template prediction and the candidate template includes: The target template is calculated based on the candidate template and the previously filtered candidate template prediction; The coefficients of the current candidate filter are obtained by minimizing the difference between the current candidate template prediction and the target template; and The current filtered candidate template prediction is calculated by applying the current candidate filter to the current candidate template prediction.

30. The method of claim 29, wherein, Predicting the at least one candidate template and obtaining the at least one candidate filter based on the at least one candidate template further includes: The coefficients of the first candidate filter are obtained by minimizing the difference between the prediction of the first candidate template and the candidate template; and The previously filtered candidate template prediction is computed by applying the first candidate filter to the first candidate template prediction.

31. The method of claim 29, wherein, Obtaining the filtered candidate prediction based on the at least one candidate filter and the candidate prediction includes: A first filtered candidate prediction is obtained by applying a first candidate filter to the first candidate prediction; A second filtered candidate prediction is obtained by applying a second candidate filter to the second candidate prediction; and The filtered candidate prediction is obtained by combining the first filtered candidate prediction and the second filtered candidate prediction.

32. The method of claim 21, wherein, The candidate list of motion vectors is applied to the Advanced Motion Vector Prediction (AMVP) - Merge mode.

33. The method of claim 32, further comprising: In response to determining that the adaptive motion compensation filter was not applied to any first reconstructed sample, the encoder obtains the bilateral matching cost for each of the plurality of first reconstructed samples; as well as The encoder obtains the target motion vector based on the bilateral matching cost of the plurality of first reconstructed samples.

34. The method of claim 33, wherein, The bilateral matching cost for obtaining each of the plurality of first reconstructed samples includes: The encoder obtains the AMVP block of the current inter-frame coded block by applying the AMVP mode; The encoder obtains the merged prediction block of the current inter-frame coded block by applying a merging mode; and The encoder obtains the bilateral matching cost based on the AMVP block and the merged prediction block.

35. A video coding method, comprising: The encoder obtains the intra-predicted block of the current inter-frame coded block; In response to determining that adaptive motion compensation filtering is applied to the current inter-frame coding block, the encoder obtains multiple inter-frame prediction blocks for the current inter-frame coding block; The encoder obtains a filtered inter-frame prediction block based on at least one template filter and the plurality of inter-frame prediction blocks, wherein the at least one template filter is obtained based on the current template of the current inter-frame coding block, and the current template includes a plurality of reconstructed samples adjacent to the current inter-frame coding block; and The encoder obtains the final prediction block by combining the intra-frame prediction block with the filtered inter-frame prediction block.

36. The method of claim 35, wherein, The at least one template filter is obtained through the following steps: The encoder obtains the current template of the current inter-frame coded block; The encoder obtains multiple template predictions of the current template, each corresponding to one of the multiple prediction blocks of the current inter-frame coded block; as well as The encoder obtains the at least one template filter based on the multiple template predictions and the current template.

37. The method of claim 35, further comprising: In response to determining that the picture order count (POC) of all reference pictures to be used in inter-frame coding is less than the POC of the current picture of the current inter-frame coding block, the encoder determines to apply the adaptive motion compensation filter to the current inter-frame coding block.

38. The method of claim 37, wherein, The Proof of Concept (POC) of all reference images to be used in inter-frame encoding and decoding is less than the POC of the current image in the current inter-frame coding block, including: The POC distance between the current image and the first reference image is equal to 1, and among all reference images, the POC distance between the first reference image and the current image is the smallest.

39. An apparatus for video decoding, the apparatus comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 1 to 19 when executing the instructions.

40. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 1 to 19.

41. An apparatus for video encoding, the apparatus comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 20 to 38 when executing the instructions.

42. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 20 to 38.

43. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1-19.

44. A non-transitory computer-readable storage medium for storing a bit stream generated by the method of any one of claims 20-38.