Motion compensation considering out-of-boundary conditions in video coding.

The video encoding method addresses OOB conditions by deriving predictor samples with motion compensation and weightings, enhancing encoding efficiency and quality maintenance.

JP7819300B2Active Publication Date: 2026-02-24BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024519468
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-30
Filing Date
2022-10-31
Publication Date
2026-02-24
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing video encoding methods struggle with out-of-bounds (OOB) conditions in video coding, leading to inefficiencies and quality degradation in video compression.

Method used

A video encoding method that addresses OOB conditions by deriving predictor samples using motion compensation processes and applying additional weightings based on OOB determinations to obtain final predictor samples.

Benefits of technology

Enhances video encoding efficiency by effectively handling OOB conditions, maintaining video quality during compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819300000062
    Figure 0007819300000062
  • Figure 0007819300000063
    Figure 0007819300000063
  • Figure 0007819300000064
    Figure 0007819300000064
Patent Text Reader

Abstract

A video encoding method, an apparatus, and a non-transitory computer-readable storage medium are provided. In one method, a decoder derives a first reference picture and a second reference picture for a current coding block. The decoder uses a motion compensation process from the first reference picture to derive a first predictor sample based on a first motion vector associated with the first reference picture. The decoder uses a motion compensation process from the second reference picture to derive a second predictor sample based on a second motion vector associated with the second reference picture. The decoder obtains a final predictor sample in the current coding block based on at least one of the first predictor sample or the second predictor sample and an out-of-bounds (OOB) condition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to Provisional Application No. 63 / 273,930, filed October 30, 2021, the entire contents of which are incorporated herein by reference for all purposes. FIELD OF THE DISCLOSURE This disclosure relates to video encoding and compression. More particularly, this disclosure relates to methods and apparatus for inter prediction in video encoding. [Background technology]

[0002] Digital video is supported by a variety of electronic devices, including digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video teleconferencing devices, and video streaming devices. Electronic devices transmit and receive or otherwise communicate digital video data over communication networks and / or store the digital video data on storage devices. Due to limited bandwidth capacity in communication networks and limited memory resources in storage devices, video coding may be used to compress video data according to one or more video coding standards before the video data is communicated or stored. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Expert Group (MPEG) coding, and others. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]

[0003] Examples of this disclosure provide a video encoding method and apparatus using intra prediction. [Means for solving the problem]

[0004] According to a first aspect of the present disclosure, a video encoding method is provided. The method may include: deriving, by a decoder, a first reference picture and a second reference picture for a current coding block; deriving, by the decoder, a first predictor sample based on a first motion vector associated with the first reference picture using a motion compensation process from the first reference picture; deriving, by the decoder, a second predictor sample based on a second motion vector associated with the second reference picture using a motion compensation process from the second reference picture; and obtaining, by the decoder, a final predictor sample in the current coding block based on at least one of the first predictor sample or the second predictor sample and an out-of-bounds (OOB) condition.

[0005] According to a second aspect of the present disclosure, a video encoding method is provided. The method may include: determining, by a decoder, whether a predictor sample of a reference picture for a currently coded block is out of bounds (OOB); in response to determining that the predictor sample is OOB, assigning, by the decoder, an additional weighting of zero to the predictor sample when combining two or more predictor samples generated by a motion compensation process to obtain a final predictor sample of the currently coded block; and in response to determining that the predictor sample is not OOB, assigning, by the decoder, an additional weighting other than zero to the predictor sample when combining the two or more predictor samples generated by the motion compensation process to obtain the final predictor sample of the currently coded block.

[0006] According to a second aspect of the present disclosure, a video encoding method is provided. The method may include: determining, by a decoder, whether a predictor sample of a reference picture for a current coding block is out of bounds (OOB) based on integer reference samples used to generate the predictor sample; assigning, by the decoder, in response to determining that the predictor sample is OOB, a first additional weighting to the predictor sample when combining two or more predictor samples generated by a motion compensation process to obtain a final predictor sample of the current coding block; and assigning, by the decoder, in response to determining that the predictor sample is not OOB, a second additional weighting to the predictor sample when combining the two or more predictor samples generated by the motion compensation process to obtain the final predictor sample of the current coding block.

[0007] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to be restrictive of the present disclosure.

[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram illustrating an example system for encoding and decoding video blocks, according to some implementations of the present disclosure. [Figure 2] 1 is a block diagram illustrating an example video encoder according to some implementations of the present disclosure. [Figure 3] 1 is a block diagram illustrating an example video decoder according to some implementations of the present disclosure. [Figure 4A] A block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some implementations of this disclosure. [Figure 4B]A block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some implementations of this disclosure. [Figure 4C] A block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some implementations of this disclosure. [Figure 4D] A block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some implementations of this disclosure. [Figure 4E] A block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some implementations of this disclosure. [Figure 5] 10A and 10B are diagrams illustrating examples of spatial merge candidate locations, according to some implementations of the present disclosure. [Figure 6] FIG. 10 illustrates an example of candidate pairs considered for redundancy check of spatial merge candidates, according to some implementations of the present disclosure. [Figure 7] FIG. 10 illustrates an example motion vector scaling of a temporal merge candidate, according to some implementations of the present disclosure. [Figure 8] FIG. 10 illustrates candidate locations for temporal merge candidates, according to some implementations of the present disclosure. [Figure 9A] FIG. 10 is a diagram illustrating an example of MMVD search points according to some implementations of the present disclosure. [Figure 9B] FIG. 10 is a diagram illustrating an example of MMVD search points according to some implementations of the present disclosure. [Figure 10A] FIG. 2 illustrates an example of a control point-based affine motion model, according to some implementations of the present disclosure. [Figure 10B] FIG. 2 illustrates an example of a control point-based affine motion model, according to some implementations of the present disclosure. [Figure 11] FIG. 10 illustrates an example of a sub-block-wise affine MVF according to some implementations of the present disclosure. [Figure 12] FIG. 10 illustrates the location of inherited affine motion predictors, according to some implementations of the present disclosure. [Figure 13] FIG. 10 illustrates the locations of candidate positions for constructed affine merge modes, in accordance with some implementations of the present disclosure. [Figure 14] FIG. 10 is a diagram illustrating an example of sub-blocks MVs and pixels according to some implementations of the present disclosure. [Figure 15A] A diagram showing an example of spatially adjacent blocks used by the ATVMP, in accordance with some implementations of the present disclosure. [Figure 15B] FIG. 1 illustrates an example of an SbTMVP process, according to some implementations of the present disclosure. [Figure 16] FIG. 10 is a diagram illustrating an example of an extended CU region used in BDOF, according to some implementations of the present disclosure. [Figure 17] FIG. 10 is a diagram illustrating an example of decoding-side motion vector refinement according to some implementations of the present disclosure. [Figure 18] FIG. 10 illustrates an example of GPM splits grouped by the same angle, according to some implementations of the present disclosure. [Figure 19] FIG. 10 is a diagram illustrating an example of uni-predictive MV selection for geometric partitioning mode, according to some implementations of the present disclosure. [Figure 20] 10A-10C are diagrams illustrating examples of upper and left neighboring blocks used in CIIP weight derivation, according to some implementations of the present disclosure. [Figure 21] FIG. 10 illustrates an example of spatially adjacent blocks used to derive spatial merge candidates, according to some implementations of the present disclosure. [Figure 22] FIG. 10 illustrates an example of template matching performed on a search area around an initial MV, according to some implementations of the present disclosure. [Figure 23] FIG. 10 illustrates an example of a diamond region within a search area, according to some implementations of the present disclosure. [Figure 24]FIG. 10 illustrates the frequency response of an interpolation filter at half pixel phase and a VVC interpolation filter according to some implementations of the present disclosure. [Figure 25] FIG. 10 is a diagram illustrating an example of a template and a reference sample of the template in a reference picture, according to some implementations of the present disclosure. [Figure 26] FIG. 10 is a diagram illustrating an example of a template and a reference sample for a block with sub-block motion using motion information of the sub-blocks of a current block, according to some implementations of the present disclosure. [Figure 27] FIG. 10 is a diagram illustrating an example of a padded reference picture for motion compensation, according to some implementations of the present disclosure. [Figure 28] FIG. 10 illustrates an example of fractional interpolation, according to some implementations of the present disclosure. [Figure 29] FIG. 1 is a block diagram illustrating a computing environment coupled with a user interface, according to some implementations of the present disclosure. [Figure 30] FIG. 1 is a block diagram illustrating a video encoding process, according to some implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations described in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Rather, the implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure as set forth in the appended claims.

[0011] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It is also to be understood that the term "and / or," as used herein, is intended to mean and include any and all possible combinations of one or more of the associated listed items.

[0012] Terms such as "first," "second," and "third" may be used herein to describe various pieces of information, but it should be understood that these terms are not intended to limit the information. These terms are used only to distinguish one category of information from another. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. The term "if," as used herein, may be understood to mean "when," "in the event of," or "at the discretion of," depending on the context.

[0013] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, today, well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its predecessor, VP9. Audio-Video Coding (AVS), referring to digital audio and digital video compression standards, is another video compression standard series developed by the China Audio and Video Coding Standards Workgroup. Most of the existing video coding standards are built on the well-known hybrid video coding framework, i.e., using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in a video image or sequence, and using transform coding to compact the energy of the prediction error. An important goal of video coding techniques is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0014] The first generation of AVS standards includes the Chinese national standards "Information Technology, Advanced Audio-Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio-Video Coding, Part 16: Radio, Television, and Video" (known as AVS+). Compared to the MPEG-2 standard, this achieves approximately 50% bitrate savings at the same perceptual quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards includes a series of Chinese national standards, "Information Technology, Efficient Multimedia Coding" (known as AVS2), primarily aimed at transmitting additional HD TV programs. AVS2's coding efficiency is twice that of AVS+. AVS2 was published as a Chinese national standard in May 2016. Meanwhile, the video portion of the AVS2 standard has been proposed by the Institute of Electrical and Electronics Engineers (IEEE) as one of the international standards for applications. The AVS3 standard is one of a new generation of video coding standards for UHD video applications, aiming to exceed the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS Conference, the AVS3-P2 baseline was finalized, achieving approximately 30% bitrate savings compared to the HEVC standard. Currently, there is one reference software, called the High Performance Model (HPM), maintained by the AVS group to demonstrate a reference implementation of the AVS3 standard.

[0015] 1 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0016] In some implementations, destination device 14 may receive encoded video data to be decoded over link 16. Link 16 may comprise any type of communication medium or device capable of moving encoded video data from source device 12 to destination device 14. In one example, link 16 may comprise a communication medium that enables source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.

[0017] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, or local disk drives. Destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0018] 1 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, implementations described herein may be applicable to video encoding generally and may be applicable to wireless and / or wired applications.

[0019] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0020] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or modem and receive encoded video data over link 16. The encoded video data communicated over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0021] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may comprise any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0022] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or future standards.

[0023] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If implemented partially in software, the electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.

[0024] 2 is a block diagram illustrating an example video encoder 20 according to some implementations described herein. Video encoder 20 may perform intra-predictive and inter-predictive coding of video blocks within video frames. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. Note that the term "frame" is sometimes used synonymously with the terms "image" or "picture" in the field of video coding.

[0025] 2, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, adder 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58 for video block reconstruction, inverse transform processing unit 60, and adder 62. An in-loop filter 63, such as a deblocking filter, may be disposed between adder 62 and DPB 64 to filter block boundaries and remove blocky artifacts from the reconstructed video. In addition to the deblocking filter, another in-loop filter, such as a Sample Adaptive Offset (SAO) filter and / or an Adaptive in-loop filter (ALF), may also be used to filter the output of summer 62. In some examples, the in-loop filter may be omitted, and the decoded video block may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.

[0026] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from video source 18 shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or reference pictures) for use in encoding video data by video encoder 20 (e.g., in an intra-predictive coding mode or an inter-predictive coding mode). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with other components of video encoder 20 or may be off-chip relative to those components.

[0027] As shown in FIG. 2, after receiving video data, partitioning unit 45 in prediction processing unit 41 partitions the video data into video blocks. This partitioning may include partitioning the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predetermined partitioning structure, such as a quad-tree (QT) structure associated with the video data. A video frame may be or be considered as a two-dimensional array or matrix of samples having sample values. The samples in the array may also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into multiple video blocks, for example, by using QT partitioning. A video block may also be or be considered as a two-dimensional array or matrix of samples having sample values, although with smaller dimensions than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further partitioned into one or more block partitions or sub-blocks (which may again form blocks), e.g., by repeatedly using QT partitioning, binary-tree (BT) partitioning, or triple-tree (TT) partitioning, or any combination thereof. Note that the term “block” or “video block” as used herein may refer to a portion of a frame or picture, particularly a rectangular (square or non-square) portion. For example, a root block or video block referring to HEVC and VVC may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or may be or correspond to a corresponding block, e.g., a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB), and / or sub-block.

[0028] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., coding rate and distortion level). Prediction processing unit 41 may provide the resulting intra-predictively coded block or inter-predictively coded block to summer 50 to generate a residual block and to summer 62 to reconstruct the coded block for later use as part of a reference frame. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy coding unit 56.

[0029] To select an appropriate intra-prediction coding mode for a current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0030] In some implementations, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating motion vectors that indicate the displacement of video blocks in the current video frame relative to predictive blocks in a reference video frame according to a predetermined pattern in the sequence of video frames. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. The motion vectors may indicate, for example, the displacement of video blocks in a current video frame or picture relative to predictive blocks in a reference frame relative to the current block being coded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to the determination of motion vectors by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine block vectors.

[0031] A prediction block for a video block may be or correspond to a block or reference block of a reference frame that is deemed to closely match the video block to be encoded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of square differences (SSD), or other difference metrics. In some implementations, video encoder 20 may calculate values ​​for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches for whole pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0032] Motion estimation unit 42 calculates a motion vector for a video block in an inter-predictively coded frame by comparing the position of the video block to the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy coding unit 56.

[0033] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may locate the predictive block pointed to by the motion vector in one of the reference frame lists, obtain the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual video block of pixel difference values ​​by subtracting pixel values ​​of the predictive block provided by motion compensation unit 44 from pixel values ​​of the current video block being coded. The pixel difference values ​​forming the residual video block may include a luma difference component, a chroma difference component, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, a syntax element defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.

[0034] In some implementations, the intra BC unit 48 may generate a vector and fetch a predictive block in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is within the same frame as the current block being coded, and the vector is referred to as a block vector rather than a motion vector. Specifically, the intra BC unit 48 may determine an intra prediction mode to use to code the current block. In some examples, the intra BC unit 48 may code the current block using various intra prediction modes, e.g., during separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 may then select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values ​​using the rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics from the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to create the coded block, as well as the bit rate (i.e., number of bits) used to create the coded block. Intra BC unit 48 may calculate ratios from the distortions and rates of the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.

[0035] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction according to implementations described herein. In either case, for intra block copying, the predictive block may be a block that is deemed to closely match the block to be coded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the predictive block may include calculating values ​​for sub-integer pixel positions.

[0036] Regardless of whether the predictive block is a block from the same frame via intra prediction or a block from a different frame via inter prediction, video encoder 20 may form a residual video block by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being coded to form pixel difference values. The pixel difference values ​​forming the residual video block may include both luma and chroma component differences.

[0037] Intra-prediction processing unit 46 may intra-predict the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra-block copy prediction performed by intra BC unit 48, as described above. Specifically, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. To do so, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.

[0038] After prediction processing unit 41 determines a predictive block for the current video block by inter-prediction or intra-prediction, adder 50 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0039] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0040] Following quantization, entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding methodology or technique. The coded bitstream may then be transmitted to video decoder 30 as shown in FIG. 1 or archived to storage device 32 as shown in FIG. 1 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also entropy code motion vectors and other syntax elements of the current video frame being coded.

[0041] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate reference blocks for prediction of other video blocks. As described above, motion compensation unit 44 may generate motion-compensated prediction blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values ​​for use in motion estimation.

[0042] Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to create a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.

[0043] 3 is a block diagram illustrating an example video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.

[0044] In some examples, units of video decoder 30 may be tasked with performing implementations of the present application. Also, in some examples, implementations of the present disclosure may be divided among one or more of the units of video decoder 30. For example, intra BC unit 85 may perform implementations of the present application alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85, and the functions of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.

[0045] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. A data buffer (DPB) 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra-predictive or inter-predictive coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous dynamic random access memory (SDRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), or other types of memory devices. 3, video data memory 79 and DPB 92 are depicted as two separate components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or may be off-chip relative to those components.

[0046] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.

[0047] When a video frame is coded as an intra-predictively coded (I) frame or for intra-coded predictive blocks within other types of frames, intra prediction unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on a signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.

[0048] When a video frame is coded as an inter-predictive (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 creates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be created from a reference frame in one of the reference frame lists. Video decoder 30 may construct reference frame lists List 0 and List 1 using a default construction technique based on the reference frames stored in DPB 92.

[0049] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 creates a predictive block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be within the same reconstructed region of the picture as the current video block defined by video encoder 20.

[0050] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by analyzing motion vectors and other syntax elements, and then use the prediction information to create a predictive block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) used to encode the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), construction information regarding one or more of the frame's reference frame lists, the motion vectors of each inter-predictively coded video block of the frame, the inter-prediction status of each inter-predictively coded video block of the frame, and other information for decoding video blocks in the current video frame.

[0051] Similarly, intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block was predicted using intra BC mode, construction information regarding which video blocks of the frame are within the reconstruction domain and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.

[0052] Motion compensation unit 82 may also perform interpolation using an interpolation filter used by video encoder 20 during encoding of the video block to calculate sub-integer pixel interpolated values ​​of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter used by video encoder 20 from the received syntax element and use the interpolation filter to create the predictive block.

[0053] Inverse quantization unit 86 uses the same quantization parameter calculated by video encoder 20 for each video block in a video frame to inverse quantize the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct residual blocks in the pixel domain.

[0054] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs a decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. To further process the decoded video block, an in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF, may be disposed between adder 90 and the DPB. In some examples, in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by adder 90. The decoded video block in a given frame is then stored in DPB 92, which stores reference frames used for subsequent motion compensation of the next video block. DPB 92 or a memory device separate from DPB 92 may store the decoded video for later presentation on a display device, such as display device 34 of FIG. 1.

[0055] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame may be monochromatic and therefore include only one two-dimensional array of luma samples.

[0056] As shown in FIG. 4A, video encoder 20 (or, more specifically, partitioning unit 45) generates a coded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, which may be one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to encode the samples of the coding tree blocks. The syntax elements describe how a video sequence may be reconstructed at video decoder 30, including characteristics of different types of units of coded blocks of pixels, as well as inter- or intra-prediction, intra-prediction modes, motion vectors, and other parameters. For monochrome pictures or pictures with three separate color planes, a CTU may include a single coding tree block and syntax elements used to code samples of the coding tree block. A coding tree block may be an N×N block of samples.

[0057] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of the CTU to divide the CTU into smaller CUs. As depicted in FIG. 4C , 64×64 CTU 400 is first partitioned into four smaller CUs, each with a block size of 32×32. Of the four smaller CUs, CU 410 and CU 420 are each partitioned into four 16×16 CUs by block size. Two 16×16 CUs, CU 430 and CU 440, are each further partitioned into four 8×8 CUs by block size. FIG. 4D depicts a quad tree data structure showing the final result of the partitioning process of CTU 400 depicted in FIG. 4C , where each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32×32 to 8×8. Similar to the CTU depicted in FIG. 4B, each CU may include a CB for luma samples, two corresponding coding blocks for chroma samples of the same size frame, and syntax elements used to encode the samples of the coding block. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures used to encode the samples of the coding block. Note that the quadtree partitioning depicted in FIGS. 4C and 4D is for illustrative purposes only, and a CTU may be divided into CUs based on quadtree / ternary tree / binary tree partitioning to accommodate various local characteristics. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by binary tree and ternary tree structures. As shown in FIG. 4E, there are five possible partition types of a coding block with width W and height H: quad-partition, horizontal 2-partition, vertical 2-partition, horizontal 3-partition, and vertical 3-partition.

[0058] In some implementations, video encoder 20 may further partition the coding blocks of a CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction, inter prediction or intra prediction, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PB. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and syntax structures used to predict the PB. Video encoder 20 may generate predictive luma blocks, predictive Cb blocks, and predictive Cr blocks for the luma, Cb, and Cr PBs of each PU of a CU.

[0059] Video encoder 20 may generate the predictive blocks of a PU using intra prediction or inter prediction. If video encoder 20 generates the predictive blocks of a PU using intra prediction, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of a frame associated with the PU. If video encoder 20 generates the predictive blocks of a PU using inter prediction, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0060] After video encoder 20 generates the predictive luma block, the predictive Cb block, and the predictive Cr block for one or more PUs of a CU, video encoder 20 may generate the luma residual block of the CU by subtracting the predictive luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate the Cb residual block and the Cr residual block of the CU, such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and such that each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0061] Further, as shown in FIG. 4C , video encoder 20 may use quadtree partitioning to decompose the luma residual block, Cb residual block, and Cr residual block of a CU into one or more luma transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and syntax structures used to transform the samples of the transform block.

[0062] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of the TU to generate a Cr coefficient block of the TU.

[0063] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to a process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby achieving further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a sequence of bits forming a representation of the encoded frame and associated data, which may be stored on a storage device 32 or transmitted to a destination device 14.

[0064] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks of PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing coding blocks for each CU of the frame, video decoder 30 may reconstruct the frame.

[0065] As mentioned above, video coding achieves video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that IBC can be considered as intra-frame prediction or a third mode. Of the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict a current video block from a reference video block.

[0066] However, as video data capture technology is constantly improving and video block sizes become finer to preserve details of the video data, the amount of data required to represent the motion vectors of the current frame is also increasing significantly. One way to overcome this challenge is to benefit from the fact that a group of adjacent CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between these adjacent CUs. Therefore, the motion information of spatially adjacent CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by investigating their spatial and temporal correlations, which is also called the "motion vector predictor (MVP)" of the current CU.

[0067] Instead of encoding the actual motion vector of the current CU determined by motion estimation unit 42 into the video bitstream as described above in connection with Figure 2, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to create a Motion Vector Difference (MVD) for the current CU. By doing so, the motion vector determined for each CU of a frame by motion estimation unit 42 does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream may be significantly reduced.

[0068] Similar to the process of choosing a predictive block in a reference frame during inter-frame prediction of a codeblock, a set of rules needs to be adopted by both video encoder 20 and video decoder 30 to construct a motion vector candidate list (also called a “merge list”) for the current CU using potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then select one element from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from video encoder 20 to video decoder 30; the index of the selected motion vector predictor in the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU.

[0069] For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index, a reference picture list usage index, and additional information required for the new coding features of VVC used to generate inter-predicted samples. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, it is associated with one PU and has no significant residual coefficients, coded motion vector deltas, or reference picture indexes. Merge mode is specified, whereby the motion parameters of the current CU, including spatial and temporal candidates and the additional schedule introduced in VVC, are obtained from neighboring CUs. Merge mode can be applied to any inter-predicted CU, not just skip mode. An alternative to merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index of each reference picture list, the reference picture list usage flag, and other required information are explicitly signaled for each CU.

[0070] Additionally, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) are studying the potential need for standardization of future video coding technologies with compression capabilities significantly beyond those of the current VVC standard. Such future standardization action could take the form of additional VVC extensions or an entirely new standard. These groups are working together on this search in a collaborative effort known as the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. The first exploratory experiment (EE) was established at the JVET meeting on January 6–15, 2021. This exploratory software model, named the Enhanced Compression Model (ECM), was published in August 2021. ECM Version 2 (ECM2) was released in August 2021. The newly developed inter-prediction scheme in ECM2 is described in detail in the next section.

[0071] The following sections provide details of VVC and its interprediction methodology as specified in the ECM model under development.

[0072] Enhanced merge prediction in VVC In VVC, a merge candidate list is constructed by including the following five types of candidates in order: 1) Spatial MVP from spatially adjacent CUs 2) Temporal MVP from collocated CUs 3) History-based MVP from FIFO tables 4) Average MVP per pair 5) Zero's music video

[0073] The size of the merge list is signaled in the Sequence Parameter Set header, and the maximum allowed size of the merge list is 6. For each CU code in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first bin of the merge index is coded using the context, while bypass coding is used for the other bins.

[0074] In this section, the derivation process for each category of merge candidates is provided. As is done in HEVC, VVC also supports parallel derivation of merge candidate lists for all CUs within an area of ​​a certain size.

[0075] Spatial candidate derivation in VVC The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Up to four merge candidates are selected from the candidates located in the positions depicted in Figure 5. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more CUs in positions B0, A0, B1, and A1 are unavailable (e.g., because they belong to another slice or tile) or are intra-coded. After the candidate in position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are removed from the list to improve coding efficiency. To reduce computational complexity, the aforementioned redundancy check does not consider all possible candidate pairs. Instead, only pairs linked by arrows in Figure 6 are considered, and a candidate is added to the list only if the corresponding candidates used in the redundancy check do not have the same motion information.

[0076] Temporal candidate derivation in VVC In this step, only one candidate is added to the list. In particular, in the derivation of this temporal merge candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list and reference index used to derive the co-located CU are explicitly signaled in the slice header. The scaled motion vector of the temporal merge candidate is obtained as shown by the dotted line in FIG. 7. This motion vector is scaled from the motion vector of the co-located CU using POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the constellation picture and the constellation picture. The reference picture index of the temporal merge candidate is set equal to zero.

[0077] The location of the temporal candidate is selected between candidate C0 and candidate C1, as depicted in Figure 8. If the CU at location C0 is unavailable, intra-coded, or outside the current row of the CTU, location C1 is used. Otherwise, location C0 is used to derive the temporal merge candidate.

[0078] History-Based Merge Candidate Derivation in VVC History-based MVP (HMVP) merge candidates are added to the merge list after spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP of the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row occurs, the table is reset (emptied). Whenever there is a non-subblock inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0079] The HMVP table size, S, is set to 6, which indicates that up to five history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, and a redundancy check is first applied to find whether there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the identical HMVP is inserted into the last entry of the table.

[0080] HMVP candidates may be used in the merge candidate list construction process. The most recent HMVP candidates in the table are checked in order and inserted after the TMVP candidate in the candidate list. Redundancy checks are applied to spatial or temporal merge candidates in the HMVP candidates.

[0081] To reduce the number of redundancy checking operations, the following simplifications are introduced. 1. The last two entries in the table are redundancy checked against the A1 and B1 space candidates, respectively. 2. The process of building a merge candidate list from HMVP ends when the total number of available merge candidates reaches the maximum allowed merge candidates minus one.

[0082] Pairwise Average Merge Candidate Derivation in VVC The pairwise average candidate is generated by averaging a predefined candidate pair in the existing merge candidate list using the first two merge candidates. The first merge candidate may be defined as p0Cand, and the second merge candidate may be defined as p1Cand. The averaged motion vector is calculated separately for each reference list depending on the availability of motion vectors in p0Cand and p1Cand. If both motion vectors are available in a list, these two motion vectors are averaged even if they point to different reference pictures, and the reference picture is set to one of p0Cand. If only one motion vector is available, that motion vector is used directly. If no motion vector is available, the list is disabled. Also, if the half-pixel interpolation filter indexes of p0Cand and p1Cand are different, they are set to 0.

[0083] If the merge list is not full after the pairwise average merge candidates are added, zero MVPs are inserted at the end until the maximum number of merge candidates is reached.

[0084] High-precision (1 / 16 pixel) motion compensation and motion vector storage in VVC. VVC increases MV accuracy to 1 / 16 luma sample to improve prediction efficiency for slow-motion video. This higher motion accuracy is particularly useful for video content with locally varying, non-translational motion, such as in affine mode. To generate fractional position samples with higher MV accuracy, HEVC's 8-tap luma interpolation filter and 4-tap chroma interpolation filter are extended to 16 phases for luma and 32 phases for chroma. This extended filter set is applied to the MC process for inter-coded CUs, except for affine mode CUs. In affine mode, a set of 6-tap luma interpolation filters with 16 phases is used to reduce computational complexity and save memory bandwidth.

[0085] In VVC, the highest precision of explicitly signaled motion vectors for non-affine CUs is 1 / 4 luma sample. In some inter prediction modes, such as affine mode, motion vectors may be signaled with 1 / 16 luma sample precision. In all inter-coded CUs with implicitly inferred MVs, the MVs are derived with 1 / 16 luma sample precision, and motion compensation prediction is performed with 1 / 16 luma sample precision. For intra motion field storage, all motion vectors are stored with 1 / 16 luma sample precision.

[0086] For the temporal motion field storage used by TMVP and SbTVMP, motion field compression is performed at a granularity of 8x8 size, as opposed to a granularity of 16x16 size in HEVC.

[0087] Merge Mode with MVD in VVC (MMVD) In addition to the merge mode where implicitly derived motion information is directly used to generate prediction samples for the current CU, a merge mode with motion vector difference (MMVD) is introduced in VVC. Immediately after sending the regular merge flag, an MMVD flag is signaled to the CU to specify whether the MMVD mode is used.

[0088] In MMVD, after a merge candidate is selected, the merge candidate is further refined by signaled MVD information. The additional information includes a merge candidate flag, an index for specifying the motion magnitude, and an index for indicating the motion direction. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV basis. An mmvd candidate flag is signaled to specify which of the first and second merge candidates is used.

[0089] The distance index specifies the magnitude of the motion and indicates a predetermined offset from the starting point. As shown in Figures 9A and 9B, the offset is added to either the horizontal or vertical component of the starting MV. Table 1 specifies the relationship between the distance index and the predetermined offset.

[0090] [Table 1]

[0091] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions shown in Table 2. Note that the meaning of the MVD code may change depending on the information of the starting MV. If the starting MV is a predicted MV or a bi-predictive MV and both lists point to the same side of the current picture (i.e., the POCs of the two references are both greater than or both less than the POC of the current picture), the code in Table 2 specifies the sign of the MV offset added to the starting MV. If the starting MV is a bi-predictive MV and the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), if the difference in POC in list 0 is greater than the difference in POC in list 1, the code in Table 2 specifies the sign of the MV offset added to the MV component of the starting MV of list 0, and the sign of the MV in list 1 has the opposite value. Otherwise, if the difference in POC in List 1 is greater than the difference in POC in List 0, the sign in Table 2 specifies the sign of the MV offset added to the MV component of the starting MV of List 1, and the sign of the MV in List 0 has the opposite value.

[0092] The MVD is scaled according to the POC difference in each direction. If the POC difference in both lists is the same, no scaling is required. Otherwise, if the POC difference in list 0 is larger than the POC difference in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb, as described in Figure 7. If the POC difference of L1 is larger than L0, the MVD of list 0 is also scaled in the same way. If the starting MV is uni-predicted, the MVD is added to the available MV.

[0093] [Table 2]

[0094] Symmetric MVD coding in VVC In VVC, in addition to the usual MVD signaling for unidirectional and bidirectional prediction modes, a symmetric MVD mode for bi-predictive MVD signaling is applied. In the symmetric MVD mode, motion information including reference picture indexes for both list 0 and list 1 and the MVD for list 1 is derived without being signaled.

[0095] The decoding process for symmetric MVD mode is as follows: 1. At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a forward-backward or backward-forward reference picture pair, BiDirPredFlag is set to 1 and the reference pictures in both list 0 and list 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2. At the CU level, if the CU is bi-predictively coded and BiDirPredFlag is equal to 1, then a symmetric mode flag is explicitly signaled indicating whether symmetric mode is used.

[0096] If the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of list 0 and list 1 are set equal to their respective reference picture pairs. MVD1 is set equal to (-MVD0). The final motion vector is shown in the following equation:

number

[0097] At the encoder, symmetric MVD motion estimation begins with initial MV evaluation. The set of initial MV candidates consists of the MVs obtained from the unipredictive search, the MVs obtained from the bipredictive search, and the MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.

[0098] Affine motion compensation prediction in VVC In HEVC, only a translational motion model is applied to motion compensation prediction (MCP). Meanwhile, in the real world, many types of motion exist, such as zoom, rotation, perspective motion, and other irregular motion. In VVC, block-based affine transformation motion compensation prediction is applied. As shown in Figures 10A and 10B, the affine motion field of a block is described by motion information of two control points (four parameters) as shown in Figure 10A or three control point motion vectors (six parameters) as shown in Figure 10B.

[0099] For a four-parameter affine motion model, the motion vector at a sample position (x,y) within a block is

number

[0100] For a six-parameter affine motion model, the motion vector at sample position (x,y) within a block is

number

[0101] where (mv 0x ,mv 0y ) is the motion vector of the control point in the upper left corner, and (mv 1x ,mv 1y ) is the motion vector of the control point in the upper right corner, and (mv 2x ,mv 2y ) is the motion vector of the control point in the bottom left corner.

[0102] To simplify motion-compensated prediction, block-based affine transformation prediction is applied. To derive the motion vector for each 4x4 luma sub-block, the motion vector of the center sample of each sub-block, as shown in Figure 11, is calculated according to the above equation and rounded to 1 / 16 decimal accuracy. Then, a motion-compensated interpolation filter is applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size of the chroma components is also set to 4x4. The MV of a 4x4 chroma sub-block is calculated as the average of the MVs of the top-left and bottom-right luma sub-blocks in the same 8x8 luma region.

[0103] As with translational motion inter prediction, there are also two affine motion inter prediction modes: affine merge mode and affine AMVP mode.

[0104] Affine merge prediction The AF_MERGE mode can be applied to CUs whose width and height are both 8 or more. In this mode, the CPVM for the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate the CPMVP candidate to be used for the current CU. The following three types of CPVM candidates are used to form the affine merge candidate list: - Inherited affine merge candidates extrapolated from the CPMVs of neighboring CUs - Constructed affine merge candidate CPMVP derived using translational MVs of neighboring CUs - Zero's MV

[0105] In VVC, there are up to two inherited affine candidates, derived from the affine motion models of the neighboring blocks: one from the left neighboring CU and the other from the neighboring CU above. The candidate blocks are shown in Figure 12. For the left predictor, the scanning order is A0->A1, and for the top predictor, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No pruning check is performed between the two inherited candidates. Once a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidate in the affine merge list of the current CU. If the neighboring bottom-left block A is coded in affine mode, motion vectors v_2, v_3, and v_4 are obtained for the top-left, top-right, and bottom-left corners of the CU containing block A. If block A is coded using a four-parameter affine model, the two CPMVs for the current CU are calculated according to v_2 and v_3. If block A is coded using a six-parameter affine model, the three CPMVs of the current CU are calculated according to v_2, v_3, and v_4.

[0106] Constructed affine candidates means that candidates are constructed by combining neighboring translational motion information for each control point. The motion information for a control point is derived from the specified spatial and temporal neighbors shown in Figure 13. CPMVk (k=1, 2, 3, 4) represents the kth control point. For CPMV1, the B2->B3->A2 block is checked, and the MV of the first available block is used. For CPMV2, the B1->B0 block is checked, and for CPMV3, the A1->A0 block is checked. TMVP is used as CPMV4 if available.

[0107] After the MVs of the four control points are obtained, an affine merge candidate is constructed based on their motion information. The following combinations of control point MVs are used for construction in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}

[0108] A combination of three CPMVs constructs a six-parameter affine merge candidate, and a combination of two CPMVs constructs a four-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the associated control point MV combination is discarded.

[0109] After the inherited and constructed affine merge candidates have been checked, if the list is not yet filled, a zero MV is inserted at the end of the list.

[0110] Affine AMVP prediction in VVC Affine AMVP mode can be applied to CUs whose width and height are both 16 or greater. A CU-level affine flag is signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predictor CPMVP is signaled in the bitstream. The size of the affine AVMP candidate list is 2, and this candidate list is generated by sequentially using the following four types of CPVM candidates: - Inherited affine AMVP candidates extrapolated from the CPMV of neighboring CUs - Constructed affine AMVP candidate CPMVP derived using translational MVs of neighboring CUs - Translational MV from adjacent CU - Zero's MV

[0111] The check order for inherited affine AMVP candidates is the same as that for inherited affine merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. No pruning process is applied when inserting inherited affine motion predictors into the candidate list.

[0112] The constructed AMVP candidate is derived from the specified spatial neighborhood shown in Figure 13. The same check order is used as in the construction of the affine merge candidate. In addition, the reference picture indexes of the neighboring blocks are also checked. The first block in the check order that is inter-coded and has the same reference picture as in the current CU is used. There is only one. If the current CU is coded in 4-parameter affine mode and both mv0 and mv1 are available, they are added as one candidate in the affine AMVP list. If the current CU is coded in 6-parameter affine mode and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set as unavailable.

[0113] After inserting valid inherited and constructed affine AMVP candidates, if the number of affine AMVP list candidates is still less than 2, mv0, mv1, and mv2 are added in order, when available, as translation MVs to predict all control point MVs for the current CU. Finally, if the affine AMVP list is not yet filled, zero MVs are used to fill the affine AMVP list.

[0114] Prediction refinement using optical flow of affine modes in VVC Subblock-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation, but at the expense of prediction accuracy. To achieve finer granularity in motion compensation, prediction refinement with optical flow (PROF) is used to refine subblock-based affine motion compensation predictions without increasing memory access bandwidth for motion compensation. In VVC, after subblock-based affine motion compensation is performed, luma prediction samples are refined by adding the difference derived by the optical flow equation. PROF can be explained as the following four steps:

[0115] Step 1) Sub-block based affine motion compensation is performed to generate the sub-block prediction I(i,j).

[0116] Step 2) Spatial gradients of subblock predictions g x (i,j) and g y (i,j) is calculated at each sample location using a 3-tap filter [-1,0,1]. The gradient calculation is exactly the same as the gradient calculation for BDOF.

number

number

[0117] SHIFT1 is used to control the gradient precision. Sub-block (i.e., 4x4) predictions are extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, the extended samples on the extension boundary are copied from the nearest integer pixel location in the reference picture.

[0118] Step 3) Luma prediction refinement is performed using the following optical flow equation:

number

[0119] Since the affine model parameters and sample positions relative to the center of the sub-block do not change for each sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. dx(i,j) and dy(i,j) are calculated from sample positions (i,j) to the center of the sub-block (x SB ,y SB ), then Δv(x,y) is given by the following equation:

number

number

[0120] To maintain accuracy, the subblock (x SB ,y SB ) input is ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block, respectively.

[0121] For a four-parameter affine model,

number

[0122] For a six-parameter affine model,

number

[0123] Step 4) Finally, the luma prediction refinement ΔI(i,j) is added to the sub-block prediction I(i,j). The final prediction I′ is generated as the following equation: I'(i,j)=I(i,j)+ΔI(i,j)

[0124] PROF does not apply in two cases for affine-coded CUs: 1) all control point MVs are the same, which indicates that the CU has only translational motion; and 2) the affine motion parameters are larger than the specified limit, since sub-block-based affine MC is downgraded to CU-based MC to avoid large memory access bandwidth requirements.

[0125] To reduce the coding complexity of affine motion estimation using PROF, a fast coding method is applied. PROF is not applied in the affine motion estimation stage in the following two situations: a) If this CU is not the root block and its parent block does not select affine mode as its best mode, PROF is not applied because the current CU is unlikely to select affine mode as its best mode. b) If the magnitudes of the four affine parameters (C, D, E, F) are all smaller than a predetermined threshold and the current picture is not a low-latency picture, PROF is not applied because the improvement brought by PROF in this case is small. In this way, affine motion estimation using PROF can be accelerated.

[0126] Sub-block-based Temporal Motion Vector Prediction (SbTMVP) in VVC VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses motion fields in aligned pictures to improve the motion vector prediction and merge mode of CUs in the current picture. The same aligned pictures used by TMVP are also used for SbTVMP. SbTMVP differs from TMVP in two main ways:

[0127] - TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level.

[0128] - TMVP fetches temporal motion vectors from an alignment block in an alignment picture (the alignment block is the bottom right or center block relative to the current CU), while SbTMVP applies a motion shift before fetching temporal motion information from the alignment picture, and the motion shift is obtained from a motion vector from one of the spatially adjacent blocks of the current CU.

[0129] The SbTVMP process is shown in Figures 15A and 15B. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 15A is examined. If A1 has a motion vector that uses the alignment picture as a reference picture, this motion vector is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0,0).

[0130] In the second step, as shown in FIG. 15B, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the alignment picture. The example in FIG. 15B assumes that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block in the alignment picture (the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After the motion information of the co-located sub-CU is identified, this motion information is converted to the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, and temporal motion scaling is applied to align the reference picture of the temporal motion vector to the picture of the current CU.

[0131] In VVC, a combined subblock-based merge list containing both SbTVMP candidates and affine merge candidates is used to signal the subblock-based merge mode. SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. When SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of subblock-based merge candidates, followed by the affine merge candidates. The size of the subblock-based merge list is signaled in the SPS, and the maximum allowed size of the subblock-based merge list is 5 in VVC.

[0132] The sub-CU size used in SbTMVP is fixed at 8x8, and like the affine merge mode, the SbTMVP mode only applies to CUs whose width and height are both 8 or greater.

[0133] The encoding logic of the additional SbTMVP merge candidate is the same as other merge candidates, i.e., for each CU in a P slice or B slice, an additional RD check is performed to determine whether to use the SbTMVP candidate.

[0134] Adaptive Motion Vector Resolution (AMVR) for VVC In HEVC, if use_integer_mv_flag in the slice header is equal to 0, the motion vector differential (MVD) (between the CU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. VVC introduces the CU-level adaptive motion vector resolution (AMVR) scheme. AMVR allows the CU's MVD to be coded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows: - Normal AMVP mode: quarter luma samples, half luma samples, integer luma samples, or four luma samples - Affine AMVP mode: quarter luma sample, integer luma sample, or 1 / 16 luma sample

[0135] If the current CU has at least one non-zero MVD component, a CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal and vertical MVDs in reference list L0 and reference list L1) are zero, a quarter-luma sample MVD resolution is inferred.

[0136] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MVD precision is used for that CU. If the first flag is 0, no further signaling is required and quarter luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate that half luma sample or other MVD precision (integer or 4 luma sample) is used for the regular AMVP CU. In the case of half luma sample, a 6-tap interpolation filter is used instead of the default 8-tap interpolation filter for half luma sample positions. Otherwise, a third flag is signaled to indicate whether integer luma sample or 4 luma sample MVD precision is used for the regular AMVP CU. For affine AMVP CUs, the second flag is used to indicate whether integer luma sample or 1 / 16 luma sample MVD precision is used. To ensure that the reconstructed MV has the intended precision (quarter luma samples, half luma samples, integer luma samples, or 4 luma samples), the motion vector predictor of a CU may be rounded to the same precision as the MVD precision before being added with the MVD. The motion vector predictor is rounded towards zero (i.e., negative motion vector predictors are rounded towards positive infinity, and positive motion vector predictors are rounded towards negative infinity).

[0137] The encoder uses the RD check to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, VTM14 only conditionally invokes the RD check for MVD precisions other than quarter-luma sample MVD precision. In normal AVMP mode, the RD costs of quarter-luma sample MVD precision and integer-luma sample MVD precision are first calculated. Then, to determine whether the RD cost of four-luma sample MVD precision needs to be further checked, the RD cost of integer-luma sample MVD precision is compared with the RD cost of quarter-luma sample MVD precision. If the RD cost of quarter-luma sample MVD precision is much smaller than the RD cost of integer-luma sample MVD precision, the RD check for four-luma sample MVD precision is skipped. If the RD cost of integer-luma sample MVD precision is significantly larger than the best RD cost of the previously tested MVD precision, the check for half-luma sample MVD precision is skipped. For Affine AMVP mode, if Affine Inter mode is not selected after checking the rate-distortion costs of Affine Merge / Skip mode, Merge / Skip mode, Regular AMVP mode with Quarter Luma Sample MVD precision, and Affine AMVP mode with Quarter Luma Sample MVD precision, Affine Inter modes with 1 / 16 Luma Sample MV precision and 1 Pixel MV precision are not checked. Additionally, the affine parameters obtained for Affine Inter mode with Quarter Luma Sample MV precision are used as the starting search points for Affine Inter modes with 1 / 16 Luma Sample and Quarter Luma Sample MV precision.

[0138] Bi-prediction with CU-level weight (BCW) in VVC In HEVC, a bi-predictive signal is generated by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. In VVC, the bi-predictive mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals.

number

[0139] Weighted average bi-prediction allows five weights w∈{-2,3,4,5,10}. For each bi-predicted CU, the weight w is determined in one of two ways: 1) For non-merged CUs, the weight index is signaled after the motion vector differential; 2) For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index; BCW is only applied to CUs with 256 or more luma samples (i.e., CU width × CU height is 256 or more); For low-latency pictures, all five weights are used; For non-low-latency pictures, only three weights (w∈{3,4,5}) are used.

[0140] - In the encoder, fast search algorithms are applied to find the weight indices without significantly increasing the encoder complexity. These algorithms are summarized as follows. For more details, see the VTM software and document JVET-L0646. When combined with AMVR, for 1-pixel and 4-pixel motion vector precision, unequal weights are only conditionally checked if the current picture is a low-latency picture.

[0141] When combined with affine, affine ME may be performed for unequal weights if and only if the affine mode is selected as the current best mode.

[0142] - When the two reference pictures for bi-prediction are the same, unequal weights are only conditionally checked.

[0143] - Depending on the POC distance between the current picture and its reference pictures, the coding QP and the temporal level, unequal weights are not searched if certain conditions are met.

[0144] The BCW weight index is coded using one context coding bin followed by a bypass coding bin. The first context coding bin indicates whether equal weights are used, and if unequal weights are used, additional bins are signaled using bypass coding to indicate which unequal weights are used.

[0145] Weighted prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficiently encoding video content with fading. WP support was also added to the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. The weights and offsets of the corresponding reference pictures are then applied during motion compensation. WP and BCW are designed for various types of video content. To avoid interactions between WP and BCW that can complicate VVC decoder design, when a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. This can be applied to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, affine motion information is constructed based on the motion information of up to three blocks. The BCW index of a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV.

[0146] In VVC, CIIP and BCW cannot be applied to a CU together. If a CU is coded in CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight.

[0147] Bi-directional optical flow (BDOF) in VVC VVC includes a bidirectional optical flow (BDOF) tool. BDOF was previously called BIO and was included in JEM. Compared to the JEM version, BDOF in VVC is a simpler version, requiring significantly less computation, especially in terms of the number of multiplications and the size of the multipliers.

[0148] BDOF is used to refine the bi-predictive signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: - The CU is coded using "true" bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other is after the current picture in display order. - The distances from the two reference pictures to the current picture (i.e., the POC difference) are the same - Both reference pictures are short-term reference pictures - The CU is not coded using affine mode or SbTMVP merge mode - CU has more than 64 luma samples - Both CU height and CU width are 8 or more luma samples - BCW Weight Index shows equal weighting - WP is not currently enabled for the CU - CIIP mode is not currently in use for the CU

[0149] BDOF is applied only to the luma component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes smooth object motion. For each 4x4 sub-block, motion refinement (v) is performed by minimizing the difference between the L0 and L1 predicted samples. x ,v y ) is calculated. Then, motion refinement is used to adjust the bi-predictive sample values ​​within the 4x4 sub-block. The following steps are applied in the BDOF process:

[0150] First, the horizontal and vertical gradients of the two prediction signals TIFF0007819300000014.tif30170(k=0, 1) is between two adjacent samples, i.e.,

number

[0151] Then the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are

number

number

[0152] Next,

number

[0153] Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block:

number

[0154] Finally, the BDOF samples for the CU are calculated by adjusting the bi-predictive samples as follows:

number

[0155] These values ​​are chosen so that the multipliers of the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters of the BDOF process is kept within 32 bits.

[0156] To derive the gradient value, we select some predicted samples I in list k (k=0, 1) outside the current CU boundary. (k)(i,j) needs to be generated. As depicted in Figure 16, BDOF in VVC uses one extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, prediction samples within the extended area (white locations) are generated by directly taking reference samples at nearby integer locations (using floor() operations on the coordinates) without interpolation, and a regular 8-tap motion compensation interpolation filter is used to generate prediction samples within the CU (gray locations). These extended sample values ​​are only used in gradient calculation. In the remaining steps of the BDOF process, if sample and gradient values ​​outside the CU boundary are needed, they are padded (i.e., repeated) from their nearest neighbors.

[0157] If the width and / or height of a CU is greater than 16 luma samples, the CU may be split into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries are treated as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process may be skipped. If the SAD between the initial L0 predicted sample and the L1 predicted sample is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W denotes the width of the sub-block and H denotes the height of the sub-block. To avoid additional complexity in SAD calculation, the SAD between the initial L0 predicted sample and the L1 predicted sample calculated in the DVMR process is reused here.

[0158] If BCW is enabled for the current block, i.e., if the BCW weight index indicates unequal weights, bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., if luma_weight_lx_flag is 1 for either of the two reference pictures, BDOF is also disabled. If the CU is coded in symmetric MVD mode or CIIP mode, BDOF is also disabled.

[0159] Decoder side motion vector refinement (DMVR) in VVC To improve the accuracy of merge mode motion vectors, VVC applies bilateral matching (BM)-based decoder-side motion vector refinement. In bi-predictive operation, refined motion vectors around the initial motion vector (MV) are searched for in reference picture lists L0 and L1. The BM method calculates the distortion between two candidate blocks in reference picture lists L0 and L1. As shown in Figure 17, the SAD between red blocks based on each MV candidate around the initial motion vector (MV) is calculated. The MV candidate with the lowest SAD becomes the refined motion vector (MV) and is used to generate the bi-predictive signal.

[0160] In VVC, the application of DMVR is limited and applies only to CUs coded with the following modes and features: - CU level merge mode with bi-predictive MV - one reference picture is in the past and the other in the future relative to the current picture - The distances from the two reference pictures to the current picture (i.e., the POC difference) are the same - Both reference pictures are short-term reference pictures - CU has more than 64 luma samples - Both CU height and CU width are 8 or more luma samples - BCW Weight Index shows equal weighting - WP is not currently enabled for the block - CIIP mode is not currently used for the block

[0161] The refined MVs derived by the DMVR process are used to generate inter prediction samples and are also used for temporal motion vector prediction for future picture encoding, while the original MVs are used in the deblocking process and are also used for spatial motion vector prediction for future CU encoding.

[0162] In the next section, additional features of DMVR are mentioned.

[0163] Search Method In DVMR, the search points are around the initial MV, and the MV offset follows the mirroring rule of the MV difference. In other words, any point checked by DMVR, indicated by a candidate MV pair (MV0, MV1), follows the following two equations:

number

number

[0164] where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage.

[0165] A full search of 25 points is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is smaller than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To mitigate the penalty due to uncertainty in DMVR refinement, it is proposed to prioritize the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value.

[0166] The integer sample search is followed by fractional sample refinement. To reduce computational complexity, the fractional sample refinement is derived by using a parametric error surface equation instead of an additional search using SAD comparison. The fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. If the integer sample search stage ends with the center with the smallest SAD in either the first or second iterative search, the fractional sample refinement is further applied.

[0167] Parametric error surface-based sub-pixel offset estimation uses the cost of the center location and the costs of the four neighboring locations from the center to form a 2D parabolic error surface equation of the form:

number

number

number

[0168] All cost values ​​are positive, and the smallest value is E(0,0), so x min and y min The value of is automatically constrained between -8 and 8, which corresponds to a half pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement delta MV.

[0169] Bilinear Interpolation and Sample Padding In VVC, the resolution of the motion vector (MV) is 1 / 16 luma sample. Fractional-position samples are interpolated using an 8-tap interpolation filter. In DMVR, search points surround the initial fractional-pixel motion vector (MV) with integer sample offsets, so samples at these fractional positions need to be interpolated in the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important advantage is that by using a bilinear filter, the DVMR does not access more reference samples than the conventional motion compensation process due to the two-sample search range. After the refined motion vector (MV) is obtained in the DMVR search process, a conventional 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the conventional MC process, samples not required for the interpolation process based on the original motion vector but required for the interpolation process based on the refined motion vector are padded from the available samples.

[0170] Geometric partitioning mode (GPM) in VVC VVC supports geometric partitioning mode for inter prediction. Geometric partitioning mode is signaled using a CU-level flag as a type of merge mode; other merge modes include normal merge mode, MMVD mode, CIIP mode, and sub-block merge mode. Geometric partitioning mode allows for a maximum possible CU size of w×h=2, except for 8×64 and 64×8. m ×2 n , m,n∈{3···6}, for each,a total of 64 partitions are supported.

[0171] When this mode is used, a CU is divided into two parts by a geometrically positioned line (Figure 18). The position of the division line is mathematically derived from the angle and offset parameters of the particular partition. Each part of the geometric partition within a CU is inter-predicted using its own motion, and only uni-prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. As with traditional bi-prediction, a uni-prediction motion constraint is applied to ensure that only two motion-compensated predictions are required for each CU. The uni-prediction motion for each partition is derived using the process described in 3.4.11.1.

[0172] If the geometric partitioning mode is used for the current CU, a geometric partition index indicating the partitioning mode (angle and offset) of the geometric partition and two merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, which specifies the syntactic binarization of the GPM merge index. After predicting each part of the geometric partition, the sample values ​​along the edges of the geometric partition are adjusted using a blending process with adaptive weights, as described in 3.4.11.2. This is the prediction signal for the entire CU, and as with other prediction modes, the transformation and quantization processes may be applied to the entire CU. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored, as described in 3.4.11.3.

[0173] Building a list of uniprediction candidates The uni-predictive candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process in 3.4.1. Let n be the index of a uni-predictive motion vector in the geometric uni-predictive candidate list. The LX motion vector (X equals the parity of n) of the nth extended merge candidate is used as the nth uni-predictive motion vector for the geometric partitioning mode. In Figure 19, these motion vectors are marked with an "x". If the corresponding LX motion vector of the nth extended merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the uni-predictive motion vector for the geometric partitioning mode.

[0174] Combined inter and intra prediction (CIIP) in VVC In VVC, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., CU width × CU height is 64 or more), and if both the CU width and CU height are less than 128 luma samples, an additional flag is signaled to indicate whether a combined inter / intra prediction (CIIP) mode is currently applied to the CU. As the name suggests, CIIP prediction combines the inter prediction signal and the intra prediction signal. The inter prediction signal P in CIIP mode inter is derived using the same inter prediction process as applied in the regular merge mode, resulting in an intra predicted signal P intra is derived according to the normal intra prediction process with planar mode. The intra prediction signal and the inter prediction signal are then combined using a weighted average, with the weight value calculated according to the coding modes of the upper and left neighboring blocks (depicted in Figure 20) as follows: - If the top neighbor is available and intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. - If the left neighbor is available and intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. - If (isIntraLeft+isIntraTop) is equal to 2, then wt is set to 3. - Otherwise, if (isIntraLeft+isIntraTop) is equal to 1, then wt is set to 2 - Otherwise, set wt to 1.

[0175] The CIIP forecast is formed as follows:

number

[0176] Intra Block Copy (IBC) in VVC Intra Block Copy (IBC) is a tool adopted in the HEVC extension of SCC. It is well known that IBC significantly improves the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block already reconstructed in the current picture. The luma block vectors of IBC-coded CUs are integer precision. Chroma block vectors are also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel and 4-pixel motion vector precision. IBC-coded CUs are treated as a third prediction mode other than intra or inter prediction modes. IBC mode is applicable to CUs whose width and height are both 64 luma samples or less.

[0177] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs an RD check for blocks whose width or height is less than or equal to 16 luma samples. In non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a block-matching-based local search is performed.

[0178] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for all locations in the current picture is based on 4x4 sub-blocks. For larger-sized current blocks, if all hash keys of all 4x4 sub-blocks match the hash key of the corresponding reference location, the hash key is determined to match the hash key of the reference block. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matched reference is calculated and the one with the smallest cost is selected.

[0179] In a block matching search, the search range is set to cover both the previous CTU and the current CTU.

[0180] At the CU level, the IBC mode is signaled by a flag, signaled as IBC AMVP mode or IBC skip / merge mode as follows:

[0181] IBC Skip / Merge Mode: The merge candidate index is used to indicate which block vector in the list from neighboring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and pairwise candidates.

[0182] IBC AMVP mode: Block vector differentials are coded in the same way as motion vector differentials. The block vector prediction method uses two candidates as predictors: one from the left neighbor and one from the above neighbor (for IBC coding). If neither neighbor is available, a default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index.

[0183] IBC reference area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predetermined area, including the area of ​​the current CTU and some areas of the left CTU. Depending on the position of the current coding CU within the current CTU, the following applies:

[0184] If the current block corresponds to the top-left 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, the current block can also refer to reference samples in the bottom-right 64x64 block of the left CTU using CPR mode.The current block can also refer to reference samples in the bottom-left 64x64 block of the left CTU and reference samples in the top-right 64x64 block of the left CTU using CPR mode.

[0185] If the current block corresponds to the top-right 64x64 block of the current CTU, and the luma position (0,64) for the current CTU has not yet been reconstructed, the current block can also refer to reference samples in the bottom-left 64x64 block and bottom-right 64x64 block of the left CTU using CPR mode, in addition to the already reconstructed samples in the current CTU. Otherwise, the current block can also refer to reference samples in the bottom-right 64x64 block of the left CTU.

[0186] If the current block corresponds to the bottom-left 64x64 block of the current CTU, and the luma position (64,0) for the current CTU has not yet been reconstructed, the current block can also refer to reference samples in the top-right 64x64 block and bottom-right 64x64 block of the left CTU using the CPR mode, in addition to the already reconstructed samples in the current CTU. Otherwise, the current block can also refer to reference samples in the bottom-right 64x64 block of the left CTU using the CPR mode.

[0187] If the current block falls within the bottom right 64x64 block of the current CTU, the CPR mode can be used to refer only to samples already reconstructed within the current CTU.

[0188] This restriction allows the hardware implementation to implement IBC mode using local on-chip memory.

[0189] Local illumination compensation (LIC) in ECM

[0190] LIC is an inter-prediction technique for modeling local illumination variation between a current block and its predicted block as a function of local illumination variation between a current block template and a reference block template. The parameters of the function can be represented by a scale α and an offset β, forming a linear equation: α*p[x]+β to compensate for illumination changes, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. Because α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for α and β, except that a LIC flag is signaled for AMVP mode to indicate the use of LIC.

[0191] The local illumination compensation proposed in JVET-O0066 is used for uni-predictive inter CUs with the following modifications: - Intra-adjacent samples may be used in the LIC parameter derivation. - LIC is disabled for blocks with less than 32 luma samples. For both non-subblock and affine modes, the LIC parameter derivation is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16x16 unit. - The samples of the reference block template are generated by using MC together with block MV without rounding to integer pixel precision.

[0192] Non-adjacent spatial candidates in ECM Non-adjacent spatial merge candidates, as in JVET-L0399, are inserted after TMVP in the regular merge candidate list. The spatial merge candidate pattern is shown in Figure 21. The distance between the non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block. No line buffer restrictions apply.

[0193] Template Matching (TM) in ECM Template matching (TM) is a decoder-side MV derivation method for refining the motion information of a current CU by finding the closest match between a template in the current picture (i.e., a block adjacent to the top and / or left of the current CU) and a block in a reference picture (i.e., the same size as the template). As shown in Figure 22, a better MV is searched for within a [-8, +8] pixel search range centered on the initial motion of the current CU. The template matching method in JVET-J0021 is used with the following modifications: the size of the search step is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in merge mode.

[0194] In AMVP mode, an MVP candidate is determined based on the template matching error to select the MVP candidate that achieves the minimum error between the current block template and the reference block template. Then, TM is performed only on this specific MVP candidate for MV refinement. TM refines this MVP candidate using an iterative diamond search, starting with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) within the [-8, +8] pixel search range. The AMVP candidate may be further refined by using a cross search of full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel MVD accuracy, depending on the AMVR mode, as specified in Table 3. This search process ensures that the MVP candidate maintains the same MVD accuracy after the TM process as exhibited in AMVR mode.

[0195] [Table 3]

[0196] In merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As Table 3 shows, TM may perform up to 1 / 8-pixel MVD accuracy or skip accuracy beyond 1 / 2-pixel MVD accuracy, depending on whether an alternative interpolation filter (used when AMVR is in half-pixel mode) is used according to the merged motion information. Furthermore, when TM mode is enabled, template matching may function as an independent process or an additional MV refinement process between the block-based bilateral matching (BM) method and the subblock-based BM method, depending on whether BM can be enabled according to the enablement condition check.

[0197] Multi-pass decoder-side motion vector refinement in ECM Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 sub-block within the coding block. In the third pass, the MVs within each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction.

[0198] First Pass – Block-Based Bilateral Matching MV Refinement In the first pass, a refined MV is derived by applying the BM to the coding block. Similar to decoder-side motion vector refinement (DMVR), in bi-predictive operation, a refined MV is searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. Based on the minimum bilateral matching cost between the two reference blocks in L0 and L1, refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MV.

[0199] BM performs a local search to derive the integer sample precision intDeltaMV. The local search applies a 3x3 square search pattern, looping over the horizontal search range [-sHor,sHor] and the vertical search range [-sVer,sVer], where the values ​​of sHor and sVer are determined by the block dimensions, and the maximum value of sHor and sVer is 8.

[0200] The bilateral matching cost is calculated as bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. When bilCost at the center point of the 3x3 search pattern has the minimum cost, the intDeltaMV local search ends. Otherwise, the current minimum-cost search point becomes the new center point of the 3x3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.

[0201] Further existing fractional sample refinement is applied to derive the final deltaMV. The refined MV after the first pass is then - MV0_pass1=MV0+deltaMV - MV1_pass1=MV1-deltaMV is derived as:

[0202] Second Pass – Subblock-Based Bilateral Matching MV Refinement In the second pass, refined MVs are derived by applying BM to a 16x16 grid of sub-blocks. For each sub-block, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.

[0203] For each sub-block, the BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a horizontal search range of [-sHor,sHor] and a vertical search range of [-sVer,sVer], where the values ​​of sHor and sVer are determined by the block dimensions, and the maximum value of sHor and sVer is 8.

[0204] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference subblocks: bilCost = satdCost * costFactor. The search area (2 * sHor + 1) * (2 * sVer + 1) is divided into up to five diamond-shaped search regions, as shown in Figure 23. Each search region is assigned a costFactor determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search area. In each region, the search points are processed in raster scan order, starting from the top-left corner of the region and moving toward the bottom-right corner. If the minimum bilCost within the current search region is less than a threshold equal to sbW * sbH, the integer-pixel full search terminates; otherwise, the integer-pixel full search proceeds to the next search region until all search points have been examined.

[0205] Further existing VVC DMVR fractional sample refinement is applied to derive the final deltaMV(sbIdx2). The refined MV in the second pass is - MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2) - MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2) is derived as:

[0206] Third Pass - Sub-block based bidirectional optical flow MV refinement In the third pass, refined MVs are derived by applying BDOF to 8x8 grid sub-blocks. For each 8x8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block in the second pass. The derived bioMv(Vx,Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.

[0207] The refined MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are - MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv - MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv is derived as:

[0208] OBMC in ECM When OBMC is applied, the top and left boundary pixels of a CU are refined using motion information from neighboring blocks along with weighted prediction, as described in JVET-L0101.

[0209] The conditions under which OBMC does not apply are as follows: - If OBMC is disabled at the SPS level - If the current block has intra or IBC mode - If the current block applies LIC - If the current Luma Block area is 32 or less

[0210] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using the motion information of neighboring sub-blocks. This is also known as sub-block-based coding tool, - Affine AMVP mode - Affine merge mode and sub-block-based temporal motion vector prediction (SbTMVP) - Subblock-based bilateral matching is valid for.

[0211] Sample-based BDOF in ECM In sample-based BDOF, instead of deriving the motion refinement (Vx, Vy) block-wise, the derivation is performed sample-by-sample.

[0212] A coding block is divided into 8x8 sub-blocks. For each sub-block, the decision to apply BDOF is made by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, a sliding 5x5 window is used for all samples in the sub-block, and the existing BDOF process is applied to all sliding windows to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predictive sample value of the center sample of the window.

[0213] Interpolation in ECM The 8-tap interpolation filter used in VVC is replaced with a 12-tap filter. The interpolation filter is derived from a sinc function, whose frequency response is cut off at the Nyquist frequency and cropped with a cosine window function. Table 4 shows the filter coefficients for all 16 phases. Figure 24 compares the frequency response of the interpolation filter with the VVC interpolation filter, all at half-pixel phase.

[0214] [Table 4]

[0215] Multi-hypothesis prediction (MHP) in ECM

[0216] In the multi-hypothesis inter-prediction mode (JVET-M0425), in addition to the conventional bi-predictive signal, one or more additional motion-compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted superposition on a sample-by-sample basis. The bi-predictive signal p bi and the first additional inter prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows: p3=(1-α)p bi +αh3

[0217] The weighting factor α is specified by the new syntax element add_hyp_weight_idx according to the following mapping: TIFF0007819300000030.tif28170

[0218] Similar to above, two or more additional prediction signals may be used, with the resulting overall prediction signal being accumulated iteratively with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1

[0219] The resulting overall predicted signal is the final p n (i.e., p with the largest index n n ) Within this EE, up to two additional prediction signals may be used (i.e., n is limited to 2).

[0220] The motion parameters of each additional prediction hypothesis are signaled either explicitly by specifying a reference index, a motion vector predictor index, and a motion vector differential, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag distinguishes between these two signaling modes.

[0221] For inter AMVP mode, MHP is only applied when unequal weights in BCW are selected in bi-predictive mode.

[0222] A combination of MHP and BDOF is possible, but BDOF is only applied to the bi-predictive signal part of the prediction signal (ie, usually the first two hypotheses).

[0223] Adaptive reordering of merge candidates with template matching in ECM (ARMC-TM)

[0224] Merge candidates are adaptively reordered by template matching (TM). This reordering method applies to the regular merge mode, the template matching (TM) merge mode, and the affine merge mode (except for SbTMVP candidates). For the TM merge mode, the merge candidates are reordered before the refinement process.

[0225] After the merge candidate list is constructed, the merge candidates are divided into several subgroups. For normal merge mode and TM merge mode, the subgroup size is set to 5. For affine merge mode, the subgroup size is set to 3. The merge candidates within each subgroup are sorted in ascending order according to their cost value based on template matching. For simplicity, merge candidates in the last subgroup (but not the first) are not sorted.

[0226] The template matching cost of a merge candidate is measured by the sum of absolute differences (SAD) between the template's samples of the current block and their corresponding reference samples. The template includes a set of reconstructed samples that neighbor the current block. The template's reference samples are located by the merge candidate's motion information.

[0227] If the merge candidate utilizes bi-prediction, the reference samples of the merge candidate's template are also generated bi-predictively, as shown in FIG.

[0228] For a subblock-based merging candidate with a subblock size equal to W×H, the top template contains several subtemplates of size W×1, and the left template contains several subtemplates of size 1×H. As shown in Figure 26, the motion information of the subblocks in the first row and first column of the current block is used to derive the reference samples for each subtemplate.

[0229] Geometric Partitioning Mode (GPM) with Merged Motion Vector Differential (MMVD) in ECM GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. First, a flag is signaled to specify whether this mode is used for the GPM CU. If this mode is used, each geometric partition of the GPM CU can further decide whether to signal MVD. If MVD is signaled for the geometric partition, after GPM merge candidates are selected, the partition motion is further refined by the signaled MVD information. All other procedures are the same as GPM.

[0230] MVD is signaled as a distance and direction pair, similar to MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (4 horizontal / vertical and 4 diagonal) involved in GPM with MMVD (GPM-MMVD). Furthermore, if pic_fpel_mmvd_enabled_flag is equal to 1, MVD is left-shifted by 2, similar to MMVD.

[0231] Geometric Partitioning Mode (GPM) using Template Matching (TM) in ECM Template matching is applied to the GPM. When GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to both geometric partitions. The motion information of each geometric partition is refined using the TM. Once a TM is chosen, a template is constructed using neighboring samples to the left, above, or left and above according to the partition angle, as shown in Table 5. The motion is then refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern in merge mode with the half-pel interpolation filter disabled.

[0232] [Table 5]

[0233] The GPM candidate list is constructed as follows:

[0234] 1. Interleaved List 0 MV candidates and List 1 MV candidates are directly derived from the normal merge candidate list. List 0 MV candidates have higher priority than List 1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.

[0235] 2. From the normal merge candidate list, the interleaved list 1 MV candidates and list 0 MV candidates are further directly derived. The list 1 MV candidates have higher priority than the list 0 MV candidates. Similarly, a pruning method using an adaptive threshold is applied to remove redundant MV candidates.

[0236] 3. The GPM candidate list is padded with zero MV candidates until it is filled.

[0237] GPM-MMVD and GPM-TM are enabled exclusively for one GPM CU. This is done by first signaling the GPM-MMVD syntax. If two GPM-MMVD control flags are both equal to false (i.e., GPM-MMVD is disabled for two GPM partitions), then a GPM-TM flag is signaled to indicate whether template matching applies to the two GPM partitions. Otherwise (if at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.

[0238] [Table 6]

[0239] In existing video codecs such as HEVC and VVC, the reference blocks of an inter-mode coded CU (hereafter referred to as inter CU) may be located partially or entirely outside the reference picture because an iterative padding process is applied to the reference picture to generate reference pixels around each reference picture.

[0240] In the VVC specification, the padding process is implemented by modifying the integer reference sample fetching process (as shown in Table 6): whenever an integer reference sample to be fetched is located outside the reference picture, the nearest integer reference sample in the reference picture is used instead.

[0241] With padded reference pictures, it is valid for inter CUs to have reference blocks that are located partially or entirely outside the reference picture, as shown in Figure 27, where bidirectional motion compensation is performed to generate inter predicted blocks for the current block. In this example, the reference blocks in list 0 are partially out of bounds (OOB), while the reference blocks in list 1 are entirely inside the reference picture.

[0242] When an inter block is bidirectionally or multidirectionally predicted, the final predictor is simply the average of two or more motion-compensated prediction blocks. When BWC is further applied, different weights are applied to the MC predictors in list 0 and list 1 to generate the final predictor, respectively. However, the out-of-bounds (OOB) portion of a motion-compensated block is less effective for prediction because the OOB portion is simply a repetitive pattern generated by boundary pixels in the reference picture. However, in existing video codecs, the less effective OOB portion of an MC block is not considered in inter prediction.

[0243] This disclosure proposes several methods for improving inter prediction by taking into account the low effectiveness of the out-of-band portion of MC blocks. Furthermore, the term "block" is used to describe the concepts of this disclosure, and "block" may be easily replaced with any specific definition used in existing codecs. For example, "block" may be a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a coding block (CB), a prediction block (PB), a transform block (TB), a coding tree block (CTB), a sub-CU, a sub-PU, or any other clearly defined term in existing codecs.

[0244] The following methods may be applied alone or in combination.

[0245] In one embodiment according to the present disclosure, when combining two or more prediction blocks generated by a motion compensation process, an additional weighting is applied to the predictor, and the additional weighting is derived based on whether the prediction sample is OOB or not. The basic concept is that when performing a weighted average of two predictor blocks, the OOB predictor sample is given less weight because it has less effect.

[0246] In one scheme, the final predictor sample of a bidirectional inter-coded block is calculated using the following equation: It is generated by weighted averaging of predictors using TIFF0007819300000033.tif10170, where: TIFF0007819300000034.tif27170 are predictor samples derived from reference pictures in list 0 and list 1, respectively, by the motion compensation process, TIFF0007819300000035.tif26170 is the weight associated with the corresponding predictor derived by the OOB condition, shift is the mean coefficient, set to 1 in the case of the average of two predictors, ο offset is the rounding offset, usually set as 1<<(shift-1).

[0247] As shown in Figure 28, the predictor sample of a block is derived by performing motion compensation from a reference picture. The motion vector of the current block is used to locate the reference block. Note that if the reference block is located at a fractional position between integer reference samples, a fractional interpolation process is further applied to derive the predicted sample. As described in Section 8.5.6.3.2, "Luma Sample Interpolation Filtering Process," of VVC, fractional interpolation first performs vertical one-dimensional interpolation, which accesses eight vertically adjacent integer samples, followed by horizontal one-dimensional interpolation, which accesses the eight vertically interpolated horizontally adjacent samples. In ECM, the 8-tap interpolation filter used in VVC is further replaced with a 12-tap filter.

[0248] Based on the motion compensation scheme described above, if at least one of the integer reference samples used to generate the predictor sample (through an interpolation process) is located outside the reference picture, the predictor sample TIFF0007819300000036.tif8170 is defined as OOB.

[0249] The weightings may be derived using the following method.

[0250] TIFF0007819300000037.tif92170TIFF0007819300000038.tif100170TIFF0007819300000039.tif100170TIFF0007819300000040.tif99170

[0251] FIG. 30 shows an example of an encoding method according to the present disclosure. In step 3001, a decoder derives a first reference picture and a second reference picture for a current coding block. In step 3002, the decoder uses a motion compensation process from the first reference picture to derive a first predictor sample based on a first motion vector associated with the first reference picture. In step 3003, the decoder uses a motion compensation process from the second reference picture to derive a second predictor sample based on a second motion vector associated with the second reference picture. In step 3004, the decoder obtains a final predictor sample within the current coding block based on at least one of the first predictor sample or the second predictor sample and an out-of-bounds (OOB) condition. This example method may also be performed by an encoder.

[0252] In another embodiment of the present disclosure, when a block is coded as a BCW mode that combines two prediction blocks generated by a motion compensation process with a BCW weighted average, an additional weighting is applied to the predictor, and the additional weighting is derived based on whether the predictor is OOB or not. In one scheme, the final predictor sample of a BCW coded block is calculated using the following equation: It is generated by weighted averaging of predictors using TIFF0007819300000041.tif12170, where w is the BCW weighting, TIFF0007819300000042.tif29170 are predictor samples derived from reference pictures in list 0 and list 1, respectively, by the motion compensation process, TIFF0007819300000043.tif26170 are the weights associated with the corresponding predictors derived by the OOB condition, shift is an averaging factor, set to 3 to normalize the BCW weights, and ο offset is the rounding offset, usually set as 1<<(shift-1).

[0253] TIFF0007819300000044.tif91170TIFF0007819300000045.tif99170TIFF0007819300000046.tif100170TIFF0007819300000047.tif100170

[0254] In another embodiment of the present disclosure, when BDOF is enabled for a block, an additional weighting is applied to the predictor, and the additional weighting is derived based on whether the predictor is OOB or not. In one scheme, the final predictor sample for a BDOF enabled block is calculated using the following equation: TIFF0007819300000048.tif12170, where B i,j is the BDOF offset of each predictor sample, TIFF0007819300000049.tif26170 are predictor samples derived from reference pictures in list 0 and list 1, respectively, by the motion compensation process, TIFF0007819300000050.tif26170 are the weights associated with the corresponding predictors derived by the OOB condition, shift is an averaging factor, set to 3 to normalize the BCW weights, and ο offset is the rounding offset, usually set as 1<<(shift-1).

[0255] TIFF0007819300000051.tif93170TIFF0007819300000052.tif99170TIFF0007819300000053.tif100170TIFF0007819300000054.tif100170

[0256] In another embodiment of the present disclosure, when BDOF is enabled for a block, an additional weighting is applied to the predictor, and the additional weighting is derived based on whether the predictor is OOB or not. In one scheme, the final predictor sample for a BDOF enabled block is calculated using the following equation: TIFF0007819300000055.tif13170, where B i,j is the BDOF offset of each predictor sample, TIFF0007819300000056.tif27170 are predictor samples derived from reference pictures in list 0 and list 1, respectively, by the motion compensation process, TIFF0007819300000057.tif27170 are the weights associated with the corresponding predictors derived by the OOB condition, shift is an averaging factor, set to 3 to normalize the BCW weights, and ο offset is the rounding offset, usually set as 1<<(shift-1).

[0257] TIFF0007819300000058.tif98170TIFF0007819300000059.tif106170TIFF0007819300000060.tif100170TIFF0007819300000061.tif103170

[0258] Based on the motion compensation scheme described above, if at least one of the integer reference samples used to generate a predictor sample (through the interpolation process) is located outside the reference picture, the predictor sample is defined as OOB.

[0259] However, in some cases, only a small number of OOB reference integer samples are used to generate a predictor sample through an interpolation process, and in this case, this predictor sample may still provide an efficient prediction (e.g., only two of eight reference integer samples are OOB). Therefore, to provide a tolerance for OOB determination, various schemes have been proposed for determining whether a predicted sample is OOB or not.

[0260] In another embodiment of the present disclosure, a predictor sample is determined to be OOB if at least one of the vertically nearest integer reference samples is OOB or at least one of the horizontally nearest integer reference samples is OOB. For example, as shown in Figure 28, the horizontally nearest integer reference samples of predictor P_0,0^Lx are samples d and e.

[0261] In another embodiment of the present disclosure, a predictor sample is determined to be OOB if at least N of the integer samples used to generate the predictor sample are OOB, where N is any integer.

[0262] In another embodiment of the present disclosure, when performing a weighted average of two or more predictor blocks, a predictor sample that uses an OOB integer reference sample to perform interpolation is given less weight. Furthermore, the weight is inversely proportional to the number of OOB integer reference samples.

[0263] In another embodiment of the present disclosure, when performing weighted averaging of two or more predictor blocks, weights are assigned to predictor samples that use OOB integer reference samples for interpolation, and the weights are signaled in the bitstream at different levels, such as the sequence level (sequence parameter set), picture level (picture parameter set), slice level (slice header), or block level.

[0264] It should be noted that the proposed MC scheme taking into account the OOB conditions is not limited to be applied to the coding method in the proposed embodiment, but can also be applied to all inter-tools such as OBMC, IBC, SMVD, DMVR, etc. as described in the previous section.

[0265] The above methods may be implemented using an apparatus including one or more circuits, including an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The apparatus may use the circuits in combination with other hardware or software components to perform the above-described methods. Each module, sub-module, unit, or sub-unit disclosed above may be implemented at least in part using one or more circuits.

[0266] 29 shows a computing environment 1610 coupled with a user interface 1650. The computing environment 1610 may be part of a data processing server. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.

[0267] The processor 1620 typically controls the overall operation of the computing environment 1610, such as operations related to display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors for executing instructions called for performing all or some of the steps in the methods described above. Additionally, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.

[0268] Memory 1630 is configured to store various types of data to support the operation of computing environment 1610. Memory 1630 may include predefined software 1632. Examples of such data include instructions for any applications or methods run on computing environment 1610, video data sets, image data, etc. Memory 1630 may be implemented using any type of volatile or non-volatile memory device, or combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0269] The I / O interface 1640 provides an interface between the processor 1620 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1640 may be coupled to an encoder and a decoder.

[0270] In one embodiment, a non-transitory computer-readable storage medium is also provided that includes, e.g., in memory 1630, a plurality of programs executable by processor 1620 in computing environment 1610 for performing the methods described above. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or datastream including encoded video information (e.g., video information including one or more syntax elements) generated by an encoder (e.g., video encoder 20 of FIG. 2) using, e.g., the encoding method described above, for use by a decoder (e.g., video decoder 30 of FIG. 3) in decoding video data. The non-transitory computer-readable storage medium may be, e.g., a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0271] In one embodiment, a computing device is also provided that includes one or more processors (e.g., processor 1620) and a non-transitory computer-readable storage medium or memory 1630 storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured, upon execution of the plurality of programs, to perform the methods described above.

[0272] In one embodiment, a computer program product is also provided that includes a plurality of programs, e.g., in memory 1630, executable by processor 1620 in computing environment 1610 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0273] In one embodiment, the computing environment 1610 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0274] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosure. Many modifications, variations and alternative implementations will be apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0275] Unless otherwise specified, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, and may be changed according to actual conditions. Furthermore, at least one of the steps of the method according to the present disclosure may be adjusted, combined, or deleted according to actual requirements.

[0276] The examples have been chosen and described to explain the principles of the disclosure, to enable those skilled in the art to understand the disclosure in various implementations, and to make best use of the underlying principles and various implementations with various modifications as suited to the particular use contemplated. Therefore, it should be understood that the scope of the disclosure is not limited to the particular examples of implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the disclosure.

Claims

1. deriving, by a decoder, a first reference picture and a second reference picture for a current coding block; deriving, by the decoder, a first predictor sample based on a first motion vector associated with the first reference picture using a motion compensation process from the first reference picture; deriving, by the decoder, a second predictor sample based on a second motion vector associated with the second reference picture using the motion compensation process from the second reference picture; obtaining, by the decoder, a final predictor sample in the current coding block based on at least one of the first predictor sample or the second predictor sample and an out-of-bounds (OOB) condition, wherein the OOB condition includes whether the first predictor sample is OOB and whether the second predictor sample is OOB; determining, by the decoder, the OOB condition by determining whether the first predictor sample and the second predictor sample are OOB; determining that at least N of the integer reference samples used to generate the predictor sample are located outside the reference picture; 11. A video decoding method comprising:

2. 2. The video decoding method of claim 1, further comprising: in response to determining that the first predictor sample is OOB and the second predictor sample is not OOB, determining the second predictor sample of the second reference picture as the final predictor sample of the current coding block.

3. 2. The video decoding method of claim 1, further comprising: in response to determining that the first predictor sample is not OOB and the second predictor sample is OOB, determining the first predictor sample of the first reference picture as the final predictor sample of the current coding block.

4. assigning a first weight to the first predictor sample and a second weight to the second predictor sample in response to determining that the first predictor sample and the second predictor sample are both OOB or that the first predictor sample and the second predictor sample are not OOB; determining the final predictor sample for the currently coded block based on a weighted average of the first predictor sample and the second predictor sample; The video decoding method of claim 1 further comprising:

5. 5. The video decoding method of claim 4, wherein the first weight assigned to the first predictor sample is equal to the second weight assigned to the second predictor sample, and the final predictor sample of the currently coded block is determined based on an average of the first predictor sample and the second predictor sample.

6. 5. The video decoding method of claim 4, further comprising: in response to determining that the current coding block is coded as a bi-predictive (BCW) mode with coding unit level weights, combining the first predictor sample and the second predictor sample to obtain the final predictor sample for the current coding block based on the BCW weighting.

7. 5. The video decoding method of claim 4, further comprising: in response to determining that bidirectional optical flow (BDOF) is enabled for the currently coded block, combining the first predictor sample and the second predictor sample to obtain the final predictor sample for the currently coded block based on the first weight assigned to the first predictor sample, the second weight assigned to the second predictor sample, and a BDOF offset.

8. setting the BDOF offset to 0 in response to determining that the first predictor sample is OOB and the second predictor sample is not OOB; setting the BDOF offset to 0 in response to determining that the first predictor sample is not OOB and the second predictor sample is OOB; The video decoding method of claim 7 further comprising:

9. determining whether the first predictor sample and the second predictor sample are OOB; 2. The video decoding method of claim 1, further comprising: determining that the predictor sample is OOB in response to determining that at least one of the vertically-closest integer reference samples of the predictor sample is OOB or that at least one of the horizontally-closest integer reference samples of the predictor sample is OOB.

10. determining whether the first predictor sample and the second predictor sample are OOB; Responsive to determining that at least one of a horizontal coordinate or a vertical coordinate of a predictor sample exceeds the boundary of the reference picture by a distance threshold. The video decoding method of claim 1 further comprising:

11. 11. The video decoding method of claim 10, wherein the distance threshold is equal to one-half sample.

12. determining, by a decoder, whether a predictor sample of a reference picture for a current coding block is out-of-bounds (OOB) based on integer reference samples used to generate said predictor sample; and in response to determining, by the decoder, that the predictor sample is OOB, assigning an additional weighting of zero to the predictor sample when combining two or more predictor samples generated by a motion compensation process to obtain a final predictor sample for the currently coded block. and in response to determining, by the decoder, that the predictor sample is not OOB, assigning a non-zero additional weighting to the predictor sample when combining the two or more predictor samples generated by the motion compensation process to obtain the final predictor sample for the currently coded block.

11. A video decoding method comprising:

13. determining whether the predictor sample is OOB based on the integer reference samples used to generate the predictor sample; determining that the predictor sample is OOB in response to determining that at least one of the integer reference samples used to generate the predictor sample is located outside a reference picture; determining that the predictor sample is not OOB in response to determining that all of the integer reference samples used to generate the predictor sample are located within the reference picture; The video decoding method of claim 12 further comprising:

14. determining, by a decoder, whether a predictor sample of a reference picture for a current coding block is out-of-bounds (OOB) based on integer reference samples used to generate said predictor sample; and in response to determining, by the decoder, that the predictor sample is OOB, assigning a first additional weight to the predictor sample when combining two or more predictor samples generated by a motion compensation process to obtain a final predictor sample for the currently coded block. and in response to determining, by the decoder, that the predictor sample is not OOB, assigning a second additional weight to the predictor sample when combining the two or more predictor samples generated by the motion compensation process to obtain the final predictor sample for the currently coded block.

11. A video decoding method comprising:

15. one or more processors; a memory configured to store instructions executable by the one or more processors, 15. Video decoding apparatus, wherein the one or more processors are configured to, upon execution of the instructions, perform the method of any of claims 1 to 14.

16. A method for receiving a bitstream, comprising: Receive the bitstream, Decoding the bitstream by a method according to any one of claims 1 to 14. How to receive the bitstream.

17. 15. A computer program for execution by a computing device having one or more processors and stored on a non-transitory computer-readable storage medium, the computer program, when executed by the one or more processors, causing the computing device to perform the method of any of claims 1 to 14.

Citation Information

Patent Citations

  • Moving image encoding device and moving image decoding device

    JP2021064819A

  • Moving image decoding apparatus and moving image encoding apparatus

    WO2020032049A1

  • Image encoding / decoding method and apparatus based on wrap-around motion compensation, and recording medium storing bitstream

    WO2021194307A1