Method and apparatus for filtered intra block copy
By improving intra-frame block copying technology and optimizing video encoding and decoding using filter coefficients and template matching scores, the problem of low efficiency in intra-frame block copying in existing technologies is solved, achieving more efficient video compression and decoding effects.
Patent Information
- Application Number
- CN202480023458.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-12
- Filing Date
- 2024-04-12
- Publication Date
- 2025-11-04
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in intra-frame block copying (FIBC) and cannot effectively utilize redundant information in video data for efficient compression.
By determining the filter shape of the reference block and template region, a set of filter coefficients is obtained. These coefficients are used to derive the predicted sample values of the current block, and the current block is reconstructed based on the predicted sample values. This improves the sorting of the merged candidate list for intra-block copy (IBC) prediction and combines intra-block copy and filtering techniques for video encoding and decoding.
It improves the efficiency of video encoding and decoding, reduces the bit rate, enhances video quality, and increases the flexibility and accuracy of the encoding and decoding process.
Smart Images

Figure CN120898418A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application is based on provisional application No. 63 / 495,677 filed on April 12, 2023 and claims priority to that provisional application. The entire contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0002] This application relates to video coding and compression. More specifically, this application relates to methods and apparatuses that improve coding efficiency of filtered intra block copy (FIBC). BACKGROUND
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video teleconferencing devices, video streaming devices, etc. An electronic device can send and receive or otherwise communicate digital video data, such as via a communication network, and / or store digital video data on a storage device. Due to limited bandwidth capacity of communication networks and limited memory resources of storage devices, video coding can be used to compress video data prior to transmitting or storing the video data to allow more efficient use of such resources. Video coding techniques can save bandwidth and storage space SUMMARY
[0004] Embodiments of the present disclosure provide methods and apparatuses that improve coding efficiency of image / video blocks to which FIBC techniques are applied.
[0005] According to an aspect of the present disclosure, a method for video decoding is provided, comprising: determining a reference block in a video frame from a bitstream for predicting a current block in the video frame; obtaining a set of filter coefficients corresponding to a filter shape based on sample values from both a template region associated with the reference block and a template region associated with the current block, wherein the template region associated with the reference block is extended based on the filter shape; deriving each of predicted sample values of the current block based on a plurality of corresponding sample values associated with the reference block using the set of filter coefficients and the filter shape; and reconstructing the current block based on the predicted sample values.
[0006] According to an aspect of the disclosure, a method for video encoding is provided, comprising: determining a reference block in a video frame for predicting a current block in the video frame; obtaining a set of filter coefficients corresponding to a filter shape based on sample values from both a template region associated with the reference block and a template region associated with the current block, wherein the template region associated with the reference block is extended based on the filter shape; deriving each of predicted sample values of the current block based on a plurality of corresponding sample values associated with the reference block using the set of filter coefficients and the filter shape; and generating a bitstream based on the predicted sample values.
[0007] According to an aspect of the disclosure, a method for video decoding is provided, comprising: obtaining a merge candidate list for intra block copy (IBC) prediction of a current block, the merge candidate list including a plurality of candidates that have been encoded with IBC; reordering the plurality of candidates in the merge candidate list based on a template matching score of each candidate in the plurality of candidates, the template matching score obtained based on a difference between sample values of a template of a candidate and corresponding reference sample values of a reference template of a reference block of the candidate, wherein the corresponding reference sample values of the reference template are obtained using filtered IBC in response to determining that the candidate is encoded with filtered IBC; and reconstructing the current block based on the reordered merge candidate list.
[0008] According to an aspect of the disclosure, a method for video encoding is provided, comprising: obtaining a merge candidate list for intra block copy (IBC) prediction of a current block, the merge candidate list including a plurality of candidates that have been encoded with IBC; reordering the plurality of candidates in the merge candidate list based on a template matching score of each candidate in the plurality of candidates, the template matching score obtained based on a difference between sample values of a template of a candidate and corresponding reference sample values of a reference template of a reference block of the candidate, wherein the corresponding reference sample values of the reference template are obtained using filtered IBC in response to determining that the candidate is encoded with filtered IBC; and encoding the current block based on the reordered merge candidate list to generate a bitstream.
[0009] According to an aspect of the disclosure, a method for video decoding is provided, comprising: determining a block vector for a chroma block from a bitstream based on a luma block that is encoded with intra block copy (IBC) and is associated with the chroma block; obtaining an IBC prediction of the chroma block based on the block vector; in response to determining that a filtered IBC prediction is used for the chroma block, filtering the IBC prediction of the chroma block to obtain the filtered IBC prediction; and reconstructing the chroma block based on the filtered IBC prediction.
[0010] According to an aspect of the disclosure, a method for video encoding is provided, comprising: determining a block vector for a chroma block based on a luma block coded with intra block copy (IBC) and associated with the chroma block; obtaining an IBC prediction of the chroma block based on the block vector; filtering the IBC prediction of the chroma block; and generating a bitstream based on the filtered IBC prediction.
[0011] According to an aspect of the disclosure, an apparatus is provided, comprising: one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of the disclosure.
[0012] According to an aspect of the disclosure, a computer program product is provided, storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of the disclosure.
[0013] According to an aspect of the disclosure, a computer-readable storage medium storing instructions that, when executed by a computing device having one or more processors, cause the one or more processors to perform the decoding method of the disclosure and store a bitstream to be decoded by the decoding method of the disclosure, or perform the encoding method of the disclosure and store a bitstream generated by the encoding method of the disclosure.
[0014] According to an aspect of the disclosure, a computer-readable medium storing a bitstream is provided, wherein the bitstream is to be decoded by performing the operations of the method of the disclosure, or the bitstream is obtained by performing the operations of the method of the disclosure.
[0015] According to an aspect of the disclosure, a method for receiving a bitstream to be decoded by the decoding method of the disclosure is provided.
[0016] According to an aspect of the disclosure, a method for transmitting a bitstream generated by the encoding method of the disclosure is provided.
[0017] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0019] Figure 1 is a block diagram showing an exemplary system for encoding and decoding a video block according to some embodiments of the present disclosure.
[0020] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.
[0021] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.
[0022] Figure 4A 、 Figure 4B 、 Figure 4C 、 Figure 4D and Figure 4E is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.
[0023] Figure 5 is a diagram illustrating the position of spatial candidates.
[0024] Figure 6 is a diagram illustrating candidate pairs considering the redundancy check for spatial candidates.
[0025] Figure 7 is a diagram illustrating scaling of motion vectors for temporal candidates.
[0026] Figure 8 is a diagram illustrating candidate positions for temporal candidates.
[0027] Figure 9 is a diagram illustrating the merge mode with motion vector difference (MMVD) search points.
[0028] Figure 10 is a diagram illustrating uni-prediction motion vector selection for geometric partition mode (GPM).
[0029] Figure 11 is a diagram illustrating top and left neighboring blocks used in CIIP weight derivation.
[0030] Figure 12 is a diagram illustrating the current CTU processing order and its available reference samples in the current CTU and the left CTU.
[0031] Figure 13 is a diagram illustrating padding candidates for replacing zero vectors in the IBC list.
[0032] Figure 14 is a diagram illustrating the reference region of IBC when CTU (m, n) is coded.
[0033] Figure 15 is a diagram illustrating the IBC reference region for camera captured content.
[0034] Figure 16A to Figure 16BA partitioning method for angular modes is shown.
[0035] Figure 17A to Figure 17C Available IPM candidates are shown, and Figure 17D An example of GPM with intra prediction and intra prediction is shown.
[0036] Figure 18 An edge on a template is shown.
[0037] Figure 19 An intra template matching search region used is shown.
[0038] Figure 20 A template for template matching based OBMC is shown.
[0039] Figure 21 A template and reference samples of the template in a reference picture are shown.
[0040] Figure 22 A template and reference samples of the template for a block with subblock motion using motion information of subblocks of the current block are shown.
[0041] Figure 23 A luma block used to derive a direct block vector is shown.
[0042] Figure 24 A partitioning method and corresponding weights for intra coded blocks for angular modes and planar modes are shown.
[0043] Figure 25 A diagram showing filter shapes and training regions for a reference block is shown.
[0044] Figure 26 A diagram showing spatial terms corresponding to neighboring luma samples is shown.
[0045] Figure 27 A diagram showing examples of different shapes / numbers of filter taps is shown.
[0046] Figure 28 A diagram showing possible locations of a candidate region is shown.
[0047] Figure 29 A diagram showing possible locations of a candidate is shown.
[0048] Figure 30 A workflow of a method for decoding video data according to one or more aspects of the disclosure is shown.
[0049] Figure 31 A workflow of a method for encoding video data according to one or more aspects of the disclosure is shown.
[0050] Figure 32A workflow of a method for decoding video data according to one or more aspects of the disclosure is shown.
[0051] Figure 33 A workflow of a method for encoding video data according to one or more aspects of the disclosure is shown.
[0052] Figure 34 A diagram showing filter shape and training region for a reference block is shown.
[0053] Figure 35 A workflow of a method for video decoding according to one or more aspects of the disclosure is shown.
[0054] Figure 36 A workflow of a method for video encoding according to one or more aspects of the disclosure is shown.
[0055] Figure 37 A workflow of a method for video decoding according to one or more aspects of the disclosure is shown.
[0056] Figure 38 A workflow of a method for video encoding according to one or more aspects of the disclosure is shown.
[0057] Figure 39 A workflow of a method for video decoding according to one or more aspects of the disclosure is shown.
[0058] Figure 40 A workflow of a method for video encoding according to one or more aspects of the disclosure is shown.
[0059] Figure 41 A diagram showing a computing environment coupled with a user interface according to some embodiments of the disclosure is shown. DETAILED DESCRIPTION
[0060] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description of embodiments of the present subject matter, numerous specific details are set forth in order to assist in understanding the subject matter presented herein. However, the subject matter of the present disclosure can be practiced without the specific details, and can involve other alternatives, and equivalent(s) and / or modifications, and can be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices having digital video capabilities.
[0061] It should be noted that the terms "first", "second", and the like used in the description, claims, and drawings of the disclosure are used to differentiate between objects and are not used to describe any particular order or sequence. It should be understood that data used in this manner can be interchanged, under appropriate conditions, such that embodiments of the disclosure described herein can be implemented in orders other than those shown in the figures or described in the disclosure.
[0062] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, in accordance with some embodiments of the disclosure. As shown in Figure 1 The system 10 includes a source device 12 that generates and encodes video data for later decoding by a destination device 14, as shown in
[0063] In some implementations, the destination device 14 can receive encoded video data to be decoded via a link 16. The link 16 can comprise any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 can comprise a communication medium to enable the source device 12 to transmit encoded video data directly to the destination device 14 in real-time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from the source device 12 to the destination device 14.
[0064] In some other implementations, encoded video data can be transmitted from output interface 22 to storage device 32. Subsequently, encoded video data in storage device 32 can be accessed by destination device 14 via input interface 28. Storage device 32 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Disc Read-Only Memories (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In a further example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Destination device 14 can access stored video data from storage device 32 via streaming or download. The file server can be any type of computer capable of storing encoded video data and transmitting that encoded video data to destination device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 14 can access the encoded video data from the file server through any standard data connection, including a wireless channel (e.g., a Bluetooth®wireless connection), a wired connection (e.g., Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both, that is appropriate for accessing encoded video data stored on a file server. The transmission of encoded video data from storage device 32 can be a streaming transmission, a download transmission, or a combination of both.
[0065] As shown in FIG. 1, source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 can include a source such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface to receive video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As one example, if video source 18 is a video camera of a security surveillance system, source device 12 and destination device 14 can form a camera phone or video phone. However, the implementations described in the present application can be applied to video coding in general, and can have application to wireless and / or wired applications. Figure 1
[0066] The captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored onto storage device 32 for later access by destination device 14 or other devices, for decoding and / or playback. Output interface 22 can further include a modem and / or a transmitter.
[0067] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 can include a receiver and / or a modem and receive the encoded video data over link 16. The encoded video data transmitted over link 16 or provided on storage device 32 can include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements can be included with the encoded video data within the communications medium being sent, on storage media, or on a file server.
[0068] In some implementations, target device 14 can include a display device 34, which can be an integrated display device and an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to a user and can include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0069] Video encoder 20 and video decoder 30 can operate according to a proprietary standard or industry standard, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions to such standards. It should be understood that the application is not limited to a specific video encoding / decoding standard and can apply to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0070] Video encoder 20 and video decoder 30 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. When implemented partially in software, an electronic device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined video encoder / decoder (CODEC) in a respective device.
[0071] In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or the components of video encoder 20 described below in reference to FIG. 2 Figure 2 and / or at least a portion of the components of destination device 14 (e.g., input interface 28, video decoder 30, or the components of video decoder 30 described below in reference to FIG. 3 Figure 3 described as included in video decoder 30, and display device 34) can operate in a cloud computing services network, which can provide software, platforms, and / or infrastructure, such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS). In some implementations, one or more components of source device 12 and / or destination device 14 that are not included in the cloud computing services network can be provided in one or more client devices, and the one or more client devices can communicate with server computers in the cloud computing services network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power-line communication (PLC) network). In an embodiment, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers implemented by at least a portion of the components of source device 12 and / or at least a portion of the components of destination device 14 in the cloud computing services network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing services network can be a private cloud, a public cloud, or a hybrid cloud. The terms such as “cloud,” “cloud computing,” “cloud-based,” and the like can be used interchangeably as appropriate herein without departing from the scope of the disclosure. It should be understood that the disclosure is not limited to implementation in the above-described cloud computing services network. Rather, the disclosure can also be implemented in any other type of computing environment currently known or developed in the future.
[0072] Figure 2 A block diagram of an example video encoder 20 according to some implementations described in this application is shown. Video encoder 20 can perform intra-predictive coding and inter-predictive coding on video blocks within a video frame. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within neighboring video frames or pictures of a video sequence. It should be noted that the term “frame” can be used as a synonym for the term “image” or “picture” in the field of video coding.
[0073] As Figure 2As shown in FIG. 1, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, summer 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block copy (IBC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58, inverse transform processing unit 60, and summer 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between summer 62 and DPB 64 to filter block boundaries to remove artifacts from reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of summer 62. It is noted that for CCSAO techniques, the present application is not limited to the embodiments described herein, but rather, the present application can be applied to the case where an offset is selected for a luma component, a Cb chroma component, and Cr chroma component from any other one of the luma component, the Cb chroma component, and the Cr chroma component to modify said any component based on the selected offset. The first component referred to herein can be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component referred to herein can be any other one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component referred to herein can be the remaining one of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filters can be omitted, and the decoded video blocks can be provided directly from summer 62 to DPB 64. Video encoder 20 can take the form of a fixed or programmable hardware encoder, or can be dispersed in one or more of the illustrated fixed or programmable hardware encoders.
[0074] Video data memory 40 can store video data to be encoded by the components of video encoder 20. The video data in video data memory 40 can be obtained, for example, from video source 18, as shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20. Video data memory 40 and DPB 64 can be formed by a variety of memory devices, such as a Figure 1 suggested in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20. Video data memory 40 and DPB 64 can be formed by a variety of memory devices, such as a
[0075] As Figure 2As shown in FIG. 1, after receiving video data, partitioning unit 45 within prediction processing unit 41 partitions the video data into video blocks. This partitioning can also include partitioning of a video frame into slices, tiles (e.g., a collection of video blocks), or other larger coding units (CUs) according to a predefined splitting structure (e.g., a quadtree (QT) structure) associated with the video data. A video frame is or can be considered as a two-dimensional array or matrix of sample values. A sample in the array can also be referred to as a pixel or pel. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame can be divided into a plurality of video blocks by using QT partitioning. A video block is or can be considered as a two-dimensional array or matrix of sample values as well, but its dimensions are smaller than those of the video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block can be further partitioned into one or more block partitions or sub-blocks (which can form blocks again) by, for example, iteratively using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or can be or correspond to a respective block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.
[0076] Prediction processing unit 41 can select one of a plurality of possible predictive encoding modes, e.g., one of a plurality of intra- or inter-predictive encoding modes, for the current video block based on error results (e.g., rate and distortion levels). Prediction processing unit 41 can provide the resulting intra- or inter-predicted block to summer 50 to generate a residual block, and to summer 62 to reconstruct the encoded block for use as part of a reference frame at a later time. Prediction processing unit 41 also provides at least one of the syntax elements (e.g., motion vectors, intra- or inter-mode indicators, partitioning information, and other such syntax information) to entropy encoding unit 56.
[0077] To select a suitable intra-prediction coding mode for the current video block, intra-prediction processing unit 46 can perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block being coded to provide a spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide a temporal prediction. Video encoder 20 can perform multiple encoding passes, e.g., to select a suitable coding mode for each block of video data.
[0078] In some implementations, motion estimation unit 42 determines an inter-prediction mode for a current video frame by generating a motion vector that indicates a displacement of a video block within the current video frame relative to a prediction block within a reference video frame according to a predetermined pattern within a sequence of video frames. Motion estimation performed by motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. For example, the motion vector can indicate a displacement of a video block within a current video frame or picture relative to a prediction block within a reference frame relative to a current block being coded within the current frame. The predetermined pattern can designate video frames in the sequence as P-frames or B-frames. Intra-BC unit 48 can determine vectors, e.g., block vectors, for intra-BC coding in a manner similar to the determination of motion vectors by motion estimation unit 42 for inter-prediction, or can utilize motion estimation unit 41 to determine the block vectors.
[0079] In terms of pixel differences, a prediction block for a video block can be or can correspond to a block or reference block of a reference frame that is deemed to closely match the video block to be coded, and the pixel differences can be determined by a sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some implementations, video encoder 20 can calculate values for sub-integer pixel positions of reference frames stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference frame. Thus, motion estimation unit 42 can perform a motion search relative to full-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.
[0080] Motion estimation unit 42 calculates motion vector information for a video block in an inter-prediction coded frame by comparing a location of the video block to a location of a prediction block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of the first and second reference frame lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the determined motion vector information to motion compensation unit 44, which is then sent to entropy encoding unit 56.
[0081] Motion compensation performed by motion compensation unit 44 can involve fetching or generating a prediction block based on motion vector information determined by motion estimation unit 42. After receiving motion vector information for a current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block can include luma component differences or chroma component differences or both. Motion compensation unit 44 can also generate syntax elements associated with a video block of a video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements can include, for example, syntax elements defining motion vectors used to identify a prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It is noted that motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are illustrated separately for conceptual purposes.
[0082] In some implementations, intra BC unit 48 can generate vectors and extract prediction blocks in a manner similar to that described above in connection with motion estimation unit 42 and motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are referred to as block vectors rather than motion vectors. In particular, intra BC unit 48 can determine an intra prediction mode to use to encode a current block. In some examples, intra BC unit 48 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and test their performance through rate-distortion analysis. Next, intra BC unit 48 can select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, intra BC unit 48 can use rate-distortion analysis for the various tested intra prediction modes to compute rate-distortion values, and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines an amount of distortion (or error) between an encoded block and the original, unencoded block used to produce the encoded block, and a bit rate (i.e., number of bits) used to produce the encoded block. Intra BC unit 48 can compute a ratio from the distortion and rate of various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion values for that block.
[0083] In other examples, intra BC unit 48 can use motion estimation unit 42 and motion compensation unit 44, in whole or in part, to perform such functions for intra BC prediction according to the implementations described herein. In either case, for intra block copy, the prediction block can be a block that is considered to closely match the block to be coded in terms of pixel differences, which can be determined by SAD, SSD, or other difference metrics, and the identification of the prediction block can include the calculation of values for sub-integer pixel positions.
[0084] Whether the prediction block is from the same frame according to intra prediction or a different frame according to inter prediction, video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being coded, thereby forming pixel difference values. The pixel difference values forming the residual video block can include both luma component differences and chroma component differences.
[0085] Intra prediction processing unit 46 can perform intra prediction on the current video block as an alternative to inter prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra block copy prediction performed by intra BC unit 48, as described above. In particular, intra prediction processing unit 46 can determine an intra prediction mode to use to encode the current block. To do so, intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and intra prediction processing unit 46 (or, in some examples, a mode selection unit) can select an appropriate intra prediction mode to use from the tested intra prediction modes. Intra prediction processing unit 46 can provide information indicative of the selected intra prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 can encode the information indicative of the selected intra prediction mode in the bitstream.
[0086] After prediction processing unit 41 determines the prediction block for the current video block via inter prediction or intra prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0087] Transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, quantization unit 54 can then perform a scan of the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 can perform the scan.
[0088] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into the video bitstream using, for example, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy encoding methodology or technique. The encoded bitstream can then be transmitted to video decoder 30 as shown in FIG. 1, or archived in storage device 32 for later transmission to or retrieval by video decoder 30 as shown in FIG. 1. Entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements of the current video frame being encoded. Figure 1 Figure 1
[0089] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual video blocks in the pixel domain to produce reference blocks for prediction of other video blocks. As mentioned above, motion compensation unit 44 can generate motion compensated predicted blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 can also apply one or more interpolation filters to a predicted block to calculate sub-integer pixel values for motion estimation.
[0090]
[0091] Figure 3 is a block diagram illustrating an exemplary video decoder 30 according to some implementations of the present application. Video decoder 30 includes video data memory 79, entropy decoding unit 80, prediction processing unit 81, inverse quantization unit 86, inverse transform processing unit 88, summer 90, and DPB 92. Prediction processing unit 81 further includes motion compensation unit 82, intra prediction unit 84, and intra BC unit 85. Video decoder 30 can perform a decoding process generally reciprocal to the encoding process described above with respect to video encoder 20. For example, motion compensation unit 82 can generate prediction data based on motion vectors received from entropy decoding unit 80, while intra prediction unit 84 can generate prediction data based on intra prediction mode indicators received from entropy decoding unit 80. Figure 2
[0092] In some examples, the components of video decoder 30 can be tasked to perform the implementations of the present application. Moreover, in some examples, the implementations of the present disclosure can be divided among one or more of the components of video decoder 30. For example, the intra BC unit 85 can perform the implementations of the present application, alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 can not include intra BC unit 85, and the functionality of intra BC unit 85 can be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0093] Video data memory 79 can store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source, such as a camera, via wired or wireless network communication of video data, or by accessing physical data storage media (e.g., a flash drive or hard disk). Video data memory 79 can include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra or inter predictive coding modes). Video data memory 79 and DPB 92 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For Figure 3 Video data memory 79 and DPB 92 are depicted as two distinct components of video decoder 30 in FIG. 3 for illustrative purposes. One of skill in the art will understand that video data memory 79 and DPB 92 can be provided by the same memory device or by separate memory devices. In some examples, video data memory 79 can be on-chip with other components of video decoder 30, or off-chip relative to those components.
[0094] During the decoding process, video decoder 30 receives an encoded video bitstream that represents encoded video frames and associated syntax elements. Video decoder 30 can receive the syntax elements at the video frame level and / or video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators, among other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra prediction mode indicators, among other syntax elements, to prediction processing unit 81.
[0095] When a video frame is encoded as an intra-predicted (I) frame or for intra-coded prediction blocks in other types of frames, intra-prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-prediction mode signaled and reference data from previously decoded blocks of the current frame.
[0096] When a video frame is encoded as an inter-predicted (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 produces one or more prediction blocks for a video block of the current video frame based on motion vector information and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks can be produced from a reference frame within one of the reference frame lists. Video decoder 30 can construct the reference frame lists, i.e., List 0 and List 1, using default construction techniques based on reference frames stored in DPB 92.
[0097] In some examples, when a video block is encoded according to the intra BC modes described herein, intra BC unit 85 produces a prediction block for the current video block based on block vector information and other syntax elements received from entropy decoding unit 80. The prediction block can be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.
[0098] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the vector information and other syntax elements, and then use the prediction information to produce a prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode used to encode the video block of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, motion vectors for each inter-predicted encoded video block of the frame, inter-prediction status for each inter-predicted encoded video block of the frame, and other information used to decode the video block in the current video frame.
[0099] Similarly, intra BC unit 85 can use some of the received syntax elements, e.g., flags, to determine that the current video block is predicted using the intra BC mode, construction information for which video blocks of the frame are within the reconstructed region and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information used to decode the video block in the current video frame.
[0100] Motion compensation unit 82 can also perform interpolation using interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks. In this case, motion compensation unit 82 can determine the interpolation filters used by video encoder 20 from the syntax elements received and use these interpolation filters when generating the prediction blocks.
[0101] Inverse quantization unit 86 inverse quantizes quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 using the same quantization parameter calculated by video encoder 20 for each video block in the video frame. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual blocks in the pixel domain.
[0102] After motion compensation unit 82 or intra BC unit 85 generates the prediction block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs the decoded video block for the current video block by summing the residual block from inverse transform processing unit 88 and the corresponding prediction block generated by motion compensation unit 82 and intra BC unit 85. In-loop filter 91, e.g., a deblock filter, a SAO filter, a CCSAO filter, and / or an ALF, can be located between adder 90 and DPB 92 to further process the decoded video block. In some examples, in-loop filter 91 can be omitted, and the decoded video block can be provided directly by adder 90 to DPB 92. The decoded video blocks in a given frame are then stored in DPB 92, which stores reference frames for subsequent motion compensation of next video blocks. DPB 92, or a separate memory device from DPB 92, can also store decoded video for later presentation on a display device (e.g., display device 34 of FIG. 1). Figure 1
[0103] In a typical video coding process, a video sequence generally includes an ordered set of frames or pictures. Each frame can include three arrays of samples, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other cases, a frame can be monochrome and thus include only one two-dimensional array of luma samples.
[0104] As Figure 4A As shown in FIG. 1, video encoder 20 (or more specifically, partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame can include an integer number of CTUs ordered consecutively in the raster scan order from left to right and from top to bottom. Each CTU is the largest logical coding unit and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set such that all CTUs in a video sequence have the same size of one of 128x128, 64x64, 32x32, and 16x16. It should be noted, however, that the present application is not necessarily limited to a particular size. As Figure 4B As shown in FIG. 3, each CTU can include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to code the samples of the coding tree blocks. The syntax elements describe properties of different types of units of the coded pixel blocks and how the video sequence, including inter-prediction or intra-prediction, intra-prediction modes, motion vectors, and / or other parameters, can be reconstructed at video decoder 30. In monochrome pictures or pictures having three separate color planes, a CTU can include a single coding tree block and syntax elements used to code the samples of the coding tree block. A coding tree block can be an NxN block of samples.
[0105] To achieve better performance, video encoder 20 can recursively perform tree partitioning, e.g., binary tree partitioning, ternary tree partitioning, quad-tree partitioning, or a combination thereof, on the coding tree blocks of a CTU and divide the CTU into smaller CUs. As Figure 4C As depicted in FIG. 4, a 64x64 CTU 400 is first divided into four smaller CUs, each having a block size of 32x32. Of the four smaller CUs, CUs 410 and 420 are each divided into four CUs having a block size of 16x16. The two 16x16 CUs 430 and 440 are each further divided into four CUs having a block size of 8x8. Figure 4D A quad-tree data structure showing the final result of the partitioning process of CTU 400 is depicted in FIG. 5. Each leaf node of the quad-tree corresponds to one CU having a respective size ranging from 32x32 to 8x8. Similar to the binary tree data structure depicted in FIG. 3, the root node of the quad-tree corresponds to the 64x64 CTU 400. Figure 4C As depicted in FIG. 4, a 64x64 CTU 400 is first divided into four smaller CUs, each having a block size of 32x32. Of the four smaller CUs, CUs 410 and 420 are each divided into four CUs having a block size of 16x16. The two 16x16 CUs 430 and 440 are each further divided into four CUs having a block size of 8x8. Figure 4B As depicted in FIG. 3, each CTU can include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to code the samples of the coding tree blocks. The syntax elements describe properties of different types of units of the coded pixel blocks and how the video sequence, including inter-prediction or intra-prediction, intra-prediction modes, motion vectors, and / or other parameters, can be reconstructed at video decoder 30. In monochrome pictures or pictures having three separate color planes, a CTU can include a single coding tree block and syntax elements used to code the samples of the coding tree block. A coding tree block can be an NxN block of samples. Figure 4C and Figure 4DThe quad-tree partitioning depicted in the middle is for illustrative purposes only, and one CTU can be split into multiple CUs based on quad-tree partitioning / triple-tree partitioning / binary-tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to a quad-tree structure, and each quad-tree leaf CU can be further partitioned according to a binary and / or triple-tree structure. As shown in Figure 4E The CB with width W and height H has five possible partition types, i.e., quad partition, horizontal binary partition, vertical binary partition, horizontal ternary partition, and vertical ternary partition.
[0106] In some implementations, video encoder 20 can further partition the coding blocks of a CU into one or more PBs of MxN. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. A PU of a CU can include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In a monochrome picture or a picture having three separate color planes, a PU can include a single PB and syntax structures used to predict the PB. Video encoder 20 can generate predicted luma, Cb, and Cr blocks for the luma, Cb, and Cr PBs of each PU of a CU.
[0107] Video encoder 20 can use intra prediction or inter prediction to generate the predicted blocks of a PU. If video encoder 20 uses intra prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0108] After video encoder 20 generates the predicted luma, Cb, and Cr blocks for one or more PUs of a CU, video encoder 20 can generate a luma residual block for the CU by subtracting the predictive luma blocks of the CU from their original luma coding blocks, such that each sample in the luma residual block for the CU indicates a difference between a luma sample in one of the predictive luma blocks for the CU and a corresponding sample in the original luma coding block for the CU. Similarly, video encoder 20 can generate Cb and Cr residual blocks for the CU, such that each sample in the Cb residual block for the CU indicates a difference between a Cb sample in one of the predictive Cb blocks for the CU and a corresponding sample in the original Cb coding block for the CU, and each sample in the Cr residual block for the CU can indicate a difference between a Cr sample in one of the predictive Cr blocks for the CU and a corresponding sample in the original Cr coding block for the CU.
[0109] Furthermore, as Figure 4CAs shown in FIG. 1, video encoder 20 can partition a picture into one or more slices, each including one or more CTUs. Video encoder 20 can partition a slice into one or more tiles, each including one or more CTUs. Video encoder 20 can partition a slice into one or more warps, each including one or more CTUs. Video encoder 20 can partition a CTU into one or more CUs. A CU including an intra-coding node in a quadtree can include one or more PUs and one or more Ts. A CU including an inter-coding node in a quadtree can include one or more PUs. Each PU can include one or more picture blocks. Each TU can include one or more transform blocks.
[0110] Video encoder 20 can apply one or more transforms to the luma transform block of a TU to produce a luma coefficient block for the TU. A coefficient block can be a two-dimensional array of transform coefficients. A transform coefficient can be a scalar. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to produce a Cr coefficient block for the TU.
[0111] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. After video encoder 20 quantizes a coefficient block, video encoder 20 can entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream that includes a sequence of bits that form a representation of coded frames and associated data, which is retained in storage device 32 or transmitted to destination device 14.
[0112] After receiving the bitstream generated by video encoder 20, video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 can reconstruct the frames of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 can perform inverse transforms on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs the coding block of the current CU by adding the samples of the prediction block of the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks of each CU of a frame, video decoder 30 can reconstruct the frame.
[0113] As described above, video coding primarily achieves video compression using two modes, namely, intra prediction (or intra prediction) and inter prediction (or inter prediction). It should be noted that IBC can be considered as a third mode of intra prediction or. Between the two modes, inter prediction contributes more coding efficiency than intra prediction because motion vectors are used for predicting the current video block from the reference video block.
[0114] But as video data capturing technology for preserving details in video data and finer video block sizes are continuously improved, the amount of data required to represent the motion vectors of the current frame also significantly increases. One way to overcome this challenge is to benefit from the fact that not only the group of neighboring CUs in both spatial and temporal domains have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, it is possible to use the motion information of the spatially neighboring CUs and / or the temporally collocated CUs as an approximation of the motion information (e.g., motion vectors) of the current CU by exploring the spatial and temporal correlation of the current CU, which is also referred to as the “motion vector predictor (MVP)” of the current CU.
[0115] Instead of encoding the actual motion vector of the current CU determined by motion estimation unit 42 into the video bitstream, as described above in connection with Figure 2 the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to generate the motion vector difference (MVD) of the current CU. By doing so, the motion vectors determined by motion estimation unit 42 for each CU of the frame do not need to be encoded into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be significantly reduced.
[0116] Similar to the process of selecting a prediction block in a reference frame during inter prediction of a coding block, both video encoder 20 and video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also referred to as a "merge list") for a current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally collocated CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from video encoder 20 to video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor for encoding and decoding the current CU.
[0117] Generally, the basic inter prediction scheme applied in VVC remains almost the same as in HEVC, except that several prediction tools are further extended, added and / or improved, such as extended merge prediction, MMVD and GPM.
[0118] Extended merge prediction
[0119] With ever-improving video data capturing techniques for preserving details in video data and finer video block sizes, the amount of data needed to represent the motion vectors of a current picture also increases significantly. One way to overcome this challenge is to use the motion information (e.g., motion vectors) of spatially neighboring CUs, temporally collocated CUs, etc. of the current CU as an approximation (e.g., prediction) of the motion information of the current CU, which is also referred to as a "motion vector predictor (MVP)" of the current CU. As used throughout this disclosure, "motion vector" includes not only the motion vectors between CUs from different frames (e.g., between temporally collocated CUs in inter prediction), but also the block vectors between CUs in the same frame (e.g., between spatially neighboring CUs in intra prediction).
[0120] As with the process of selecting a prediction block in a reference picture during inter prediction of a coding block, both video encoder 20 and video decoder 30 need to employ a set of rules to construct a MVP candidate list for a current CU, and then select one MVP candidate from the MVP candidate list as the MVP for the current CU. By doing so, the MVP candidate list itself does not need to be transmitted between video encoder 20 and video decoder 30, and the index of the selected MVP candidate from the MVP candidate list is sufficient for video encoder 20 and video decoder 30 to use the same MVP candidate selected from the MVP candidate list to encode and decode the current CU.
[0121] In VVC, the MVP candidate list is constructed by including the following five types of MVPs in order:
[0122] —The spatial MVP of spatially adjacent CUs (i.e., spatial candidates);
[0123] —The temporal MVP of the CU that is temporally co-located (i.e., the temporal candidate);
[0124] —History-based MVP (HMVP) from a first-in-first-out (FIFO) table;
[0125] —Paired average MVP; and
[0126] —Zero MVP.
[0127] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded and decoded in merge mode, a truncated unary binary representation is used to encode the index of the best MVP candidate. The first binary number of the index is encoded using the context, and the other binary numbers used for the index are bypassed.
[0128] The export process for each type of MVP is shown below. For example, in HEVC, VVC also supports exporting the MVP candidate list of all CUs in parallel within a certain size region.
[0129] Deriving MVP from Spatial Candidates
[0130] In VVC, from spatial candidates (e.g., Figure 5 The MVP derived from the CU adjacent to the current CU 101 is the same as that in HEVC, except that the positions of the first two space candidates are swapped. Figure 5 Up to four spatial candidates are selected from the spatial candidates at the locations depicted: top position B0, left position A0, upper right position B1, lower left position A1, and upper left position B2. Derivation is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because said one or more CUs belong to other slices or tiles) or if intra-frame encoding / decoding is performed.
[0131] After adding the CU at position B0 as a candidate to the merged candidate list, the remaining candidates are added to the merged candidate list for redundancy checking. This ensures that candidates with the same motion information are excluded from the merged candidate list, thus improving encoding and decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only those using the same motion information are considered. Figure 6 The arrows in the diagram link the pairs of candidates, and a candidate is added to the merged candidate list only if the motion information of the candidate in the pair used for redundancy check is different from that of the candidate to be added. The spatial MVP derived from the candidates in the merged candidate list is added to the MVP candidate list.
[0132] Deriving MVP from temporal candidate
[0133] During derivation of MVP from temporal candidate, only one temporal candidate is added to the merge candidate list. Specifically, when deriving MVP from the temporal candidate, the collocated CU (e.g., col_CU 301 in Figure 7 ) is used as a temporal candidate of the collocated picture (e.g., col_pic 302) belonging to the current CU (e.g., curr_CU 303) to derive a scaled motion vector, and the scaled motion vector is added to the MVP candidate list as a temporal MVP candidate. The reference picture list and the reference picture index for deriving the collocated CU are explicitly signaled in the slice header. As shown in Figure 7 , the scaled motion vector is obtained (i.e., scaled) from the motion vector of the collocated CU using the picture order count (POC) distance (i.e., tb and td), where tb is defined as the POC difference between the reference picture (e.g., curr_ref 305 in Figure 7 ) of the current picture (e.g., curr_pic 304) and the current picture, and td is defined as the POC difference between the reference picture (e.g., col_ref 306 in Figure 7 ) of the collocated picture and the collocated picture. The reference picture index of the temporal candidate is set to be equal to zero. Figure 7 Figure 7 As depicted in Figure 7 , the position of the temporal candidate (i.e., collocated CU) in the current CU 401 is selected between positions C0 and C1. If the CU at position C0 in the collocated picture is unavailable, intra coded, or outside the current CTU row, the CU at position C1 is used as the collocated CU for deriving the temporal MVP candidate. Otherwise, the CU at position C0 is used as the collocated CU for deriving the temporal MVP candidate.
[0134] As depicted in Figure 8 , the position of the temporal candidate (i.e., collocated CU) in the current CU 401 is selected between positions C0 and C1. If the CU at position C0 in the collocated picture is unavailable, intra coded, or outside the current CTU row, the CU at position C1 is used as the collocated CU for deriving the temporal MVP candidate. Otherwise, the CU at position C0 is used as the collocated CU for deriving the temporal MVP candidate.
[0135] Derivation of HMVP candidate
[0136] After spatial and temporal MVPs, the HMVP candidate is added to the MVP candidate list. The motion information of previously coded blocks is stored in the HMVP table and used as the MVP of the current CU. The table with multiple HMVP candidates is kept during the encoding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is a non-subblock inter coded CU, the associated motion information is added to the last entry of the HMVP table as a new HMVP candidate.
[0137] The size of the HMVP table is set to 6. When a new HMVP candidate is inserted into the HMVP table, a constrained FIFO rule is utilized, where a redundancy check is first applied to find if there is an identical HMVP in the HMVP table. If found, the identical HMVP is removed from the HMVP table, and after that all HMVP candidates are moved forward and the identical HMVP is added to the last entry of the HMVP table.
[0138] The HMVP candidates can be used in the MVP candidate list construction process. The latest few HMVP candidates in the HMVP table are checked in order and inserted into the MVP candidate list after the temporal MVP candidate. A redundancy check is applied to the HMVP candidates with respect to the spatial candidates and / or the temporal MVP candidate.
[0139] To reduce the number of redundancy check operations, the following simplifications are introduced:
[0140] - The last two entries in the HMVP table are redundancy checked with respect to the spatial MVP candidates derived from the spatial candidates at positions Al and Bl, respectively; and
[0141] - The MVP candidate list construction process from the HMVP candidates is terminated once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1.
[0142] Derivation of pair-wise average MVP candidate
[0143] The pair-wise average MVP candidate is generated by averaging the MVPs derived using the first two predefined merge candidates in the existing merge candidate list pair. The first merge candidate in the predefined pair can be defined as pOcand and the second merge candidate in the predefined pair can be defined as plcand. The average motion vector is calculated separately for each reference picture list depending on the availability of the motion vectors of pOcand and plcand. If two motion vectors are available for one reference picture list, these two motion vectors are averaged even if they point to different reference pictures and the reference picture of the average motion vector is set to the reference picture of pOcand; if only one motion vector is available for one reference picture list, the motion vector is used directly; if no motion vector is available for one reference picture list, the motion vector and the reference picture index for this reference picture list remain invalid.
[0144] Zero MVP
[0145] When the MVP candidate list is not full after the pair-wise average MVP candidate is added, a zero MVP is inserted at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0146] MMVD
[0147] As mentioned above, in merge mode, the motion information (i.e., MVP candidate) is implicitly derived from the MVP candidate list constructed for the current CU and directly used as the MV of the current CU for generating the prediction samples of the current CU, which can cause some error between the actual MV of the current CU and the implicitly derived MVP. To improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, in which the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. After the regular merge flag is signaled, the MMVD flag is signaled to specify whether the MMVD mode is used for the current CU.
[0148] In MMVD mode, after the MVP candidate is selected from the first two MVP candidates in the MVP candidate list, MMVD information is signaled, in which the MMVD information includes an MMVD candidate flag, which is used to specify which one of the first two MVP candidates is selected as the MV base, a distance index for the indication of the motion magnitude information of the MVD, and a direction index for the indication of the motion direction information of the MVD.
[0149] The distance index for specifying the motion magnitude information of the MVD indicates a predefined offset from a starting point (e.g., indicated by the dashed circle in Figure 9 L0 reference picture 501 or L1 reference picture 503 in Figure 9 from the selected MVP candidate, and the MVD can be derived from the offset and added to the selected MVP candidate. The relationship between the distance index and the predefined offset is specified in Table 1 below. Table 1
[0150] The direction index specifies the sign of the MVD, which indicates the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined sign. It should be noted that the meaning of the sign of the MVD can change depending on the information of the selected MVP candidate. When the selected MVP candidate is a non-predicted MV or a bi-predicted MV with two MVs pointing to the same side of the current picture (i.e., both reference pictures of the current picture (e.g., the reference picture of List 0 and the reference picture of List 1, which are also called L0 reference picture and L1 reference picture, respectively) have POCs greater than the POC of the current picture or both have POCs less than the POC of the current picture), the sign in Table 2 specifies the sign of the MVD added to the MVD of the selected MVP candidate. When the selected MVP candidate is a bi-predicted MV with two MVs pointing to different sides of the current picture (i.e., one reference picture of the current picture has a POC greater than the POC of the current picture and the other reference picture of the current picture has a POC less than the POC of the current picture), if the POC distance of the L0 reference picture (i.e., the POC distance between the L0 reference picture and the current picture) is greater than the POC distance of the L1 reference picture (i.e., the POC distance between the L1 reference picture and the current picture), then the sign in Table 2 specifies the sign of the MVD of the List 0 MVD0 of the MVP of the List 0 MVP0 added to the selected MVP candidate and the sign of the MVD of the List 1 MVD1 of the MVP of the List 1 MVP1 added to the selected MVP candidate is opposite to the sign in Table 2; otherwise, if the POC distance of the L1 reference picture is greater than the POC distance of the L0 reference picture, then the sign in Table 2 specifies the sign of the MVD1 added to the selected MVP candidate and the sign of the MVD0 added to the selected MVP candidate is opposite to the sign in Table 2. Table 2
[0151] The MVD is scaled according to the POC distance. If the POC distance of the L0 reference picture and the L1 reference picture are the same, the MVD does not need to be scaled. Otherwise, if the POC distance of the L0 reference picture is greater than the POC distance of the L1 reference picture, the MVD1 is scaled. If the POC distance of the L1 reference picture is greater than the POC distance of the L0 reference picture, the MVD0 is scaled.
[0152] GPM
[0153] In VVC, GPM is supported for inter prediction. GPM is signaled using a CU-level flag as a kind of merge mode, other merge modes include regular merge mode, MMVD mode, CIIP mode and subblock merge mode. For each CU size except 8x8 64 and 64 8 ( ), GPM supports 64 partitions in total.
[0154] When GPM is used, a CU is split into two parts by a geometrically positioned straight line. The position of the split line is mathematically derived from a specific partition’s angle and offset parameters. Each part of the CU obtained by the geometric split uses its own motion for inter prediction; and only uni-prediction is allowed for each split, i.e., each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that similar to regular bi-prediction, each CU only needs two motion- compensated predictions.
[0155] If GPM is used for the current CU, the geometry partition index is further signaled to indicate the partition mode of the geometric partition (indicating the angle and offset of the geometric partition) and two merge indices (one for each partition).
[0156] The uni-prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process described above. Let n denote the index of a uni-prediction motion vector in the uni-prediction candidate list. The LX motion vector of the n-th merge candidate in the merge candidate list, where X equals the parity of n, is used as the n-th uni-prediction motion vector of GPM. These motion vectors are marked with “x” in Figure 10 . In case the corresponding LX motion vector of the n-th merge candidate in the merge candidate list is not present, the L(l-X) motion vector of the same merge candidate is used instead of the uni-prediction motion vector of GPM.
[0157] CIIP
[0158] In VVC, when a CU is coded in merge mode, an additional flag is signaled to indicate whether the CIIP mode is applied to the current CU if the CU includes at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if both the width and height of the CU are smaller than 128 luma samples. In the CIIP mode, the prediction signal is obtained by combining the inter prediction signal with the intra prediction signal. The inter prediction signal in the CIIP mode is derived using the same inter prediction process as applied in the regular merge mode; and the intra prediction signal in the CIIP mode is derived following the regular intra prediction process with the planar mode. Then, the intra prediction signal and the inter prediction signal are combined using a weighted average, where the weight values are calculated according to the coding modes of the top and left neighboring blocks of the current CU 1601 (as shown in Figure 11
[0159] - isIntraTop is set to 1 if the top neighboring block is available and is intra coded, otherwise isIntraTop is set to 0;
[0160] - if the left neighboring block is available and is intra coded, then isIntraLeft is set to 1, otherwise isIntraLeft is set to 0;
[0161] - if (isIntraLeft + islntraTop) is equal to 2, then the weight value is set to 3;
[0162] - otherwise, if (islntraLeft + islntraTop) is equal to 1, then the weight value is set to 2;
[0163] - otherwise, the weight value is set to 1.
[0164] Prediction signal in CIIP mode is derived as follows:
[0165] wherein, is the inter prediction signal in CIIP mode, is the intra prediction signal in CIIP mode, is the weight value, and denotes a right shift operation.
[0166] Intra block copy in Versatile Video Coding (VVC)
[0167] Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency for screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has been reconstructed within the current picture. The luma block vector of an IBC-coded CU is in integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1 pel and 4 pel motion vector precision. An IBC-coded CU is considered as a third prediction mode in addition to the intra prediction mode or the inter prediction mode. IBC mode is applicable to a CU whose width and height are both less than or equal to 64 luma samples.
[0168] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check for blocks whose width or height is not larger than 16 luma samples. For non-merge mode, block vector search is first performed using hash-based search. If hash search does not return a valid candidate, a local search based on block matching will be performed.
[0169] In hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4x4 sub-blocks. For a current block of larger size, the hash key is determined to match the hash key of the reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found to have hash keys matching the hash key of the current block, the block vector cost of each matching reference is calculated, and the block vector cost with the smallest cost is selected.
[0170] In block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0171] At the CU level, the IBC mode is signaled with a flag, and it can be signaled as IBC AMVP mode or IBC skip / merge mode, as follows:
[0172] IBC skip / merge mode: The merge candidate index is used to indicate which block vectors from the list of neighboring candidate IBC coded blocks are used to predict the current block. The merge list includes spatial, HMVP, and pairwise candidates.
[0173] IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if IBC coded). When either neighbor is not available, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0174] IBC reference region
[0175] To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined region, including the region of the current CTU and some regions of the left CTU. Figure 12 The reference region of IBC mode is shown, where each block represents a 64x64 luma sample unit.
[0176] Depending on the location of the current coded CU position within the current CTU, the following applies:
[0177] If the current block falls into the top-left 64x64 block of the current CTU, the current block can also use the CPR mode to reference the reference samples in the bottom-right 64x64 block of the left CTU, in addition to the samples in the current CTU that have already been reconstructed. The current block can also use the CPR mode to reference the reference samples in the bottom-left 64x64 block of the left CTU and the reference samples in the top-right 64x64 block of the left CTU.
[0178] If the current block falls into the top-right 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, the current block can also use the CPR mode to refer to the reference samples in the bottom-left 64x64 block and the bottom-right 64x64 block of the left-side CTU if the luma position (0, 64) relative to the current CTU has not been reconstructed; otherwise, the current block can also refer to the reference samples in the bottom-right 64x64 block of the left-side CTU.
[0179] If the current block falls into the bottom-left 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, the current block can also use the CPR mode to refer to the reference samples in the top-right 64x64 block and the bottom-right 64x64 block of the left-side CTU if the luma position (64, 0) relative to the current CTU has not been reconstructed; otherwise, the current block can also use the CPR mode to refer to the reference samples in the bottom-right 64x64 block of the left-side CTU.
[0180] If the current block falls into the bottom-right 64x64 block of the current CTU, it can only use the CPR mode to refer to the already reconstructed samples in the current CTU.
[0181] This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementation.
[0182] Interaction of IBC with other coding tools
[0183] Interaction of IBC mode with other inter coding tools in VVC (such as paired merge candidate, history-based motion vector predictor (HMVP), combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partition mode (GPM)) is as follows:
[0184] IBC can be used together with paired merge candidate and HMVP. A new paired IBC merge candidate can be generated by averaging the two IBC merge candidates. For HMVP, the IBC motion is inserted into the history buffer for future reference.
[0185] IBC cannot be used in combination with the following inter tools: affine motion, CIIP, MMVD, and GPM.
[0186] When DUAL_TREE partitioning is used, IBC is not allowed for chroma coding blocks.
[0187] Unlike in the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in the reference picture list 0 for IBC prediction. The derivation process of the motion vector for IBC mode excludes all neighboring blocks in inter mode and vice versa. The following IBC design aspects are applied:
[0188] IBC shares the same process as regular MV merge, including pair-wise merge candidates and history-based motion predictor, but does not allow TMVP and zero vector as they are not valid for IBC mode.
[0189] A separate HMVP buffer (5 candidates per HMVP buffer) is used for regular MV and IBC.
[0190] The block vector constraints are implemented in the form of bitstream conformance constraints, the encoder needs to make sure that no invalid vector exists in the bitstream, and merge should not be used if the merge candidate is invalid (out of range or zero). Such bitstream conformance constraints are expressed in terms of a virtual buffer as described below.
[0191] For deblocking, IBC is treated as inter mode.
[0192] If the current block is coded using IBC prediction mode, AMVR does not use quarter-pel; instead, AMVR is signaled to indicate only whether the MV is inter-pixel or 4 integer-pixel.
[0193] The number of IBC merge candidates can be signaled in the slice header separately from the number of regular, subblock and geometric merge candidates.
[0194] The virtual buffer concept is used to describe the allowable reference region for IBC prediction mode and the valid block vectors. Denote the CTU size as ctbSize, the virtual buffer ibcBuf has width wiBCBuf = 128x128 / ctbSize and height hiBCBuf = ctbSize. For example, for CTU size 128x128, the size of ibcBuf is also 128x128; for CTU size 64x64, the size of ibcBuf is 256x64; and for CTU size 32x32, the size of ibcBuf is 512x32.
[0195] The size of VPDU is min(ctbSize, 64) in each dimension. v = min(ctbSize, 64).
[0196] The virtual IBC buffer ibcBuf is maintained as follows.
[0197] At the beginning of decoding each CTU row, flush the entire ibcBuf with invalid value -1.
[0198] ibcBuf[x][y] = -1, where x = xVPDU%wIbcBuf,..., xVPDU%wIbcBuf+w-1; y = yVPDU%ctbSize,..., yVPDU%ctbSize+w-1 v -1; y = y VPDU %ctbSize,..., yVPDU %ctbSize+w-1 v -1.
[0199] After decoding, the CU includes (x, y) relative to the top-left corner of the picture, set
[0200] ibcBuf[x%wIbcBuf][y%ctbSize] = recSample[x][y]
[0201] For a block covering coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid:
[0202] ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1.
[0203] Intra block copy in enhanced compression model (ECM)
[0204] In ECM, IBC is improved from the following aspects.
[0205] IBC merge / AMVP list construction
[0206] IBC merge / AMVP list construction is modified as follows:
[0207] Only if an IBC merge / AMVP candidate is valid, it can be inserted into the IBC merge / AMVP candidate list.
[0208] Top-right, bottom-left, and top-left spatial candidates, as well as one pair-wise average candidate, can be added to the IBC merge / AMVP candidate list.
[0209] Template-based adaptive reordering (ARMC-TM) is applied in the IBC merge list.
[0210] The HMVP table size of IBC is increased to 25. After deriving up to 20 IBC merge candidates with full pruning, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0211] The zero vector filling the IBC merge / AMVP list is replaced by a set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in IBC merge mode and thus, it is discarded as a BVP in the IBC candidate list.
[0212] Three candidates are located on the nearest corners of the reference region and three additional candidates are determined in the middle of the three sub-regions (A, B and C) with coordinates determined by the width and height of the current block and the ΔΧ and ΔΥ parameters as shown in Figure 13
[0213] IBC with template matching
[0214] Template matching is used in IBC for both IBC merge mode and IBC AMVP mode.
[0215] The IBC-TM merge list is modified compared to the merge list used by regular IBC merge mode so that the candidates are selected according to a pruning method with the motion distance between the candidates in regular TM merge mode. The zero motion realization at the end is replaced by the motion vectors of the left side (-W, 0), the top (0, -H) and the top-left corner (-W, -H) where W is the width of the current CU and H is the height of the current CU.
[0216] In IBC-TM merge mode, the selected candidate is refined with a template matching method before the RDO or the decoding process. The IBC-TM merge mode competes with regular IBC merge mode and a TM merge flag is signaled.
[0217] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of the 3 selected candidates is refined using a template matching method and ordered according to the resulting template matching cost. Then, only the first 2 candidates are usually considered in the motion estimation process.
[0218] The template matching refinement in both IBC-TM merge mode and AMVP mode is very simple because IBC motion vectors are constrained (i) to be integer and (ii) to be within the reference region as shown in Figure 12 Thus, in IBC-TM merge mode, all refinements are performed with integer precision and in IBC-TM AMVP mode, the refinements are performed with integer or 4 image element precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vector and the template used in each refinement step must respect the reference region constraints.
[0219] IBC reference region
[0220] The reference region of IBC is extended to the two CTU rows above. Figure 14 The reference region for encoding a CTU (m, n) is shown. Specifically, for a CTU (m, n) to be encoded, the reference region includes CTUs with indices (m-2, n-2)...(W, n-2), (0, n-1)...(W, n-1), (0, n)...(m, n), where W denotes the maximum horizontal index within the current tile, slice, or picture. This setup ensures that for a CTU size of 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is limited to [- (C « 1), C » 2] in the horizontal direction and [-C, C » 2] in the vertical direction to accommodate the reference region extension, where C denotes the CTU size.
[0221] IBC merge mode with block vector difference
[0222] The IBC merge mode with block vector difference is employed in ECM. The distance set is {1-image element, 2-image element, 4-image element, 8-image element, 12-image element, 16-image element, 24-image element, 32-image element, 40-image element, 48-image element, 56-image element, 64-image element, 72-image element, 80-image element, 88-image element, 96-image element, 104-image element, 112-image element, 120-image element, 128-image element} and the BVD direction is two horizontal directions and two vertical directions.
[0223] The base candidate is selected from the top five candidates in the reordered IBC merge list. And based on the SAD cost between the template (one row above and one column left of the current block) and the reference of each refinement location, all possible MBVD refinement locations (20x4) for each base candidate are reordered. Finally, the top 8 refinement locations with the lowest template SAD cost are kept as available locations, thus for MBVD index coding.
[0224] IBC adaptation for camera captured content
[0225] When adapting IBC for camera captured content, the IBC reference range is reduced from 2 CTU rows to 2x128 rows, as shown in Figure 15 At the encoder side, to reduce complexity, the local search range is set to be centered at the first block vector predictor of the current CU, [-8, 8] in the horizontal direction and [-8, 8] in the vertical direction. This encoder modification does not apply to SCC sequences.
[0226] Combination of CIIP with TIMD and TM merge
[0227] In the CIIP mode, the prediction samples are generated by weighting the inter prediction signal using the CIIP-TM merge candidate and the intra prediction signal predicted using the TIMD derived intra prediction mode. This method is applied only for coding blocks with an area smaller than or equal to 1024.
[0228] The TIMD derivation method is used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of the 67 regular intra prediction modes.
[0229] Furthermore, if the derived intra prediction mode is an angular mode, it is also proposed to modify the weights (wIntra, wInter) of the two tests. For near horizontal modes (2 <= angle mode index < 34), the current block is vertically partitioned as shown in Figure 16A For near vertical modes (34 <= angle mode index < 66), the current block is horizontally partitioned as shown in Figure 16B
[0230] The different sub-blocks (wIntra, wInter) are shown in Table 3. The modified weights for angular modes are shown in Table 3.
[0231] With CIIP-TM, a CIIP-TM merge candidate list is constructed for the CIIP-TM mode. The merge candidates are refined by template matching. The ARMC method also reorders the CIIP-TM merge candidates as regular merge candidates. The maximum number of CIIP-TM merge candidates is equal to two.
[0232] Multi-hypothesis prediction (MHP)
[0233] In the multi-hypothesis inter prediction mode, one or more additional motion- compensated prediction signals are signaled in addition to the traditional bi-prediction signal. The resulting overall prediction signal is obtained by sample-wise weighted addition. With the two prediction signals and the first additional inter prediction signal / hypothesis the resulting prediction signal is obtained by: (2)
[0234] According to the mapping given in Table 4, the weighting factors are specified by the new syntax element add_hyp_weight_idx: Table 4 Mapping between add_hyp_weight_idx and
[0235] Similarly to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. (3)
[0236] The resulting overall prediction signal is obtained as the final (i.e., the prediction signal with the largest index n ). Within this mode, up to two additional prediction signals can be used (i.e., n is limited to 2).
[0237] The motion parameters of each additional prediction hypothesis can be explicitly signaled by specifying the reference index, the motion vector predictor index and the motion vector difference or implicitly signaled by specifying the merge index. A separate multi-hypothesis merge flag distinguishes between the two signaling modes.
[0238] For the inter AMVP mode, MHP is applied only in the case of non-equal weights in BCW, selected in bi-prediction mode.
[0239] A combination of MHP and BDOF is possible, however BDOF is applied only to the bi-predicted signal part of the prediction signal (i.e., the first two hypotheses in normal).
[0240] Geometric partition mode (GPM) in ECM
[0241] GPM with merge motion vector difference (MMVD)
[0242] GPM in VVC is extended by applying motion vector refinement on top of the existing GPM uni-directional MV. A flag is first signaled for the GPM CU to specify whether this mode is used or not. If this mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD or not. If the MVD is signaled for a geometric partition, the motion of the partition is further refined by the signaled MVD information after the GPM merge candidate is selected. All other procedures remain the same as in GPM.
[0243] Similar to MMVD, MVD signals a pair of distance and direction. Nine candidate distances (¼-pel, ½-pel, 1-pel, 2-pel, 3-pel, 4-pel, 6-pel, 8-pel, 16-pel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are involved in GPM with MMVD (GPM-MMVD). In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, MVD is left-shifted by 2 as MMVD.
[0244] GPM with template matching (TM)
[0245] Template matching is applied to GPM. When GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to two geometric partitions. TM is used to refine the motion information for each geometric partition. When TM is selected, a template is constructed using the left, top, or both left and top neighboring samples depending on the partition angle as shown in Table 5. Then the motion is refined by using the same search mode of merge mode with disabled half-pel interpolation filter to minimize the difference between the current template and the template in the reference picture. Table 5
[0246] Table 5 shows the template for the first and second geometric partitions, where A denotes using the top sample, L denotes using the left sample, and L+A denotes using both the left and top samples.
[0247] The GPM candidate list is constructed as follows:
[0248] 1. Interleaved list-0 and list-1 MV candidates are derived directly from the regular merge candidate list, where the priority of list-0 MV candidates is higher than that of list-1 MV candidates. A pruning method with adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0249] 2. Interleaved list-1 and list-0 MV candidates are derived directly from the regular merge candidate list, where the priority of list-1 MV candidates is higher than that of list-0 MV candidates. The same pruning method with adaptive threshold is also applied to remove redundant MV candidates.
[0250] 3. Zero MV candidates are padded until the GPM candidate list is full.
[0251] GPM-MMVD and GPM-TM are specific to one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.
[0252] GPM with inter and intra prediction
[0253] In GPM with inter and intra prediction, the final prediction samples are generated by weighting the inter and intra prediction samples of each GPM-separate region. The inter prediction samples are derived by inter GPM, while the intra prediction samples are derived by an intra prediction mode (IPM) candidate list and an index signaled from the encoder. The IPM candidate list size is predefined to be 3. The available IPM candidates are the parallel angular mode (parallel mode) for GPM block boundary, the vertical angular mode (vertical mode) for GPM block boundary, and the planar mode, respectively, as shown in Figure 17A to Figure 17C Figure 17D GPM with inter and intra prediction is restricted to reduce the signaling overhead of IPM, and to avoid the size increase of intra prediction circuit on hardware decoder, as shown in
[0254] In DIMD and neighbor mode based IPM derivation, the parallel mode is registered first. Therefore, up to two IPM candidates derived from decoder side intra mode (DIMD) method and / or neighbor block derivation can be registered if the same IPM candidate does not exist in the list. For neighbor mode derivation, there are at most five available neighbor block positions, but they are restricted by the angles of GPM block boundary as shown in Table 6, which has already been used for GPM with template matching (GPM-TM). Table 6
[0255] Table 6 shows the available neighbor block positions for IPM candidate derivation based on the angles of GPM block boundary. A and L denote the top and left of the prediction block.
[0256] GPM-intra can be combined with GPM (GPM-MMVD) that merges with motion vector difference. TIMD is used for IPM candidate of GPM-intra to further improve the coding performance. The IPM candidates of TIMD, DIMD, and neighbor block can be registered first after the parallel mode.
[0257] Template matching based reordering of GPM split modes
[0258] In template matching based reordering of GPM split modes, given the motion information of the current GPM block, the corresponding TM cost value of GPM split modes is calculated. Then, all GPM split modes are reordered in ascending order based on the TM cost value. Instead of sending GPM split modes, Golomb-Rice code is signaled to indicate the index of where the exact GPM split mode is located in the reordered list.
[0259] The reordering method of GPM partition modes is a two-step process performed after the generation of the corresponding reference templates of the two GPM partitions in a coding unit, as follows:
[0260] • The GPM partition edges are extended into the reference templates of the two GPM partitions, resulting in 64 reference templates and calculating the corresponding TM cost of each of the 64 reference templates;
[0261] • The GPM split modes are reordered in ascending order based on their TM cost values, and the best 32 are marked as available partition modes.
[0262] As shown in Figure 18 , the edges on the template are extended from the edges of the current CU, but the GPM blending process is not used in the template area that crosses the edges.
[0263] After the incremental reordering using TM cost, the index is signaled.
[0264] Intra template matching
[0265] Intra template matching prediction (Intra-TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode and the same prediction operation is performed at the decoder side.
[0266] The prediction signal is generated by matching the L-shaped causal neighbor of the current block with another block in the search area predefined in Figure 19 , which consists of the following parts:
[0267] R1 : the current CTU
[0268] R2: the top-left CTU
[0269] R3: the top CTU
[0270] R4: the left CTU
[0271] The sum of absolute differences (SAD) is used as cost function.
[0272] Within each region, the decoder searches for the template having the smallest SAD with respect to the current template and uses its corresponding block as prediction block.
[0273] The size of all regions (SearchRange_w, SearchRange_h) is set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:
[0274] SearchRange_w = a BlkW
[0275] SearchRange_h = a BlkH
[0276] where 'a' is a constant controlling the gain / complexity trade-off. In practice, 'a' is equal to 5.
[0277] The intra template matching tool is enabled for CUs with width and height smaller or equal to 64. The maximum CU size for intra template matching is configurable.
[0278] When the current CU does not use DIMD, the intra template matching prediction mode is signaled at the CU level by a dedicated flag.
[0279] Fusion for template-based intra mode derivation (TIMD)
[0280] For each intra prediction mode in the MPM, the SATD between the prediction samples of the template and the reconstructed samples is computed. The two intra prediction modes with the smallest SATD are first selected as TIMD modes. These two TIMD modes are fused with a weight after applying the PDPC process and this weighted intra prediction is used to encode the current CU. The position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0281] The cost of the two selected modes is compared with a threshold, in the test, a cost factor 2 is applied as follows:
[0282] costMode2 < 2 costMode1
[0283] If this condition is true, the fusion is applied, otherwise only mode 1 is used.
[0284] The weight of the modes is computed according to their SATD cost as follows:
[0285] Weight 1 = costMode2 / (costMode1 + costMode2)
[0286] Weight 2 = 1 - Weight 1
[0287] The division operation is performed using the same lookup table (LUT)-based integration scheme used by CCLM.
[0288] Local illumination compensation (LIC)
[0289] LIC is an inter prediction technique that models the local illumination change between the current block and its prediction block according to the local illumination change between the current block template and the reference block template. The parameters of the function can be represented by a scale a and an offset b, which form a linear equation, i.e., a p[x] + b to compensate the illumination change, where p[x] is the reference sample at position x on the reference picture pointed by the MV. When wraparound motion compensation is enabled, the wraparound offset should be considered to clip the MV. Since a and b can be derived based on the current block template and the reference block template, they do not require signaling overhead except for signaling an LIC flag for AMVP mode to indicate the usage of LIC.
[0290] The local illumination compensation proposed in JVET-O0066 is used for uni-prediction inter CUs with the following modifications:
[0291] Intra neighboring samples can be used for LIC parameter derivation;
[0292] LIC is disabled for blocks with less than 32 luma samples;
[0293] For both non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16x16 unit; and
[0294] The samples of the reference block template are generated by using MC with block MV without rounding it to integer pixel precision.
[0295] Overlapped block motion compensation (OBMC)
[0296] When OBMC is applied, the motion information of neighboring blocks with weighted prediction as described in JVET-L0101 is used to refine the top and left boundary pixels of a CU.
[0297] The condition under which OBMC is not applied is as follows:
[0298] When OBMC is disabled at SPS level;
[0299] When the current block has intra mode or IBC mode;
[0300] LIC is applied for the current block; and
[0301] The current luma block area is smaller than or equal to 32.
[0302] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. It enables for sub-block based coding tools:
[0303] Affine AMVP mode;
[0304] Affine merge mode and sub-block based temporal motion vector prediction (SbTMVP); and
[0305] Sub-block based bilateral matching.
[0306] When OBMC mode is used in CIIP mode with LMCS, inter blending is performed before LMCS mapping of inter samples. LMCS is applied to the blended inter samples combined with intra samples in CIIP mode where LMCS is applied,
[0307] wherein represents a sample predicted by motion of the current block in the original domain, represents a sample predicted in the mapping domain, represents a sample predicted by motion of a neighboring block in the original domain, and and are weights.
[0308] Template matching based OBMC
[0309] In the template matching based OBMC scheme, instead of directly using weighted prediction, the prediction value of the CU boundary sample derivation method is decided according to the template matching cost, including using motion information of the current block only, or using motion information of neighboring blocks and one of blending modes.
[0310] In this scheme for each block with 4x4 size at the top CU boundary, the above template size is equal to 4x1. If N neighboring blocks have the same motion information, the above template size is enlarged to 4Nx1 because the MC operation can be processed at one time. For each left block with 4x4 size at the left CU boundary, the left template size is equal to 1x4 or 1x4N (N is the number of neighboring blocks with the same motion information). Figure 20 ).
[0311] For each 4x4 top block (or N 4x4 block group), the prediction value of the boundary sample is derived after the following steps.
[0312] Take block A as the current block, and its above neighbor block AboveNeighbor_A as an example. The operation of the left block is performed in the same way.
[0313] First, measure three template matching costs (Cost1, Cost2, Cost3) by SAD between the reconstructed samples of the template and their corresponding reference samples derived by the MC process according to the following three types of motion information:
[0314] Calculate Cost1 according to the motion information of A.
[0315] Calculate Cost2 according to the motion information of AboveNeighbor_A.
[0316] Calculate Cost3 according to the weighted prediction of the motion information of A and AboveNeighbor_A, with the weighting factors being 3⁄4 and 1⁄4 respectively.
[0317] Second, select one method to calculate the final prediction result of the boundary samples by comparing Cost1, Cost2 and Cost3.
[0318] The original MC result using the motion information of the current block is denoted as Pixel1, and the MC result using the motion information of the neighboring block is denoted as Pixel2. The final prediction result is denoted as NewPixel.
[0319] If Cost1 is the smallest, then NewPixel(i, j) = Pixel1(i, j).
[0320] If (Cost2 + (Cost2 » 2) + (Cost2 » 3)) <= Cost1, then use hybrid mode 1.
[0321] For luma blocks, the number of mixed pixel rows is 4.
[0322] NewPixel(i, 0) = (26 * Pixel1(i, 0) + 6 * Pixel2(i, 0) + 16) » 5
[0323] NewPixel(i, 1) = (7 * Pixel1(i, 1) + Pixel2(i, 1) + 4) » 3
[0324] NewPixel(i, 2) = (15 * Pixel1(i, 2) + Pixel2(i, 2) + 8) » 4
[0325] NewPixel(i, 3) = (31 x Pixel1(i, 3) + Pixel2(i, 3) + 16) » 5
[0326] For chroma blocks, the number of mixed pixel rows is 1.
[0327] NewPixel(i, 0) = (26 x Pixel1(i, 0) + 6 x Pixel2(i, 0) + 16) » 5
[0328] If Cost1<= Cost2, use blending mode 2.
[0329] For luma blocks, the number of mixed pixel rows is 2.
[0330] NewPixel(i, 0) = (15 x Pixel1(i, 0) + Pixel2(i, 0) + 8) » 4
[0331] NewPixel(i, 1) = (31 x Pixel1(i, 1) + Pixel2(i, 1) + 16) » 5
[0332] For chroma blocks, the number of mixed pixel rows / columns is 1.
[0333] NewPixel(i, 0) = (15 x Pixel1(i, 0) + Pixel2(i, 0) + 8) » 4
[0334] Otherwise, use blending mode 3.
[0335] For luma blocks, the number of mixed pixel rows is 4.
[0336] NewPixel(i, 1) = (7 x Pixel1(i, 1) + Pixel2(i, 1) + 4) » 3
[0337] NewPixel(i, 2) = (15 x Pixel1(i, 2) + Pixel2(i, 2) + 8) » 4
[0338] NewPixel(i, 3) = (31 x Pixel1(i, 3) + Pixel2(i, 3) + 16) » 5
[0339] For chroma blocks, the number of mixed pixel rows is 1.
[0340] NewPixel(i, 0) = (7 x Pixel1(i, 0) + Pixel2(i, 0) + 4) » 3
[0341] Merge candidate adaptive reordering based on template matching (ARMC-TM)
[0342] The reordering method is applied to regular merge mode, template matching (TM) merge mode and affine merge mode (excluding SbTMVP candidates). For TM merge mode, the merge candidates are reordered before the refinement process.
[0343] After constructing the merge candidate list, the merge candidates are partitioned into several subgroups. For regular merge mode and TM merge mode, the subgroup size is set to 5. For affine merge mode, the subgroup size is set to 3. The merge candidates in each subgroup are incrementally reordered according to the template matching based cost value. For simplicity, the merge candidates in the last subgroup, instead of the first subgroup, are not reordered.
[0344] The template matching cost of a merge candidate is measured by the sum of absolute difference (SAD) between the samples of the template of the current block and their corresponding reference samples. The template includes a set of reconstructed samples neighboring the current block. The reference samples of the template are located by the motion information of the merge candidate.
[0345] When a merge candidate utilizes bi-prediction, the reference samples of the template of the merge candidate are also generated by bi-prediction, as shown in Figure 21 .
[0346] For a subblock-based merge candidate with subblock size equal to WsubxHsub, the above template includes several sub-templates with size Wsubx1, and the left template includes several sub-templates with size 1xHsub. As shown in Figure 22 , the reference samples of each sub-template are derived using the motion information of the subblocks in the first row and the first column of the current block.
[0347] Direct block vector for chroma blocks
[0348] Direct block vector is used for chroma blocks in dual tree slices. When chroma dual tree is activated, a flag is signaled to indicate whether IBC mode is used to code the chroma blocks. If Figure 23 one of the five positions in the luma blocks shown in is coded in IBC or intra TM P mode, its block vector is scaled and used as the block vector for the chroma blocks. Template matching is used to perform the block vector scaling.
[0349] While the existing IBC scheme can provide significant improvement of intra coding in ECM, there is room to further improve its performance. At the same time, some parts of the existing Convolutional Cross-Component Model (CCCM) mode also need to be simplified for efficient codec hardware implementation or improved for better coding efficiency. Furthermore, there is a need to further improve the trade-off between its implementation complexity and its coding efficiency benefits.
[0350] In the present disclosure, to address the above problems, methods to further improve the existing design of IBC are provided. Generally, the main features of the techniques proposed in the present disclosure are summarized as follows.
[0351] Filtering IBC prediction with CCCM tools. Filtered Intra Block Copy (FIBC) is a special intra prediction mode that applies filters to IBC-based prediction blocks to increase prediction accuracy and adapt the characteristics of the copied block to the local neighborhood.
[0352] In FIBC, the training samples can be adjacent to the current block. Knowing the reference from the local area can improve the accuracy of the prediction in the prediction.
[0353] In FIBC, the training samples can not be adjacent to the current block. Knowing the reference from the non-local area can also improve the accuracy of the prediction in the prediction.
[0354] In FIBC, only one hypothesis can be utilized, i.e., the best matching block that results in the minimum matching cost is selected as the final prediction.
[0355] In FIBC, multiple hypotheses can also be utilized.
[0356] It should be understood that the figures in the present disclosure can be combined with all examples mentioned in the present disclosure, and the disclosed methods can be applied independently or jointly.
[0357] Filtered Intra Block Copy (FIBC)
[0358] According to one or more embodiments of the present disclosure, IBC prediction is filtered with CCCM tools. Different approaches can be used to achieve this goal. The existing CCCM mode applies various filters for predicting chroma sample values based on corresponding luma sample values. Unlike CCCM, Filtered Intra Block Copy (FIBC) is a special intra prediction mode that applies filters on IBC-based prediction blocks to predict target luma or chroma samples of the current block based on corresponding luma or chroma samples of the reference block, respectively, in order to increase prediction accuracy and adapt the characteristics of the copied block to the local neighborhood.
[0359] According to one or more embodiments of the present disclosure, IBC prediction is further filtered. Different approaches can be used to achieve this goal. Filtered Intra Block Copy (FIBC) is a special intra prediction mode that applies filters to IBC-based prediction blocks to increase prediction accuracy and adapt the characteristics of the copied block to the local neighborhood.
[0360] According to one or more embodiments of the disclosure, the reconstructed luma / chroma samples on the template region of the reference block are used as inputs of the filter during the training stage, and the corresponding reconstructed luma / chroma samples in the template region of the current block are the targets. In one example, Figure 25 One filter shape (cross shape) and the training region of the reference block are shown. It should be understood that for this filter shape, both the template region and the boundary region of the template region can be part of the training region of the reference block. The reconstructed samples in the boundary region can be used for training when available, and they are filled with the closest available samples when not available. On the other hand, the training region of the current block can be determined as the template region of the current block. In the filtering stage where the filter coefficients of the filter have been trained / determined through the training stage, the filter can be applied to the corresponding sample values of the reference block and the boundary region of the reference block to predict each of the sample values of the current block.
[0361] According to one or more embodiments of the disclosure, the filter coefficients (i.e., parameters) are derived using the regression-based MSE minimization technique (i.e., LDL decomposition) that exists in ECM and is utilized by other tools such as CCCM.
[0362] According to one or more embodiments of the disclosure, a convolutional N-tap (N is an integer and greater than 1) filter can include (N-1-M) tap (M is an integer) spatial terms, M non-linear terms, and a bias term. The (N-1-M) tap spatial terms correspond to neighboring sample values such as luma samples (i.e., L0, L1, …, L8) from the reconstructed reference block as shown. Figure 26 In this example, the formula for each new predicted luma sample is as follows:
[0363] wherein, is a coefficient associated with and is an offset (i.e., 1 « (bitDepth - 1)). The reference luma sample value of the top-left sample neighboring the current block is available as the offsetLuma value. The positions and number of the spatial terms and the non-linear terms can be different. Examples of different shapes / numbers of filter taps are shown as follows. Figure 27 For another example, different positions and numbers are used as shown in the following table.
[0364] According to one or more embodiments of the disclosure, the number of filter taps can be predefined or signaled / switched at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / Region / CTU / CU / Subblock / Sample level.
[0365] According to one or more embodiments of the present disclosure, the template size and shape can be the same as in intra-TMP, the template size for training depends on their availability 4 lines above and left of the current block.
[0366] According to one or more embodiments of the present disclosure, the template size for training depends on their availability up to 5 lines above and left of the current block.
[0367] According to one or more embodiments of the present disclosure, the template size and shape can be the same as in CCCM, the template size for training is 6 lines above and left of the current block, depending on their availability.
[0368] According to one or more embodiments of the present disclosure, the template size for training can be N lines above and left of the current block, depending on their availability, N is an integer.
[0369] According to one or more embodiments of the present disclosure, the template size for training can be N lines above of the current block, depending on their availability, N is an integer.
[0370] According to one or more embodiments of the present disclosure, the template size for training can be N lines left of the current block, depending on their availability, N is an integer.
[0371] According to one or more embodiments of the present disclosure, the reference samples of the reference block / template region of the current block can be predefined or signaled / switched in different coding levels (such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / samples level).
[0372] According to one or more embodiments of the disclosure, the position information can be used to compute the model parameters, including with the horizontal / vertical / diagonal distance and its non-linear terms, one or more position information can be used for this purpose. In one example, the position-based parameters are related to the vertical and horizontal coordinates of the center luma sample (Xc, Yc), and are computed relative to the top-left coordinates of the block (Xtl, Ytl), e.g., Xc-Xtl+Yc-Ytl. In another example, the position-based parameters are related to the vertical and horizontal coordinates of the center luma sample (Xc, Yc), and are computed relative to the top-left coordinates of the block (Xtl, Ytl), e.g., Xc-Xtl+Yc-Ytl, Xc-Xtl, Yc-Ytl. In yet another example, the position-based parameters are related to the vertical and horizontal coordinates of the center luma sample (Xc, Yc), and are computed relative to the top-left coordinates of the block (Xtl, Ytl), e.g., (Xc-Xtl+Yc-Ytl) / N, where N is a predefined number, e.g., 2. In yet another example, the position-based parameters are related to the vertical and horizontal coordinates of the center luma sample (Xc, Yc), and are computed relative to the top-left coordinates of the block (Xtl, Ytl), e.g., (Xc-Xtl+Yc-Ytl) / N1, (Xc-Xtl) / N2, (Yc-Ytl) / N3, where N1~N3 are predefined numbers, such as 2, 3, and 4. In yet another example, the position-based non-linear terms are expressed as the power of two of the horizontal / vertical / diagonal distance, e.g., (Xc-Xtl+Yc-Ytl) (Xc-Xtl+Yc-Ytl), (Xc-Xtl) (Xc-Xtl), (Yc-Ytl) (Yc-Ytl), where (Xc, Yc) are the vertical and horizontal coordinates of the center luma sample, and (Xtl, Ytl) are the top-left coordinates.
[0373] According to one or more embodiments of the disclosure, one enabling flag can be signaled in the bitstream to indicate the used FIBC mode. The enabling flag can be signaled at different coding levels, such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0374] According to one or more embodiments of the disclosure, instead of explicitly signaling the selected mode flag, the mode flag can be derived at the decoder to save the bit overhead.
[0375] According to one or more embodiments of the present disclosure, no additional control flag is needed and the FIBC mode will be derived under certain predefined condition (e.g., certain mode, certain block size, certain partition). When the predefined condition is matched, the FIBC mode will be derived based on the previously decoded information.
[0376] According to one or more embodiments of the present disclosure, samples in a region that is not adjacent to the current block can be used to derive the model of the current block. In one embodiment, a candidate region list with N candidates can be constructed by checking potential MxM regions in sequence. If a checking region is available, it will be put into the candidate region list. For example, a candidate region list with 6 candidates is constructed by checking potential 8x8 regions in sequence. The top-left positions of the potential 8x8 regions are pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (2xStep, 0), (0, -2yStep), (-2xStep, -2yStep), (-xStep / 2, 0), (0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep = Max(width, 16), yStep = Max(height, 16). Some possible positions of the candidate regions are shown. Figure 28
[0377] According to one or more embodiments of the present disclosure, one non-adjacent neighboring candidate with N candidates can be constructed from the positions and the inclusion order of the two sets of spatial non-adjacent neighboring candidates from the inter merge mode. If a checking region is available, it will be put into the candidate region list. Figure 29 Some possible positions of the candidate are shown.
[0378] According to one or more embodiments of the present disclosure, inherited parameters from previously decoded TB / CB / slice / picture / sequence level FIBC can be used in the current block. According to one or more embodiments of the present disclosure, one control flag is signaled in the TB / CB / slice / picture / sequence level to indicate whether the signaling of inherited FIBC is enabled or disabled. When the control flag is signaled as enabled, further flags of inherited FIBC are signaled to the decoder to indicate whether inherited FIBC is used at the signaling level.
[0379] According to one or more embodiments of the present disclosure, derived parameters from previously decoded TB / CB / slice / picture / sequence level FIBC can be stored and used as the current FIBC, which is referred to as inherited FIBC. In one embodiment, a history-based FIBC (H-FIBC) table can be maintained similar to the HMVP table. In one embodiment, an index value can be signaled in the bitstream to indicate which candidate model in the H-FIBC table is selected. In one embodiment, the corresponding table can be updated after decoding a FIBC coded block. In one embodiment, the size of the H-FIBC table is N. N is an integer (e.g., 4, 5, 6, 7).
[0380] Multiple hypothesis FIBC
[0381] According to one or more embodiments of the present disclosure, more than one prediction block candidate is used and weighted to generate the final prediction of the current block. N prediction block candidates are used.
[0382] Prediction block candidate derivation
[0383] In one embodiment, the prediction block candidates are searched and selected according to the criterion of minimizing the template matching cost, i.e., the top N candidates that result in the minimum template matching cost are selected. The template matching cost can not be limited to SAD (sum of absolute difference) and SSE (sum of squared error).
[0384] In one embodiment, the prediction block candidates can be selected according to a predefined mode (i.e., planar mode).
[0385] In one embodiment, the prediction block candidates can be selected according to a neighboring predefined mode (i.e., top predefined mode, left predefined mode).
[0386] Fixed multiple hypothesis FIBC
[0387] In this embodiment, the weighting factors to generate the final prediction block are predefined and fixed at both the encoder side and the decoder side. As an example, equal weighting factors can be used, i.e., 1 / N for all candidate blocks.
[0388] Adaptive multi-hypothesis intra FIBC
[0389] To accommodate different characteristics of video content, an adaptive multi-hypothesis intra FIBC method is also proposed.
[0390] In one embodiment, the weighting factors can be derived based on the template matching cost. The template matching cost of N candidates is denoted as , , The weighting factors are calculated as follows. (4)
[0391] It should be noted that the template matching cost can be measured by (but not limited to) SAD and SSE.
[0392] In yet another embodiment, the weighting factors can be derived / switched based on the block size or syntax elements signaled in different coding levels (such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level).
[0393] In yet another embodiment, the weighting factors can be derived at the encoder side and then signaled to the decoder in the bitstream. The N prediction block candidates are denoted as , , and the current block is denoted as The weighting factors can then be solved by the following equation: (5)
[0394] Equation (5) can be solved using Wiener-Hopf equation as ALF. The derived filter coefficients are then quantized to integer type and signaled at the block level.
[0395] In yet another embodiment, the weighting factors are derived based on the template and applied to the prediction block candidates to generate the final prediction block. The template of the prediction candidates is denoted as , , and the current block is denoted as The weighting factors can then be derived using the following equation: (6)
[0396] Equation (6) can be solved using Wiener-Hopf equation. The final prediction block can then be calculated as where denotes the i-th prediction block candidate.
[0397] The FIBC mode exploits the non-local correlation to improve the prediction accuracy, where similar blocks are searched and used to generate the final prediction block. In this embodiment, a combination of non-local mean filtering and multi-hypothesis FIBC is proposed, as described below. In the first step, N prediction block candidates are searched and identified to be performed in FIBC. In the second step, the weighting factors are calculated as follows. (7)
[0398] where is used to measure the distance between the template of the i-th prediction block candidate and the template of the current block, is used as the weighting degree, and is a normalization constant: (8)
[0399] To compute the weighting factors in equation (7), the strength of the weighting should be determined first. In this disclosure, several methods are proposed to decide the weighting strength.
[0400] In the first method, a list of weighting strength candidates including some typical weighting strength values is defined and fixed at both the encoder and the decoder side. At the encoder side, the rate-distortion optimization is used to examine the weighting strength values, and the best weighting strength value is identified and signaled to the decoder side in the bitstream.
[0401] In the second method, the template of the prediction block candidates and the template of the current block are used to estimate the weighting strength value. The template of the prediction candidates is denoted as , , and the current block is denoted as . Then, the weighting strength value can be solved using the following equation: (9)
[0402] In the third method, the QP value and the variance of the template of the current block can be used to estimate the weighting strength value, i.e., the relationship between the weighting strength value, the QP value and the template variance can be fitted offline.
[0403] To better exploit the non-local correlation in FIBC, in this embodiment, the singular value decomposition (SVD) is used to generate the final prediction block from the prediction block candidates. The width and height of the current block are denoted as W and H, and the area of the current block is denoted as .
[0404] Step 1. Search and identify K prediction block candidates as performed in FIBC .
[0405] Step 2. The current block K prediction block candidate construction block groups and arranged as a matrix: (10)
[0406] where, is a matrix of size by arranging each candidate in group as a column vector.
[0407] Step 3. Perform SVD decomposition on the matrix . (11)
[0408] Step 4. Apply a soft-thresholding operation on the singular value matrix . (12)
[0409] where is the diagonal element of shrunken by a threshold . For the kth diagonal element in , it is shrunk by a non-linear function at level : (13)
[0410] is a matrix consisting of singular values shrunken at the diagonal positions.
[0411] Step 5. Perform inverse SVD to obtain the filtered patch group. (14)
[0412] One of the key steps is to determine the threshold for each diagonal element in Step 4. In this disclosure, the threshold is calculated as follows. Estimate the threshold for each group of image blocks with the following equation: (15)
[0413] where is the standard deviation of the noise, and is the standard deviation of the original block in the kth dimension of the SVD space of group . The deviation of the original block in the SVD space is estimated as follows. (16)
[0414] where is the kth singular value of . When When the value is zero, the soft-thresholding operation is skipped. In addition, the power function with the parameterization is used to estimate the deviation of the noise with the deviation of the prediction block. and The power function with the parameterization estimates the deviation of the noise with the deviation of the prediction block. (17)
[0415] where is calculated as follows, (18)
[0416] Here, denotes the i-th pixel of the prediction block candidate vector .
[0417] Multiple hypothesis FIBC signaling
[0418] In the present disclosure, the proposed multiple hypothesis FIBC can be used as a replacement of the current FIBC mode, or the encoder can adaptively choose the FIBC mode or the multiple hypothesis FIBC mode.
[0419] In one embodiment, the proposed multiple hypothesis FIBC is used as a replacement of the current FIBC mode, i.e. always use multiple hypotheses for prediction.
[0420] In yet another embodiment, one of the multiple hypothesis FIBC methods in the above sections is used jointly with the current FIBC mode. A flag is signaled in the bitstream to indicate whether the multiple hypothesis FIBC mode is applied to the CU.
[0421] In yet another embodiment, more than one of the multiple hypothesis FIBC methods in the above sections is used jointly with the current FIBC mode. First a flag is signaled in the bitstream to indicate whether the multiple hypothesis FIBC mode is applied. Then an index is signaled to indicate which of the multiple hypothesis FIBC methods is applied to the CU.
[0422] Coordination of filters for FIBC and FTMP modes
[0423] TMP prediction can also be filtered with the CCCM tool, which is referred to as filtered template matching prediction (FTMP) mode. The process of FTMP mode is the same as that of FIBC mode, except that FTMP mode does not need a signaled block vector from the encoder to find the reference block. Instead, in FTMP mode, the reference block can be determined at the decoder side by searching for the L-shaped template that is most similar to the current template in the reconstructed part of the current frame, and using the corresponding block as the reference block for the current block to be predicted. In other words, the L-shaped template associated with the reference block is the template that is most similar to the L-shaped template associated with the current block in the reconstructed part of the frame. This operation for determining the reference block is the same as that of the intra-TMP mode. After the reference block is determined, the same filtering process as FIBC mode is applied to predict the target luma or chroma samples of the current block based on the corresponding luma or chroma samples of the reference block, respectively. For example, as shown in FIG. 18, a cross-shaped filter can be applied to the corresponding sample values (the sample values of the reference block and the boundary regions of the reference block) to predict each of the sample values of the current block. Figure 25
[0424] According to one or more embodiments of the disclosure, the same filter shape and / or template region can be applied to both prediction in FIBC mode and prediction in FTMP mode. For example, before deciding which mode to apply, the decoder and / or the encoder can try both FIBC mode and FTMP mode with the same filter shape and / or template region to achieve better performance or reduced cost. Different approaches can be used to achieve this goal.
[0425] In a first example, it is proposed to apply the filter operation used in FTMP mode to FIBC mode as well. In one example, the 6-tap filter used for FTMP mode (cross shape with 5 spatial components and a bias term) and the template region used for training (4 lines above and to the left of the current block in terms of samples, depending on their availability) can also be applied to FIBC mode in the same CU.
[0426] In a second example, it is proposed to apply the filter operation used in FIBC mode to FTMP mode as well. In one example, the 2-tap filter used for FIBC mode (single-sample filter with 1 spatial component and a bias term) and the template region used for training (1 line above and to the left of the current block, depending on their availability) can also be applied to FTMP mode in the same CU.
[0427] Coordination of filters for FIBC mode, FTMP mode, and CCCM mode
[0428] The filter is used in the Convolutional Cross-Component Model (CCCM) mode to predict chroma sample values based on corresponding luma sample values. In the CCCM mode, a set of chroma sample values in the reconstructed region of the current block to be predicted and corresponding luma sample values of the set of chroma sample values are used to determine filter coefficients of the CCCM filter. In one example, the chroma sample values to be predicted and their corresponding luma sample values are co-located sample values. Although the training results of the filter coefficients can be different, the same filter shape and / or template region can be reused in CCCM, FIBC, and FTMP for better performance / reduced cost. In one example, the filter shape and / or template region can be signaled by the encoder to the decoder. In another example, the filter shape and / or template region can be derived by the encoder based on a predetermined rule (e.g., from a predefined candidate set).
[0429] According to one or more embodiments of the disclosure, the same filter shape and / or template area can be applied to at least two of the FIBC mode, the FTMP mode, and the CCCM mode. Different approaches can be used to achieve this goal.
[0430] In a first example, it is proposed to apply the filter operation used in the CCCM mode to the FIBC mode as well. In one example, the 7-tap filter used for the CCCM mode (cross shape with 5 spatial components, a non-linear term, and a bias term) and the template region used for training (6 lines above and to the left of the current block, depending on their availability) can also be applied to the FIBC mode in the same CU.
[0431] In a second example, it is proposed to apply one of the filter operations used in the CCCM mode to both the FIBC mode and the FTMP mode. In one example, the 11-tap filter used for the CCCM mode (3 squares with 9 spatial components, a non-linear term, and a bias term) and the template region used for training (6 lines above and to the left of the current block, depending on their availability) can also be applied to the FIBC mode and the FTMP mode. 3 squares with 9 spatial components, a non-linear term, and a bias term) and the template region used for training (6 lines above and to the left of the current block, depending on their availability) can also be applied to the FIBC mode and the FTMP mode.
[0432] Figure 30 A workflow of a method 3000 for video decoding according to one or more aspects of the disclosure is shown.
[0433] At step 3010, the method 3000 includes determining a reference block in a reconstructed portion of a video frame for predicting a current block in the video frame, wherein an L-shaped template associated with the reference block is a most similar template to an L-shaped template associated with the current block in the reconstructed portion of the video frame.
[0434] At step 3020, the method 3000 includes obtaining a set of filter coefficients corresponding to a filter shape based on sample values from both a training region associated with the reference block and a training region associated with the current block.
[0435] At step 3030, the method 3000 includes deriving, with the set of filter coefficients and the filter shape, predicted sample values of the current block based on a plurality of corresponding sample values associated with the reference block.
[0436] At step 3040, the method 3000 includes reconstructing the current block based on the predicted sample values.
[0437] Figure 31 A workflow of a method 3100 for video encoding in accordance with one or more aspects of the present disclosure is shown.
[0438] At step 3110, the method 3100 includes partitioning a video frame into a plurality of blocks.
[0439] At step 3120, the method 3100 includes determining a reference block in a reconstructed portion of the video frame for predicting a current block in the video frame, wherein an L-shaped template associated with the reference block is a most similar template to an L-shaped template associated with the current block in the reconstructed portion of the video frame.
[0440] At step 3130, the method 3100 includes obtaining a set of filter coefficients corresponding to a filter shape based on sample values from both a training region associated with the reference block and a training region associated with the current block.
[0441] At step 3140, the method 3100 includes deriving, with the set of filter coefficients and the filter shape, predicted sample values of the current block based on a plurality of corresponding sample values associated with the reference block.
[0442] At step 3150, the method 3100 includes generating a bitstream based on the predicted sample values.
[0443] Figure 32 A workflow of a method 3200 for video decoding in accordance with one or more aspects of the present disclosure is shown.
[0444] At step 3210, the method 3200 includes determining at least one of a filter shape and a template region, wherein the at least one of the filter shape and the template region is to be used in at least two of a filtered intra block copy (FIBC) mode, a filtered template matching prediction (FTMP) mode, and a convolution cross component model (CCCM) mode for predicting sample values of a current block in a video frame.
[0445] At step 3220, the method 3200 includes training each filter coefficient of a set of filter coefficients corresponding to a filter shape for at least two of the FIBC mode, the FTMP mode, and the CCCM mode, respectively, with a template region.
[0446] At step 3230, the method 3200 includes deriving prediction sample values for the current block with the set of filter coefficients.
[0447] At step 3240, the method 3200 includes reconstructing the current block based on the prediction sample values.
[0448] In one example, training the set of filter coefficients for the FIBC mode includes determining a reference block in the reconstructed portion of the video frame for predicting the current block, wherein the reference block is determined based on a block vector received from the bitstream, and obtaining the set of filter coefficients for the FIBC mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined based at least in part on the template region for the FIBC mode.
[0449] In one example, training the set of filter coefficients for the FTMP mode includes determining a reference block in the reconstructed portion of the video frame for predicting the current block, wherein an L-shaped template associated with the reference block is the most similar template to the L-shaped template associated with the current block in the reconstructed portion of the video frame, and obtaining the set of filter coefficients for the FTMP mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined based at least in part on the template region for the FTMP mode.
[0450] In one example, training the set of filter coefficients for the CCCM mode includes determining a set of chroma sample values in the template region, and obtaining the set of filter coefficients for the CCCM mode based on the set of chroma sample values and corresponding luma sample values for the set of chroma sample values.
[0451] In one example, the filter shape includes at least one of: a cross-shaped filter shape corresponding to 5 spatial terms, a non-linear term, and a bias term; a single-sample filter shape corresponding to 1 spatial term and a bias term; or a 3 3 square filter shape.
[0452] In one example, the template region includes at least one of: 4 rows above and to the left of the current block; 1 row above and to the left of the current block; or 6 rows above and to the left of the current block.
[0453] Figure 33 A workflow of a method 3300 for video coding is shown in accordance with one or more aspects of the present disclosure.
[0454] At step 3310, the method 3300 includes partitioning a video frame into a plurality of blocks.
[0455] At step 3320, the method 3300 includes determining at least one of a filter shape and a template region, wherein the at least one of the filter shape and the template region is to be used in at least two of a filtered intra block copy (FIBC) mode, a filtered template matching prediction (FTMP) mode, and a convolution cross component model (CCCM) mode for predicting sample values of a current block in the video frame.
[0456] At step 3330, the method 3300 includes training each filter coefficient of a set of filter coefficients corresponding to the filter shape for the at least two of the FIBC mode, the FTMP mode, and the CCCM mode, respectively, with the template region.
[0457] At step 3340, the method 3300 includes deriving predicted sample values of the current block with the set of filter coefficients.
[0458] At step 3350, the method 3300 includes generating a bitstream based on the predicted sample values.
[0459] In one example, training the set of filter coefficients for the FIBC mode includes determining a reference block in a reconstructed portion of the video frame for predicting the current block, wherein the reference block is determined based on a block vector to be transmitted via the bitstream, and obtaining the set of filter coefficients for the FIBC mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined based at least in part on the template region for the FIBC mode.
[0460] In one example, training the set of filter coefficients for the FTMP mode includes determining a reference block in a reconstructed portion of the video frame for predicting the current block, wherein an L-shaped template associated with the reference block is a most similar template to an L-shaped template associated with the current block in the reconstructed portion of the video frame, and obtaining the set of filter coefficients for the FTMP mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined based at least in part on the template region for the FTMP mode.
[0461] In one example, training the set of filter coefficients for the CCCM mode includes determining a set of chroma sample values in the template region; and obtaining the set of filter coefficients for the CCCM mode based on the set of chroma sample values and corresponding luma sample values for the set of chroma sample values.
[0462] In one example, the filter shape includes at least one of: a cross filter shape corresponding to 5 spatial terms, a non-linear term, and a bias term; a single sample filter shape corresponding to 1 spatial term and a bias term; or a 3 3 square filter shape.
[0463] In one example, the template region includes at least one of: 4 lines above and to the left of the current block; 1 line above and to the left of the current block; or 6 lines above and to the left of the current block.
[0464] Reference region and padding process in filtered intra block copy (FIBC)
[0465] According to one or more embodiments of the disclosure, the filter coefficients are calculated by minimizing the MSE between the predicted and reconstructed luma and / or chroma samples in the reference region. In one example, Figure 25 A reference region is shown including luma / chroma samples above and to the left of the CU. An extension of the region shown in blue (diagonal portion) is needed to support the "side samples" of the plus spatial filter, and different approaches can be used to achieve this goal.
[0466] In a first approach, it is proposed to pad with the closest available sample when unavailable.
[0467] In a second approach, it is proposed to pad with the closest available sample regardless of whether it is available or not.
[0468] According to one or more embodiments of the disclosure, the reference region can be extended one CU width to the right and one CU height below the CU boundary.
[0469] According to one or more embodiments of the disclosure, the reference region can be adjusted to include only available samples. In one example, the reference region includes N lines of luma / chroma samples above and to the left of the CU. N is an integer and / or has a maximum upper limit (e.g., 4, 5, 6, 7).
[0470] Adaptive reordering of merge candidates with filtered intra block copy (FIBC)
[0471] According to one or more embodiments of the disclosure, in the process of extending ARM C-TM to IBC merge list, when the merge candidate utilizes FIBC prediction, the reference samples of the template of the merge candidate are also generated by FIBC.
[0472] According to one or more embodiments of the disclosure, the filter coefficients are calculated by minimizing the MSE between the predicted and reconstructed luma and / or chroma samples in the reference region.
[0473] According to one or more embodiments of the disclosure, the template size and shape can not be included in the reference samples of the template of the merge candidate. In one example, Figure 34 A reference region including 3 rows of luma samples above and to the left of the CU is shown.
[0474] Merge candidate with filtered intra block copy
[0475] According to one or more embodiments of the disclosure, the IBC prediction from the merge candidate is further filtered. Different methods can be used to achieve this goal.
[0476] According to one or more embodiments of the disclosure, an enabling flag can be signaled in the bitstream to indicate the FIBC merge mode used. The enabling flag can be signaled at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0477] According to one or more embodiments of the disclosure, the mode flag can be derived at the decoder to save bit overhead instead of explicitly signaling the selected mode flag.
[0478] Direct block vector for chroma block with filtered intra block copy
[0479] According to one or more embodiments of the disclosure, the direct block vector for chroma block is further filtered. Different methods can be used to achieve this goal.
[0480] According to one or more embodiments of the disclosure, an enabling flag can be signaled in the bitstream to indicate the FIBC merge mode used. The enabling flag can be signaled at SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0481] According to one or more embodiments of the disclosure, the mode flag can be inherited from the luma block at the decoder to save bit overhead instead of explicitly signaling the selected mode flag.
[0482] Figure 35A workflow of a method 3500 for video decoding according to one or more aspects of the disclosure is shown. The method 3500 can be performed by a decoder (e.g., video decoder 30) of Figure 3 FIG. 1.
[0483] At step 3510, a bitstream including a video frame can be received, and to predict a current block in the video frame, a reference block in the same video frame can be determined. For example, the reference block can be determined by using a block vector indicated in the bitstream for the current block. Here, the block vector is used to indicate a displacement from the current block to the reference block, which has been reconstructed within the current video frame.
[0484] At step 3520, a set of filter coefficients corresponding to a filter shape can be obtained by training sample values of a template region associated with the reference block and sample values of a template region associated with the current block. The template region associated with the reference block can be extended based on the filter shape, for example, to further include sample values of edge points of the filter shape.
[0485] In one example, the template region associated with the reference block including samples to the left and above of the reference block can be extended to further include one or more rows to the right and one or more rows below of the reference block. For example, a row can correspond to a row or column of a picture or frame in which the samples are located.
[0486] In another example, the template region associated with the reference block can extend from the reference block in four directions, i.e., left and right, up and down, to further include additional one or more rows from four sides around the reference block.
[0487] In one or more examples, the extended one or more rows can correspond to one CU width or one CU height. The sample values in the extended rows can be padded with the closest available sample values. In another example, the samples padded with the closest available sample values are unavailable samples, e.g., samples to the right or below of the reference block that have not been reconstructed yet.
[0488] In another example, the template region associated with the reference block can be extended to include only available samples. For example, the template region associated with the reference block is extended to include up to N rows to the left of the reference block and up to N rows above of the reference block, and N is an integer (e.g., 4, 5, 6, 7).
[0489] At step 3530, each of the prediction sample values of the current block can be derived based on a plurality of corresponding sample values associated with the reference block by using the set of filter coefficients and the filter shape. For example, the plurality of corresponding sample values associated with the reference block can be used as input to the obtained filter to output each of the prediction sample values of the current block. The input sample values associated with the reference block can include sample values in the extended template region.
[0490] At step 3540, the current block can be reconstructed based on the prediction sample values. For example, the current block can be reconstructed by combining the prediction sample values with a residual carried in the bitstream.
[0491] Figure 36 A workflow of a method 3600 for video encoding according to one or more aspects of the disclosure is shown. The method 3600 can be performed by an encoder (e.g., the video encoder 20 of Figure 2 FIG. 3). The steps of the method 3600 can be corresponding steps to the steps of the method 3800.
[0492] At step 3610, a reference block in a video frame captured by a camera can be determined for predicting a current block in the same video frame. For example, the determination can be made based on block matching (BM) performed at the encoder.
[0493] At step 3620, a filter having a set of filter coefficients and a filter shape can be obtained by training sample values of a template region associated with the reference block and sample values of a template region associated with the current block, where the template region associated with the reference block is extended based on the filter shape. For example, the template region associated with the reference block is extended in the same manner as described with the reference method 3800.
[0494] At step 3630, each of the prediction sample values of the current block can be derived based on a plurality of corresponding sample values associated with the reference block by using the obtained filter.
[0495] At step 3640, a bitstream can be generated based on the prediction sample values.
[0496] Figure 37 A workflow of a method 3700 for video decoding according to one or more aspects of the disclosure is shown. The method 3700 can be performed by a decoder (e.g., the video decoder 30 of Figure 3 FIG. 3).
[0497] At step 3710, a merge candidate list for intra block copy (IBC) prediction of the current block can be obtained. For example, the merge candidate list for IBC prediction can be constructed according to the description above or other criteria. The merge candidate list can include a plurality of candidates that have been encoded with IBC, each candidate having, for example, a block vector.
[0498] At step 3720, the plurality of candidates in the merge candidate list can be reordered based on a template matching score of each of the plurality of candidates, the template matching score being computed based on a difference between sample values of a template of a candidate and corresponding reference sample values of a reference template of a reference block pointed to by a block vector of the candidate. In the computation of the template matching score, in response to determining that the candidate is coded with filtered IBC (i.e., FIBC), the corresponding reference sample values of the reference template are also obtained using filtered IBC (i.e., FIBC).
[0499] At step 3730, the current block can be reconstructed based on the reordered merge candidate list.
[0500] In one example, the IBC prediction of the current block can be obtained from a candidate of the merge candidate list, for example, by using the block vector of the candidate. A determination is made as to whether the IBC prediction of the current block is filtered.
[0501] In one example, the determination is made based on a syntax element transmitted in the bitstream or an inference derived at the decoder. For example, the determination that the current block is reconstructed by using filtered IBC can be made based on the syntax element or based on the inference without the syntax element. In response to determining that the current block is reconstructed by using filtered IBC, FIBC is performed on the IBC prediction.
[0502] In one example, at least a portion of the template and at least a portion of the reference template are not used to obtain the template matching score. For example, as shown by the grid in Figure 34 , the samples that are immediately adjacent to the reference block and the candidate are not used to compute the template matching score.
[0503] Figure 38 A workflow of a method 3800 for video encoding according to one or more aspects of the disclosure is shown. The method 3800 can be performed by an encoder, for example, the video encoder 20 of FIG. 2. The steps of the method 3800 can be corresponding steps to the steps of the method 3700. Figure 2
[0504] At step 3810, a merge candidate list for intra block copy (IBC) prediction of the current block can be obtained. For example, the merge candidate list for IBC prediction can be constructed according to the description above or other criteria. The merge candidate list can include a plurality of candidates that have been coded with IBC, each candidate having, for example, a block vector.
[0505] At step 3820, the plurality of candidates in the merge candidate list can be reordered based on a template matching score of each of the plurality of candidates, the template matching score being computed based on a difference between sample values of a template of a candidate and corresponding reference sample values of a reference template of a reference block pointed to by a block vector of the candidate. In the computation of the template matching score, in response to determining that a candidate is coded with filtered IBC (i.e., FIBC), the corresponding reference sample values of the reference template are also obtained using filtered IBC (i.e., FIBC).
[0506] At step 3830, a bitstream can be generated by encoding the current block based on the reordered merge candidate list.
[0507] In one example, an IBC prediction of the current block can be obtained from a candidate of the merge candidate list, for example, by using a block vector of the candidate. A determination is made as to whether the IBC prediction of the current block is filtered. If a determination is made that the IBC prediction of the current block is to be filtered, FIBC is performed on the IBC prediction.
[0508] In one example, a syntax element indicating the determination that the IBC prediction of the current block is to be filtered can be transmitted in the bitstream.
[0509] Figure 39 A workflow of a method 3900 for video decoding in accordance with one or more aspects of the present disclosure is shown. The method 3900 can be performed by a decoder (e.g., video decoder 30 of FIG. 3) in accordance with one or more aspects of the present disclosure. Figure 3
[0510] At step 3910, a block vector of a chroma block from a bitstream can be determined by using a block vector of a luma block associated with the chroma block, wherein the luma block has been coded with IBC. By way of example, the block vector of the chroma block can be directly inherited from the luma block.
[0511] At step 3920, an IBC prediction of the chroma block can be obtained based on the inherited block vector.
[0512] At step 3930, in response to determining that a filtered IBC prediction is used for the chroma block, the IBC prediction of the chroma block can be filtered to obtain a filtered IBC prediction, i.e., using FIBC.
[0513] In one example, the determination that filtered IBC prediction is used for the chroma block can be made based on a syntax element in the bitstream.
[0514] In another example, the mode to be used by the chroma block (e.g., FIBC or IBC) can be directly inherited from the mode used by the associated luma block (e.g., FIBC or IBC), such that explicit signaling can be saved.
[0515] At step 3940, the chroma block can be reconstructed based on the filtered IBC prediction.
[0516] In one example, the filter shape and / or filter coefficients for the chroma block can be calculated by using a template of the chroma block at the decoder (e.g., in a similar manner as for the luma block).
[0517] In another example, the filter shape and / or filter coefficients for the chroma block can be directly inherited from the filter shape and / or filter coefficients of the associated luma block.
[0518] Figure 40 A workflow of a method 4000 for video encoding according to one or more aspects of the disclosure is shown. The method 4000 can be performed by an encoder (e.g., the video encoder 20 of FIG. 1). Figure 2 The steps of the method 4000 can be corresponding steps to the steps of the method 4200.
[0519] At step 4010, a block vector for a chroma block can be determined based on a luma block associated with the chroma block, wherein the luma block has been coded with intra block copy (IBC).
[0520] At step 4020, an IBC prediction for the chroma block can be obtained based on the block vector.
[0521] At step 4030, the IBC prediction for the chroma block can be filtered.
[0522] In one example, a syntax element indicating that the IBC prediction for the chroma block is to be filtered can be generated.
[0523] In another example, a FIBC mode used by the associated luma block indicating that the IBC prediction for the chroma block is to be filtered can be generated.
[0524] At step 4040, a bitstream can be generated based on the filtered IBC prediction.
[0525] In one example, a syntax element indicating that the IBC prediction for the chroma block is to be filtered can be transmitted in the bitstream.
[0526] In another example, if a syntax element indicating that IBC prediction of a chroma block is to be filtered can be inherited directly from the mode used by the associated luma block, then this signaling can not be transmitted in the bitstream.
[0527] Figure 41 A computing environment 4110 coupled with a user interface 4150 is shown. The computing environment 4110 can be part of a data processing server. The computing environment 4110 includes a processor 4120, a memory 4130, and an input / output (I / O) interface 4140.
[0528] The processor 4120 generally controls the overall operation of the computing environment 4110, such as operations associated with displaying, data acquisition, data communication, and image processing. The processor 4120 can include one or more processors to execute instructions to perform all or some of the steps in the above-described methods. In addition, the processor 4120 can include one or more modules that facilitate interaction between the processor 4120 and other components. The processor can be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.
[0529] The memory 4130 is configured to store various types of data to support the operation of the computing environment 4110. The memory 4130 can include predetermined software 4132. Examples of such data include any application or method instructions for operation on the computing environment 4110, video data sets, image data, etc. The memory 4130 can be implemented by using any type of volatile or non-volatile memory devices, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0530] The I / O interface 4140 provides an interface between the processor 4120 and peripheral interface modules (e.g., a keyboard, a click wheel, a button, etc.). The buttons can include, but are not limited to, a home button, a start scanning button, and a stop scanning button. The I / O interface 4140 can be coupled with an encoder and a decoder.
[0531] In an embodiment, a non-transitory computer-readable storage medium including, for example, a plurality of programs in the memory 4130 and / or storing a bitstream generated by the above-described encoding method or to be decoded by the above-described decoding method, the plurality of programs can be executed by the processor 4120 in the computing environment 4110 for performing the above-described methods. In one example, the plurality of programs can be executed by the processor 4120 in the computing environment 4110 to, for example, (e.g., from Figure 2The video encoder 20 in the computing environment 4110 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 4120 in the computing environment 4110 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 4120 in the computing environment 4110 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 4120 in the computing environment 4110 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.
[0532] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.
[0533] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 4120); and a non-transitory computer-readable storage medium or memory 4130 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0534] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 4130 is also provided, the plurality of programs being executable by a processor 4120 in a computing environment 4110 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0535] In an embodiment, the computing environment 4110 can be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, micro-controllers, microprocessors, or other electronic components for performing the above-described methods.
[0536] In an embodiment, a method of storing a bitstream is also provided, including: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-mentioned encoding method or encoded video information to be decoded by the above-mentioned decoding method.
[0537] In an embodiment, a method for transmitting a bitstream generated by the above-mentioned encoder is also provided. In an embodiment, a method for receiving a bitstream to be decoded by the above-mentioned decoder is also provided.
[0538] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosure as set forth in the above description and the associated drawings. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art based on the teachings from the above description and associated drawings.
[0539] Unless otherwise specified, the order of steps of the method according to the present disclosure is only intended to be illustrative, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to the actual situation. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined or deleted according to actual needs.
[0540] Examples are chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to best utilize the various embodiments of the present disclosure, and to best utilize the various embodiments for the intended purposes and with various modifications as are suited to the particular use or implementation. Accordingly, it is to be understood that the scope of the present disclosure is not to be limited to the specific examples disclosed and that modifications and other embodiments are intended to be included within the scope of the present disclosure.
Claims
1. A method for decoding video data, comprising: A reference block is determined in a video frame from a bitstream to be used to predict the current block in the video frame; A set of filter coefficients corresponding to the filter shape is obtained based on sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the template region associated with the reference block is extended based on the filter shape; Using the filter coefficient set and the filter shape, each of the predicted sample values for the current block is derived based on multiple corresponding sample values associated with the reference block; and The current block is reconstructed based on the predicted sample values.
2. The method according to claim 1, wherein, The template region associated with the reference block is expanded to include one or more rows to the right of the reference block and one or more rows below the reference block.
3. The method according to claim 2, wherein, The template region associated with the reference block is expanded to obtain an expanded template region that further includes one or more rows to the left of the reference block and one or more rows above the reference block.
4. The method according to claim 3, wherein, The samples in the expanded template area are filled with the closest available sample values.
5. The method according to claim 4, wherein, The sample point that is filled with the closest available sample point value is an unavailable sample point.
6. The method according to claim 1, wherein, The template region associated with the reference block is expanded to include up to N rows to the left of the reference block and up to N rows above the reference block, where N is an integer.
7. The method according to claim 1, wherein, The plurality of corresponding sample values associated with the reference block include sample values in the extended template region.
8. A method for decoding video data, comprising: Obtain a merge candidate list of intra-block copy (IBC) predictions for the current block, the merge candidate list including multiple candidates that have been encoded using IBC; Based on the template matching score of each candidate in the merged candidate list, the candidates in the merged candidate list are reordered. The template matching score is obtained based on the difference between the sample value of a candidate's template and the corresponding reference sample value of the reference template of the candidate's reference block. This is done in response to determining that a candidate is encoded using filtered IBC, the corresponding reference sample value of the reference template is obtained using filtered IBC. The current block is reconstructed based on the reordered list of merge candidates.
9. The method according to claim 8, further comprising: In response to determining that the current block was reconstructed using filtered IBC, the IBC prediction of the current block based on the reordered merge candidate list is filtered.
10. The method of claim 9, further comprising: Determine whether to reconstruct the current block using filtered IBC based on syntax elements or, in the absence of syntax elements, based on reasoning.
11. The method according to claim 8, wherein, At least a portion of the template and at least a portion of the reference template are not used to obtain the template matching score.
12. A method for decoding video data, comprising: A block vector for the chroma block is determined based on the luma block associated with the chroma block from the bitstream, wherein the luma block is encoded using intra-block copy (IBC); IBC prediction of the chroma block is obtained based on the block vector; In response to determining that the filtered IBC prediction is used for the chroma block, the IBC prediction of the chroma block is filtered to obtain the filtered IBC prediction; and The chroma block is reconstructed based on the filtered IBC prediction.
13. The method of claim 12, further comprising: Based on syntax elements or the patterns used by the luminance blocks associated with the chroma blocks, filtered IBC predictions are determined for the chroma blocks.
14. A method for encoding video data, comprising: A reference block in the video frame is determined for use in predicting the current block in the video frame; A set of filter coefficients corresponding to the filter shape is obtained based on sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the template region associated with the reference block is extended based on the filter shape; Using the filter coefficient set and the filter shape, each of the predicted sample values for the current block is derived based on multiple corresponding sample values associated with the reference block; and A bitstream is generated based on the predicted sample values.
15. The method according to claim 14, wherein, The template region associated with the reference block is expanded to include one or more rows to the right of the reference block and one or more rows below the reference block.
16. The method according to claim 15, wherein, The template region associated with the reference block is expanded to obtain an expanded template region that further includes one or more rows to the left of the reference block and one or more rows above the reference block.
17. The method according to claim 16, wherein, The samples in the expanded template area are filled with the closest available sample values.
18. The method according to claim 17, wherein, The sample point that is filled with the closest available sample point value is an unavailable sample point.
19. The method of claim 14, wherein, The template region associated with the reference block is expanded to include up to N rows to the left of the reference block and up to N rows above the reference block, where N is an integer.
20. The method of claim 14, wherein, The plurality of corresponding sample values associated with the reference block include sample values in the extended template region.
21. A method for encoding video data, comprising: Obtain a merge candidate list of intra-block copy (IBC) predictions for the current block, the merge candidate list including multiple candidates that have been encoded using IBC; Based on the template matching score of each candidate in the merged candidate list, the candidates in the merged candidate list are reordered. The template matching score is obtained based on the difference between the sample value of the candidate's template and the corresponding reference sample value of the reference template of the candidate's reference block. In response to determining that the candidate is encoded using filtered IBC, the corresponding reference sample value of the reference template is obtained using filtered IBC. as well as A bitstream is generated by encoding the current block based on a reordered list of merge candidates.
22. The method of claim 21, further comprising: The IBC prediction of the current block is filtered according to the reordered list of merge candidates.
23. The method of claim 22, further comprising: Syntax elements are sent in the bitstream to indicate that the filtered IBC is used for the current block.
24. The method according to claim 21, wherein, At least a portion of the template and at least a portion of the reference template are not used to obtain the template matching score.
25. A method for encoding video data, comprising: The block vector for the chroma block is determined based on the luma block associated with the chroma block, wherein the luma block is encoded using intra-block copy (IBC); IBC prediction of the chroma block is obtained based on the block vector; Filter the IBC prediction of the chroma block; and The bitstream is generated based on the filtered IBC prediction.
26. The method of claim 25, further comprising: Syntax elements for indicating the use of filtered IBCs for the chroma blocks are transmitted in the bitstream.
27. An apparatus comprising: One or more processors; as well as One or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operation of the method described in any one of claims 1-26.
28. A computer program product storing computer-executable instructions, which, when executed, cause one or more processors to perform the operations of the method described in any one of claims 1-26.
29. A computer-readable storage medium storing instructions, which, when executed by a computing device having one or more processors, cause the one or more processors to perform the following operations: Perform the decoding method according to any one of claims 1-13, and store the bitstream to be decoded by the decoding method according to any one of claims 1-13.
30. A computer-readable storage medium storing instructions, which, when executed by a computing device having one or more processors, cause the one or more processors to perform the following operations: Perform the encoding method according to any one of claims 14-26, and store the bit stream generated by the encoding method according to any one of claims 14-26.
31. A computer-readable medium storing a bitstream, wherein, The bitstream will be decoded by performing the operation of the method described in any one of claims 1-13.
32. A computer-readable medium storing a bitstream, wherein, The bitstream is obtained by performing the method described in any one of claims 14-26.
33. A method for receiving a bitstream to be decoded by the decoding method according to any one of claims 1-13.
34. A method for transmitting a bit stream generated by the encoding method according to any one of claims 14-26.