Method and apparatus for intra block copy and intra template matching
By introducing a combination mode in video encoding, combining intra-template matching prediction and other prediction modes, the problem of increasing the amount of motion vector representation data in high-detail video data is solved, which improves the encoding efficiency, especially in screen content encoding, which significantly improves the encoding efficiency.
Patent Information
- Application Number
- CN202380079744.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-20
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing video encoding and decoding technology processes high-detail video data, the amount of data represented by motion vectors has increased significantly, resulting in a decrease in encoding efficiency, and the efficiency improvement of intra-block copy mode in screen content encoding is limited.
A combination mode is introduced, combining intra-template matching prediction (TMP) with other prediction modes, such as inter-frame prediction, local lighting compensation (LIC) and overlapping block motion compensation (OBMC), to optimize the prediction of the current encoding unit through the combination mode and improve coding efficiency.
Through combined mode optimization, the amount of data represented by motion vectors is reduced, and the efficiency of video encoding is improved, especially in screen content encoding, which is significantly improved.
Smart Images

Figure CN120303935A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 426,712, filed on November 18, 2022, entitled "Methods and Devices For Intra Block Copy and Intra Template Matching", the entire content of which is incorporated herein by reference for all purposes. Technical Field
[0003] The present disclosure relates to video encoding and decoding and compression, and more specifically but not limited to, methods and apparatuses for improving the encoding and decoding efficiency of intra - block copy (IBC) and intra - template matching prediction (intra - TMP). Background Art
[0004] Various electronic devices (such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc.) support digital video. Electronic devices send and receive or otherwise transmit digital video data via a communication network and / or store digital video data on a storage device. Since the bandwidth capacity of the communication network is limited and the storage resources of the storage device are limited, video data can be compressed using video encoding and decoding according to one or more video encoding and decoding standards before the video data is transmitted or stored. For example, video encoding and decoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Experts Group (MPEG) coding, etc. Video encoding and decoding typically employs prediction methods (such as inter - frame prediction, intra - frame prediction, etc.) that utilize the redundancy inherent in video data. Video encoding and decoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. Summary of the Invention
[0005] The present disclosure provides examples of techniques related to improving the intra - block copy method in a video encoding or decoding process.
[0006] According to a first aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain a current coding unit (CU) encoded based on a combined mode, where the combined mode combines at least one of the following modes with an intra-template matching prediction (TMP) mode: an intra prediction mode, an inter prediction mode, an intra mode derivation based on a template (TIMD) mode, a local illumination compensation (LIC) mode, or an overlapped block motion compensation (OBMC) mode. Additionally, the decoder may obtain a final prediction for the current CU based on the combined mode.
[0007] According to a second aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may encode a current CU based on a combined mode, where the combined mode combines at least one of the following modes with an intra TMP mode: an intra prediction mode, an inter prediction mode, a TIMD mode, a LIC mode, or an OBMC mode. Additionally, the encoder may send the current CU encoded based on the combined mode to a decoder.
[0008] According to a third aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain a current CU encoded based on an intra TMP mode combined with a GPM mode. Additionally, the decoder may obtain a final prediction for the current CU based on the intra TMP mode combined with the GPM mode.
[0009] According to a fourth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may encode a current CU based on an intra TMP mode combined with a GPM mode. Additionally, the encoder may send the current CU encoded based on the intra TMP mode combined with the GPM mode to a decoder.
[0010] According to a fifth aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain a current CU encoded based on a combined mode, where the combined mode combines at least one of a TIMD mode or an OBMC mode with an IBC mode. Additionally, the decoder may obtain a final prediction for the current CU based on the combined mode.
[0011] According to a sixth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may encode a current CU based on a combined mode, where the combined mode combines at least one of a TIMD mode or an OBMC mode with an IBC mode. Additionally, the encoder may send the current CU encoded based on the combined mode to a decoder.
[0012] According to a seventh aspect of the present disclosure, there is provided an apparatus for video decoding. The apparatus may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. Additionally, the one or more processors, when executing the instructions, are configured to perform the method according to the first aspect, the third aspect, or the fifth aspect described above.
[0013] According to an eighth aspect of the present disclosure, there is provided an apparatus for video encoding. The apparatus may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. Additionally, the one or more processors, when executing the instructions, are configured to perform the method according to the second aspect, the fourth aspect, or the sixth aspect described above.
[0014] According to a ninth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions, when executed by one or more computer processors, causing the one or more computer processors to perform the method according to the first aspect, the third aspect, or the fifth aspect described above.
[0015] According to a tenth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions, when executed by one or more computer processors, causing the one or more computer processors to perform the method according to the second aspect, the fourth aspect, or the sixth aspect described above.
[0016] According to an eleventh aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to the first aspect, the third aspect, or the fifth aspect.
[0017] According to a twelfth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing a bitstream generated by the method according to the second aspect, the fourth aspect, or the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] A more specific description of the examples of the present disclosure will be presented by referring to specific examples shown in the accompanying drawings. Given that these drawings only depict some examples and are thus not considered to be a limitation of the scope, the examples will be described and explained with additional features and details by using the drawings.
[0019] Figure 1A is a block diagram showing a system for encoding and decoding video blocks according to some examples of the present disclosure.
[0020] Figure 1Bshows a quadtree data structure showing the final result of the partitioning process of CTU 400 as depicted in Figure 1E the following.
[0021] Figure 1C shows an encoded representation of a frame by first partitioning the frame into a group of CTUs according to some examples of the present disclosure.
[0022] Figure 1D shows a CTU including one CTB of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements for encoding and decoding the samples of the coding tree blocks according to some examples of the present disclosure.
[0023] Figure 2 is a block diagram showing an exemplary video encoder according to some examples of the present disclosure.
[0024] Figure 3 is a block diagram showing an exemplary video decoder according to some examples of the present disclosure.
[0025] Figure 4A is a diagram showing block partitioning in a multi-type tree structure according to some examples of the present disclosure.
[0026] Figure 4B is a diagram showing block partitioning in a multi-type tree structure according to some examples of the present disclosure.
[0027] Figure 4C is a diagram showing block partitioning in a multi-type tree structure according to some examples of the present disclosure.
[0028] Figure 4D is a diagram showing block partitioning in a multi-type tree structure according to some examples of the present disclosure.
[0029] Figure 4E is a diagram showing block partitioning in a multi-type tree structure according to some examples of the present disclosure.
[0030] Figure 5 is a diagram showing the positions of spatial candidates according to some examples of the present disclosure.
[0031] Figure 6 is a diagram showing candidate pairs considering redundancy checks for spatial candidates according to some examples of the present disclosure.
[0032] Figure 7 is a diagram showing scaling of motion vectors for temporal candidates according to some examples of the present disclosure.
[0033] Figure 8 is a diagram showing candidate positions for temporal candidates according to some examples of the present disclosure.
[0034] Figure 9 A diagram showing merge mode with motion vector difference (MMVD) search points according to some examples of the present disclosure.
[0035] Figure 10 Showing unidirectional prediction motion vector selection for geometric partitioning mode (GPM) according to some examples of the present disclosure.
[0036] Figure 11 Showing top neighboring blocks and left neighboring blocks in CIIP weight derivation according to some examples of the present disclosure.
[0037] Figure 12 Showing the current CTU processing order and available reference sample points in the current CTU and the left CTU according to some examples of the present disclosure.
[0038] Figure 13 Showing padding candidates for replacing zero vectors in the IBC list according to some examples of the present disclosure.
[0039] Figure 14 Showing the reference region for IBC when CTU (m, n) is encoded according to some examples of the present disclosure.
[0040] Figure 15 Showing the IBC reference region for camera-captured content according to some examples of the present disclosure.
[0041] Figure 16A - Figure 16B Showing a partitioning method for an angular mode according to some examples of the present disclosure.
[0042] Figure 17A - Figure 17D Showing GPM with inter-frame prediction and intra-frame prediction according to some examples of the present disclosure.
[0043] Figure 18 Showing edges on a template according to some examples of the present disclosure.
[0044] Figure 19 Showing the intra-frame template matching search region used according to some examples of the present disclosure.
[0045] Figure 20 Showing a template for template matching-based OBMC according to some examples of the present disclosure.
[0046] Figure 21 A diagram showing a computing environment coupled to a user interface according to some embodiments of the present disclosure.
[0047] Figure 22 A flowchart showing a method for video decoding according to some examples of the present disclosure.
[0048] Figure 23 is a flowchart showing a method for video encoding corresponding to the method for video decoding as shown in Figure 22 the following with respect to some examples of the present disclosure.
[0049] Figure 24 is a flowchart showing a method for video decoding according to some examples of the present disclosure.
[0050] Figure 25 is a flowchart showing a method for video encoding corresponding to the method for video decoding as shown in Figure 24 the following with respect to some examples of the present disclosure.
[0051] Figure 26 is a flowchart showing a method for video decoding according to some examples of the present disclosure.
[0052] Figure 27 is a flowchart showing a method for video encoding corresponding to the method for video decoding as shown in Figure 26 the following with respect to some examples of the present disclosure. DETAILED DESCRIPTION
[0053] Reference will now be made in detail to the detailed description, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, various alternative solutions may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.
[0054] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. In the present disclosure and the appended claims, the singular forms (such as "a / an", "the", etc.) are also intended to include the plural forms unless clearly stated otherwise throughout the disclosure. It should also be understood that the term "and / or" used in the present disclosure indicates and includes any or all possible combinations of the listed multiple related items.
[0055] References throughout this specification to "one embodiment", "an embodiment", "an example", "some embodiments", "some examples" or similar language mean that the particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise clearly stated, features, structures, elements, or characteristics described in connection with one or some embodiments are also applicable to other embodiments.
[0056] Throughout the disclosure, unless otherwise expressly stated, the terms "first", "second", "third", etc. are used only as names for references to related elements (e.g., devices, components, compositions, steps, etc.) and do not imply any spatial or temporal order. For example, "a first device" and "a second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be named arbitrarily.
[0057] The terms "module", "sub-module", "circuit", "sub-circuit", "circuitry", "sub-circuitry", "unit", or "sub-unit" may include a memory (shared, dedicated, or grouped) that stores code or instructions executable by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components connected directly or indirectly. These components may or may not be physically attached to each other or located adjacent to each other.
[0058] As used herein, depending on the context, the terms "if" or "when" may be understood to mean "based on" or "in response to". If these terms appear in a claim, they may not indicate that the related limitation or feature is conditional or optional. For example, a method may include the following steps: i) performing function or action X' when or if condition X exists, and ii) performing function or action Y' when or if condition Y exists. The method may be implemented using the ability to perform function or action X' and the ability to perform function or action Y'. Thus, both function X' and Y' may be performed at different times during multiple executions of the method.
[0059] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components linked directly or indirectly to perform a specific function.
[0060] Figure 1A is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. As Figure 1A shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a destination device 14. The source device 12 and the destination device 14 may include any of a variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.
[0061] In some embodiments, the target device 14 may receive encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In one example, the link 16 may include a communication medium that enables the source device 12 to send the encoded video data directly to the target device 14 in real time. The encoded video data may be modulated according to a communication standard (such as a wireless communication protocol) and sent to the target device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include routers, switches, base stations, or any other device that may facilitate communication from the source device 12 to the target device 14.
[0062] In some other embodiments, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the target device 14 via the input interface 28. The storage device 32 may include any data storage medium in a variety of distributed or local access data storage media, such as a hard disk drive, a Blu-ray disc, a Digital Versatile Disk (DVD), a Compact Disc Read-Only Memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device that can store the encoded video data generated by the source device 12. The target device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and sending the encoded video data to the target device 14. Exemplary file servers include web servers (e.g., for websites), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, or local disk drives. The target device 14 may access the encoded video data through any standard data connection suitable for accessing the encoded video data stored on the file server, and the standard data connection includes a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from the storage device 32 may be streaming, download transmission, or a combination of both streaming and download transmission.
[0063] As Figure 1A shown, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include a source such as the following or a combination of such sources: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video. As an example, if the video source 18 is a camera of a security monitoring system, the source device 12 and the target device 14 may form a camera phone or a video phone. However, the embodiments described in the present application are generally applicable to video coding and decoding and may be applied to wireless and / or wired applications.
[0064] Video encoder 20 can encode captured, pre-captured, or computer-generated video. The encoded video data can be sent directly to target device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 can also include a modem and / or a transmitter.
[0065] Target device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 can include a receiver and / or a modem, and receives the encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 can include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements can be included within the encoded video data sent on a communication medium, stored on a storage medium, or stored on a file server.
[0066] In some embodiments, target device 14 can include display device 34, which can be an integrated display device and an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to the user, and can include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0067] Video encoder 20 and video decoder 30 can operate according to proprietary standards or industry standards (e.g., VVC, HEVC, MPEG-4, Part 10, AVC) or extensions of such standards. It should be understood that the present application is not limited to specific video coding / decoding standards and can be applicable to other video coding / decoding standards. It is generally considered that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally considered that video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0068] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.
[0069] In some embodiments, at least some components of the source device 12 (e.g., the video source 18, the video encoder 20, or components included in the video encoder 20 as described below with reference to FIG. 1G, and the output interface 22) and / or at least some components of the destination device 14 (e.g., the input interface 28, the video decoder 30, or components included in the video decoder 30 as described below with reference Figure 3The described components, as well as the display device 34), may operate in a cloud computing service network, which may provide software, platform, and / or infrastructure, such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS). In some embodiments, one or more components not included in the cloud computing service network in the source device 12 and / or the target device 14 may be provided in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In an embodiment, at least a part of the operations described herein may be implemented as a cloud-based service provided by one or more server computers in the cloud computing service network, and the one or more server computers are implemented by at least part of the components of the source device 12 and / or at least part of the components of the target device 14; one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Terms such as "cloud", "cloud computing", "cloud-based", etc. in this document may be used interchangeably adaptively without departing from the scope of the present disclosure. It should be understood that the present disclosure is not limited to being implemented in the above cloud computing service network. Alternatively, the present disclosure may also be implemented in any other type of computing environment known currently or developed in the future.
[0070] Figure 4A - Figure 4E is a schematic diagram showing a multi-type tree partitioning pattern according to some embodiments of the present disclosure. Figure 4A - Figure 4E respectively show five partitioning types, including quaternary partitioning ( Figure 4A ), vertical binary partitioning ( Figure 4B ), horizontal binary partitioning ( Figure 4C ), vertical ternary partitioning ( Figure 4D ), and horizontal ternary partitioning ( Figure 4E ).
[0071] Figure 2 is a block diagram showing another exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 may perform intra prediction coding and inter prediction coding on video blocks within a video frame. Intra prediction coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter prediction coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within adjacent video frames or pictures of a video sequence. It should be noted that the term "frame" may be used as a synonym for the term "image" or "picture" in the field of video coding and decoding.
[0072] AsFigure 2 As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a Decoded Picture Buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and a Block Copy (BC) unit 48. In some embodiments, the video encoder 20 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63 (such as a deblocking filter) may be located between the adder 62 and the DPB 64 to filter block boundaries to remove block effect artifacts from the reconstructed video. In addition to the deblocking filter that may be used, other loop filters such as a Sample Adaptive Offset (SAO) filter, a Cross Component Sample Adaptive Offset (CCSAO) filter, and / or an Adaptive in-Loop Filter (ALF) may be used to filter the output of the adder 62. It should be noted that for the CCSAO technology, this application is not limited to the embodiments described herein, and alternatively, this application can be applied to the following cases: an offset for any one of the luminance component, the Cb chrominance component, and the Cr chrominance component is selected according to any other one of the luminance component, the Cb chrominance component, and the Cr chrominance component, and the any other component is modified based on the selected offset. In addition, it should also be noted that the first component mentioned herein can be any one of the luminance component, the Cb chrominance component, and the Cr chrominance component, the second component mentioned herein can be any other one of the luminance component, the Cb chrominance component, and the Cr chrominance component, and the third component mentioned herein can be the remaining one of the luminance component, the Cb chrominance component, and the Cr chrominance component. In some examples, the loop filter may be omitted, and the decoded video block may be directly provided to the DPB 64 by the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed in one or more of the illustrated fixed or programmable hardware units.
[0073] The video data memory 40 may store video data to be encoded by the components of the video encoder 20. The video data in the video data memory 40 may, for example, be from such as Figure 1Aobtained from the video source 18 shown. The DPB 64 is a buffer that stores reference video data (reference frames or pictures) for use by the video encoder 20 (e.g., in an intra-frame or inter-frame prediction coding mode) when encoding video data. The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20 or off-chip relative to those components.
[0074] As Figure 2 shown, after receiving the video data, the partitioning unit 45 within the prediction processing unit 41 partitions the video data into video blocks. This partitioning operation can also include partitioning the video frame into strips, tiles (e.g., a collection of video blocks), or other larger coding units (CUs) according to a predefined partitioning structure associated with the video data (e.g., a quad-tree (QT) structure). A video frame is or can be considered a two-dimensional array or matrix of samples having sample values. The samples in the array can also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame can be partitioned into multiple video blocks by (e.g.) using QT partitioning. A video block is also or can be considered a two-dimensional array or matrix of samples having sample values, but the size of the video block is smaller than that of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block can be further partitioned into one or more block partitions or sub-blocks (which can again form blocks) by, for example, iteratively using QT partitioning, binary-tree (BT) partitioning, or triple-tree (TT) partitioning or any combination thereof. It should be noted that as used herein, the term "block" or "video block" can be a part of a frame or picture, particularly a rectangular (square or non-square) part. Referring, for example, to HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or correspond to a corresponding block, such as a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB) and / or correspond to a sub-block.
[0075] The prediction processing unit 41 may select one of multiple feasible prediction coding modes for the current video block based on error results (e.g., bit rate and distortion level), such as one of multiple intra prediction coding modes or one of multiple inter prediction coding modes. The prediction processing unit 41 may provide the resulting intra prediction coded block (e.g., prediction block) or inter prediction coded block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coded block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (such as motion vectors, intra mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.
[0076] To select a suitable intra prediction coding mode for the current video block, the intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may execute multiple coding channels, e.g., select a suitable coding mode for each block of video data.
[0077] In some embodiments, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating motion vectors according to a predetermined pattern within the video frame sequence, where the motion vectors indicate the displacement of video blocks within the current video frame relative to prediction blocks within a reference video frame. The motion estimation performed by the motion estimation unit 42 is a process of generating motion vectors that estimate the motion for video blocks. For example, the motion vectors may indicate the displacement of video blocks within the current video frame or picture relative to prediction blocks within a reference frame, and the prediction blocks within the reference frame relative to the current block being coded within the current frame. The predetermined pattern may designate video frames in the sequence as P frames or B frames. The intra BC unit 48 may determine vectors (e.g., block vectors) for intra BC coding in a manner similar to the motion vectors determined by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vectors.
[0078] In terms of pixel difference, the predicted block of a video block can be or can correspond to a block or reference block of a reference frame that closely matches the video block to be encoded, and the pixel difference can be determined by the Sum of Absolute Difference (SAD), the Sum of Square Difference (SSD), or other difference metrics. In some embodiments, video encoder 20 can calculate values for sub-integer pixel positions of the reference frames stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 can perform a motion search relative to full-pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.
[0079] Motion estimation unit 42 calculates a motion vector for a video block in an inter-predicted coded frame by comparing the position of the video block with the position of a predicted block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), where each of the first reference frame list and the second reference frame list identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44, and then to entropy coding unit 56.
[0080] The motion compensation performed by motion compensation unit 44 can involve extracting or generating a predicted block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for the current video block, motion compensation unit 44 can locate the predicted block pointed to by the motion vector in one of the reference frame lists in the reference frame list, retrieve the predicted block from DPB 64, and forward the predicted block to adder 50. Then, adder 50 forms a residual video block of pixel differences by subtracting the pixel values of the predicted block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel differences forming the residual video block can include luminance component differences or chrominance component differences or both. Motion compensation unit 44 can also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 when decoding the video blocks of the video frame. The syntax elements can include, for example, syntax elements defining the motion vectors for identifying the predicted blocks, any flags indicating the prediction mode, or any other syntax information described herein. Note that motion estimation unit 42 and motion compensation unit 44 can be highly integrated but are shown separately for conceptual purposes.
[0081] In some embodiments, the intra BC unit 48 may generate vectors and extract prediction blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine an intra prediction mode to be used for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during different coding passes, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select a suitable intra prediction mode to use from among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the suitable intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra BC unit 48 may calculate a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0082] In other examples, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction according to the embodiments described herein. In either case, for intra block copy, in terms of pixel differences, the prediction block may be a block that is considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and the identification of the prediction block may include calculating values at sub-integer pixel positions.
[0083] Regardless of whether the prediction block is from the same frame according to intra prediction or a different frame according to inter prediction, the video encoder 20 may form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values for forming the residual video block may include both luminance component differences and chrominance component differences.
[0084] As an alternative to the inter - frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above or the intra - block copy prediction performed by the intra - BC unit 48, the intra - prediction processing unit 46 may perform intra - prediction on the current video block. Specifically, the intra - prediction processing unit 46 may determine an intra - prediction mode for encoding the current block. To this end, the intra - prediction processing unit 46 may, for example, use various intra - prediction modes to encode the current block during different coding passes, and the intra - prediction processing unit 46 (or in some examples, the mode selection unit) may select a suitable intra - prediction mode from the tested intra - prediction modes for use. The intra - prediction processing unit 46 may provide information indicating the intra - prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra - prediction mode into the bitstream.
[0085] After the prediction processing unit 41 determines a prediction block for the current video block via inter - frame prediction or intra - frame prediction, the adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0086] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may subsequently perform a scan on the matrix including the quantized transform coefficients. Optionally, the entropy coding unit 56 may perform the scan.
[0087] After quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding methods or techniques. Then, the encoded bitstream can be sent to video decoder 30 as shown in Figure 1A or archived in a storage device 32 as shown in Figure 1A for later sending to or extraction by video decoder 30. Entropy encoding unit 56 can also entropy encode the motion vectors and other syntax elements for the current video frame being encoded.
[0088] Inverse quantization unit 58 and inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual video block in the pixel domain for generating a reference block for predicting other video blocks. As noted above, motion compensation unit 44 can generate a motion compensated prediction block from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0089] Adder 62 adds the reconstructed residual block to the motion compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. Then, the reference block can be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter prediction of another video block in a subsequent video frame.
[0090] Figure 3 is a block diagram showing another exemplary video decoder 30 according to some embodiments of the present application. Video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. Prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. Video decoder 30 can perform operations in combination with those described above with respect to Figure 2A decoding process that is substantially inverse to the encoding process described for video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, while intra prediction unit 84 may generate prediction data based on intra prediction mode indicators received from entropy decoding unit 80.
[0091] In some examples, units of video decoder 30 may be assigned tasks to perform embodiments of the present application. Additionally, in some examples, embodiments of the present disclosure may be distributed across one or more of multiple units of video decoder 30. For example, intra BC unit 85 may perform embodiments of the present application either alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85, and the functionality of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0092] Video data memory 79 may store video data to be decoded by other components of video decoder 30, such as an encoded video bitstream. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source (such as a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a Coded Picture Buffer (CPB) that stores encoded video data from the encoded video bitstream. DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra or inter prediction decoding modes) when decoding video data. Video data memory 79 and DPB 92 may be formed of any of a variety of memory devices, such as Dynamic Random Access Memory (DRAM) (including Synchronous DRAM (SDRAM)), Magnetoresistive RAM (MRAM), Resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are depicted in Figure 3 as two different components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or off-chip relative to those components.
[0093] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or at the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators, and other syntax elements. Then, entropy decoding unit 80 forwards the motion vectors or intra prediction mode indicators and other syntax elements to prediction processing unit 81.
[0094] When a video frame is encoded as an intra prediction encoded (I) frame or for intra-coded prediction blocks in other types of frames, intra prediction unit 84 of prediction processing unit 81 may generate prediction data for video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.
[0095] When a video frame is encoded as an inter prediction encoded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be generated from a reference frame within one of the reference frame lists. Video decoder 30 may construct the reference frame lists, list 0 and list 1, using default construction techniques based on the reference frames stored in DPB 92.
[0096] In some examples, when decoding a video block according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within the reconstructed region of the same picture as the current video block defined by video encoder 20.
[0097] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for decoding video blocks of the video frame, an inter prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, motion vectors for each inter prediction encoded video block of the frame, an inter prediction state for each inter prediction encoded video block of the frame, and other information for decoding video blocks in the current video frame.
[0098] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as a flag determining that the current video block is predicted using the intra BC mode, construction information on which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, a block vector for each intra BC predicted video block of the frame, an intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.
[0099] The motion compensation unit 82 may also perform interpolation using an interpolation filter as used by the video encoder 20 during the encoding of the video block to calculate interpolation for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use these interpolation filters to generate a prediction block.
[0100] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameters used by the video encoder 20 to determine the degree of quantization for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0101] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (such as a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. Then, the decoded video blocks in a given frame are stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device (e.g., Figure 1A display device 34).
[0102] In a typical video encoding and decoding process, a video sequence generally includes an ordered collection of frames or pictures. Each frame may include three arrays of samples, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other cases, a frame may be monochrome and thus include only one two-dimensional array of luminance samples.
[0103] As Figure 1C shown, video encoder 20 (or more specifically, the partitioning unit in the prediction processing unit of video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs that are consecutively ordered from left to right and top to bottom in a raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size of one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 1D shown, each CTU may include one CTB of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements for encoding and decoding the samples of the coding tree blocks. The syntax elements describe the nature of different types of units of the encoded pixel blocks and how the video sequence can be reconstructed at video decoder 30, including inter prediction or intra prediction, intra prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding and decoding the samples of the coding tree block. The coding tree block can be an N×N sample block.
[0104] To achieve better performance, video encoder 20 may recursively perform tree partitioning on the coding tree blocks of a CTU, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, and partition the CTU into smaller CUs. Figure 1B - Figure 1E is a block diagram showing how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure. As Figure 1E depicted, a 64×64 CTU 400 is first partitioned into four smaller CUs, each CU having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are respectively partitioned into four CUs with a block size of 16×16. Two 16×16 CUs 430 and 440 are respectively further partitioned into four CUs with a block size of 8×8. Figure 1B depicts as shown in Figure 1EThe quadtree data structure of the final result of the partitioning process of CTU 400 depicted therein, where each leaf node of the quadtree corresponds to a CU of various sizes ranging from 32×32 to 8×8. Similar to Figure 1D the CTU depicted therein, each CU may include two corresponding coded blocks of the luminance samples and chrominance samples of a frame of the same size, and syntax elements for encoding and decoding the samples of the coded blocks. In a monochrome picture or a picture with three separate color planes, the CU may include a single coded block and a syntax structure for encoding and decoding the samples of the coded block. It should be noted that Figure 1E and Figure 1B the quadtree partitioning depicted therein is for illustrative purposes only, and a CTU may be partitioned into CUs based on quadtree / tritree / binary tree partitioning to adapt to varying local characteristics. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by a binary tree structure and a tritree structure. As Figure 4A - Figure 4E shown, there are five possible partitioning types for a coded block with width W and height H, namely quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0105] In some embodiments, the video encoder 20 may further partition the coded block of the CU into one or more M×NPBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB of the luminance samples, two corresponding PBs of the chrominance samples, and syntax elements for predicting the PB. In a monochrome picture or a picture with three separate color planes, the PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luminance block, a predicted Cb block, and a predicted Cr block for the luminance PB, Cb PB, and Cr PB of each PU of the CU.
[0106] The video encoder 20 may use intra-frame prediction or inter-frame prediction to generate a predicted block for the PU. If the video encoder 20 uses intra-frame prediction to generate a predicted block for the PU, the video encoder 20 may generate the predicted block for the PU based on the decoded samples of the frame associated with the PU. If the video encoder 20 uses inter-frame prediction to generate a predicted block for the PU, the video encoder 20 may generate the predicted block for the PU based on the decoded samples of one or more frames other than the frame associated with the PU.
[0107] After the video encoder 20 generates prediction luminance blocks, prediction Cb blocks, and prediction Cr blocks for one or more PUs of a CU, the video encoder 20 may generate a luminance residual block for the CU by subtracting the prediction luminance block of the CU from the original luminance coded block of the CU, such that each sample in the luminance residual block of the CU indicates the difference between the luminance sample in one of the prediction luminance blocks of the CU and the corresponding sample in the original luminance coded block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the prediction Cb blocks of the CU and the corresponding sample in the original Cb coded block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between the Cr sample in one of the prediction Cr blocks of the CU and the corresponding sample in the original Cr coded block of the CU.
[0108] In addition, as Figure 1E shown, the video encoder 20 may use quadtree partitioning to decompose the luminance residual block, the Cb residual block, and the Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of the CU may include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Thus, each TU of the CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with the TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, the TU may include a single transform block and a syntax structure for transforming the samples of the transform block.
[0109] The video encoder 20 may apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalars. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0110] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a sequence of bits that form a representation of an encoded frame and associated data, and the bitstream is stored in storage device 32 or sent to target device 14.
[0111] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of the current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding samples of the prediction block for the PU of the current CU to corresponding samples of the transform block of the TU of the current CU. After reconstructing the coding block for each CU of the frame, the video decoder 30 may reconstruct the frame.
[0112] As mentioned above, video codecs mainly use two modes, namely, intra-frame prediction (or intra-frame prediction) and inter-frame prediction (or inter-frame prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to codec efficiency than intra-frame prediction because motion vectors are used to predict the current video block from the reference video block.
[0113] However, with the ever-improving video data capture technology and more refined video block sizes for retaining details in the video data, the amount of data required to represent the motion vector of the current frame has also increased significantly. One way to overcome this challenge is to benefit from the fact that not only a group of neighboring CUs in both the spatial domain and the temporal domain have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by exploring the spatial and temporal correlation of the spatially neighboring CUs and / or the temporally co-located CUs, which is also referred to as the "motion vector predictor (MVP)) of the current CU.
[0114] Instead of encoding the actual motion vector of the current CU determined by the motion estimation unit as described above in conjunction with Figure 2 into the video bitstream, the actual motion vector of the current CU is subtracted from the motion vector predictor of the current CU to generate a Motion Vector Difference (MVD) of the current CU. By doing so, it is not necessary to encode the motion vectors determined by the motion estimation unit for each CU of a frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0115] As in the process of selecting a prediction block in a reference frame during inter prediction of an encoded block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally collocated CUs of the current CU to construct a motion vector candidate list (also referred to as a "merge list") of the current CU, and then select a member from the motion vector candidate list as the motion vector predictor of the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the motion vector predictor selected within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list to encode and decode the current CU.
[0116] Generally, the basic inter prediction scheme applied in VVC remains almost the same as the basic inter prediction scheme of HEVC, except for further expanding, adding, and / or improving several prediction tools, such as expanded merge prediction, MMVD, and GPM.
[0117] Extended Merge Prediction
[0118] As video data capture technologies continue to improve and the video block sizes for preserving details in video data become finer, the amount of data required to represent the motion vectors of the current picture also increases significantly. One way to overcome this challenge is to use the motion information (e.g., motion vectors) of spatially neighboring CUs, temporally collocated CUs, etc. of the current CU as an approximation (e.g., prediction) of the motion information of the current CU, which is also referred to as the "Motion Vector Predictor (MVP)" of the current CU.
[0119] Similar to the process of selecting a prediction block in a reference picture during inter - prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules to construct an MVP candidate list for the current CU, and then select an MVP candidate from the MVP candidate list as the MVP for the current CU. By doing so, it is not necessary to send the MVP candidate list itself between the video encoder 20 and the video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same MVP candidate selected from the MVP candidate list to encode and decode the current CU.
[0120] In VVC, an MVP candidate list is constructed by sequentially including the following five types of MVPs:
[0121] - Spatial MVPs from spatially neighboring CUs (i.e., spatial candidates);
[0122] - Temporal MVPs from temporally collocated CUs (i.e., temporal candidates);
[0123] - History - based MVPs (HMVP) from a first - in - first - out (FIFO) table;
[0124] - Paired - average MVPs; and
[0125] - Zero MVPs.
[0126] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded in merge mode, the index of the best MVP candidate is encoded using truncated unary binarization. The first binary bit of the index is coded and decoded with context, and bypass coding is used for the other binary bits of the index.
[0127] The derivation process for each type of MVP is provided as follows. Similar to HEVC, VVC also supports parallel derivation of the MVP candidate list for all CUs within a region of a specific size.
[0128] Deriving MVP from Spatial Candidates
[0129] Except for swapping the positions of the first two spatial candidates, the derivation of MVPs from spatial candidates (e.g., Figure 5 the CU adjacent to the current CU 101 in Figure 5Select up to four spatial candidates from the spatial candidates at the positions shown (i.e., the top position B0, the left position A0, the upper right position B1, the lower left position A1, and the upper left position B2). The derivation is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only if one or more of the CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because one or more of the CUs belong to another strip or tile) or are intra-coded.
[0130] After adding the CU at position B0 as a candidate to the merge candidate list, a redundancy check is performed on the remaining candidates added to the merge candidate list, which ensures that candidates with the same motion information are excluded from the merge candidate list, thereby improving the coding and decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only the pairs linked by the lines with arrows in Figure 6 are considered, and a candidate is added to the merge candidate list only if the candidates in the corresponding pair used for the redundancy check do not have the same motion information as the candidate to be added. The spatial MVP derived from the candidates in the merge candidate list is added to the MVP candidate list.
[0131] Deriving MVP from Temporal Candidates
[0132] During the derivation of the MVP from the temporal candidates, only one temporal candidate is added to the merge candidate list. Specifically, when deriving the MVP from this temporal candidate, a scaled motion vector is derived based on the co-located CU (e.g., col_CU 301 in Figure 7 ), which is a temporal candidate, of the co-located picture (e.g., col_pic 302 in Figure 7 ), which belongs to the current CU (e.g., curr_CU 303 in Figure 7 ), and it is added to the MVP candidate list as a temporal MVP candidate. The reference picture list and the reference picture index to be used for deriving the co-located CU are explicitly signaled in the strip header. The scaled motion vector is obtained (i.e., scaled) from the motion vector of the co-located CU using the picture order count (POC) distance (i.e., tb and td), as shown in Figure 7 where tb is defined as the POC difference between the reference picture (e.g., curr_ref 305 in Figure 7 ), which is of the current picture (e.g., curr_pic 304 in Figure 7 ), and the current picture, and td is defined as the POC difference between the reference picture (e.g., col_ref 306 in Figure 7 ), which is of the co-located picture, and the co-located picture. The reference picture index of the temporal candidate is set to be equal to zero.
[0133] As shown inFigure 8 As shown, a position for a temporal candidate (i.e., a co-located CU) in the current CU 401 is selected between positions C0 and C1. If the CU at position C0 in the co-located picture is unavailable, intra-coded, or outside the current row of the CTU, the CU at position C1 is used as the co-located CU for deriving the temporal MVP candidate. Otherwise, the CU at position C0 is used as the co-located CU for deriving the temporal MVP candidate.
[0134] Derivation of HMVP Candidates
[0135] After the spatial MVP and the temporal MVP, the HMVP candidates are added to the MVP candidate list. The motion information of previously encoded blocks is stored in the HMVP table and is used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a non-sub-block inter-coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the HMVP table.
[0136] The size of the HMVP table can be set to 6. When inserting a new HMVP candidate into the HMVP table, a constrained FIFO rule is used, where first a redundancy check is applied to find if there is the same HMVP in the HMVP table. If found, the same HMVP is deleted from the HMVP table, and all subsequent HMVP candidates are moved forward, and the same HMVP is added to the last entry of the HMVP table.
[0137] The HMVP candidates can be used in the MVP candidate list construction process. The last few HMVP candidates in the HMVP table are checked in order and inserted into the MVP candidate list after the temporal MVP candidates. A redundancy check is applied to the HMVP candidates relative to the spatial candidates and / or the temporal MVP candidates.
[0138] To reduce the number of redundancy check operations, the following simplifications are introduced:
[0139] - A redundancy check is performed on the last two entries in the HMVP table relative to the spatial MVP candidates derived from the spatial candidates at positions A1 and B1 respectively; and
[0140] - Once the total number of available MVP candidates reaches one less than the maximum allowed size of the MVP candidate list, the process of constructing the MVP candidate list from the HMVP candidates is terminated.
[0141] Derivation of Pairwise Averaged MVP Candidates
[0142] A paired-average MVP candidate is generated by averaging MVPs derived from predefined pairs that use the first two merge candidates in the existing merge candidate list. The first merge candidate in the predefined pair can be defined as p0Cand, and the second merge candidate in the predefined pair can be defined as p1Cand. Average motion vectors are calculated for the availability of each reference picture list according to the motion vectors of p0Cand and p1Cand. If for a reference picture list, both motion vectors are available, the two motion vectors are averaged even when the two motion vectors point to different reference pictures, and the reference picture of the average motion vector is set to the reference picture of p0Cand; if for a reference picture list, only one motion vector is available, that motion vector is directly used; if for a reference picture list, no motion vector is available, the motion vector and reference picture index of that reference picture list remain invalid.
[0143] Zero MVP
[0144] When the MVP candidate list is not full after adding the paired-average MVP candidate, zero MVPs are inserted at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0145] MMVD
[0146] As described above, in the merge mode, motion information (i.e., MVP candidates) is implicitly derived from the MVP candidate list constructed for the current CU and is directly used as the MV of the current CU for generating the predicted samples of the current CU, which may result in a specific error between the actual MV of the current CU and the implicitly derived MVP. To improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, where the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. The MMVD flag is signaled after sending the regular merge flag to specify whether Pairwise Averaging MVP Candidates .
[0147] In the MMVD mode, after selecting an MVP candidate from the first two MVP candidates in the MVP candidate list, the MMVD information is signaled, where the MMVD information includes an MMVD candidate flag, a distance index, and a direction index, where the MMVD candidate flag is used to specify which one of the first two MVP candidates is selected as the MV basis, the distance index is used to indicate the motion amplitude information of the MVD, and the direction index is used to indicate the motion direction information of the MVD.
[0148] The distance index that specifies the motion amplitude information of the MVD indicates the distance from the reference picture of the current CU pointed to by the selected MVP candidate (e.g.,Figure 9 a predefined offset of a starting point (e.g., indicated by a dashed circle in L0 reference picture 501 or L1 reference picture 503 in Figure 9 ), and the MVD can be derived from this offset and added to the selected MVP candidate. The relationship between the distance index and the predefined offset is specified in Table 1 below.
[0149]
[0150] Table 1
[0151] The direction index specifies the sign of the MVD, which represents the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined signs. In some examples, the meaning of the sign of the MVD may vary according to the information of the selected MVP candidate. When the selected MVP candidate is a uni - directional predicted MV or a bi - directional predicted MV where two MVs point to the same side of the current picture (i.e., the POCs of the two reference pictures of the current picture (e.g., the reference pictures of list 0 and list 1, which are also referred to as L0 reference picture and L1 reference picture respectively) are both greater than the POC of the current picture, or both less than the POC of the current picture), the signs in Table 2 specify the signs of the MVDs added to the MVs of the selected MVP candidate. When the selected MVP candidate is a bi - directional predicted MV where two MVs point to different sides of the current picture (i.e., the POC of one reference picture of the current picture is greater than the POC of the current picture, and the POC of the other reference picture of the current picture is less than the POC of the current picture), if the POC distance of the L0 reference picture (i.e., the POC distance between the L0 reference picture and the current picture) is greater than the POC distance of the L1 reference picture (i.e., the POC distance between the L1 reference picture and the current picture), the signs in Table 2 specify the signs of the MVDs added to the MV of list 0 MVD0 for the list 0 MVP0 in the selected MVP candidate, and the signs of the MVDs added to the MV of list 1 MVD1 for the list 1 MVP1 in the selected MVP candidate are opposite to the signs in Table 2; otherwise, if the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, the signs in Table 2 specify the signs of the MVDs added to the MV of list 1 MVD1, and the signs of the MVDs added to the MV of list 0 MVD0 are opposite to the signs in Table 2.
[0152] Direction Index 00 01 10 11 x - axis + - N / A N / A y - axis N / A N / A + -
[0153] Table 2
[0154] Scale the MVD according to the POC distance. If the POC distances for both the L0 reference picture and the L1 reference picture are the same, the MVD does not need to be scaled. Otherwise, if the POC distance for the L0 reference picture is greater than the POC distance for the L1 reference picture, MVD1 is scaled. If the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, MVD0 is scaled.
[0155] GPM
[0156] In VVC, GPM is supported for inter prediction. GPM is signaled using a CU-level flag as a merge mode, and other merge modes include the regular merge mode, the MMVD mode, the CIIP mode, and the sub-block merge mode. For each possible CU size W×H (W = 2 m and H = 2 n , where m,n ∈ {3,4,5,6}) excluding 8×64 and 64×8, GPM supports a total of 64 partitions.
[0157] When using GPM, the CU is divided into two parts by a geometrically positioned line. The position of the dividing line is derived mathematically based on the angle and offset parameters of a specific partition. Each part of the CU obtained by the geometric division is inter predicted using its own motion; and only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure that, like regular bi-directional prediction, only two motion-compensated predictions are required for each CU.
[0158] If GPM is used for the current CU, a geometric partition index indicating the partition mode (indicating the angle and offset of the geometric partition) and two merge indices (one for each partition) are further signaled.
[0159] The unidirectional prediction candidate list is directly derived from the merge candidate list constructed according to the above extended merge prediction process. Let n denote the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector of the nth merge candidate in the merge candidate list (where X is equal to the parity of n) is used as the nth unidirectional prediction motion vector for GPM. These motion vectors are marked with "x" in Figure 10 . In the case where there is no corresponding LX motion vector of the nth merge candidate in the merge candidate list, the L(1 - x) motion vector of the same merge candidate is used as the unidirectional prediction motion vector for GPM.
[0160] CIIP
[0161] In VVC, when a CU is encoded in merge mode, if the CU contains at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if both the width and height of the CU are less than 128 luma samples, an extra flag is signaled to indicate whether the CIIP mode is applied to the current CU. In the CIIP mode, a prediction signal is obtained by combining an inter prediction signal and an intra prediction signal. The same inter prediction process as that applied in the regular merge mode is used to derive the inter prediction signal in the CIIP mode; and the intra prediction signal in the CIIP mode is derived after the regular intra prediction process with a planar mode. Then, weighted averaging is used to combine the intra prediction signal and the inter prediction signal, where the weight value is calculated according to the coding modes of the top neighboring block and the left neighboring block of the current CU 1601 (as Figure 11 shown) as follows:
[0162] - If the top neighboring block is available and is intra-coded, isIntraTop is set to 1, otherwise isIntraTop is set to 0;
[0163] - If the left neighboring block is available and is intra-coded, isIntraLeft is set to 1, otherwise isIntraLeft is set to 0;
[0164] - If (isIntraLeft + isIntraTop) equals 2, the weight value is set to 3;
[0165] - Otherwise, if (isIntraLeft + isIntraTop) equals 1, the weight value is set to 2;
[0166] - Otherwise, the weight value is set to 1.
[0167] - The prediction signal P in the CIIP mode is derived as follows CIIP :
[0168] P CIIP = ((4 - wt)*P inter + wt*P intra + 2) >> 2 (1)
[0169] where P inter is the inter prediction signal in the CIIP mode, P intra is the intra prediction signal in the CIIP mode, wt is the weight value, and >> represents a right shift operation.
[0170] Intra - Block Copy in Versatile Video Coding (VVC)
[0171] Intra Block Copy (IBC) is a tool adopted in the HEVC extension for SCC. IBC significantly improves the coding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture. The luminance block vector of the IBC-encoded CU is of integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The IBC-encoded CU is regarded as a third prediction mode in addition to the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs with both width and height less than or equal to 64 luminance samples.
[0172] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height not greater than 16 luminance samples. For non-merge modes, a block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a block-matching based local search will be performed.
[0173] In the hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For a current block of a larger size, when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions, the hash key is determined to match the hash key of the reference block. If it is found that the hash keys of multiple reference blocks match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference block with the minimum cost is selected.
[0174] In the block-matching search, the search range is set to cover both the previous CTU and the current CTU.
[0175] At the CU level, the IBC mode is signaled using a flag, and the IBC mode can be signaled as the IBCAMVP mode or the IBC skip / merge mode:
[0176] IBC skip / merge mode: The merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC-encoded blocks is used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and pairwise candidates.
[0177] IBC AMVP mode: The block vector difference is coded and decoded in the same way as the motion vector difference. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (in the case of IBC coding). When either neighbor is not available, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0178] IBC Reference Region
[0179] To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstructed part of the predefined region to include the region of the current CTU and a specific region of the left CTU. Figure 12 The reference region of the IBC mode is shown, where each block represents a 64×64 luma sample unit.
[0180] According to the position of the current coded CU position within the current CTU, the following operations are applied:
[0181] In the case where the current block falls within the upper-left 64×64 block of the current CTU, in addition to referring to the reconstructed samples in the current CTU, the current block can also use the CPR mode to refer to the reference samples in the lower-right 64×64 block of the left CTU. The current block can also use the CPR mode to refer to the reference samples in the lower-left 64×64 block of the left CTU and the reference samples in the upper-right 64×64 block of the left CTU.
[0182] In the case where the current block falls within the upper-right 64×64 block of the current CTU, in addition to referring to the reconstructed samples in the current CTU, if the luma position (0,64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower-left 64×64 block and the lower-right 64×64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower-right 64×64 block of the left CTU.
[0183] In the case where the current block falls within the lower-left 64×64 block of the current CTU, in addition to referring to the reconstructed samples in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the upper-right 64×64 block and the lower-right 64×64 block of the left CTU. Otherwise, the current block can also use the CPR mode to refer to the reference samples in the lower-right 64×64 block of the left CTU.
[0184] In the case where the current block falls within the lower-right 64×64 block of the current CTU, the current block can use the CPR mode to only refer to the reconstructed samples in the current CTU.
[0185] This restriction allows the use of local on-chip memory to implement the IBC mode for hardware implementation.
[0186] Interaction of IBC with Other Coding Tools
[0187] The interaction between the IBC mode and other inter-frame coding tools in VVC, such as paired merge candidates, history-based motion vector predictors (HMVP), combined intra / inter-frame prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM), is as follows:
[0188] IBC can be used together with paired merge candidates and HMVP. New paired IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference.
[0189] IBC cannot be combined with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM.
[0190] When using the DUAL_TREE partition, IBC is not allowed for chrominance coding blocks.
[0191] Different from the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The process of deriving motion vectors for the IBC mode does not include all neighboring blocks in the inter-frame mode, and vice versa. The following IBC design aspects are applied:
[0192] IBC shares the same process as regular MV merge, including paired merge candidates and history-based motion predictors, but does not allow TMVP and zero vectors because they are invalid for the IBC mode.
[0193] Separate HMVP buffers (5 candidates each) are used for regular MV and IBC.
[0194] Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream, and if the merge candidate is invalid (out of range or 0), the merge should not be used. This bitstream consistency constraint is expressed according to the virtual buffer described below.
[0195] For deblocking, IBC is treated as an inter-frame mode.
[0196] If the current block is encoded using the IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to indicate only whether the MV is an integer pixel or 4 integer pixels.
[0197] The number of IBC merge candidates can be signaled separately in the strip header from the number of regular merge candidates, sub-block merge candidates, and geometric merge candidates.
[0198] The virtual buffer concept is used to describe the admissible reference regions and valid block vectors for IBC prediction modes. Represent the CTU size as ctbSize. The width wIbcBuf of the virtual buffer ibcBuf = 128×128 / ctbSize and the height hIbcBuf = ctbSize. For example, for a CTU size of 128×128, the size of ibcBuf is also 128×128; for a CTU size of 64×64, the size of ibcBuf is 256×64; for a CTU size of 32×32, the size of ibcBuf is 512×32.
[0199] The size of the VPDU is min(ctbSize,64) in each dimension, Wv = min(ctbSize,64).
[0200] The virtual IBC buffer ibcBuf is maintained as follows.
[0201] When starting to decode each CTU row, the entire ibcBuf is flushed with an invalid value of -1.
[0202] When starting to decode the VPDU(xVPDU,yVPDU) relative to the upper left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU%wIbcBuf,..., xVPDU%wIbcBuf + Wv - 1; y = yVPDU%ctbSize,..., yVPDU%ctbSize + Wv - 1.
[0203] After decoding, for the CU including (x,y) relative to the upper left corner of the picture, set:
[0204] ibcBuf[x%wIbcBuf][y%ctbSize] = recSample[x][y]
[0205] For a block covering the coordinates (x,y), the block vector is valid if the following condition for the block vector bv = (bv[0],bv[1]) is true; otherwise, the block vector is invalid:
[0206] ibcBuf[(x + bv[0])%wIbcBuf][(y + bv[1])%ctbSize] should not be equal to -1.
[0207] Intra - Block Copy in Enhanced Compression Model (ECM)
[0208] In the ECM, IBC is improved in the following aspects.
[0209] IBC Merge List / AMVP List Construction
[0210] Modify the IBC merge list / AMVP list construction as follows:
[0211] Only if the IBC merge candidate / AMVP candidate is valid can it be inserted into the IBC merge candidate list / AMVP candidate list.
[0212] The top - right spatial candidate, bottom - left spatial candidate, top - left spatial candidate, and a pairwise average candidate can be added to the IBC merge candidate list / AMVP candidate list.
[0213] Template - based adaptive re - ordering (ARMC - TM) is applied to the IBC merge list.
[0214] The size of the HMVP table for IBC is increased to 25. After exporting at most 20 IBC merge candidates by using full pruning, they are re - ordered together. After re - ordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0215] The candidate for the zero vector used to fill the IBC merge list / AMVP list is replaced by the set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in the IBC merge mode, and thus, it is discarded as a BVP in the IBC candidate list.
[0216] Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of three sub - regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Figure 13 shown.
[0217] IBC with Template Matching
[0218] Template matching is used in both the IBC merge mode and the IBC AMVP mode in IBC.
[0219] Compared with the IBC - TM merge list used by the conventional IBC merge mode, the IBC - TM merge list is modified such that candidates are selected according to a pruning method with the motion distance between candidates, as in the conventional TM merge mode. The end - zero motion is replaced by motion vectors to the left (-W, 0), top (0, -H), and top - left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0220] In the IBC-TM merge mode, a template matching method is used to refine the selected candidates before the RDO or decoding process. The IBC-TM merge mode competes with the regular IBC merge mode and signals the TM merge flag.
[0221] In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. A template matching method is used to refine each of the 3 selected candidates and they are sorted according to their resulting template matching cost. Then only the first 2 candidates are considered as normal during the motion estimation process.
[0222] The template matching refinement for both the IBC-TM merge mode and the AMVP mode is very simple because the IBC motion vectors are constrained to be (i) integers and (ii) within the reference region, as Figure 12 shown. Therefore, in the IBC-TM merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, they are performed with integer precision or 4-pixel precision according to the AMVR value. Such refinements only access samples without interpolation. In both cases, the motion vectors refined in each refinement step and the templates used must comply with the constraints of the reference region.
[0223] IBC Reference Region
[0224] The reference region for IBC is extended to the upper two CTU rows. Figure 14 The reference region for encoding the CTU(m,n) is shown. Specifically, for the CTU(m,n) to be encoded, the reference region includes CTUs with indices (m-2,n-2)…(W,n-2), (0,n-1)…(W,n-1), (0,n)…(m,n), where W represents the maximum horizontal index within the current tile, slice, or picture. This setting ensures that for a CTU of size 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is limited to [-(C<<1),C>>2] in the horizontal direction and to [-C,C>>2] in the vertical direction to accommodate the reference region extension, where C represents the CTU size.
[0225] IBC Merge Mode with Block Vector Difference
[0226] The IBC merge mode with block vector difference is adopted in the ECM. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.
[0227] The basic candidate is selected from the first five candidates in the re-ordered IBC merge list. And, all possible MBVD refinement positions (20×4) for each basic candidate are re-ordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. Finally, the first 8 refinement positions with the lowest template SAD cost are retained as available positions for MBVD index coding.
[0228] IBC Adaptation for Camera - Captured Content
[0229] When adapting IBC for camera-captured content, the IBC reference range is reduced from 2 CTU rows to 2×128 rows, as Figure 15 shown. On the encoder side, to reduce complexity, the local search range is centered on the first block vector predictor of the current CU and is set to [-8,8] in the horizontal direction and [-8,8] in the vertical direction. This encoder modification does not apply to SCC sequences.
[0230] Combination of CIIP with TIMD and TM Merging
[0231] In the CIIP mode, the predicted sample points are generated by weighting the inter-frame prediction signal predicted using the CIIP-TM merge candidate and the intra-frame prediction signal predicted using the intra-frame prediction mode derived by TIMD. This method is only applied to coding blocks with an area less than or equal to 1024.
[0232] The TIMD derivation method is used to derive the intra-frame prediction mode in CIIP. Specifically, the intra-frame prediction mode with the minimum SATD value in the TIMD mode list is selected and mapped to one of the 67 regular intra-frame prediction modes.
[0233] In addition, it is also proposed to modify the weights (wIntra, wInter) for the two tests when the derived intra-frame prediction mode is an angular mode. For the near-horizontal mode (2 <= angular mode index < 34), the current block is divided vertically as Figure 16A shown; for the near-vertical mode (34 <= angular mode index <= 66), the current block is divided horizontally as Figure 16B shown.
[0234] (wIntra, wInter) for different sub - blocks are shown in Table 3.
[0235] Sub - Block Index (wIntra, wInter) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6)
[0236] Table 3. Modified weights for angular modes.
[0237] Using CIIP - TM, a CIIP - TM merge candidate list is constructed for the CIIP - TM mode. The merge candidates are refined by template matching. The CIIP - TM merge candidates are also reordered into regular merge candidates by the ARMC method. The maximum number of CIIP - TM merge candidates is equal to two.
[0238] Multiple Hypothesis Prediction (MHP)
[0239] In the multi - hypothesis inter - prediction mode, in addition to the traditional bi - directionally predicted signal, one or more additional motion - compensated prediction signals are signaled. The resulting overall predicted signal is obtained by per - sample weighted superposition. Using the bi - directionally predicted signal p bi and the first additional inter - prediction signal / hypothesis h3, the resulting predicted signal p3 is obtained as follows:
[0240] p3=(1 - α)p bi +αh3 (2)
[0241] According to the mapping presented in Table 4, the weighting factor α is specified by the new syntax element add_hyp_weight_idx:
[0242] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0243] Table 4. Mapping between add_hyp_weight_idx and α
[0244] Similar to the above, more than one additional prediction signal can be used. The resulting overall predicted signal is iteratively accumulated with each additional prediction signal.
[0245] p n+1 =(1 - α n+1 )p n +α n+1 h n+1 (3)
[0246] The resulting overall predicted signal is obtained as the final p n (i.e., the p with the largest index n n ). Within this mode, at most two additional prediction signals can be used (i.e., n is restricted to 2).
[0247] The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag differentiates between these two signaling modes.
[0248] For the inter-frame AMVP mode, MHP is applied only when unequal weights in the BCW are selected in the bi-prediction mode.
[0249] The combination of MHP and BDOF is possible; however, BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses).
[0250] Geometric Partitioning Mode (GPM) in ECM
[0251] Generalized Partitioned Motion (GPM) with Merged Motion Vector Difference (MMVD)
[0252] The GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MVs. First, a flag is signaled for the GPM CU to specify whether to use this mode. If this mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD. If the MVD is signaled for a geometric partition, after selecting the GPM merge candidate, the motion of that partition is further refined by the signaled MVD information. All other procedures remain the same as GPM.
[0253] The MVD is signaled as a pair of distance and direction, similar to MMVD. In GPM with MMVD (GPM-MMVD), nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are involved. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, similar to MMVD, the MVD is left-shifted by 2.
[0254] Generalized Partitioned Motion (GPM) with Template Matching (TM)
[0255] Template matching is applied to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to two geometric partitions. TM is used to refine the motion information for each geometric partition. When TM is selected, according to the partitioning angle, left neighboring samples, upper neighboring samples, or both left and upper neighboring samples are used to construct a template as shown in Table 5. Then the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern as the merge mode (where the half-pixel interpolation filter is disabled).
[0256]
[0257] Table 5. Templates for the first geometric partition and the second geometric partition, where A indicates using upper samples, L indicates using left samples, and L+A indicates using both left samples and upper samples.
[0258] Construct the GPM candidate list as follows:
[0259] 1. Directly derive the interleaved list-0MV candidates and list-1MV candidates from the regular merge candidate list, where the list-0MV candidates have a higher priority than the list-1MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0260] 2. Further directly derive the interleaved list-1MV candidates and list-0MV candidates from the regular merge candidate list, where the list-1MV candidates have a higher priority than the list-0MV candidates. The same pruning method with an adaptive threshold is also applied to remove redundant MV candidates.
[0261] 3. Fill in zero MV candidates until the GPM candidate list is full.
[0262] GPM-MMVD and GPM-TM are enabled for only one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.
[0263] GPM with Inter-Frame Prediction and Intra-Frame Prediction
[0264] Under GPM with inter-frame prediction and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame prediction samples and intra-frame prediction samples for each GPM partition region. The inter-frame prediction samples are derived through inter-frame GPM, while the intra-frame prediction samples are derived through the intra-frame prediction mode (IPM) candidate list and the index signaled from the encoder. The IPM candidate list size is predefined as 3. The available IPM candidates are the parallel angle mode (parallel mode) relative to the GPM block boundary, the vertical angle mode (vertical mode) relative to the GPM block boundary, and the planar mode, respectively, as Figure 17A - Figure 17D shown. Additionally, for as Figure 17DThe GPM with intra prediction and intra prediction is restricted to reduce the signaling overhead for IPM and avoid an increase in the size of the intra prediction circuit on the hardware decoder. In addition, a direct motion vector and IPM storage on the GPM hybrid region are introduced to further improve the coding and decoding performance.
[0265] In the IPM derivation based on DIMD and neighboring modes, the parallel mode is registered first. Therefore, if there are no identical IPM candidates in the list, up to two IPM candidates derived from the decoder-side intra mode (DIMD) method and / or neighboring block derivation can be registered. For neighboring mode derivation, there are up to five positions of available neighboring blocks, but they are restricted by the angle of the GPM block boundary. As shown in Table 6, these positions have been used for GPM with template matching (GPM-TM).
[0266]
[0267] Table 6. Positions of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary. A and L represent above and to the left of the prediction block.
[0268] GPM-intra can be combined with GPM with motion vector difference merging (GPM-MMVD). TIMD is used for IPM candidates in GPM-intra to further improve the coding and decoding performance. The parallel mode can be registered first, followed by IPM candidates of TIMD, DIMD, and neighboring blocks.
[0269] Template-matching-based reordering for GPM partition modes
[0270] In the template-matching-based reordering for GPM partition modes, given the motion information of the current GPM block, the respective TM cost values for the GPM partition modes are calculated. Then, all GPM partition modes are reordered in ascending order based on the TM cost values. Instead of sending the GPM partition modes, an index using the Golomb-Rice code is signaled to indicate the position of the exact GPM partition mode in the reordered list.
[0271] The reordering method for GPM partition modes is a two-step process performed after generating the respective reference templates for the two GPM partitions in the coding unit, as follows:
[0272] · Expand the GPM partition edges into the reference templates of the two GPM partitions to obtain 64 reference templates, and calculate the respective TM costs for each of the 64 reference templates;
[0273] · Reorder the GPM partition modes in ascending order based on the TM cost values of the GPM partition modes, and mark 32 best modes as available partition modes.
[0274] The edges on the template extend from the edges of the current CU, as Figure 18 shown, but the GPM mixing process is not used in the template region across the edges.
[0275] After reordering the TM costs in ascending order, the indices are signaled.
[0276] Intra-template matching
[0277] Intra-template matching prediction (Intra-TMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed part of the current frame, where the best prediction block is the block where the L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed part of the current frame for the template that is most similar to the current template and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0278] A prediction signal is generated by matching the L-shaped causal neighbor of the current block with Figure 19 another block in the predefined search region in
[0279] R1: the current CTU
[0280] R2: the top-left CTU
[0281] R3: the top CTU
[0282] R4: the left CTU
[0283] The sum of absolute differences (SAD) is used as the cost function.
[0284] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses the corresponding block of that template as the prediction block.
[0285] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. That is:
[0286] SearchRange_w = a × BlkW
[0287] SearchRange_h = a × BlkH
[0288] where "a" is a constant that controls the gain / complexity trade-off. In practice, "a" is equal to 5.
[0289] For CUs with width and height less than or equal to 64, the intra-template matching tool is enabled. The maximum CU size for intra-template matching is configurable.
[0290] When DIMD is not used for the current CU, the intra-template matching prediction mode is signaled at the CU level through a dedicated flag.
[0291] Fusion of Template-based Intra-mode Derivation (TIMD)
[0292] For each intra-prediction mode in MPM, the SATD between the predicted samples of the template and the reconstructed samples is calculated. First, the two intra-prediction modes with the minimum SATD are selected as TIMD modes. After applying the PDPC process, these two TIMD modes are fused using weights, and this weighted intra-prediction is used to encode and decode the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of TIMD modes.
[0293] The costs of the two selected modes are compared with a threshold. At test time, cost factor 2 is applied as follows:
[0294] Cost of mode 2 (costMode2) < 2 × cost of mode 1 (costMode1).
[0295] If this condition is true, fusion is applied; otherwise, only mode 1 is used.
[0296] The weights of the two modes are calculated based on the SATD costs of the two modes as follows:
[0297] Weight 1 = costMode2 / (costMode1 + costMode2)
[0298] Weight 2 = 1 - Weight 1
[0299] The same lookup table (LUT)-based integerization scheme used by CCLM is used for the division operation.
[0300] Local Illumination Compensation (LIC)
[0301] LIC is an inter - prediction technique that models the local illumination change between the current block template and the reference block template as the local illumination change between the current block and its predicted block. The parameters of this function can be represented by a scaling factor α and an offset β, which form a linear equation, i.e., α×p[x]+β, to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When loop - based motion compensation is enabled, the loop offset should be considered to clip the MV. Since α and β can be derived based on the current block template and the reference block template, there is no signaling overhead for α and β other than signaling the LIC flag for the AMVP mode to indicate the use of LIC.
[0302] The local illumination compensation proposed in JVET - O0066 is used for unidirectional - predicted inter - CU and is modified as follows.
[0303] Intra - neighboring samples can be used for LIC parameter derivation;
[0304] For blocks with fewer than 32 luma samples, LIC is disabled;
[0305] For both non - sub - block mode and affine mode, LIC parameter derivation is performed based on the samples of the modulo - block corresponding to the current CU instead of the partial modulo - block samples corresponding to the top - left first 16×16 unit.
[0306] The samples of the reference block template are generated by using MC with a block MV that does not need to be rounded to integer - pixel precision.
[0307] OBMC
[0308] As described in JVET - L0101, when OBMC is applied, the motion information of neighboring blocks is used for weighted prediction to refine the top - boundary pixels and left - boundary pixels of the CU.
[0309] The conditions for not applying OBMC are as follows:
[0310] When OBMC is disabled at the SPS level
[0311] When the current block has an intra - mode or IBC mode
[0312] When the current block applies LIC
[0313] When the current luma block area is less than or equal to 32
[0314] Sub - block boundary OBMC is performed by applying the same blending to the top - sub - block boundary pixels, left - sub - block boundary pixels, bottom - sub - block boundary pixels, and right - sub - block boundary pixels by using the motion information of neighboring sub - blocks. Sub - block boundary OBMC is enabled for the following sub - block - based coding tools:
[0315] Affine AMVP mode;
[0316] Affine merge mode and sub-block based temporal motion vector prediction (SbTMVP);
[0317] Sub-block based bilateral matching.
[0318] When the OBMC mode is used together with LMCS in the CIIP mode, inter-frame blending is performed before the LMCS mapping of inter-frame samples. In the CIIP mode, LMCS is applied to the blended inter-frame samples and combined with the intra-frame samples to which LMCS is applied.
[0319]
[0320] where Inter predY denotes the samples predicted by the motion of the current block in the original domain, Intra predY denotes the samples predicted in the mapped domain, OBMC predY denotes the samples predicted by the motion of neighboring blocks in the original domain, and w0 and w1 are weights.
[0321] Template matching based OBMC
[0322] In the template matching based OBMC scheme, the predicted value of the CU boundary sample derivation method is not directly determined using weighted prediction, but is determined according to the template matching cost, including using only the motion information of the current block or using the motion information of neighboring blocks and one of the blending modes.
[0323] In this scheme, for each block of size 4×4 at the top CU boundary, the upper template size is equal to 4×1. If N adjacent blocks have the same motion information, since the MC operation can be processed at one time, the upper template size is enlarged to 4N×1. For each left block of size 4×4 at the left CU boundary, the left template size is equal to 1×4 or 1×4N( Figure 20 ).
[0324] For each 4×4 top block (or group of N 4×4 blocks), the predicted value of the boundary sample is derived following the steps below.
[0325] Taking block A as the current block and its upper neighboring block AboveNeighbor_A as an example. The operation for the left block is performed in the same way.
[0326] First, according to the following three types of motion information, three template matching costs (Cost1, Cost2, Cost3) are measured by the sum of absolute differences (SAD) between the reconstructed samples of the template and the corresponding reference samples derived by the MC process of the template:
[0327] Calculate Cost1 based on the motion information of A.
[0328] Calculate Cost2 based on the motion information of AboveNeighbor_A.
[0329] Calculate Cost3 using weighted prediction with weighting factors of 3 / 4 and 1 / 4 respectively based on the motion information of A and the motion information of AboveNeighbor_A.
[0330] Secondly, by comparing Cost1, Cost2, and Cost3, select a method to calculate the final prediction result of the boundary sample points.
[0331] The original MC result using the motion information of the current block is represented as Pixel1, and the MC result using the motion information of the neighboring block is represented as Pixel2. The final prediction result is represented as NewPixel.
[0332] If Cost1 is the smallest, then NewPixel(i,j) = Pixel1(i,j).
[0333] If (Cost2+(Cost2>>2)+(Cost2>>3)) <= Cost1, then use Hybrid Mode 1.
[0334] For the luma block, the number of hybrid pixel rows is 4.
[0335] NewPixel(i,0) = (26×Pixel1(i,0)+6×Pixel2(i,0)+16) >> 5
[0336] NewPixel(i,1) = (7×Pixel1(i,1)+Pixel2(i,1)+4) >> 3
[0337] NewPixel(i,2) = (15×Pixel1(i,2)+Pixel2(i,2)+8) >> 4
[0338] NewPixel(i,3) = (31×Pixel1(i,3)+Pixel2(i,3)+16) >> 5
[0339] For the chroma block, the number of hybrid pixel rows is 1.
[0340] NewPixel(i,0) = (26×Pixel1(i,0)+6×Pixel2(i,0)+16) >> 5
[0341] If Cost1 <= Cost2, then use Hybrid Mode 2.
[0342] For the luminance block, the number of hybrid pixel rows is 2.
[0343] NewPixel(i,0) = (15 × Pixel1(i,0) + Pixel2(i,0) + 8) >> 4
[0344] NewPixel(i,1) = (31 × Pixel1(i,1) + Pixel2(i,1) + 16) >> 5
[0345] For the chrominance block, the number of hybrid pixel rows / columns is 1.
[0346] NewPixel(i,0) = (15 × Pixel1(i,0) + Pixel2(i,0) + 8) >> 4
[0347] Otherwise, use Hybrid Mode 3.
[0348] For the luminance block, the number of hybrid pixel rows is 4.
[0349] NewPixel(i,1) = (7 × Pixel1(i,1) + Pixel2(i,1) + 4) >> 3
[0350] NewPixel(i,2) = (15 × Pixel1(i,2) + Pixel2(i,2) + 8) >> 4
[0351] NewPixel(i,3) = (31 × Pixel1(i,3) + Pixel2(i,3) + 16) >> 5
[0352] For the chrominance block, the number of hybrid pixel rows is 1.
[0353] NewPixel(i,0) = (7 × Pixel1(i,0) + Pixel2(i,0) + 4) >> 3
[0354] Currently, the IBC tool is not combined with the GPM tool. It is simple to combine them together, which can improve the prediction accuracy and the coding / decoding performance.
[0355] Currently, the coded blocks encoded in the IBC mode are not combined with the coded blocks encoded in the intra mode or the inter mode. It is simple to combine them together, which can improve the prediction accuracy and the coding / decoding performance.
[0356] Currently, the number of block vectors (BV) in the IBC tool is odd. Increasing the number of block vectors (BV) is simple and the predicted results can be combined, which can improve the prediction accuracy and codec performance.
[0357] Currently, the coded blocks encoded in the intra-TMP mode are not combined with the coded blocks encoded in the intra mode or the inter mode. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0358] Currently, the intra-TMP tool is not combined with the GPM tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0359] Currently, the IBC tool is not combined with the TIMD tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0360] Currently, the intra-TMP tool is not combined with the TIMD tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0361] Currently, the intra-TMP tool is not combined with the LIC tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0362] Currently, the IBC tool is not combined with the OBMC tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0363] Currently, the intra-TMP tool is not combined with the OBMC tool. Combining them together is simple, which can improve the prediction accuracy and codec performance.
[0364] In the present disclosure, to solve the above problems, a method for further improving the existing design of IBC is provided. Generally, the main features of the technology proposed in the present disclosure are summarized as follows.
[0365] 1. Combine the IBC tool with the GPM tool. The combination form can be GPM with IBC prediction and IBC prediction, GPM with IBC prediction and intra prediction, or GPM with IBC prediction and inter prediction.
[0366] 2. As a simplified version of the combination of the IBC tool and the GPM tool, for a predefined direction (such as 45 degrees), the upper left part is predicted in the intra mode, the lower right part is predicted in the IBC mode, and then they are averaged and weighted to obtain the final prediction signal.
[0367] 3. Combine the IBC tool with the CIIP tool, where the IBC prediction is combined with the intra prediction mode or the IBC prediction is combined with the inter prediction mode.
[0368] 4. Combine the IBC tool with the MHP tool, where more than one BV prediction is obtained and they are weighted averaged to obtain a final prediction signal.
[0369] 5. Combine the intra TMP tool with the CIIP tool, where the intra TMP is combined with the intra prediction mode or the intra TMP is combined with the inter prediction mode.
[0370] 6. Combine the intra TMP tool with the GPM tool, and the combination form can be GPM with intra TMP prediction and intra TMP prediction, GPM with intra TMP prediction and intra prediction, or GPM with intra TMP prediction and inter prediction.
[0371] 7. As a simplified version of the combination of the intra TMP tool and the GPM tool, for a predefined direction (such as 45 degrees), the upper left part is predicted with the intra mode, the lower right part is predicted with the intra TMP mode, and then they are averaged and weighted to obtain a final prediction signal.
[0372] 8. Combine the IBC tool with the TIMD tool, where the IBC mode is used together with the intra prediction mode in the MPM for TIMD fusion.
[0373] 9. Combine the intra TMP tool with the TIMD tool, where the intra TMP mode is used together with the intra prediction mode in the MPM for TIMD fusion.
[0374] 10. Combine the intra TMP tool with the LIC tool, where the LIC tool is used to compensate for the local illumination change between the current block and its intra TMP prediction block.
[0375] 11. Combine the IBC tool with the OBMC tool, where the top boundary pixels and left boundary pixels of the current block predicted by IBC are refined by the OBMC tool.
[0376] 12. Combine the intra TMP tool with the OBMC tool, where the OBMC tool is used to refine the top boundary pixels and left boundary pixels of the current block predicted with the intra TMP.
[0377] In some examples, the disclosed methods can be applied independently or jointly.
[0378] GPM with IBC Prediction and IBC Prediction
[0379] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of GPM with IBC prediction and IBC prediction. Different methods can be used to achieve this goal.
[0380] In the first method, two "inter - frame" parts of the GPM with inter - frame prediction methods in VVC are replaced by IBC. This means that the combined prediction results of the two IBCs are weighted - averaged with each other according to the dividing line in the coding block. The weights can be obtained by referring to the GPM with inter - frame prediction methods in VVC.
[0381] In the second method, two "inter - frame" parts of the GPM with inter - frame prediction methods in ECM are replaced by IBC, where some template - matching tools can be utilized to further improve the encoding and decoding performance.
[0382] GPM with IBC Prediction and Intra Prediction
[0383] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC prediction and intra - frame prediction. Different methods can be used to achieve this goal.
[0384] In the first method, the "inter - frame" part of the GPM with inter - frame prediction method and intra - frame prediction method in ECM is replaced by IBC, where the combined prediction result of IBC is weighted - averaged with the intra - frame prediction result to obtain the final prediction signal.
[0385] GPM with IBC Prediction and Inter Prediction
[0386] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC prediction and inter - frame prediction. Different methods can be used to achieve this goal.
[0387] In the first method, one "inter - frame" part of the GPM with inter - frame prediction methods in VVC is replaced by IBC, where the combined prediction result of IBC is weighted - averaged with the inter - frame combined prediction result to obtain the final prediction signal.
[0388] In the second method, one "inter - frame" part of the GPM with inter - frame prediction methods in ECM is replaced by IBC, where some template - matching tools can be utilized to further improve the encoding performance.
[0389] Simplified Combination of IBC Prediction and Intra Prediction in the Form of GPM
[0390] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a simplified GPM with IBC prediction and intra - frame prediction, such as IBC prediction and intra - frame prediction are combined in a specific partitioning mode, which can save the bit overhead for representing the partitioning mode. Different methods can be used to achieve this goal.
[0391] In the first method, for a dividing line, such as 45 degrees, the upper left part of the coding block is encoded using an intra prediction mode, and the lower right part of the coding block is encoded using an IBC prediction mode, and then they are averaged in the form of GPM to obtain a final prediction signal.
[0392] Combined IBC - Intra / Inter Prediction
[0393] According to one or more embodiments of the present disclosure, a coding block encoded in the IBC mode and a coding block encoded in the intra mode or the inter mode are combined. Different methods can be used to achieve this goal.
[0394] In the first method, the encoder / decoder can combine a coding block encoded in the IBC mode and a coding block encoded in the intra mode. Various methods can be utilized in this combination. In one example, similar to the CIIP technique in VVC, a coding block encoded in the IBC merge mode is regarded as a coding block encoded in the inter merge mode, and a coding block encoded in the IBC merge mode and a coding block encoded in the planar intra prediction mode are combined. In another example, similar to the combination of CIIP, TIMD, and TM merge techniques in ECM, a coding block encoded in the IBC merge-TM mode and a coding block encoded in the intra prediction mode derived from TIMD are combined.
[0395] In the second method, the encoder / decoder can combine a coding block encoded in the IBC mode and a coding block encoded in the inter mode. Various methods can be utilized in this combination. In one example, similar to the CIIP technique in VVC, a coding block encoded in the IBC merge mode is regarded as a coding block encoded in the planar intra mode, and a coding block encoded in the IBC merge mode and a coding block encoded in the inter merge mode are combined. In another example, a coding block encoded in the IBC merge mode is regarded as a coding block encoded in the inter merge mode, and a coding block encoded in the IBC merge mode and a coding block encoded in the inter merge mode are combined by equally averaging.
[0396] In the third method, the encoder / decoder can combine a coding block encoded in the IBC mode, a coding block encoded in the intra mode, and a coding block encoded in the inter mode. Various methods can be utilized in this combination. In one example, the coding block encoded in the IBC mode, the coding block encoded in the intra mode, and the coding block encoded in the inter mode are directly combined by equally averaging. In another example, first, the coding block encoded in the IBC mode is separately combined with the coding block encoded in the intra mode and the coding block encoded in the inter mode, as presented in the first method and the second method. Then, the separate combination results are combined by equally averaging.
[0397] Multiple Hypothesis IBC Prediction
[0398] According to one or more embodiments of the present disclosure, the number of block vectors (BVs) in the IBC tool is increased to 2 or more, and 2 or more hypotheses are combined to obtain a final prediction result. Different methods can be used to achieve this goal.
[0399] In the first method, the encoder / decoder can combine 2 hypotheses corresponding to 2 BVs to obtain a final prediction result. Various methods can be utilized to achieve this goal. In one example, the 2 BVs corresponding to the minimum rate distortion metric and the second minimum rate distortion metric in the IBC AMVP mode are equally averaged to obtain a final prediction result. In another example, the prediction result corresponding to the IBC AMVP mode and the prediction result corresponding to the IBC merge mode are equally averaged to obtain a final prediction result.
[0400] In the second method, the encoder / decoder can combine more hypotheses corresponding to more BVs to obtain a final prediction result. Various methods can be utilized to achieve this goal. In one example, the iterative accumulation method proposed in the multi-hypothesis prediction (MHP) technique is used to obtain a final prediction result. In another example, all the BVs corresponding to the minimum rate distortion metric, the second minimum rate distortion metric, the third minimum rate distortion metric,... in the IBC AMVP mode are equally averaged to obtain a final prediction result.
[0401] Combined Intra TMP - Intra / Inter Prediction
[0402] According to one or more embodiments of the present disclosure, an encoded block encoded in the intra TMP mode is combined with an encoded block encoded in the intra mode or the inter mode. Different methods can be used to achieve this goal.
[0403] In the first method, the encoder / decoder can combine an encoded block encoded in the intra TMP mode with an encoded block encoded in the intra mode. Various methods can be utilized in this combination. In one example, similar to the CIIP technique in VVC, the encoded block encoded in the intra TMP mode is regarded as an encoded block encoded in the inter merge mode, and it is combined with an encoded block encoded in the planar intra prediction mode. In another example, similar to the combination of CIIP, TIMD, and TM merge techniques in ECM, the encoded block encoded in the intra TMP mode is combined with an encoded block encoded in the intra prediction mode derived from TIMD.
[0404] In a second method, the encoder / decoder may combine an encoded block encoded in the intra TMP mode with an encoded block encoded in an inter mode. Various methods may be utilized in this combination. In one example, similar to the CIIP technique in VVC, the encoded block encoded in the intra TMP mode is regarded as an encoded block encoded in the planar intra mode, and it is combined with the encoded block encoded in the inter merge mode. In another example, the encoded block encoded in the intra TMP mode is regarded as an encoded block encoded in the inter merge mode, and it is combined with the encoded block encoded in the inter merge mode by equally averaging.
[0405] In a third method, the encoder / decoder may combine an encoded block encoded in the intra TMP mode with an encoded block encoded in the intra mode and an encoded block encoded in the inter mode. Various methods may be utilized in this combination. In one example, the encoded block encoded in the intra TMP mode, the encoded block encoded in the intra mode, and the encoded block encoded in the inter mode are directly combined by equally averaging. In another example, first, as presented in the first method and the second method, the encoded block encoded in the intra TMP mode is respectively combined with the encoded block encoded in the intra mode and the encoded block encoded in the inter mode. Then, the separate combination results are combined by equally averaging.
[0406] GPM with Intra TMP Prediction and Intra TMP Prediction
[0407] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the GPM tool in the form of GPM having intra TMP prediction and intra TMP prediction. Different methods may be used to achieve this goal.
[0408] In a first method, two "inter" parts of the GPM having an inter prediction method and an inter prediction method in VVC are replaced with intra TMP. This means that two intra TMP prediction results are weighted and averaged with each other according to the dividing line in the encoded block. The weights can be obtained by referring to the GPM having an inter prediction method and an inter prediction method in VVC.
[0409] In a second method, two "inter" parts of the GPM having an inter prediction method and an inter prediction method in ECM are replaced with intra TMP, where some template matching tools may be utilized to further improve the coding and decoding performance.
[0410] GPM with Intra TMP Prediction and Intra Prediction
[0411] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the GPM tool in the form of GPM having intra TMP prediction and intra prediction. Different methods may be used to achieve this goal.
[0412] In the first method, the "inter - frame" part of the ECM with both inter - frame prediction method and intra - frame prediction method is replaced by intra - frame TMP, where the intra - frame TMP prediction result and the intra - frame prediction result are weighted and averaged to obtain the final prediction signal.
[0413] GPM with Intra TMP Prediction and Inter Prediction
[0414] According to one or more embodiments of the present disclosure, the intra - frame TMP tool is combined with the GPM tool in the form of a GPM with intra - frame TMP prediction and inter - frame prediction. Different methods can be used to achieve this goal.
[0415] In the first method, one "inter - frame" part of the GPM in VVC with inter - frame prediction method and inter - frame prediction method is replaced by intra - frame TMP. The intra - frame TMP prediction result and the inter - frame merge prediction result are weighted and averaged to obtain the final prediction signal.
[0416] In the second method, one "inter - frame" part of the GPM in ECM with inter - frame prediction method and inter - frame prediction method is replaced by intra - frame TMP, where some template - matching tools can be utilized to further improve the coding and decoding performance.
[0417] Simplified Combination of Intra TMP Prediction and Intra Prediction in the Form of GPM
[0418] According to one or more embodiments of the present disclosure, the intra - frame TMP tool is combined with the GPM tool in the form of a simplified GPM with intra - frame TMP prediction and intra - frame prediction. For example, intra - frame TMP prediction and intra - frame prediction are combined in a specific partitioning mode, which can save the bit overhead of the partitioning representation. Different methods can be used to achieve this goal.
[0419] In the first method, for a dividing line, such as 45 degrees, the upper - left part of the coding block is encoded with an intra - frame prediction mode, the lower - right part of the coding block is encoded with an intra - frame TMP prediction mode, and then they are averaged in the form of GPM to obtain the final prediction signal.
[0420] Combining IBC Mode with TIMD Mode
[0421] According to one or more embodiments of the present disclosure, the IBC tool is combined with the TIMD tool. Different methods can be used to achieve this goal.
[0422] In the first method, the IBC mode is regarded as an intra - frame prediction mode added to the MPM list, then the IBC mode is compared with other intra - frame prediction modes in the MPM list using the template - matching cost, and finally the TIMD method is used to fuse the two modes with the minimum cost and the second - minimum cost to obtain the final prediction result.
[0423] In the second method, first obtain the conventional TIMD prediction result, then calculate the template matching cost between the IBC mode and the conventional TIMD prediction result, and finally use the TIMD method to fuse the IBC mode and the conventional TIMD prediction result to obtain the final prediction result.
[0424] Combining Intra TMP Mode with TIMD Mode
[0425] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the TIMD tool. Different methods can be used to achieve this goal.
[0426] In the first method, the intra TMP mode is regarded as an intra prediction mode added to the MPM list, then the template matching cost is used to compare the intra TMP mode with other intra prediction modes in the MPM list, and finally the TIMD method is used to fuse the two modes with the minimum cost and the second minimum cost to obtain the final prediction result.
[0427] In the second method, first obtain the conventional TIMD prediction result, then calculate the template matching cost between the intra TMP mode and the conventional TIMD prediction result, and finally use the TIMD method to fuse the intra TMP mode and the conventional TIMD prediction result to obtain the final prediction result.
[0428] Combining Intra TMP with LIC
[0429] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the LIC tool. Different methods can be used to achieve this goal.
[0430] In the first method, the intra TMP mode is regarded as an inter mode, and LIC is used to model the local illumination change between the current block template and the reference block template as a function of the local illumination change between the current block and its intra TMP prediction block. This function is a linear equation as used in the conventional LIC method.
[0431] Combining IBC with OBMC
[0432] According to one or more embodiments of the present disclosure, the IBC tool is combined with the OBMC tool. Different methods can be used to achieve this goal.
[0433] In the first method, the IBC mode is regarded as an inter mode, and the conventional OBMC method is applied to refine the top boundary pixels and left boundary pixels of the CU encoded by IBC by using the block vector information of neighboring blocks for weighted prediction.
[0434] In a second method, the IBC mode is regarded as an inter-frame mode, and an OBMC method based on template matching is applied to refine the top boundary pixels and the left boundary pixels of the CU encoded by IBC using the template matching-based method.
[0435] Combining Intra TMP with OBMC
[0436] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the OBMC tool. Different methods can be used to achieve this goal.
[0437] In a first method, the intra TMP mode is regarded as an inter-frame mode, and a conventional OBMC method is applied to perform weighted prediction using the block vector information of neighboring blocks to refine the top boundary pixels and the left boundary pixels of the CU encoded by intra TMP.
[0438] In a second method, the intra TMP mode is regarded as an inter-frame mode, and an OBMC method based on template matching is applied to refine the top boundary pixels and the left boundary pixels of the CU encoded by intra TMP using the template matching-based method.
[0439] Figure 21 A computing environment (or computing device) 1610 coupled to a user interface 1650 is shown. The computing environment 1610 can be part of a data processing server. In some embodiments, the computing device 1610 can execute any of the various methods or processes (such as encoding / decoding methods or processes) described above according to various examples of the present disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.
[0440] The processor 1620 generally controls the overall operation of the computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors to execute instructions to perform all or some of the steps of the above methods. In addition, the processor 1620 may include one or more modules that facilitate the interaction between the processor 1620 and other components. The processor can be a Central Processing Unit (CPU), a microprocessor, a single-chip microcomputer, a Graphical Processing Unit (GPU), etc.
[0441] The memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1630 may include predetermined software 1632. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, and the like. The memory 1630 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0442] The I / O interface 1640 provides an interface between the processor 1620 and peripheral interface modules (such as a keyboard, click wheel, buttons, etc.). The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1640 may be coupled to an encoder and a decoder.
[0443] Figure 22 is a flowchart showing a method for video decoding according to an example of the present disclosure.
[0444] In step 2201, on the decoder side, the processor 1620 may obtain a current CU encoded based on a combined mode that combines at least one of the following modes with an intra TMP mode: an intra prediction mode, an inter prediction mode, a TIMD mode, a LIC mode, or an OBMC mode.
[0445] In step 2202, the processor 1620 may obtain a final prediction for the current CU based on the combined mode.
[0446] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with an intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with an intra prediction mode; and obtain a final prediction for the current CU based on the first prediction and the second prediction. In some examples, the intra prediction mode may include one of the following modes: a planar intra prediction mode or an intra prediction mode derived from TIMD.
[0447] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with an intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with an inter prediction mode; and obtain a final prediction for the current CU based on the first prediction and the second prediction.
[0448] In some examples, the inter prediction mode may include an inter merge mode, and the processor 1620 may obtain a final prediction for the current CU by equally averaging the first prediction and the second prediction.
[0449] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with an intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with an intra prediction mode; obtain a third prediction for the current CU, where the third prediction is associated with an inter prediction mode; and obtain a final prediction for the current CU based on the first prediction, the second prediction, and the third prediction. In some examples, the processor 1620 may obtain a final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.
[0450] In some examples, the processor 1620 may obtain a first intermediate prediction based on the first prediction and the second prediction, obtain a second intermediate prediction based on the first prediction and the third prediction, and obtain a final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.
[0451] In some examples, the processor 1620 may obtain a most probable mode (MPM) list including an intra TMP mode and one or more other intra prediction modes, calculate a plurality of TM costs by comparing the intra TMP mode and the one or more other intra prediction modes, and select a first TM cost and a second TM cost from the plurality of TM costs, where the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list. In addition, the processor 1620 may obtain a final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
[0452] In some examples, the processor 1620 may obtain an intra TMP prediction for the current CU, where the intra TMP prediction is associated with an intra TMP mode; obtain a conventional TIMD prediction for the current CU, where the conventional TIMD prediction is associated with TIMD; and fuse the intra TMP prediction and the conventional TIMD prediction using a template-based intra mode derivation (TIMD) mode to obtain a final prediction for the current CU.
[0453] In some examples, the processor 1620 may obtain an intra TMP prediction for the current CU, where the intra TMP prediction is associated with an intra TMP mode; obtain a modeling function by modeling the local illumination change between the current CU and the intra TMP prediction using a function of the local illumination change between the current CU template and the reference block template, where the function is a linear equation; and obtain a locally illumination compensated intra TMP prediction based on the modeling function.
[0454] In some examples, the processor 1620 may obtain the current CU encoded based on the intra TMP mode and perform weighted prediction using the block vector information of neighboring blocks. For example, perform weighted prediction using the block vector information of neighboring blocks to refine the top boundary pixels and left boundary pixels of the current CU encoded based on the intra TMP mode. For example, assume that the current CU and the CU above it both use the intra TMP mode for encoding but use different block vectors. Then, to refine the top boundary pixels of the current CU, another prediction result of the top boundary pixels of the current CU is obtained using the block vector of the CU above, and then a weighted average of the original prediction result and the another prediction result is performed to obtain the final prediction.
[0455] In some examples, the processor 1620 may obtain the current CU encoded based on the intra TMP mode and use a template matching based method to refine the top boundary pixels and left boundary pixels of the current CU encoded based on the intra TMP mode.
[0456] Figure 23 is a flowchart showing a method for video encoding corresponding to the method for video decoding as shown in Figure 22 The flowchart shows a method for video encoding corresponding to the method for video decoding as shown in the figure.
[0457] In step 2301, on the encoder side, the processor 1620 may encode the current CU based on a combined mode that combines at least one of the following modes with the intra TMP mode: intra prediction mode, inter prediction mode, TIMD mode, LIC mode, or OBMC mode.
[0458] In step 2302, on the encoder side, the processor 1620 may send the current CU encoded based on the combined mode to the decoder.
[0459] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with the intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with the intra prediction mode; and obtain a final prediction for the current CU based on the first prediction and the second prediction. In some examples, the intra prediction mode may include one of the following modes: planar intra prediction mode or TIMD derived intra prediction mode.
[0460] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with an intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with an inter prediction mode; and obtain a final prediction for the current CU based on the first prediction and the second prediction.
[0461] In some examples, the inter prediction mode may include an inter merge mode, and the processor 1620 may obtain a final prediction for the current CU by equally averaging the first prediction and the second prediction.
[0462] In some examples, the processor 1620 may obtain a first prediction for the current CU, where the first prediction is associated with an intra TMP mode; obtain a second prediction for the current CU, where the second prediction is associated with an intra prediction mode; obtain a third prediction for the current CU, where the third prediction is associated with an inter prediction mode; and obtain a final prediction for the current CU based on the first prediction, the second prediction, and the third prediction. In some examples, the processor 1620 may obtain a final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.
[0463] In some examples, the processor 1620 may obtain a first intermediate prediction based on the first prediction and the second prediction, obtain a second intermediate prediction based on the first prediction and the third prediction, and obtain a final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.
[0464] In some examples, the processor 1620 may obtain a most probable mode (MPM) list including an intra TMP mode and one or more other intra prediction modes, calculate a plurality of TM costs by comparing the intra TMP mode and the one or more other intra prediction modes, and select a first TM cost and a second TM cost from the plurality of TM costs, where the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list. Additionally, the processor 1620 may obtain a final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
[0465] In some examples, the processor 1620 may obtain an intra TMP prediction for a current CU, where the intra TMP prediction is associated with an intra TMP mode; obtain a conventional TIMD prediction for the current CU, where the conventional TIMD prediction is associated with TIMD; and fuse the intra TMP prediction and the conventional TIMD prediction using a template-based intra mode derivation (TIMD) mode to obtain a final prediction for the current CU.
[0466] In some examples, the processor 1620 may obtain an intra TMP prediction for a current CU, where the intra TMP prediction is associated with an intra TMP mode; obtain a modeling function by modeling a local illumination change between the current CU and the intra TMP prediction using a function of a local illumination change between a current CU template and a reference block template, where the function is a linear equation; and obtain a locally illumination-compensated intra TMP prediction based on the modeling function.
[0467] In some examples, the processor 1620 may obtain a current CU encoded based on an intra TMP mode and perform weighted prediction using block vector information of neighboring blocks, e.g., perform weighted prediction using block vector information of neighboring blocks to refine top boundary pixels and left boundary pixels of the current CU encoded based on the intra TMP mode. For example, assuming that the current CU and the CU above it are both encoded using the intra TMP mode but use different block vectors, in order to refine the top boundary pixels of the current CU, another prediction result of the top boundary pixels of the current CU is obtained using the block vector of the CU above, and then a weighted average of the original prediction result and the another prediction result is performed to obtain a final prediction.
[0468] In some examples, the processor 1620 may obtain a current CU encoded based on an intra TMP mode and refine top boundary pixels and left boundary pixels of the current CU encoded based on the intra TMP mode using a template matching-based method.
[0469] Figure 24 is a flowchart showing a method for video decoding according to an example of the present disclosure.
[0470] In step 2401, at the decoder side, the processor 1620 may obtain a current CU encoded based on an intra template matching prediction (TMP) mode combined with a geometric partitioning mode (GPM) mode.
[0471] In step 2402, at the decoder side, the processor 1620 may obtain a final prediction for the current CU based on the intra TMP mode combined with the GPM mode.
[0472] In some examples, the current CU is divided into a first intra-TMP prediction part and a second intra-TMP prediction part, and the processor 1620 can obtain a first intra-TMP prediction for the first intra-TMP prediction part, obtain a second intra-TMP prediction for the second intra-TMP prediction part, and obtain a final prediction for the current CU based on the first intra-TMP prediction and the second intra-TMP prediction. In some examples, the processor 1620 can also obtain a reordered GPM partition pattern using a TM-based method, as discussed in the template matching-based reordering section for the GPM partition pattern.
[0473] In some examples, the processor 1620 can obtain a final prediction for the current CU by performing a weighted average of the first intra-TMP prediction and the second intra-TMP prediction.
[0474] In some examples, the current CU is divided into a first intra-TMP prediction part and a second intra prediction part. The processor 1620 can also obtain a first intra-TMP prediction for the first intra-TMP prediction part, obtain a second intra prediction for the second intra prediction part, and obtain a final prediction for the current CU based on the first intra-TMP prediction and the second intra prediction.
[0475] In some examples, the decoder can obtain a final prediction for the current CU by performing a weighted average of the first intra-TMP prediction and the second intra prediction.
[0476] In some examples, the current CU is divided into a first intra-TMP prediction part and a second inter prediction part. The processor 1620 can also obtain a first intra-TMP prediction for the first intra-TMP prediction part, obtain a second inter merge prediction for the second inter prediction part, and obtain a final prediction for the current CU by performing a weighted average of the first intra-TMP prediction and the second inter merge prediction. In some examples, the processor 1620 can also obtain a reordered GPM partition pattern using a TM-based method, as discussed in the template matching-based reordering section for the GPM partition pattern.
[0477] In some examples, the current CU is divided into a first part and a second part based on a predefined direction. Additionally, the processor 1620 can obtain a first intra prediction for the first part, obtain a second intra-TMP prediction for the second part, and obtain a final prediction for the current CU by performing a weighted average of the first intra prediction and the second intra-TMP prediction.
[0478] In some examples, the predefined direction is 45 degrees, the first intra prediction is located in the upper left part of the current CU, and the second intra-TMP prediction is located in the lower right part of the current CU.
[0479] Figure 25 is a flowchart showing a method for video encoding corresponding to the method for video decoding as Figure 24 shown therein.
[0480] In step 2501, on the encoder side, the processor 1620 may encode the current CU based on the intra TMP mode combined with the GPM mode.
[0481] In step 2502, on the encoder side, the processor 1620 may send the current CU encoded based on the intra TMP mode combined with the GPM mode to the decoder.
[0482] In some examples, the current CU is divided into a first intra TMP prediction part and a second intra TMP prediction part. Additionally, the processor 1620 may obtain a first intra TMP prediction for the first intra TMP prediction part, obtain a second intra TMP prediction for the second intra TMP prediction part, and obtain a final prediction for the current CU based on the first intra TMP prediction and the second intra TMP prediction.
[0483] In some examples, the current CU is divided into a first intra TMP prediction part and a second intra TMP prediction part, and the processor 1620 may obtain a first intra TMP prediction for the first intra TMP prediction part, obtain a second intra TMP prediction for the second intra TMP prediction part, and obtain a final prediction for the current CU based on the first intra TMP prediction and the second intra TMP prediction. In some examples, the processor 1620 may also obtain a reordered GPM partitioning mode using a TM-based method, as discussed in the template matching-based reordering section for the GPM partitioning mode.
[0484] In some examples, the processor 1620 may obtain a final prediction for the current CU by performing a weighted average of the first intra TMP prediction and the second intra TMP prediction.
[0485] In some examples, the current CU is divided into a first intra TMP prediction part and a second intra prediction part. The processor 1620 may also obtain a first intra TMP prediction for the first intra TMP prediction part, obtain a second intra prediction for the second intra prediction part, and obtain a final prediction for the current CU based on the first intra TMP prediction and the second intra prediction.
[0486] In some examples, on the encoder side, the processor 1620 may obtain a final prediction for the current CU by performing a weighted average of the first intra TMP prediction and the second intra prediction.
[0487] In some examples, the current CU is partitioned into a first intra-frame TMP prediction part and a second inter-frame prediction part. The processor 1620 may also obtain a first intra-frame TMP prediction for the first intra-frame TMP prediction part, obtain a second inter-frame merge prediction for the second inter-frame prediction part, and obtain a final prediction for the current CU by performing a weighted average of the first intra-frame TMP prediction and the second inter-frame merge prediction. In some examples, the processor 1620 may also obtain a reordered GPM partition mode using a TM-based method, as discussed in the template matching-based reordering section for the GPM partition mode.
[0488] In some examples, the current CU is partitioned into a first part and a second part based on a predefined direction. Additionally, the processor 1620 may obtain a first intra-frame prediction for the first part, obtain a second intra-frame TMP prediction for the second part, and obtain a final prediction for the current CU by performing a weighted average of the first intra-frame prediction and the second intra-frame TMP prediction.
[0489] In some examples, the predefined direction is 45 degrees, the first intra-frame prediction is located in the upper left part of the current CU, and the second intra-frame TMP prediction is located in the lower right part of the current CU.
[0490] Figure 26 is a flowchart showing a method for video decoding according to an example of the present disclosure.
[0491] In step 2601, on the decoder side, the processor 1620 may obtain a current CU encoded based on a combined mode that combines at least one of a TIMD mode or an OBMC mode with an IBC mode.
[0492] In step 2602, on the decoder side, the processor 1620 may obtain a final prediction for the current CU based on the combined mode.
[0493] In some examples, the processor 1620 may obtain a most probable mode (MPM) list including an IBC mode and one or more other intra-frame prediction modes, calculate a plurality of template matching (TM) costs by comparing the IBC mode and the one or more other intra-frame prediction modes, select a first TM cost and a second TM cost from the plurality of TM costs, where the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list, and obtain a final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
[0494] In some examples, the processor 1620 may obtain an IBC prediction for the current CU, obtain a conventional TIMD prediction for the current CU, and fuse the IBC mode and the conventional TIMD prediction using the TIMD mode to obtain a final prediction for the current CU, where the IBC prediction is associated with the IBC mode and the conventional TIMD prediction is associated with the TIMD mode.
[0495] In some examples, the processor 1620 may obtain a current CU encoded based on the IBC mode and perform weighted prediction using the block vector information of neighboring blocks, e.g., perform weighted prediction using the block vector information of neighboring blocks to refine the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode. For example, assume that the current CU and the CU above it are both encoded using the IBC mode but use different block vectors. To refine the top boundary pixels of the current CU, another prediction result of the top boundary pixels of the current CU is obtained using the block vector of the CU above, and then a weighted average of the original prediction result and the other prediction result is performed to obtain the final prediction result.
[0496] In some examples, the processor 1620 may obtain a current CU encoded based on the IBC mode and refine the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode using a template matching-based method.
[0497] Figure 27 is a flowchart showing a method for video encoding corresponding to the method for video decoding as Figure 26 shown.
[0498] In step 2701, on the encoder side, the processor 1620 may encode the current CU based on a combined mode that combines at least one of the TIMD mode or the OBMC mode with the IBC mode.
[0499] In step 2702, on the encoder side, the processor 1620 may send the current CU encoded based on the combined mode to the decoder.
[0500] In some examples, the processor 1620 may obtain a list of most probable modes (MPM) including the IBC mode and one or more other intra prediction modes, calculate a plurality of template matching (TM) costs by comparing the IBC mode and the one or more other intra prediction modes, select a first TM cost and a second TM cost from the plurality of TM costs, where the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list, and obtain a final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
[0501] In some examples, the processor 1620 may obtain an IBC prediction for the current CU, obtain a conventional TIMD prediction for the current CU, and fuse the IBC mode and the conventional TIMD prediction using the TIMD mode to obtain a final prediction for the current CU, where the IBC prediction is associated with the IBC mode, and the conventional TIMD prediction is associated with the TIMD mode.
[0502] In some examples, the processor 1620 may obtain a current CU encoded based on the IBC mode and perform weighted prediction using the block vector information of neighboring blocks, for example, perform weighted prediction using the block vector information of neighboring blocks to refine the top boundary pixels and the left boundary pixels of the current CU encoded based on the IBC mode. For example, assume that the current CU and the CU above it are both encoded using the IBC mode but use different block vectors. To refine the top boundary pixels of the current CU, another prediction result of the top boundary pixels of the current CU is obtained using the block vector of the CU above, and then a weighted average of the original prediction result and the another prediction result is performed to obtain the final prediction result.
[0503] In some examples, the processor 1620 may obtain a current CU encoded based on the IBC mode and refine the top boundary pixels and the left boundary pixels of the current CU encoded based on the IBC mode using a template matching-based method.
[0504] In some examples, a device for video coding and decoding is provided. The device includes a processor 1620 and a memory 1640 configured to store instructions executable by the processor; where the processor is configured to execute any method as shown in Figure 22 - Figure 27 as described.
[0505] In an embodiment, a non-transitory computer-readable storage medium including a plurality of programs is also provided. The plurality of programs are, for example, in a memory 1630 and can be executed by a processor 1620 in a computing environment 1610 to perform the above method and / or store the bitstream generated by the above encoding method or the bitstream decoded by the above decoding method. In one example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to receive (e.g., from Figure 2 the video encoder 20 therein) a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above decoding method according to the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to perform the above encoding method to encode video information (e.g., video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to send the bitstream or data stream (e.g., send to Figure 3 the video decoder 30 therein). Optionally, the non-transitory computer-readable storage medium can store therein a bitstream or data stream, and the bitstream or data stream includes encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.) generated by an encoder (e.g., Figure 2 the video encoder 20 therein) using (e.g.) the encoding method described above for use by a decoder (e.g., Figure 3 the video decoder 30 therein) when decoding video data. The non-transitory computer-readable storage medium can be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0506] In an embodiment, a bitstream generated by the above encoding method or a bitstream decoded by the above decoding method is provided. In an embodiment, a bitstream is provided that includes encoded video information generated by the above encoding method or encoded video information decoded by the above decoding method.
[0507] In an embodiment, a computing device is also provided, which includes one or more processors (e.g., the processor 1620); and a non-transitory computer-readable storage medium or a memory 1630 storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the above method when executing the plurality of programs.
[0508] In an embodiment, a computer program product with instructions is also provided, where the instructions are for storing or transmitting a bitstream, and the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product including multiple programs is also provided, where the multiple programs, for example, in the memory 1630, can be executed by the processor 1620 in the computing environment 1610 for performing the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0509] In an embodiment, the computing environment 1610 can be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0510] In an embodiment, a method for storing a bitstream is also provided, including storing the bitstream on a digital storage medium, where the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.
[0511] In an embodiment, a method for transmitting the bitstream generated by the above encoder is also provided. In an embodiment, a method for receiving the bitstream to be decoded by the above decoder is also provided.
[0512] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art who benefit from the teachings presented in the foregoing description and the associated drawings.
[0513] Unless otherwise specifically stated, the order of the method steps according to the present disclosure is only illustrative, and the method steps according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual conditions. In addition, at least one of the method steps according to the present disclosure can be adjusted, combined, or deleted according to actual requirements.
[0514] The examples are selected and described to explain the principles of the present disclosure and enable other technicians in the art to understand various embodiments of the present disclosure and best utilize the basic principles and various embodiments with various modifications suitable for the intended specific purposes. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the present disclosure.
Claims
1. A method for video decoding, comprising: obtaining, by a decoder, a current coding unit (CU) encoded based on a combined mode, the combined mode combining at least one of the following modes with an intra-template matching prediction (TMP) mode: an intra prediction mode, an inter prediction mode, a template-based intra mode derivation (TIMD) mode, a local illumination compensation (LIC) mode, or an overlapping block motion compensation (OBMC) mode; and obtaining, by the decoder, a final prediction for the current CU based on the combined mode.
2. The method according to claim 1, further comprising: obtaining, by the decoder, a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; obtaining, by the decoder, a second prediction for the current CU, wherein the second prediction is associated with the intra prediction mode; and obtaining, by the decoder, the final prediction for the current CU based on the first prediction and the second prediction.
3. The method according to claim 2, wherein the intra prediction mode includes one of the following modes: a planar intra prediction mode or a TIMD-derived intra prediction mode.
4. The method according to claim 1, further comprising: obtaining, by the decoder, a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; obtaining, by the decoder, a second prediction for the current CU, wherein the second prediction is associated with the inter prediction mode; and obtaining, by the decoder, the final prediction for the current CU based on the first prediction and the second prediction.
5. The method according to claim 4, wherein the inter prediction mode includes an inter merge mode, and the method further comprises: obtaining, by the decoder, the final prediction for the current CU by equally averaging the first prediction and the second prediction.
6. The method according to claim 1, further comprising: obtaining, by the decoder, a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; obtaining, by the decoder, a second prediction for the current CU, wherein the second prediction is associated with the intra prediction mode; obtaining, by the decoder, a third prediction for the current CU, wherein the third prediction is associated with the inter prediction mode; and obtaining, by the decoder, the final prediction for the current CU based on the first prediction, the second prediction, and the third prediction.
7. The method according to claim 6, further comprising: obtaining, by the decoder, the final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.
8. The method according to claim 6, further comprising: obtaining, by the decoder, a first intermediate prediction based on the first prediction and the second prediction; obtaining, by the decoder, a second intermediate prediction based on the first prediction and the third prediction; and The decoder obtains the final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.
9. The method according to claim 1, further comprising: The decoder obtains a most probable mode MPM list including the intra TMP mode and one or more other intra prediction modes; The decoder calculates a plurality of template matching TM costs by comparing the intra TMP mode and the one or more other intra prediction modes; The decoder selects a first TM cost and a second TM cost from the plurality of TM costs, wherein the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list; and The decoder obtains the final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
10. The method according to claim 1, further comprising: The decoder obtains an intra TMP prediction for the current CU, wherein the intra TMP prediction is associated with the intra TMP mode; The decoder obtains a conventional TIMD prediction for the current CU, wherein the conventional TIMD prediction is associated with the TIMD mode; and The decoder fuses the intra TMP prediction and the conventional TIMD prediction using a template-based intra mode derivation TIMD mode to obtain the final prediction for the current CU.
11. The method according to claim 1, further comprising: The decoder obtains an intra TMP prediction for the current CU, wherein the intra TMP prediction is associated with the intra TMP mode; and The decoder models the local illumination change between the current CU and the intra TMP prediction by using a function of the local illumination change between the current CU template and the reference block template, wherein the function is a linear equation; and The decoder obtains a locally illumination-compensated intra TMP prediction based on the modeling function.
12. The method according to claim 1, further comprising: The decoder obtains the current CU encoded based on the intra TMP mode; and The decoder refines the top boundary pixels and the left boundary pixels of the current CU encoded based on the intra TMP mode by performing weighted prediction using the block vector information of neighboring blocks.
13. The method according to claim 1, further comprising: The decoder obtains the current CU encoded based on the intra TMP mode; and The decoder refines the top boundary pixels and the left boundary pixels of the current CU encoded based on the intra TMP mode by using a template matching-based method.
14. A method for video coding, comprising: The encoder encodes the current coding unit (CU) based on a combined mode that combines at least one of the following modes with an intra-template matching prediction (TMP) mode: an intra prediction mode, an inter prediction mode, a template-based intra mode derivation (TIMD) mode, a local illumination compensation (LIC) mode, or an overlapped block motion compensation (OBMC) mode; and the encoder sends the current CU encoded based on the combined mode to the decoder.
15. The method according to claim 14, further comprising: the encoder obtaining a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; the encoder obtaining a second prediction for the current CU, wherein the second prediction is associated with the intra prediction mode; and the encoder obtaining the final prediction for the current CU based on the first prediction and the second prediction.
16. The method according to claim 15, wherein the intra prediction mode comprises one of the following modes: a planar intra prediction mode or an intra prediction mode derived from TIMD.
17. The method according to claim 14, further comprising: the encoder obtaining a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; the encoder obtaining a second prediction for the current CU, wherein the second prediction is associated with the inter prediction mode; and the encoder obtaining the final prediction for the current CU based on the first prediction and the second prediction.
18. The method according to claim 17, wherein the inter prediction mode comprises an inter merge mode, and the method further comprises: the encoder obtaining the final prediction for the current CU by equally averaging the first prediction and the second prediction.
19. The method according to claim 14, further comprising: the encoder obtaining a first prediction for the current CU, wherein the first prediction is associated with the intra TMP mode; the encoder obtaining a second prediction for the current CU, wherein the second prediction is associated with the intra prediction mode; the encoder obtaining a third prediction for the current CU, wherein the third prediction is associated with the inter prediction mode; and the encoder obtaining the final prediction for the current CU based on the first prediction, the second prediction, and the third prediction.
20. The method according to claim 19, further comprising: the encoder obtaining the final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.
21. The method according to claim 19, further comprising: the encoder obtaining a first intermediate prediction based on the first prediction and the second prediction; the encoder obtaining a second intermediate prediction based on the first prediction and the third prediction; and The encoder obtains the final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.
22. The method according to claim 14, further comprising: The encoder obtains a most probable mode MPM list including the intra TMP mode and one or more other intra prediction modes; The encoder calculates a plurality of template matching TM costs by comparing the intra TMP mode and the one or more other intra prediction modes; The encoder selects a first TM cost and a second TM cost from the plurality of TM costs, wherein the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list; and The encoder obtains the final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
23. The method according to claim 14, further comprising: The encoder obtains an intra TMP prediction for the current CU, wherein the intra TMP prediction is associated with the intra TMP mode; The encoder obtains a conventional TIMD prediction for the current CU, wherein the conventional TIMD prediction is associated with the TIMD mode; and The encoder fuses the intra TMP prediction and the conventional TIMD prediction using a template-based intra mode derivation TIMD mode to obtain the final prediction for the current CU.
24. The method according to claim 14, further comprising: The encoder obtains an intra TMP prediction for the current CU, wherein the intra TMP prediction is associated with the intra TMP mode; The encoder obtains a modeling function by modeling the local illumination change between the current CU and the intra TMP prediction using a function of the local illumination change between the current CU template and the reference block template, wherein the function is a linear equation; and The encoder obtains a locally illumination-compensated intra TMP prediction based on the modeling function.
25. The method according to claim 14, further comprising: The encoder obtains the current CU encoded based on the intra TMP mode; and The encoder refines the top boundary pixels and the left boundary pixels of the current CU encoded based on the intra TMP mode by performing weighted prediction using the block vector information of neighboring blocks.
26. The method according to claim 14, further comprising: The encoder obtains the current CU encoded based on the intra TMP mode; and The encoder refines the top boundary pixels and the left boundary pixels of the current CU encoded based on the intra TMP mode using a template matching-based method.
27. A method for video decoding, comprising: Obtain, by a decoder, a current coding unit (CU) encoded based on an intra-template matching prediction (TMP) mode combined with a geometric partitioning mode (GPM) mode; and Obtain, by the decoder, a final prediction for the current CU based on the intra-TMP mode combined with the GPM mode.
28. The method according to claim 27, wherein the current CU is partitioned into a first intra-TMP prediction part and a second intra-TMP prediction part, and wherein the method further comprises: Obtain, by the decoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the decoder, a second intra-TMP prediction for the second intra-TMP prediction part; and Obtain, by the decoder, the final prediction for the current CU based on the first intra-TMP prediction and the second intra-TMP prediction.
29. The method according to claim 28, further comprising: Obtain, by the decoder, the final prediction for the current CU by performing a weighted average on the first intra-TMP prediction and the second intra-TMP prediction.
30. The method according to claim 28, further comprising: Obtain, by the decoder, a reordered GPM partitioning mode by using a method based on template matching (TM).
31. The method according to claim 27, wherein the current CU is partitioned into a first intra-TMP prediction part and a second inter prediction part, and wherein the method further comprises: Obtain, by the decoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the decoder, a second inter prediction for the second inter prediction part; and Obtain, by the decoder, the final prediction for the current CU based on the first intra-TMP prediction and the second inter prediction.
32. The method according to claim 31, further comprising: Obtain, by the decoder, the final prediction for the current CU by performing a weighted average on the first intra-TMP prediction and the second inter prediction.
33. The method according to claim 27, wherein the current CU is partitioned into a first intra-TMP prediction part and a second inter prediction part, and wherein the method further comprises: Obtain, by the decoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the decoder, a second inter merge prediction for the second inter prediction part; and Obtain, by the decoder, the final prediction for the current CU by performing a weighted average on the first intra-TMP prediction and the second inter merge prediction.
34. The method according to claim 33, further comprising: Obtain, by the decoder, a reordered GPM partitioning mode by using a method based on template matching (TM).
35. The method according to claim 27, wherein the current CU is partitioned into a first part and a second part based on a predefined direction, wherein the method further comprises: Obtain, by the decoder, a first intra prediction for the first part; Obtain, by the decoder, an intra-TMP prediction for the second part; and Obtain, by the decoder, the final prediction for the current CU by performing a weighted average of the intra prediction of the first frame and the intra-TMP prediction of the second frame.
36. The method according to claim 35, wherein the predefined direction is 45 degrees, the intra prediction of the first frame is located in the upper left part of the current CU, and the intra-TMP prediction of the second frame is located in the lower right part of the current CU.
37. A method for video coding, comprising: Encode, by an encoder, a current coding unit CU based on an intra-template matching prediction TMP mode combined with a geometric partitioning mode GPM mode; and Send, by the encoder, the current CU encoded based on the intra-TMP mode combined with the GPM mode to a decoder.
38. The method according to claim 37, wherein the current CU is divided into a first intra-TMP prediction part and a second intra-TMP prediction part, and wherein the method further comprises: Obtain, by the encoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the encoder, a second intra-TMP prediction for the second intra-TMP prediction part; and Obtain, by the encoder, the final prediction for the current CU based on the first intra-TMP prediction and the second intra-TMP prediction.
39. The method according to claim 38, further comprising: Obtain, by the encoder, the final prediction for the current CU by performing a weighted average of the first intra-TMP prediction and the second intra-TMP prediction.
40. The method according to claim 38, further comprising: Obtain, by the encoder, a reordered GPM partitioning mode using a method based on template matching TM.
41. The method according to claim 37, wherein the current CU is divided into a first intra-TMP prediction part and a second intra prediction part, and wherein the method further comprises: Obtain, by the encoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the encoder, a second intra prediction for the second intra prediction part; and Obtain, by the encoder, the final prediction for the current CU based on the first intra-TMP prediction and the second intra prediction.
42. The method according to claim 41, further comprising: Obtain, by the encoder, the final prediction for the current CU by performing a weighted average of the first intra-TMP prediction and the second intra prediction.
43. The method according to claim 37, wherein the current CU is divided into a first intra-TMP prediction part and a second inter prediction part, and wherein the method further comprises: Obtain, by the encoder, a first intra-TMP prediction for the first intra-TMP prediction part; Obtain, by the encoder, a second inter prediction for the second inter prediction part; and The encoder obtains the final prediction for the current CU by performing a weighted average of the in - frame TMP prediction in the first frame and the inter - frame merge prediction in the second frame.
44. The method according to claim 43, further comprising: The encoder obtains a reordered GPM partition mode using a method based on template matching TM.
45. The method according to claim 37, wherein the current CU is partitioned into a first part and a second part based on a predefined direction, wherein the method further comprises: The encoder obtains an in - frame prediction for the first part; The encoder obtains an in - frame TMP prediction for the second part; and The encoder obtains the final prediction for the current CU by performing a weighted average of the in - frame prediction and the in - frame TMP prediction.
46. The method according to claim 45, wherein the predefined direction is 45 degrees, the in - frame prediction is located in the upper - left part of the current CU, and the in - frame TMP prediction is located in the lower - right part of the current CU.
47. A method for video decoding, comprising: The decoder obtains a current coding unit CU encoded based on a combined mode, the combined mode combining at least one of an in - frame mode derived based on template TIMD mode or an overlapping - block motion - compensation OBMC mode with an in - frame block - copy IBC mode; and The decoder obtains a final prediction for the current CU based on the combined mode.
48. The method according to claim 47, further comprising: The decoder obtains a most - probable - mode MPM list, the MPM list including the IBC mode and one or more other in - frame prediction modes; The decoder calculates a plurality of template - matching TM costs by comparing the IBC mode with the one or more other in - frame prediction modes; The decoder selects a first TM cost and a second TM cost from the plurality of TM costs, wherein the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second - minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list; and The decoder obtains the final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
49. The method according to claim 47, further comprising: The decoder obtains an IBC prediction for the current CU, wherein the IBC prediction is associated with the IBC mode; The decoder obtains a conventional TIMD prediction for the current CU, wherein the conventional TIMD prediction is associated with the TIMD mode; The decoder fuses the IBC mode and the conventional TIMD prediction using the TIMD mode to obtain the final prediction for the current CU.
50. The method according to claim 47, further comprising: Obtaining, by the decoder, the current CU encoded based on the IBC mode; and Refining, by the decoder, the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode by performing weighted prediction using block vector information of neighboring blocks.
51. The method according to claim 47, further comprising: Obtaining, by the decoder, the current CU encoded based on the IBC mode; and Refining, by the decoder, the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode using a method based on template matching.
52. A method for video coding, comprising: Encoding, by an encoder, a current coding unit (CU) based on a combined mode, the combined mode combining at least one of a template-based intra mode derivation (TIMD) mode or an overlapping block motion compensation (OBMC) mode with an intra block copy (IBC) mode; and Sending, by the encoder, the current CU encoded based on the combined mode to a decoder.
53. The method according to claim 52, further comprising: Obtaining, by the encoder, a most probable mode (MPM) list, the MPM list including the IBC mode and one or more other intra prediction modes; Calculating, by the encoder, a plurality of template matching (TM) costs by comparing the IBC mode with the one or more other intra prediction modes; Selecting, by the encoder, a first TM cost and a second TM cost from the plurality of TM costs, wherein the first TM cost is the minimum TM cost among the plurality of TM costs and is associated with a first mode in the MPM list, and the second TM cost is the second minimum TM cost among the plurality of TM costs and is associated with a second mode in the MPM list; and Obtaining, by the encoder, a final prediction for the current CU by fusing the first mode and the second mode using the TIMD mode.
54. The method according to claim 52, further comprising: Obtaining, by the encoder, an IBC prediction for the current CU, wherein the IBC prediction is associated with the IBC mode; Obtaining, by the encoder, a conventional TIMD prediction for the current CU, wherein the conventional TIMD prediction is associated with the TIMD mode; and Fusing, by the encoder, the IBC mode and the conventional TIMD prediction using the TIMD mode to obtain the final prediction for the current CU.
55. The method according to claim 52, further comprising: Obtaining, by the encoder, the current CU encoded based on the IBC mode; and Refining, by the encoder, the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode by performing weighted prediction using block vector information of neighboring blocks.
56. The method according to claim 52, further comprising: Obtaining, by the encoder, the current CU encoded based on the IBC mode; and The encoder refines the top boundary pixels and left boundary pixels of the current CU encoded based on the IBC mode using a template matching-based method.
57. An apparatus for video decoding, comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method according to any one of claims 1-13, 27-36, and 47-51.
58. An apparatus for video encoding, comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method according to any one of claims 14-26, 37-46, and 52-56.
59. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to any one of claims 1-13, 27-36, and 47-51.
60. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to any one of claims 14-26, 37-46, and 52-56.
61. A non-transitory computer-readable storage medium for storing a bitstream decoded by the method according to any one of claims 1-13, 27-36, and 47-51.
62. A non-transitory computer-readable storage medium for storing a bitstream generated by the method according to any one of claims 14-26, 37-46, and 52-56.