Method and apparatus for intra block copy and intra template matching
By improving the methods of intra-frame block copying and intra-frame template matching, the video encoding and decoding efficiency is improved, the problem of low prediction efficiency of intra-frame block copying and intra-frame template matching in the existing technology is solved, and more efficient video data compression and quality preservation are achieved.
Patent Information
- Application Number
- CN202480009695.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-03
- Filing Date
- 2024-02-03
- Publication Date
- 2025-09-05
AI Technical Summary
Existing video coding and decoding technologies are inefficient in intra-frame block copying and intra-frame template matching prediction, resulting in poor video data compression effects.
By improving the intra-block copy method, the decoder and encoder obtain the final prediction of the current block based on multiple prediction block candidates, use the intra-block copy mode to improve encoding and decoding efficiency, and combine the intra-frame template matching technology to optimize the prediction process.
It improves the efficiency of video encoding and decoding, reduces the bit rate of video data, and maintains or improves video quality.
Smart Images

Figure CN120604510A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based upon and claims the benefit of U.S. Provisional Application No. 63 / 443,322, filed on February 3, 2023, entitled “Methods and Devices For Intra Block Copy and Intra Template Matching,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] The present disclosure relates to video coding and compression, and in particular, but not limited to, methods and apparatus for improving coding efficiency of intra block copy (IBC) and intra template matching prediction (Intra TMP). Background Art
[0004] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data through a communication network, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited storage resources of the storage device, before the video data is transmitted or stored, video codecs can be used to compress the video data according to one or more video codec standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Experts Group (MPEG) encoding, etc. Video codecs typically utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilize the redundancy inherent in video data. Video codecs are intended to compress video data into a form using a lower bit rate while avoiding or minimizing the decline in video quality. Summary of the Invention
[0005] This disclosure provides examples of techniques for improving intra-block copying methods in video encoding or decoding processes.
[0006] According to a first aspect of the present disclosure, a method for video decoding is provided. In this method, a decoder may obtain multiple prediction block candidates based on an intra-block copy (IBC) mode for a current block. Furthermore, the decoder may obtain a final prediction of the current block based on the multiple prediction block candidates.
[0007] According to a second aspect of the present disclosure, a method for video encoding is provided. In this method, an encoder may obtain multiple prediction block candidates based on an intra block copy (IBC) mode for a current block. Furthermore, the encoder may obtain a final prediction for the current block based on the multiple prediction block candidates. Furthermore, the encoder may encode the final prediction for the current block. Furthermore, the encoder may send the final prediction encoded based on the IBC mode to a decoder.
[0008] According to a third aspect of the present disclosure, a device for video decoding is provided. The device may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, when executing the instructions, the one or more processors are configured to perform the method described in the first aspect.
[0009] According to a fourth aspect of the present disclosure, a device for video encoding is provided. The device may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, when executing the instructions, the one or more processors are configured to perform the method described in the second aspect.
[0010] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium for storing computer-executable instructions is provided, wherein the computer-executable instructions, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect.
[0011] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium for storing computer-executable instructions is provided, wherein the computer-executable instructions, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to the second aspect.
[0012] According to a seventh aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first aspect.
[0013] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by the method according to the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] A more particular description of examples of the present disclosure will be presented by reference to specific examples shown in the accompanying drawings. Given that these drawings depict only some examples and are therefore not to be considered limiting of scope, the examples will be described and explained with additional specificity and detail through use of the accompanying drawings.
[0015] Figure 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some examples of the present disclosure.
[0016] Figure 2 is a block diagram illustrating an exemplary video encoder according to some examples of the present disclosure.
[0017] Figure 3 is a block diagram illustrating an exemplary video decoder according to some examples of the present disclosure.
[0018] Figures 4A to 4E is a block diagram illustrating how a frame may be recursively partitioned into multiple video blocks of different sizes and shapes according to some examples of the present disclosure.
[0019] Figure 5 A schematic diagram illustrating locations of spatial candidates according to some examples of the present disclosure.
[0020] Figure 6 A schematic diagram illustrating candidate pairs considered for redundancy checking of spatial candidates according to some examples of the present disclosure is shown.
[0021] Figure 7 A schematic diagram illustrating scaling of motion vectors for temporal candidates according to some examples of the present disclosure is shown.
[0022] Figure 8 A schematic diagram illustrating candidate positions of temporal candidates according to some examples of the present disclosure is shown.
[0023] Figure 9 A schematic diagram illustrating a merge mode using motion vector difference (MMVD) search point according to some examples of the present disclosure is shown.
[0024] Figure 10 Unidirectional prediction motion vector selection for geometric partitioning mode (GPM) according to some examples of the present disclosure is shown.
[0025] Figure 11 Shown are top and left neighboring blocks used in CIIP weight derivation according to some examples of the present disclosure.
[0026] Figure 12 The current CTU processing order and its available reference samples in the current and left CTUs according to some examples of the present disclosure are shown.
[0027] Figure 13 Padding candidates for replacing zero vectors in an IBC list according to some examples of the present disclosure are shown.
[0028] Figure 14The reference region used for IBC when CTU(m,n) is encoded is shown. According to some examples of the present disclosure, a blue block represents a current CTU; a green block represents a reference region; and a white block represents an invalid reference region.
[0029] Figure 15 IBC reference areas for camera-captured content according to some examples of the present disclosure are shown.
[0030] 16A to 16B A method of dividing angle patterns according to some examples of the present disclosure is shown.
[0031] 17A to 17D A GPM with inter and intra prediction is shown. 17A to 17C The available IPM candidates are shown. Figure 17D Examples of GPM with intra-frame and intra-frame prediction are shown according to some examples of the present disclosure.
[0032] Figure 18 Edges on a template according to some examples of the present disclosure are shown.
[0033] Figure 19 The intra-frame template matching search area used according to some examples of the present disclosure is shown.
[0034] Figure 20 Templates for template matching based OBMC according to some examples of the present disclosure are shown.
[0035] Figure 21 Partitioning methods and corresponding weights of intra-coded blocks for angular and planar modes according to some examples of the present disclosure are shown.
[0036] Figure 22 is a schematic diagram illustrating a computing environment coupled with a user interface according to some examples of the present disclosure.
[0037] Figure 23 is a workflow illustrating a method for video decoding according to some examples of the present disclosure.
[0038] Figure 24 is a workflow illustrating a method for video encoding according to some examples of the present disclosure, which corresponds to Figure 23 The method for video decoding is shown. DETAILED DESCRIPTION
[0039] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to facilitate understanding of the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices with digital video capabilities.
[0040] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a / an," "the," and "the" in this disclosure and the appended claims are intended to include the plural forms, unless otherwise clearly indicated throughout this disclosure. It should also be understood that the term "and / or" as used in this disclosure refers to and includes one or any or all possible combinations of the listed multiple associated items.
[0041] Reference throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language means that the particular feature, structure, or characteristic being described is included in at least one embodiment or example. Unless expressly stated otherwise, a feature, structure, element, or characteristic described in conjunction with one or some embodiments may also be applicable to other embodiments.
[0042] Throughout the disclosure, the terms "first," "second," "third," and the like are used as nomenclature to refer only to related elements, such as devices, components, compositions, steps, and the like, without implying any spatial or temporal order, unless expressly stated otherwise. For example, "first device" and "second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be arbitrarily named.
[0043] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to or located adjacent to each other.
[0044] As used herein, the terms "if" or "when" may be understood to mean "when" or "in response to" depending on the context. If these terms appear in a claim, they may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include the following steps: i) when or if condition X exists, perform function or action X', and ii) when or if condition Y exists, perform function or action Y'. The method may be implemented with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, both functions X' and Y' may be performed at different times during multiple executions of the method.
[0045] A unit or module may be implemented purely by software, purely by hardware, or a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components that are linked together directly or indirectly to perform a specific function.
[0046] Figure 1 FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1 As shown in , system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including a cloud server, a server computer, a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0047] In some embodiments, target device 14 can receive the encoded video data to be decoded via link 16. Link 16 can include any type of communication medium or device capable of moving the encoded video data from source device 12 to target device 14. In one example, link 16 can include a communication medium that enables source device 12 to send the encoded video data directly to target device 14 in real time. The encoded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to target device 14. The communication medium can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form a portion of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium can include a router, a switch, a base station, or any other device that can facilitate communication from source device 12 to target device 14.
[0048] In some other embodiments, the encoded video data can be sent from the output interface 22 to a storage device 32. The encoded video data in the storage device 32 can then be accessed by the target device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending the encoded video data to the target device 14. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data through any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both streaming and download transmissions.
[0049] like Figure 1 As shown in , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include a source such as a video capture device (e.g., a video camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application may be generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.
[0050] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be sent directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may also include a modem and / or a transmitter.
[0051] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0052] In some implementations, target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0053] The video encoder 20 and the video decoder 30 can operate according to a proprietary standard or an industry standard (e.g., VVC, HEVC, Part 10 of MPEG-4, AVC) or an extension of such a standard. It should be understood that the present application is not limited to a specific video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally believed that the video encoder 20 of the source device 12 can be configured to encode the video data according to any of these current standards or future standards. Similarly, it is also generally believed that the video decoder 30 of the target device 14 can be configured to decode the video data according to any of these current standards or future standards.
[0054] The video encoder 20 and the video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuits, respectively, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and either encoder or decoder can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.
[0055] In some embodiments, components of source device 12 (e.g., video source 18, video encoder 20, or the like) may be configured to: Figure 2 The components included in the video encoder 20 and the output interface 22) and / or the components of the target device 14 (for example, the input interface 28, the video decoder 30 or the following reference Figure 3The components included in the video decoder 30 and at least a portion of the components in the display device 34 may be operated in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), wherein the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be set in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In one embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network can be a private cloud, a public cloud, or a hybrid cloud. Terms such as "cloud," "cloud computing," and "cloud-based" may be used interchangeably herein without departing from the scope of this disclosure. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service network. Instead, this disclosure may also be implemented within any other type of computing environment currently known or developed in the future.
[0056] Figure 2 FIG2 is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described herein. Video encoder 20 can perform intra-frame prediction coding and inter-frame prediction coding on video blocks within a video frame. Intra-frame prediction coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that in the field of video coding, the term "frame" can be used as a synonym for the term "image" or "picture."
[0057] like Figure 2As shown in FIG, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 also includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58 for video block reconstruction, an inverse transform processing unit 60, and an adder 62. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can also be used to filter the output of the adder 62. It should be noted that with respect to the CCSAO technique, the present application is not limited to the embodiments described herein, but may also be applied to the case where an offset is selected for any other of the luma component, the Cb chroma component, and the Cr chroma component based on any of the luma component, the Cb chroma component, and the Cr chroma component to modify the other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein may be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component mentioned herein may be any other one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component mentioned herein may be the remaining component of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filter may be omitted, and the decoded video block may be provided directly by the adder 62 to the DPB 64. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the fixed or programmable hardware units described.
[0058] Video data memory 40 may store video data to be encoded by the components of video encoder 20. Figure 1 The video source 18 shown obtains video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by the video encoder 20 (e.g., in intra-frame or inter-frame prediction coding mode) when encoding the video data. The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.
[0059] like Figure 2As shown in , after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined splitting structure associated with the video data, such as a quadtree (QT) structure. A video frame is or can be viewed as a two-dimensional array or matrix of sample values. The samples in the array may also be referred to as pixels or picture elements (pel). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame may be divided into multiple video blocks by using QT segmentation. A video block is again or can be viewed as a two-dimensional array or matrix of sample values, but its dimensions are smaller than the dimensions of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block may be further partitioned into one or more block partitions or sub-blocks (which may again form blocks) by, for example, iteratively using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the terms "block" or "video block" as used herein may be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or may be or correspond to a corresponding block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.
[0060] The prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block based on the error results (e.g., coding rate and distortion level), such as one of one or more inter-frame prediction coding modes among a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block or inter-frame prediction coding block to the adder 50 to generate a residual block and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.
[0061] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.
[0062] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for the current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, where the motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. For example, a motion vector may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded within the current frame. The predetermined pattern may designate the video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors (e.g., block vectors) for intra BC coding in a manner similar to the motion vectors determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vectors.
[0063] In terms of pixel differences, the prediction block for a video block may be or may correspond to a block or reference block of a reference frame that is considered to closely match the video block to be encoded, and the pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches relative to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.
[0064] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing the position of the video block to the position of a prediction block of a reference frame selected from either a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.
[0065] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block may include luma component differences, chroma component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are described separately for conceptual purposes.
[0066] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example during separate encoding passes, and test their performance using rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original, uncoded block that was coded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate ratios based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0067] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the prediction block may include calculating values for sub-integer pixel positions.
[0068] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from a different frame according to inter-frame prediction, video encoder 20 can form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values forming the residual video block may include both luma component differences and chroma component differences.
[0069] As an alternative to the inter-frame prediction performed by motion estimation unit 42 and motion compensation unit 44 or the intra-frame block copy prediction performed by intra BC unit 48 as described above, intra-frame prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, intra-frame prediction processing unit 46 can determine an intra-frame prediction mode to use for encoding the current block. To do so, intra-frame prediction processing unit 46 can use various intra-frame prediction modes to encode the current block, for example, during separate encoding passes, and intra-frame prediction processing unit 46 (or in some examples, mode selection unit) can select an appropriate intra-frame prediction mode to use from the tested intra-frame prediction modes. Intra-frame prediction processing unit 46 can provide information indicating the intra-frame prediction mode selected for the block to entropy coding unit 56. Entropy coding unit 56 can encode the information indicating the selected intra-frame prediction mode into the bitstream.
[0070] After prediction processing unit 41 determines a prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0071] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan on the matrix comprising the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0072] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to a video bitstream such as Figure 1 The video decoder 30 shown, or archived as Figure 1 The video frame is shown in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.
[0073] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for use in generating a reference block for predicting other video blocks. As noted above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0074] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block to inter-predict another video block in a subsequent video frame.
[0075] Figure 3 is a block diagram illustrating another exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform the operations described above in conjunction with the above. Figure 2 The encoding process is essentially the inverse of the decoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.
[0076] In some examples, units of the video decoder 30 may be tasked with performing embodiments of the present application. Furthermore, in some examples, embodiments of the present disclosure may be dispersed across one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (e.g., the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functionality of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., the motion compensation unit 82).
[0077] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (e.g., a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 (e.g., in intra-frame or inter-frame prediction coding mode) when decoding the video data. The video data memory 79 and the DPB 92 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are stored in Figure 3 92 as two distinct components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, video data memory 79 may be on-chip with the other components of video decoder 30, or off-chip relative to those components.
[0078] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators, and other syntax elements to the prediction processing unit 81.
[0079] When a video frame is encoded as an intra-frame prediction coded (I) frame or for intra-frame coded prediction blocks in other types of frames, intra-frame prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the signaled intra-frame prediction mode and reference data from a previously decoded block of the current frame.
[0080] When the video frame is encoded as an inter-frame prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vector and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can use a default construction technique to construct the reference frame lists, i.e., List 0 and List 1, based on the reference frames stored in the DPB 92.
[0081] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.
[0082] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode the video block of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-prediction-encoded video block of the frame, the inter-prediction state for each inter-prediction-encoded video block of the frame, and other information used to decode the video block in the current video frame.
[0083] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine whether the current video block is predicted using intra BC mode, construction information of which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.
[0084] Motion compensation unit 82 may also perform interpolation to calculate interpolated values for sub-integer pixels of a reference block using interpolation filters, such as those used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from received syntax elements and use these interpolation filters to produce the prediction block.
[0085] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters that were calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.
[0086] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (e.g., a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, may also store the decoded video for later presentation on a display device (e.g., Figure 1 on the display device 34).
[0087] In a typical video encoding process, a video sequence typically consists of an ordered set of frames or pictures. Each frame can include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame can be monochrome and therefore only include a two-dimensional array of luma samples.
[0088] like Figure 4AAs shown in , the video encoder 20 (or more specifically, a partitioning unit in the prediction processing unit of the video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size of one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. As Figure 4B As shown in , each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding the samples of the coding tree block. The coding tree block may be an N×N block of samples.
[0089] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding treeblock of the CTU and divide the CTU into smaller CUs. Figures 4B to 4E FIG is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure. Figure 4C As depicted in FIG, a 64×64 CTU 400 is first divided into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. Two 16×16 CUs 430 and CU 440 are each further divided into four CUs with a block size of 8×8. Figure 4D Depicted is a diagram showing Figure 4C The quadtree data structure is the final result of the partitioning process of the CTU 400 depicted in FIG. 4 , with each leaf node of the quadtree corresponding to a CU of a corresponding size ranging from 32×32 to 8×8. Figure 4B Each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements for encoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. It should be noted that Figure 4C and Figure 4D The quadtree partitioning depicted in FIG is for illustrative purposes only, and one CTU can be split into multiple CUs based on quadtree partitioning / ternary tree partitioning / binary tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to the quadtree structure, and each quadtree leaf CU can be further partitioned according to the binary and ternary tree structures. Figure 4E As shown, there are five possible partition types for a coding block with width W and height H, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0090] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.
[0091] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0092] After the video encoder 20 generates the predicted luma block, the predicted Cb block, and the predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, such that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0093] In addition, if Figure 4C As shown in , the video encoder 20 can use quadtree partitioning to decompose the luma residual block, Cb residual block and Cr residual block of a CU into one or more luma transform blocks, Cb transform blocks and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples and syntax elements for transforming the transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.
[0094] Video encoder 20 may apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block may be a two-dimensional array of transform coefficients. A transform coefficient may be a scalar. Video encoder 20 may apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.
[0095] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream comprising a sequence of bits forming a representation of the encoded frame and associated data, which is stored in the storage device 32 or sent to the target device 14.
[0096] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. The video decoder 30 also reconstructs the coding blocks of the current CU by adding samples of the prediction blocks for the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.
[0097] As mentioned above, video coding mainly uses two modes: intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on the reference video block.
[0098] However, with ever-improving video data capture technologies and finer video block sizes for preserving details in video data, the amount of data required to represent the motion vector for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a group of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU (which is also called the "motion vector predictor" (MVP) of the current CU) by exploiting their spatial and temporal correlations.
[0099] Instead of combining as above Figure 2 As described, the actual motion vector of the current CU determined by the motion estimation unit 42 is encoded into the video bitstream, and the motion vector prediction value of the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0100] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also called a "merge list") for the current CU using those potential candidate motion vectors associated with the spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then selecting one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be sent from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list to encode and decode the current CU.
[0101] In general, the basic inter prediction scheme applied in VVC remains almost the same as that of HEVC, except that several prediction tools (eg, extended merge prediction, MMVD, and GPM) are further extended, added, and / or improved.
[0102] Extended Merger Forecast
[0103] As video data acquisition technology continues to improve and the size of video blocks used to retain details of video data becomes more refined, the amount of data required to represent the motion vector of the current picture has also increased significantly. One way to overcome this challenge is to use the motion information (e.g., motion vectors) of the current CU's spatially neighboring CUs, temporally co-located CUs, etc. as an approximation (e.g., prediction) of the current CU's motion information, which is also called the "motion vector predictor (MVP)" of the current CU.
[0104] Just like the process of selecting a prediction block in a reference picture during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules to construct the MVP candidate list of the current CU, and then select one MVP candidate from the MVP candidate list as the MVP of the current CU. By doing so, there is no need to transmit the MVP candidate list itself between the video encoder 20 and the video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same MVP candidate selected from the MVP candidate list to encode and decode the current CU.
[0105] In VVC, the MVP candidate list is constructed by including the following five MVPs in sequence:
[0106] - spatial MVP from spatially neighboring CUs (i.e., spatial candidates);
[0107] - Temporal MVP from the temporally co-located CU (i.e., temporal candidate);
[0108] -- History-based MVP (HMVP) from a first-in-first-out (FIFO) table;
[0109] - Average MVP for each pair; and
[0110] ——Zero MVP.
[0111] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded in merge mode, the index of the best MVP candidate is encoded using truncated unary binarization. The first binary bit of the index is encoded using context, and bypass coding is used for the remaining binary bits of the index.
[0112] The following provides the derivation process of each type of MVP. Like HEVC, VVC also supports parallel derivation of MVP candidate lists for all CUs in a certain size area.
[0113] Derive MVP based on spatial candidates
[0114] In VVC, based on spatial candidates (e.g. Figure 5 The MVP derived from the CU adjacent to the current CU 101 in HEVC is the same as the MVP derived from the spatial candidates in HEVC, except that the positions of the first two spatial candidates are swapped. Figure 5 Up to four spatial candidates are selected from the spatial candidates at the positions shown (i.e., top position B0, left position A0, upper right position B1, lower left position A1, and upper left position B2). The derivation process is performed in the order of CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is only considered if one or more CUs at positions B0, A0, B1, and A1 are not available (for example, because the one or more CUs belong to other slices or tiles) or are intra-coded.
[0115] After adding the CU at position B0 as a candidate to the merge candidate list, a redundancy check is performed on the remaining candidates added to the merge candidate list. This ensures that candidates with the same motion information are excluded from the merge candidate list, thereby improving the encoding and decoding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only Figure 6 Only pairs linked by arrows are considered, and a candidate is added to the merge candidate list only if the motion information of the candidate in the corresponding pair used for redundancy check is different from the motion information of the candidate to be added. The spatial MVP derived from the candidates in the merge candidate list is added to the MVP candidate list.
[0116] Export MVP based on time candidate
[0117] In the process of deriving MVP from temporal candidates, only one temporal candidate is added to the merge candidate list. In particular, when deriving MVP from the temporal candidate, for the current CU (e.g., Figure 7 curr_CU 303 in ) is based on the CU belonging to the same-position picture (e.g., Figure 7 col_pic 302 in ) of the same CU (e.g., Figure 7 The col_CU 301 in the slice header is used as a temporal candidate to derive the scaled motion vector and the scaled motion vector is added to the MVP candidate list as a temporal MVP candidate. The reference picture list and reference picture index used to derive the co-located CU are explicitly signaled in the slice header. Figure 7 As shown, the scaled motion vector is obtained (i.e., scaled) based on the motion vector of the co-located CU using the picture order count (POC) distance (i.e., tb and td), where tb is defined as the current picture (e.g., Figure 7 curr_pic 304) of the reference picture (e.g., Figure 7 305 in ) and the POC difference between the current picture, and td is defined as the reference picture of the co-located picture (e.g., Figure 7 The POC difference between the col_ref 306 in the temporal candidate and the co-located picture. The reference picture index of the temporal candidate is set equal to zero.
[0118] like Figure 8 As shown, the position for the temporal candidate (i.e., the co-located CU) in the current CU 401 is selected from positions C0 and C1. If the CU at position C0 in the co-located picture is not available, is intra-coded, or is located outside the current CTU row, the CU at position C1 is used as the co-located CU for deriving the temporal MVP candidate. Otherwise, the CU at position C0 is used as the co-located CU for deriving the temporal MVP candidate.
[0119] Export HMVP candidates
[0120] After spatial MVP and temporal MVP, the HMVP candidate is added to the MVP candidate list. The motion information of the previously coded block is stored in the HMVP table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the HMVP table.
[0121] The size of the HMVP table is set to 6. When a new HMVP candidate is inserted into the HMVP table, a constrained FIFO rule is used, in which a redundancy check is first applied to find whether the same HMVP exists in the HMVP table. If found, the same HMVP is removed from the HMVP table, and all subsequent HMVP candidates are moved forward, and the same HMVP is added to the last entry of the HMVP table.
[0122] HMVP candidates can be used in the MVP candidate list construction process. The latest HMVP candidates in the HMVP table are checked in turn and inserted into the MVP candidate list after the temporal MVP candidate. Redundancy check is applied to the HMVP candidates relative to the spatial candidates and / or temporal MVP candidates.
[0123] To reduce the number of redundancy check operations, the following simplifications are introduced:
[0124] - performing redundancy checks on the last two entries in the HMVP table against the spatial MVP candidates derived from the spatial candidates at positions A1 and B1, respectively; and
[0125] - Once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1, the process of constructing the MVP candidate list from the HMVP candidates is terminated.
[0126] Export pairwise average MVP candidates
[0127] A pairwise average MVP candidate is generated by averaging the MVPs derived from a predefined pair of the first two merge candidates in the existing merge candidate list. The first merge candidate in the predefined pair may be defined as p0Cand, and the second merge candidate in the predefined pair may be defined as p1Cand. Separately for each reference picture list, an average motion vector is calculated based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available for a reference picture list, even if the two motion vectors point to different reference pictures, the two motion vectors are averaged, and the reference picture of the averaged motion vector is set as the reference picture of p0Cand. If only one motion vector is available for a reference picture list, that motion vector is used directly. If no motion vector is available for a reference picture list, the motion vector and reference picture index for that reference picture list remain invalid.
[0128] Zero MVP
[0129] When the MVP candidate list is not full after adding the pairwise average MVP candidates, zero MVPs are inserted at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0130] MMVD
[0131] As mentioned above, in merge mode, motion information (i.e., MVP candidates) is implicitly derived based on the MVP candidate list constructed for the current CU and directly used as the MV of the current CU to generate the prediction sample of the current CU, which may result in a certain error between the actual MV of the current CU and the implicitly derived MVP. In order to improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, in which the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. After sending the regular merge flag, the MMVD flag is signaled to specify whether the MMVD mode is used for the current CU.
[0132] In the MMVD mode, after an MVP candidate is selected from the top two MVP candidates in the MVP candidate list, MMVD information is signaled, wherein the MMVD information includes an MMVD candidate flag for specifying which of the top two MVP candidates is selected as a MV basis, a distance index for indicating motion magnitude information of the MVD, and a direction index for indicating motion direction information of the MVD.
[0133] The distance index of the motion magnitude information of the specified MVD indicates the distance from the reference picture of the current CU (e.g., Figure 9 The L0 reference picture 501 or L1 reference picture 503 in FIG. 50 is pointed to by the selected MVP candidate (eg, Figure 9 The distance index and the predefined offset are specified in Table 1 below.
[0134]
[0135] Table 1
[0136] The direction index specifies the sign of the MVD, which represents the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined symbols. In some examples, the meaning of the sign of the MVD may vary depending on the information of the selected MVP candidate. When the selected MVP candidate is a unidirectionally predicted MV or a bidirectionally predicted MV (wherein both MVs point to the same side of the current picture (i.e., the POCs of the two reference pictures of the current picture (e.g., the reference pictures of list 0 and list 1, also referred to as L0 reference pictures and L1 reference pictures, respectively) are both greater than the POC of the current picture, or both are less than the POC of the current picture)), the symbol in Table 2 specifies the sign of the MVD added to the selected MVP candidate. When the selected MVP candidate is a bidirectional prediction MV (wherein the two MVs point to different sides of the current picture (i.e., the POC of one reference picture of the current picture is greater than the POC of the current picture, and the POC of the other reference picture of the current picture is smaller than the POC of the current picture), if the POC distance for the L0 reference picture (i.e., the POC distance between the L0 reference picture and the current picture) is greater than the POC distance for the L1 reference picture (i.e., the POC distance between the L1 reference picture and the current picture), the signs in Table 2 specify the sign of the MVD for List 0 (MVD0) added to the MVP for List 0 (MVP0) of the selected MVP candidate, and the sign of the MVD for List 1 (MVD1) added to the MVP for List 1 (MVP1) of the selected MVP candidate is opposite to the sign in Table 2; otherwise, if the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, the signs in Table 2 specify the sign of the MVD1 added to MVP1, and the sign of the MVD0 added to MVP0 is opposite to the sign in Table 2.
[0137]
[0138] Table 2
[0139] MVD is scaled based on the POC distance. If the POC distances for the L0 and L1 reference pictures are the same, no MVD scaling is required. Otherwise, if the POC distance for the L0 reference picture is greater than the POC distance for the L1 reference picture, MVD1 is scaled. If the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, MVD0 is scaled.
[0140] GPM
[0141] In VVC, GPM is supported for inter-frame prediction. A CU-level flag is used to signal GPM as a merge mode. Other merge modes include normal merge mode, MMVD mode, CIIP mode, and sub-block merge mode. For each possible CU size W×H (W=2 m And H=2 n , where m,n∈{3,4,5,6}), GPM supports a total of 64 partitions, and the possible CU size W×H does not include 8×64 and 64×8.
[0142] When using GPM, the CU is split into two parts by a geometrically positioned straight line. The position of the split line is mathematically derived from the angle and offset parameters of the specific split. Each part of the CU obtained by geometric partitioning is inter-predicted using its own motion; and only unidirectional prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion-compensated predictions are required for each CU, just like traditional bidirectional prediction.
[0143] If GPM is used for the current CU, a geometric partitioning index indicating a partitioning mode of the geometric partitioning (indicating the angle and offset of the geometric partitioning) and two merge indexes (one merge index for each partition) are further signaled.
[0144] The unidirectional prediction candidate list is directly derived from the merge candidate list constructed by the extended merge prediction process described above. Let n be the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth merge candidate in the merge candidate list is used as the nth unidirectional prediction motion vector of the GPM. Figure 10 In the merging candidate list, these motion vectors are marked with "x". In the case that the corresponding LX motion vector of the n-th merging candidate does not exist, the L(1-X) motion vector of the same merging candidate is used as the unidirectional prediction motion vector of GPM instead.
[0145] CIIP
[0146] In VVC, when a CU is encoded in merge mode, if the CU contains at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if the width of the CU and the height of the CU are less than 128 luma samples, an additional flag is signaled to indicate whether the CIIP mode is applied to the current CU. In CIIP mode, a prediction signal is obtained by combining an inter-frame prediction signal with an intra-frame prediction signal. The inter-frame prediction signal in CIIP mode is derived using the same inter-frame prediction process as that applied in the conventional merge mode; and the intra-frame prediction signal in CIIP mode is derived according to the conventional intra-frame prediction process using the planar mode. The intra-frame prediction signal and the inter-frame prediction signal are then combined using a weighted average, where the weighted average is calculated based on the top neighboring block and the left neighboring block (such as Figure 11 The weight value is calculated by the encoding mode shown in:
[0147] -- If the top neighboring block is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0;
[0148] -- If the left neighboring block is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0;
[0149] If (isIntraLeft + isIntraTop) is equal to 2, then set the weight value to 3.
[0150] ——Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, set the weight value to 2;
[0151] Otherwise, set the weight value to 1.
[0152] ——The prediction signal P in CIIP mode is derived as follows CIIP :
[0153] P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (1)
[0154] Among them, P inter is the inter-frame prediction signal in CIIP mode, P intra It is the intra prediction signal in CIIP mode, wt is the weight value, and >> represents the right shift operation.
[0155] Intra-block copying in Versatile Video Coding (VVC)
[0156] Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. IBC significantly improves the coding efficiency of screen content material. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture. The luminance block vector of the IBC-encoded CU has integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The IBC-encoded CU is regarded as a third prediction mode in addition to the intra or inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples.
[0157] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed.
[0158] In a hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4x4 sub-block. For larger current block sizes, a hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash key in the corresponding reference position. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference block with the smallest cost is selected.
[0159] In the block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0160] At CU level, IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows:
[0161] IBC skip / merge mode: The merge candidate index is used to indicate which block vectors from a list of neighboring candidate IBC coded blocks are used to predict the current block. The merge list consists of: spatial candidates, HMVP candidates, and pairwise candidates.
[0162] IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (if IBC encoded). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0163] IBC Reference Area
[0164] To reduce memory consumption and decoder complexity, IBC in VVC only allows reconstruction of predefined areas, including the current CTU area and some areas of the left CTU. Figure 12 The reference area of the IBC mode is shown, where each block represents a 64x64 luma sample unit.
[0165] Depending on where the current coded CU position is within the current CTU, the following applies:
[0166] If the current block falls within the upper left 64x64 block of the current CTU, in addition to the reconstructed samples in the current CTU, the CPR mode can also be used to reference the reference samples in the lower right 64x64 block of the left CTU. The current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU.
[0167] If the current block falls into the upper right 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (0 and 64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU.
[0168] If the current block falls into the lower left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU. Otherwise, the current block can also use the CPR mode to refer to the reference samples in the lower right 64x64 block of the left CTU.
[0169] If the current block falls into the lower right 64x64 block of the current CTU, it can only use the CPR mode to refer to the reconstructed samples in the current CTU.
[0170] For hardware implementations, this restriction allows the use of local on-chip memory to implement the IBC mode.
[0171] Interaction of IBC with other coding tools
[0172] The interaction between IBC mode and other inter-coding tools in VVC, such as pairwise merge candidates, history-based motion vector predictor (HMVP), combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM), is as follows:
[0173] IBC can be used with pairwise merge candidates and HMVP. A new pairwise IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference.
[0174] IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM.
[0175] When DUAL_TREE partitioning is used, IBC is not allowed for chroma-coded blocks.
[0176] Unlike in the HEVC screen content coding extension, the current picture is no longer included in reference picture list 0 as one of the reference pictures for IBC prediction. The derivation process of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects apply:
[0177] IBC shares the same process as in regular MV merging, including having paired merge candidates and history-based motion prediction values, but TMVP and zero vectors are not allowed since they are invalid for IBC mode.
[0178] Separate HMVP buffers (5 candidates each) are used for regular MV and IBC.
[0179] The block vector constraint is implemented as a bitstream consistency constraint. The encoder needs to ensure that there are no invalid vectors in the bitstream, and if the merge candidate is invalid (out of range or 0), then the merge should not be used. This bitstream consistency constraint is expressed in terms of a virtual buffer as described below.
[0180] For deblocking, IBC is handled as inter mode.
[0181] If the current block is encoded using IBC prediction mode, AMVR does not use quarter pels; instead, AMVR is signaled to only indicate whether the MV is inter-pel or 4-integer-pel.
[0182] The number of IBC merge candidates may be signaled in the slice header separately from the number of regular, sub-block, and geometric merge candidates.
[0183] The virtual buffer concept is used to describe the allowable reference area and valid block vectors for IBC prediction mode. Denoting the CTU size as ctbSize, the virtual buffer ibcBuf has a width of wIbcBuf = 128x128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32.
[0184] The size of VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64).
[0185] The virtual IBC buffer ibcBuf is maintained as follows.
[0186] At the beginning of decoding each CTU line, flush the entire ibcBuf with an invalid value of -1.
[0187] At the start of decoding the VPDU (xVPDU, yVPDU) relative to the upper left corner of the picture, set ibcBuf[x][y]=-1, where x=xVPDU%wIbcBuf, ..., xVPDU%wIbcBuf+W v -1;y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1.
[0188] After decoding, the CU contains the (x, y) relative to the upper left corner of the picture, set
[0189] ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]
[0190] For a block covering coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid:
[0191] ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1.
[0192] Intra-block copying in the Enhanced Compression Model (ECM)
[0193] In ECM, IBC is improved in the following aspects.
[0194] IBC merging / AMVP list construction
[0195] The IBC merge / AMVP list construction is modified as follows:
[0196] An IBC merge / AMVP candidate can be inserted into the IBC merge / AMVP candidate list only if it is valid.
[0197] The top-right, bottom-left, and top-left spatial candidates and one pairwise average candidate may be added to the IBC merge / AMVP candidate list.
[0198] Adaptive reordering based on template (ARMC-TM) is applied to the IBC merge list.
[0199] The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC merge candidates using full pruning, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0200] The zero vector candidates that populate the IBC merge / AMVP list are replaced with the BVP candidate set in the IBC reference region. The zero vector is invalid as a block vector in IBC merge mode, so it is discarded as a BVP in the IBC candidate list.
[0201] Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Figure 13 shown.
[0202] IBC with template matching
[0203] Template matching is used for IBC for both IBC merge mode and IBC AMVP mode.
[0204] Compared to the IBC-TM merge list used by the conventional IBC merge mode, the IBC-TM merge list is modified so that candidates are selected according to a pruning method with motion distances between candidates, as in the conventional TM merge mode. The ending zero motion implementation is replaced by motion vectors pointing to the left (-W, 0), up (0, -H), and left-up (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0205] In IBC-TM merge mode, template matching methods are utilized to refine the selected candidates before the RDO or decoding process. The IBC-TM merge mode has been made competitive with the regular IBC merge mode, and the TM merge flag is signaled.
[0206] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of these 3 selected candidates is refined using the template matching method and ranked according to their resulting template matching cost. Then, only the first 2 are considered in the motion estimation process as usual.
[0207] Template matching refinement for both IBC-TM merging and AMVP modes is quite simple, since IBC motion vectors are constrained to be (i) integers and (ii) within the reference region, e.g. Figure 12 As shown. Therefore, in IBC-TM merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision, depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the template used in each refinement step must respect the constraints of the reference region.
[0208] IBC Reference Area
[0209] The reference area of the IBC extends to the upper two CTU rows. Figure 14 The reference region for encoding CTU(m,n) is shown. Specifically, for the CTU(m,n) to be encoded, the reference region contains CTUs with indices (m-2,n-2)…(W,n-2), (0,n-1)…(W,n-1), (0,n)…(m,n), where W represents the maximum horizontal index within the current tile, slice, or picture. This setting ensures that for a CTU size of 128, IBC does not require additional memory in current ETM platforms. The per-sample block vector search (or local search) range is limited to [-(C<<1), C>>2] horizontally and [-C, C>>2] vertically to accommodate the reference region expansion, where C represents the CTU size.
[0210] IBC merge mode with block vector difference
[0211] In ECM, an IBC merge mode with block vector differences is used. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.
[0212] A base candidate is selected from the first five candidates in the reordered IBC merge list. All possible MBVD refinement positions (20x4) for each base candidate are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. Finally, the top eight refinement positions with the lowest template SAD cost are kept as available positions and are therefore used for MBVD index encoding.
[0213] IBC adaptation for camera-captured content
[0214] When adapting IBC for camera-captured content, the IBC reference range is reduced from 2 CTU rows to 2x128 rows, e.g. Figure 15 As shown. On the encoder side, to reduce complexity, the local search range is set to [-8, 8] horizontally and [-8, 8] vertically, centered on the first block vector prediction value of the current CU. This encoder modification does not apply to SCC sequences.
[0215] Combination of CIIP with TIMD and TM merger
[0216] In CIIP mode, prediction samples are generated by weighting the inter prediction signal predicted using CIIP-TM merged candidate predictions and the intra prediction signal predicted using TIMD derived intra prediction modes. This method is only applied to coding blocks with an area less than or equal to 1024.
[0217] The TIMD derivation method is used to derive intra prediction modes in CIIP. Specifically, the intra prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of the 67 conventional intra prediction modes.
[0218] In addition, it is proposed to modify the weights of the two tests (wIntra, wInter) if the derived intra prediction mode is an angular mode. For near-horizontal mode (2 <= angular mode index < 34), as Figure 16A As shown, the current block is divided vertically; for the near vertical mode (34 <= angle mode index <= 66), as Figure 16B The shown horizontal division of the current block.
[0219] The (wIntra, wInter) of different sub-blocks are shown in Table 3.
[0220] Sub-block index (wIntra, wInter) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6)
[0221] Table 3 Modified weights for angle mode
[0222] Using CIIP-TM, a list of CIIP-TM merge candidates is constructed for the CIIP-TM mode. Merge candidates are refined through template matching. CIIP-TM merge candidates are also reordered using the ARMC method, just like conventional merge candidates. The maximum number of CIIP-TM merge candidates is 2.
[0223] Multiple Hypothesis Prediction (MHP)
[0224] In the multi-hypothesis inter prediction mode, in addition to the conventional bi-directional prediction signal, one or more additional motion compensated prediction signals are signaled. The resulting total prediction signal is obtained by sample-by-sample weighted superposition. bi and the first additional inter-frame prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows:
[0225] p3=(1-α)p bi +αh3 (2)
[0226] According to the mapping presented in Table 4, the weighting factor α is specified by a new syntax element add_hyp_weight_idx:
[0227] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0228] Table 4 Mapping between add_hyp_weight_idx and α
[0229] Similar to the above, more than one additional prediction signal may be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0230] p n+1 =(1-α n+1 )p n +α n+1 h n+1 (3)
[0231] The resulting overall prediction signal is obtained as the last p n (i.e., the p with the largest index n n ). In this mode, up to two additional prediction signals can be used (ie, n is limited to 2).
[0232] The motion parameters for each additional prediction hypothesis can be signaled explicitly by specifying the reference index, motion vector predictor index and motion vector difference, or implicitly by specifying the merge index. A separate multi-hypothesis merge flag distinguishes these two signaling modes.
[0233] For inter-AMVP mode, MHP is applied only if unequal weights in BCW are selected in bi-prediction mode.
[0234] A combination of MHP and BDOF is possible, however BDOF is only applied to the bidirectional prediction signal part of the prediction signal (ie the ordinary first two hypotheses).
[0235] Geometric Partitioning Mode (GPM) in ECM
[0236] GPM with Merged Motion Vector Difference (MMVD)
[0237] GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. First, a flag is signaled for the GPM CU to specify whether to use this mode. If the mode is used, each geometric partition of the GPM CU can further decide whether to signal MVD. If MVD is signaled for a geometric partition, then after selecting a GPM merge candidate, the partition's motion is further refined using the signaled MVD information. All other procedures remain the same as in GPM.
[0238] MVD is signaled as a pair of distance and direction, similar to MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels), and eight candidate directions (four horizontal / vertical directions and four diagonal directions) involved in GPM with MMVD (GPM-MMVD). In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, MVD is left-shifted by 2 as in MMVD.
[0239] GPM with Template Matching(TM)
[0240] Template matching is applied to GPM. When GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether template matching is applied to both geometric partitions. Template matching is used to refine the motion information for each geometric partition. When template matching is selected, a template is constructed using the left, top, or left and top neighboring samples depending on the partition angle, as shown in Table 5. Motion is then refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern with the merge mode of the half-pixel interpolation filter disabled.
[0241] Split Angle 0 2 3 4 5 8 11 12 13 14 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A Split Angle 16 18 19 20 21 24 27 28 29 30 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A
[0242] Table 5
[0243] Templates for the first and second geometric partitions, where A indicates using the top sample, L indicates using the left sample, and L+A indicates using both the left and top samples.
[0244] The GPM candidate list is constructed as follows:
[0245] 1. Interleaved list 0 MV candidates and list 1 MV candidates are directly derived from the regular merge candidate list, where list 0 MV candidates have higher priority than list 1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0246] 2. Interleaved list 1 MV candidates and list 0 MV candidates are further derived directly from the regular merge candidate list, where list 1 MV candidates have higher priority than list 0 MV candidates. The same pruning method with adaptive threshold is also applied to remove redundant MV candidates.
[0247] 3. Fill in zero MV candidates until the GPM candidate list is full.
[0248] GPM-MMVD and GPM-TM are dedicated to a GPM CU. This is accomplished by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.
[0249] GPM with inter-frame and intra-frame prediction
[0250] In GPM with inter and intra prediction, the final prediction samples are generated by weighting the inter prediction samples and intra prediction samples for each GPM separation area. The inter prediction samples are derived from the inter GPM, while the intra prediction samples are derived from the intra prediction mode (IPM) candidate list and the index signaled from the encoder. The IPM candidate list size is predefined as 3. The available IPM candidates are parallel angle mode (parallel mode) for GPM block boundaries, vertical angle mode (vertical mode) for GPM block boundaries, and planar mode, as shown in Figure 2. 17A to 17D In addition, Figure 17D The GPM with intra and intra prediction shown in
[15] is restricted to reduce the signaling overhead for IPM and avoid increasing the size of the intra prediction circuit on the hardware decoder. In addition, direct motion vector and IPM storage on the GPU hybrid region is introduced to further improve the encoding and decoding performance.
[0251] In IPM derivation based on DIMD and neighboring modes, parallel modes are registered first. Therefore, if the same IPM candidate does not exist in the list, a maximum of two IPM candidates derived from the decoder-side intra mode derivation (DIMD) method and / or neighboring block derivation can be registered. For neighboring mode derivation, there are up to five available neighboring block positions, but they are limited by the angle of the GPM block boundary as shown in Table 6, which has been used for GPM with template matching (GPM-TM).
[0252] GPM angle 0 2 3 4 5 8 11 12 13 14 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A GPM angle 16 18 19 20 21 24 27 28 29 30 First Division A A A A L+A L+A L+A L+A A A Second Division L+A L+A L+A L L L L L+A L+A L+A
[0253] Table 6
[0254] The positions of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary. A and L represent the above and left sides of the prediction block.
[0255] GPM intra can be combined with GPM with merged motion vector difference (GPM-MMVD). TIMD is used for IPM candidates within GPM to further improve codec performance. Parallel mode can be registered first, followed by TIMD, DIMD, and IPM candidates for neighboring blocks.
[0256] Template matching based reordering for GPM segmentation mode
[0257] In the template matching-based reordering for GPM partitioning modes, given the motion information of the current GPM block, the corresponding TM cost value of the GPM partitioning mode is calculated. Then, all GPM partitioning modes are reordered in ascending order based on the TM cost value. Instead of sending the GPM partitioning mode, an index is signaled using a Golomb-Rice code to indicate where the exact GPM partitioning mode is located in the reordering list.
[0258] The reordering method for the GPM partitioning mode is a two-step process performed after generating the corresponding reference templates for the two GPM partitions in the coding unit, as follows:
[0259] Extend the GPM partition edge into the reference templates of the two GPM partitions, generate 64 reference templates and calculate the corresponding TM cost of each of the 64 reference templates;
[0260] The GPM partitioning patterns are re-sorted in ascending order based on their TM cost values, and the best 32 are marked as usable partitioning patterns.
[0261] The edges on the template extend from the edges of the current CU, such as Figure 18 As shown, but the GPM blending process is not used in the template area across the edge.
[0262] After reordering using TM cost in ascending order, the index is signaled.
[0263] Intra-frame template matching
[0264] Intra Template Matching (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, using an L-shaped template that matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0265] By combining the L-shaped causal neighbors of the current block with Figure 19 Generate a prediction signal by matching another block in a predefined search area in the image, including:
[0266] R1: Current CTU
[0267] R2: Upper left CTU
[0268] R3: Upper CTU
[0269] R4: Left CTU
[0270] The sum of absolute differences (SAD) is used as the cost function.
[0271] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0272] The sizes of all regions (SearchRange_w, SearchRange_h) are set proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:
[0273] SearchRange_w=a*BlkW
[0274] SearchRange_h=a*BlkH
[0275] Where 'a' is a constant that controls the gain / complexity tradeoff. In practice, 'a' is equal to 5.
[0276] The intra template matching tool is enabled for CUs with width and height less than or equal to 64. This maximum CU size for intra template matching is configurable.
[0277] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag.
[0278] Fusion for Template-based Intra Mode Derivation (TIMD)
[0279] For each intra prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The first two intra prediction modes with the smallest SATD are selected as TIMD modes. These two TIMD modes are fused with weights after applying the PDPC process, and this weighted intra prediction is used to encode the current CU. Position-dependent intra prediction combining (PDPC) is included in the derivation of TIMD modes.
[0280] The costs of the two selected patterns are compared to a threshold, and in the test a cost factor of 2 is applied as follows:
[0281] costMode2<2*costMode1.
[0282] If this condition is true, then fusion is applied, otherwise only mode 1 is used.
[0283] The weight of a mode is calculated from its SATD cost as follows:
[0284] weight1=costMode2 / (costMode1+costMode2)
[0285] weight2=1-weight1
[0286] The division operation is performed using the same lookup table (LUT) based integration scheme used by CCLM.
[0287] Local Illumination Compensation (LIC)
[0288] LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its prediction block based on the local illumination variation between the current block template and the reference block template. The parameters of the function can be expressed as a scale α and an offset β, which form a linear equation: α*p[x]+β to compensate for illumination variation, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When surround motion compensation is enabled, the MV is cropped to account for the surround offset. Because α and β can be derived based on the current block template and the reference block template, they require no signaling overhead, other than signaling the LIC flag for AMVP mode to indicate its use.
[0289] The local illumination compensation proposed in JVET-O0066 is used for unidirectionally predicted inter CUs with the following modifications:
[0290] Neighbor samples within the frame can be used to derive LIC parameters;
[0291] Disable LIC for blocks with less than 32 luma samples;
[0292] For both non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16x16 unit;
[0293] The samples of the reference block template are generated by using MC with the block MV without rounding them to integer pixel precision.
[0294] OBMC
[0295] When OBMC is applied, the top and left boundary pixels of the CU are refined using neighboring block motion information with weighted prediction, as described in JVET-L0101.
[0296] The conditions under which OBMC should not be applied are as follows:
[0297] When OBMC is disabled at the SPS level;
[0298] When the current block has intra mode or IBC mode;
[0299] When the current block applies LIC;
[0300] When the current luminance block area is less than or equal to 32.
[0301] Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using the motion information of neighboring sub-blocks. For sub-block based coding tools, enable:
[0302] Affine AMVP mode;
[0303] Affine merge mode and sub-block based temporal motion vector prediction (SbTMVP);
[0304] Sub-block based bilateral matching.
[0305] When using OBMC mode in CIIP mode with LMCS, inter-mixing is performed before LMCS mapping of inter samples. LMCS is applied to the mixed inter samples, which are combined with intra samples to which LMCS was applied in CIIP mode.
[0306]
[0307] Among them, Inter predY Intra represents the sample points predicted by the motion of the current block in the original domain. predY Represents the sample points predicted in the mapping domain, OBMC predY represents the sample points predicted by the motion of the neighboring blocks in the original domain, and w0 and w1 are weights.
[0308] OBMC based on template matching
[0309] In the template matching-based OBMC scheme, instead of using weighted prediction directly, the prediction value of the CU boundary sample derivation method is determined according to the template matching cost, including using only the motion information of the current block, or using the motion information of the neighboring blocks and one of the hybrid modes.
[0310] In this scheme, for each 4x4 block at the top CU boundary, the template size is equal to 4x1. If N neighboring blocks have the same motion information, the template size is enlarged to 4N×1 because the MC operation can be processed at one time. For each left block of 4x4 at the left CU boundary, the left template size is equal to 1x4 or 1x4N ( Figure 20 ).
[0311] For each 4x4 top block (or N 4x4 block groups), the following steps are followed to derive the prediction values for the boundary samples.
[0312] Take block A as the current block and its above neighbor block AboveNeighbor_A as an example. The operation on the left block is performed in the same way.
[0313] First, three template matching costs (Cost1, Cost2, Cost3) are measured by the SAD between the reconstructed samples of the template derived by the MC process and their corresponding reference samples according to the following three types of motion information:
[0314] Calculate Cost1 based on A's motion information.
[0315] Cost2 is calculated based on the motion information of AboveNeighbor_A.
[0316] Cost3 is calculated based on the weighted prediction of the motion information of A and AboveNeighbor_A, where the weighting factors are 3 / 4 and 1 / 4 respectively.
[0317] Secondly, a method is selected to calculate the final prediction results of the boundary samples by comparing Cost1, Cost2 and Cost3.
[0318] The original MC result using the motion information of the current block is denoted as Pixel 1, and the MC result using the motion information of the neighboring block is denoted as Pixel 2. The final prediction result is denoted as NewPixel.
[0319] If Cost1 is the smallest, then NewPixel(i,j)=Pixel1(i,j).
[0320] If (Cost2+(Cost2>>2)+(Cost2>>3))<=Cost1, then blending mode 1 is used.
[0321] For luma blocks, the number of mixed pixel rows is 4.
[0322] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5
[0323] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3
[0324] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4
[0325] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5
[0326] For chroma blocks, the number of mixed pixel rows is 1.
[0327] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5
[0328] If Cost1<=Cost2, then blending mode 2 is used.
[0329] For luma blocks, the number of mixed pixel rows is 2.
[0330] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4
[0331] NewPixel(i,1)=(31×Pixel1(i,1)+Pixel2(i,1)+16)>>5
[0332] For chroma blocks, the number of mixed pixel rows / columns is 1.
[0333] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4
[0334] Otherwise, use blend mode 3.
[0335] For luma blocks, the number of mixed pixel rows is 4.
[0336] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3
[0337] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4
[0338] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5
[0339] For chroma blocks, the number of mixed pixel rows is 1.
[0340] NewPixel(i,0)=(7×Pixel1(i,0)+Pixel2(i,0)+4)>>3
[0341] Currently, IBC tools are not combined with GPM tools. Therefore, this disclosure provides an example of combining them, which can improve prediction accuracy and improve encoding and decoding performance.
[0342] Currently, coding blocks encoded in IBC mode are not combined with coding blocks encoded in intra or inter modes. Therefore, this disclosure provides an example of combining them, which can improve prediction accuracy and improve encoding and decoding performance.
[0343] Currently, the number of block vectors (BVs) in IBC tools is singular. Therefore, the present disclosure provides examples to increase the number of block vectors (BVs) and combine prediction results, which can improve prediction accuracy and improve encoding and decoding performance.
[0344] Currently, coding blocks encoded using intra-frame TMP mode are not combined with coding blocks encoded using intra-frame mode or inter-frame mode. Therefore, the present disclosure provides an example of combining them together, which can improve prediction accuracy and improve encoding and decoding performance.
[0345] Currently, the intra-frame TMP tool is not combined with the GPM tool. Therefore, this disclosure provides an example of combining them together, which can improve prediction accuracy and improve encoding and decoding performance.
[0346] Currently, IBC tools are not combined with TIMD tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and improve encoding and decoding performance.
[0347] Currently, the intra-frame TMP tool is not combined with the TIMD tool. Therefore, this disclosure provides an example of combining them together, which can improve prediction accuracy and improve encoding and decoding performance.
[0348] Currently, the intra-frame TMP tool is not combined with the LIC tool. Therefore, this disclosure provides an example of combining them together, which can improve prediction accuracy and improve encoding and decoding performance.
[0349] Currently, IBC tools are not combined with OBMC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and improve encoding and decoding performance.
[0350] Currently, the intra-frame TMP tool is not combined with the OBMC tool. Therefore, this disclosure provides an example of combining them together, which can improve prediction accuracy and improve encoding and decoding performance.
[0351] In the present disclosure, in order to solve the above problems, a method is provided to further improve the existing design of IBC. Generally, the main features of the technology proposed in this disclosure are summarized as follows.
[0352] The IBC tool is combined with the GPM tool, and the combination can be GPM with IBC and IBC prediction, GPM with IBC and intra-frame prediction, or GPM with IBC and inter-frame prediction.
[0353] The IBC tool, which is a simplified version combined with the GPM tool, predicts the upper left part using the intra mode and the lower right part using the IBC mode for a predefined direction (such as 45 degrees), and then weights them averagely to obtain the final prediction signal.
[0354] The IBC tool is combined with the CIIP tool, where IBC prediction is combined with intra prediction mode, or IBC prediction is combined with inter prediction mode.
[0355] The IBC tool is combined with the MHP tool, where more than one BV prediction is obtained and they are weighted averaged to obtain the final prediction signal.
[0356] The intra TMP tool is combined with the CIIP tool, where intra TMP is combined with intra prediction mode, or intra TMP is combined with inter prediction mode.
[0357] The intra TMP tool is combined with the GPM tool, and the combination can be GPM with intra TMP and intra TMP prediction, GPM with intra TMP and intra prediction, or GPM with intra TMP and inter prediction.
[0358] As a simplified version of the Intra TMP tool combined with the GPM tool, for a predefined direction (such as 45 degrees), the upper left part is predicted using the Intra mode, the lower right part is predicted using the Intra TMP mode, and then they are weighted averagely to obtain the final prediction signal.
[0359] The IBC tool is combined with the TIMD tool, where the IBC mode is used together with the intra prediction mode in the MPM for TIMD fusion.
[0360] The intra TMP tool is combined with the TIMD tool, where the intra TMP mode is used together with the intra prediction mode in MPM for TIMD fusion.
[0361] The intra TMP tool is combined with the LIC tool, where the LIC tool is used to compensate for local illumination variations between the current block and its intra TMP prediction block.
[0362] The IBC tool is combined with the OBMC tool, where the OBMC tool is used to refine the top and left boundary pixels of the current block predicted by IBC.
[0363] The intra TMP tool is combined with the OBMC tool, where the OBMC tool is used to refine the top and left boundary pixels of the current block predicted by the intra TMP.
[0364] In some examples, the disclosed methods can be applied independently or in combination.
[0365] GPM with IBC and IBC forecast
[0366] According to one or more embodiments of the present disclosure, an IBC tool is combined with a GPM tool in the form of a GPM with IBC and IBC predictions.Different approaches can be used to achieve this goal.
[0367] In the first approach, the two "inter" parts of the GPM with inter and inter prediction in VVC are replaced with IBC. This means that the two IBC merge prediction results are weighted averaged with respect to the dividing line in the coding block. The weights can be obtained by referring to the GPM with inter and inter prediction in VVC.
[0368] In the second approach, the two "inter" parts of the GPM with inter and inter prediction methods in the ECM are replaced by IBC, where some template matching tools can be utilized to further improve the encoding and decoding performance.
[0369] GPM with IBC and intra prediction
[0370] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC and intra prediction.Different approaches can be used to achieve this goal.
[0371] In the first method, the "inter" part of the GPM with inter and intra prediction methods in the ECM is replaced by IBC, where the IBC merged prediction result is weighted averaged with the intra prediction result to obtain the final prediction signal.
[0372] GPM with IBC and inter-frame prediction
[0373] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC and inter prediction.Different approaches can be used to achieve this goal.
[0374] In the first method, an "inter" part of the GPM with inter and inter prediction methods in VVC is replaced by IBC, wherein the IBC merged prediction result is weighted averaged with the inter merged prediction result to obtain the final prediction signal.
[0375] In the second method, an "inter" part of the GPM with inter and inter prediction methods in the ECM is replaced by IBC, where some template matching tools can be used to further improve the encoding and decoding performance.
[0376] Simplified IBC and intra prediction combination in GPM form
[0377] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a simplified GPM with IBC and intra prediction, such as combining IBC and intra prediction in a certain partitioning mode, which can save the bit overhead of the partitioning representation. Different methods can be used to achieve this goal.
[0378] In the first method, for a divided line, such as 45 degrees, the upper left part of the coding block is encoded using the intra-frame prediction mode, and the lower right part of the coding block is encoded using the IBC prediction mode, and then they are averaged in the form of GPM to obtain the final prediction signal.
[0379] Combined IBC - Intra / Inter Prediction
[0380] According to one or more embodiments of the present disclosure, a coding block encoded using the IBC mode is combined with a coding block encoded using the intra mode or the inter mode. Different methods can be used to achieve this goal.
[0381] In the first method, the decoder / encoder can combine the coding blocks encoded using the IBC mode with the coding blocks encoded using the intra mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, the coding blocks encoded in the IBC merge mode are regarded as coding blocks encoded in the inter-frame merge mode, and are combined with the coding blocks encoded in the planar intra prediction mode. In another example, similar to the combination of CIIP with TIMD and TM merge technology in ECM, the coding blocks encoded in the IBC merge-TM mode are combined with the coding blocks encoded in the intra prediction mode derived from TIMD.
[0382] When a coding block encoded in IBC mode is combined with a coding block encoded in intra mode, the weights can be designed similarly to the CIIP technology in VVC and the combination of CIIP with TIMD and TM merging technology in ECM, namely: 1) the weights of both the IBC coding block and the intra coding block are greater than zero and less than one, or the weight of the intra coding block gradually changes from one to zero from one area to another in the current block (and vice versa for the weight of the IBC coding block); 2) the weights of the IBC coding block and the intra coding block can be determined based on the coding mode of the neighboring block and the intra mode of the current block; 3) the weights of the IBC coding block and the intra coding block can be uniform in the entire current block or different in different positions of the current block.
[0383] For example, the weights of IBC-coded blocks and intra-coded blocks can be determined as follows: when the upper and left neighboring blocks of the current block are both intra-coded and the intra-frame mode of the current block is planar mode, the weights of the IBC-coded blocks and intra-frame coded blocks in the entire current block are 1 / 4 and 3 / 4. When the upper and left neighboring blocks of the current block are both IBC-coded and the intra-frame mode of the current block is planar mode, the weights of the IBC-coded blocks and intra-frame coded blocks in the entire current block are 3 / 4 and 1 / 4. When one upper or left neighboring block is IBC-coded and the other neighboring block is intra-coded, and the intra-frame mode of the current block is planar mode, the weights of the IBC-coded blocks and intra-frame coded blocks in the entire current block are 1 / 2 and 1 / 2.
[0384] When the intra mode of the current block is close to the horizontal angle mode (2<=angle mode index<34), as shown in Figure 16A As shown, the current block is divided vertically; when the intra mode of the current block is close to the vertical angle mode (34<=angle mode index<=66), as shown Figure 16B As shown, the current block is divided horizontally. Table 7 shows the weights of the IBC coding block (wIBC) and intra coding block (wIntra) for different sub-blocks. In addition, the weights for IBC coding blocks and intra coding blocks can be determined in the CIIP-PDPC version. In this version, the intra mode of the current block is set to planar mode, and when the combination position moves from the upper left to the lower right in the current block, the weight for the intra coding block gradually decreases, and vice versa for the weight for the IBC coding block.
[0385] Sub-block index (wIntra,wIBC) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6)
[0386] Table 7 Modified weights for angle mode.
[0387] When combining a coded block coded in IBC mode with a coded block coded in intra mode, the weights can also be designed in the mask version, i.e., the weights of the IBC coded block and the intra coded block can be one or zero for different areas of the current block. The specific weights of the IBC coded block and the intra coded block can be determined based on the coding mode of the neighboring blocks and the intra mode of the current block. For example, the weights for the IBC coded block and the intra coded block can be determined as follows: when the intra mode of the current block is close to the horizontal angle mode (2 <= angle mode index < 34), if the upper and left neighboring blocks of the current block are both intra coded, then Figure 21 As shown in (a), the weight for the intra-coded block is one in the left 3 / 4 area of the current block and zero in the right 1 / 4 area of the current block, and vice versa for the IBC-coded block; if only one neighboring block is intra-coded, then Figure 21 As shown in (b), the weight for the intra-coded block is one in the left 1 / 2 area of the current block and zero in the right 1 / 2 area of the current block, and vice versa for the weight of the IBC-coded block; if neither the upper nor the left neighboring blocks of the current block are intra-coded, then Figure 21 As shown in (c), the weight for the intra-coded block is one in the left 1 / 4 region of the current block and zero in the right 3 / 4 region of the current block, and vice versa for the IBC-coded block.
[0388] When the intra mode of the current block is close to the vertical angle mode (34 <= angle mode index <= 66), if both the upper and left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the top 3 / 4 area of the current block and zero in the bottom 1 / 4 area of the current block, as shown in FIG. Figure 21 (d), and vice versa for the weights of IBC coded blocks; if only one neighboring block is intra coded, the weight of the intra coded block is one in the top 1 / 2 region of the current block and zero in the bottom 1 / 2 region of the current block, as shown in Figure 21 (e), and vice versa for the weights of IBC coded blocks; if neither the upper nor the left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the top 1 / 4 region of the current block and zero in the bottom 3 / 4 region of the current block, as shown in Figure 21 (f) and vice versa for the weights of IBC coded blocks.
[0389] When the intra mode of the current block is planar mode, if the upper and left neighboring blocks of the current block are both intra-coded, the weight for the intra-coded block is one in the upper left 3 / 4 area of the current block (horizontal index is less than 1 / 2 width of the current block or vertical index is less than 1 / 2 height of the current block) and is zero in the lower right 1 / 4 area of the current block (horizontal index is equal to or greater than 1 / 2 width of the current block and vertical index is equal to or greater than 1 / 2 height of the current block), as shown in FIG. Figure 21 (g) and vice versa for the weights of IBC coded blocks; if only the upper neighboring block is intra coded, the weight of the intra coded block is one in the top 1 / 2 region of the current block and zero in the bottom 1 / 2 region of the current block, as shown in Figure 21 (e) and vice versa for the weights of IBC coded blocks; if only the left neighboring block is intra coded, the weight of the intra coded block is one in the left 1 / 2 region of the current block and zero in the right 1 / 2 region of the current block, as shown in Figure 21 As shown in (b), and vice versa for the weights of IBC coded blocks; if neither the upper nor the left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the upper left 1 / 4 area of the current block (the horizontal index is less than 1 / 2 of the width of the current block, and the vertical index is less than 1 / 2 of the height of the current block), and is zero in the lower right 3 / 4 area of the current block (the horizontal index is equal to or greater than 1 / 2 of the width of the current block, or the vertical index is equal to or greater than 1 / 2 of the height of the current block), as shown in Figure 21 (h) and vice versa for the weights of IBC coded blocks.
[0390] The above two weight design methods can be used independently or in combination. For example, when the intra mode of the current block is planar mode, if the upper and left neighboring blocks of the current block are not intra-coded, the weights for IBC coded blocks and intra-coded blocks can be designed similar to the CIIP technology in VVC. Under other conditions, the weights for IBC coded blocks and intra-coded blocks can be designed in the mask version.
[0391] In the second method, the decoder / encoder may combine a coding block encoded in IBC mode with a coding block encoded in inter mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, a coding block encoded in IBC merge mode is treated as a coding block encoded in planar intra mode and combined with a coding block encoded in inter merge mode. In another example, a coding block encoded in IBC merge mode is treated as a coding block encoded in inter merge mode and combined with a coding block encoded in inter merge mode by equal averaging.
[0392] In the third method, the decoder / encoder can combine the coding blocks encoded in IBC mode with the coding blocks encoded in intra mode and the coding blocks encoded in inter mode. Various methods can be used in this combination. In one example, the coding blocks encoded in IBC mode, the coding blocks encoded in intra mode, and the coding blocks encoded in inter mode are directly combined by equal averaging. In another example, the coding blocks encoded in IBC mode are first combined separately with the coding blocks encoded in intra mode and inter mode, as shown in the first and second methods. Then, the individual combination results are combined by equal averaging.
[0393] Multi-hypothesis IBC prediction
[0394] According to one or more embodiments of the present disclosure, the number of block vectors (BVs) in the IBC tool is increased to 2 or more, and 2 or more hypotheses are combined to obtain the final prediction result. Different methods can be used to achieve this goal.
[0395] In the first approach, the decoder / encoder can combine two hypotheses corresponding to two BVs to obtain the final prediction result. Various methods can be used to achieve this goal. In one example, the two BVs corresponding to the minimum and second minimum rate-distortion metrics in the IBC AMVP mode are equally averaged to obtain the final prediction result. In another example, the prediction result corresponding to the IBC AMVP mode and the prediction result corresponding to the IBC merge mode are equally averaged to obtain the final prediction result.
[0396] In the second method, the decoder / encoder can combine more hypotheses corresponding to more BVs to obtain the final prediction result. Various methods can be used to achieve this goal. In one example, the iterative accumulation method proposed in the multi-hypothesis prediction (MHP) technology is used to obtain the final prediction result. In another example, all BVs corresponding to the minimum rate distortion metric, the second minimum rate distortion metric, the third minimum rate distortion metric, etc. in the IBC AMVP mode are equally averaged to obtain the final prediction result.
[0397] Prediction block candidate derivation
[0398] In some embodiments, prediction block candidates are searched and selected based on the criterion of minimizing template matching cost, i.e., the top N candidates that result in the smallest block vector (BV) matching cost are selected. The BV matching cost may not be limited to SAD (sum of absolute differences) and SSE (sum of squared errors). In some embodiments, the template matching cost or BV matching cost may be calculated based on a comparison of the distortion of a corresponding reference with a neighboring block (e.g., a top neighboring block or a left neighboring block).
[0399] In some embodiments, prediction block candidates may be selected according to a predefined mode (ie, planar mode).
[0400] In some embodiments, prediction block candidates may be selected based on neighboring predefined patterns (i.e., top predefined pattern, left predefined pattern). In some examples, the top predefined pattern of the current block inherits the predefined pattern of the top neighbor block of the current block, and the left predefined pattern inherits the predefined pattern of the left neighbor block of the current block.
[0401] Fixed multi-hypothesis IBC
[0402] In this embodiment, the weighting factors used to generate the final prediction block are predefined and fixed at both the encoder side and the decoder side.As an example, equal weighting factors, ie 1 / N, may be used for all candidate blocks.
[0403] Adaptive Multi-Hypothesis IBC
[0404] In order to adapt to the diverse characteristics of video content, an adaptive multi-hypothesis IBC method is also proposed.
[0405] In some embodiments, the weighting factors can be derived based on the BV matching costs. Denote the BV matching costs of N candidates as C1, C2, ..., C N , the weighting factors are calculated as follows.
[0406]
[0407] It should be noted that the BV matching cost can be measured by (but not limited to) SAD and SSE.
[0408] In yet another embodiment, the weighting factors may be derived / switched based on the block size or syntax elements signaled in the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / sub-block / sample level.
[0409] In yet another embodiment, the weighting factors can be derived at the encoder side and then signaled to the decoder in the bitstream. Denote the N prediction block candidates as P1, P2, ..., P N , and denote the current block as X, the weighting factor can be solved by the following equation:
[0410]
[0411] Equation (5) can be solved using the Wiener-Hopf equation such as ALF. The derived filter coefficients are then quantized to integer type and signaled in the block level.
[0412] In yet another embodiment, the weighting factors may be derived at the encoder side and then signaled to the decoder in the bitstream. Denote the N prediction block candidates as P1, P2, ..., PN , and denote the current block as X, the weighting factor can be solved by the following equation:
[0413]
[0414] Equation (6) can be solved using LDL decomposition or Gaussian elimination.
[0415] In yet another embodiment, weighting factors are derived based on the templates, and the derived weighting factors are applied to the prediction block candidates to generate the final prediction block. The templates of the prediction candidates are represented as T1, T2, ..., T N , the current block is T, then the following equation can be used to derive the weighting factors:
[0416]
[0417] The Wiener-Hopf equation can be used to solve equation (7). Then, the final prediction block can be calculated as Among them, P i represents the i-th prediction block candidate.
[0418] The IBC mode exploits non-local correlation to improve prediction accuracy, searching for similar blocks and using them to generate the final prediction block. In this embodiment, a combination of non-local mean filtering and multi-hypothesis IBC is proposed, as described below. In some examples, non-local mean filtering is an image processing algorithm used for image denoising. When executed, the algorithm takes the mean of all pixels in the image, weighted by their similarity to a target pixel. In the first step, N prediction block candidates are searched and identified, as in IBC. In the second step, weighting factors are calculated as follows.
[0419]
[0420] Among them, D i Used to measure the distance between the template of the i-th prediction block candidate and the template of the current block, h is used as the weighted degree or weighted strength, and Z[i] is the normalization constant:
[0421]
[0422] In order to calculate the weighting factors in equation (8), the strength of the weighting should be determined first. In this disclosure, several methods are proposed to determine the strength of the weighting or the strength of the weighting.
[0423] In the first approach, a weighted strength candidate list containing some typical weighted strength values is defined and fixed at both the encoder and decoder sides. At the encoder side, the weighted strength values are checked using rate-distortion optimization, and the best weighted strength value is identified and signaled to the decoder side in the bitstream.
[0424] In the second method, the weighted strength value is estimated using the template of the prediction block candidate and the template of the current block. The templates of the prediction candidates are represented as T1, T2, ..., T N , and denote the current block as T. The weighted intensity value can then be solved using the following equation:
[0425]
[0426] In a third method, the weighted intensity value may be estimated using the QP value and variance of the template of the current block, ie, the relationship among the weighted intensity value, the QP value, and the template variance may be fitted offline.
[0427] In order to better utilize the non-local correlation in IBC, in this embodiment, singular value decomposition (SVD) is used to generate the final prediction block from the prediction block candidates. The width and height of the current block are denoted as W and H, and the area of the current block is denoted as d=H×W.
[0428] Step 1. Search and identify K prediction block candidates y as done in FIBC i .
[0429] Step 2. The K prediction block candidates of the current block y constitute a block group G and are arranged as a matrix:
[0430]
[0431] where Y G is a matrix with size d×K (by arranging each candidate in group G as a column vector).
[0432] Step 3. Matrix Y G Perform SVD decomposition.
[0433]
[0434] Step 4. For the singular value matrix Λ G Applies a soft thresholding operation.
[0435]
[0436] Where softTh() is the threshold τ to reduce Λ G The diagonal elements of . G The kth diagonal element in is given by the nonlinear function D τ(k)Zoom out at level τ(k):
[0437] D τ(k) :λ k,τ(k) =max(|λ k |-τ(k),0) (14)
[0438] Λ G,τ is the singular value λ at the diagonal position reduced by k,τ(k) The matrix composed of .
[0439] Step 5. Perform inverse SVD to obtain the filtered patch set.
[0440]
[0441] One of the key steps is to determine the threshold value of each diagonal element in step 4. In the present disclosure, the threshold value is calculated as follows. The threshold value is estimated for each group of image blocks using the following equation:
[0442]
[0443] where σ n,G is the standard deviation of the noise, and σ x,G,k is the standard deviation of the original block in the kth dimension of the SVD space of group G.
[0444] The deviation of the original block in the SVD space is estimated as follows.
[0445]
[0446] in yes The kth singular value of . When σ x,G,k When α is zero, the soft threshold operation is skipped. In addition, the deviation of the noise is estimated using the deviation of the prediction block using a power function parameterized by α and β.
[0447] σ n =α×σ y β (18)
[0448] where σ y The calculation is as follows:
[0449]
[0450] Here y k (i) represents the prediction block candidate vector y k The i-th pixel of .
[0451] Multi-hypothesis IBC signaling
[0452] In the present disclosure, the proposed multi-hypothesis IBC can be used as a replacement for the current IBC mode, or the encoder can adaptively select the IBC mode or the multi-hypothesis IBC mode.
[0453] In some embodiments, multi-hypothesis IBC can be used as an alternative to the current IBC mode, ie, always using multiple hypotheses for prediction.
[0454] In yet another embodiment, one of the multi-hypothesis IBC methods in the above section is used in conjunction with the current IBC mode.A flag is signaled in the bitstream to indicate whether the multi-hypothesis IBC mode is applied to a CU.
[0455] In yet another embodiment, more than one of the multi-hypothesis IBC methods described in the previous section is used in conjunction with the current IBC mode. First, a flag is signaled in the bitstream to indicate whether the multi-hypothesis IBC mode is applied. Then, an index is signaled to indicate which of the multi-hypothesis IBC methods is applied to the CU.
[0456] In yet another embodiment, the multi-hypothesis IBC method in the above section is used in conjunction with the current IBC mode. Multi-hypothesis IBC can be used as an alternative to the current IBC mode based on certain coded information of the current block (e.g., SAD (sum of absolute difference), SSE (sum of squared error), quantization parameter (QP) associated with TB / CB and / or slice, CU's neighbor prediction mode (e.g., IBC mode or intra or inter), and / or slice type (e.g., I slice, P slice, or B slice)).
[0457] Combined intra-frame TMP intra-frame / inter-frame prediction
[0458] According to one or more embodiments of the present disclosure, a coding block encoded using the intra-frame TMP mode is combined with a coding block encoded using the intra-frame mode or the inter-frame mode. Different methods can be used to achieve this goal.
[0459] In the first method, the decoder / encoder may combine a coding block encoded in intra-frame TMP mode with a coding block encoded in intra-frame mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, a coding block encoded in intra-frame TMP mode is regarded as a coding block encoded in inter-frame merge mode, and it is combined with a coding block encoded in planar intra-frame prediction mode. In another example, similar to the combination of CIIP with TIMD and TM merge technology in ECM, a coding block encoded in intra-frame TMP mode is combined with a coding block encoded in intra-frame prediction mode derived from TIMD.
[0460] In the second method, the decoder / encoder may combine a coding block encoded in intra-frame TMP mode with a coding block encoded in inter-frame mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, a coding block encoded in intra-frame TMP mode is regarded as a coding block encoded in plane intra-frame mode, and it is combined with a coding block encoded in inter-frame merge mode. In another example, a coding block encoded in intra-frame TMP mode is regarded as a coding block encoded in inter-frame merge mode, and it is combined with a coding block encoded in inter-frame merge mode by equal averaging.
[0461] In a third method, the decoder / encoder may combine coding blocks encoded in intra-frame TMP mode with coding blocks encoded in intra-frame mode and coding blocks encoded in inter-frame mode. Various methods can be used in this combination. In one example, coding blocks encoded in intra-frame TMP mode, coding blocks encoded in intra-frame mode, and coding blocks encoded in inter-frame mode are directly combined by equal averaging. In another example, coding blocks encoded in intra-frame TMP mode are first combined separately from coding blocks encoded in intra-frame mode and inter-frame mode, as presented in the first and second methods. Then, the results of the separate combinations are combined by equal averaging.
[0462] GPM with intra-frame TMP and intra-frame TMP prediction
[0463] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the GPM tool in the form of a GPM with intra TMP and intra TMP prediction.Different approaches can be used to achieve this goal.
[0464] In the first approach, the two "inter" parts of the GPM with inter and inter prediction methods in VVC are replaced with intra TMP. This means that the two intra TMP prediction results are weighted averaged with each other according to the dividing line in the coding block. The weights can be obtained by referring to the GPM with inter and inter prediction methods in VVC.
[0465] In the second approach, the two "inter" parts of the GPM with inter and intra prediction methods in the ECM are replaced with intra TMP, where some template matching tools can be utilized to further improve the encoding and decoding performance.
[0466] GPM with intra-frame TMP and intra-frame prediction
[0467] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the GPM tool in the form of a GPM with intra TMP and intra prediction.Different approaches can be used to achieve this goal.
[0468] In the first method, the "inter" part of the GPM with inter and intra prediction methods in the ECM is replaced with the intra TMP, where the intra TMP prediction result is weighted averaged with the intra prediction result to obtain the final prediction signal.
[0469] GPM with intra-frame TMP and inter-frame prediction
[0470] According to one or more embodiments of the present disclosure, the intra TMP tool is combined with the GPM tool in the form of a GPM with intra TMP and inter prediction.Different approaches can be used to achieve this goal.
[0471] In the first method, an "inter" part of the GPM with inter and inter prediction methods in VVC is replaced by intra TMP, where the intra TMP prediction result is weighted averaged with the inter merged prediction result to obtain the final prediction signal.
[0472] In the second approach, an "inter" portion of the GPM with inter and inter prediction methods in the ECM is replaced with an intra TMP, where some template matching tools can be utilized to further improve the encoding and decoding performance.
[0473] Simplified intra TMP and intra prediction combination in GPM form
[0474] According to one or more embodiments of the present disclosure, the intra-frame TMP tool is combined with the GPM tool in the form of a simplified GPM with intra-frame TMP and intra-frame prediction, such as intra-frame TMP and intra-frame prediction combined in a specific partitioning mode, which can save the bit overhead of the partitioning representation. Different methods can be used to achieve this goal.
[0475] In the first method, for a dividing line, such as 45 degrees, the upper left part of the coding block is encoded using the intra-frame prediction mode, and the lower right part of the coding block is encoded using the intra-frame TMP prediction mode, and then they are averaged in the GPM form to obtain the final prediction signal.
[0476] Combined IBC with TIMD mode
[0477] According to one or more embodiments of the present disclosure, an IBC tool is combined with a TIMD tool. Different approaches can be used to achieve this goal.
[0478] In the first approach, the IBC mode is regarded as an intra prediction mode added to the MPM list, and then the IBC mode is compared with other intra prediction modes in the MPM list using template matching cost. Finally, the two modes with the minimum and second minimum costs are fused using the TIMD method to obtain the final prediction result.
[0479] In the second method, the conventional TIMD prediction result is first obtained, then the template matching cost of the IBC mode and the conventional TIMD prediction result is calculated, and finally the TIMD method is used to fuse the IBC mode and the conventional TIMD prediction result to obtain the final prediction result.
[0480] According to one or more embodiments of the present disclosure, the intra-frame TMP tool is combined with the TIMD tool. Different methods can be used to achieve this goal.
[0481] In the first method, the Intra TMP mode is regarded as an intra prediction mode added to the MPM list, and then the Intra TMP mode is compared with other intra prediction modes in the MPM list using template matching cost. Finally, the two modes with the minimum and second minimum costs are fused using the TIMD method to obtain the final prediction result.
[0482] In the second method, the conventional TIMD prediction result is first obtained, then the template matching cost of the Intra TMP mode and the conventional TIMD prediction result is calculated, and finally the Intra TMP mode and the conventional TIMD prediction result are fused using the TIMD method to obtain the final prediction result.
[0483] Combined intra-frame TMP and LIC
[0484] According to one or more embodiments of the present disclosure, the intra-frame TMP tool is combined with the LIC tool. Different approaches can be used to achieve this goal.
[0485] In the first approach, the intra TMP mode is treated as an inter mode, and LIC is used to model the local illumination variation between the current block and its intra TMP prediction block according to the local illumination variation between the current block template and the reference block template. The function is a linear equation as used in the conventional LIC method.
[0486] Combining IBC and OBMC
[0487] According to one or more embodiments of the present disclosure, an IBC tool is combined with an OBMC tool. Different approaches can be used to achieve this goal.
[0488] In the first approach, IBC mode is considered as inter mode, and conventional OBMC method is applied to refine (using block vector information of neighboring blocks with weighted prediction) the top and left boundary pixels of the IBC-coded CU.
[0489] In the second method, the IBC mode is regarded as an inter mode, and the template matching based OBMC method is applied to refine (using the template matching based method) the top and left boundary pixels of the IBC coded CU.
[0490] Combined intra-frame TMP and OBMC
[0491] According to one or more embodiments of the present disclosure, the intra-frame TMP tool is combined with the OBMC tool. Different approaches can be used to achieve this goal.
[0492] In the first approach, the intra TMP mode is treated as an inter mode, and the conventional OBMC method is applied to refine (using block vector information of neighboring blocks with weighted prediction) the top and left boundary pixels of the intra TMP coded CU.
[0493] In the second method, the intra TMP mode is treated as the inter mode, and the template matching based OBMC method is applied to refine (using the template matching based method) the top and left boundary pixels of the intra TMP coded CU.
[0494] Figure 22 A computing environment (or computing device) 1610 is shown coupled to a user interface 1650. The computing environment 1610 may be part of a data processing server. In some embodiments, the computing device 1610 may perform any of the various methods or processes (such as encoding / decoding methods or processes) described above according to various examples of the present disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.
[0495] The processor 1620 generally controls the overall operation of the computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the above-described method. In addition, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.
[0496] The memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1630 may include predetermined software 1632. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, etc. The memory 1630 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0497] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (e.g., keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 may be coupled to an encoder and a decoder.
[0498] Figure 23 is a flowchart illustrating a method for video decoding according to an example of the present disclosure.
[0499] In step 2301 , at the decoder side, the processor 1620 may obtain a plurality of prediction block candidates, wherein at least one prediction block candidate among the plurality of prediction block candidates is encoded based on an intra block copy (IBC) mode for a current block.
[0500] In step 2302 , the processor 1620 may obtain a final prediction for a current block based on a plurality of prediction block candidates.
[0501] In some examples, processor 1620 may also obtain multiple weights corresponding to multiple prediction block candidates, and in step 2302, the processor may obtain a final prediction for the current block based on the multiple prediction block candidates and the multiple weights.
[0502] In some examples, the processor 1620 may search for multiple block candidates in step 2301, wherein the block vector (BV) matching cost of the multiple block candidates is lower than the BV matching cost of other block candidates. In addition, in step 2301, the processor 1620 may select the multiple block candidates as multiple prediction block candidates. In some examples, the BV matching cost may be measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE). In some examples, the BV matching cost may be calculated based on a comparison of the distortion of the corresponding reference block with a neighboring block (such as a top neighboring block or a left neighboring block).
[0503] In some examples, processor 1620 may select multiple block candidates based on a predefined mode in step 2301, where the predefined mode may include a planar mode.
[0504] In some examples, in step 2301, processor 1620 may select multiple block candidates based on predefined patterns of neighboring blocks, where the predefined patterns of neighboring blocks may include one of a top predefined pattern or a left predefined pattern. In some examples, the top predefined pattern of the current block inherits the predefined pattern of the top neighboring block of the current block, and the left predefined pattern inherits the predefined pattern of the left neighboring block of the current block.
[0505] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may set each of the multiple weights to the inverse of an integer N, where N is the number of the multiple prediction block candidates.
[0506] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may derive multiple weights based on multiple block vector (BV) matching costs. In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may also obtain multiple predicted BV matching costs for the multiple prediction block candidates as multiple BV matching costs. In some examples, the multiple BV matching costs are measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE).
[0507] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, the processor 1620 may derive multiple weights based on a block size of the current block or a syntax element signaled by the encoder, wherein the syntax element is signaled at one of the following levels: sequence parameter set (SPS), decoded picture set (DPS), video parameter set (VPS), supplemental enhancement information (SEI), adaptation parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, coding tree unit (CTU), coding unit (CU), sub-block, or sample level.
[0508] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may receive the multiple weights signaled in a bitstream by an encoder.
[0509] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may obtain multiple weights based on multiple templates. In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may also obtain multiple prediction block templates of the multiple prediction block candidates as multiple templates. Furthermore, to obtain multiple weights corresponding to the multiple prediction block candidates, processor 1620 may obtain a current block template of the current block. Furthermore, to obtain multiple weights based on multiple templates, processor 1620 may obtain multiple weights based on multiple prediction block templates and the current block template.
[0510] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may further use non-local mean filtering to obtain multiple weights. In some examples, to obtain multiple means using non-local mean filtering, processor 1620 may obtain multiple distances based on multiple prediction block candidates and the current block; further, to obtain multiple means using non-local mean filtering, processor 1620 may obtain multiple weights based on multiple distances and weighted strengths. In some examples, processor 1620 may further determine the weighted strength. In some examples, non-local mean filtering is an algorithm in image processing used for image denoising, wherein the algorithm, when executed, uses the mean of all pixels in the image (weighted by the degree of similarity between the pixels and the target pixel).
[0511] In some examples, to determine the weighted strength, processor 1620 may define and fix a weighted strength candidate list including a plurality of typical weighted strength values, and select the typical weighted strength value as the optimal weighted strength value based on a flag signaled in the bitstream. Alternatively or additionally, processor 1620 may receive the optimal weighted strength value signaled in the bitstream.
[0512] In some examples, to determine the weighted strength, processor 1620 may obtain multiple predicted CU templates of multiple predicted CU candidates, obtain a current CU template of the current CU, and estimate the weighted strength based on the multiple predicted CU templates and the current CU template.
[0513] In some examples, to determine the weighted strength, processor 1620 may estimate the weighted strength using a quantization parameter (QP) value and a variance of a current CU template for the current CU.
[0514] In some examples, in step 2302, processor 1620 may perform singular value decomposition (SVD) on a matrix generated from a plurality of prediction block candidates; further, in step 2302, processor 1620 may obtain a final prediction based on a result of performing SVD.
[0515] In some examples, processor 1620 may also receive a flag signaled in the bitstream, wherein the flag indicates whether the multi-hypothesis IBC mode is applied to the current block. In some examples, processor 1620 may further receive an index signaled in the bitstream in response to determining that the flag indicates that the multi-hypothesis IBC mode is applied to the current block, wherein the index indicates the multi-hypothesis IBC method applied to the current block.
[0516] In some examples, processor 1620 may further determine that a multi-hypothesis IBC mode is applied to the current block based on encoded information, wherein the encoded information includes at least one of the following information: sum of absolute differences (SAD), sum of squared errors (SSE), quantization parameter (QP), neighbor prediction mode of the current block, or slice type.
[0517] Figure 24 is a flowchart illustrating a method for video encoding, which corresponds to Figure 23 The method for video decoding shown in .
[0518] In step 2401 , at the encoder side, the processor 1620 may obtain a plurality of prediction block candidates, wherein at least one prediction block candidate among the plurality of prediction block candidates is encoded based on an intra block copy (IBC) mode for a current block.
[0519] In step 2402 , the processor 1620 may obtain a final prediction for a current block based on a plurality of prediction block candidates.
[0520] In step 2403 , the processor 1620 may obtain a final prediction for the current block based on the IBC mode.
[0521] In step 2404 , the processor 1620 may generate a bitstream based on the final prediction.
[0522] In some examples, at least one prediction block candidate from the plurality of prediction block candidates is encoded based on the IBC mode for the current block. In some examples, processor 1620 may also obtain a plurality of weights corresponding to the plurality of prediction block candidates, and in step 2402, the processor may obtain a final prediction for the current block based on the plurality of prediction block candidates and the plurality of weights.
[0523] In some examples, the processor 1620 may search for multiple block candidates in step 2401, wherein the block vector (BV) matching cost of the multiple block candidates is lower than the BV matching cost of other block candidates. In addition, in step 2401, the processor 1620 may select the multiple block candidates as multiple prediction block candidates. In some examples, the BV matching cost may be measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE). In some examples, the BV matching cost may be calculated based on a comparison of the distortion of the corresponding reference block with a neighboring block (such as a top neighboring block or a left neighboring block).
[0524] In some examples, processor 1620 may select a plurality of block candidates based on a predefined mode in step 2401 , where the predefined mode may include a planar mode.
[0525] In some examples, in step 2401, processor 1620 may select multiple block candidates based on predefined patterns of neighboring blocks, where the predefined patterns of neighboring blocks may include one of a top predefined pattern or a left predefined pattern. In some examples, the top predefined pattern of the current block inherits the predefined pattern of the top neighboring block of the current block, and the left predefined pattern inherits the predefined pattern of the left neighboring block of the current block.
[0526] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may set each of the multiple weights to the inverse of an integer N, where N is the number of the multiple prediction block candidates.
[0527] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may derive multiple weights based on multiple block vector (BV) matching costs. In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may also obtain multiple predicted BV matching costs for the multiple prediction block candidates as multiple BV matching costs. In some examples, the multiple BV matching costs are measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE).
[0528] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, the processor 1620 may also signal multiple weights to the decoder based on a block size of the current block or a syntax element signaled by the encoder, where the syntax element is signaled at one of the following levels: sequence parameter set (SPS), decoded picture set (DPS), video parameter set (VPS), supplemental enhancement information (SEI), adaptation parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, coding tree unit (CTU), coding unit (CU), sub-block, or sample level.
[0529] In some examples, in order to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 can obtain multiple weights based on multiple prediction block candidates and the current block; in addition, in order to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 can signal multiple weights in the bitstream.
[0530] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may obtain multiple weights based on multiple templates. In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may also obtain multiple prediction block templates of the multiple prediction block candidates as multiple templates. Furthermore, to obtain multiple weights corresponding to the multiple prediction block candidates, processor 1620 may obtain a current block template of the current block. Furthermore, to obtain multiple weights based on multiple templates, processor 1620 may obtain multiple weights based on multiple prediction block templates and the current block template.
[0531] In some examples, to obtain multiple weights corresponding to multiple prediction block candidates, processor 1620 may further use non-local mean filtering to obtain multiple weights. In some examples, to obtain multiple means using non-local mean filtering, processor 1620 may obtain multiple distances based on multiple prediction block candidates and the current block; further, to obtain multiple means using non-local mean filtering, processor 1620 may obtain multiple weights based on multiple distances and weighted strengths. In some examples, processor 1620 may further determine the weighted strength. In some examples, non-local mean filtering is an algorithm in image processing used for image denoising, wherein the algorithm, when executed, uses the mean of all pixels in the image (weighted by the degree of similarity between the pixels and the target pixel).
[0532] In some examples, to determine the weighted strength, processor 1620 may define and fix a weighted strength candidate list comprising multiple typical weighted strength values, examine the multiple typical weighted strength values using rate-distortion optimization, identify an optimal weighted strength value, and signal the optimal weighted strength value in the bitstream.
[0533] In some examples, to determine the weighted strength, processor 1620 may obtain multiple predicted CU templates of multiple predicted CU candidates, obtain a current CU template of the current CU, and estimate the weighted strength based on the multiple predicted CU templates and the current CU template.
[0534] In some examples, to determine the weighted strength, processor 1620 may estimate the weighted strength using a quantization parameter (QP) value and a variance of a current CU template for the current CU.
[0535] In some examples, in step 2402, processor 1620 may perform singular value decomposition (SVD) on a matrix generated based on multiple prediction block candidates; further, in step 2402, processor 1620 may obtain a final prediction based on a result of performing SVD.
[0536] In some examples, processor 1620 may further set a flag signaled in the bitstream, wherein the flag indicates whether the multi-hypothesis IBC mode is applied to the current block. In some examples, processor 1620 may further set an index signaled in the bitstream in response to determining that the flag indicates that the multi-hypothesis IBC mode is applied to the current block, wherein the index indicates the multi-hypothesis IBC method applied to the current block.
[0537] In some examples, processor 1620 may further determine that a multi-hypothesis IBC mode is applied to the current block based on encoded information, wherein the encoded information includes at least one of the following information: sum of absolute differences (SAD), sum of squared errors (SSE), quantization parameter (QP), neighbor prediction mode of the current block, or slice type.
[0538] In some examples, a device for video encoding and decoding is provided. The device includes a processor 1620 and a memory 1640 coupled to the processor 1620, wherein the memory 1640 is configured to store instructions executable by the processor; wherein the processor 1620 is configured to perform the following when executing the instructions: Figures 23 to 24 In order to perform any of the methods shown in Figure 23 The method shown in FIG. 16 is also configured to store the bit stream. Figure 24 In the method shown, the processor 1620 is further configured to generate a bitstream and store the bitstream in a memory when executing the instructions.
[0539] In an embodiment, a non-transitory computer-readable storage medium including, for example, a plurality of programs in a memory 1630 and / or storing a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method is also provided. The plurality of programs can be executed by a processor 1620 in a computing environment 1610 to perform the above method. In one example, the plurality of programs can be executed by a processor 1620 in a computing environment 1610 to (for example, from Figure 2 The video encoder 20 in the computing environment 1610 receives a bit stream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned decoding method according to the received bit stream or data stream. In another example, the multiple programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned encoding method to encode the video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bit stream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to Figure 3Alternatively, a non-transitory computer readable storage medium may store the bit stream or data stream generated by the encoder (e.g., Figure 2 The video encoder 20 in FIG. 1 generates a video signal for use by a decoder (eg, Figure 3 A bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) used by the video decoder 30 in the video decoder 30 when decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0540] In an embodiment, a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method is provided. In an embodiment, a bit stream including coded video information generated by the above encoding method or coded video information to be decoded by the above decoding method is provided.
[0541] In an embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 having stored therein a plurality of programs that can be executed by the one or more processors, wherein the one or more processors are configured to perform the above-described method when executing the plurality of programs.
[0542] In an embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, wherein the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product is also provided, including, for example, a plurality of programs in a memory 1630, which can be executed by a processor 1620 in a computing environment 1610 to perform the above method. For example, the computer program product can include a non-transitory computer-readable storage medium.
[0543] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.
[0544] In an embodiment, a method for storing a bitstream is further provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.
[0545] In an embodiment, a method for transmitting a bit stream generated by the above encoder is also provided. In an embodiment, a method for receiving a bit stream to be decoded by the above decoder is also provided.
[0546] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.
[0547] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual circumstances. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined, or deleted according to actual needs.
[0548] The examples are chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present disclosure is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present disclosure.
Claims
1. A method for video decoding, comprising: Obtaining, by a decoder, a plurality of prediction block candidates, wherein at least one prediction block candidate among the plurality of prediction block candidates is encoded based on an intra block copy (IBC) mode for a current block; and A final prediction for the current block is obtained by the decoder based on the multiple prediction block candidates.
2. The method according to claim 1, further comprising: Obtaining, by the decoder, a plurality of weights corresponding to the plurality of prediction block candidates; and The obtaining, by the decoder, a final prediction for the current block based on the multiple prediction block candidates includes: A final prediction for the current block is obtained by the decoder based on the multiple prediction block candidates and the multiple weights.
3. The method according to claim 1, wherein Obtaining, by the decoder, the plurality of prediction block candidates comprises: searching, by the decoder, for a plurality of block candidates, wherein block vector (BV) matching costs of the plurality of block candidates are lower than BV matching costs of other block candidates; and The plurality of block candidates are selected by the decoder as the plurality of prediction block candidates.
4. The method according to claim 3, wherein: The BV matching cost is measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE).
5. The method according to claim 1, wherein Obtaining, by the decoder, the plurality of prediction block candidates comprises: selecting, by the decoder, the plurality of block candidates based on a predefined pattern; Wherein, the predefined mode includes a plane mode.
6. The method according to claim 1, wherein Obtaining, by the decoder, the plurality of prediction block candidates comprises: selecting, by the decoder, the plurality of block candidates based on a predefined pattern of neighboring blocks; The predefined pattern of the neighboring block includes one of a top predefined pattern and a left predefined pattern.
7. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: Each of the plurality of weights is set by the decoder to a reciprocal of an integer N, where N is the number of the plurality of prediction block candidates.
8. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: The plurality of weights are derived by the decoder based on a plurality of block vector (BV) matching costs.
9. The method according to claim 8, wherein Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates further comprises: A plurality of prediction BV matching costs of the plurality of prediction block candidates are obtained by the decoder as the plurality of BV matching costs.
10. The method according to claim 8, wherein The plurality of BV matching costs are measured using one of a sum of absolute differences (SAD) or a sum of squared errors (SSE).
11. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: deriving, by the decoder, the plurality of weights based on a block size or a syntax element of the current block; wherein the syntax element is signaled at one of the following levels: Sequence parameter set (SPS), decoded picture set (DPS), video parameter set (VPS), supplemental enhancement information (SEI), adaptation parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, coding tree unit (CTU), coding unit (CU), sub-block, or sample level.
12. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: The plurality of weights signaled in a bitstream is received by the decoder.
13. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: The plurality of weights are obtained by the decoder based on a plurality of templates.
14. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: Obtaining, by the decoder, a plurality of prediction block templates of the plurality of prediction block candidates as the plurality of templates; and Obtaining, by the decoder, a current block template of the current block; and The plurality of weights are obtained by the decoder based on the plurality of prediction block templates and the current block template.
15. The method according to claim 2, wherein: Obtaining, by the decoder, the plurality of weights corresponding to the plurality of prediction block candidates further comprises: The plurality of weights are obtained by the decoder using non-local means filtering.
16. The method according to claim 15, wherein Obtaining the plurality of weights by using the non-local means filtering includes: obtaining a plurality of distances based on the plurality of prediction block candidates and the current block; and The plurality of weights are obtained based on the plurality of distances and the weighted intensities.
17. The method according to claim 16, further comprising: The decoder determines the weighted strength by one of the following steps: defining and fixing, by the decoder, a weighted intensity candidate list comprising a plurality of typical weighted intensity values; and selecting, by the decoder, the typical weighted intensity value as the optimal weighted intensity value based on a flag signaled in a bitstream; receiving, by the decoder, an optimal weighted strength value signaled in a bitstream; Obtaining, by the decoder, a plurality of prediction block templates of the plurality of prediction block candidates, obtaining, by the decoder, a current block template of the current block, and estimating, by the decoder, the weighted strength based on the plurality of prediction block templates and the current block template; or, The weighted strength is estimated by the decoder using a quantization parameter (QP) value and a variance of a current CU template of the current CU.
18. The method according to claim 1, wherein Obtaining, by the decoder, a final prediction for the current block based on the multiple prediction block candidates includes: performing, by the decoder, singular value decomposition (SVD) on a matrix generated according to the plurality of prediction block candidates; and The final prediction is obtained by the decoder based on a result of performing the SVD.
19. The method of claim 1, further comprising: A flag signaled in a bitstream is received by the decoder, wherein the flag indicates whether a multi-hypothesis (IBC) mode is applied to the current block.
20. The method according to claim 19, further comprising: In response to determining that the flag indicates that the multi-hypothesis IBC mode is applied to the current block, receiving, by the decoder, an index signaled in the bitstream, wherein the index indicates a multi-hypothesis IBC method applied to the current block.
21. The method of claim 1 , further comprising: Determining, by the decoder based on the encoded information, that a multi-hypothesis IBC mode is applied to the current block; The encoded information includes at least one of the following information: Sum of Absolute Difference (SAD), Sum of Squared Error (SSE), Quantization Parameter (QP), neighbor prediction mode of the current block, or slice type.
22. A method for video encoding, comprising: Obtaining, by an encoder, a plurality of prediction block candidates, wherein at least one prediction block candidate among the plurality of prediction block candidates is encoded based on an intra block copy (IBC) mode for a current block; Obtaining, by the encoder, a final prediction for the current block based on the multiple prediction block candidates; Obtaining, by the encoder, a final prediction for the current block based on the IBC mode; and A bitstream is generated by the encoder based on the final prediction.
23. The method according to claim 22, further comprising: Obtaining, by the encoder, a plurality of weights corresponding to the plurality of prediction block candidates; and The obtaining, by the encoder, a final prediction for the current block based on the multiple prediction block candidates includes: A final prediction for the current block is obtained by the encoder based on the multiple prediction block candidates and the multiple weights.
24. The method according to claim 22, wherein Obtaining, by the encoder, the plurality of prediction block candidates comprises: searching, by the encoder, for a plurality of block candidates, wherein block vector (BV) matching costs of the plurality of block candidates are lower than BV matching costs of other block candidates; and The plurality of block candidates are selected by the encoder as the plurality of prediction block candidates.
25. The method according to claim 24, wherein The BV matching cost is measured using one of the sum of absolute differences (SAD) or the sum of squared errors (SSE).
26. The method according to claim 22, wherein Obtaining, by the encoder, the plurality of prediction block candidates comprises: selecting, by the encoder, the plurality of block candidates based on a predefined pattern; Wherein, the predefined mode includes a plane mode.
27. The method according to claim 22, wherein Obtaining, by the encoder, the plurality of prediction block candidates comprises: selecting, by the encoder, the plurality of block candidates based on a predefined pattern of neighboring blocks; The predefined pattern of the neighboring block includes one of a top predefined pattern and a left predefined pattern.
28. The method according to claim 23, wherein Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: Each of the plurality of weights is set by the encoder to a reciprocal of an integer N, where N is the number of the plurality of prediction block candidates.
29. The method according to claim 23, wherein Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: The plurality of weights are derived by the encoder based on a plurality of block vector (BV) matching costs.
30. The method according to claim 29, wherein Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates further comprises: A plurality of prediction BV matching costs of the plurality of prediction block candidates are obtained by the encoder as the plurality of BV matching costs.
31. The method according to claim 29, wherein The plurality of BV matching costs are measured using one of a sum of absolute differences (SAD) or a sum of squared errors (SSE).
32. The method of claim 23, wherein: Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates further comprises: signaling, by the encoder, a block size of the current block or a syntax element indicating the plurality of weights; wherein the syntax element is signaled at one of the following levels: Sequence parameter set (SPS), decoded picture set (DPS), video parameter set (VPS), supplemental enhancement information (SEI), adaptation parameter set (APS), picture parameter set (PPS), picture header (PH), slice header (SH), region, coding tree unit (CTU), coding unit (CU), sub-block, or sample level.
33. The method according to claim 23, wherein Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: Obtaining, by the encoder, the plurality of weights based on the plurality of prediction block candidates and the current block; and The plurality of weights are signaled in a bitstream by the encoder.
34. The method of claim 23, wherein: Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: The plurality of weights are obtained by the encoder based on a plurality of templates.
35. The method of claim 23, wherein: Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates comprises: Obtaining, by the encoder, a plurality of prediction block templates of the plurality of prediction block candidates as the plurality of templates; Obtaining, by the encoder, a current block template of the current block; and The plurality of weights are obtained by the encoder based on the plurality of templates and the current block template.
36. The method of claim 23, wherein: Obtaining, by the encoder, the plurality of weights corresponding to the plurality of prediction block candidates further comprises: The plurality of weights are obtained by the encoder using non-local means filtering.
37. The method according to claim 36, wherein Obtaining, by the encoder, the plurality of weights by using the non-local means filtering comprises: obtaining a plurality of distances based on the plurality of prediction block candidates and the current block; and The plurality of weights are obtained based on the plurality of distances and the weighted intensities.
38. The method of claim 37, further comprising: The encoder determines the weighted strength by one of the following steps: defining and fixing, by the encoder, a weighted intensity candidate list comprising a plurality of typical weighted intensity values, the plurality of typical weighted intensity values being checked by the encoder using rate-distortion optimization; identifying, by the encoder, an optimal weighted intensity value; and signaling, by the encoder, the optimal weighted intensity value in a bitstream; Obtaining, by the encoder, a plurality of prediction block templates of the plurality of prediction block candidates; Obtaining, by the encoder, a current block template of the current block; and estimating, by the encoder, the weighted strength based on the plurality of prediction block templates and the current block template; or The weighted strength is estimated by the encoder using a quantization parameter (QP) value and a variance of a current block template of the current block.
39. The method of claim 22, wherein: Obtaining, by the encoder, a final prediction for the current block based on the plurality of prediction block candidates includes: performing, by the encoder, singular value decomposition (SVD) on a matrix generated according to the plurality of prediction block candidates; and The final prediction is obtained by the encoder based on a result of performing the SVD.
40. The method of claim 22, further comprising: A flag is signaled by the encoder in a bitstream, wherein the flag indicates whether a multi-hypothesis (IBC) mode is applied to the current block.
41. The method of claim 40, further comprising: On condition that the flag indicates that the multi-hypothesis IBC mode is applied to the current block, an index is signaled by the encoder in the bitstream, wherein the index indicates the multi-hypothesis IBC method applied to the current block.
42. The method of claim 22, further comprising: Determining, by the encoder based on the encoded information, that a multi-hypothesis IBC mode is applied to the current block; The encoded information includes at least one of the following information: Sum of Absolute Difference (SAD), Sum of Squared Error (SSE), Quantization Parameter (QP), neighbor prediction mode of the current block, or slice type.
43. An apparatus for video decoding, comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions and a bitstream, the instructions being executable by the one or more processors, Wherein, when executing the instructions, the one or more processors are configured to use the bitstream to perform the method according to any one of claims 1-21.
44. An apparatus for video encoding, comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions, the instructions being executable by the one or more processors, Wherein, when executing the instructions, the one or more processors are configured to perform the method according to any one of claims 22 to 42 to generate a bitstream and store the bitstream in the memory.
45. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to any one of claims 1-21.
46. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 22-42.
47. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to any one of claims 1-21.
48. A non-transitory computer-readable storage medium for storing a bitstream generated by the method according to any one of claims 22-42.