Method and apparatus using intra block copy

By combining the geometric division mode with the intra-block copy mode, the problem of intra-block copying efficiency in the prior art is solved, and more efficient video decoding and encoding performance is achieved.

CN120036000APending Publication Date: 2025-05-23BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072280.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-10
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are inefficient in intra-block copying (IBC), resulting in poor encoding and decoding performance.

Method used

Using an IBC mode combined with a geometric division mode (GPM), the current encoding unit (CU) encoded based on the GPM combination of IBC mode is obtained by a decoder, and prediction is made based on this mode.

Benefits of technology

Improve the efficiency of video decoding, reduce the bit rate, and maintain video quality, and improve the encoding and codec performance of intra-block copying.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036000A_ABST
    Figure CN120036000A_ABST
Patent Text Reader

Abstract

Methods, apparatus, and non-transitory computer-readable storage media for video decoding and encoding are provided. In a method for video decoding, a decoder may obtain a current coding unit (CU) encoded based on an intra block copy (IBC) mode combined with a geometric partition mode (GPM). In addition, the decoder may obtain a prediction for the current CU based on the IBC mode combined with the GPM.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 414,895, filed on October 10, 2022, entitled "Methods and Devices with Intra Block Copy", the entire content of which is incorporated herein by reference for all purposes. Technical Field

[0003] The present disclosure relates to video encoding and compression, and more particularly but not limited to, methods and apparatuses for improving the encoding and decoding efficiency of Intra Block Copy (IBC). Background Art

[0004] Various electronic devices (such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc.) support digital video. The electronic devices send and receive or otherwise transmit digital video data via a communication network and / or store the digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited storage resources of the storage device, video data can be compressed using video encoding and decoding according to one or more video encoding and decoding standards before the video data is transmitted or stored. For example, video encoding and decoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Experts Group (MPEG) coding, etc. Video encoding and decoding generally employs prediction methods (such as inter - frame prediction, intra - frame prediction, etc.) that utilize the redundancy inherent in video data. Video encoding and decoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. Summary of the Invention

[0005] The present disclosure provides examples of techniques related to improving the intra - block copy method in a video encoding or decoding process.

[0006] According to a first aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain a current coding unit (CU) encoded based on an IBC mode combined with a Geometric Partitioning Mode (GPM). Additionally, the decoder may obtain a prediction for the current CU based on the IBC mode combined with the GPM.

[0007] According to a second aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may encode a current CU based on an IBC mode combined with a GPM. In addition, the encoder may send the current CU encoded based on the IBC mode combined with the GPM to a decoder.

[0008] According to a third aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain a first prediction for a current CU, wherein the first prediction is associated with an IBC mode. In addition, the decoder may obtain a second prediction for the current CU, wherein the second prediction is associated with one of an intra mode or an inter mode. In addition, the decoder may obtain a final prediction for the current CU based on the first prediction and the second prediction.

[0009] According to a fourth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may obtain a first prediction for a current CU, wherein the first prediction is associated with an IBC mode. In addition, the encoder may obtain a second prediction for the current CU, wherein the second prediction is associated with one of an intra mode or an inter mode. In addition, the encoder may obtain a final prediction for the current CU based on the first prediction and the second prediction.

[0010] According to a fifth aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder may obtain multiple block vectors for a current CU based on an IBC mode. In addition, the decoder may obtain a final prediction for the current CU based on the multiple block vectors.

[0011] According to a sixth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder may obtain multiple block vectors for a current CU based on an IBC mode. In addition, the encoder may obtain a final prediction for the current CU based on the multiple block vectors.

[0012] According to a seventh aspect of the present disclosure, a device for video decoding is provided. The device may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. In addition, the one or more processors are configured to execute the method according to the first aspect, the third aspect, or the fifth aspect when executing the instructions.

[0013] According to an eighth aspect of the present disclosure, a device for video encoding is provided. The device may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. In addition, the one or more processors are configured to execute the method according to the second aspect, the fourth aspect, or the sixth aspect when executing the instructions.

[0014] According to the ninth aspect of the present disclosure, a non-transitory computer-readable storage medium for storing computer-executable instructions is provided, and when the computer-executable instructions are executed by one or more computer processors, the one or more computer processors execute the method according to the first aspect, the third aspect or the fifth aspect mentioned above.

[0015] According to the tenth aspect of the present disclosure, a non-transitory computer-readable storage medium for storing computer-executable instructions is provided, and when the computer-executable instructions are executed by one or more computer processors, the one or more computer processors execute the method according to the second aspect, the fourth aspect or the sixth aspect above.

[0016] According to an eleventh aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first aspect, the third aspect or the fifth aspect.

[0017] According to a twelfth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by a method according to the second aspect, the fourth aspect or the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] A more particular description of examples of the present disclosure will be presented by reference to specific examples shown in the accompanying drawings. Given that these drawings depict only some examples and are therefore not to be considered limiting in scope, the examples will be described and explained with additional specificity and detail through use of the accompanying drawings.

[0019] Figure 1A is a block diagram illustrating a system for encoding and decoding video blocks according to some examples of the present disclosure.

[0020] Figure 1B Examples of some examples according to the present disclosure are shown. Figure 1E 4. The quadtree data structure of the final result of the partitioning process of CTU 400 depicted in FIG.

[0021] Figure 1C An encoded representation of a frame by first dividing the frame into a set of CTUs according to some examples of the present disclosure is shown.

[0022] Figure 1D A CTU including one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements for encoding and decoding samples of the coding tree blocks according to some examples of the present disclosure is shown.

[0023] Figure 2 is a block diagram illustrating an exemplary video encoder according to some examples of the present disclosure.

[0024] Figure 3 is a block diagram illustrating an exemplary video decoder according to some examples of the present disclosure.

[0025] Figure 4A is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0026] Figure 4B is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0027] Figure 4C is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0028] Figure 4D is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0029] Figure 4E is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0030] Figure 5 Diagram showing locations of spatial candidates according to some examples of the present disclosure.

[0031] Figure 6 A diagram showing candidate pairs considered for redundancy checking of spatial candidates according to some examples of the present disclosure.

[0032] Figure 7 Diagram showing scaling of motion vectors for temporal candidates according to some examples of the present disclosure.

[0033] Figure 8 A diagram showing candidate positions for temporal candidates according to some examples of the present disclosure.

[0034] Fig. 9 Diagram showing merge mode with motion vector difference (MMVD) search points according to some examples of the present disclosure.

[0035] Fig.10 Unidirectional prediction motion vector selection for geometric partitioning mode (GPM) according to some examples of the present disclosure is shown.

[0036] Fig.11 Top and left neighboring blocks used in CIIP weight derivation according to some examples of the present disclosure are shown.

[0037] Fig.12 The current CTU processing order and its available reference samples in the current CTU and the left CTU according to some examples of the present disclosure are shown.

[0038] Fig.13 Filling candidates for replacing zero vectors in an IBC list according to some examples of the present disclosure are shown.

[0039] Fig.14 Reference regions for IBC when CTU(m,n) is encoded according to some examples of the present disclosure are shown.

[0040] Fig.15 IBC reference areas for camera captured content are shown according to some examples of the present disclosure.

[0041] Figure 16A-16B A division method for angle modes according to some examples of the present disclosure is shown.

[0042] Figures 17A-17D A GPM with inter prediction and intra prediction is shown according to some examples of the present disclosure.

[0043] Fig.18 Edges on a template are shown according to some examples of the present disclosure.

[0044] Fig.19 is a diagram illustrating a computing environment coupled to a user interface according to some implementations of the present disclosure.

[0045] Fig. 20 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0046] Fig.21 is a diagram showing some examples according to the present disclosure. Fig. 20 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding.

[0047] Fig. 22 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0048] Fig.23 is a diagram showing some examples according to the present disclosure. Fig. 22 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding.

[0049] Fig.24 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0050] Fig.25 is a diagram showing some examples according to the present disclosure. Fig.24 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding. DETAILED DESCRIPTION

[0051] Reference will now be made in detail to specific embodiments, examples of which are shown in the accompanying drawings. In the following detailed description, many non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be implemented without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0052] The terms used in this disclosure are only suitable for the purpose of describing specific embodiments and are not intended to limit the present disclosure. In the present disclosure and the appended claims, singular forms ("a / an", "the", "said", etc.) are also intended to include plural forms, unless other meanings are clearly indicated throughout the disclosure. It should also be understood that the term "and / or" used in this disclosure indicates and includes one or any or all possible combinations of the listed multiple related items.

[0053] References throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that the particular feature, structure, or characteristic being described is included in at least one embodiment or example. Unless explicitly stated otherwise, features, structures, elements, or characteristics described in conjunction with one or some embodiments are also applicable to other embodiments.

[0054] Throughout the disclosure, unless otherwise explicitly stated, the terms "first", "second", "third", etc. are used only as references to related elements (e.g., devices, components, compositions, steps, etc.), and do not imply any spatial or temporal order. For example, "first device" and "second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be named arbitrarily.

[0055] The terms "module", "sub-module", "circuit", "sub-circuit", "circuitry", "sub-circuitry", "unit", or "sub-unit" may include a memory (shared, dedicated, or group) storing code or instructions that may be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components connected directly or indirectly. These components may or may not be physically attached to each other or located adjacent to each other.

[0056] As used herein, the terms "if" or "when" may be understood to mean "based on" or "in response to" depending on the context. These terms, if they appear in a claim, may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include the following steps: i) when or if condition X exists, perform function or action X', and ii) when or if condition Y exists, perform function or action Y'. The method may be implemented with the ability to perform function or action X' and the ability to perform function or action Y'. Therefore, both functions X' and Y' may be performed at different times during multiple executions of the method.

[0057] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components that are linked together directly or indirectly to perform a specific function.

[0058] Figure 1A is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1A As shown in , system 10 includes a source device 12 that generates and encodes video data to be later decoded by a target device 14. Source device 12 and target device 14 may include any of a wide variety of electronic devices, including a cloud server, a server computer, a desktop or laptop computer, a tablet computer, a smart phone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and target device 14 are equipped with wireless communication capabilities.

[0059] In some embodiments, the target device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In one example, the link 16 may include a communication medium that enables the source device 12 to send the encoded video data directly to the target device 14 in real time. The encoded video data may be modulated according to a communication standard (such as a wireless communication protocol) and sent to the target device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a portion of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, a switch, a base station, or any other device that may be useful in facilitating communication from the source device 12 to the target device 14.

[0060] In some other embodiments, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the target device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disk (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device that can store the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing encoded video data and sending the encoded video data to the target device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The target device 14 may access the encoded video data through any standard data connection suitable for accessing the encoded video data stored on the file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both a streaming transmission and a download transmission.

[0061] like Figure 1A As shown in , source device 12 includes video source 18, video encoder 20 and output interface 22. Video source 18 may include sources such as or a combination of such sources: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or a video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.

[0062] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be sent directly to the target device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on a storage device 32 for later access by the target device 14 or other devices for decoding and / or playback. The output interface 22 may also include a modem and / or a transmitter.

[0063] Target device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included in the encoded video data sent over a communication medium, stored on a storage medium, or stored on a file server.

[0064] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the target device 14. The display device 34 displays the decoded video data to the user and may include any of a variety of display devices, such as a Liquid Crystal Display (LCD), a plasma display, an Organic Light Emitting Diode (OLED) display, or another type of display device.

[0065] The video encoder 20 and the video decoder 30 may operate according to a proprietary standard or an industry standard (e.g., VVC, HEVC, MPEG-4, Part 10, AVC) or an extension of such a standard. It should be understood that the present application is not limited to a specific video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally believed that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current standards or future standards. Similarly, it is also generally believed that the video decoder 30 of the target device 14 may be configured to decode video data according to any of these current standards or future standards.

[0066] The video encoder 20 and the video decoder 30 may be implemented as any of a variety of suitable encoder and / or decoder circuits, respectively, such as one or more microprocessors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), discrete logic, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for the software in a suitable non-volatile computer-readable medium, and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and either of the encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0067] In some embodiments, at least some components of the source device 12 (e.g., the video source 18, the video encoder 20 or the components included in the video encoder 20 as described below with reference to FIG. 1G, and the output interface 22) and / or at least some components of the target device 14 (e.g., the input interface 28, the video decoder 30 or the components included in the video decoder 30 as described below with reference to FIG. 1G) are configured to generate a signal. Figure 3The components described herein, and the display device 34) may be run in a cloud computing service network, which may provide software, platforms and / or infrastructure, such as software as a service (SaaS), platform as a service (PaaS) or infrastructure as a service (IaaS). In some embodiments, one or more components in the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be set in one or more client devices, and the one or more client devices may communicate with the server computer in the cloud computing service network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In an embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers in the cloud computing service network, the one or more server computers being implemented by at least some components of the source device 12 and / or at least some components of the target device 14; one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud or a hybrid cloud. The terms "cloud", "cloud computing", "cloud-based", etc. in this document may be used interchangeably without departing from the scope of the present disclosure. It should be understood that the present disclosure is not limited to being implemented in the above-mentioned cloud computing service network. Alternatively, the present disclosure may also be implemented in any other type of computing environment currently known or developed in the future.

[0068] Figure 4A-4E is a schematic diagram illustrating a multi-type tree partitioning mode according to some embodiments of the present disclosure. Figure 4A-4E Five types of partitioning are shown, including four-element partitioning ( Figure 4A ), vertical binary division ( Figure 4B ), horizontal binary partition ( Figure 4C ), vertical ternary division ( Figure 4D ) and horizontal ternary partitioning ( Figure 4E ).

[0069] Figure 2 is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 may perform intra-frame prediction encoding and inter-frame prediction encoding of video blocks within a video frame. Intra-frame prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that the term "frame" may be used as a synonym for the term "image" or "picture" in the field of video coding and decoding.

[0070] like Figure 2 As shown in FIG. 1 , the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 also includes a motion estimation unit 42, a motion compensation unit 44, a partition unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63 (such as a deblocking filter) may be located between the adder 62 and the DPB 64 to filter the block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter, such as a sample adaptive offset (SAO) filter, a cross component sample adaptive offset (CCSAO) filter and / or an adaptive loop filter (ALF) may be used to filter the output of the adder 62. It should be noted that for the CCSAO technology, the present application is not limited to the embodiments described herein, and alternatively, the present application may be applied to the following case: according to any one of the luminance component, the Cb chrominance component and the Cr chrominance component, an offset for any other component of the luminance component, the Cb chrominance component and the Cr chrominance component is selected to modify the any other component based on the selected offset. In addition, it should be noted that the first component mentioned herein may be any one of the luminance component, the Cb chrominance component, and the Cr chrominance component, the second component mentioned herein may be any other component of the luminance component, the Cb chrominance component, and the Cr chrominance component, and the third component mentioned herein may be the remaining one of the luminance component, the Cb chrominance component, and the Cr chrominance component. In some examples, the loop filter may be omitted, and the decoded video block may be provided directly to the DPB 64 by the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be dispersed in one or more of the fixed or programmable hardware units described.

[0071] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from Figure 1AThe video data memory 40 and the DPB 64 are obtained from the video source 18 shown. The DPB 64 is a buffer that stores reference video data (reference frames or pictures) for use by the video encoder 20 (e.g., in intra-frame or inter-frame prediction coding mode) when encoding video data. The video data memory 40 and the DPB 64 may be formed by any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip relative to those components.

[0072] like Figure 2 As shown in , after receiving the video data, the partitioning unit 45 within the prediction processing unit 41 partitions the video data into video blocks. This partitioning operation may also include partitioning the video frame into strips, tiles (e.g., a set of video blocks) or other larger coding units (Coding Unit, CU) according to a predefined partitioning structure associated with the video data (e.g., a quadtree (Quad-Tree, QT) structure). A video frame is or can be considered to be a two-dimensional array or matrix of samples with sample values. The samples in the array may also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into multiple video blocks by (for example) using QT partitioning. A video block is also or can be considered to be a two-dimensional array or matrix of samples with sample values, but the size of the video block is smaller than the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block may be further divided into one or more block partitions or sub-blocks (which may again form blocks) by, for example, iteratively using QT partitioning, binary-tree (BT) partitioning or ternary-tree (TT) partitioning or any combination thereof. It should be noted that the term "block" or "video block" as used herein may be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU) or a transform unit (TU) and / or may be or correspond to a corresponding block, such as a coding tree block (CTB), a coding block (CB), a prediction block (PB) or a transform block (TB) and / or correspond to a sub-block.

[0073] The prediction processing unit 41 may select one of a plurality of feasible prediction coding modes for the current video block based on the error results (e.g., code rate and distortion level), such as one of one or more inter-frame prediction coding modes in a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block (e.g., prediction block) or inter-frame prediction coding block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (such as motion vectors, intra-frame mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.

[0074] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, e.g., to select an appropriate coding mode for each block of video data.

[0075] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating a motion vector according to a predetermined pattern within a sequence of video frames, the motion vector indicating the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating a motion vector that estimates the motion for a video block. For example, a motion vector may indicate the displacement of a video block within a current video frame or picture relative to a prediction block within a reference frame, the prediction block in the reference frame relative to the current block being encoded in the current frame. The predetermined pattern may designate video frames in a sequence as P frames or B frames. Intra BC unit 48 may determine a vector (e.g., a block vector) for intra BC coding in a manner similar to the motion vectors determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vector.

[0076] In terms of pixel differences, the prediction block of the video block may be or may correspond to a block or reference block of a reference frame that closely matches the video block to be encoded, and the pixel difference may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values ​​for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, the motion estimation unit 42 may perform a motion search relative to the full pixel positions and the fractional pixel positions and output a motion vector with fractional pixel precision.

[0077] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing the position of the video block to the position of a prediction block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.

[0078] The motion compensation performed by the motion compensation unit 44 may involve extracting or generating a prediction block based on the motion vector determined by the motion estimation unit 42. After receiving the motion vector for the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from the DPB 64, and forward the prediction block to the adder 50. The adder 50 then forms a residual video block of pixel difference values ​​by subtracting the pixel values ​​of the prediction block provided by the motion compensation unit 44 from the pixel values ​​of the current video block being encoded. The pixel difference values ​​forming the residual video block may include luma component differences or chroma component differences or both. The motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by the video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for identifying the prediction block, any flag indicating a prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.

[0079] In some embodiments, the intra BC unit 48 may generate vectors and extract prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during different encoding passes, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select a suitable intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values ​​for the various tested intra prediction modes using rate-distortion analysis, and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the suitable intra prediction mode to use. The rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was encoded to generate the coded block, as well as the bit rate (i.e., the number of bits) used to generate the coded block. Intra BC unit 48 may calculate ratios from the distortions and rates for the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0080] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identification of the prediction block may include calculating values ​​for sub-integer pixel positions.

[0081] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from a different frame according to inter-frame prediction, video encoder 20 can form pixel difference values ​​by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being encoded, thereby forming a residual video block. The pixel difference values ​​forming the residual video block may include both luma component differences and chroma component differences.

[0082] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 or the intra-frame block copy prediction performed by the intra BC unit 48 as described above, the intra-frame prediction processing unit 46 may perform intra-frame prediction on the current video block. Specifically, the intra-frame prediction processing unit 46 may determine an intra-frame prediction mode for encoding the current block. To this end, the intra-frame prediction processing unit 46 may use various intra-frame prediction modes to encode the current block, for example, during different encoding passes, and the intra-frame prediction processing unit 46 (or in some examples, the mode selection unit) may select a suitable intra-frame prediction mode from the tested intra-frame prediction modes to use. The intra-frame prediction processing unit 46 may provide information indicating the intra-frame prediction mode selected for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra-frame prediction mode into the bitstream.

[0083] After prediction processing unit 41 determines a prediction block for the current video block via inter-prediction or intra-prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0084] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan on the matrix including the quantized transform coefficients. Optionally, entropy encoding unit 56 may perform the scan.

[0085] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to a video bitstream such as Figure 1A The video decoder 30 shown, or archived as Figure 1A The video frame may be stored in storage device 32 as shown for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0086] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for use in generating reference blocks for predicting other video blocks. As noted above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values ​​for use in motion estimation.

[0087] Adder 62 adds the reconstructed residual block to the motion compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block may then be used as a prediction block by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 to inter-predict another video block in a subsequent video frame.

[0088] Figure 3 30 is a block diagram showing another exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform the above-mentioned Figure 2The encoding process is substantially the inverse of the decoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.

[0089] In some examples, units of the video decoder 30 may be tasked to perform embodiments of the present application. In addition, in some examples, embodiments of the present disclosure may be dispersed in one or more of the multiple units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (such as the motion compensation unit 82).

[0090] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (such as a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding the video data (e.g., in an intra-frame or inter-frame prediction decoding mode). The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown in FIG. Figure 3 9 as two different components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30, or off-chip relative to those components.

[0091] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators and other syntax elements to the prediction processing unit 81.

[0092] When a video frame is encoded as an intra-prediction coded (I) frame or for intra-coded prediction blocks in other types of frames, intra-prediction unit 84 of prediction processing unit 81 may generate prediction data for a video block of a current video frame based on a signaled intra-prediction mode and reference data from a previously decoded block of the current frame.

[0093] When the video frame is encoded as an inter-frame prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks may be generated from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists, List 0 and List 1, using a default construction technique based on the reference frames stored in the DPB 92.

[0094] In some examples, when a video block is decoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block as defined by video encoder 20.

[0095] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for the video block of the current video frame by parsing the motion vector and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) for decoding the video block of the video frame, the inter prediction frame type (e.g., B or P), the construction information for one or more of the reference frame lists for the frame, the motion vector for each inter prediction encoded video block of the frame, the inter prediction state for each inter prediction encoded video block of the frame, and other information for decoding the video block in the current video frame.

[0096] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as a flag to determine whether the current video block is predicted using intra BC mode, construction information of which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, a block vector for each intra BC predicted video block of the frame, an intra BC prediction state for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.

[0097] Motion compensation unit 82 may also perform interpolation using interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values ​​for sub-integer pixels of reference blocks. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from the received syntax elements and use these interpolation filters to generate the prediction blocks.

[0098] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.

[0099] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (such as a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device (e.g., Figure 1A on a display device 34).

[0100] In a typical video encoding and decoding process, a video sequence usually includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other cases, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.

[0101] like Figure 1C As shown in , the video encoder 20 (or more specifically, a partitioning unit in a prediction processing unit of the video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs sequentially ordered from left to right and from top to bottom in a raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. It should be noted, however, that the present application is not necessarily limited to a particular size. As Figure 1D As shown in , each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding and decoding samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding and decoding samples of the coding tree block. The coding tree block may be an N×N sample block.

[0102] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding treeblock of the CTU and divide the CTU into smaller CUs. Figure 1B-Figure 1E is a block diagram showing how to recursively divide a frame into a plurality of video blocks of different sizes and shapes according to some embodiments of the present disclosure. Figure 1E As depicted in , a 64×64 CTU 400 is first divided into four smaller CUs, each CU having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are respectively divided into four CUs of a block size of 16×16. Two 16×16 CUs 430 and CU 440 are further divided into four CUs of a block size of 8×8. Figure 1B Depicted is a diagram showing Figure 1EThe quadtree data structure of the final result of the partitioning process of the CTU 400 depicted in FIG. 4 is shown in FIG. 4 , where each leaf node of the quadtree corresponds to a CU of each size ranging from 32×32 to 8×8. Figure 1D Each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements for encoding and decoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding and decoding the samples of the coding block. It should be noted that Figure 1E and Figure 1B The quadtree partitioning depicted in FIG is for illustrative purposes only, and one CTU may be divided into CUs based on quadtree / ternary tree / binary tree partitioning to adapt to varying local characteristics. In a multi-type tree structure, one CTU is divided by a quadtree structure, and each quadtree leaf CU may be further divided by a binary tree structure and a ternary tree structure. Figure 4A-4E As shown, there are five possible partition types for a coding block with a width of W and a height of H, namely, quadruple partition, horizontal binary partition, vertical binary partition, horizontal ternary partition and vertical ternary partition.

[0103] In some embodiments, the video encoder 20 may further divide the coding block of the CU into one or more M×NPBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.

[0104] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0105] After the video encoder 20 generates the predicted luma block, the predicted Cb block, and the predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, so that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, so that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0106] In addition, if Figure 1E As shown in , the video encoder 20 may use quadtree partitioning to decompose the luma residual block, Cb residual block and Cr residual block of the CU into one or more luma transform blocks, Cb transform blocks and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0107] The video encoder 20 may apply one or more transforms to the luma transform block of the TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficient may be a scalar. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0108] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a sequence of bits that form a representation of an encoded frame and associated data, and the bitstream is stored in storage device 32 or sent to target device 14.

[0109] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of the current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding samples of the prediction block for the PU of the current CU to corresponding samples of the transform block of the TU of the current CU. After reconstructing the coding block for each CU of the frame, the video decoder 30 may reconstruct the frame.

[0110] As mentioned above, video codecs mainly use two modes, namely, intra-frame prediction (or intra-frame prediction) and inter-frame prediction (or inter-frame prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to codec efficiency than intra-frame prediction because motion vectors are used to predict the current video block from the reference video block.

[0111] However, with the ever-improving video data capture technology and more refined video block sizes for retaining details in the video data, the amount of data required to represent the motion vector of the current frame has also increased significantly. One way to overcome this challenge is to benefit from the fact that not only a group of neighboring CUs in both the spatial domain and the temporal domain have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by exploring the spatial and temporal correlation of spatially neighboring CUs and / or temporally co-located CUs, which is also referred to as the "motion vector predictor (MVP)) of the current CU.

[0112] Instead of combining as above Figure 2 The actual motion vector of the current CU determined by the motion estimation unit is encoded into the video bitstream, and the motion vector prediction factor of the current CU is subtracted from the actual motion vector of the current CU to generate the motion vector difference (MVD) of the current CU. By doing so, the motion vector determined by the motion estimation unit for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0113] As in the process of selecting a prediction block in a reference frame during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also called a "merge list") of the current CU using those potential candidate motion vectors associated with the spatial neighboring CUs and / or the temporal co-located CUs of the current CU, and then selecting a member from the motion vector candidate list as the motion vector predictor of the current CU. By doing so, the motion vector candidate list itself does not need to be sent from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor in the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU.

[0114] In general, the basic inter prediction scheme applied in VVC remains almost the same as that of HEVC, except that several prediction tools are further extended, added and / or improved, for example, extended merge prediction, MMVD and GPM.

[0115] Extended Consolidated Forecast

[0116] As video data capture technology continues to improve and the size of video blocks used to retain details in video data becomes finer, the amount of data required to represent the motion vector of the current picture has also increased significantly. One way to overcome this challenge is to use the motion information (e.g., motion vector) of the current CU's spatial neighboring CUs, temporal co-located CUs, etc. as an approximation (e.g., prediction) of the current CU's motion information, which is also called the "motion vector predictor (MVP)" of the current CU.

[0117] Similar to the process of selecting a prediction block in a reference picture during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules to construct an MVP candidate list for the current CU, and then select one MVP candidate from the MVP candidate list as the MVP for the current CU. By doing so, there is no need to send the MVP candidate list itself between the video encoder 20 and the video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for the video encoder 20 and the video decoder 30 to encode and decode the current CU using the same MVP candidate selected from the MVP candidate list.

[0118] In VVC, the MVP candidate list is constructed by including the following five types of MVPs in order:

[0119] - Spatial MVP from spatially neighboring CUs (i.e., spatial candidates);

[0120] - Temporal MVP from the temporally co-located CU (i.e., temporal candidate);

[0121] - History-based MVP (HMVP) from a First-In-First-Out (FIFO) table;

[0122] - Average MVP in pairs; and

[0123] -Zero MVP.

[0124] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded in merge mode, the index of the best MVP candidate is encoded using truncated unary binarization. The first binary bit of the index is encoded and decoded with context, and the other binary bits for the index are bypassed.

[0125] The derivation process of each type of MVP is provided as follows. As in HEVC, VVC also supports parallel derivation of MVP candidate lists for all CUs in a region of a certain size.

[0126] Deriving MVP from spatial candidates

[0127] In addition to swapping the positions of the first two spatial candidates, the slave spatial candidates in VVC (e.g. Figure 5 The MVP derived from the CU adjacent to the current CU 101 in HEVC is the same as the MVP derived from the spatial candidate in HEVC. Figure 5Up to four spatial candidates are selected from the spatial candidates at the positions shown (i.e., top position B0, left position A0, upper right position B1, lower left position A1, and upper left position B2). The derivation is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are not available (e.g., because the one or more CUs belong to another slice or tile) or are intra-coded.

[0128] After the CU at position B0 is added as a candidate to the merge candidate list, the remaining candidates added to the merge candidate list are subjected to redundancy check, which ensures that candidates with the same motion information are excluded from the merge candidate list, thereby improving the encoding and decoding efficiency. In order to reduce the computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only the pairs using Figure 6 The pairs linked by the arrowed lines in , and only when the candidate in the corresponding pair for redundancy check does not have the same motion information as the motion information of the candidate to be added, the candidate is added to the merge candidate list. The spatial MVP derived from the candidates in the merge candidate list is added to the MVP candidate list.

[0129] Derive MVP from temporal candidates

[0130] During the MVP derivation from the temporal candidate, only one temporal candidate is added to the merge candidate list. Specifically, when MVP is derived from the temporal candidate, based on the CUs belonging to the current CU (e.g., Figure 7 curr_CU 303) of the same location picture (e.g., Figure 7 col_pic 302 in ) as a temporal candidate co-located CU (eg, Figure 7 The scaled motion vector is derived from the col_CU 301 in the slice header and added to the MVP candidate list as a temporal MVP candidate. The reference picture list and reference picture index to be used to derive the co-located CU are explicitly signaled in the slice header. The scaled motion vector is obtained (i.e., scaled) from the motion vector of the co-located CU using the picture order count (POC) distance (i.e., tb and td), as Figure 7 As shown in , where tb is defined as the current picture (e.g., Figure 7 curr_pic 304) of the reference picture (e.g., Figure 7 The POC difference between curr_ref 305 in td and the current picture, and td is defined as the reference picture of the co-located picture (e.g., Figure 7 The POC difference between col_ref 306 in ( ) and the co-located picture. The reference picture index of the temporal candidate is set equal to zero.

[0131] like Figure 8 As shown, at position C 0 and C 1 The position of the temporal candidate (ie, the co-located CU) in the current CU 401 is selected between. If the position C in the co-located picture 0 The CU at position C is unavailable, intra-coded, or outside the current row of the CTU, then position C 1 The CU at position C is used as the co-located CU for deriving the temporal MVP candidate. 0 The CU at is used as the co-located CU for deriving the temporal MVP candidate.

[0132] Derivation of HMVP candidates

[0133] After spatial MVP and temporal MVP, the HMVP candidate is added to the MVP candidate list. The motion information of the previously coded block is stored in the HMVP table and is used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the HMVP table as a new HMVP candidate.

[0134] The size of the HMVP table can be set to 6. When a new HMVP candidate is inserted into the HMVP table, a constrained FIFO rule is used, where a redundancy check is first applied to find if the same HMVP exists in the HMVP table. If found, the same HMVP is deleted from the HMVP table, and all subsequent HMVP candidates are moved forward, and the same HMVP is added to the last entry of the HMVP table.

[0135] HMVP candidates can be used in the MVP candidate list construction process. The last few HMVP candidates in the HMVP table are checked in order and inserted into the MVP candidate list after the temporal MVP candidate. Redundancy checks are applied to HMVP candidates relative to spatial candidates and / or temporal MVP candidates.

[0136] In order to reduce the number of redundant check operations, the following simplified method is introduced:

[0137] - redundancy checking the last two entries in the HMVP table with respect to the spatial MVP candidates derived from the spatial candidates at positions A1 and B1 respectively; and

[0138] - Once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1, the process of building the MVP candidate list from the HMVP candidates is terminated.

[0139] Derivation of pairwise average MVP candidates

[0140] A pairwise average MVP candidate is generated by averaging the MVPs derived from a predefined pair of the first two merge candidates in an existing merge candidate list. The first merge candidate in the predefined pair may be defined as p0Cand, and the second merge candidate in the predefined pair may be defined as p1Cand. The average motion vector is calculated for the availability of each reference picture list according to the motion vectors of p0Cand and p1Cand, respectively. If both motion vectors are available for one reference picture list, the two motion vectors are averaged even when they point to different reference pictures, and the reference picture of the average motion vector is set to the reference picture of p0Cand; if only one motion vector is available for one reference picture list, the motion vector is used directly; if no motion vector is available for one reference picture list, the motion vector and reference picture index of the reference picture list remain invalid.

[0141] Zero MVP

[0142] When the MVP candidate list is not full after adding the pairwise average MVP candidates, a zero MVP is inserted at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.

[0143] MMVD

[0144] As described above, in merge mode, motion information (i.e., MVP candidates) is implicitly derived from the MVP candidate list constructed for the current CU and is directly used as the MV of the current CU for generating prediction samples of the current CU, which may result in a certain error between the actual MV of the current CU and the implicitly derived MVP. In order to improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, where the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. The MMVD flag is signaled after the regular merge flag is sent to specify whether to use the MVP for the current CU. Pairwise average MVP Candidate .

[0145] In the MMVD mode, after an MVP candidate is selected from the first two MVP candidates in the MVP candidate list, MMVD information is sent using a signal, wherein the MMVD information includes an MMVD candidate flag, a distance index, and a direction index, wherein the MMVD candidate flag is used to specify which of the first two MVP candidates is selected as the MV basis, the distance index is used to indicate the motion amplitude information of the MVD, and the direction index is used to indicate the motion direction information of the MVD.

[0146] The distance index of the motion magnitude information of the specified MVD indicates the distance from the reference picture of the current CU pointed to by the selected MVP candidate (for example, Fig. 9 The starting point in the L0 reference picture 501 or the L1 reference picture 503 in FIG. Fig. 9 The dashed circle in FIG. 1 represents a predefined offset of the distance index (indicated by the dashed circle in FIG. 1 ), and the MVD can be derived from the offset and can be added to the selected MVP candidate. The relationship between the distance index and the predefined offset is specified in Table 1 below.

[0147]

[0148] Table 1

[0149] The direction index specifies the sign of the MVD, which represents the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined symbols. In some examples, the meaning of the sign of the MVD may vary depending on the information of the selected MVP candidate. When the selected MVP candidate is a unidirectional prediction MV or a bidirectional prediction MV in which two MVs point to the same side of the current picture (i.e., the POCs of the two reference pictures of the current picture (e.g., the reference pictures of list 0 and list 1, which are also referred to as L0 reference pictures and L1 reference pictures, respectively) are both greater than the POC of the current picture, or are both less than the POC of the current picture), the symbol in Table 2 specifies the sign of the MVD added to the selected MVP candidate. When the selected MVP candidate is a bidirectional prediction MV in which two MVs point to different sides of the current picture (i.e., a POC of one reference picture of the current picture is greater than the POC of the current picture, and a POC of another reference picture of the current picture is less than the POC of the current picture), if the POC distance of the L0 reference picture (i.e., the POC distance between the L0 reference picture and the current picture) is greater than the POC distance of the L1 reference picture (i.e., the POC distance between the L1 reference picture and the current picture), the sign in Table 2 specifies the sign of the MVD for list 0 MVD0 added to the MVP for list 0 MVP0 in the selected MVP candidate, and the sign of the MVD for list 1 MVD1 added to the MVP for list 1 MVP1 in the selected MVP candidate is opposite to the sign in Table 2; otherwise, if the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, the sign in Table 2 specifies the sign of MVD1 added to MVP1, and the sign of MVD0 added to MVP0 is opposite to the sign in Table 2.

[0150] Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + -

[0151] Table 2

[0152] The MVD is scaled according to the POC distance. If the POC distance for both the L0 reference picture and the L1 reference picture is the same, then the MVD does not need to be scaled. Otherwise, if the POC distance for the L0 reference picture is greater than the POC distance for the L1 reference picture, then MVD1 is scaled. If the POC distance for the L1 reference picture is greater than the POC distance for the L0 reference picture, then MVD0 is scaled.

[0153] GPM

[0154] In VVC, GPM is supported for inter prediction. GPM is signaled as a merge mode using a CU level flag. Other merge modes include normal merge mode, MMVD mode, CIIP mode, and sub-block merge mode. For each possible CU size W×H (W=2 m and H = 2 n , where m,n∈{3,4,5,6}), GPM supports a total of 64 partitions.

[0155] When using GPM, the CU is divided into two parts by a geometrically positioned straight line. The position of the partition line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the CU obtained by geometric partitioning is inter-predicted using its own motion; and only unidirectional prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion compensated predictions are required for each CU, just like conventional bidirectional prediction.

[0156] If GPM is used for the current CU, a geometric partitioning index indicating a partitioning mode of the geometric partitioning (indicating an angle and an offset of the geometric partitioning) and two merge indexes (one merge index for each partition) are further signaled.

[0157] The unidirectional prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process described above. Denote n as the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector (where X equals the parity of n) of the nth merge candidate in the merge candidate list is used as the nth unidirectional prediction motion vector for GPM. These motion vectors are Fig.10 In the case where there is no corresponding LX motion vector of the nth merge candidate in the merge candidate list, the L(1-x) motion vector of the same merge candidate is used as the unidirectional prediction motion vector for GPM.

[0158] CIIP

[0159] In VVC, when a CU is encoded in merge mode, if the CU contains at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if both the width and height of the CU are less than 128 luma samples, an additional flag is signaled to indicate whether the CIIP mode applies to the current CU. In CIIP mode, a prediction signal is obtained by combining an inter prediction signal with an intra prediction signal. The inter prediction signal in CIIP mode is derived using the same inter prediction process as the inter prediction process applied in the conventional merge mode; and the intra prediction signal in CIIP mode is derived after the conventional intra prediction process with the planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weighted average is calculated based on the current CU 1601 (such as Fig.11 The weight value is calculated based on the encoding mode of the top neighboring block and the left neighboring block (shown), as shown below:

[0160] - if the top neighboring block is available and is intra-coded, isIntraTop is set to 1, otherwise isIntraTop is set to 0;

[0161] - if the left neighboring block is available and is intra-coded, isIntraLeft is set to 1, otherwise isIntraLeft is set to 0;

[0162] - If (isIntraLeft+isIntraTop) is equal to 2, the weight value is set to 3;

[0163] - Otherwise, if (isIntraLeft+isIntraTop) is equal to 1, the weight value is set to 2;

[0164] Otherwise, the weight value is set to 1.

[0165] The prediction signal P in CIIP mode is derived as follows CIIP :

[0166] P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (1)

[0167] Where P inter is the inter-frame prediction signal in CIIP mode, P intra is the intra prediction signal in CIIP mode, wt is the weight value, and >> indicates a right shift operation.

[0168] Intra-block copying in Versatile Video Codec (VVC)

[0169] Intra Block Copy (IBC) is a tool adopted in the HEVC extension of SCC. IBC significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the IBC-encoded CU is integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The IBC-encoded CU is regarded as a third prediction mode in addition to the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples.

[0170] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed.

[0171] In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For larger-sized current blocks, the hash key is determined to match the hash key of the reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference block with the smallest cost is selected.

[0172] In the block matching search, the search range is set to cover both the previous CTU and the current CTU.

[0173] At the CU level, the IBC mode is signaled using a flag and can be signaled as either IBC AMVP mode or IBC Skip / Merge mode:

[0174] IBC skip / merge mode: Use the merge candidate index to indicate which block vector from the list of neighboring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and pairwise candidates.

[0175] IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (in case of IBC encoding). When either neighbor is not available, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.

[0176] IBC Reference Area

[0177] In order to reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of predefined areas including the area of ​​the current CTU and a specific area of ​​the left CTU. Fig.12 The reference area of ​​the IBC mode is shown, where each block represents a 64×64 luma sample unit.

[0178] Depending on the location of the current coded CU position within the current CTU, the following operations are applied:

[0179] When the current block falls into the upper left 64×64 block of the current CTU, the current block can use the CPR mode to refer to the reference samples in the lower right 64×64 block of the left CTU in addition to the reconstructed samples in the current CTU. The current block can also use the CPR mode to refer to the reference samples in the lower left 64×64 block of the left CTU and the reference samples in the upper right 64×64 block of the left CTU.

[0180] When the current block falls into the upper right 64×64 block of the current CTU, in addition to referring to the reconstructed samples in the current CTU, if the brightness position (0, 64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower left 64×64 block and the lower right 64×64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64×64 block of the left CTU.

[0181] In the case where the current block falls into the lower left 64×64 block of the current CTU, in addition to referring to the reconstructed samples in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block may also use the CPR mode to refer to the reference samples in the upper right 64×64 block and the lower right 64×64 block of the left CTU. Otherwise, the current block may also use the CPR mode to refer to the reference samples in the lower right 64×64 block of the left CTU.

[0182] In the case where the current block falls into the lower right 64×64 block of the current CTU, the current block may use the CPR mode to refer only to the reconstructed samples in the current CTU.

[0183] This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementation.

[0184] Interaction between IBC and other coding tools

[0185] The interaction between IBC mode and other inter-coding tools in VVC, such as paired merge candidates, history-based motion vector predictor (HMVP), combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM), is as follows:

[0186] IBC can be used with paired merge candidates and HMVP. A new paired IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference.

[0187] IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM.

[0188] When dual-tree (DUAL_TREE) partitioning is used, IBC is not allowed for chroma coding blocks.

[0189] Unlike the HEVC screen content codec extension, the current picture is no longer included as one of the reference pictures in the reference picture list 0 for IBC prediction. The derivation process of motion vectors for IBC mode does not include all neighboring blocks in inter mode, and vice versa. The following IBC design aspects apply:

[0190] IBC shares the same process as regular MV merging, including pairwise merging candidates and history-based motion predictors, but TMVPs and zero vectors are not allowed since they are invalid for IBC mode.

[0191] Separate HMVP buffers (5 candidates each) are used for regular MV and IBC.

[0192] The block vector constraint is implemented as a bitstream consistency constraint, the encoder needs to ensure that there are no invalid vectors in the bitstream, and if the merge candidate is invalid (out of range or 0), then the merge should not be used. This bitstream consistency constraint is expressed in terms of a virtual buffer as described below.

[0193] For deblocking, IBC is handled as an inter mode.

[0194] If the current block is encoded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to only indicate whether the MV is an integer pixel or 4 integer pixels.

[0195] The number of IBC merge candidates may be signaled in the slice header separately from the number of regular merge candidates, sub-block merge candidates, and geometric merge candidates.

[0196] The virtual buffer concept is used to describe the allowable reference area and valid block vectors for the IBC prediction mode. Denoting the CTU size as ctbSize, the width of the virtual buffer ibcBuf is wIbcBuf=128×128 / ctbSize and the height is hIbcBuf=ctbSize. For example, for a CTU size of 128×128, the size of ibcBuf is also 128×128; for a CTU size of 64×64, the size of ibcBuf is 256×64; for a CTU size of 32×32, the size of ibcBuf is 512×32.

[0197] The size of the VPDU is min(ctbSize, 64) in each dimension, Wv = min(ctbSize, 64).

[0198] The virtual IBC buffer ibcBuf is maintained as follows.

[0199] At the start of decoding of each CTU row, the entire ibcBuf is flushed with an invalid value of -1.

[0200] When starting to decode the VPDU (xVPDU, yVPDU) relative to the upper left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU% wIbcBuf, ..., xVPDU% wIbcBuf + Wv-1; y = yVPDU% ctbSize, ..., yVPDU% ctbSize + Wv-1.

[0201] After decoding, the CU contains (x,y) relative to the top left corner of the picture, set:

[0202] ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]

[0203] For a block covering coordinates (x, y), a block vector bv = (bv[0], bv[1]) is valid if the following condition is true for the block vector; otherwise, the block vector is invalid:

[0204] ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1.

[0205] Intra-frame Block Copying in Enhanced Compression Model (ECM)

[0206] In ECM, IBC is improved in the following aspects.

[0207] IBC Merged List / AMVP List Construction

[0208] Modify the IBC merge list / AMVP list construction as follows:

[0209] Only when an IBC merge candidate / AMVP candidate is valid, it can be inserted into the IBC merge candidate list / AMVP candidate list.

[0210] The upper right spatial candidate, the lower left spatial candidate, and the upper left spatial candidate and one pairwise average candidate may be added to the IBC merge candidate list / AMVP candidate list.

[0211] Adaptive reordering based on template (ARMC-TM) is applied to the IBC merge list.

[0212] The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC merge candidates by using full pruning, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.

[0213] Candidates for zero vectors used to populate the IBC merge list / AMVP list are replaced by a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in IBC merge mode, and therefore, they are discarded as BVPs in the IBC candidate list.

[0214] Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Fig.13 shown.

[0215] IBC with Template Matching

[0216] Template matching is used in IBC for both IBC merge mode and IBC AMVP mode.

[0217] Compared to the IBC-TM merge list used by the regular IBC merge mode, the IBC-TM merge list is modified so that candidates are selected according to a pruning method with motion distances between candidates, as in the regular TM merge mode. The ending zero motion realization is replaced by motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU.

[0218] In IBC-TM merge mode, a template matching method is used to refine the selected candidates before RDO or decoding process. IBC-TM merge mode competes with regular IBC merge mode and signals the TM merge flag.

[0219] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of these 3 selected candidates is refined using a template matching method and sorted according to their resulting template matching cost. Only the top 2 candidates are then considered in the motion estimation process as usual.

[0220] Template matching refinement for both IBC-TM merge mode and AMVP mode is very simple because IBC motion vectors are constrained to be (i) integers and (ii) within the reference region, as Fig.12 . Therefore, in IBC-TM merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with integer precision or 4-pixel precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the template used in each refinement step must respect the constraints of the reference region.

[0221] IBC Reference Area

[0222] The reference area for IBC is extended to the upper two CTU rows. Fig.14 The reference region for encoding CTU(m,n) is shown. Specifically, for the CTU(m,n) to be encoded, the reference region includes CTUs with indices (m-2,n-2)...(W,n-2), (0,n-1)...(W,n-1), (0,n)...(m,n), where W represents the maximum horizontal index within the current tile, slice, or picture. This setting ensures that for CTUs of size 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is limited to [-(C<<1), C>>2] in the horizontal direction and [-C, C>>2] in the vertical direction to accommodate reference region expansion, where C represents the CTU size.

[0223] IBC merge mode with block vector difference

[0224] The IBC merge mode with block vector difference is adopted in ECM. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 ​​pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.

[0225] The base candidate is selected from the first five candidates in the reordered IBC merge list. And all possible MBVD refinement positions (20×4) for each base candidate are reordered based on the SAD cost between the template (one row above the current block and one column to the left) and its reference for each refinement position. Finally, the first 8 refinement positions with the lowest template SAD cost are retained as available positions for MBVD index encoding.

[0226] IBC adaptation for camera captured content

[0227] When adapting IBC for camera captured content, the IBC reference range is reduced from 2 CTU rows to 2×128 rows, e.g. Fig.15 As shown in . On the encoder side, to reduce complexity, the local search range is centered on the first block vector predictor of the current CU and is set to [-8,8] in the horizontal direction and [-8,8] in the vertical direction. This encoder modification is not applicable to SCC sequences.

[0228] Combination of CIIP with TIMD and TM

[0229] In CIIP mode, prediction samples are generated by weighting the inter prediction signal using CIIP-TM merge candidate prediction and the intra prediction signal predicted using TIMD derived intra prediction mode. This method is only applied to coding blocks with an area less than or equal to 1024.

[0230] The TIMD derivation method is used to derive intra prediction modes in CIIP. Specifically, the intra prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of the 67 conventional intra prediction modes.

[0231] Furthermore, it is proposed to modify the weights (wIntra, wInter) for both tests in case the derived intra prediction mode is an angular mode. For near-horizontal modes (2 <= angular mode index < 34), Fig.16A The current block is divided vertically as shown in ; for near-vertical mode (34<=angle mode index<=66), as Fig. 16B The current block is divided horizontally as shown.

[0232] Table 3 shows (wIntra, wInter) for different sub-blocks.

[0233] Sub-block index (wIntra,wInter) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6)

[0234] Table 3. Modified weights for angle mode.

[0235] Using CIIP-TM, a CIIP-TM merge candidate list is constructed for the CIIP-TM mode. The merge candidates are refined by template matching. The CIIP-TM merge candidates are also reordered into regular merge candidates by the ARMC method. The maximum number of CIIP-TM merge candidates is equal to two.

[0236] Multiple Hypothesis Prediction (MHP)

[0237] In the multi-hypothesis inter-frame prediction mode, in addition to the traditional bidirectional prediction signal, one or more additional motion compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted superposition on a sample-by-sample basis. bi and the first additional inter-frame prediction signal / hypothesis h 3 , the resulting prediction signal p is obtained as follows 3 :

[0238] p 3 =(1-α)p bi +αh 3 (2)

[0239] According to the mapping presented in Table 4, the weighting factor α is specified by a new syntax element add_hyp_weight_idx:

[0240] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8

[0241] Table 4. Mapping between add_hyp_weight_idx and α

[0242] Similar to the above, more than one additional prediction signal may be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.

[0243] p n+1 =(1-α n+1 ) n +α n+1 h n+1 (3)

[0244] The resulting overall prediction signal is obtained as the final p n (i.e., the p with the largest index n n ). In this mode, at most two additional prediction signals may be used (ie, n is limited to 2).

[0245] The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference, or implicitly by specifying a merge index. A separate multi-hypothesis merge flag differentiates between these two signaling modes.

[0246] For the inter-frame AMVP mode, MHP is applied only when unequal weights in the BCW are selected in the bi-prediction mode.

[0247] A combination of MHP and BDOF is possible; however, BDOF is applied only to the bi-prediction signal part of the prediction signal (i.e., the ordinary first two hypotheses).

[0248] Geometric Partitioning Mode (GPM) in ECM

[0249] GPM with Merged Motion Vector Difference (MMVD)

[0250] The GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. First, a flag is signaled for the GPM CU to specify whether to use this mode. If this mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD. If the MVD is signaled for a geometric partition, after selecting the GPM merge candidate, the motion of the partition is further refined by the signaled MVD information. All other procedures remain the same as GPM.

[0251] The MVD is signaled as a pair of distance and direction, similar to MMVD. In the GPM with MMVD (GPM-MMVD), nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are involved. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, similar to MMVD, the MVD is left-shifted by 2.

[0252] GPM with Template Matching (TM)

[0253] Template matching is applied to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to two geometric partitions. TM is used to refine the motion information for each geometric partition. When TM is selected, the template is constructed using the left neighboring samples, the upper neighboring samples, or the left and upper neighboring samples according to the partitioning angle, as shown in Table 5. Then the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern as the merge mode (where the half-pixel interpolation filter is disabled).

[0254]

[0255] Table 5. Templates for the first and second geometric partitions, where A indicates using the top sample, L indicates using the left sample, and L+A indicates using both the left and top samples.

[0256] The GPM candidate list is constructed as follows:

[0257] 1. Interleaved list-0 MV candidates and list-1 MV candidates are derived directly from the regular merge candidate list, where list-0 MV candidates have higher priority than list-1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.

[0258] 2. Further derive interleaved list-1 MV candidates and list-0 MV candidates directly from the regular merge candidate list, where list-1 MV candidates have higher priority than list-0 MV candidates. The same pruning method with adaptive threshold is also applied to remove redundant MV candidates.

[0259] 3. Fill in zero MV candidates until the GPM candidate list is full.

[0260] GPM-MMVD and GPM-TM are enabled for only one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), a GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is inferred to be false.

[0261] GPM with inter-frame prediction and intra-frame prediction

[0262] Under GPM with inter prediction and intra prediction, the final prediction samples are generated by weighting the inter prediction samples and intra prediction samples for each GPM partition region. The inter prediction samples are derived through the inter GPM, while the intra prediction samples are derived through the intra prediction mode (IPM) candidate list and the index signaled from the encoder. The IPM candidate list size is predefined as 3. The available IPM candidates are parallel angle mode (parallel mode) relative to the GPM block boundary, vertical angle mode (vertical mode) relative to the GPM block boundary, and planar mode, as shown in Figure 2, respectively. Figures 17A-17D In addition, Fig.17DThe GPM with intra prediction and intra prediction shown is restricted to reduce the signaling overhead for IPM and avoid the increase in the size of the intra prediction circuit on the hardware decoder. In addition, direct motion vector and IPM storage on the GPM mixed region is introduced to further improve the encoding and decoding performance.

[0263] In the IPM derivation based on DIMD and neighboring modes, the parallel mode is registered first. Therefore, if the same IPM candidate does not exist in the list, up to two IPM candidates derived from the decoder side intra mode derivation (DIMD) method and / or neighboring block derivation can be registered. For neighboring mode derivation, there are up to five available neighboring block positions, but they are limited by the angle of the GPM block boundary, as shown in Table 6, which have been used for GPM with template matching (GPM-TM).

[0264]

[0265] Table 6. Location of available neighboring blocks for IPM candidate derivation based on the angle of GPM block boundary. A and L represent above and left of the prediction block.

[0266] GPM-Intra can be combined with GPM with motion vector difference merging (GPM-MMVD). TIMD is used for IPM candidates within GPM frames to further improve codec performance. Parallel mode can be registered first, followed by TIMD, DIMD, and IPM candidates of neighboring blocks.

[0267] Template matching based reordering for GPM partitioning pattern

[0268] In the template matching based reordering for GPM partitioning modes, given the motion information of the current GPM block, the respective TM cost values ​​of the GPM partitioning modes are calculated. Then, all GPM partitioning modes are reordered in ascending order based on the TM cost values. Instead of sending the GPM partitioning mode, an index using a Golomb-Rice code is signaled to indicate the position of the exact GPM partitioning mode in the reordering list.

[0269] The reordering method for the GPM partitioning mode is a two-step process performed after generating respective reference templates for the two GPM partitions in the coding unit, as follows:

[0270] Extend the GPM partition edge into the reference templates of the two GPM partitions to obtain 64 reference templates, and calculate the respective TM cost for each of the 64 reference templates;

[0271] Re-sort the GPM partitioning patterns in ascending order based on their TM cost values, and mark the 32 best patterns as available partitioning patterns.

[0272] The edges on the template extend from the edges of the current CU, such as Fig.18 As shown, however, the GPM blending process is not used in the template region across the edge.

[0273] After reordering in ascending order using the TM cost, the index is signaled.

[0274] Currently, the IBC tool is not combined with the GPM tool, but it is straightforward to combine them, which can improve prediction accuracy and improve encoding and decoding performance.

[0275] Currently, coding blocks encoded in IBC mode are not combined with coding blocks encoded in intra mode or inter mode, it is simple to combine them together, which can improve prediction accuracy and improve encoding and decoding performance.

[0276] Currently, the number of block vectors (BV) in the IBC tool is singular, and it is simple to increase the number of block vectors (BV), and the prediction results can be combined, which can improve the prediction accuracy and improve the encoding and decoding performance.

[0277] In the present disclosure, in order to solve the above problems, a method of further improving the existing design of IBC is provided. In general, the main features of the technology proposed in the present disclosure are summarized as follows.

[0278] 1. The IBC tool is combined with the GPM tool. The combination can be GPM with IBC prediction and IBC prediction, GPM with IBC prediction and intra-frame prediction, or GPM with IBC prediction and inter-frame prediction.

[0279] 2. As a simplified version of the combination of the IBC tool and the GPM tool, for a predefined direction (such as 45 degrees), the upper left part is predicted using the intra mode and the lower right part is predicted using the IBC mode, and then they are averagely weighted to obtain the final prediction signal.

[0280] 3. The IBC tool is combined with the CIIP tool, where the IBC prediction is combined with the intra-frame prediction mode, or the IBC prediction is combined with the inter-frame prediction mode.

[0281] 4. The IBC tool is combined with the MHP tool, where more than one BV prediction is obtained and they are weighted averaged to obtain the final prediction signal.

[0282] In some examples, the disclosed methods may be applied independently or in combination.

[0283] GPM with IBC Forecast and IBC Forecast

[0284] According to one or more embodiments of the present disclosure, an IBC tool is combined with a GPM tool in the form of a GPM with IBC prediction and an IBC predicted IBC tool.Different approaches can be used to achieve this goal.

[0285] In the first method, the two "inter" parts of the GPM with inter prediction method and inter prediction method in VVC are replaced by IBC. This means that the two IBC merge prediction results are weighted averaged with each other according to the dividing line in the coding block. The weights can be obtained by referring to the GPM with inter prediction method and inter prediction method in VVC.

[0286] In the second method, the two "inter" parts of the GPM with the inter prediction method and the inter prediction method in the ECM are replaced by IBC, where some template matching tools can be used to further improve the encoding and decoding performance.

[0287] GPM with IBC prediction and intra prediction

[0288] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC prediction and intra prediction.Different approaches can be used to achieve this goal.

[0289] In the first method, the "inter" part of the GPM having the inter prediction method and the intra prediction method in the ECM is replaced by IBC, where the IBC merged prediction result is weighted averaged with the intra prediction result to obtain the final prediction signal.

[0290] GPM with IBC prediction and inter prediction

[0291] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a GPM with IBC prediction and inter prediction.Different approaches can be used to achieve this goal.

[0292] In the first method, an "inter" part of the GPM with an inter prediction method and an inter prediction method in VVC is replaced by IBC, where the IBC merged prediction result is weighted averaged with the inter merged prediction result to obtain the final prediction signal.

[0293] In the second method, an "inter" part of the GPM with an inter prediction method and an inter prediction method in the ECM is replaced by IBC, where some template matching tools can be used to further improve the coding performance.

[0294] Simplified IBC prediction and intra prediction combination in GPM form

[0295] According to one or more embodiments of the present disclosure, the IBC tool is combined with the GPM tool in the form of a simplified GPM with IBC prediction and intra prediction, such as combining IBC prediction and intra prediction in a specific partition mode, which can save the bit overhead of the partition mode representation. Different methods can be used to achieve this goal.

[0296] In the first method, for a partition line, such as 45 degrees, the upper left part of the coding block is encoded in intra prediction mode, and the lower right part of the coding block is encoded in IBC prediction mode, and then they are averaged in GPM form to obtain the final prediction signal.

[0297] Combined IBC - Intra / Inter Prediction

[0298] According to one or more embodiments of the present disclosure, a coding block encoded in the IBC mode and a coding block encoded in the intra mode or the inter mode are combined. Different methods can be used to achieve this goal.

[0299] In the first method, the encoder / decoder may combine the coding blocks encoded in the IBC mode and the coding blocks encoded in the intra-frame mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, the coding blocks encoded in the IBC merge mode are regarded as coding blocks encoded in the inter-frame merge mode, and the coding blocks encoded in the IBC merge mode and the coding blocks encoded in the planar intra-frame prediction mode are combined. In another example, similar to the combination of CIIP in ECM with TIMD and TM merge technology, the coding blocks encoded in the IBC merge-TM mode and the coding blocks encoded in the intra-frame prediction mode derived from TIMD are combined.

[0300] In the second method, the encoder / decoder may combine the coding blocks encoded in the IBC mode and the coding blocks encoded in the inter-frame mode. Various methods can be used in this combination. In one example, similar to the CIIP technology in VVC, the coding blocks encoded in the IBC merge mode are regarded as coding blocks encoded in the planar intra-frame mode, and the coding blocks encoded in the IBC merge mode and the coding blocks encoded in the inter-frame merge mode are combined. In another example, the coding blocks encoded in the IBC merge mode are regarded as coding blocks encoded in the inter-frame merge mode, and the coding blocks encoded in the IBC merge mode and the coding blocks encoded in the inter-frame merge mode are combined by equal averaging.

[0301] In the third method, the encoder / decoder may combine the coding blocks encoded in the IBC mode, the coding blocks encoded in the intra-frame mode, and the coding blocks encoded in the inter-frame mode. Various methods can be used in this combination. In one example, the coding blocks encoded in the IBC mode, the coding blocks encoded in the intra-frame mode, and the coding blocks encoded in the inter-frame mode are directly combined by averaging equally. In another example, first, the coding blocks encoded in the IBC mode are separately combined with the coding blocks encoded in the intra-frame mode and the coding blocks encoded in the inter-frame mode, as presented in the first method and the second method. Then, the separate combination results are combined by averaging equally.

[0302] Multi-hypothesis IBC prediction

[0303] According to one or more embodiments of the present disclosure, the number of block vectors (BVs) in the IBC tool is increased to 2 or more, and 2 or more hypotheses are combined to obtain the final prediction result. Different methods can be used to achieve this goal.

[0304] In a first approach, the encoder / decoder may combine two hypotheses corresponding to two BVs to obtain a final prediction result. Various methods may be used to achieve this goal. In one example, two BVs corresponding to the minimum rate distortion metric in the IBC AMVP mode and the second minimum rate distortion metric are averaged equally to obtain a final prediction result. In another example, the prediction result corresponding to the IBC AMVP mode and the prediction result corresponding to the IBC merge mode are averaged equally to obtain a final prediction result.

[0305] In the second method, the encoder / decoder can combine more hypotheses corresponding to more BVs to obtain the final prediction result. Various methods can be used to achieve this goal. In one example, the iterative accumulation method proposed in the multi-hypothesis prediction (MHP) technology is used to obtain the final prediction result. In another example, all BVs corresponding to the minimum rate distortion metric, the second minimum rate distortion metric, the third minimum rate distortion metric, ..., in the IBC AMVP mode are averaged equally to obtain the final prediction result.

[0306] Fig.19 A computing environment (or computing device) 1610 coupled to a user interface 1650 is shown. The computing environment 1610 may be part of a data processing server. In some embodiments, the computing device 1610 may perform any of the various methods or processes (such as encoding / decoding methods or processes) described above according to various examples of the present disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.

[0307] The processor 1620 generally controls the overall operation of the computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors to execute instructions to perform all or some steps in the above method. In addition, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0308] The memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1630 may include predetermined software 1632. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, etc. The memory 1630 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0309] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules such as a keyboard, click wheel, buttons, etc. Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 may be coupled to an encoder and a decoder.

[0310] Fig. 20 is a flowchart illustrating a method for video decoding according to an example of the present disclosure. Fig. 20 Methods are shown for GPM with IBC prediction and IBC prediction, GPM with IBC prediction and intra prediction, GPM with IBC prediction and inter prediction, and / or a simplified combination of IBC prediction and intra prediction in the form of GPM.

[0311] In step 2001 , at the decoder side, the processor 1620 may obtain a current CU encoded based on the IBC mode combined with the GPM.

[0312] In step 2002, the processor 1620 may obtain a prediction for the current CU based on the IBC mode combined with the GPM.

[0313] In some examples, the current CU is divided into a first IBC prediction portion and a second IBC prediction portion.

[0314] In one or more examples, following the GPM with an inter prediction method and an inter prediction method in VVC, the current CU is divided into two IBC prediction parts. The two "inter" parts of the GPM with an inter prediction method and an inter prediction method in VVC are replaced with IBC. This means that the two IBC merge prediction results are weighted averaged with each other according to the dividing line in the coding block. The weights can be obtained by referring to the GPM with an inter prediction method and an inter prediction method in VVC. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second IBC merge prediction for the second IBC prediction part, and obtain a prediction for the current CU based on the first IBC merge prediction and the second IBC merge prediction. The prediction for the current CU can be obtained by weighted averaging the first IBC merge prediction and the second IBC merge prediction.

[0315] In one or more examples, the current CU is divided into two IBC prediction parts following the GPM with an inter prediction method and an inter prediction method in the ECM. The two "inter" parts of the GPM with an inter prediction method and an inter prediction method in the ECM are replaced with IBC, where some template matching tools can be used to further improve the encoding and decoding performance. For example, the processor 1620 can obtain a first IBC-TM prediction for the first IBC prediction part and a second IBC-TM prediction for the second IBC prediction part, and obtain a prediction for the current CU based on the first IBC-TM prediction and the second IBC-TM prediction.

[0316] In some examples, the current CU is divided into a first IBC prediction part and a second intra prediction part. In these examples, the "inter" part of the GPM having an inter prediction method and an intra prediction method in the ECM is replaced with IBC, where the IBC merge prediction result and the intra prediction result are weighted averaged to obtain a final prediction signal. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second intra prediction for the second intra prediction part, and obtain a prediction for the current CU by weighted averaging the first IBC merge prediction and the second intra prediction.

[0317] In some examples, the current CU is divided into a first IBC prediction part and a second inter prediction part. In these examples, an "inter" part of a GPM having an inter prediction method and an inter prediction method in a VVC or ECM is replaced with IBC. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second inter prediction for the second inter prediction part, and obtain a prediction for the current CU by weighted averaging the first IBC merge prediction and the second inter prediction. For another example, the processor 1620 may obtain a first IBC-TM prediction for the first IBC prediction part and a second inter prediction for the second inter prediction part, and obtain a prediction for the current CU based on the first IBC-TM prediction and the second inter prediction.

[0318] In some examples, the IBC tool can be combined with the GPM tool in the form of a simplified GPM with IBC prediction and intra-frame prediction. For example, the current CU is divided into a first part and a second part based on a predefined direction, corresponding to a specific partitioning mode used in the GPM. In these examples, the processor 1620 may obtain a first IBC prediction for the first part and a second intra-frame prediction for the second part, and obtain a prediction for the current CU by averaging the first IBC prediction and the second intra-frame prediction. In some examples, the predefined direction may be a dividing line at 45 degrees, the upper left part of the coding block is encoded in the intra-frame prediction mode, and the lower right part of the coding block is encoded in the IBC prediction mode, and then they are averaged in the form of GPM to obtain the final prediction signal.

[0319] Fig.21 is shown as Fig. 20 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding.

[0320] In step 2101 , at the encoder side, the processor 1620 may encode the current CU based on the IBC mode combined with the GPM.

[0321] In step 2102 , the processor 1620 may send the current CU encoded based on the IBC mode combined with the GPM to the decoder.

[0322] In some examples, the current CU is divided into a first IBC prediction portion and a second IBC prediction portion.

[0323] In one or more examples, the current CU is divided into two IBC prediction parts according to the GPM with the inter prediction method and the inter prediction method in VVC. The two "inter" parts of the GPM with the inter prediction method and the inter prediction method in VVC are replaced with IBC. This means that the two IBC merge prediction results are weighted averaged with each other according to the dividing line in the coding block. The weights can be obtained by referring to the GPM with the inter prediction method and the inter prediction method in VVC. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second IBC merge prediction for the second IBC prediction part, and obtain a prediction for the current CU based on the first IBC merge prediction and the second IBC merge prediction. The prediction for the current CU can be obtained by weighted averaging the first IBC merge prediction and the second IBC merge prediction.

[0324] In one or more examples, the current CU is divided into two IBC prediction parts according to the GPM with the inter prediction method and the inter prediction method in the ECM. The two "inter" parts of the GPM with the inter prediction method and the inter prediction method in the ECM are replaced with IBC, where some template matching tools can be used to further improve the encoding and decoding performance. For example, the processor 1620 can obtain a first IBC-TM prediction for the first IBC prediction part and a second IBC-TM prediction for the second IBC prediction part, and obtain a prediction for the current CU based on the first IBC-TM prediction and the second IBC-TM prediction.

[0325] In some examples, the current CU is divided into a first IBC prediction part and a second intra prediction part. In these examples, the "inter" part of the GPM having an inter prediction method and an intra prediction method in the ECM is replaced with IBC, where the IBC merge prediction result is weighted averaged with the intra prediction result to obtain a final prediction signal. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second intra prediction for the second intra prediction part, and obtain a prediction for the current CU by weighted averaging the first IBC merge prediction and the second intra prediction.

[0326] In some examples, the current CU is divided into a first IBC prediction part and a second inter prediction part. In these examples, an "inter" part of a GPM having an inter prediction method and an inter prediction method in a VVC or ECM is replaced with IBC. For example, the processor 1620 may obtain a first IBC merge prediction for the first IBC prediction part and a second inter prediction for the second inter prediction part, and obtain a prediction for the current CU by weighted averaging the first IBC merge prediction and the second inter prediction. For another example, the processor 1620 may obtain a first IBC-TM prediction for the first IBC prediction part and a second inter prediction for the second inter prediction part, and obtain a prediction for the current CU based on the first IBC-TM prediction and the second inter prediction.

[0327] In some examples, the IBC tool can be combined with the GPM tool in the form of a simplified GPM with IBC prediction and intra-frame prediction. For example, the current CU is divided into a first part and a second part based on a predefined direction, corresponding to a specific partitioning pattern used in the GPM. In these examples, the processor 1620 may obtain a first IBC prediction for the first part and a second intra-frame prediction for the second part, and obtain a prediction for the current CU by averaging the first IBC prediction and the second intra-frame prediction. In some examples, the predefined direction may be a dividing line at 45 degrees, the upper left part of the coding block is encoded in the intra-frame prediction mode, and the lower right part of the coding block is encoded in the IBC prediction mode, and then they are averaged in the form of GPM to obtain the final prediction signal.

[0328] Fig. 22 is a flowchart illustrating a method for video decoding according to an example of the present disclosure. Fig. 22 Combined IBC - intra / inter prediction is shown.

[0329] In step 2201 , at the decoder side, the processor 1620 may obtain a first prediction for a current CU, wherein the first prediction is associated with an intra block copy (IBC) mode.

[0330] In step 2202 , the processor 1620 may obtain a second prediction for the current CU, where the second prediction is associated with one of an intra mode or an inter mode.

[0331] In step 2203 , the processor 1620 may obtain a final prediction for the current CU based on the first prediction and the second prediction.

[0332] In some examples, similar to the CIIP technique in VVC, the first prediction is associated with the IBC merge mode and the second prediction is obtained based on the planar intra prediction mode.

[0333] In some examples, similar to the combination of CIIP in ECM with TIMD and TM merge techniques, the first prediction is obtained based on the IBC merge-TM mode, and the second prediction is obtained based on the intra prediction mode derived from template-based intra mode derivation (TIMD).

[0334] In some examples, similar to the CIIP technique in VVC, the first prediction is associated with the IBC merge mode and the second prediction is obtained based on the inter merge mode, and the final prediction is obtained by averaging the first prediction and the second prediction equally.

[0335] In some examples, processor 1620 may combine a coding block encoded in IBC mode with a coding block encoded in intra mode and a coding block encoded in inter mode.

[0336] For example, the processor 1620 may also obtain a third prediction for the current CU, where the third prediction is associated with another mode of the intra mode or the inter mode, and obtain a final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.

[0337] For another example, the processor 1620 may obtain a first intermediate prediction based on the first prediction and the second prediction, and obtain a second intermediate prediction based on the first prediction and the third prediction, and obtain a final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.

[0338] Fig.23 is shown as Fig. 22 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding.

[0339] In step 2301 , at the encoder side, the processor 1620 may obtain a first prediction for a current CU, wherein the first prediction is associated with an intra block copy (IBC) mode.

[0340] In step 2302 , the processor 1620 may obtain a second prediction for the current CU, where the second prediction is associated with one of an intra mode or an inter mode.

[0341] In step 2303 , the processor 1620 may obtain a final prediction for the current CU based on the first prediction and the second prediction.

[0342] In some examples, similar to the CIIP technique in VVC, the first prediction is associated with the IBC merge mode and the second prediction is obtained based on the planar intra prediction mode.

[0343] In some examples, similar to the combination of CIIP in ECM with TIMD and TM merge techniques, the first prediction is obtained based on the IBC merge-TM mode, and the second prediction is obtained based on the intra prediction mode derived from template-based intra mode derivation (TIMD).

[0344] In some examples, similar to the CIIP technique in VVC, the first prediction is associated with the IBC merge mode and the second prediction is obtained based on the inter merge mode, and the final prediction is obtained by averaging the first prediction and the second prediction equally.

[0345] In some examples, processor 1620 may combine a coding block encoded in IBC mode with a coding block encoded in intra mode and a coding block encoded in inter mode.

[0346] For example, the processor 1620 may also obtain a third prediction for the current CU, where the third prediction is associated with another mode of the intra mode or the inter mode, and obtain a final prediction for the current CU by equally averaging the first prediction, the second prediction, and the third prediction.

[0347] For another example, the processor 1620 may obtain a first intermediate prediction based on the first prediction and the second prediction, and obtain a second intermediate prediction based on the first prediction and the third prediction, and obtain a final prediction for the current CU by equally averaging the first intermediate prediction and the second intermediate prediction.

[0348] Fig.24 is a flowchart illustrating a method for video decoding according to an example of the present disclosure. Fig.24 Multi-hypothesis IBC predictions are shown.

[0349] In step 2401 , at the decoder side, the processor 1620 may obtain a plurality of block vectors for a current CU based on the IBC mode.

[0350] In step 2402 , at the decoder side, the processor 1620 may obtain a final prediction for the current CU based on a plurality of block vectors.

[0351] In some examples, the plurality of block vectors may include two block vectors, namely, a first block vector and a second block vector. Processor 1620 may obtain the first block vector based on a minimum rate-distortion metric in the IBC AMVP mode, and obtain the second block vector based on a second minimum rate-distortion metric in the IBC AMVP mode, and obtain a final prediction for the current CU by equally averaging the prediction results of the first block vector and the second block vector.

[0352] In some examples, the plurality of block vectors may include two block vectors, namely, a first block vector and a second block vector. Processor 1620 may obtain the first block vector based on the minimum rate-distortion metric in the IBC AMVP mode, and obtain the second block vector based on the minimum rate-distortion metric in the IBC merge mode, and obtain a final prediction for the current CU by equally averaging the prediction result of the first block vector and the prediction result of the second block vector.

[0353] In some examples, processor 1620 may obtain multiple block vectors based on the distortion metric in the IBC AMVP mode, and obtain a final prediction for the current CU by equally averaging prediction results corresponding to the multiple block vectors.

[0354] In some examples, the processor 1620 may obtain a plurality of block vectors. An iterative accumulation method in a multiple hypothesis prediction (MHP) technique may be used to obtain a final prediction result.

[0355] Fig.25 is shown as Fig.24 The method for video decoding shown in FIG. 1 corresponds to a flowchart of a method for video encoding.

[0356] In step 2501 , at the encoder side, the processor 1620 may obtain a plurality of block vectors for a current CU based on the IBC mode.

[0357] In step 2502 , at the encoder side, the processor 1620 may obtain a final prediction for the current CU based on a plurality of block vectors.

[0358] In some examples, the plurality of block vectors may include two block vectors, namely, a first block vector and a second block vector. Processor 1620 may obtain the first block vector based on a minimum rate-distortion metric in the IBC AMVP mode, and obtain the second block vector based on a second minimum rate-distortion metric in the IBC AMVP mode, and obtain a final prediction for the current CU by equally averaging the prediction results of the first block vector and the second block vector.

[0359] In some examples, the plurality of block vectors may include two block vectors, namely, a first block vector and a second block vector. Processor 1620 may obtain the first block vector based on the minimum rate-distortion metric in the IBC AMVP mode, and obtain the second block vector based on the minimum rate-distortion metric in the IBC merge mode, and obtain a final prediction for the current CU by equally averaging the prediction result of the first block vector and the prediction result of the second block vector.

[0360] In some examples, processor 1620 may obtain multiple block vectors based on the distortion metric in the IBC AMVP mode, and obtain a final prediction for the current CU by equally averaging prediction results corresponding to the multiple block vectors.

[0361] In some examples, the processor 1620 may obtain a plurality of block vectors. An iterative accumulation method in a multiple hypothesis prediction (MHP) technique may be used to obtain a final prediction result.

[0362] In some examples, a device for video encoding and decoding is provided. The device includes a processor 1620 and a memory 1640 configured to store instructions executable by the processor; wherein the processor is configured to perform the following when executing the instructions: Figure 20-Figure 25 Any of the methods shown in .

[0363] In an embodiment, a non-transitory computer-readable storage medium including a plurality of programs is also provided, wherein the plurality of programs are, for example, in the memory 1630 and can be executed by the processor 1620 in the computing environment 1610 to perform the above method and / or store a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method. In one example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to receive (for example, from Figure 2 The video encoder 20 in the computing environment 1610 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned decoding method according to the received bitstream or data stream. In another example, multiple programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned encoding method, encode the video information (e.g., video blocks representing video frames and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to send a bitstream or data stream (e.g., to a Figure 3 Optionally, a non-transitory computer-readable storage medium may store therein a bit stream or a data stream, the bit stream or the data stream comprising a video decoder 30 in the encoder (e.g., Figure 2 The video encoder 20 in FIG. 1 generates a video signal for a decoder (eg, Figure 3The video decoder 30 in the video decoder 30) in decoding the video data (e.g., video blocks representing encoded video frames and / or associated one or more syntax elements, etc.). The non-transitory computer-readable storage medium can be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0364] In an embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In an embodiment, a bitstream is provided, which includes the encoded video information generated by the above encoding method or the encoded video information to be decoded by the above decoding method.

[0365] In an embodiment, a computing device is also provided, which includes one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 having stored therein a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the above-mentioned method when executing the plurality of programs.

[0366] In an embodiment, a computer program product having instructions is also provided, wherein the instructions are used to store or send a bitstream, the bitstream including the encoded video information generated by the above encoding method or the encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product including multiple programs is also provided, wherein the multiple programs are, for example, in the memory 1630 and can be executed by the processor 1620 in the computing environment 1610 to perform the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0367] In an embodiment, the computing environment 1610 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0368] In an embodiment, a method for storing a bitstream is also provided, comprising storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.

[0369] In an embodiment, a method for transmitting a bit stream generated by the above encoder is also provided. In an embodiment, a method for receiving a bit stream to be decoded by the above decoder is also provided.

[0370] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.

[0371] Unless otherwise specifically stated, the order of the method steps according to the present disclosure is intended to be illustrative only, and the method steps according to the present disclosure are not limited to the order specifically described above, but may be changed according to actual conditions. In addition, at least one of the method steps according to the present disclosure may be adjusted, combined or deleted according to actual requirements.

[0372] The examples are chosen and described in order to explain the principles of the present disclosure and to enable other persons skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications suitable for the intended specific use. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the present disclosure.

Claims

1. A method for video decoding, include: A current coding unit CU encoded based on the intra block copy IBC mode combined with the geometric partitioning mode GPM is obtained by the decoder; as well as A prediction for the current CU is obtained by the decoder based on the IBC mode combined with the GPM.

2. The method according to claim 1, wherein the current CU is divided into a first IBC prediction part and a second IBC prediction part, and The method also include: obtaining, by the decoder, a first IBC merged prediction for the first IBC prediction portion; obtaining, by the decoder, a second IBC merged prediction for the second IBC prediction portion; as well as The prediction for the current CU is obtained by the decoder based on the first IBC merge prediction and the second IBC merge prediction.

3. The method according to claim 2, further comprising: include: The prediction for the current CU is obtained by the decoder by weighted averaging the first IBC merge prediction and the second IBC merge prediction.

4. The method according to claim 1, wherein the current CU is divided into a first IBC prediction part and a second IBC prediction part, and The method also include: obtaining, by the decoder, a first IBC-Template Matching (TM) prediction for the first IBC prediction portion; obtaining, by the decoder, a second IBC-TM prediction for the second IBC prediction portion; as well as The prediction for the current CU is obtained by the decoder based on the first IBC-TM prediction and the second IBC-TM prediction.

5. The method according to claim 1, wherein the current CU is divided into a first IBC prediction part and a second intra prediction part, and The method also include: obtaining, by the decoder, a first IBC merged prediction for the first IBC prediction portion; obtaining, by the decoder, a second intra prediction for the second intra prediction portion; as well as The prediction for the current CU is obtained by the decoder by weighted averaging the first IBC merge prediction and the second intra prediction.

6. The method according to claim 1, wherein the current CU is divided into a first IBC prediction part and a second inter-frame prediction part, and The method also include: obtaining, by the decoder, a first IBC merged prediction for the first IBC prediction portion; obtaining, by the decoder, a second inter-frame prediction for the second inter-frame prediction portion; as well as The prediction for the current CU is obtained by the decoder by weighted averaging the first IBC merge prediction and the second inter-frame prediction.

7. The method according to claim 1, wherein the current CU is divided into a first IBC prediction part and a second inter-frame prediction part, and The method also include: Obtaining, by the decoder, a first IBC-Template Matching (TM) prediction for the first IBC prediction portion; obtaining, by the decoder, a second inter-frame prediction for the second inter-frame prediction portion; as well as The prediction for the current CU is obtained by the decoder based on the first IBC-TM prediction and the second inter-frame prediction.

8. The method according to claim 1, wherein the current CU is divided into a first part and a second part based on a predefined direction, wherein the method further comprises: obtaining, by the decoder, a first IBC prediction for the first part; obtaining, by the decoder, a second intra-frame prediction for the second part; and obtaining, by the decoder, the prediction for the current CU by averaging the first IBC prediction and the second intra-frame prediction.

9. The method according to claim 8, wherein the predefined direction is 45 degrees, the first IBC prediction is located in the lower right part of the current CU and the second intra-frame prediction is located in the upper left part of the current CU.

10. A method for video coding, comprising: encoding, by an encoder, a current coding unit CU based on an intra-block copy IBC mode combined with a geometric partitioning mode GPM; and sending, by the encoder, the current CU encoded based on the IBC mode combined with GPM to a decoder.

11. The method according to claim 10, wherein the current CU is divided into a first IBC prediction part and a second IBC prediction part, and wherein the method further comprises: obtaining, by the encoder, a first IBC merge prediction for the first IBC prediction part; obtaining, by the encoder, a second IBC merge prediction for the second IBC prediction part; and obtaining, by the encoder, the prediction for the current CU based on the first IBC merge prediction and the second IBC merge prediction.

12. The method according to claim 11, further comprises: obtaining, by the encoder, the prediction for the current CU by performing weighted averaging on the first IBC merge prediction and the second IBC merge prediction.

13. The method according to claim 10, wherein the current CU is divided into a first IBC prediction part and a second IBC prediction part, and wherein the method further comprises: obtaining, by the encoder, a first IBC-template matching TM prediction for the first IBC prediction part; obtaining, by the encoder, a second IBC-TM prediction for the second IBC prediction part; and obtaining, by the encoder, the prediction for the current CU based on the first IBC-TM prediction and the second IBC-TM prediction.

14. The method according to claim 10, wherein the current CU is divided into a first IBC prediction part and a second intra-frame prediction part, and wherein the method further comprises: obtaining, by the encoder, a first IBC merge prediction for the first IBC prediction part; obtaining, by the encoder, a second intra-frame prediction for the second intra-frame prediction part; and obtaining, by the encoder, the prediction for the current CU by performing weighted averaging on the first IBC merge prediction and the second intra-frame prediction.

15. The method according to claim 10, wherein the current CU is divided into a first IBC prediction part and a second inter prediction part, and The method also include: obtaining, by the encoder, a first IBC merged prediction for the first IBC prediction portion; obtaining, by the encoder, a second inter-frame prediction for the second inter-frame prediction portion; as well as The encoder obtains a prediction for the current CU by weighted averaging the first IBC merge prediction and the second inter-frame prediction.

16. The method according to claim 10, wherein the current CU is divided into a first IBC prediction part and a second inter-frame prediction part, and The method also include: Obtaining, by the encoder, a first IBC-Template Matching (TM) prediction for the first IBC prediction portion; obtaining, by the encoder, a second inter-frame prediction for the second inter-frame prediction portion; as well as A prediction for the current CU is obtained by the encoder based on the first IBC-TM prediction and the second inter prediction.

17. The method according to claim 10, wherein the current CU is divided into a first part and a second part based on a predefined direction, The method also include: obtaining, by the encoder, a first IBC prediction for the first portion; The encoder obtains a second intra prediction for the second portion; as well as The encoder obtains a prediction for the current CU by weighted averaging the first IBC prediction and the second intra prediction.

18. The method according to claim 17, Wherein the predefined direction is 45 degrees, the first IBC prediction is located in a lower right portion of the current CU and the second intra prediction is located in an upper left portion of the current CU.

19. A method for video decoding, include: Obtaining, by a decoder, a first prediction for a current coding unit CU, wherein the first prediction is associated with an intra block copy (IBC) mode; Obtaining, by the decoder, a second prediction for the current CU, wherein the second prediction is associated with one of an intra mode or an inter mode; as well as A final prediction for the current CU is obtained by the decoder based on the first prediction and the second prediction.

20. The method of claim 19, wherein the first prediction is associated with an IBC merge mode, and The method also include: The second prediction is obtained by the decoder based on a planar intra prediction mode.

21. The method according to claim 19, further comprising: include: Obtaining, by the decoder, the first prediction based on an IBC Merge-Template Matching TM mode; as well as The second prediction is obtained by the decoder from a template-based intra-mode derived TIMD-derived intra-prediction mode.

22. The method of claim 19, wherein the first prediction is associated with an IBC merge mode, and The method also include: The second prediction is obtained by the decoder based on an inter-frame merge mode.

23. The method according to claim 22, further comprising: include: The final prediction for the current CU is obtained by the decoder by equally averaging the first prediction and the second prediction.

24. The method according to claim 19, further comprising: include: Obtaining, by the decoder, a third prediction for the current CU, wherein the third prediction is associated with another mode of an intra mode or an inter mode; as well as The final prediction for the current CU is obtained by the decoder based on the first prediction, the second prediction, and the third prediction.

25. The method according to claim 24, further comprising: include: The final prediction for the current CU is obtained by the decoder by equally averaging the first prediction, the second prediction and the third prediction.

26. The method according to claim 24, further comprising: include: Obtaining, by the decoder, a first intermediate prediction based on the first prediction and the second prediction; Obtaining, by the decoder, a second intermediate prediction based on the first prediction and the third prediction; as well as The final prediction for the current CU is obtained by the decoder by equally averaging the first intermediate prediction and the second intermediate prediction.

27. A method for video encoding, include: Obtaining, by an encoder, a first prediction for a current coding unit CU, wherein the first prediction is associated with an intra block copy (IBC) mode; Obtaining, by the encoder, a second prediction for the current CU, wherein the second prediction is associated with one of an intra mode or an inter mode; as well as A final prediction for the current CU is obtained by the encoder based on the first prediction and the second prediction.

28. The method of claim 27, wherein the first prediction is associated with an IBC merge mode, and The method also include: The second prediction is obtained by the encoder based on a planar intra prediction mode.

29. The method according to claim 27, further comprising: include: Obtaining, by the encoder, the first prediction based on an IBC Merge-Template Matching (TM) mode; as well as The second prediction is obtained by the encoder from a template-based intra-mode derived TIMD-derived intra-prediction mode.

30. The method of claim 27, wherein the first prediction is associated with an IBC merge mode, and The method also include: The second prediction is obtained by the encoder based on an inter-frame merging mode.

31. The method according to claim 30, further comprising: include: The final prediction for the current CU is obtained by the encoder by equally averaging the first prediction and the second prediction.

32. The method according to claim 27, further comprising: include: Obtaining, by the encoder, a third prediction for the current CU, wherein the third prediction is associated with another mode of the intra mode or the inter mode; as well as The final prediction for the current CU is obtained by the encoder based on the first prediction, the second prediction, and the third prediction.

33. The method according to claim 32, further comprising: include: The final prediction for the current CU is obtained by the encoder by equally averaging the first prediction, the second prediction and the third prediction.

34. The method according to claim 32, further comprising: include: Obtaining, by the encoder, a first intermediate prediction based on the first prediction and the second prediction; Obtaining, by the encoder, a second intermediate prediction based on the first prediction and the third prediction; as well as The final prediction for the current CU is obtained by the encoder by equally averaging the first intermediate prediction and the second intermediate prediction.

35. A method for video decoding, include: The decoder obtains multiple block vectors for the current coding unit CU based on the intra block copy IBC mode; as well as A final prediction for the current CU is obtained by the decoder based on the multiple block vectors.

36. The method of claim 35, wherein the plurality of block vectors comprises a first block vector and a second block vector, and The method also include: The first block vector is obtained by the decoder based on the minimum rate distortion metric in the IBC advanced motion vector prediction AMVP mode; Obtaining, by the decoder, the second block vector based on a second minimum rate-distortion metric in the IBC AMVP mode; as well as The final prediction for the current CU is obtained by the decoder by equally averaging the prediction result of the first block vector and the prediction result of the second block vector.

37. The method of claim 35, wherein the plurality of block vectors comprises a first block vector and a second block vector, and The method also include: The first block vector is obtained by the decoder based on the minimum rate distortion metric in the IBC advanced motion vector prediction AMVP mode; Obtaining, by the decoder, the second block vector based on a minimum rate-distortion metric in an IBC merge mode; And the final prediction for the current CU is obtained by the decoder by equally averaging the prediction result of the first block vector and the prediction result of the second block vector.

38. The method according to claim 35, further comprising: include: The decoder obtains the plurality of block vectors based on a distortion metric in an IBC advanced motion vector prediction AMVP mode; And the decoder obtains the final prediction for the current CU by equally averaging the prediction results corresponding to the multiple block vectors.

39. The method according to claim 35, further comprising: include: The decoder uses iterative accumulation in multi-hypothesis prediction (MHP) to obtain the final prediction for the current CU based on the multiple block vectors.

40. A method for video encoding, include: The encoder obtains multiple block vectors for the current coding unit CU based on the intra block copy IBC mode; as well as A final prediction for the current CU is obtained by the encoder based on the multiple block vectors.

41. The method of claim 40, wherein the plurality of block vectors comprises a first block vector and a second block vector, and The method also include: The encoder obtains the first block vector based on a minimum rate distortion metric in an IBC advanced motion vector prediction AMVP mode; Obtaining, by the encoder, the second block vector based on a second minimum rate-distortion metric in the IBC AMVP mode; as well as The encoder obtains a final prediction for the current CU by equally averaging a prediction result of the first block vector and a prediction result of the second block vector.

42. The method of claim 40, wherein the plurality of block vectors comprises a first block vector and a second block vector, and The method also include: The encoder obtains the first block vector based on a minimum rate distortion metric in an IBC advanced motion vector prediction AMVP mode; Obtaining, by the encoder, the second block vector based on a minimum rate-distortion metric in an IBC merge mode; as well as The encoder obtains a final prediction for the current CU by equally averaging a prediction result of the first block vector and a prediction result of the second block vector.

43. The method according to claim 40, further comprising: include: The encoder obtains the plurality of block vectors based on a distortion metric in an IBC advanced motion vector prediction AMVP mode; And the encoder obtains the final prediction for the current CU by equally averaging the prediction results corresponding to the multiple block vectors.

44. The method according to claim 40, further comprising: include: The final prediction for the current CU is obtained by the decoder based on the multiple block vectors using iterative accumulation in multi-hypothesis prediction (MHP).

45. An apparatus for video decoding, include: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein the one or more processors are configured to perform the method according to any one of claims 1-9, 19-26 and 35-39 when executing the instructions.

46. ​​An apparatus for video encoding, include: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein the one or more processors are configured to perform the method of any one of claims 10-18, 27-34 and 40-44 when executing the instructions.

47. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1-9, 19-26, and 35-39.

48. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 10-18, 27-34, and 40-44.

49. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1-9, 19-26, and 35-39.

50. A non-transitory computer-readable storage medium for storing a bitstream generated by the method of any one of claims 10-18, 27-34, and 40-44.