Method and apparatus for intra block copy

By adopting the intra-block copy mode during the video encoding and decoding process, and using fractional motion information and refinement technology, the intra-block copy process is optimized, which solves the problem of inefficiency in the prior art and achieves more efficient video compression and quality maintenance.

CN120457696APending Publication Date: 2025-08-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089812.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-27
Filing Date
2023-12-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has problems of inefficiency and quality loss in intra-block replication, and it is difficult to effectively use the redundant information in the video image for efficient compression.

Method used

Intra-block copy (IBC) mode is adopted to obtain fractional motion information, half-pixel and quarter-pixel refine, and combine weighted average and template distortion cost calculation to optimize the intra-block copy process.

Benefits of technology

Improves the efficiency of video encoding and decoding, reduces bit rate requirements, while maintaining or improving video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457696A_ABST
    Figure CN120457696A_ABST
Patent Text Reader

Abstract

Methods and apparatus for video decoding and encoding are provided. In a video decoding method, a decoder can acquire fractional motion information (MV) in an IBC mode, where the fractional MV is signaled by an encoder and determined by: determining a first number of integer BVs having a minimum distortion cost; applying half-pixel refinement around each of the integer BV; obtaining a second number of optimal half-pixel positions for each of the integer BVs, wherein the second number of optimal half-pixel positions indicates the second number of half-pixel BV differences with the lowest rate distortion cost; obtaining a quarter-pixel refinement result by applying a quarter-pixel refinement around the second number of optimal half-pixel positions for each of the integer BVs; and obtaining the score motion information based on the quarter pixel refinement result. In addition, the decoder can obtain a final BV based on the fractional motion information; and obtaining a final prediction block based on the final BV.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based upon and claims the benefit of U.S. Provisional Application No. 63,435,369, filed on December 27, 2022, entitled “Method and Apparatus for Intra-Frame Block Copying,” the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0002] The present invention relates to video coding and compression, and more particularly to, but not limited to, methods and apparatus for improving intra-block copying schemes in video coding and decoding processes. Background Art

[0003] Various video coding and decoding technologies can be used to compress video data. Video coding and decoding are performed according to one or more video coding and decoding standards. For example, video coding and decoding standards include Versatile Video Codec (VVC), High Efficiency Video Codec (H.265 / HEVC), Advanced Video Codec (H.264 / AVC), Moving Picture Experts Group (MPEG) codec, etc. Video coding and decoding generally adopts prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilize redundancy in video images or sequences. An important goal of video coding and decoding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the Invention

[0004] This disclosure provides examples of techniques related to methods for improving intra-block copying during video encoding and decoding.

[0005] According to a first aspect of the present disclosure, a video decoding method is provided. In the method, a decoder is capable of obtaining fractional motion information of a current block in an intra-block copy (IBC) mode, wherein the fractional motion information is transmitted by an encoder using a signal and is determined by: searching for a first number of integer BVs with minimum distortion cost; applying half-pixel refinement around each of the first number of integer BVs; obtaining a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates the second number of half-pixel BV differences with the lowest rate-distortion cost; obtaining a quarter-pixel refinement result by applying quarter-pixel refinement around the second number of best half-pixel positions for each of the first number of integer BVs; and obtaining the fractional motion information based on the quarter-pixel refinement result. In addition, the decoder is capable of obtaining a final BV of the current block based on the fractional motion information; and obtaining a final prediction block of the current block based on the final BV.

[0006] According to a second aspect of the present disclosure, a video encoding method is provided. In the method, an encoder can determine a first number of integer BVs with minimum distortion costs for a current block in IBC mode. Furthermore, the encoder can apply half-pixel refinement around each of the first number of integer BVs and obtain a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates the second number of half-pixel BV differences with the lowest rate-distortion cost.

[0007] Furthermore, the encoder can obtain quarter-pixel refinement results by applying quarter-pixel refinement around the second number of best half-pixel positions for each of the first number of integer BVs, and obtain the fractional motion information based on the quarter-pixel refinement results. Furthermore, the encoder can encode the current block based on the fractional motion information.

[0008] According to a third aspect of the present disclosure, a video decoding method is provided. In this method, a decoder can obtain fractional motion information of a current block in IBC mode and obtain multiple BVs for the current block based on the fractional motion information. Furthermore, the decoder can obtain multiple motion-compensated prediction blocks associated with the multiple BVs and obtain a final prediction block for the current block by weighted averaging the multiple motion-compensated prediction blocks.

[0009] According to a fourth aspect of the present disclosure, a video encoding method is provided. In this method, an encoder can obtain fractional motion information of a current block in IBC mode and obtain multiple BVs for the current block based on the fractional motion information. In addition, the encoder can obtain multiple motion-compensated prediction blocks associated with the multiple BVs and obtain a final prediction block for the current block by weighted averaging the multiple motion-compensated prediction blocks.

[0010] According to a fifth aspect of the present disclosure, a video decoding method is provided. In the method, a decoder can obtain one or more block vectors of a current block based on fractional motion information in IBC mode. Furthermore, the decoder can calculate a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion. Furthermore, the decoder can reorder the one or more block vectors according to the template-based distortion cost.

[0011] According to a sixth aspect of the present disclosure, a video encoding method is provided. In the method, an encoder can obtain one or more block vectors of a current block based on fractional motion information in IBC mode. Furthermore, the encoder can calculate a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion. Furthermore, the encoder can reorder the one or more block vectors according to the template-based distortion cost.

[0012] According to a seventh aspect of the present disclosure, a video decoding method is provided. In the method, a decoder can obtain a BV prediction value for a current block in IBC mode. Furthermore, the decoder can receive one or more syntax elements to obtain multiple precisions of the BV prediction value. Furthermore, the decoder can determine whether to obtain a BV difference for the current block based on the multiple precisions.

[0013] According to an eighth aspect of the present disclosure, a video encoding method is provided. In the method, an encoder can obtain a BV prediction value for a current block in IBC mode. Furthermore, the encoder can signal one or more syntax elements to obtain multiple precisions of the BV prediction value. Furthermore, the encoder can determine whether to obtain a BV difference for the current block based on the multiple precisions.

[0014] According to a ninth aspect of the present disclosure, a video decoding method is provided. In the method, a decoder is capable of obtaining multiple motion vector candidate lists. Furthermore, the decoder is capable of obtaining an updated motion vector candidate list by grouping multiple motion vector candidates in the multiple motion vector candidate lists into different groups based on a grouping criterion. Furthermore, the decoder is capable of obtaining at least one of the group index or the candidate list index from the updated motion vector candidate list. Furthermore, the decoder is capable of obtaining a motion vector index of a motion vector for a current block to be predicted based on one of the group index or the candidate list index.

[0015] According to a tenth aspect of the present disclosure, a video encoding method is provided. In the method, an encoder can obtain multiple motion vector candidate lists. Furthermore, the encoder can obtain an updated motion vector candidate list by grouping multiple motion vector candidates in the multiple motion vector candidate lists into different groups based on a grouping criterion. Furthermore, the encoder can obtain at least one of the group index or the candidate list index from the updated motion vector candidate list. Furthermore, the encoder can obtain a motion vector index of a motion vector for a current block to be predicted based on one of the group index or the candidate list index.

[0016] According to an eleventh aspect of the present disclosure, a video decoding method is provided. In this method, a decoder can obtain at least one block vector of a current block in IBC mode or through intra-frame template matching (ITM). Furthermore, the decoder can obtain a final prediction block based on the at least one block vector and both the IBC mode and the ITM.

[0017] According to a twelfth aspect of the present disclosure, a video encoding method is provided. In this method, an encoder can obtain at least one block vector of a current block in an intra block copy (IBC) mode or by intra template matching (ITM). Furthermore, the encoder can obtain a final prediction block based on the at least one block vector and both the IBC mode and the ITM.

[0018] According to a thirteenth aspect of the present disclosure, a video encoding apparatus is provided. The apparatus may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to, when executing the instructions, perform the method according to the first, third, fifth, seventh, ninth, or eleventh aspect.

[0019] According to a fourteenth aspect of the present disclosure, a video encoding apparatus is provided. The apparatus may include one or more processors and a memory, the memory being coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the second, fourth, sixth, eighth, tenth, or twelfth aspects when executing the instructions.

[0020] According to the fifteenth aspect of the present disclosure, a non-temporary computer-readable storage medium is provided for storing computer-executable instructions, which, when executed by one or more computer processors, causes the one or more computer processors to execute the method according to the first aspect, the third aspect, the fifth aspect, the seventh aspect, the ninth aspect or the eleventh aspect.

[0021] According to the sixteenth aspect of the present disclosure, a non-temporary computer-readable storage medium is provided for storing computer-executable instructions, which, when executed by one or more computer processors, causes the one or more computer processors to execute the method according to the second aspect, the fourth aspect, the sixth aspect, the eighth aspect, the tenth aspect or the twelfth aspect.

[0022] According to a seventeenth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first, third, fifth, seventh, ninth or eleventh aspect.

[0023] According to an eighteenth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a bitstream generated by the method according to the second aspect, the fourth aspect, the sixth aspect, the eighth aspect, the tenth aspect or the twelfth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Examples of the present disclosure will be described in more detail with reference to specific examples shown in the accompanying drawings, which are intended to describe and explain these examples with additional specificity and detail, as these drawings depict only some examples and are not to be considered limiting in scope.

[0025] Figure 1A is a block diagram illustrating a system for encoding and decoding video blocks according to some examples of the present disclosure.

[0026] Figure 1B A quadtree data structure is shown, which illustrates some examples according to the present disclosure. Figure 1E The final result of the segmentation process for CTU 400 is shown.

[0027] Figure 1C An encoded representation of a frame by first partitioning the frame into a set of CTUs is shown according to some examples of the present disclosure.

[0028] Figure 1D A CTU according to some examples of the present disclosure is shown, which includes one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements for encoding the samples of the coding tree blocks.

[0029] Figure 2A is a block diagram illustrating an exemplary video encoder according to some examples of the present disclosure.

[0030] Figure 2B is a block diagram illustrating an exemplary video decoder according to some examples of the present disclosure.

[0031] Figure 3A is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0032] Figure 3B is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0033] Figure 3Cis a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0034] Figure 3D is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0035] Figure 3E is a diagram illustrating block partitioning in a multi-type tree structure according to some examples of the present disclosure.

[0036] Figures 4A-4B An example of a 4-parameter affine model according to some examples of the present disclosure is shown.

[0037] Figure 5 An example of a 6-parameter affine model according to some examples of the present disclosure is shown.

[0038] Figure 6 Examples of adjacent neighboring blocks that inherit affine merging candidates according to some examples of the present disclosure are shown.

[0039] Figure 7 Examples of neighboring neighboring blocks constructing affine merging candidates according to some examples of the present disclosure are shown.

[0040] Figure 8 The current CTU processing order and available reference samples thereof in the current CTU and the left CTU according to some examples of the present disclosure are shown.

[0041] Figure 9 Padding candidates for replacing zero vectors in an IBC list according to some examples of the present disclosure are shown.

[0042] Figure 10 1 shows a reference area of IBC when encoding CTU(m,n) according to some examples of the present disclosure.

[0043] Figure 11 IBC reference areas for camera captured content according to some examples of the present disclosure are shown.

[0044] Figures 12A-12B A division method for angular mode according to some examples of the present disclosure is shown.

[0045] Figure 13A Spatially neighboring blocks used by ATVMP according to some examples of the present disclosure are shown.

[0046] Figure 13B An example of deriving a sub-CU motion field by applying motion offsets from spatially neighboring blocks and scaling motion information from corresponding co-located sub-CUs according to some examples of the present disclosure is shown.

[0047] Figure 14 is a flowchart of decoding a binary bit (bin) according to some examples of the present disclosure.

[0048] Figure 15 is a diagram illustrating a computing environment coupled with a user interface according to some examples of the present disclosure.

[0049] Figure 16 An intra-frame template matching search area used in some examples of the present disclosure is shown.

[0050] Figure 17 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0051] Figure 18 is a diagram showing some examples according to the present disclosure corresponding to Figure 17 The video decoding method shown is a flowchart of a video encoding method.

[0052] Figure 19 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0053] Figure 20 is a diagram showing some examples according to the present disclosure corresponding to Figure 19 The video decoding method shown is a flowchart of a video encoding method.

[0054] Figure 21 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0055] Figure 22 is a diagram showing some examples according to the present disclosure corresponding to Figure 21 The video decoding method shown is a flowchart of a video encoding method.

[0056] Figure 23 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0057] Figure 24 is a diagram showing some examples according to the present disclosure corresponding to Figure 23 The video decoding method shown is a flowchart of a video encoding method.

[0058] Figure 25 is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0059] Figure 26 is a diagram showing some examples according to the present disclosure corresponding to Figure 25 The video decoding method shown is a flowchart of a video encoding method.

[0060] Figure 27is a flowchart illustrating a video decoding method according to some examples of the present disclosure.

[0061] Figure 28 is a diagram showing some examples according to the present disclosure corresponding to Figure 27 The video decoding method shown is a flowchart of a video encoding method. DETAILED DESCRIPTION

[0062] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to facilitate understanding of the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices with digital video capabilities.

[0063] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. In the disclosure and appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless otherwise expressly indicated throughout the disclosure. It should also be understood that the term "and / or" as used in this disclosure refers to and includes one or any or all possible combinations of the listed items.

[0064] Throughout the specification, references to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that the particular feature, structure, or characteristic being described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in conjunction with one or some embodiments may also apply to the other embodiments unless expressly stated otherwise.

[0065] Throughout this disclosure, the terms "first," "second," "third," and the like are used solely to refer to related elements, such as devices, components, ingredients, steps, and the like, and do not imply any spatial or temporal order, unless expressly stated otherwise. For example, "first device" and "second device" can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be arbitrarily named.

[0066] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" can include memory (shared, dedicated, or grouped) that stores code or instructions that can be executed by one or more processors. A module can include one or more circuits with or without stored code or instructions. A module or circuit can include one or more directly or indirectly connected components. These components may or may not be physically connected or adjacent to each other.

[0067] As used herein, the terms "if" or "when" can be understood to mean "according to" or "in response to," depending on the context. These terms, if appearing in a claim, may not indicate that the associated limitation or feature is conditional or optional. For example, a method can include the steps of: i) when or if condition X exists, performing function or action X', and ii) when or if condition Y exists, performing function or action Y'. The method can be implemented with the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' can be performed at different times in multiple executions of the method.

[0068] A unit or module may be implemented entirely by software, entirely by hardware, or a combination of hardware and software. For example, in a pure software implementation, a unit or module may include functionally related code blocks or software components that are linked together directly or indirectly to perform a specific function.

[0069] Figure 1A FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1A As shown in , system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including a cloud server, a server computer, a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0070] In some embodiments, target device 14 can receive the coded video data to be decoded via link 16. Link 16 can include any type of communication medium or device capable of moving the coded video data from source device 12 to target device 14. In one example, link 16 can include a communication medium that enables source device 12 to send the coded video data directly to target device 14 in real time. The coded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to target device 14. The communication medium can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form a part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium can include a router, a switch, a base station, or any other device that helps to achieve communication from source device 12 to target device 14.

[0071] In some other embodiments, the encoded video data can be sent from the output interface 22 to a storage device 32. The encoded video data in the storage device 32 can then be accessed by the target device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending the encoded video data to the target device 14. Exemplary file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data through any standard data connection suitable for accessing encoded video data stored on a file server, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both. The transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both streaming and download transmissions.

[0072] like Figure 1A As shown in , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include a source such as a video capture device (e.g., a video camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application may be generally applicable to video encoding and decoding, and may be applied to wireless and / or wired applications.

[0073] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be sent directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0074] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0075] In some implementations, target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0076] The video encoder 20 and the video decoder 30 can operate according to a proprietary standard or an industry standard (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or an extension of such a standard. It should be understood that the present application is not limited to a specific video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally believed that the video encoder 20 of the source device 12 can be configured to encode the video data according to any of these current standards or future standards. Similarly, it is also generally believed that the video decoder 30 of the target device 14 can be configured to decode the video data according to any of these current standards or future standards.

[0077] The video encoder 20 and the video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuits, respectively, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and either encoder or decoder can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0078] In some embodiments, components of source device 12 (e.g., video source 18, video encoder 20, or the like) may be configured to: Figure 2A The components included in the video encoder 20 and the output interface 22) and / or the components of the target device 14 (for example, the input interface 28, the video decoder 30 or the following reference Figure 2BThe components included in the video decoder 30 and at least a portion of the components in the display device 34 may operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), wherein the cloud computing service network may provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be provided in one or more client devices, and the one or more client devices may communicate with a server computer in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In one embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in a cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of the present disclosure, terms such as "cloud," "cloud computing," and "cloud-based" may be used interchangeably herein as appropriate. It should be understood that the present disclosure is not limited to implementation in the above-mentioned cloud computing service network. Instead, the present disclosure may also be implemented in any other type of computing environment currently known or developed in the future.

[0079] Figures 3A-3E is a schematic diagram illustrating multiple types of tree splitting modes according to some embodiments of the present disclosure. Figures 3A-3E Five segmentation types are shown, including four-element segmentation ( Figure 3A ), vertical binary segmentation ( Figure 3B ), horizontal binary segmentation ( Figure 3C ), vertical ternary split ( Figure 3D ) and horizontal ternary split ( Figure 3E ).

[0080] Figure 2AFIG2 is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described herein. Video encoder 20 can perform intra-frame prediction coding and inter-frame prediction coding on video blocks within a video frame. Intra-frame prediction coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or image. Inter-frame prediction coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or images of a video sequence. It should be noted that in the field of video coding, the term "frame" can be used as a synonym for the term "image" or "picture."

[0081] like Figure 2A As shown in FIG, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, adder 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, segmentation unit 45, intra prediction processing unit 46, and intra block copy (BC) unit 48. In some embodiments, video encoder 20 also includes an inverse quantization unit 58 for video block reconstruction, an inverse transform processing unit 60, and adder 62. A loop filter 63, such as a deblocking filter, can be located between adder 62 and DPB 64 to filter block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can also be used to filter the output of adder 62. It should be noted that with respect to the CCSAO technique, the present application is not limited to the embodiments described herein, but may also be applied to the case where an offset is selected for any one of the luma component, the Cb chroma component, and the Cr chroma component based on any one of the luma component, the Cb chroma component, and the Cr chroma component to modify the other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein may be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component mentioned herein may be any one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component mentioned herein may be the remaining component of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filter may be omitted, and the decoded video block may be provided directly by the adder 62 to the DPB 64. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the fixed or programmable hardware units described.

[0082] Video data memory 40 may store video data to be encoded by the components of video encoder 20. Figure 1AThe video source 18 shown obtains video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by the video encoder 20 when encoding the video data (e.g., in intra-frame or inter-frame prediction coding mode). The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with the other components of the video encoder 20, or off-chip relative to those components.

[0083] like Figure 2A As shown in FIG, after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into slices, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined splitting structure associated with the video data, such as a quadtree (QT) structure. A video frame is or can be viewed as a two-dimensional array or matrix of samples having sample values. The samples in the array may also be referred to as pixels or picture elements (pels). The number of samples in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame may be divided into multiple video blocks using QT segmentation. A video block is also or can be viewed as a two-dimensional array or matrix of samples having sample values, but at a smaller scale than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further segmented into one or more block partitions or sub-blocks (which may again form blocks) by, for example, iteratively using QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof. It should be noted that the term "block" or "video block" used herein may be a portion of a frame or image, in particular a rectangular (square or non-square) portion. With reference to, for example, HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or may be or correspond to a corresponding block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.

[0084] The prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block based on the error results (e.g., coding rate and distortion level), such as one of one or more inter-frame prediction coding modes among a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block or inter-frame prediction coding block to the adder 50 to generate a residual block and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, partition information, and other such syntax information) to the entropy coding unit 56.

[0085] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.

[0086] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for the current video frame by generating motion vectors according to a predetermined pattern within the sequence of video frames, where the motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. For example, a motion vector may indicate the displacement of a video block within the current video frame or image relative to a prediction block within a reference frame associated with the current block being encoded within the current frame. The predetermined pattern may designate the video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors (e.g., block vectors) for intra BC coding in a manner similar to the motion vectors determined by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine the block vectors.

[0087] In terms of pixel differences, the prediction block for a video block may be or may correspond to a block or reference block of a reference frame that is considered to closely match the video block to be encoded, and the pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches relative to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0088] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing the position of the video block to the position of a prediction block of a reference frame selected from either a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.

[0089] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block may include luma component differences, chroma component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are described separately for conceptual purposes.

[0090] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used to encode the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example during separate encoding passes, and test their performance using rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original, uncoded block that was coded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate ratios based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0091] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the prediction block may include calculating values for sub-integer pixel positions.

[0092] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from a different frame according to inter-frame prediction, video encoder 20 can form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values forming the residual video block may include both luma component differences and chroma component differences.

[0093] As an alternative to the inter-frame prediction performed by motion estimation unit 42 and motion compensation unit 44 or the intra-frame block copy prediction performed by intra BC unit 48 as described above, intra-frame prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, intra-frame prediction processing unit 46 can determine an intra-frame prediction mode to use for encoding the current block. To do so, intra-frame prediction processing unit 46 can use various intra-frame prediction modes to encode the current block, for example, during separate encoding passes, and intra-frame prediction processing unit 46 (or in some examples, mode selection unit) can select an appropriate intra-frame prediction mode to use from the tested intra-frame prediction modes. Intra-frame prediction processing unit 46 can provide information indicating the intra-frame prediction mode selected for the block to entropy coding unit 56. Entropy coding unit 56 can encode the information indicating the selected intra-frame prediction mode into the bitstream.

[0094] After prediction processing unit 41 determines a prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0095] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan on the matrix comprising the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0096] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to a video bitstream such as Figure 1A The video decoder 30 shown, or archived as Figure 1A The video frame is shown in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0097] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain for use in generating a reference block for predicting other video blocks. As noted above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0098] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block to inter-predict another video block in a subsequent video frame.

[0099] Figure 2B 3 is a block diagram illustrating another exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform the above-mentioned operations in combination with the above-mentioned operations. Figure 2A The encoding process is essentially the inverse of the decoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.

[0100] In some examples, units of the video decoder 30 may be tasked with performing embodiments of the present application. Furthermore, in some examples, embodiments of the present disclosure may be dispersed across one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (e.g., the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functionality of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., the motion compensation unit 82).

[0101] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (e.g., a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding the video data (e.g., in intra-frame or inter-frame prediction coding modes). The video data memory 79 and the DPB 92 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown in FIG. Figure 2B 92 as two distinct components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, video data memory 79 may be on-chip with the other components of video decoder 30, or off-chip relative to those components.

[0102] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-frame prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0103] When a video frame is encoded as an intra-frame prediction coded (I) frame or for intra-frame coded prediction blocks in other types of frames, intra-frame prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-frame prediction mode transmitted by the signal and reference data from a previously decoded block of the current frame.

[0104] When the video frame is encoded as an inter-frame prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vector and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can use a default construction technique to construct the reference frame lists, i.e., List 0 and List 1, based on the reference frames stored in the DPB 92.

[0105] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within the same reconstructed region of the image as defined by video encoder 20 as the current video block.

[0106] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode the video block of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-prediction-encoded video block of the frame, the inter-prediction state for each inter-prediction-encoded video block of the frame, and other information used to decode the video block in the current video frame.

[0107] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine whether the current video block is predicted using intra BC mode, construction information of which video blocks of the frame are within the reconstruction region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.

[0108] Motion compensation unit 82 may also perform interpolation to calculate interpolated values for sub-integer pixels of a reference block using interpolation filters, such as those used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from received syntax elements and use these interpolation filters to produce the prediction block.

[0109] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters that were calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.

[0110] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (e.g., a deblocking filter, an SAO filter, a CCSAO filter, and / or an ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be provided directly to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, may also store the decoded video for later presentation on a display device (e.g., Figure 1A on the display device 34).

[0111] In a typical video encoding process, a video sequence typically consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.

[0112] like Figure 1C As shown in , the video encoder 20 (or more specifically, the partitioning unit in the prediction processing unit of the video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. As Figure 1DAs shown in , each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding pixel blocks and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In monochrome images or images with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding the samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0113] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, on the coding treeblock of the CTU and divide the CTU into smaller CUs. Figures 1B-1E FIG is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure. Figure 1E As depicted in FIG, a 64×64 CTU 400 is first divided into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. Two 16×16 CUs 430 and CU 440 are each further divided into four CUs with a block size of 8×8. Figure 1B Depicted is a diagram showing Figure 1E The quadtree data structure is the final result of the partitioning process of the CTU 400 depicted in FIG. 4 , with each leaf node of the quadtree corresponding to a CU of a corresponding size ranging from 32×32 to 8×8. Figure 1D Each CU may include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size frame, and syntax elements for encoding the samples of the coding blocks. In a monochrome image or an image with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. It should be noted that Figure 1E and Figure 1B The quadtree partitioning depicted in FIG is for illustrative purposes only, and one CTU can be split into multiple CUs based on quadtree partitioning / ternary tree partitioning / binary tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to the quadtree structure, and each quadtree leaf CU can be further partitioned according to the binary and ternary tree structures. Figures 3A-3E As shown, there are five possible partition types for a coding block with width W and height H, namely, quadruple partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0114] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.

[0115] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0116] After the video encoder 20 generates the predicted luma block, the predicted Cb block, and the predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, such that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0117] In addition, if Figure 1EAs shown in , the video encoder 20 can use quadtree partitioning to decompose the luma residual block, Cb residual block and Cr residual block of a CU into one or more luma transform blocks, Cb transform blocks and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. The TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome image or an image with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0118] Video encoder 20 may apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block may be a two-dimensional array of transform coefficients. A transform coefficient may be a scalar. Video encoder 20 may apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0119] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream comprising a sequence of bits forming a representation of the encoded frame and associated data, which is stored in the storage device 32 or sent to the target device 14.

[0120] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. The video decoder 30 also reconstructs the coding blocks of the current CU by adding samples of the prediction blocks for the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0121] As mentioned above, video coding mainly uses two modes: intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction) to achieve video compression. It should be noted that IBC can be regarded as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on the reference video block.

[0122] However, with the continuous improvement of video data capture technology and finer video block sizes for preserving details in video data, the amount of data required to represent the motion vector for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that a group of neighboring CUs in the spatial and temporal domains not only have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU (which is also called the "motion vector predictor" (MVP) of the current CU) by exploiting their spatial and temporal correlations.

[0123] Instead of combining as above Figure 2A The actual motion vector of the current CU determined by the motion estimation unit is encoded into the video bitstream, and the motion vector prediction value of the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0124] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also called a "merged list") for the current CU using those potential candidate motion vectors associated with the spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then selecting a member from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, there is no need to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value in the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector prediction value in the motion vector candidate list to encode and decode the current CU.

[0125] The main focus of the present disclosure is to further enhance the intra block copy method by improving the coding efficiency and / or reducing its coding complexity.

[0126] Affine model

[0127] In HEVC, only the translational motion model is applied to motion compensation prediction. In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, affine motion compensation prediction is applied by signaling a flag for each inter-frame coding block to indicate whether the translational motion model or the affine motion model is applied to inter-frame prediction. In the current VVC, an affine coding block supports two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode.

[0128] The 4-parameter affine model has the following parameters: two parameters for translation motion in the horizontal and vertical directions, one parameter for scaling motion, and one parameter for rotation motion in both directions. In this model, the horizontal scaling parameter is equal to the vertical scaling parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. In order to better fit the motion vector and affine parameters, these affine parameters are derived from the two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. Figures 4A-4B As shown, the affine motion field of a block is described by two CPMVs (V0, V1). Based on the motion of the control points, an affine coded block motion field (v x , v y ) is described as

[0129] The 6-parameter affine model has the following parameters: two parameters for translation in the horizontal and vertical directions, two parameters for scaling and rotation in the horizontal direction, and two parameters for scaling and rotation in the vertical direction. The 6-parameter affine motion model is encoded using three CPMVs. Figure 5 As shown, the three control points of a 6-parameter affine block are located at the upper left corner, upper right corner, and lower left corner of the block. The movement of the upper left control point is related to the translation movement, the movement of the upper right control point is related to the horizontal rotation and scaling movement, and the movement of the lower left control point is related to the vertical rotation and scaling movement. Compared with the 4-parameter affine motion model, the horizontal rotation and scaling movement of the 6-parameter can be different from the vertical movement. Assume that (V0, V1, V2) is Figure 5 The MVs of the upper left, upper right, and lower left corners of the current block are used to derive the motion vector (v) of each sub-block using the three MVs at the control points. x , v y ),as follows:

[0130] Affine merge mode

[0131] In affine merge mode, the CPMV of the current block is not explicitly signaled, but derived from neighboring blocks. Specifically, in this mode, the motion information of spatially neighboring blocks is used to generate the CPMV of the current block. The affine merge mode candidate list has a limited size. For example, in the current VVC design, there can be at most five candidates. The encoder is able to evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. There are three ways to decide on the affine merge candidate: Inherited from neighboring affine encoding blocks Constructed from the translation MV of neighboring blocks Zero MV

[0132] For the inherited method, there are at most two candidates. If available, the next block to the left and below the current block (e.g. Figure 6 As shown, the scanning order is from A0 to A1) and the adjacent block located to the upper right of the current block (e.g., Figure 6 As shown, the scanning order is from B0 to B2) to obtain candidates.

[0133] For the constructed method, a candidate is a combination of neighboring block translation MVs, which is generated in two steps. • Step 1: Get four translation MVs from available neighboring blocks. οMV1: The MV of one of the three neighboring blocks near the upper left corner of the current block. Figure 7As shown, the scanning order is B2, B3 and A2. οMV2: The MV of one of the two neighboring blocks near the upper right corner of the current block. Figure 7 As shown, the scanning order is B1 and B0. οMV3: The MV of one of the two neighboring blocks near the lower left corner of the current block. Figure 7 As shown, the scanning order is A1 and A0. οMV4: The MV of the temporally co-located block of the neighboring block near the lower right corner of the current block. Figure 7 As shown, the neighboring blocks are T. Step 2: Derive a combination based on the four translation MVs in step 1. οCombination 1: MV1, MV2, MV3 ο Combination 2: MV1, MV2, MV4 ο Combination 3: MV1, MV3, MV4 ο Combination 4: MV2, MV3, MV4 ο Combination 5: MV1, MV2 ο Combination 6: MV1, MV3

[0134] When the merge candidate list is not full after being populated with inherited and constructed candidates, a zero MV is inserted at the end of the list.

[0135] Affine AMVP mode

[0136] Affine AMVP (Advanced Motion Vector Prediction) mode can be applied to CUs with width and height both greater than or equal to 16. A CU-level affine flag is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and their predicted value CPMVP is signaled in the bitstream. The affine AMVP candidate list size is 2, and the affine AMVP candidate list is generated by using the following four types of CPMV candidates in sequence: - Inherited affine AMVP candidate extrapolated from the CPMV of neighboring CUs -Affine AMVP candidate CPMVP constructed using the translation MV of the neighboring CU -Translated MV from neighboring CU -Time MV from the same CU -Zero MV

[0137] The order in which inherited affine AMVP candidates are checked is the same as the order in which inherited affine merge candidates are checked. The only difference is that for AMVP candidates, only affine CUs with the same reference image as the current block are considered. When the inherited affine motion prediction value is inserted into the candidate list, the pruning process is not applied.

[0138] The constructed AMVP candidates are derived from the same spatially neighboring blocks as the affine merge mode. The same check order as in the affine merge candidate construction is used. In addition, the reference picture index of the neighboring blocks is also checked. The first block in the check order is used, which is inter-coded and has the same reference image as the current CU. When the current CU is encoded with a 4-parameter affine mode and both mv0 and mv1 are available, mv0 and mv1 are added to the affine AMVP candidate list as a candidate. When the current CU is encoded with a 6-parameter affine mode and all three CPMVs are available, they are added to the affine AMVP candidate list as a candidate. Otherwise, the constructed AMVP candidate is set to unavailable.

[0139] If after inserting the valid inherited affine AMVP candidates and constructed AMVP candidates, the number of affine AMVP list candidates is still less than 2, when mv0, mv1 and mv2 are available, then mv0, mv1 and mv2 will be added in sequence as translation MVs to predict all control point MVs of the current CU. Finally, if the affine AMVP list is still not full, it will be filled with zero MVs.

[0140] Intra-block copying in Versatile Video Coding (VVC)

[0141] Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current image. The luminance block vector of the IBC-encoded CU is integer precision. The chrominance block vector is also rounded to integer precision. When used in conjunction with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The IBC-encoded CU is regarded as a third prediction mode different from the intra or inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples.

[0142] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed.

[0143] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. The hash key calculation for each location in the current image is based on a 4×4 sub-block. For larger current block sizes, a hash key is determined to match the hash key of a reference block when all hash keys of all 4×4 sub-blocks match the hash key in the corresponding reference location. If multiple reference blocks are detected whose hash keys match the hash key of the current block, the block vector cost is calculated for each matching reference block, and the reference block with the smallest cost is selected.

[0144] In the block matching search, the search range is set to cover the previous and current CTUs.

[0145] At the CU level, the IBC mode is signaled through a flag, which can be signaled as IBC AMVP mode or IBC skip / merge mode as shown below: -IBC skip / merge mode: The merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial, HMVP and pairwise candidates. - IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if IBC coding). When either neighbor is unavailable, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.

[0146] IBC Reference Area

[0147] To reduce memory consumption and decoder complexity, VVC's IBC only allows the reconstruction of predefined areas including the area of the current CTU and some areas of the left CTU. Figure 8 Describes the reference area for IBC mode, where each block represents a 64×64 luma sample unit.

[0148] Depending on the position of the current coding CU within the current CTU, the following applies: -If the current block falls into the upper left 64×64 block of the current CTU, in addition to the reconstructed samples in the current CTU, the CPR mode can also be used to refer to the reference samples in the lower right 64×64 block of the left CTU. Using the CPR mode, the current block can also refer to the reference samples in the lower left 64×64 block of the left CTU and the reference samples in the upper right 64×64 block of the left CTU. -If the current block falls into the upper right 64×64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (0, 64) relative to the current CTU has not been reconstructed, the CPR mode is used, and the current block can also refer to the reference samples in the lower left 64×64 block and the lower right 64×64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU. - If the current block falls into the lower left 64×64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64, 0) relative to the current CTU has not been reconstructed, the CPR mode is used and the current block can also refer to the reference samples in the upper right 64×64 block and the lower right 64×64 block of the left CTU. Otherwise, the CPR mode is used and the current block can also refer to the reference samples in the lower right 64×64 block of the left CTU. - If the current block falls into the lower right 64×64 block of the current CTU, only the CPR mode can be used to refer to the reconstructed samples in the current CTU.

[0149] This restriction allows the IBC mode to be implemented in hardware using local on-chip memory.

[0150] Interaction of IBC with other coding tools

[0151] The interaction between IBC mode and other inter coding tools of VVC, such as pairwise merge candidates, history-based motion vector prediction (HMVP), combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM), is as follows: -IBC can be used with pairwise merge candidates and HMVP. A new pairwise IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference. - IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM. - When using dual-tree (DUAL_TREE) partitioning, IBC is not allowed for chroma coded blocks.

[0152] Unlike the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in the reference picture list 0 for IBC prediction. The derivation process of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects are applied: - IBC shares the same process as in regular MV merging, including pairwise merging candidates and history-based motion prediction, but TMVP and zero vectors are not allowed since they are not valid for IBC mode. - Separate HMVP buffers (5 candidates each) for regular MV and IBC. -Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream, and if the merge candidate is invalid (out of range or 0), then the merge should not be used. As described below, this bitstream consistency constraint is expressed in terms of virtual buffers. -For deblocking, IBC is processed in inter mode. - If the current block is encoded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to only indicate whether the MV is inter pixels or 4 integer pixels. - The number of IBC merge candidates may be signaled in the slice header separately from the number of regular, sub-block and geometric merge candidates.

[0153] The concept of a virtual buffer is used to describe the permissible reference area for IBC prediction modes and valid block vectors. Denoting the CTU size as ctbSize, the width of the virtual buffer ibcBuf is wIbcBuf = 128x128 / ctbSize, and the height is hIbcBuf = ctbSize. For example, for a 128x128 CTU size, the ibcBuf size is also 128x128; for a 64x64 CTU size, the ibcBuf size is 256x64, and for a 32x32 CTU size, the ibcBuf size is 512x32.

[0154] In each dimension, the size of the VPDU is min(ctbSize, 64), Wv=min(ctbSize, 64).

[0155] For the virtual IBC buffer, ibcBuf is maintained as follows. - Flushes the entire ibcBuf with an invalid value of -1 at the start of decoding each CTU row. -When starting to decode the VPDU (xVPDU, yVPDU) relative to the upper left corner of the image, set ibcBuf[x][y] = -1, x = xVPDU% wIbcBuf, ..., xVPDU% wIbcBuf + Wv-1; y = yVPDU% ctbSize, ..., yVPDU% ctbSize + Wv-1. - After decoding the CU containing (x, y) relative to the top left corner of the picture, set ibcBuf[x % wIbcBuf][y % ctbSize] = recSample[x][y]

[0156] For a block covering coordinates (x, y), for a block vector bv = (bv[0], bv[1]), if the following holds, it is valid; otherwise, it is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1.

[0157] Intra-block copying in the Enhanced Compression Model (ECM)

[0158] In ECM, IBC has been improved in the following aspects.

[0159] IBC merger / AMVP list construction

[0160] IBC merge / AMVP list construction has been modified as follows: An IBC merge / AMVP candidate can be inserted into the IBC merge / AMVP candidate list only if it is valid. • Top-right, bottom-left, and top-left spatial candidates and one pairwise average candidate may be added to the IBC merge / AMVP candidate list. • Adaptive Reordering Based on Template (ARMC-TM) is applied to the IBC merge list.

[0161] The HMVP table size for IBC is increased to 25. After obtaining up to 20 IBC merge candidates through full pruning, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.

[0162] The zero vector candidates that populate the IBC merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. In IBC merge mode, the zero vector is invalid as a block vector, so it is discarded as a BVP in the IBC candidate list.

[0163] Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Figure 9 shown.

[0164] IBC with template matching

[0165] In IBC, template matching is used in IBC merge mode and IBC AMVP mode.

[0166] Compared to the list used by the regular IBC merge mode, the IBC-TM merge list is modified so that the motion distance between candidates in the regular TM merge mode is used to select candidates according to the pruning method. The ending zero motion is replaced by the motion vectors to the left (-W, 0), up (0, -H), and left-up (-W, -H), where W is the width of the current CU and H is its height.

[0167] In IBC-TM merge mode, a template matching method is used to refine the selected candidates before the RDO or decoding process.The IBC-TM merge mode has been made competitive with the conventional IBC merge mode and the TM-Merge flag is signaled.

[0168] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of these 3 selected candidates is refined using the template matching method and ranked according to their final template matching cost. Then, as usual, only the top two are considered in the motion estimation process.

[0169] Template matching refinement for IBC-TM merging and AMVP mode is very simple because the IBC motion vectors are constrained to be (i) integer and (ii) within the reference region, as Figure 8 As shown. Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision, depending on the AMVR value. This refinement only accesses samples without interpolation. In both cases, the refined motion vectors and the template used in each refinement step must take into account the constraints of the reference area.

[0170] IBC Reference Area

[0171] The reference area of the IBC extends to the upper two CTU rows. Figure 10The reference region used to encode a CTU (m, n) is shown. Specifically, for a CTU (m, n) to be encoded, the reference region includes CTUs with indices (m–2, n–2)…(W, n–2), (0, n–1)…(W, n–1), (0, n)…(m, n), where W represents the maximum horizontal index within the current partition, slice, or picture. This setup ensures that IBC does not require additional memory in current ETM platforms for CTUs of size 128. The range of the per-sample block vector search (or local search) is limited horizontally to [–(C<<1), C>>2] and vertically to [–C, C>>2] to accommodate the expansion of the reference region, where C represents the CTU size.

[0172] IBC merge mode with block vector difference

[0173] In ECM, IBC merging mode with block vector difference is adopted. The distance set is {1-pixel, 2-pixel, 4-pixel, 8-pixel, 12-pixel, 16-pixel, 24-pixel, 32-pixel, 40-pixel, 48-pixel, 56-pixel, 64-pixel, 72-pixel, 80-pixel, 88-pixel, 96-pixel, 104-pixel, 112-pixel, 120-pixel, 128-pixel}, and the BVD directions are two horizontal directions and two vertical directions.

[0174] The base candidate is selected from the first five candidates in the reordered IBC merge list. And all possible MBVD refinement positions (20×4) for each base candidate are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. Finally, the first 8 refinement positions with the lowest template SAD cost are retained as available positions and are therefore used for MBVD index encoding.

[0175] IBC adaptation of camera-captured content

[0176] like Figure 11 As shown in Figure 1, when IBC is adjusted for camera-captured content, the IBC reference range is reduced from 2 CTU rows to 2 × 128 rows. On the encoder side, to reduce complexity, the local search range is set to [–8, 8] horizontally and [–8, 8] vertically, centered on the first block vector prediction value of the current CU. This encoder modification does not apply to SCC sequences.

[0177] Sub-block based temporal motion vector prediction (SbTMVP)

[0178] VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to HEVC's temporal motion vector prediction (TMVP), SbTMVP uses the motion field in the co-located image to improve the motion vector prediction and merge mode of the CU in the current image. The same co-located image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: -TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; -While TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right or center block relative to the current CU), SbTMVP applies a motion offset before obtaining the temporal motion information from the co-located picture, where the motion offset is obtained from the motion vector from one of the spatially neighboring blocks of the current CU.

[0179] Figures 13A-13B The SbTMVP process is shown in Figure 2. The SbTMVP predicts the motion vector of the sub-CU in the current CU in two steps. In the first step, the motion vector of the sub-CU is detected. Figure 13A The spatial neighbor A1 in . If A1 has a motion vector that uses the co-located image as its reference image, then that motion vector is selected as the motion offset to be applied. If no such motion is identified, then the motion offset is set to (0, 0).

[0180] In the second step, the motion offset identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located image, as Figure 13B shown. Figure 13B The example in assumes that the motion offset is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block in the co-located image (the minimum motion grid covering the center sample) is used to derive the motion information of the sub-CU. After the motion information of the co-located sub-CU is identified, it is converted into a motion vector and reference index for the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.

[0181] In VVC, a combined sub-block based merge list is used for signaling of sub-block based merge mode, where the combined sub-block based merge list contains SbTVMP candidates and affine merge candidates. SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the sub-block based merge candidate list, followed by the affine merge candidate. The size of the sub-block based merge list is signaled in the SPS, and the maximum allowed size of the sub-block based merge list in VVC is 5.

[0182] The sub-CU size used in SbTMVP is fixed to 8×8. Like the affine merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.

[0183] The encoding logic for the additional SbTMVP merge candidates is the same as that for other merge candidates, ie, for each CU in a P slice or a B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.

[0184] Intra-frame template matching

[0185] Intra Template Matching (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side.

[0186] The prediction signal is generated by matching the L-shaped causal neighborhood of the current block with another block in the predefined search area in Figure 4, which includes: R1: Current CTU R2: Upper left CTU R3: Upper CTU R4: Left CTU

[0187] The sum of absolute differences (SAD) is used as the cost function.

[0188] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.

[0189] The sizes of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW SearchRange_h=a*BlkH Where "a" is a constant that controls the gain / complexity tradeoff. In this example, "a" is equal to 5.

[0190] The intra template matching tool is enabled for CUs with width and height less than or equal to 64. The maximum CU size for intra template matching is configurable.

[0191] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag.

[0192] Probability Estimation Techniques for CABAC in AVC and HEVC

[0193] CABAC (Context-Adaptive Binary Arithmetic Coding) was originally introduced in the H.264 / AVC standard as one of two supported entropy coding schemes. In CABAC, arithmetic coding consists of two modules: codeword mapping (also known as binarization) and probability estimation. During codeword mapping, syntax elements are mapped into binary bit strings. This mapping is achieved by a binarizer, which translates syntax elements into multiple groups of bits based on different binarization schemes. In practice, various binarization schemes can be used for this translation, such as fixed-length codes, unary codes, truncated unary codes, and k-order exponential Golomb codes. One purpose of the probability estimation module is to determine the probability that a bit has the value 1 or 0. In AVC, the probability of a bit is calculated based on an exponential aging model, where the probability of a current bit being equal to 1 or 0 depends on the value of the previously encoded bit. Furthermore, according to public data statistics, the influence of the bit immediately preceding the current bit is generally greater than that of the bit encoded more recently. Taking this into account, a parameter α is introduced in CABAC, which controls the number N of previously coded bins used to estimate the probability of the current bin, that is, N = 1 / α. This parameter translates into an adaptive speed at which the probability is updated as the number of coded bins increases. Specifically, by adapting the parameter α, the probability that a bin is the least likely symbol (LPS) is recursively calculated as p(t+1)=p(t)·(1-α)+x(t)·α (3) where p(t) is the probability of the LPS symbol at time instant t; p(t+1) is the updated probability of the LPS symbol at time instant t+1; x(t) is equal to 1 when the current bin is an LPS symbol and equal to 0 when the current bin is the most probable symbol (MPS). In the CABAC engine of AVC and HEVC, the probability is updated independently according to (3) for each syntax element with a fixed value α≈1 / 19.69, that is, when estimating the probability of a current bin, approximately 19.69 previously coded bins are considered. In addition, to avoid multiplication operations during probability estimation, the probability p(t) in (3) is quantized into a set of fixed probability states, which are real numbers ranging from 0 to 1. For example, in AVC and HEVC, the probability has 7 bits of precision, corresponding to 128 probability states.

[0194] In AVC and HEVC, a video bitstream typically consists of one or more independently decodable segments. At the beginning of each segment, the probabilities of all contexts are initialized to some predefined values. In theory, if the statistical properties of a given context are known, a uniform distribution (i.e., p init = 0.5) to initialize the context probabilities. However, in order to make the probability of a context catch up with its corresponding statistical distribution more quickly, it is found that it is beneficial to provide some appropriate initial probability values (which may not be equally probable) for each context. Specifically, in AVC and HEVC, given a slice SliceQP Y The initial QP of a context, the initial probability state InitProbState is calculated as follows: m=SlopeIdx·5-45 (4) n=(OffsetIdx<<3)-16 InitProbState=Clip3(1,127,(m·SliceQP Y )>>4+n) Where SlopeIdx and OffsetIdx (both ranging from 0 to 15) are two initialization parameters that are predefined and stored as a lookup table (LUT) to calculate the initial probability of a context. As shown in (4), the initial probability state is modeled by a linear function of the segment QP with a slope equal to (m>>4) and an offset equal to n.

[0195] Probability estimation technology of CABAC in VVC

[0196] The probability estimation module applied in VVC is almost the same as that in AVC and HEVC, except for the following key differences: 1. VVC maintains two probability estimates for each context, where each context has its own probability adaptation rate α in (3). The final probability actually used for arithmetic coding is the average of the two estimates. 2. In VVC, multiple probability LUTs are predefined and used to initialize the probabilities of different contexts within a clip. Also, similar to AVC and HEVC, the initial estimate of the probabilities is based on a linear model that takes the clip QP as input. However, in VVC, the derived values represent the actual probability values, while in AVC / HEVC, the derived values represent the index of the probability state.

[0197] Multiple hypothesis probability estimation

[0198] Obviously, using a fixed adaptive parameter for all syntax elements may not be optimal because they have different statistical properties. On the other hand, several scientific studies have shown that using multiple probability estimators can achieve better estimation accuracy than a single estimator. Therefore, a multi-hypothesis probability estimation scheme is applied in the CABAC design of VVC, in which two different adaptive parameters α0 and α1 are utilized, which corresponds to a slow speed and a fast speed of probability adaptation. In this way, two different probabilities can be calculated for each binary bit using two adaptive parameters, and then averaged to generate the final probability of the binary bit, that is, Where α0 and α1 are two adaptive parameters associated with two probability hypotheses. In VVC, the values of α0 and α1 are selected independently for each context using a training algorithm designed to jointly optimize the adaptive parameters and the initial probabilities. Specifically, according to the current design, each context is allowed to select α0 from a set of predefined values {1 / 4, 1 / 8, 1 / 16, 1 / 32} and α1 from another set of predefined values {1 / 32, 1 / 64, 1 / 128, 1 / 256, 1 / 512}.

[0199] Initial probability calculation

[0200] As in AVC / HEVC, VVC's CABCA process also calls a QP-related probability initialization process at the beginning of each segment. However, compared to AVC / HEVC, which initializes the state of a probability state machine, the actual value of the initial probability is directly derived as shown below SlopeIdx and OffsetIdx are two initialization parameters used to calculate the slope and offset of the linear model, and each parameter is represented with 3 bits of precision; and are the two initial probabilities calculated for the two probability estimators.

[0201] Entropy Coding in ECM

[0202] Extended Precision

[0203] The intermediate precision used in the arithmetic coding engine has been increased, including three elements. First, the precision of both probability states has been increased to 15 bits compared to 10 bits and 14 bits in VVC. Second, the LPS range update process has been modified as follows, If q>=16384 q=2 15 –1–q R LPS=((range*(q>>6))>>9)+1, Among them, range is a 9-bit variable representing the width of the current interval, q is a 15-bit variable representing the probability state of the current context model, R LPS is the update range of LPS. This operation can also be implemented by looking up 512×256 entries in a 9-bit lookup table. Third, at the encoder side, the 256-entry lookup table for VTM bit estimation is expanded to 512 entries.

[0204] Window size based on fragment type

[0205] Since statistics vary for different segment types, it is desirable to update the context probability state at a rate that provides more accurate probability estimates (e.g., more accurate predictions of the likelihood of a binary bit having a value of 1 or 0) for a given segment type. Therefore, for each context model, three window sizes are predefined for I-segments, B-segments, and P-segments, respectively, as initialization parameters.

[0206] The context initialization parameters and window size are kept unchanged.

[0207] Improved CABAC probability estimation

[0208] Multi-hypothesis probability estimation with adaptive weights

[0209] The probability based on multiple hypotheses is estimated based on adaptive weights (MHP-AW). Specifically, two independent probability estimates p0 and p1 are maintained for each context and updated according to their own adaptation rates. However, instead of using a simple average, multiple weights are introduced to derive the resulting probability p for binary arithmetic coding as follows: p=(ω0·p0+ω1·p1)>>s Where ω0 and ω1 are weights selected from a predefined set {10, 12, 16, 20, 22}; s is a bitwise right-shift value equal to 5 when (ω0 + ω1) ≤ 32, and equal to 6 otherwise. Three different sets of weights are predefined for each context model of I, B, and P slice types. The weights for I slice types are only allowed for intra slices, while the weights for B and P slice types can be switched at the slice level for inter slices.

[0210] CABAC initialization based on previous inter-frame segments and window adjustment

[0211] After encoding the last CTU, the context initialization stored in the previously coded picture can be used to initialize inter-frame segments with the same segment type, QP, and temporal ID. For each segment type, the buffer size used to store the previous initialization is set equal to 5. When the buffer is full, the entry with the smallest QP and temporal ID is removed first before storing the initialization.

[0212] CABAC uses two probability states, updated with short and long window sizes respectively. The predefined window size for each context model is not optimal for the varying statistics in different regions, so the window size is adjusted based on the previously coded bins of each context.

[0213] The short and long window sizes used in the CABAC update are adjusted by two delta parameters stored in a lookup table for each context and retrieved by the previously encoded bin used as an index. The previously encoded bin is used as an index to obtain the adjustment parameters from the lookup table: delta0 for the short window and delta1 for the long window. The original short and long window sizes stored in the existing initialization table and defined for the context model are denoted as shift0 and shift1, respectively. After the adjustment, the actual window sizes used to encode the current bin are (shift0+delta0) and (shift1+delta1), respectively, where shift0 and shift1 are the existing predefined window sizes stored in the context initialization table.

[0214] Problem Statement

[0215] In video coding, intra-block copying (IBCC) is well known for its ability to accurately predict both screen content and artificially generated content, where patterns and edges may repeat within a frame. Intra-block copying is also beneficial for natural content prediction in situations where the current frame has repetitive textures. For coding scenarios without much repetitive content, IBC mode can be deselected while still transmitting the minimum signaling bits. In order to further improve the coding efficiency of IBC, it is desirable to provide a more flexible enabling / disabling control mechanism at different granularities.

[0216] In inter-frame prediction coding, fractional motion vectors are used to improve prediction accuracy. However, in the current intra block copy (IBC) mode, only integer motion vectors are used. It is desirable to explore the beneficial effects of fractional motion vectors for intra block copy coding. When using fractional motion in intra block copy, several subsequent issues need to be addressed: fractional motion derivation, signaling, interpolation padding, interpolation filter selection, and interaction with other coding tools.

[0217] The present invention improves the intra-block copy coding tool from the following aspects: Flexible enable / disable control mechanism CABAC context window Interpolation-based fractional intra block copying ο Score Sports Search ο Fractional motion refinement o Conditional sample / pixel padding for fractional interpolation οInterpolation filter switching o Multi-hypothesis fractional intra block copy Signal transmission of motion information IBC merger / AMVP campaign candidate list construction Combined with intra-frame template matching Combination with IBC merge mode with block vector difference

[0218] Flexible enable / disable control mechanism

[0219] In this section, several methods are proposed to enable / disable control of the application of IBC mode. The enable / disable control indicates whether IBC mode is allowed to be enabled for the current sequence, frame, slice, CTU, or block at different granularities. If IBC mode is enabled, further flags (e.g., whether IBC mode is enabled or disabled for a particular block) and / or information (e.g., block vectors) can be signaled. If IBC mode is disabled, no flags or information are signaled.

[0220] In some embodiments, the enable / disable control of intra block copy can be based on an explicit signaling method.

[0221] In one embodiment, the enable / disable control is based on one or more sequence level, or frame level, or slice level, or coding tree unit (CTU) level, or block level flags, or any combination of different level flags. When any combination of different level flags is used, the transmission of lower level flags depends on the enable / disable of higher level flags. In one example, if the frame level flag indicates that IBC mode is disabled, no flag is transmitted at the slice or block level. Otherwise, the lower level flag is further transmitted.

[0222] In another embodiment, the enable / disable control is based on different zones.One purpose of the zone concept is to provide a more flexible granularity for IBC enable / disable control.

[0223] In one embodiment, a region can be defined as a non-overlapping region within a frame, slice, or CTU. For all blocks within a specific region, a single enable / disable control flag can be set to indicate whether IBC mode is disabled for all these blocks. The size of the region can be predefined as a set of fixed values, such as M×N, or a set of signaled values.

[0224] In some other embodiments, the enable / disable control of intra block copy can be based on local information and does not require explicit signaling.

[0225] In some embodiments, the enable / disable control is based on prediction information. In one embodiment, IBC mode is always disabled for inter-frame predicted blocks. In another embodiment, IBC mode is always disabled for inter-frame unidirectional predicted blocks and / or inter-frame bidirectional predicted blocks.

[0226] In another embodiment, IBC mode is always disabled for blocks encoded in sub-block mode. Sub-block mode is a mode that divides the current block into sub-blocks, and each sub-block can have its own motion information. Examples include affine mode and SbTMVP mode. In another embodiment, IBC mode is always disabled for blocks not encoded in sub-block mode.

[0227] In some other embodiments, the enable / disable control is based on other coding information. In one embodiment, the IBC mode is always disabled when one or more other coding modes are applied to the current block. For example, when the affine mode is enabled, the IBC mode is always disabled.

[0228] In some other embodiments, the enable / disable control is based on the frame type. In one embodiment, IBC mode is always disabled for B frames and / or P frames.

[0229] In some other embodiments, the enable / disable control is based on block information. In one embodiment, IBC mode is always disabled for coding blocks smaller than a certain size (e.g., 8×8 blocks) or larger than a certain size (e.g., 64×64). In one embodiment, IBC mode is always disabled for wide blocks (e.g., blocks whose width is M times longer than their height) or long blocks (e.g., blocks whose height is N times longer than their width), and the values of M and N can be fixed values (e.g., M=2, N=3) or values signaled at the sequence level or frame level.

[0230] CABAC context window

[0231] In current IBC designs, one or more IBC mode-related flags may be CABAC context-encoded. For example, the block-level IBC enable flag is context-encoded. Since statistics may vary for different segment or frame types, it is desirable to update the context probability state at a rate that provides more accurate probability estimates (e.g., more accurately predicting the likelihood of a binary bit having a value of 1 or 0) for a given segment / frame type.

[0232] In some embodiments, for each context model associated with the IBC mode, three windows can be predefined for three different segments, including an I segment, a B segment, and a P segment, respectively.

[0233] In some embodiments, for each context model associated with the IBC mode, two windows can be predefined for different segments with two different prediction modes, including intra-frame prediction segments (I segments) and inter-frame prediction segments (B segments and P segments).

[0234] When multiple windows are defined for different segments or frames, the context window size and initialization parameters can also be kept unchanged, either individually or jointly.

[0235] Interpolation-based fractional intra block copying

[0236] Score Sports Search

[0237] In one embodiment, fractional motion search can be performed on the encoder side and the final motion can be signaled to the decoder side. The signaled motion can be in the form of motion difference after subtracting a motion prediction value known to both the encoder and decoder. The motion search can be performed in three steps: • In step 1, the best N integer motion vectors with the minimum distortion cost (eg, sum of absolute differences (SAD)) may be searched first. In step 2, half-pixel refinement is applied around each of the N integer motion vectors. In this step 2, the M best half-pixel positions can be obtained (the best M positions can indicate M half-pixel motion differences with the lowest rate-distortion cost). For example, the encoder or decoder can obtain the M best half-pixel positions with the lowest rate-distortion cost. If K of the N integer motion vectors are selected, the output can be a total of K*M half-pixel positions. In step 3, quarter-pixel refinement is applied around the best half-pixel position of the N integer motion vectors. In this step 3, for each of the K*M half-pixel positions obtained in step 2, a set of Q quarter-pixel positions can be obtained. And the best R quarter-pixel positions among all K*M*Q candidate positions can be generated. The best position among the R positions (e.g., the position with the minimum rate distortion) can be determined by a full rate-distortion calculation and signaled to the decoder. For example, the encoder can generate the best R quarter-pixel positions and then select the best position among the R positions with the minimum rate distortion and signal it to the decoder. The values of N, M, K, Q, and R are integer numbers of positions.

[0238] After the three steps, the best refined motion vector (after half-pel and / or quarter-pel refinement) is signaled (e.g., in the format of a motion vector difference). In this disclosure, motion vectors are used interchangeably with block vectors, which identify reference / prediction blocks in the same picture / frame in this and subsequent sections.

[0239] In another embodiment, the fractional motion search may be performed on both the encoder side and the decoder side so that the final fractional motion does not need to be signaled. In this method, a template matching based method may be used to find the best fractional motion.

[0240] In one or more embodiments, an inverted L-shaped sample / pixel area adjacent to a coding block may be used as a matching template, and the pixel / sample width may be preset, configurable, or signaled at a sequence level, and / or a picture level, and / or a fragment level, and / or a CTU level.

[0241] Within a constrained search area (defined by a preset, configurable or signaled number of CTUs, CTU lines or samples from the above spatial region, the left spatial region and / or the top-left spatial region), the template similarity between any adjacent / non-adjacent reference blocks and the current coding block is calculated, and the best N reference blocks with the closest similarity are selected as candidates in the template list.

[0242] An additional flag is signaled to indicate whether the template matching method is used. If the flag is true, another index value is further signaled to indicate which candidate in the template list is used.

[0243] In another embodiment, the encoder search method and the template matching method are used in conjunction. For example, integer motion and fractional refinement methods are first employed at the encoder, and then further template refinement is applied at both the encoder and decoder sides. Because the encoder search method is already sufficiently precise, template refinement can be performed with higher precision and in smaller regions. For example, motion refinement on the encoder side is performed with half-pixel or quarter-pixel precision, while template refinement can be further performed with quarter-pixel, eighth-pixel, or sixteenth-pixel precision.

[0244] Fractional Motion Refinement

[0245] With or without the use of the fractional motion search process, it is possible to identify the starting motion vector (MV). The adjustment of the starting MV can be based on two reasons: • For smaller signaling overhead, the starting Mv may be rounded to a specific precision or value such that the difference in mv between the starting Mv and the selected mv prediction value is minimized. • For smaller signaling overhead, a few least significant bits of the start MV may be discarded.

[0246] With or without the above adjustments, the starting Mv may need to be refined at the decoder side.

[0247] In one or more embodiments, a method based on template matching can be used. In one example, an inverted L-shaped sample / pixel area adjacent to a coding block can be used as a matching template. The starting Mv can be refined at integer pixel or / and fractional pixel levels. Potential refinement sets can be {1 / 4 pixel, 2 / 4 pixel, 3 / 4 pixel} or / and {1 / 8 pixel, 3 / 8 pixel, 5 / 8 pixel, 7 / 8 pixel}, and the refinement directions are two horizontal directions and two vertical directions (positive and negative values). The refined Mv that produces the prediction block with the most similar template is selected as the final Mv. In particular, if the most similar template is selected, the selected refinement can be implicitly derived by the decoder, and if multiple refinement Mvs with N most similar templates are derived, they can be explicitly derived by the encoder.

[0248] In one or more embodiments, an additional flag can be signaled to indicate whether the fractional motion refinement is applied.The additional flag can be transmitted at the sequence level, picture level, slice level or CTU level.

[0249] Conditional sample / pixel padding for fractional interpolation

[0250] When using fractional-based Mv, the interpolation operation may require a greater number of pixels / samples than the current block. The actual difference depends on the interpolation filter tap length. In cases where some pixels / samples are unavailable, a pixel / sample padding process may be required. Different padding schemes can be used.

[0251] In one or more embodiments, a type of repeated filling can be used. Unavailable pixel / sample positions can be filled based on the same value of the nearest available pixel / sample in the same row or column. This repeated filling can be performed first in the horizontal direction (left and right border filling), and then in the vertical direction (top and bottom border filling). Alternatively, this repeated filling can be performed first in the vertical direction (top and bottom border filling), and then in the horizontal direction (left and right border filling).

[0252] In one or more embodiments, a type of symmetric padding can be used. Unavailable pixel / sample positions can be filled based on pixels at positions symmetric to the padding boundary. This padding can be performed first in the horizontal direction (left or right border padding) and then in the vertical direction (top or bottom border padding). Alternatively, this repeated padding can be performed first in the vertical direction (top or bottom border padding) and then in the horizontal direction (left or right border padding).

[0253] Interpolation filter switching

[0254] Interpolation filters may need to be switched for different reasons. For example, if the image / video content is rich in noise and a smooth filtering effect is desired, a longer tap length of the filter may be preferred. On the other hand, if padding complexity needs to be reduced or the image / video content has rich texture edges, a shorter tap length of the filter may be preferred.

[0255] In one or more embodiments, the switching of filters can be decided at the decoder side by analyzing the image / video content (eg gradient histogram), which does not require signaling bits.

[0256] In one or more other embodiments, filter switching can be evaluated at the encoder side and signaled at different granularities (sequence level, picture level, slice level, CTU level, or region based).

[0257] Multiple Hypothesis Fractional Intra Block Copy

[0258] When multiple motion vectors (from motion search and / or motion refinement) are available, multiple prediction blocks can be generated.Multiple hypothesis intra block copying can be used when the average of multiple similar blocks can generate a better block prediction.

[0259] In one or more embodiments, the number of multiple hypotheses may be predefined, configured, or signaled. In addition, the weights for averaging the multiple prediction hypotheses may also be predefined, configured, or signaled.

[0260] In one or more other embodiments, the number of multiple hypotheses can be determined implicitly at the decoder side. For example, if N prediction blocks can be generated, and the value signaled is N, which is outside the range of valid single prediction blocks (0 to N-1), then it indicates that multiple hypotheses are enabled and the average of all N prediction blocks can be used.

[0261] Multiple hypotheses can be generated from N motion / block vectors, each of which can generate a specific motion compensated prediction block, where N is a positive integer. The N motion / block vectors can be obtained from the same candidate list or different candidate lists. In one example, the N motion / block vectors can be obtained from the same IBC merge candidate list or AMVP list, or partially from the IBC merge candidate list and partially from the IBC AMVP candidate list. In another example, the N motion / block vectors can be obtained in whole or in part from an intra-frame template matching method.

[0262] Template matching interpolation processing

[0263] In case that template-based adaptive reordering (ARMC™) is applied to IBC merging and / or AMVP mode, or / and template matching-based motion refinement is applied to IBC merging and / or AMVP mode, it may be necessary to adaptively use fractional motion / block vectors.

[0264] For Template-Based Adaptive Reordering (ARMC™), a template-based distortion cost may need to be calculated for each candidate motion / block vector. Since this template-based distortion cost is only used for candidate reordering and not for the final compensated prediction, the fractional portion of each motion / block vector candidate may or may not need to be considered for the template-based distortion calculation if there is a non-zero fractional portion. Specifically, in one or more examples, for each motion / block vector candidate that does not have a non-zero fractional portion, fractional motion-based interpolation may or may not be performed.

[0265] Similarly, when computing the template distortion cost for each motion refinement position, the fractional part of each position may or may not need to be considered.

[0266] Signal transmission of motion information

[0267] When motion vectors in intra block copy support multiple precisions, the allowed signaling methods may be defined accordingly.

[0268] In one or more embodiments, only one precision is allowed to have zero motion vector difference. This precision can be predefined, configurable, or signaled. For example, the precision can be predefined to be the highest precision supported by the motion vector, such as 1 / 4 pixel or 1 / 8 pixel.

[0269] In one or more embodiments, multiple precisions are allowed to have zero motion vector difference. These multiple precisions can be predefined, configurable, or signaled. For example, the precision can be predefined as the highest precision supported by the motion vector or the second highest precision, such as 1 / 4 pixel and 1 pixel. In the case of allowing multiple precisions to have zero motion vector difference, after the signaled indication of zero motion vector difference (1 flag or 1 bit), one or more other flags are sent to indicate which precision to use.

[0270] In the case of supporting multiple precisions, the current precision flag can be signaled in different ways. In one example, a flag indicating whether the current precision is greater than 0 is first signaled. If so, another flag indicating whether the current precision is greater than 1 is further signaled. Alternatively, a second flag indicating whether the current precision is greater than 1 can be implicitly derived at the decoder without explicit signaling. In one example, the value of the motion / block vector difference can be used to achieve this purpose (for example, an even or odd motion / block vector difference can indicate a specific motion precision value). Here, values of 0, 1, or other values greater than 1 can be predefined or configured to represent different motion vector (or motion vector difference) precisions (for example, 0 represents 1-pixel precision, 1 represents 1 / 2-pixel precision, 2 represents 1 / 4-pixel precision, and 3 represents 1 / 8-pixel precision).

[0271] In one or more embodiments, multiple MV candidates in the IBC merge / AMVP motion candidate list are grouped into different groups. In one example, the grouping criterion can be MV precision, where MV candidates in the same group have the same actual MV precision. The actual MV precision is defined as the MV precision after all the least significant zero bits of the MV are right-shifted.

[0272] IBC merger / AMVP movement candidate list construction

[0273] In one or more embodiments, multiple MV candidates in the IBC merge / AMVP motion candidate list are grouped into different groups. In one example, the grouping criterion can be MV precision, where MV candidates in the same group have the same actual MV precision. The actual MV precision is defined as the MV precision after all the least significant zero bits of the MV are right-shifted.

[0274] In one or more other embodiments, multiple IBC merge / AMVP motion candidate lists are created in addition to the existing list. For each list, only Mv candidates with the same actual Mv precision are added. Similarly, the actual Mv precision is defined as the Mv precision after all the least significant zero bits of Mv are right-shifted.

[0275] In the case of generating multiple groups of candidate lists or / and multiple candidate lists, before the actual Mv candidate index can be determined, the group index or / and candidate list index need to be determined first. In one or more other examples, the group index or / and candidate list index can be evaluated at the encoder side and then signaled to the decoder. In still another one or more other examples, the group index or / and candidate list index can be inherited from a specific neighboring block without explicit signaling.

[0276] Combined with intra-frame template matching

[0277] When a motion vector is determined for intra block copying, a prediction block can be generated based on the motion vector. When combined with intra template matching, the prediction block generated by intra block copying is further refined through intra template matching. Specifically, for a predefined search range around the prediction block generated by intra block copying, the encoder searches for the template that is most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. In this method, the prediction block generated by intra block copying is treated as the starting block position to guide the subsequent block search process in the intra template matching method.

[0278] The combination of intra block copy and intra template matching prediction modes can be signaled at the CU level through a dedicated flag. Alternatively, the original intra block copy flag on top of the original intra template matching mode flag can be used to indicate the combination of intra block copy and intra template matching prediction modes. Alternatively, the original intra template matching mode flag on top of the original intra block copy mode flag can be used to indicate the combination of intra block copy and intra template matching prediction modes.

[0279] In one or more other examples, a combination of intra block copy and intra template matching prediction can generate improved motion vectors.

[0280] In one example, the intra block copy (IBC) mode provides an initial motion vector, which can be further refined by an intra template matching method.

[0281] In another example, the motion / block vectors obtained from the intra template matching method can be reused to generate the IBC merge candidate list or the IBC AMVP candidate list. In this case, the motion / block vectors generated for spatially adjacent or non-adjacent neighboring blocks can be cached or saved. In one example, the motion / block vectors of neighboring blocks encoded using the intra template matching method can be saved in a historical motion vector table. In addition, or alternatively, the motion / block vectors of neighboring blocks encoded using the intra template matching method can be saved in the encoder's local cache and then reused by the encoder during the motion search process.

[0282] In other examples, a combination of intra block copy and intra template matching prediction can generate an improved prediction block. In one instance, two prediction blocks can be generated by intra block copy and intra template matching prediction methods, respectively, and a weighted average of the two prediction blocks can be generated to represent the final prediction block of the current coding block. In another example, multiple prediction blocks can be generated separately (e.g., N>1), where M of the N prediction blocks (e.g., M is less than or equal to N) can be generated by intra block copy, and S of the N prediction blocks (e.g., S is less than or equal to N) can be generated by intra template matching prediction. When combining N prediction blocks (e.g., N>1), the weight value can be derived by using a matching cost value (e.g., one example of matching cost calculation can be based on an L-shaped template, and a higher matching cost value can indicate a lower weight value, while a lower matching cost value can indicate a higher weight value) or a least squares method.

[0283] Combination with IBC merge mode with block vector difference

[0284] With the support of fractional Mv, the IBC merge mode with block vector differences can be expanded by adopting more candidate distance values. In one or more other embodiments, there can be two distance sets. The first set is the existing integer distance set, and the second set is additionally added for fractional distances. In one example, the fractional distance set can be: {1 / 8-pixel, 2 / 8-pixel, 3 / 8-pixel, 4 / 8-pixel, 5 / 8-pixel, 6 / 8-pixel, 7 / 8-pixel}. In another example, the fractional distance set can be: {1 / 8-pixel, 2 / 8-pixel, 4 / 8-pixel}. The BVD directions of the second set are also two horizontal directions and two vertical directions.

[0285] Figure 15A computing environment (or computing device) 1610 is shown coupled to a user interface 1650. The computing environment 1610 may be part of a data processing server. In some embodiments, the computing device 1610 may perform any of the various methods or processes (e.g., encoding / decoding methods or processes) described above, according to various examples of the present disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.

[0286] The processor 1620 generally controls the overall operation of the computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the above-described method. In addition, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0287] The memory 1630 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1630 may include predetermined software 1632. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, etc. The memory 1630 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0288] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (e.g., keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 may be coupled to an encoder and a decoder.

[0289] Figure 17 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 17 The method shown in is also described in the "Fractional Motion Search" section.

[0290] In step 1701, at the decoder side, the processor 1620 can obtain fractional motion information of the current block in IBC mode. Specifically, the fractional motion information is signaled by the encoder and determined by: searching for a first number of integer BVs with minimum distortion cost; applying half-pixel refinement around each of the first number of integer BVs by obtaining a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates a second number of half-pixel BV differences with the lowest rate-distortion cost; obtaining quarter-pixel refinement results by applying quarter-pixel refinement around the second number of best half-pixel positions for each of the first number of integer BVs; and obtaining the fractional motion information based on the quarter-pixel refinement results.

[0291] In step 1702 , at the decoder side, the processor 1620 can obtain the final BV of the current block based on the fractional motion information.

[0292] In step 1703 , at the decoder side, the processor 1620 can obtain a final prediction block of the current block based on the final BV.

[0293] In some examples, processor 1620 may obtain a second number of optimal half-pixel positions for each of the first number of integer BVs by selecting a third number of integer BVs from the first number of integer BVs and obtaining a fourth number of half-pixel positions, where the fourth number is equal to the product of the second number and the third number. For example, the second number (M) of optimal half-pixel positions with the lowest rate-distortion cost may be obtained. If a third number (K) of the first number (N) of integer motion vectors is selected, the output may be a total of K*M half-pixel positions.

[0294] In some examples, processor 1620 may obtain a quarter-pixel refinement result by applying quarter-pixel refinement around a second number of best half-pixel positions for each of the first number of integers BV, by: obtaining a fifth number (Q) of quarter-pixel positions for each of the fourth number of half-pixel positions, obtaining a sixth number (R) of best quarter-pixel positions for all candidate positions, where the number (K*M*Q) of all candidate positions is equal to the product of the second number, the third number, and the fifth number, and selecting a best quarter-pixel position with minimum rate distortion from the sixth number of best quarter-pixel positions.

[0295] Figure 18 is shown corresponding to Figure 17 The video decoding method shown is a flowchart of a video encoding method. Figure 18 The method shown in is also described in the "Fractional Motion Search" section.

[0296] In step 1801 , at the encoder side, the processor 1620 can determine a first number of integer BVs with minimum distortion cost for a current block in IBC mode.

[0297] In step 1802, on the encoder side, the processor 1620 can apply half-pixel refinement around each of the first number of integer BVs by obtaining a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates a second number of half-pixel BV differences having a lowest rate-distortion cost.

[0298] In step 1803 , at the encoder side, the processor 1620 may obtain quarter-pixel refinement results by applying quarter-pixel refinement around a second number of best half-pixel positions for each of the first number of integer BVs.

[0299] In step 1804 , at the encoder side, the processor 1620 can obtain fractional motion information based on the quarter refinement result.

[0300] In step 1805 , at the encoder side, the processor 1620 can encode the current block based on the fractional motion information.

[0301] In some examples, processor 1620 may obtain a second number (M) of best half-pixel positions for each of the first number of integer BVs by selecting a third number of integer BVs from the first number of integer BVs and obtaining a fourth number of half-pixel positions, where the fourth number is equal to the product of the second number and the third number. For example, the second number (M) of best half-pixel positions with the lowest rate-distortion cost may be obtained. If a third number (K) of the first number (N) of integer motion vectors is selected, the output may be a total of K*M half-pixel positions.

[0302] In some examples, processor 1620 may obtain a quarter-pixel refinement result by applying quarter-pixel refinement around a second number of best half-pixel positions for each of the first number of integers BV, by: obtaining a fifth number (Q) of quarter-pixel positions for each of the fourth number of half-pixel positions, obtaining a sixth number (R) of best quarter-pixel positions for all candidate positions, where the number (K*M*Q) of all candidate positions is equal to the product of the second number, the third number, and the fifth number, and selecting a best quarter-pixel position with minimum rate distortion from the sixth number of best quarter-pixel positions.

[0303] Figure 19 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 19The method shown in is also described in the section “Multiple Hypothesis Fractional Intra Block Copying”.

[0304] In step 1901 , at the decoder side, the processor 1620 can obtain fractional motion information of a current block in an intra block copy (IBC) mode.

[0305] In step 1902 , at the decoder side, the processor 1620 can obtain multiple BVs of the current block based on the fractional motion information.

[0306] In step 1903 , at the decoder side, the processor 1620 can obtain a plurality of motion compensated prediction blocks associated with a plurality of BVs.

[0307] In step 1904 , at the decoder side, the processor 1620 can obtain a final prediction block of the current block by performing weighted averaging on multiple motion compensated prediction blocks.

[0308] In some examples, processor 1620 may obtain multiple BVs for the current block based on the fractional motion information by obtaining multiple BVs from the same candidate list or different candidate lists.

[0309] In some examples, the processor 1620 can obtain multiple BVs for the current block based on the fractional motion information by one of the following steps: obtaining multiple BVs from the same IBC merge candidate list or the same advanced motion vector prediction (AMVP) list; obtaining multiple BVs from the IBC merge candidate list and the AMVP list; or obtaining multiple BVs based on intra template matching (ITM).

[0310] Figure 20 is shown corresponding to Figure 19 The video decoding method shown is a flowchart of a video encoding method. Figure 20 The method shown in is also described in the section “Multiple Hypothesis Fractional Intra Block Copying”.

[0311] In step 2001 , at the encoder side, the processor 1620 can obtain fractional motion information of a current block in an intra block copy (IBC) mode.

[0312] In step 2002 , at the encoder side, the processor 1620 can obtain multiple BVs of the current block based on the fractional motion information.

[0313] In step 2003 , at the encoder side, the processor 1620 can obtain a plurality of motion compensated prediction blocks associated with a plurality of BVs.

[0314] In step 2004 , at the encoder side, the processor 1620 can obtain a final prediction block of the current block by performing weighted averaging on multiple motion compensated prediction blocks.

[0315] In some examples, processor 1620 may obtain multiple BVs for the current block based on the fractional motion information by obtaining multiple BVs from the same candidate list or different candidate lists.

[0316] In some examples, the processor 1620 can obtain multiple BVs for the current block based on the fractional motion information by one of the following steps: obtaining multiple BVs from the same IBC merge candidate list or the same advanced motion vector prediction (AMVP) list; obtaining multiple BVs from the IBC merge candidate list and the AMVP list; or obtaining multiple BVs based on intra template matching (ITM).

[0317] Figure 21 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 21 The method shown in is also described in the section “Interpolation Processing of Template Matching”.

[0318] In step 2101 , at the decoder side, the processor 1620 can obtain one or more block vectors of a current block based on fractional motion information in an intra block copy (IBC) mode.

[0319] In step 2102 , at the decoder side, the processor 1620 can calculate a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion.

[0320] In step 2103 , at the decoder side, the processor 1620 can reorder one or more block vectors according to the template-based distortion cost.

[0321] In some examples, processor 1620 can calculate a template-based distortion cost for one or more block vectors based on determining whether one or more block vectors include a non-zero fractional portion by: in response to determining that the first block vector does not include a non-zero fractional portion, performing, by the decoder, a fractional motion-based interpolation process on the first block vector, or in response to determining that the second block vector does not include a non-zero fractional portion, determining not to perform a fractional motion-based interpolation process on the second block vector. For Template-Based Adaptive Reordering (ARMC™), a template-based distortion cost may need to be calculated for each candidate motion / block vector. Because this template-based distortion cost is only used for candidate reordering and not for the final compensated prediction, if a non-zero fractional portion is present, the fractional portion of each motion / block vector candidate may or may not need to be considered for template-based distortion calculation. Specifically, in one or more examples, fractional motion-based interpolation may or may not be performed for each motion / block vector candidate that does not have a non-zero fractional portion.

[0322] In some examples, processor 1620 can calculate a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion by: calculating a template-based distortion cost for each fractional motion refinement position of the one or more block vectors based on the determination of whether each fractional motion refinement position includes a non-zero fractional portion. When calculating the template distortion cost for each motion refinement position, the fractional portion for each position may or may not be considered.

[0323] Figure 22 is shown corresponding to Figure 21 The video decoding method shown is a flowchart of a video encoding method. Figure 22 The method shown in is also described in the section “Interpolation Processing of Template Matching”.

[0324] In step 2201 , at the encoder side, the processor 1620 can obtain one or more block vectors of a current block based on fractional motion information in an intra block copy (IBC) mode.

[0325] In step 2202 , at the encoder side, the processor 1620 can calculate a template-based distortion cost for one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion.

[0326] In step 2203 , at the decoder side, the processor 1620 can reorder one or more block vectors according to the template-based distortion cost.

[0327] In some examples, processor 1620 can calculate a template-based distortion cost for one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion by: in response to determining that the first block vector does not include the non-zero fractional portion, performing fractional motion-based interpolation processing on the first block vector, or in response to determining that the second block vector does not include the non-zero fractional portion, determining not to perform fractional motion-based interpolation processing on the second block vector.

[0328] In some examples, processor 1620 can calculate a template-based distortion cost for one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion, by calculating a template-based distortion cost for each fractional motion refinement position of the one or more block vectors based on a determination of whether each fractional motion refinement position includes a non-zero fractional portion.

[0329] Figure 23 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 23 The method shown in is also described in the section "Signaling Motion Information".

[0330] In step 2301 , at the decoder side, the processor 1620 can obtain a block vector (BV) prediction value of a current block in an intra block copy (IBC) mode.

[0331] In step 2302 , at the decoder side, the processor 1620 can receive one or more syntax elements to obtain multiple precisions of a BV prediction value.

[0332] In step 2303 , at the decoder side, the processor 1620 can determine whether to obtain the BV difference of the current block based on multiple precisions.

[0333] In some examples, processor 1620 can receive one or more syntax elements to obtain multiple precisions of the BV prediction value by: receiving a first flag by the decoder indicating whether the current precision is greater than 0, and in response to determining that the first flag indicates that the current precision is greater than 0, obtaining a second flag by the decoder indicating whether the current precision is greater than 1.

[0334] In some examples, processor 1620 can receive one or more syntax elements to obtain multiple precisions of a BV prediction value by: receiving a first flag indicating whether a current precision is greater than 0, and in response to determining that the first flag indicates that the current precision is greater than 0, deriving a second flag indicating whether the current precision is greater than 1.

[0335] In some examples, processor 1620 may derive a second flag indicating whether the current precision is greater than 1 by deriving a second flag indicating whether the current precision is greater than 1 based on the BV difference.

[0336] In some examples, processor 1620 may derive a second flag indicating whether the current precision is greater than 1 based on a value predefined or configured for the BV difference. For example, a value of 0, 1, or other value greater than 1 may be predefined or configured to represent different motion vector (or motion vector difference) precisions (e.g., 0 represents 1-pixel precision, 1 represents 1 / 2-pixel precision, 2 represents 1 / 4-pixel precision, and 3 represents 1 / 8-pixel precision).

[0337] Figure 24 is shown corresponding to Figure 23 The video decoding method shown is a flowchart of a video encoding method. Figure 23 The method shown in is also described in the section "Signaling Motion Information".

[0338] In step 2401 , the processor 1620 on the encoder side can obtain a block vector (BV) of a current block in an intra block copy (IBC) mode.

[0339] In step 2402 , at the encoder side, the processor 1620 can signal one or more syntax elements to obtain multiple precisions of the BV prediction value.

[0340] In step 2403 , on the encoder side, the processor 1620 can determine whether to obtain the BV difference of the current block based on multiple precisions.

[0341] In some examples, processor 1620 can signal one or more syntax elements to obtain multiple precisions of a BV prediction value by: the encoder signaling a first flag indicating whether the current precision is greater than 0, and in response to determining that the first flag indicates that the current precision is greater than 0, the encoder signaling a second flag indicating whether the current precision is greater than 1.

[0342] In some examples, processor 1620 may signal a second flag indicating whether the current precision is greater than 1 based on a value predefined or configured for the BV difference. For example, values of 0, 1, or other values greater than 1 may be predefined or configured to represent different motion vector (or motion vector difference) precisions (e.g., 0 for 1-pixel precision, 1 for 1 / 2-pixel precision, 2 for 1 / 4-pixel precision, and 3 for 1 / 8-pixel precision).

[0343] Figure 25 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 25 The method shown in is also described in the section “IBC merging / AMVP motion candidate list construction”.

[0344] In step 2501 , at the decoder side, the processor 1620 can obtain a plurality of motion vector candidate lists.

[0345] In step 2502 , at the decoder side, the processor 1620 can obtain an updated motion vector candidate list by grouping a plurality of motion vector candidates in a plurality of motion vector candidate lists into different groups based on a grouping criterion.

[0346] In step 2503 , at the decoder side, the processor 1620 can obtain at least one of a group index or a candidate list index from the updated motion vector candidate list.

[0347] In step 2504 , at the decoder side, the processor 1620 can obtain a motion vector index of a motion vector of a current block for prediction based on one of a group index or a candidate list index.

[0348] In some examples, processor 1620 can obtain at least one of the group index or candidate list index from the updated motion vector candidate list by: obtaining at least one of the group index or candidate list index from the updated motion vector candidate list of the decoder; or inheriting at least one of the group index or candidate list index from the updated motion vector candidate list of a particular neighboring block.

[0349] Figure 26 is shown corresponding to Figure 25 The video decoding method shown is a flowchart of a video encoding method. Figure 26 The method shown in is also described in the section “IBC merging / AMVP motion candidate list construction”.

[0350] In step 2601 , at the encoder side, the processor 1620 can obtain a plurality of motion vector candidate lists.

[0351] In step 2602 , at the encoder side, the processor 1620 can obtain an updated motion vector candidate list by grouping multiple motion vector candidates in the multiple motion vector candidate lists into different groups based on a grouping criterion.

[0352] In step 2603 , at the encoder side, the processor 1620 can acquire at least one of a group index or a candidate list index from the updated motion vector candidate list.

[0353] In step 2604 , at the encoder side, the processor 1620 can obtain a motion vector index of a motion vector of a current block for prediction based on one of a group index or a candidate list index.

[0354] In some examples, on the encoder side, processor 1620 may obtain at least one of the group index or candidate list index from the updated motion vector candidate list by signaling at least one of the group index or candidate list index from the updated motion vector candidate list.

[0355] Figure 27 is a flowchart illustrating a video decoding method according to an example of the present disclosure. Figure 27 The method shown in is also described in the section “Combination with internal template matching”.

[0356] In step 2701 , at the decoder side, the processor 1620 can obtain at least one block vector of a current block in an intra block copy (IBC) mode or through intra template matching (ITM).

[0357] In step 2702 , at the decoder side, the processor 1620 can obtain a final prediction block based on at least one block vector and both the IBC mode and the ITM mode.

[0358] In some examples, at the encoder side, processor 1620 may obtain an initial block vector of the current block in IBC mode, and refine the initial block vector through ITM to obtain a final prediction block.

[0359] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining an ITM block vector based on ITM, and generating an IBC merge candidate list or an IBC AMVP candidate list by using the ITM block vector.

[0360] In some examples, processor 1620 may generate an IBC merge candidate list or an IBC AMVP candidate list using the ITM block vectors in one of the following ways: saving one or more block vectors of spatially adjacent neighboring blocks or non-spatially adjacent neighboring blocks of the current block; saving the block vectors of the ITM-encoded neighboring blocks in a historical motion vector table; or saving the block vectors of the ITM-encoded neighboring blocks in a local cache of the encoder for reuse in fractional motion search.

[0361] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining a first block vector in IBC mode and obtaining a first prediction block based on the first block vector; and obtaining a second block vector through ITM and obtaining a second prediction block based on the second block vector. Furthermore, processor 1620 may obtain a final prediction block based on the at least one block vector and the IBC mode and ITM by: performing a weighted average of the first prediction block and the second prediction block to obtain the final prediction block.

[0362] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining one or more first block vectors through IBC mode and obtaining one or more first prediction blocks based on the one or more first block vectors; and obtaining one or more second block vectors through ITM and obtaining one or more second prediction blocks based on the one or more second block vectors. Furthermore, processor 1620 may obtain a final prediction block based on the at least one block vector and both IBC mode and ITM by: deriving weights for the one or more first prediction blocks and the one or more second prediction blocks using a matching cost value or a least squares method; and obtaining the final prediction block by weighted averaging the one or more first prediction blocks and the one or more second prediction blocks based on the weights. For example, multiple prediction blocks may be generated separately (e.g., N>1), wherein M of the N prediction blocks (e.g., M is less than or equal to N) may be generated by intra block copying, and S of the N prediction blocks (e.g., S is less than or equal to N) may be generated by intra template matching prediction. When combining N prediction blocks (e.g., N>1), the value of the weight can be derived by using a matching cost value (e.g., one example of matching cost calculation can be based on an L-shaped template, and a higher matching cost value can indicate a lower weight value, while a lower matching cost value can indicate a higher weight value) or a least squares method.

[0363] Figure 28 is shown corresponding to Figure 27 The video decoding method shown is a flowchart of a video encoding method. Figure 28 The method shown in is also described in the section “Combination with internal template matching”.

[0364] In step 2701 , at the encoder side, the processor 1620 can obtain at least one block vector of a current block in an intra block copy (IBC) mode or through intra template matching (ITM).

[0365] In step 2702 , at the encoder side, the processor 1620 can obtain a final prediction block based on at least one block vector and both the IBC mode and the ITM mode.

[0366] In some examples, at the encoder side, processor 1620 may obtain an initial block vector of the current block in IBC mode, and refine the initial block vector through ITM to obtain a final prediction block.

[0367] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining an ITM block vector based on ITM, and generating an IBC merge candidate list or an IBC AMVP candidate list by using the ITM block vector.

[0368] In some examples, processor 1620 may generate an IBC merge candidate list or an IBC AMVP candidate list using the ITM block vectors in one of the following ways: saving one or more block vectors of spatially adjacent neighboring blocks or non-spatially adjacent neighboring blocks of the current block; saving the block vectors of the ITM-encoded neighboring blocks in a historical motion vector table; or saving the block vectors of the ITM-encoded neighboring blocks in a local cache of the encoder for reuse in fractional motion search.

[0369] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining a first block vector in IBC mode and obtaining a first prediction block based on the first block vector; and obtaining a second block vector through ITM and obtaining a second prediction block based on the second block vector. Furthermore, processor 1620 may obtain a final prediction block based on the at least one block vector and the IBC mode and ITM by: performing a weighted average of the first prediction block and the second prediction block to obtain the final prediction block.

[0370] In some examples, processor 1620 may obtain at least one block vector of the current block in IBC mode or through ITM by: obtaining one or more first block vectors through IBC mode and obtaining one or more first prediction blocks based on the one or more first block vectors; and obtaining one or more second block vectors through ITM and obtaining one or more second prediction blocks based on the one or more second block vectors. Furthermore, processor 1620 may obtain a final prediction block based on the at least one block vector and both IBC mode and ITM by: deriving weights for the one or more first prediction blocks and the one or more second prediction blocks using a matching cost value or a least squares method; and obtaining the final prediction block by weighted averaging the one or more first prediction blocks and the one or more second prediction blocks based on the weights. For example, multiple prediction blocks may be generated separately (e.g., N>1), wherein M of the N prediction blocks (e.g., M is less than or equal to N) may be generated by intra block copying, and S of the N prediction blocks (e.g., S is less than or equal to N) may be generated by intra template matching prediction. When combining N prediction blocks (e.g., N>1), the value of the weight can be derived by using a matching cost value (e.g., one example of matching cost calculation can be based on an L-shaped template, and a higher matching cost value can indicate a lower weight value, while a lower matching cost value can indicate a higher weight value) or a least squares method.

[0371] In some examples, a video encoding apparatus is provided. The apparatus includes a processor 1620 and a memory 1640, wherein the memory 1640 is configured to store instructions executable by the processor; wherein the processor is configured to perform the following when executing the instructions: Figures 17-28 Any method shown.

[0372] In some embodiments, a non-transitory computer-readable storage medium is further provided, including, for example, a plurality of programs in the memory 1630 that can be executed by the processor 1620 in the computing environment 1610 to implement the above-mentioned method, and / or store a bit stream generated by the above-mentioned encoding method or a bit stream to be decoded by the above-mentioned decoding method. In one example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to (for example, from Figure 2A The video encoder 20 in the computing environment 1610 receives a bit stream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned decoding method according to the received bit stream or data stream. In another example, the multiple programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-mentioned encoding method to encode the video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bit stream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to Figure 2B Alternatively, a non-transitory computer readable storage medium may store the bit stream or data stream generated by the encoder (e.g., Figure 2A The video encoder 20 in FIG. 1 generates a video signal for use by a decoder (eg, Figure 2B A bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) used by the video decoder 30 in the video decoder 30 in the video decoder 30 when decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0373] In an embodiment, a bit stream generated by the above encoding method or a bit stream to be decoded by the above decoding method is provided. In an embodiment, a bit stream including coded video information generated by the above encoding method or coded video information to be decoded by the above decoding method is provided.

[0374] In an embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 having stored therein a plurality of programs that can be executed by the one or more processors, wherein the one or more processors are configured to perform the above-described method when executing the plurality of programs.

[0375] In an embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, wherein the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method. In an embodiment, a computer program product is also provided, including, for example, a plurality of programs in a memory 1630, which can be executed by a processor 1620 in a computing environment 1610 to perform the above method. For example, the computer program product can include a non-transitory computer-readable storage medium.

[0376] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0377] In an embodiment, a method for storing a bitstream is further provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the encoding method or encoded video information to be decoded by the decoding method.

[0378] In an embodiment, a method for transmitting a bit stream generated by the above encoder is also provided. In an embodiment, a method for receiving a bit stream to be decoded by the above decoder is also provided.

[0379] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.

[0380] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual circumstances. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined, or deleted according to actual needs.

[0381] The examples are chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present disclosure is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present disclosure.

Claims

1. A video decoding method, comprising: The decoder obtains fractional motion information of the current block in intra block copy (IBC) mode, where the fractional motion information is signaled by the encoder and is determined by: searching for a first number of integers BV having a minimum distortion cost; applying half-pixel refinement around each of the first number of integer BVs by obtaining a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates the second number of half-pixel BV differences having the lowest rate-distortion cost; Obtaining quarter-pixel refinement results by applying quarter-pixel refinement around the second number of best half-pixel positions to each of the first number of integers BV; as well as Obtaining the fractional motion information based on the quarter-pixel refinement result; The decoder obtains a final BV of the current block based on the fractional motion information; as well as The decoder obtains a final prediction block of the current block based on the final BV.

2. The method of claim 2 , wherein obtaining the second number of best half-pixel positions for each of the first number of integer BVs comprises: selecting a third number of integers BV from said first number of integers BV; as well as A fourth number of half-pixel positions is obtained, wherein the fourth number is equal to a product of the second number and the third number.

3. The method of claim 2 , wherein obtaining the quarter-pixel refinement result by applying quarter-pixel refinement around the second number of best half-pixel positions to each of the first number of integer BVs comprises: obtaining a fifth number of quarter-pixel positions for each of the fourth number of half-pixel positions; Obtaining a sixth number of best quarter-pixel positions of all candidate positions, wherein the number of all candidate positions is equal to the product of the second number, the third number, and the fifth number; as well as A best quarter-pixel position with minimum rate distortion is selected from the sixth number of best quarter-pixel positions.

4. A video encoding method, comprising: The encoder determines, for a current block in an intra block copy (IBC) mode, a first number of integers BV having a minimum distortion cost; the encoder applying half-pixel refinement around each of the first number of integer BVs by obtaining a second number of best half-pixel positions for each of the first number of integer BVs, wherein the second number of best half-pixel positions indicates the second number of half-pixel BV differences having the lowest rate-distortion cost; The encoder obtains a quarter-pixel refinement result by applying quarter-pixel refinement around the second number of best half-pixel positions to each of the first number of integer BVs; The encoder obtains fractional motion information based on the quarter-pixel refinement result; as well as The encoder encodes the current block based on the fractional motion information.

5. The method of claim 4 , wherein obtaining the second number of best half-pixel positions for each of the first number of integer BVs comprises: selecting a third number of integers BV from said first number of integers BV; as well as A fourth number of half-pixel positions is obtained, wherein the fourth number is equal to a product of the second number and the third number.

6. The method of claim 4 , wherein obtaining the quarter-pixel refinement result by applying quarter-pixel refinement around the second number of best half-pixel positions to each of the first number of integer BVs comprises: obtaining a fifth number of quarter-pixel positions for each of the fourth number of half-pixel positions; Obtaining a sixth number of best quarter-pixel positions of all candidate positions, wherein the number of all candidate positions is equal to the product of the second number, the third number, and the fifth number; as well as A best quarter-pixel position with minimum rate distortion is selected from the sixth number of best quarter-pixel positions.

7. A video decoding method, comprising: The decoder obtains the fractional motion information of the current block in intra-block copy (IBC) mode. The decoder obtains a plurality of block vectors BV of the current block based on the fractional motion information; The decoder obtains a plurality of motion compensated prediction blocks associated with the plurality of BVs; as well as The decoder obtains a final prediction block of the current block by performing weighted averaging on the multiple motion compensated prediction blocks.

8. The method of claim 7, wherein obtaining the multiple BVs of the current block based on the fractional motion information comprises: The multiple BVs are obtained from the same candidate list or different candidate lists.

9. The method of claim 8, wherein obtaining the multiple BVs of the current block based on the fractional motion information comprises one of the following: Obtain the multiple BVs from the same IBC merge candidate list or the same advanced motion vector prediction AMVP list; Obtaining the plurality of BVs from both the IBC merge candidate list and the AMVP list; or The multiple BVs are obtained based on intra-frame template matching (ITM).

10. A video encoding method, comprising: The encoder obtains the fractional motion information of the current block in intra-block copy (IBC) mode. The encoder obtains a plurality of block vectors BV of the current block based on the fractional motion information; The encoder obtains a plurality of motion compensated prediction blocks associated with the plurality of BVs; as well as The encoder obtains a final prediction block of the current block by performing weighted averaging on the multiple motion compensation prediction blocks.

11. The method of claim 10 , wherein obtaining the plurality of block vectors of the current block based on the fractional motion information comprises: The multiple BVs are obtained from the same candidate list or different candidate lists.

12. The method of claim 11 , wherein obtaining the multiple BVs of the current block based on the fractional motion information comprises one of the following: Obtain the multiple BVs from the same IBC merge candidate list or the same advanced motion vector prediction AMVP list; Obtaining the plurality of BVs from both the IBC merge candidate list and the AMVP list; or The multiple BVs are obtained based on intra-frame template matching (ITM).

13. A video decoding method, comprising: The decoder obtains one or more block vectors based on the fractional motion information for a current block in an intra block copy (IBC) mode; The decoder calculates a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion; as well as The decoder reorders the one or more block vectors according to the template-based distortion cost.

14. The method of claim 13 , wherein the decoder calculating the template-based distortion cost for the one or more block vectors based on determining whether the one or more block vectors include the non-zero fractional portion comprises: In response to determining that the first block vector does not include a non-zero fractional portion, the decoder performs a fractional motion-based interpolation process on the first block vector; or In response to determining that the second block vector does not include a non-zero fractional portion, the decoder determines not to perform the fractional motion-based interpolation process on the second block vector.

15. The method of claim 14 , wherein the decoder calculating the template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion comprises: The decoder calculates the template-based distortion cost for each fractional motion refinement position of the one or more block vectors based on a determination of whether each fractional motion refinement position includes a non-zero fractional portion.

16. A video encoding method, comprising: The encoder obtains one or more block vectors based on the fractional motion information for a current block in an intra block copy (IBC) mode; The encoder calculates a template-based distortion cost for the one or more block vectors based on a determination of whether the one or more block vectors include a non-zero fractional portion; as well as The encoder reorders the one or more block vectors according to the template-based distortion cost.

17. The method of claim 16, wherein the encoder calculating the template-based distortion cost for the one or more block vectors based on determining whether the one or more block vectors include the non-zero fractional portion comprises: In response to determining that the first block vector does not include a non-zero fractional portion, the encoder performs a fractional motion-based interpolation process on the first block vector; or In response to determining that the second block vector does not include a non-zero fractional portion, the encoder determines not to perform the fractional motion-based interpolation process on the second block vector.

18. The method of claim 17, wherein the encoder calculating the template-based distortion cost for the one or more block vectors based on determining whether the one or more block vectors include a non-zero fractional portion comprises: The encoder calculates the template-based distortion cost for each fractional motion refinement position of the one or more block vectors based on determining whether each fractional motion refinement position includes a non-zero fractional portion.

19. A video decoding method, comprising: The decoder obtains the block vector BV prediction value of the current block in intra-frame block copy IBC mode; The decoder receives one or more syntax elements to obtain a plurality of precisions of the BV prediction value; as well as The decoder determines whether to obtain a BV difference of the current block based on the multiple precisions.

20. The method of claim 19, wherein the decoder receives the one or more syntax elements to obtain the plurality of precisions of the BV prediction value comprises: The decoder receives a first flag indicating whether the current precision is greater than 0; as well as In response to determining that the first flag indicates that the current precision is greater than 0, the decoder obtains a second flag indicating whether the current precision is greater than 1.

21. The method of claim 19, wherein the decoder receives the one or more syntax elements to obtain the plurality of precisions of the BV prediction value comprises: The decoder receives a first flag indicating whether the current precision is greater than 0; as well as In response to determining that the first flag indicates that the current precision is greater than 0, the decoder derives a second flag indicating whether the current precision is greater than 1.

22. The method of claim 21 , wherein the decoder deriving the second flag indicating whether the current precision is greater than 1 comprises: The decoder derives the second flag indicating whether the current precision is greater than 1 based on the BV difference.

23. The method of claim 22, wherein the decoder deriving the second flag indicating whether the current precision is greater than 1 based on the BV difference comprises: The decoder derives the second flag indicating whether the current precision is greater than 1 based on a value predefined or configured for the BV difference.

24. A video encoding method, comprising: The encoder obtains the block vector BV prediction value of the current block in intra-frame block copy IBC mode; The encoder signals one or more syntax elements to obtain a plurality of precisons for the BV prediction value; as well as The encoder determines whether to obtain a BV difference of the current block based on the multiple precisions.

25. The method of claim 24, wherein the encoder signals the one or more syntax elements to obtain the plurality of precisions for the BV prediction value comprises: The encoder transmits a first flag with a signal indicating whether the current precision is greater than 0; as well as In response to determining that the first flag indicates that the current precision is greater than 0, the encoder signals a second flag indicating whether the current precision is greater than 1.

26. The method of claim 25, wherein the encoder signaling the second flag indicating whether the current precision is greater than 1 comprises: The encoder signals the second flag indicating whether the current precision is greater than 1 based on a value predefined or configured for the BV difference.

27. A video decoding method, comprising: The decoder obtains multiple motion vector candidate lists; The decoder obtains an updated motion vector candidate list by grouping a plurality of motion vector candidates in the plurality of motion vector candidate lists into different groups based on a group criterion; The decoder obtains at least one of a group index or a candidate list index from the updated motion vector candidate list; and The decoder obtains a motion vector index of a motion vector of a current block for prediction based on one of the group index or the candidate list index.

28. The method of claim 27, wherein obtaining at least one of the group index or the candidate list index from the updated motion vector candidate list comprises: obtaining, from an encoder, at least one of the group index or the candidate list index from the updated motion vector candidate list; or At least one of the group index or the candidate list index from the updated motion vector candidate list is inherited from a specific neighboring block.

29. A video encoding method, comprising: The encoder obtains multiple motion vector candidate lists; The encoder obtains an updated motion vector candidate list by grouping the plurality of motion vector candidates in the plurality of motion vector candidate lists into different groups based on a group criterion; The encoder obtains at least one of a group index or a candidate list index from the updated motion vector candidate list; as well as The encoder obtains a motion vector index of a motion vector of a current block for prediction based on one of the group index or the candidate list index.

30. The method of claim 29, wherein obtaining at least one of the group index or the candidate list index from the updated motion vector candidate list comprises: At least one of the group index or the candidate list index from the updated motion vector candidate list is signaled.

31. A video decoding method, comprising: The decoder obtains at least one block vector of the current block in an intra block copy (IBC) mode or by intra template matching (ITM); as well as The decoder obtains a final prediction block based on the at least one block vector and both the IBC mode and the ITM.

32. The method according to claim 31, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtaining an initial block vector of the current block in the IBC mode; as well as Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: The final prediction block is obtained by refining the initial block vector using ITM.

33. The method of claim 31 , wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM comprises: Obtaining an ITM block vector based on the ITM; as well as An IBC merge candidate list or an IBC AMVP candidate list is generated by using the ITM block vector.

34. The method of claim 33, wherein generating the IBC merge candidate list or the IBC AMVP candidate list by using the ITM block vector comprises one of the following: Saving one or more block vectors of spatially adjacent neighboring blocks or non-spatially adjacent neighboring blocks of the current block; Save the block vectors of the ITM-coded neighboring blocks in the historical motion vector table; or The block vectors of ITM-coded neighboring blocks are saved in the encoder's local cache for reuse in fractional motion search.

35. The method according to claim 31, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtaining a first block vector through the IBC mode, and obtaining a first prediction block based on the first block vector; as well as Obtaining a second block vector through the ITM, and obtaining a second prediction block based on the second block vector; Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: The final prediction block is obtained by performing weighted averaging on the first prediction block and the second prediction block.

36. The method according to claim 31, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtain one or more first block vectors through the IBC mode, and based on the one or more first block vectors A block vector obtains one or more first prediction blocks; as well as One or more second block vectors are obtained by the ITM, and based on the one or more second block vectors vector to obtain one or more second prediction blocks; Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: deriving weights of the one or more first prediction blocks and the one or more second prediction blocks using a matching cost value or a least squares method; as well as The final prediction block is obtained by performing weighted averaging on the one or more first prediction blocks and the one or more second prediction blocks based on the weights.

37. A video encoding method, comprising: The encoder obtains at least one block vector of the current block in an intra block copy (IBC) mode or by intra template matching (ITM); as well as The encoder obtains a final prediction block based on the at least one block vector and both the IBC mode and the ITM.

38. The method according to claim 37, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtaining an initial block vector of the current block in the IBC mode; as well as Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: The final prediction block is obtained by refining the initial block vector using ITM.

39. The method of claim 37, wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM comprises: Obtaining an ITM block vector based on the ITM; as well as An IBC merge candidate list or an IBC AMVP candidate list is generated by using the ITM block vector.

40. The method of claim 39, wherein generating the IBC merge candidate list or the IBC AMVP candidate list by using the ITM block vector comprises one of the following: Saving one or more block vectors of spatially adjacent neighboring blocks or non-spatially adjacent neighboring blocks of the current block; Save the block vectors of the ITM-coded neighboring blocks in the historical motion vector table; or The block vectors of ITM-coded neighboring blocks are saved in the encoder's local cache for reuse in fractional motion search.

41. The method according to claim 37, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtaining a first block vector through the IBC mode, and obtaining a first prediction block based on the first block vector; as well as Obtaining a second block vector through the ITM, and obtaining a second prediction block based on the second block vector; Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: The final prediction block is obtained by performing weighted averaging on the first prediction block and the second prediction block.

42. The method according to claim 37, Wherein obtaining the at least one block vector of the current block in the IBC mode or through the ITM includes: Obtain one or more first block vectors through the IBC mode, and based on the one or more first block vectors A block vector obtains one or more first prediction blocks; as well as One or more second block vectors are obtained by the ITM, and based on the one or more second block vectors vector to obtain one or more second prediction blocks; Wherein obtaining the final prediction block based on the at least one block vector and both the IBC mode and the ITM mode comprises: deriving weights of the one or more first prediction blocks and the one or more second prediction blocks using a matching cost value or a least squares method; as well as The final prediction block is obtained by performing weighted averaging on the one or more first prediction blocks and the one or more second prediction blocks based on the weights.

43. A video decoding device comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein, when executing the instructions, the one or more processors are configured to perform the method according to any one of claims 1-3, 7-9, 13-15, 19-23, 27-28 and 31-36.

44. A video encoding apparatus, comprising: one or more processors; as well as a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, Wherein, when executing the instructions, the one or more processors are configured to perform the method according to any one of claims 4-6, 10-12, 16-18, 24-26, 29-30 and 37-42.

45. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1-3, 7-9, 13-15, 19-23, 27-28, and 31-36.

46. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 4-6, 10-12, 16-18, 24-26, 29-30, and 37-42.

47. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to any one of claims 1-3, 7-9, 13-15, 19-23, 27-28 and 31-36.

48. A non-transitory computer-readable storage medium for storing a bitstream generated by the method of any one of claims 4-6, 10-12, 16-18, 24-26, 29-30, and 37-42.