Method and apparatus for intra block copy and intra template matching
By improving the methods of intra-block copying and template matching prediction, candidates are obtained from multiple non-adjacent neighboring blocks. Combined with historical motion vector prediction values, the intra-block copying and template matching prediction are optimized, solving the problem of low efficiency in the existing technology and achieving more efficient video data compression and quality improvement.
Patent Information
- Application Number
- CN202480026510.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-18
- Filing Date
- 2024-04-18
- Publication Date
- 2025-12-12
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in terms of intra-frame block copying and intra-frame template matching prediction, making it difficult to effectively utilize redundant information in video data for efficient compression.
By improving the intra-block copying method in the video encoding and decoding process, IBC candidates are obtained from multiple non-adjacent neighboring blocks using scanning rules. By combining historical motion vector prediction values and intra-prediction modes, intra-block copying and template matching prediction are optimized, and a list of most probable modes is constructed to improve prediction accuracy and efficiency.
It improves the efficiency of intra-frame block copying and intra-frame template matching prediction in the video encoding and decoding process, reduces the bit rate requirement of video data, and improves video quality and compression effect.
Smart Images

Figure CN121128180A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 460,297, filed April 18, 2023, entitled “Methods and Devices for IntraBlock Copy and Intra Template Matching”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to video encoding and decoding and compression, and specifically, but not limited to, methods and apparatus for improving the encoding and decoding efficiency of intra block copy (IBC) and intra template matching prediction (intra TMP). Background Technology
[0003] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data into a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0004] This disclosure provides examples of techniques related to improving intra-frame block copying methods in video encoding or decoding processes.
[0005] According to a first aspect of this disclosure, a video decoding method is provided. The method includes: obtaining one or more IBC candidates from a plurality of non-adjacent neighboring blocks according to a scanning rule corresponding to the IBC mode, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block; and obtaining a prediction of the current block by the decoder applying the IBC mode based on the one or more IBC candidates.
[0006] According to a second aspect of this disclosure, a video decoding method is provided. The method includes: obtaining one or more IBC candidates from a plurality of neighboring blocks based on the size of the current block by a decoder for applying an intra-block copy (IBC) mode to a current block; and obtaining a prediction of the current block by the decoder applying the IBC mode based on the one or more IBC candidates.
[0007] According to a third aspect of this disclosure, a video decoding method is provided. The method includes: obtaining a list of historical motion vector prediction values (HMVPs) by a decoder based on at least one block attribute of a current block; and obtaining a prediction of the current block by the decoder applying an intra-block copy (IBC) mode based on the HMVP list.
[0008] According to a fourth aspect of this disclosure, a video decoding method is provided. The method includes: in response to determining that neighboring blocks of a current block are encoded and decoded in either an IBC mode or an intra template matching prediction (TMP) mode, obtaining by a decoder an intra prediction mode associated with a block vector of the neighboring blocks; and constructing by the decoder a list of intra most probable modes (MPMs) for the current block based on the intra prediction modes.
[0009] According to a fifth aspect of this disclosure, a video coding method is provided. The method includes: obtaining one or more IBC candidates from a plurality of non-adjacent neighboring blocks according to a scanning rule corresponding to the IBC mode, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block; and obtaining a prediction of the current block by the encoder applying the IBC mode based on the one or more IBC candidates.
[0010] According to a sixth aspect of this disclosure, a video coding method is provided. The method includes: obtaining one or more IBC candidates from a plurality of neighboring blocks based on the size of the current block by an encoder for applying an intra-block copy (IBC) mode to a current block; and obtaining a prediction of the current block by the encoder by applying the IBC mode based on the one or more IBC candidates.
[0011] According to a seventh aspect of this disclosure, a video coding method is provided. The method includes: obtaining a list of historical motion vector prediction values (HMVPs) by an encoder based on at least one block attribute of a current block; and obtaining a prediction of the current block by the encoder by applying an intra-block copy (IBC) mode based on the HMVP list.
[0012] According to an eighth aspect of this disclosure, a video coding method is provided. The method includes: in response to determining that neighboring blocks of a current block are encoded and decoded in either an IBC mode or an intra template matching prediction (TMP) mode, obtaining an intra prediction mode associated with a block vector of the neighboring blocks by an encoder; and constructing an intra most probable mode (MPM) list for the current block by the encoder based on the intra prediction mode.
[0013] According to a ninth aspect of this disclosure, a video decoding apparatus is provided. The apparatus may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, when executing the instructions, are configured to perform a method according to the first, second, third, or fourth aspect.
[0014] According to a tenth aspect of this disclosure, a video encoding apparatus is provided. The apparatus may include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors, when executing the instructions, are configured to perform a method according to a fifth, sixth, seventh, or eighth aspect.
[0015] According to the eleventh aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect, the second aspect, the third aspect, or the fourth aspect.
[0016] According to a twelfth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the fifth, sixth, seventh, or eighth aspect.
[0017] According to the thirteenth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by a method according to the first aspect, the second aspect, the third aspect, or the fourth aspect.
[0018] According to the fourteenth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by the method according to the fifth, sixth, seventh or eighth aspect. Attached Figure Description
[0019] A more specific description of the examples disclosed herein will be presented by reference to the specific examples illustrated in the accompanying drawings. Given that these drawings depict only a few examples and should therefore not be considered as limiting the scope, these examples will be described and explained in more specific and detailed manner using the accompanying drawings.
[0020] Figure 1 A block diagram of an exemplary system for encoding and decoding video blocks, according to some examples of this disclosure, is shown.
[0021] Figure 2 A block diagram of an exemplary video encoder according to some examples of this disclosure is shown.
[0022] Figure 3 A block diagram of an exemplary video decoder according to some examples of this disclosure is shown.
[0023] Figure 4A - 4E shows a block diagram illustrating, according to some examples of this disclosure, how a frame can be recursively partitioned into multiple video blocks of different sizes and shapes.
[0024] Figure 5 A diagram showing the locations of spatial candidates according to some examples of this disclosure is provided.
[0025] Figure 6 A diagram showing candidate pairs considered for redundancy checks of spatial candidates, according to some examples of this disclosure.
[0026] Figure 7 A scaled graph of motion vectors for time candidates according to some examples of this disclosure is shown.
[0027] Figure 8 A diagram showing the candidate positions of time candidates according to some examples of this disclosure is provided.
[0028] Figure 9 A graph showing a method for searching points using the merged mode of motion vector difference (MMVD) according to some examples of this disclosure is shown.
[0029] Figure 10 The following are examples of unidirectional predictive motion vector selection for Geometric Partitioning Pattern (GPM) according to this disclosure.
[0030] Figure 11The top neighbor block and left neighbor block used in CIIP weight derivation are shown in some examples according to this disclosure.
[0031] Figure 12 The current CTU processing order and its available reference points in the current CTU and the left CTU are shown in some examples according to this disclosure.
[0032] Figure 13 The following are examples of filling candidates for replacing zero vectors in the IBC list, according to some examples of this disclosure.
[0033] Figure 14 The reference region of the IBC is shown when the CTU(m,n) is encoded or decoded. According to some examples of this disclosure, dotted shaded blocks represent the current CTU(m,n); shaded blocks filled with " / " represent reference regions; and no shaded blocks represent invalid reference regions.
[0034] Figure 15 An IBC reference area of camera-captured content is shown as some examples according to this disclosure.
[0035] Figures 16A-16B The following are examples of methods for dividing angular patterns according to this disclosure.
[0036] Figures 17A-17D The GPM using inter-frame and intra-frame prediction is shown. Figures 17A-17C Available IPM candidates are shown. Figure 17D Examples of GPM utilizing intra-frame and intra-frame prediction are shown according to some examples of this disclosure.
[0037] Figure 18 The edges on the template are shown as some examples according to this disclosure.
[0038] Figure 19 The intra-frame template matching search region is shown in some examples according to this disclosure.
[0039] Figure 20 Templates for template-matching OBMCs are shown as examples of some of the features provided in this disclosure.
[0040] Figure 21 The present disclosure illustrates some examples of intra-coded block partitioning methods and corresponding weights for angular and planar modes.
[0041] Figure 22 Templates for intra-frame coded blocks and their reference samples are shown as examples of some of the features of this disclosure.
[0042] Figure 23 Templates for IBC coded blocks and their reference samples are shown as examples of some of the features of this disclosure.
[0043] Figure 24 Exemplary non-adjacent blocks for IBC AMVP candidates or IBC merge candidates are shown in accordance with some examples of this disclosure.
[0044] Figure 25 shows non-adjacent neighboring blocks of different sizes according to some examples of this disclosure: (a) neighboring blocks of the same size as the current block; (b) neighboring blocks of different sizes (e.g., 4 × 4 or 8 × 8) of the current block.
[0045] Figure 26 An example of spatially non-adjacent candidate spatial neighbor blocks for deriving IBC patterns is shown, according to some examples of this disclosure, wherein the numbers in the non-adjacent neighbor blocks represent the scan order.
[0046] Figure 27 Another example of spatially non-adjacent candidate spatial neighbor blocks for deriving IBC patterns is shown, according to some examples of this disclosure, wherein the numbers in the non-adjacent neighbor blocks indicate the scan order and the numbers after the arrows indicate the degree values of the angles.
[0047] Figure 28 Examples of some of the examples according to this disclosure are shown, in which non-adjacent areas of space are confined within half the size of the current CTU above and to the left.
[0048] Figure 29 illustrates motion storage of spatially non-adjacent neighbors (IBC-adjacent CUs or non-IBC-adjacent CUs) according to some examples of this disclosure: (a) the permissible spatially non-adjacent area beyond the current CTU and (b) motion storage in the line buffer (A is an IBC CU; B is a non-IBC CU).
[0049] Figure 30 Examples of non-adjacent neighboring locations projected / cropped when the scanned non-adjacent neighboring locations exceed the permissible space area are shown according to this disclosure.
[0050] Figure 31 Another example of a non-adjacent neighboring location projected / truncated when the scanned non-adjacent neighboring location exceeds the allowed space area (e.g., beyond the current CTU and available line buffer) is shown, according to some examples of this disclosure.
[0051] Figure 32 The granularity of IBC motion storage, as shown in some examples according to this disclosure, differs from the minimum IBC block size.
[0052] Figure 33An example of a sub-block-based IBC mode according to some examples of this disclosure is shown, wherein the BV of a sub-block in the current block is obtained by reusing the BV of a sub-block of a co-block in a co-picture.
[0053] Figure 34 An example of a sub-block-based IBC mode is shown, where the BV of the left or upper sub-block in the current block is obtained by refining the BV of the current block using a template matching method.
[0054] Figure 35 A diagram illustrating a computing environment coupled to a user interface, according to some examples of this disclosure, is shown.
[0055] Figure 36 A plot of the ramp function based on the displacement (d) from the predicted sample location to the GPM partition boundary and the size of the mixed region (τ) is shown according to some examples of this disclosure.
[0056] Figure 37 shows a diagram of spatial GPM candidates according to some examples of this disclosure.
[0057] Figure 38 A diagram of a GPM template according to some examples of this disclosure is shown.
[0058] Figure 39 A graph of GPM mixing according to some examples of this disclosure is shown.
[0059] Figure 40 Additional directions along the diagonal angle of k × π / 8 are shown in some examples according to this disclosure (position 4010 is used in the anchor point).
[0060] Figure 41 Another example of spatially non-adjacent candidate spatial neighbor blocks for deriving IBC patterns is shown, according to some examples of this disclosure, wherein the numbers in the non-adjacent neighbor blocks represent the scan order.
[0061] Figure 42 A flowchart of a video decoding method according to some examples of this disclosure is shown.
[0062] Figure 43 Examples of such disclosure are shown. Figure 42 The flowchart shows the video encoding method corresponding to the video decoding method shown.
[0063] Figure 44 A flowchart of a video decoding method according to some examples of this disclosure is shown.
[0064] Figure 45 Examples of such disclosure are shown. Figure 44The flowchart shows the video encoding method corresponding to the video decoding method shown.
[0065] Figure 46 A flowchart of a video decoding method according to some examples of this disclosure is shown.
[0066] Figure 47 Examples of such disclosure are shown. Figure 46 The flowchart shows the video encoding method corresponding to the video decoding method shown.
[0067] Figure 48 A flowchart of a video decoding method according to some examples of this disclosure is shown.
[0068] Figure 49 Examples of such disclosure are shown. Figure 47 The flowchart shows the video encoding method corresponding to the video decoding method shown. Detailed Implementation
[0069] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. Numerous non-limiting details are set forth in the following detailed description to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many classes of electronic devices with digital video capabilities.
[0070] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular forms “a / an,” “the,” and “the” in this disclosure and the appended claims are also intended to include the plural forms unless otherwise expressly indicated throughout the disclosure. It should also be understood that the term “and / or” as used in this disclosure refers to and includes one or any one or all possible combinations of the listed related items.
[0071] Throughout this specification, references to "an embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise expressly stated, the features, structures, elements, or characteristics described in connection with one or more embodiments also apply to other embodiments.
[0072] Throughout this disclosure, the terms “first,” “second,” “third,” etc., are used as a nomenclature and are used only to refer to relevant elements, such as equipment, components, parts, steps, etc., and do not imply any spatial or temporal order unless otherwise expressly stated. For example, “first equipment” and “second equipment” can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be named arbitrarily.
[0073] The terms "module," "submodule," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "subunit" can include memory (shared, dedicated, or grouped) that stores code or instructions executable by one or more processors. A module can include one or more circuits, with or without stored code or instructions. A module or circuit can include one or more components that are directly or indirectly connected. These components may or may not be physically attached to each other or adjacent to each other.
[0074] As used herein, depending on the context, the terms "if" or "when" can be understood to mean "at the time of" or "in response to". If these terms appear in the claims, they may not indicate that the relevant limitation or feature is conditional or optional. For example, a method may include the steps of: i) performing a function or action X' when or if condition X exists, and ii) performing a function or action Y' when or if condition Y exists. The method may simultaneously have the capability to perform both function or action X' and function or action Y'. Therefore, functions X' and Y' may be performed at different times in multiple executions of the method.
[0075] Units or modules can be implemented purely in software, purely in hardware, or a combination of both. For example, in a purely software implementation, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.
[0076] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.
[0077] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, switch, base station, or any other means that may facilitate communication from source device 12 to target device 14.
[0078] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.
[0079] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.
[0080] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be sent directly to the target device 14 via the output interface 22 of the source device 12. Alternatively, the encoded video data can be stored on the storage device 32 for later access by the target device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or transmitter.
[0081] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0082] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0083] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0084] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in the hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0085] In some implementations, components of source device 12 (e.g., video source 18, video encoder 20, or the following references) Figure 1 The components described include at least a portion of the components in the video encoder 20 and the output interface 22) and / or the components of the target device 14 (e.g., the input interface 28, the video decoder 30, or the following references). Figure 3At least a portion of the components included in the video decoder 30 and the display device 34 described herein can operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), where the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 not included in the cloud computing service network may be located in one or more client devices, and these client devices may communicate with server computers in the cloud computing service network via wireless communication networks (e.g., cellular communication networks, short-range wireless communication networks, or Global Navigation Satellite System (GNSS) communication networks) or wired communication networks (e.g., local area network (LAN) communication networks or power line communication (PLC) networks). In embodiments, at least a portion of the operations described herein may be implemented as cloud-based services provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” are used interchangeably as appropriate. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service networks. Rather, this disclosure may be implemented in any other type of computing environment currently known or developed in the future.
[0086] Figure 2 This is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".
[0087] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that, regarding the CCSAO technology, this application is not limited to the embodiments described herein, but can also be applied to situations where an offset is selected for any other component among the luminance, Cb, and Cr chrominance components based on any one of the luminance, Cb, and Cr chrominance components to modify that other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein can be any one of the luminance, Cb, and Cr chrominance components, the second component mentioned herein can be any other one of the luminance, Cb, and Cr chrominance components, and the third component mentioned herein can be the remaining components among the luminance, Cb, and Cr chrominance components. In some examples, the loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or can be distributed among one or more of the described fixed or programmable hardware units.
[0088] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding modes) when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.
[0089] like Figure 2As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of samples with sample values. Samples in the array may also be referred to as pixels or image elements (pel). The number of samples in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of samples with sample values, but its dimension is smaller than that of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. By iteratively using, for example, QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof, a video block can be further segmented into one or more block partitions or sub-blocks (which can again form blocks). It should be noted that the term "block" or "video block" as used herein can refer to a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring to, for example, HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.
[0090] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.
[0091] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.
[0092] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.
[0093] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0094] The motion estimation unit 42 calculates the motion vector for a video block in an inter-frame predictive coding frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1), each of the first and second reference frame lists identifying one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.
[0095] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation unit 42. Upon receiving motion vectors for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel differences forming the residual video block may include luminance component differences or chrominance component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.
[0096] In some implementations, the intra-BC unit 48 can generate vectors and acquire prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-BC unit 48 can, for example, use various intra-prediction modes to encode the current block during individual encoding passes and test their performance through rate-distortion analysis. Next, the intra-BC unit 48 can select a suitable intra-prediction mode from the various tested intra-prediction modes to use and generate an intra-mode indicator accordingly. For example, the intra-BC unit 48 can use rate-distortion analysis to calculate rate-distortion values for the various tested intra-prediction modes and select the intra-prediction mode with the best rate-distortion characteristics from the tested modes as the suitable intra-prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-frame prediction mode exhibits the optimal rate-distortion value for the block.
[0097] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction according to the embodiments described herein. In any case, for intra-frame block copying, in terms of pixel differences, the predicted block may be a block considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values for sub-integer pixel positions.
[0098] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values of the predicted block from the pixel values of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.
[0099] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra-block copy prediction performed by the intra-BC unit 48, the intra-prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. The intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.
[0100] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.
[0101] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0102] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1 The video decoder 30 shown, or archived in, for example Figure 1 The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.
[0103] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values for use in motion estimation.
[0104] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.
[0105] Figure 3 This is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 2 The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.
[0106] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).
[0107] Video data memory 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding modes) when decoding video data. Video data memory 79 and DPB 92 can be formed of any memory device from a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown in... Figure 3 The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip along with other components of the video decoder 30, or off-chip relative to those components.
[0108] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.
[0109] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-predictive unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-predictive mode transmitted by the signal and reference data from the previous decoded block of the current frame.
[0110] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.
[0111] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.
[0112] Motion compensation unit 82 and / or intra-frame prediction (BC) unit 85 determine prediction information for video blocks in the current video frame by parsing motion vectors and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) used to encode video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the reference frame list for the frame, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.
[0113] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.
[0114] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.
[0115] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0116] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1 On the display device 34).
[0117] In a typical video encoding and decoding process, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and therefore consist of only a two-dimensional array of luminance samples.
[0118] like Figure 4AAs shown, the video encoder 20 (or more specifically, the partitioning unit in the predictive processing unit of the video encoder 20) generates a coded representation of a frame by first dividing the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for the luma sample, two corresponding coding tree blocks for the chroma sample, and syntax elements for encoding the samples of the coding tree block. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block can be an N×N sample block.
[0119] To achieve better performance, the video encoder 20 can recursively perform tree segmentation on the coding tree blocks of the CTU, such as binary tree segmentation, ternary tree segmentation, quadtree segmentation, or combinations thereof, and divide the CTU into smaller CUs. Figures 4B-4E are block diagrams illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes according to some embodiments of this disclosure. Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image can include, for example, two corresponding coded blocks (CBs) of luminance and chrominance samples of the same size frame, as well as syntax elements for encoding the samples of the coded blocks. In monochrome images or images with three separate color planes, a CU can include a single coded block and a syntax structure for encoding the samples of the coded block. It should be noted that... Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / ternary / binary partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and ternary tree structures. Figure 4E As shown, a coded block with width W and height H has five possible segmentation types: quad segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation.
[0120] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB for luma samples, two corresponding PBs for chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.
[0121] Video encoder 20 can generate prediction blocks for a PU using intra-frame prediction or inter-frame prediction. If video encoder 20 uses intra-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for a PU, then video encoder 20 can generate prediction blocks for a PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0122] After the video encoder 20 generates predicted luminance blocks, predicted Cb blocks, and predicted Cr blocks for one or more PUs of the CU, the video encoder 20 can generate luminance residual blocks for the CU by subtracting the predicted luminance blocks of the CU from the original luminance coding blocks of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0123] In addition, such as Figure 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of the luminance samples, two corresponding transform blocks of the chrominance samples, and syntax elements for transforming the transform block samples. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.
[0124] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0125] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.
[0126] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding samples of the predicted blocks of the PU for the current CU to corresponding samples of the transformed blocks of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.
[0127] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.
[0128] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "Motion Vector Prediction" (MVP) of the current CU) by exploring their spatial and temporal correlations.
[0129] Instead of the above combination Figure 2 The method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit 42, into the video bitstream, and subtracting the predicted motion vector value of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0130] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to encode and decode the current CU using the same motion vector prediction value from the motion vector candidate list.
[0131] Generally, the basic inter-frame prediction scheme used in VVC is almost the same as that in HEVC, except that several prediction tools have been further extended, added and / or improved, such as extended merge prediction, MMVD and GPM.
[0132] Extended merge forecast With continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vector of the current image has also increased significantly. One way to overcome this challenge is to use motion information (e.g., motion vectors) from spatially neighboring CUs, temporally co-located CUs, etc., of the current CU as an approximation (e.g., prediction) of the motion information of the current CU, also known as the "Motion Vector Prediction Value (MVP)" of the current CU.
[0133] Similar to the process of selecting a prediction block from a reference picture during inter-frame prediction of a coded block, both video encoder 20 and video decoder 30 need to employ a set of rules to construct an MVP candidate list for the current CU, and then select an MVP candidate from the MVP candidate list as the MVP of the current CU. This avoids the need to transmit the MVP candidate list itself between video encoder 20 and video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for both video encoder 20 and video decoder 30 to use the same MVP candidate selected from the MVP candidate list for encoding and decoding the current CU.
[0134] In VVC, the MVP candidate list is constructed by including the following five types of MVPs in order: — Spatial MVP (i.e., spatial candidate) from spatially neighboring CUs; — The time MVP (i.e., time candidate) from the time-parallel CU. — A history-based MVP (HMVP) derived from a first-in, first-out (FIFO) table; — Paired average MVP; and — Zero MVP.
[0135] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded in merge mode, the index of the best MVP candidate is encoded using truncated unary binarization. The first binary bit of the index is encoded and decoded using the context, and the remaining binary bits of the index are encoded and decoded using bypass.
[0136] The derivation process for each type of MVP is provided below. Like HEVC, VVC also supports parallel derivation of the MVP candidate list for all CUs within a certain size region.
[0137] Deriving MVP from Spatial Candidates In addition to swapping the positions of the first two spatial candidates, VVC also selects from spatial candidates (e.g., Figure 5 The MVP derivation for the CU (which is adjacent to the current CU101) is the same as that in HEVC. From the location... Figure 5 Up to four spatial candidates are selected from the spatial candidates for the positions depicted (i.e., top position B0, left position A0, top right position B1, bottom left position A1, and top left position B2). The derivation is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because said one or more CUs belong to other stripes or tiles) or are intra-coded.
[0138] After adding the CU at position B0 as a candidate to the merged candidate list, a redundancy check is needed for adding the remaining candidates to the merged candidate list. This ensures that candidates with the same motion information are excluded from the merged candidate list, thereby improving encoding and decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only... Figure 6 Pairs are linked using arrowed lines, and a candidate is added to the merged candidate list only if the motion information of the candidate in the pair used for redundancy checking is different from the motion information of the candidate to be added. The spatial MVP derived from the candidates in the merged candidate list is added to the MVP candidate list.
[0139] MVP Derivation from Time Candidates During the MVP derivation from the time candidate, only one time candidate is added to the merge candidate list. Specifically, when deriving the MVP from that time candidate, it is based on whether it belongs to the current CU (e.g., Figure 7 The corresponding image of curr_CU 303 in (e.g., Figure 7 The time candidate CU (e.g., col_pic 302) in the col_pic 302) is the same as the time candidate CU (e.g., Figure 7 The scaled motion vector is derived using col_CU 301 and added as a temporal MVP candidate to the MVP candidate list. The list of reference images and their indices used to derive the co-located CUs are explicitly transmitted via signals in the strip header. Figure 7 As shown, the scaled motion vector is obtained (i.e., scaled) from the motion vector of the co-located CU using the picture order count (POC) distance (i.e., tb and td), where tb is defined as the current picture (e.g., Figure 7 Reference image for curr_ref 304 (e.g., Figure 7 The difference between curr_ref305 in the image and the current image in the POC, and td is defined as the reference image of the co-located image (e.g., Figure 7 The POC difference between col_ref 306 and the corresponding image. The reference image index for the time candidate is set to zero.
[0140] Select the position of the current time candidate (i.e., the corresponding CU) in CU 401 between positions C0 and C1, such as... Figure 8 As described. If the CU at position C0 in the co-occurring frame is unavailable, intra-coded, or outside the current CTU line, then the CU at position C1 is used as the co-occurring CU to derive the temporal MVP candidate. Otherwise, the CU at position C0 is used as the co-occurring CU to derive the temporal MVP candidate.
[0141] Derivation of HMVP Candidates HMVP candidates are added to the MVP candidate list after the spatial MVP and temporal MVP. Motion information of previously encoded blocks is stored in the HMVP table and used as the MVP of the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-coded CU exists, the associated motion information is added to the last entry of the HMVP table as a new HMVP candidate.
[0142] The size of the HMVP table is set to 6. When a new HMVP candidate is inserted into the HMVP table, a constrained FIFO rule is used, where a redundancy check is first applied to find if a duplicate HMVP already exists in the HMVP table. If found, the duplicate HMVP is removed from the HMVP table, all subsequent HMVP candidates are shifted forward, and the duplicate HMVP is added to the last entry in the HMVP table.
[0143] HMVP candidates can be used in the MVP candidate list construction process. The most recent few HMVP candidates in the HMVP table are checked sequentially and inserted into the MVP candidate list after the temporal MVP candidates. Redundancy checks are applied to the HMVP candidates relative to spatial and / or temporal MVP candidates.
[0144] To reduce the number of redundant check operations, the following simplifications are introduced: — Perform a redundancy check on the last two entries in the HMVP table relative to the spatial MVP candidates derived from the spatial candidates at positions A1 and B1, respectively; and — Once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1, the process of building the MVP candidate list from the HMVP candidates is terminated.
[0145] Derivation of Paired Average MVP Candidates Pairwise averaged MVP candidates are generated by averaging the MVPs derived using the first two merge candidates from an existing merge candidate list using a predefined pair. The first merge candidate in the predefined pair can be defined as p0Cand, and the second merge candidate in the predefined pair can be defined as p1Cand. For each list of reference images, the average motion vector is calculated separately based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available for a list of reference images, they are averaged even if they point to different reference images, and the reference image of the averaged motion vector is set as the reference image of p0Cand; if only one motion vector is available for a list of reference images, that motion vector is used directly; if no motion vector is available for a list of reference images, the motion vectors and reference image indices for that list remain invalid.
[0146] Zero MVP If the MVP candidate list is not full after adding pairwise average MVP candidates, insert zero MVP at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0147] MMVD As described above, in merge mode, motion information (i.e., MVP candidates) is implicitly derived from the MVP candidate list constructed for the current CU and directly used as the MV of the current CU to generate predicted samples for the current CU. This may result in a certain error between the actual MV of the current CU and the implicitly derived MVP. To improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, where the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. After transmitting the regular merge flag, the MMVD flag is signaled to specify whether the MMVD mode is used for the current CU.
[0148] In MMVD mode, after selecting an MVP candidate from the first two MVP candidates in the MVP candidate list, MMVD information is transmitted by signaling. The MMVD information includes an MMVD candidate flag for specifying which of the first two MVP candidates is selected as the basis for the MV, a distance index for indicating the motion amplitude information of the MVD, and a direction index for indicating the motion direction information of the MVD.
[0149] The distance index indicating the motion amplitude information of the specified MVD is compared with the reference image of the current CU pointed to by the selected MVP candidate (e.g., Figure 9 The starting point in the L0 reference image 501 or L1 reference image 503 (e.g., by) Figure 9 The dashed circles in the table represent predefined offsets, and MVD can be derived from these offsets and added to selected MVP candidates. Table 1 below specifies the relationship between distance indices and predefined offsets.
[0150]
[0151] Table 1 The direction index specifies the sign of the MVD, which represents the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined signs. In some examples, the meaning of the MVD sign can vary depending on the information of the selected MVP candidate. When the selected MVP candidate is a non-predictive MV or a bidirectional predictive MV (where both MVs point to the same side of the current image) (i.e., the POC of both reference images of the current image (e.g., the reference images of List 0 and List 1, which are also referred to as the L0 reference image and L1 reference image, respectively) is greater than the POC of the current image, or both are less than the POC of the current image), the signs in Table 2 specify the signs to be added to the MVD of the selected MVP candidate. When the selected MVP candidate is a bidirectional prediction MV (where both MVs point to different sides of the current image) (i.e., the POC of one reference image of the current image is greater than the POC of the current image, and the POC of the other reference image of the current image is less than the POC of the current image), if the POC distance of the L0 reference image (i.e., the POC distance between the L0 reference image and the current image) is greater than the POC distance of the L1 reference image (i.e., the POC distance between the L1 reference image and the current image), then the symbols in Table 2 specify the symbols to be added to List 0 MVP (MVP0) MVD (MVD0) of the selected MVP candidate, and the symbols to be added to List 1 MVP (MVP1) MVD (MVD1) of the selected MVP candidate are opposite to the symbols in Table 2; otherwise, if the POC distance of the L1 reference image is greater than the POC distance of the L0 reference image, then the symbols in Table 2 specify the symbols to be added to MVD1 of MVP1, and the symbols to be added to MVD0 of MVP0 are opposite to the symbols in Table 2.
[0152]
[0153] Table 2 MVD is scaled based on the POC distance. If the POC distances of the L0 and L1 reference images are the same, then MVD does not need to be scaled. Otherwise, if the POC distance of the L0 reference image is greater than the POC distance of the L1 reference image, then MVD1 is scaled. If the POC distance of the L1 reference image is greater than the POC distance of the L0 reference image, then MVD0 is scaled.
[0154] GPM In VVC, GPM is supported for inter-frame prediction. GPM is signaled using CU-level flags as one merging mode; other merging modes include regular merging, MMVD, CIIP, and sub-block merging. For each possible CU size... ( in, (excluding 8) 64 and 64 8) GPM supports a total of 64 partitions.
[0155] When using GPM, the CU is divided into two parts by a geometrically positioned straight line. The position of the dividing line is mathematically derived based on the angle and offset parameters of the specific partition. Inter-frame prediction is performed using the motion of each part of the CU obtained through geometric partitioning; and each partition allows only unidirectional prediction, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, each CU requires only two motion-compensated predictions.
[0156] If GPM is used for the current CU, then the geometric partition index (indicating the angle and offset of the geometric partition) and two merge indexes (one for each partition) are further transmitted using signal transmission.
[0157] The unidirectional prediction candidate list is derived directly from the merging candidate list constructed according to the extended merging prediction process described above. Let n denote the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector of the nth merging candidate in the merging candidate list (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for GPM. These motion vectors are... Figure 10 The symbol is marked with "x". If the corresponding LX motion vector of the nth merge candidate in the merge candidate list does not exist, the L(1 - X) motion vector of the same merge candidate is used as the unidirectional predicted motion vector of GPM.
[0158] CIIP In VVC, when encoding and decoding a CU in merge mode, if the CU contains at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64) and if both the width and height of the CU are less than 128 luma samples, an additional flag indicating whether CIIP mode is applied to the current CU is transmitted via signaling. In CIIP mode, the prediction signal is obtained by combining the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal in CIIP mode is derived using the same inter-frame prediction process applied in regular merge mode; and the intra-frame prediction signal in CIIP mode is derived after utilizing the regular intra-frame prediction process in planar mode. Then, a weighted average is used to combine the intra-frame prediction signal and the inter-frame prediction signal, where, according to (e.g.) Figure 11 (As shown) The weight values are calculated based on the encoding patterns of the top and left neighboring blocks of the current CU 1601 as follows: — If the top neighboring block is available and intra-coded, isIntraTop is set to 1; otherwise, isIntraTop is set to 0. — If the left neighboring block is available and has been intra-coded, isIntraLeft is set to 1; otherwise, isIntraLeft is set to 0. — If (isIntraLeft + isIntraTop) equals 2, then set the weight value to 3; — Otherwise, if (isIntraLeft + isIntraTop) equals 1, then set the weight value to 2; — Otherwise, set the weight value to 1.
[0159] — Prediction signal in CIIP mode The derivation is as follows:
[0160] in, It is the inter-frame prediction signal in CIIP mode. It is an intra-frame prediction signal in CIIP mode. It represents the weight value, and >> indicates a right shift operation.
[0161] Intra-block copying in Universal Video Codec (VVC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. IBC significantly improves the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of the IBC-encoded CU has integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-prediction or inter-prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.
[0162] On the encoder side, hash-based motion estimation is performed on the IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-merging modes, a block vector search is first performed using a hash-based search. If the hash search does not return any valid candidates, a local search based on block matching is performed.
[0163] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4 × 4 sub-blocks. For larger current blocks, a hash key is determined to match the hash key of a reference block when all hash keys of all 4 × 4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference with the minimum cost is selected.
[0164] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.
[0165] At the CU level, IBC mode is transmitted using flag signals, and it can be transmitted as either IBCAMVP mode or IBC skip / merge mode: IBC Skip / Merge Mode: The merge candidate index is used to indicate which block vectors from the list of neighboring candidate IBC encoded blocks are used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and paired candidates.
[0166] IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left neighboring block and one from the upper neighboring block (if IBC encoding and decoding is used). When neither neighboring block is available, the default block vector is used as the prediction value. A flag indicating the index of the block vector prediction value is transmitted via signal transmission.
[0167] IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of predefined regions that include the current CTU region and a certain region of the left CTU. Figure 12 The reference area for the IBC mode is shown, where each block represents a 64 × 64 luminance sample unit.
[0168] Based on the current location of the encoding CU within the current CTU, the following applies: If the current block falls within the top-left 64 × 64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, using CPR mode, the current block can also reference reference samples in the bottom-right 64 × 64 block of the left CTU. Using CPR mode, the current block can also reference reference samples in the bottom-left 64 × 64 block of the left CTU and the top-right 64 × 64 block of the left CTU.
[0169] If the current block falls within the upper right 64 × 64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, the current block can also reference the reference samples in the lower left 64 × 64 block and the lower right 64 × 64 block of the left CTU when the brightness position (0, 64) has not yet been reconstructed relative to the current CTU. Otherwise, the current block can also reference the reference samples in the lower right 64 × 64 block of the left CTU.
[0170] If the current block falls within the lower left 64 × 64 block of the current CTU, then in CPR mode, besides the samples already reconstructed in the current CTU, the current block can also reference reference samples in the upper right and lower right 64 × 64 blocks of the left CTU, provided the brightness position (64, 0) has not yet been reconstructed relative to the current CTU. Otherwise, in CPR mode, the current block can also reference reference samples in the lower right 64 × 64 block of the left CTU.
[0171] If the current block falls within the lower right 64 × 64 block of the current CTU, then CPR mode is used, and the current block can only reference samples that have already been reconstructed in the current CTU.
[0172] This restriction allows the IBC mode to be implemented using local on-chip memory in a hardware implementation.
[0173] Interaction between IBC and other codec tools The interaction between IBC mode and other inter-frame coding / decoding tools in VVC (such as Paired Merge Candidate, History-Based Motion Vector Prediction (HMVP), Combined Intra / Inter-Frame Prediction Mode (CIIP), Merge Mode Utilizing Motion Vector Difference (MMVD), and Geometric Partitioning Mode (GPM)) is as follows: IBC can be used with pairwise merge candidates and HMVP. New pairwise IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC movements are inserted into a history buffer for future reference.
[0174] IBC cannot be used in combination with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM.
[0175] When using DUAL_TREE partitions, IBC is not allowed for chroma-coded blocks.
[0176] Unlike in HEVC screen content codec extensions, the current image is no longer included as one of the reference images in the IBC prediction reference image list 0. The derivation of motion vectors in IBC mode excludes all neighboring blocks in inter-frame mode, and vice versa. The following IBC design aspects are applied: IBC uses the same process as regular MV merging, including pairwise merging candidates and historical motion predictions, but does not allow TMVP and zero vectors because they are invalid for IBC mode.
[0177] Separate HMVP buffers (5 candidates each) are used for traditional MV and IBC.
[0178] Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream, and that merging should not be used if a merging candidate is invalid (out of range or 0). This bitstream consistency constraint is represented by a virtual buffer, as described below.
[0179] For deblocking, IBC is treated as an inter-frame mode.
[0180] If the current block is encoded or decoded using IBC prediction mode, AMVR does not use quarter-pixels; instead, AMVR is signaled to indicate only whether the MV is inter-pel or 4-integer-pel.
[0181] The number of IBC merge candidates can be transmitted separately in the strip header from the number of regular merge candidates, sub-block merge candidates, and geometric merge candidates.
[0182] The concept of a virtual buffer is used to describe the permissible reference region and effective block vector of the IBC prediction mode. Representing the CTU size as ctbSize, the virtual buffer ibcBuf has a width of wIbcBuf = 128 × 128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128 × 128, the size of ibcBuf is also 128 × 128; for a CTU size of 64 × 64, the size of ibcBuf is 256 × 64; and for a CTU size of 32 × 32, the size of ibcBuf is 512 × 32.
[0183] The size of the VPDU is min(ctbSize, 64) in each dimension, W v = min(ctbSize, 64).
[0184] The virtual IBC buffer ibcBuf is maintained as follows.
[0185] When you begin decoding each CTU line, refresh the entire ibcBuf with an invalid value of -1.
[0186] When starting to decode the VPDU (xVPDU, yVPDU) relative to the top left corner of the image, set ibcBuf[x][y] = -1, where x = xVPDU%wIbcBuf, …, xVPDU%wIbcBuf + W v - 1;y = yVPDU%ctbSize, …,yVPDU%ctbSize + W v - 1.
[0187] After decoding, the CU contains (x, y) values relative to the top-left corner of the image, set... ibcBuf[ x % wIbcBuf ][ y % ctbSize ]= recSample[ x ][ y ] For a block covering coordinates (x, y), if for the block vector bv = (bv[0], bv[1]) The block is valid if the following conditions are met; otherwise, it is invalid: ibcBuf[ (x + bv[0])% wIbcBuf][ (y + bv[1]) % ctbSize ] should not be equal to -1.
[0188] Intra-block copying in Enhanced Compression Model (ECM) In ECM, IBC has been improved in the following aspects.
[0189] IBC Merge / AMVP List Construction The IBC merge / AMVP list build modifications are as follows: An IBC merge / AMVP candidate can only be inserted into the IBC merge / AMVP candidate list if it is valid.
[0190] Candidates in the upper right, lower left, and upper left spaces, as well as a pairwise average candidate, can be added to the IBC merge / AMVP candidate list.
[0191] Apply template-based adaptive reordering (ARMC-TM) to the IBC merge list.
[0192] The HMVP table size for IBC was increased to 25. After deriving up to 20 IBC merge candidates through full pruning, they were reordered together. Following reordering, the top 6 candidates with the lowest template matching cost were selected as the final candidates in the IBC merge list.
[0193] The zero vector candidate used to populate the IBC merge / AMVP list is replaced with a set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in IBC merge mode, and therefore, it is discarded as a BVP in the IBC candidate list.
[0194] Three candidates are located at the nearest corner of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as follows. Figure 13 As depicted in the text.
[0195] Block vector candidates for intra-frame TMP derivation for IBC In this method, the block vector (BV) derived from IntraTMP (Intra-Temporal Template Matching Prediction) is used for Intra-Block Copy (IBC). The stored IntraTMP BVs of neighboring blocks, together with the IBC BVs, are used as spatial BV candidates in the construction of the IBC candidate list.
[0196] The IntraTMP block vector is stored in the IBC block vector buffer, and the current IBC block can use both the IBCBV and the IntraTMP BV of neighboring blocks as BV candidates for the IBC BV candidate list. The IntraTMP block vector is added as a spatial candidate to the IBC block vector candidate list.
[0197] IBC using template matching Template matching is used in both IBC merge mode and IBC AMVP mode in IBC.
[0198] Compared to the list used in the regular IBC merge mode, the IBC-TM merge list has been modified to select candidates based on a pruning method, with the same motion distance between candidates as in the regular TM merge mode. The zero motion supplement at the end is replaced with motion vectors for the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0199] In IBC-TM merging mode, the selected candidates are refined using a template matching method before the RDO or decoding process. IBC-TM merging mode competes with the regular IBC merging mode and uses a signal transmission TM-merging flag.
[0200] In IBC-TM AMVP mode, a maximum of three candidates are selected from the IBC-TM merge list. Each of these three selected candidates is refined using a template matching method and ranked according to its resulting template matching cost. Then, during motion estimation, only the first two are considered as usual.
[0201] Template matching refinement for IBC-TM merging mode and AMVP mode is very simple because the IBC motion vector is constrained to (i) integers and (ii) within the reference region, such as Figure 12 As shown. Therefore, in IBC-TM merge mode, all thinning is performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision depending on the AMVR value. This thinning only accesses samples that are not interpolated. In both cases, the thinning motion vector and the template used in each thinning step must adhere to the constraints of the reference region.
[0202] IBC Reference Area The IBC reference area extends to the two CTU rows above. Figure 14 The reference region used for encoding CTU(m, n) is shown. Specifically, for the CTU(m, n) to be encoded, the reference region includes CTUs with indices (m–2, n–2)…(W, n–2), (0, n–1)…(W, n–1), (0, n)…(m, n), where W represents the maximum horizontal index within the current tile, strip, or image. This setting ensures that IBC does not require additional memory on the current ETM platform when the CTU size is 128. The per-sample block vector search (or local search) range is limited to: horizontal direction [–(C<<1), C>>2], vertical direction [–C, C>>2] to accommodate the expansion of the reference region, where C represents the CTU size.
[0203] IBC merging mode using block vector difference The ECM employs an IBC merging mode utilizing block vector differences. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.
[0204] Basic candidates are selected from the top five candidates in the reordered IBC merge list. Then, all possible MBVD refinement positions (20 × 4) for each basic candidate are reordered based on the SAD cost between the template (the row above the current block and the column to the left of the current block) and its reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are reserved as available positions and thus used for MBVD index encoding.
[0205] IBC adaptation for camera-captured content When adapting to IBC for camera-captured content, the IBC reference range decreases from 2 CTU lines to 2 × 128 lines, such as Figure 15 As shown. On the encoder side, to reduce complexity, the local search range is set to a horizontal [-8,8] and a vertical [-8,8] range centered on the first block vector prediction value of the current CU. This encoder modification is not applicable to SCC sequences.
[0206] CIIP combined with TIMD and TM In CIIP mode, prediction samples are generated by weighting the inter-prediction signals obtained by combining candidate predictions using CIIP-TM and the intra-prediction signals obtained by using intra-prediction modes derived using TIMD. This method is only applicable to coded blocks with an area of 1024 or less.
[0207] The TIMD derivation method is used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value is selected from the TIMD mode list and mapped to one of 67 regular intra-prediction modes.
[0208] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights of these two tests (wIntra, wInter) should be modified. For near-horizontal mode (2 <= angle mode index < 34), the current block is as follows: Figure 16A The block is shown as being vertically divided; for near-vertical mode (34 <= angle mode index <= 66), the current block is as follows: Figure 16B The diagram shows a horizontal division.
[0209] Table 3 shows the different sub-blocks (wIntra, wInter).
[0210]
[0211] Table 3. Modified weights for angle mode.
[0212] Using CIIP-TM, a CIIP-TM merge candidate list is established for the CIIP-TM pattern. The merge candidates are refined through template matching. The CIIP-TM merge candidates are also reordered into regular merge candidates using the ARMC method. The maximum number of CIIP-TM merge candidates is two.
[0213] Multiple Hypothesis Prediction (MHP) In multi-hypothesis inter-frame prediction mode, in addition to the traditional bidirectional prediction signal, one or more additional motion-compensated prediction signals are transmitted via signal transmission. The resulting overall prediction signal is obtained by sample-by-sample weighted superposition. The bidirectional prediction signal is utilized. and the first additional inter-frame prediction signal / hypothesis The predicted signal is obtained as follows. : (2) Weighting factor The new syntax element add_hyp_weight_idx specifies the weight based on the mapping presented in Table 4:
[0214] Table 4. add_hyp_weight_idx and The mapping between them.
[0215] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0216] (3) The resulting overall prediction signal is obtained as the final (i.e., has the largest index) n of Within this mode, up to two additional prediction signals can be used (i.e., n Limited to 2).
[0217] The motion parameters for each additional prediction hypothesis can be explicitly signaled by specifying a reference index, a motion vector prediction value index, and a motion vector difference, or implicitly signaled by specifying a merging index. A separate multiple hypothesis merging flag distinguishes between these two signaling modes.
[0218] For inter-frame AMVP mode, MHP is applied only when unequal weights are selected in BCW in bidirectional prediction mode.
[0219] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0220] Geometric Partitioning (GPM) in ECM GPM using combined motion vector difference (MMVD) The GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. First, a flag for the GPM CU is transmitted via signal transmission to specify whether a mode is used. If the mode is used, each geometric partition of the GPM CU can further determine whether to use signal transmission MVD. If signal transmission MVD is used for a geometric partition, the motion of the partition is further refined using the signal transmission MVD information after selecting a GPM merging candidate. All other procedures remain the same as in the GPM.
[0221] Similar to MMVD, MVD is transmitted as a pair of distance and direction signals. GPM (GPM-MMVD) utilizing MMVD involves nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 16 pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, MVD is shifted left by 2 bits, as in MMVD.
[0222] GPM using template matching Template matching is applied to GPM. When GPM mode is enabled for CU, a CU-level flag is signaled to indicate whether TM is applied to both geometric partitions. TM is used to refine the motion information for each geometric partition. When TM is selected, a template is constructed using left neighbor samples, top neighbor samples, or left neighbor samples and top neighbor samples, based on the partition angles shown in Table 5. Motion is then refined by minimizing the difference between the current template and the template in the reference image using the same search mode with the merging mode having the half-pixel interpolation filter disabled.
[0223]
[0224] Table 5. Templates used for the first and second geometric partitions, where A indicates the use of the top sample point, L indicates the use of the left sample point, and L+A indicates the use of both the left and top sample points.
[0225] The GPM candidate list is constructed as follows: 1. Derive interleaved list 0 MV candidates and list 1 MV candidates directly from the regular merge candidate list, where list 0 MV candidates have higher priority than list 1 MV candidates. Apply an adaptive threshold pruning method based on the current CU size to remove redundant MV candidates.
[0226] 2. Directly derive interleaved List 1 MV candidates and List 0 MV candidates from the regular merge candidate list, where List 1 MV candidates have higher priority than List 0 MV candidates. The same pruning method using adaptive thresholds is also applied to remove redundant MV candidates.
[0227] 3. Fill the zero MV candidate list until the GPM candidate list is full.
[0228] GPM-MMVD and GPM-TM are enabled only for one GPM CU. This is achieved by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is true), the value of the GPM-TM flag is inferred to be false.
[0229] GPM using inter-frame and intra-frame prediction In GPM utilizing inter-frame and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame and intra-frame prediction samples of each GPM-separated region. Inter-frame prediction samples are obtained from inter-frame GPM, while intra-frame prediction samples are obtained from the intra-frame prediction mode (IPM) candidate list and the index transmitted from the encoder signal. The IPM candidate list size is predefined as 3. Available IPM candidates are the parallel angle mode (parallel mode) for GPM block boundaries, the vertical angle mode (vertical mode) for GPM block boundaries, and others, etc. Figures 17A to 17D The planar pattern shown. Furthermore, as... Figure 17D The GPM utilization shown is limited to reduce the signal transmission overhead of IPM and avoid increasing the size of the intra-prediction circuitry on the hardware decoder. Furthermore, direct motion vectors and IPM storage are introduced in the GPM mixing region to further improve encoding and decoding performance.
[0230] In IPM derivation based on DIMD and neighboring modes, parallel modes are registered first. Therefore, if no identical IPM candidates exist in the list, a maximum of two IPM candidates can be registered using the decoder-side intra-frame mode derivation (DIMD) method and / or neighboring block derivation. As for neighboring mode derivation, a maximum of five neighboring block locations are available, but these locations are limited by the GPM block boundary angles (as shown in Table 6 below), which have already been used to utilize template-matched GPM (GPM-TM).
[0231]
[0232] Table 6. Locations of available neighboring blocks derived from IPM candidate derivation based on the angles of the GPM block boundaries. A and L represent the top and left sides of the predicted block.
[0233] GPM-intraframe can be combined with GPM using motion vector difference merging (GPM-MMVD). TIMD is used on IPM candidates within GPM-intraframe to further improve encoding / decoding performance. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates for neighboring blocks.
[0234] Template matching-based GPM segmentation pattern reordering In template-match-based GPM segmentation pattern reordering, given the motion information of the current GPM block, the corresponding TM generation value of the GPM segmentation pattern is calculated. Then, all GPM segmentation patterns are reordered in ascending order based on their TM generation values. Instead of sending GPM segmentation patterns, a signaling method using Golomb-Rice codes is used to indicate the exact index of the GPM segmentation pattern within the reordering list.
[0235] The GPM partition reordering method is a two-step process performed after generating the corresponding reference templates for the two GPM partitions in the coding unit, as follows: • Extend the GPM partition edge to the reference templates of the two GPM partitions to generate 64 reference templates, and compute the corresponding TM cost for each of the 64 reference templates. • The TM generation values based on the GPM splitting pattern are reordered in ascending order, and the top 32 are marked as available splitting patterns.
[0236] like Figure 18 As shown, the edges on the template extend from the edges of the current CU, but the GPM blending process does not apply to the template region on that edge.
[0237] After reordering in ascending order using TM cost, the index is transmitted using semaphores.
[0238] Utilizing an adaptive blending geometric partitioning model (GPM) In VVC, the final predicted samples are generated by mixing the predictions of the two predicted signals using a weighted average. Two integer mixing matrices (W0 and W1) are used. The weights in the GPM mixing matrix are derived from the ramp function based on the displacement from the predicted sample location to the GPM partition boundary. The mixing region size is fixed at two (two samples on each side of the GPM partition boundary).
[0239] The blending process in ECM is improved by adding four additional blending region sizes (one-quarter, half, twice, and four times the existing region size), such as Figure 36 As shown, the CU-level flag, encoded by signal transmission, represents the selected mixing region size. Furthermore, extended weighted precision is utilized, where the maximum weighted value changes from 8 (in VVC) to 32 to accommodate the extended mixing region size.
[0240] Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode of an inter-frame coding / decoding tool similar to GPM, where two prediction parts are generated based on the intra-frame prediction process. In this mode, a candidate list is built, where each entry contains one partitioning mode and two intra-frame prediction modes, as shown in Figure 37. 26 partitioning modes and 3 intra-frame prediction modes are used to form combinations. The candidate list length is set to 16. The selected candidate index is transmitted via signaling.
[0241] Reorder the list using a template. Figure 38 The SAD (Search Allocation) between the template's prediction and reconstruction is used for sorting. In one example, the template size is fixed at 1. In some other examples, the template size can be set differently.
[0242] For each partition mode, the same intra-to-inter-frame GPM list derivation is used to derive the IPM list for each partition. The IPM list size is set to 3. In this list, the TIMD derivation mode is replaced with two derivation modes with horizontal and vertical orientations.
[0243] The SGPM mode is applied to restricted block sizes: 4 <= width <= 64, 4 <= height <= 64, width < height × 8, height < width × 8, width × height >= 32.
[0244] Adaptive blending has also been used in spatial GPM, where, Figure 39 The mixing depth τ shown is derived as follows: If min(width, height) == 4, then choose 1 / 2τ; Otherwise, if min(width, height) == 8, then choose τ; Otherwise, if min(width, height) == 16, then choose 2τ; Otherwise, if min(width, height) == 32, then choose 4τ; Otherwise, choose 8τ.
[0245] Intra-frame template matching Intra-frame template matching prediction (intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame that matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and performs the same prediction operation on the decoder side.
[0246] By comparing the L-shaped causal neighbors of the current block with Figure 19The prediction signal is generated by matching another block in a predefined search region, which consists of the following: R1: Current CTU R2: Top Left CTU R3: Above CTU R4: Left CTU The sum of absolute differences (SAD) is used as the cost function.
[0247] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0248] The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportionally to the block dimensions (BlkW, BlkH) to ensure a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a × BlkW SearchRange_h = a × BlkH in,' ' is a constant that controls the trade-off between gain and complexity. In fact, ' 'Equals 5.'
[0249] For CUs with width and height dimensions less than or equal to 64, enable the intra-frame template matching tool. This maximum CU size used for intra-frame template matching is configurable.
[0250] When DIMD is not used for the current CU, the intra-frame template matching prediction mode is transmitted at the CU level using a dedicated flag.
[0251] Template-based intra-frame mode derivation (TIMD) fusion For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The top two intra-prediction modes with the minimum SATD are selected as TIMD modes. After applying the PDPC procedure, these two TIMD modes are fused with weights, and this weighted intra-prediction is used to encode the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0252] The costs of the two selected modes are compared with a threshold. In the test, the cost factor 2 is applied as follows: costMode2 < 2 × costMode1.
[0253] If the condition is true, then apply fusion; otherwise, use only mode 1.
[0254] The weights of the patterns are calculated based on their SATD costs as follows: Weight 1 = costMode2 / (costMode1 + costMode2) Weight 2 = 1 - Weight 1 Division operations are performed using the same lookup table (LUT)-based integerization scheme used by CCLM.
[0255] Local illumination compensation (LIC) LIC is an inter-frame prediction technique used to model the local illumination variation between the current block and its predicted block as a function of the local illumination variation between the current block template and the reference block template. The parameters of this function can be represented by a scaling factor α and an offset β, which form a linear equation for compensating for illumination variations (i.e., α × p[x] + β), where p[x] is the reference sample to which the MV points at position x on the reference image. When surround motion compensation is enabled, the surround offset should be considered to truncate the MV. Since α and β can be derived based on the current block template and the reference block template, they require no signaling overhead except for the LIC flag indicating the use of LIC for AMVP mode.
[0256] The following modifications are used to apply the local illumination compensation proposed in JVET-O0066 to unidirectional prediction of inter-frame CU.
[0257] Intra-frame neighbor samples can be used for LIC parameter derivation; Disable LIC for blocks with fewer than 32 luminance samples; For both non-subblock mode and affine mode, the LIC parameter derivation is performed based on the template block sample corresponding to the current CU, rather than based on the partial template block sample corresponding to the first top-left 16 × 16 element; Samples of the reference block template are generated by using MC and block MV without rounding them to integer pixel precision.
[0258] TM-based MMVD and Affine MMVD Reordering Extended MMVD offset for MMVD mode and affine MMVD mode. Figure 40The diagram illustrates additional refinement positions along a diagonal angle of k × π / 8, increasing the number of directions from 4 to 16. Next, all possible MMVD refinement positions (16 × 6) for each basic candidate are reordered based on the SAD cost between the template (the row above the current block and the column to the left of the current block) and its reference for each refinement position. Finally, the top 1 / 8 refinement positions with the minimum template SAD cost are retained as available positions and thus used for MMVD index encoding. The MMVD index is binarized using rice codes with a parameter equal to 2. The extended affine MMVD is reordered, with additional refinement positions added along a diagonal angle of k × π / 4. After reordering, the top 1 / 2 refinement positions with the minimum template SAD cost are retained.
[0259] The top N motion candidates in the candidate list are used as the base candidates for both MMVD and affine MMVD before being reordered. For MMVD, N equals 3, and for affine MMVD, N is [1, 3], depending on the affine flag of the neighboring blocks. Two ways are allowed to add MMVD offsets, including 'both sides' and 'one side', depending on whether the offset of the other reference picture list is mirrored or directly set to zero. Which method is applied to a block depends on the TM cost.
[0260] OBMC When applying OBMC, as described in JVET-L0101, motion information from neighboring blocks is used to refine the top and left boundary pixels of the CU using weighted prediction.
[0261] The following conditions should not be used for OBMC: When OBMC is disabled at SPS level When the current block has intra-frame mode or IBC mode When the current block applies a LIC When the current brightness block area is less than or equal to 32 Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using motion information from neighboring sub-blocks. It is enabled for the following sub-block-based codec tools: Affine AMVP mode; Affine merging mode and sub-block-based temporal motion vector prediction (SbTMVP). Bilateral matching based on sub-blocks.
[0262] When using OBMC mode in CIIP mode utilizing LMCS, inter-frame blending is performed before LMCS mapping of inter-frame samples. LMCS is applied to the blended inter-frame samples, which are then combined with intra-frame samples to which LMCS has been applied in CIIP mode.
[0263]
[0264] ,in, This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by the motion of neighboring blocks in the original domain, and and It's the weight.
[0265] OBMC based on template matching In the template matching-based OBMC scheme, instead of directly using weighted prediction, the predicted value of the CU boundary sample derivation method is determined based on the template matching cost, including using only the motion information of the current block, or using the motion information of neighboring blocks, or using one of the hybrid modes.
[0266] In this scheme, for each block with a size of 4 × 4 at the top CU boundary, the upper template size is equal to 4 × 1. If N If two adjacent blocks have the same motion information, the size of the template above is enlarged to 4. N × 1, because the MC operation can be processed in one go. For each left block with a size of 4 × 4 at the left CU boundary, the left stencil size is equal to 1 × 4 or 1 × 4. N ( Figure 20 ).
[0267] For each 4 × 4 top block (or N (4 × 4 blocks), derive the predicted values of the boundary samples by following these steps.
[0268] For example, take block A as the current block and take its upper neighbor block AboveNeighbor_A. The operations on the left-hand blocks are performed in the same way.
[0269] First, based on the following three types of motion information, the three template matching costs (Cost1, Cost2, Cost3) are measured by the SAD between the template reconstruction samples obtained by the MC process and their corresponding reference samples: Calculate Cost1 based on the motion information of A.
[0270] Cost2 is calculated based on the motion information of AboveNeighbor_A.
[0271] Cost3 is calculated based on the weighted prediction of the motion information of A and AboveNeighbor_A, with weighting factors of 3 / 4 and 1 / 4 respectively.
[0272] Secondly, by comparing Cost1, Cost2, and Cost3, a method is selected to calculate the final prediction result of the boundary sample points.
[0273] The original MC result using the current block ring motion information is represented as pixel 1, and the MC result using the motion information of neighboring blocks is represented as pixel 2. The final prediction result is represented as NewPixel.
[0274] If Cost1 is the minimum, then NewPixel(i,j) = Pixel1(i,j).
[0275] If (Cost2 + (Cost2>>2) + (Cost2>>3))<= Cost1, then use mixed mode 1.
[0276] For a luminance block, the number of mixed pixel rows is 4.
[0277] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5 NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3 NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4 NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5 For a chroma block, the number of mixed pixel rows is 1.
[0278] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5 If Cost1 <= Cost2, then use Mixed Mode 2.
[0279] For a luminance block, the number of mixed pixel rows is 2.
[0280] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4 NewPixel(i,1)=(31×Pixel1(i,1)+Pixel2(i,1)+16)>>5 For chroma blocks, the number of mixed pixel rows / columns is 1.
[0281] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4 Otherwise, use mixed mode 3.
[0282] For a luminance block, the number of mixed pixel rows is 4.
[0283] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3 NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4 NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5 For a chroma block, the number of mixed pixel rows is 1.
[0284] NewPixel(i,0)=(7×Pixel1(i,0)+Pixel2(i,0)+4)>>3 Currently, IBC tools are not combined with GPM tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0285] Currently, coded blocks encoded in IBC mode are not combined with coded blocks encoded in intra-frame or inter-frame modes. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0286] Currently, the weights of intra-frame and inter-frame coded blocks in CIIP are predefined in a fixed manner. Therefore, this disclosure provides an example of adaptively determining weights based on a template matching method, which can improve prediction accuracy and enhance encoding / decoding performance.
[0287] Currently, the number of block vectors (BVs) in IBC tools is singular. Therefore, this disclosure provides examples for increasing the number of block vectors (BVs) and for combining prediction results, which can improve prediction accuracy and enhance encoding / decoding performance.
[0288] Currently, coded blocks encoded using intra-frame TMP mode are not combined with coded blocks encoded using intra-frame or inter-frame modes. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0289] Currently, intra-frame TMP tools are not combined with GPM tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0290] Currently, IBC tools are not combined with TIMD tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0291] Currently, intra-frame TMP tools are not combined with TIMD tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0292] Currently, intra-frame TMP tools are not combined with LIC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0293] Currently, IBC tools are not combined with OBMC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0294] Currently, intra-frame TMP tools are not combined with OBMC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and enhance encoding / decoding performance.
[0295] Currently, the candidate derivation process for IBC merge mode and IBC AMVP mode only uses adjacent neighbor blocks and lateral non-adjacent neighbor blocks in the upper left region. Therefore, this disclosure provides a further expansion to include more non-adjacent neighbor blocks, which can improve prediction accuracy and improve encoding and decoding performance.
[0296] Currently, IBC mode typically uses block-level BV for motion compensation. Therefore, this disclosure provides a further introduction of sub-block-based IBC mode, which can improve prediction accuracy and enhance encoding / decoding performance.
[0297] Currently, TM IBC mode and TM regular inter-frame mode use both left and top templates for motion refinement. Therefore, this disclosure provides a further extended template mode, which can improve prediction accuracy and codec performance.
[0298] Currently, when obtaining boundary strength, the deblocking filter treats blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode differently. Therefore, this disclosure provides a unification of these two modes, which can improve encoding and decoding performance.
[0299] Currently, the intra-frame TMP tool is not combined with the IBC tool. Therefore, this disclosure provides a combination of these two tools in the form of GPM or CIIP, which can improve encoding and decoding performance.
[0300] Currently, similar to MMVD designs in VVC, signal transmission utilizes MMVD candidates of GPM in MMVD mode. Therefore, this disclosure provides a template matching method to reorder MMVD candidates of GPM utilizing MMVD mode, which can improve encoding and decoding performance.
[0301] Currently, HMVP candidates are utilized only based on the coding order of IBC / regular inter-frame / affine inter-frame HMVP candidates. Therefore, this disclosure provides a method to utilize HMVP candidates based on their relative position to the current coding block, which can improve encoding and decoding performance.
[0302] Currently, the intra-prediction modes pointed to by the block vectors of IBC or intra-TMP are stored in memory, so the stored intra-prediction modes can be used directly to construct the intra-MPM list.
[0303] In this disclosure, to address the problems identified above, methods are provided for further improving existing designs of IBCs. Generally, the main features of the techniques proposed in this disclosure are summarized below.
[0304] The IBC tool can be combined with the GPM tool in the following ways: using IBC and GPM predicted by IBC, using IBC and GPM predicted intra-frame, or using IBC and GPM predicted inter-frame.
[0305] As a simplified version of the combination of the IBC and GPM tools, for a predefined direction (such as 45 degrees), the upper left part is predicted using the intra-frame mode and the lower right part is predicted using the IBC mode. Then, they are averaged and weighted to obtain the final predicted signal.
[0306] The IBC tool is combined with the CIIP tool, where IBC prediction is combined with intra-frame prediction mode, or IBC prediction is combined with inter-frame prediction mode.
[0307] The template matching method adaptively determines the weights of intra-frame and inter-frame coded blocks in CIIP.
[0308] The IBC tool is combined with the MHP tool, where more than one BV prediction is obtained and they are weighted and averaged to obtain the final prediction signal.
[0309] The intra-frame TMP tool is combined with the CIIP tool, where intra-frame TMP is combined with intra-frame prediction mode, or intra-frame TMP is combined with inter-frame prediction mode.
[0310] The intra-frame TMP tool can be combined with the GPM tool in the following ways: using intra-frame TMP and GPM predicted by intra-frame TMP, using intra-frame TMP and GPM predicted by intra-frame TMP, or using intra-frame TMP and GPM predicted by inter-frame TMP.
[0311] As a simplified version of the combination of the intra-frame TMP tool and the GPM tool, for a predefined direction (such as 45 degrees), the upper left part is predicted using the intra-frame mode, the lower right part is predicted using the intra-frame TMP mode, and then they are averaged and weighted to obtain the final predicted signal.
[0312] The IBC tool is combined with the TIMD tool, where the IBC mode is used in conjunction with the intra-prediction mode in MPM for TIMD fusion.
[0313] The intra-frame TMP tool is combined with the TIMD tool, where the intra-frame TMP mode is used in conjunction with the intra-frame prediction mode in MPM for TIMD fusion.
[0314] The intra-frame TMP tool is combined with the LIC tool, where the LIC tool is used to compensate for local illumination variations between the current block and its intra-frame TMP prediction block.
[0315] The IBC tool is combined with the OBMC tool, where the OBMC tool is used to refine the top and left boundary pixels of the current block predicted by IBC.
[0316] The intra-frame TMP tool is combined with the OBMC tool, where the OBMC tool is used to refine the top and left boundary pixels of the current block predicted by the intra-frame TMP.
[0317] The candidate derivation process for the IBC merge pattern or IBCAMVP pattern is extended by using not only adjacent neighboring blocks but also non-adjacent neighboring blocks.
[0318] The IBC mode is extended to the sub-block level, where sub-blocks within the current block have their own BV for motion compensation.
[0319] The template modes of TM IBC mode and TM regular inter-frame mode are extended, with left-side template only, top-side template only, etc., being used for motion refinement.
[0320] When the boundary strength is obtained, the deblocking filter treats blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode equally.
[0321] The IBC tool can be combined with the intra-frame TMP tool. The combination can be GPM predicted using IBC and intra-frame TMP, or CIIP predicted using IBC and intra-frame TMP.
[0322] The template matching method is used to reorder the MMVD candidates of GMP using the MMVD pattern.
[0323] HMVP candidates are utilized based on their relative position to the current coding block in IBC / regular inter-frame / affine inter-frame HMVP candidates.
[0324] During the construction of the intra MPM list for the current block, when a neighboring block is IBC-coded or intra TMP-coded, the intra prediction mode pointed to by the block vector of the IBC or intra TMP is used to construct the intra MPM list for the current block.
[0325] In some examples, the disclosed methods can be applied independently or in combination.
[0326] Using IBC and IBC-predicted GPM According to one or more embodiments of this disclosure, the IBC tool and the GPM tool are combined in the form of utilizing IBC and the GPM predicted by IBC. Different methods can be used to achieve this objective.
[0327] In the first approach, both "inter-frame" portions of the GPM (Gross Frame Calculation) using inter-frame and inter-frame prediction methods in VVC are replaced with IBCs (Inter-Block Frames). This means that the combined prediction results of the two IBCs are weighted and averaged according to the dividing lines in the coded blocks. The weights can be obtained by referring to the GPM in VVC using inter-frame and inter-frame prediction methods.
[0328] In the second approach, both “inter-frame” parts of the GPM in the ECM that utilize inter-frame and inter-frame prediction methods are replaced with IBC, where template matching tools can be used to further improve encoding and decoding performance.
[0329] When the IBC tool is combined with the GPM tool in the form of a GPM utilizing both IBC and IBC-predicted IBC, the IBC prediction results can come from regular merge candidates, TM refined merge candidates, or merge candidates utilizing block vector difference (MBVD). In some examples, regular merge candidates, TM refined merge candidates, or MBVD candidates can exist in the ECM. In the first approach, the two IBC prediction results come from the same type of merge candidate. For example, both IBC prediction results come from regular merge candidates, or both from TM refined merge candidates, or both from MBVD candidates, where the merge indices of the two IBC prediction results are different. In the second approach, the two IBC prediction results come from different types of merge candidates. For example, one IBC prediction result comes from a regular merge candidate, and the other IBC prediction result comes from a TM refined merge candidate. When a TM refined merge candidate is used for a GPM utilizing both IBC and IBC, the TM refined merge candidate can be directly reused for a GPM utilizing both IBC and IBC-predicted IBC, or, similar to a GPM utilizing TM, can utilize different templates for different parts of a GPM partition with a predefined GPM partitioning pattern. When merged candidates using Block Vector Difference (MBVD) are used for GPM using IBC and IBC, the reordered IBC MBVD candidates can be directly reused for GPM using IBC and IBC prediction, or, similar to GPM using MMVD patterns, the selected MMVD candidates can be signaled using distance and orientation information. Alternatively, different templates can be used to reorder the MMVD candidates for different parts of the GPM partition of a predefined GPM segmentation pattern. When the MMVD candidates for different parts of the GPM partition are reordered, the selected MMVD candidate is represented by a signaling transfer index.
[0330] When combining IBC tools with GPM tools in the form of GPM using IBC and IBC-predicted GPM, different methods can be used to encode GPM segmentation patterns. In the first method, similar to GPM in VVC, all allowed segmentation patterns are encoded with equal probability. In the second method, all allowed GPM segmentation patterns are divided into groups, and two indices are encoded to identify the transmitted GPM segmentation pattern. The first index indicates which group is used, and the second index indicates a specific index within the selected group. For example, all allowed GPM segmentation patterns are divided into two groups, where GPM segmentation patterns along the horizontal or vertical direction form one group, and GPM segmentation patterns along other directions form another. First, a context-coded or bypass-coded flag is encoded to indicate which group the transmitted GPM segmentation pattern belongs to, and then the indices encoded with equal probability are encoded to indicate which index within the selected group the transmitted GPM segmentation pattern belongs to. In the third method, a TM-based method is used to encode GPM segmentation patterns. In one example, similar to GPM in ECM, a TM-based method is used to reorder all allowed GPM partition patterns, and then signal transmission using Golomb-Rice codes indicates the exact index of the GPM partition pattern in the reordered list. In some examples, reordering may be based on comparing template matching costs. In another example, after all allowed GPM partition patterns are divided into groups, a TM-based method is used to reorder the GPM partition patterns in the selected groups, and then signal transmission using Golomb-Rice codes indicates the exact index of the GPM partition pattern in the reordered list. In a fourth method, since both the GPM partition pattern and the two merge indices of the two GPM partitions need to be transmitted in the bitstream, similar to spatial GPM, a TM-based method is used to reorder all combinations of the GPM partition pattern and the two merge indices of the two GPM partitions, and then signal transmission using Golomb-Rice codes indicates the exact index of the combination of the GPM partition pattern and the corresponding two merge indices of the two GPM partitions in the reordered list.
[0331] When combining the IBC tool with the GPM tool in the form of GPM utilizing both IBC and IBC-predicted GPM, different methods can be used to blend the two GPM partitions. In the first method, adaptive blending is employed. For CUs encoded using GPM utilizing both IBC and IBC-predicted GPM, comparisons are made during the RDO process as follows: Figure 36The different blending widths are shown, and an index representing the selected blending width is transmitted in the bitstream. In the second method, the blending width can be selected via a predefined criterion, and then the index does not need to be transmitted in the bitstream. In one example, similar to spatial GPM, the blending width is selected based on the width and height of the current CU. In another example, similar to GPM in VVC, only one blending width is used for all CU sizes. In the third method, different blending methods are used for different types of content. For example, hard blending with a blending width of zero is used for screen content; adaptive blending is used for natural content.
[0332] When the IBC tool is combined with the GPM tool in the form of IBC and GPM predicted by IBC, different methods can be used to cover the motion information of the CU encoded using IBC and GPM predicted by IBC. In the first method, if the center of a 4 × 4 block in the current CU is located in GPM partition A, the block is filled with the block vector of the merge index of GPM partition A; the opposite is true for GPM partition B. It should be noted that in this method, partitions A and B are only divided by GPM partition lines, meaning that both partitions A and B can contain some mixed regions. In the second method, if the center of a 4 × 4 block in the current CU is located in a non-mixed region of GPM partition A or B, the block is filled with the block vector of the merge index of the corresponding GPM partition. If the center of a 4 × 4 block in the current CU is located in a mixed region of a GPM partition, the block is filled with the weighted average of the block vectors of the merge index of the two GPM partitions. In the third method, regardless of where the GPM partition line is located, the motion information of the current CU is filled entirely with block vectors from either GPM partition A or GPM partition B.
[0333] GPM using IBC and intra-frame prediction According to one or more embodiments of this disclosure, the IBC tool and the GPM tool are combined in the form of GPM utilizing IBC and intra-frame prediction. Different methods can be used to achieve this objective.
[0334] In the first method, the “inter-frame” portion of GPM in ECM, which utilizes inter-frame and intra-frame prediction methods, is replaced with IBC, where the IBC combines the prediction results with the intra-frame prediction results and takes a weighted average to obtain the final prediction signal.
[0335] Unified GPM using intra-frame and intra-frame prediction, GPM using IBC and intra-frame prediction, and GPM using IBC and IBC prediction. According to one or more embodiments of this disclosure, the codec tools utilizing intra-frame and intra-frame prediction GPM, utilizing IBC and intra-frame prediction GPM, and utilizing IBC and IBC prediction GPM are unified into a set of syntax elements.
[0336] Specifically, first, a flag indicating whether the CU is GPM-encoded is encoded into the bitstream. If this flag is true, the GPM split mode is transmitted in the bitstream. For a GPM split partition, first, a flag indicating whether the GPM split partition is intra-coded is encoded into the bitstream. If this flag is true, the intra-mode index is transmitted in the bitstream; otherwise, the IBC merge index is transmitted in the bitstream. For another GPM split partition, the same syntax elements are transmitted. It should be noted that when both GPM split partitions are intra-coded or both are IBC-coded, the intra-mode index or IBC merge index of the two GPM split partitions will be different.
[0337] GPM using IBC and inter-frame prediction According to one or more embodiments of this disclosure, the IBC tool and the GPM tool are combined in the form of GPM utilizing IBC and inter-frame prediction. Different methods can be used to achieve this objective.
[0338] In the first method, an “inter-frame” portion of the GPM in the VVC, which utilizes inter-frame and inter-frame prediction methods, is replaced with IBC, where the IBC merged prediction results are weighted and averaged with the inter-frame merged prediction results to obtain the final prediction signal.
[0339] In the second approach, an "inter-frame" portion of the GPM in the ECM that utilizes inter-frame and inter-frame prediction methods is replaced with IBC, where template matching tools can be used to further improve encoding and decoding performance.
[0340] Simplified GPM form of IBC and intra-frame prediction combination According to one or more embodiments of this disclosure, the IBC tool and the GPM tool are combined in a simplified form that utilizes IBC and intra-frame prediction in GPM, such as by combining IBC and intra-frame prediction in a certain segmentation mode, which can save bit overhead in segmentation representation. Different methods can be used to achieve this goal.
[0341] In the first method, for a dividing line, such as 45 degrees, the upper left part of the coded block is encoded and decoded using intra-prediction mode, and the lower right part of the coded block is encoded and decoded using IBC prediction mode. They are then averaged in GPM form to obtain the final prediction signal.
[0342] IBC Prediction - Intra-Frame / Inter-Frame Prediction Combination According to one or more embodiments of this disclosure, coded blocks encoded in IBC mode are combined with coded blocks encoded in intra-frame mode or inter-frame mode. Different methods can be used to achieve this objective.
[0343] In the first approach, the decoder / encoder can combine a coded block encoded using IBC mode with a coded block encoded using intra-frame mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, a coded block encoded using IBC merge mode is treated as a coded block encoded using inter-frame merge mode and combined with a coded block encoded using planar intra-frame prediction mode. In another example, similar to the combination technique of CIIP with TIMD and TM merge in ECM, a coded block encoded using IBC merge-TM mode is combined with a coded block encoded using TIMD-derived intra-frame prediction mode.
[0344] When combining IBC-coded blocks with intra-frame-coded blocks, the weights can be designed similarly to the CIIP technique in VVC and the combination of CIIP with TIMD and TM in ECM. Specifically, 1) the weights of both the IBC-coded block and the intra-frame-coded block are greater than zero and less than one, or the weight of the intra-frame-coded block gradually changes from one to zero from one region to another in the current block (the opposite is true for the weight of the IBC-coded block); 2) the weights of the IBC-coded block and the intra-frame-coded block can be determined based on the coding modes of neighboring blocks and the intra-frame mode of the current block; 3) the weights of the IBC-coded block and the intra-frame-coded block can be uniform throughout the current block, or they can be different in different positions within the current block.
[0345] For example, the weights of IBC-coded blocks and intra-coded blocks can be determined as follows: When both the upper and left neighboring blocks of the current block are intra-coded and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 1 / 4 and 3 / 4, respectively. When both the upper and left neighboring blocks of the current block are IBC-coded and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 3 / 4 and 1 / 4, respectively. When one upper or left neighboring block is IBC-coded, another neighboring block is intra-coded, and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 1 / 2 and 1 / 2, respectively.
[0346] When the intra-frame mode of the current block is close to the horizontal angle mode (2 <= angle mode index < 34), the current block... Figure 16A The block is vertically divided; when the intra-frame mode of the current block is close to the vertical angle mode (34 <= angle mode index <= 66), the current block is as follows: Figure 16BThe blocks are horizontally divided. Table 7 shows the weights of the IBC coded blocks (wIBC) and intra-coded blocks (wIntra) for different sub-blocks. Furthermore, the weights of the IBC and intra-coded blocks can be determined in the CIIP-PDPC version. In this version, the intra-block mode is set to planar mode, and the weight of the intra-coded block gradually decreases as the combination position moves from the upper left to the lower right within the current block, and vice versa for the IBC coded block.
[0347]
[0348] Table 7. Modification weights used for angle mode.
[0349] When combining IBC-coded blocks with intra-frame-coded blocks, the weights can be designed as mask versions. That is, the weights of the IBC and intra-frame-coded blocks can be one or zero for different regions of the current block. The specific weights of the IBC and intra-frame-coded blocks can be determined based on the coding modes of neighboring blocks and the intra-frame mode of the current block. For example, the weights of the IBC and intra-frame-coded blocks can be determined as follows: when the intra-frame mode of the current block is close to the horizontal angle mode (2 <= angle mode index < 34), if the upper and left neighboring blocks of the current block are both intra-frame-coded, then the weight of the intra-frame-coded block is one in the left 3 / 4 region of the current block and zero in the right 1 / 4 region of the current block. Figure 21 As shown in (a), and conversely for the weights of IBC coded blocks; if only one neighboring block is intra-coded, the weight of the intra-coded block is one in the left half region of the current block and zero in the right half region of the current block, as shown in (a). Figure 21 As shown in (b), and conversely for the weights of IBC-coded blocks; if neither the upper nor left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the left 1 / 4 region of the current block and zero in the right 3 / 4 region of the current block, as shown in (b). Figure 21 As shown in (c), and conversely for the weights of IBC coded blocks.
[0350] When the intra-frame mode of the current block is close to the vertical angle mode (34 <= angle mode index <= 66), if the upper and left neighboring blocks of the current block are both intra-coded, then the weight of the intra-coded block is one in the top 3 / 4 region of the current block and zero in the bottom 1 / 4 region of the current block. Figure 21 As shown in (d), and conversely for the weights of IBC coded blocks; if only one neighboring block is intra-coded, the weight of the intra-coded block is one in the top half region of the current block and zero in the bottom half region of the current block, as shown in (d). Figure 21 As shown in (e), and conversely for the weights of IBC-coded blocks; if neither the upper nor left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the top 1 / 4 region of the current block and zero in the bottom 3 / 4 region of the current block, as shown in (e). Figure 21 As shown in (f), and conversely for the weights of IBC coded blocks.
[0351] When the intra-frame mode of the current block is planar mode, if the upper and left neighboring blocks of the current block are both intra-coded, then the weight of the intra-coded block is one in the upper left 3 / 4 region of the current block (horizontal index less than 1 / 2 the width of the current block or vertical index less than 1 / 2 the height of the current block), and zero in the lower right 1 / 4 region of the current block (horizontal index equal to or greater than 1 / 2 the width of the current block and vertical index equal to or greater than 1 / 2 the height of the current block). Figure 21 As shown in (g), and conversely for the weights of IBC-coded blocks; if only the upper neighboring block is intra-coded, the weight of the intra-coded block is one in the top half region of the current block and zero in the bottom half region of the current block, as shown in (g). Figure 21 As shown in (e), and conversely for the weights of IBC coded blocks; if only the left neighboring block is intra-coded, the weight of the intra-coded block is one in the left half region of the current block and zero in the right half region of the current block, as shown in (e). Figure 21 As shown in (b), and conversely for the weights of IBC-coded blocks; if neither the upper nor left neighboring blocks of the current block are intra-coded, the weight of the intra-coded block is one in the upper left quarter region of the current block (the horizontal index is less than half the width of the current block and the vertical index is less than half the height of the current block), and zero in the lower right three-quarter region of the current block (the horizontal index is equal to or greater than half the width of the current block or the vertical index is equal to or greater than half the height of the current block), as shown in (b). Figure 21 As shown in (h), and conversely for the weights of IBC coded blocks.
[0352] The two weight design methods described above can be used independently or combined. For example, when the intra-frame mode of the current block is planar mode, if neither the upper nor left neighboring block of the current block is intra-coded, the weights of the IBC coded block and the intra-coded block can be designed similarly to the CIIP technique in VVC. Under other conditions, the weights of the IBC coded block and the intra-coded block can be designed as a masked version.
[0353] When combining IBC-coded blocks with intra-frame-coded blocks, weights can be designed based on template matching methods. The sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template can be used to calculate the weights of the IBC and intra-frame-coded blocks. SATD is a widely used block matching standard in scenarios such as fractional motion estimation in video compression. For example, the weights of IBC and intra-frame-coded blocks can be determined as follows: For intra-frame-coded blocks, such as... Figure 22 As shown, the SATD between the predicted samples and the reconstructed samples of the current block template is calculated as follows: In this context, the predicted samples of the current block template are obtained by intra-prediction using the reference samples of the template in the intra-frame mode of the current block. For IBC coded blocks, such as... Figure 23 As shown, the SATD between the predicted samples and the reconstructed samples of the current block template is calculated as follows: In this context, the predicted samples of the current block template are predicted using the reference samples pointed to by the block vector of the current block. (IBC encoded block) and intra-frame coded blocks The weights are determined as follows:
[0354]
[0355] When using the template of the current block to calculate the weights of IBC and intra-coded blocks, both the left and top templates can be used if both are available, or only the left or top template if only is available. The use of both the left and top templates, only the top template, or only the left template can be determined during Rate-Distortion Optimization (RDO) or according to predefined criteria. RDO techniques typically minimize the amount of distortion (video quality loss) relative to the amount of data required to encode the video. For predefined criteria, for example, if the intra-mode of the current block is planar or DC, both the left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0356] When combining a coded block encoded in IBC mode with a coded block encoded in intra-frame mode, the weights derived using the template matching method can be compared with the weights derived using CIIP in the reference ECM (or weights designed as a masked version) during the RDO process. This means that a block-level flag needs to be transmitted in the bitstream to indicate which method is used; alternatively, weights derived using the template matching method can replace all or part of the weights derived using CIIP in the reference ECM (or weights designed as a masked version). For example, if the intra-frame mode of the current block is planar mode or DC mode, the weights of the intra-frame coded block and the IBC coded block are determined based on the template matching method; otherwise, the weights are determined using CIIP in the reference ECM (or in a masked version).
[0357] In the second approach, the decoder / encoder can combine a coded block encoded in IBC mode with a coded block encoded in inter-frame mode. Various methods can be used in this combination. In one example, similar to CIIP in VVC, a coded block encoded in IBC merge mode is treated as a coded block encoded in planar intra-frame mode and combined with a coded block encoded in inter-frame merge mode. In another example, a coded block encoded in IBC merge mode is treated as a coded block encoded in inter-frame merge mode and combined with a coded block encoded in inter-frame merge mode by equal averaging.
[0358] In the third approach, the decoder / encoder can combine coded blocks encoded in IBC mode with coded blocks encoded in intra-frame mode and inter-frame mode. Various methods can be used in this combination. In one example, the coded blocks encoded in IBC mode, intra-frame mode, and inter-frame mode are directly combined by equal averaging. In another example, the coded blocks encoded in IBC mode are first combined separately with the coded blocks encoded in intra-frame mode and inter-frame mode, as presented in the first and second approaches. Then, the individual combination results are combined by equal averaging.
[0359] CIIP Improvements When designing the weights of intra-coded blocks and inter-coded blocks in CIIP, the weights can be designed based on the template matching method. The sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template can be used to calculate the weights of intra-coded blocks and inter-coded blocks. For example, the weights of intra-coded blocks and inter-coded blocks can be determined as follows: For intra-coded blocks, such as... Figure 22 As shown, the SATD between the predicted samples and the reconstructed samples of the current block template is calculated as follows: In this context, the predicted samples of the current block template are obtained by intra-frame prediction using the reference samples of the template in the intra-frame mode of the current block. For inter-coded blocks, such as... Figure 23 As shown (replace "IBC merge candidate BV" with "Inter-frame merge candidate MV" in the figure), the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: In this context, the predicted samples of the current block template are predicted using the motion vectors of the current block and the reference samples of the template. (Inter-frame coded block) and intra-frame coded blocks The weights are determined as follows:
[0360]
[0361] When using the template of the current block to calculate the weights of inter-coded blocks and intra-coded blocks, if both the left and top templates are available, both can be used; otherwise, if only the left or top template is available, only that template can be used. The decision to use both the left and top templates, only the top template, or only the left template can be made during the RDO process or based on predefined criteria. For example, if the intra-mode of the current block is planar or DC mode, both the left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0362] When using template matching-based methods to derive the weights of intra- and inter-coded blocks in CIIP, the weights derived using template matching can be compared with those derived using the original method during the RDO process. This means that block-level flags need to be transmitted in the bitstream to indicate which method was used; alternatively, weights derived using template matching can replace all or part of the weights derived using the original method. For example, if the intra-mode of the current block is planar or DC, the weights of intra- and inter-coded blocks in CIIP are determined using template matching; otherwise, the weights are determined using the original method.
[0363] Multiple Hypothesis IBC Prediction According to one or more embodiments of this disclosure, the number of block vectors (BVs) in the IBC tool is increased to two or more, and two or more hypotheses are combined to obtain the final prediction result. Different methods can be used to achieve this goal.
[0364] In the first approach, the decoder / encoder can combine two hypotheses corresponding to the two BVs to obtain the final prediction. Various methods can be used to achieve this. In one example, the two BVs corresponding to the minimum and second minimum rate distortion metrics in the IBC AMVP mode are averaged equally to obtain the final prediction. In another example, the predictions corresponding to the IBC AMVP mode and the predictions corresponding to the IBC merge mode are averaged equally to obtain the final prediction.
[0365] In the second approach, the decoder / encoder can combine more hypotheses corresponding to more BVs to obtain the final prediction. Various methods can be used to achieve this. In one example, the iterative accumulation method proposed in the Multiple Hypothesis Prediction (MHP) technique is used to obtain the final prediction. In another example, all BVs corresponding to the minimum, second smallest, third smallest, ... rate distortion metrics in the IBC AMVP mode are averaged equally to obtain the final prediction.
[0366] Predicted block candidate derivation In some embodiments, candidate prediction blocks are searched and selected based on a criterion of minimizing template matching cost; that is, the top N candidates that result in the minimum BV matching cost are selected. The BV matching cost may not be limited to SAD (sum of absolute differences) and SSE (sum of squared errors).
[0367] In some embodiments, candidate prediction blocks can be selected based on a predefined pattern (i.e., a planar pattern).
[0368] In some embodiments, candidate prediction blocks can be selected based on a neighboring predefined pattern (i.e., a top predefined pattern and a left predefined pattern).
[0369] Fixed Multiple Hypothesis IBC In this embodiment, the weighting factors used to generate the final prediction block are predefined and fixed on both the encoder and decoder sides. As an example, equal weighting factors can be used, i.e., all candidate blocks can have a weighting factor of 1 / N.
[0370] Adaptive Multiple Hypothesis IBC To adapt to the different characteristics of video content, an adaptive multiple hypothesis (IBC) method is also proposed.
[0371] In some embodiments, the weighting factor can be derived based on the BV matching cost. The BV matching costs of N candidates are expressed as... , ,…, The weighting factors are calculated as follows.
[0372] (4) It should be noted that SAD and SSE can be used (but are not limited to) to measure the cost of BV matching.
[0373] In yet another embodiment, the weighting factor can be derived / switched based on the block size or syntax element of the signal transmission at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / Region / CTU / CU / subblock / sample level.
[0374] In yet another embodiment, the weighting factors can be derived on the encoder side and then transmitted as signals to the decoder in the bitstream. The N candidate prediction blocks are represented as... , ,…, And represent the current block as The weighting factors can then be solved using the following equation: (5) Equation (5) can be solved using the Wiener-Hopf equation as an ALF. The derived filter coefficients are then quantized to integer type and transmitted as signals at the block level.
[0375] In yet another embodiment, the weighting factors can be derived on the encoder side and then transmitted as signals to the decoder in the bitstream. The N candidate prediction blocks are represented as... , ,…, And represent the current block as The weighting factors can then be solved using the following equation: (6) Equation (6) can be solved using LDL decomposition or Gaussian elimination.
[0376] In yet another embodiment, a weighting factor is derived based on a template, and the derived weighting factor is applied to prediction block candidates to generate the final prediction block. The template for the prediction candidate is represented as... , ,…, And represent the current block as The weighting factors can then be derived using the following equation: (7) Equation (7) can be solved using the Wiener-Hopf equation. Then, the final predicted block can be calculated as follows: ,in, This represents the i-th prediction block candidate.
[0377] The IBC model leverages nonlocal correlation to improve prediction accuracy, where similar blocks are searched and used to generate the final prediction block. In this embodiment, a combination of nonlocal mean filtering and multiple hypothesis IBC is proposed, as described below. In the first step, N prediction block candidates are searched and identified, as performed in IBC. In the second step, weighting factors are calculated as follows.
[0378] (8) in, Used to measure the distance between the template of the i-th prediction block candidate and the template of the current block. Used as a weighting degree, and It is a normalization constant: (9) In order to calculate the weighting factors in equation (8), the weighting strengths should first be determined. Several methods for determining the weighting strengths are proposed in this disclosure.
[0379] In the first approach, a candidate list of weighted strength values, including some typical weighted strength values, is defined and fixed on both the encoder and decoder sides. On the encoder side, rate-distortion optimization is used to examine the weighted strength values, and the optimal weighted strength value is identified and transmitted as a signal to the decoder side in the bitstream.
[0380] In the second method, the templates of the predicted block candidates and the template of the current block are used to estimate the weighted intensity value. The template of the predicted candidate is represented as... , ,…, And represent the current block as Then, the weighted intensity value can be solved using the following equation: (10) In the third method, the weighted intensity value can be estimated using the QP value and the variance of the template of the current block. That is, the relationship between the weighted intensity value, the QP value and the template variance can be fitted offline.
[0381] To better utilize the nonlocal correlations in IBC, this embodiment employs Singular Value Decomposition (SVD) to generate the final prediction block from the prediction block candidates. The width and height of the current block are represented by W and H, and the area of the current block is represented by... .
[0382] Step 1. Search and identify K candidate prediction blocks Such as in FIBC.
[0383] Step 2. Current block K predicted block candidate building block groups And it is arranged as a matrix: (11) in, It is by grouping Each candidate arrangement in the array is a column vector with dimensions. The matrix.
[0384] Step 3. For the matrix Perform SVD decomposition.
[0385] (12) Step 4. For the singular value matrix Apply soft threshold operation.
[0386] (13) in, It uses a threshold shrink A function of the diagonal elements. For The k-th diagonal element in the array is obtained through a nonlinear function. At the level Contraction occurs at this location: (14) It is a contracted singular value located on the diagonal. The matrix formed by these elements.
[0387] Step 5. Perform inverse SVD to obtain the filtered patch group.
[0388] (15) One of the key steps is determining the threshold for each diagonal element in step 4. In this disclosure, the threshold is calculated as follows. For each group of image patches, the threshold is estimated using the following equation: (16) in, It is the standard deviation of the noise, and The original block is in the group The standard deviation in the k-th dimension of the SVD space is calculated. The deviation of the original block in the SVD space is estimated as follows.
[0389] (17) in, yes The k-th singular value. When When the value is zero, the soft threshold operation is skipped. Additionally, using... and A parameterized power function is used to estimate the noise bias by using the bias of the prediction block.
[0390] (18) in The calculation is as follows: (19) here, Represents the candidate vector of the prediction block The i-th pixel.
[0391] Multiple Hypothesis IBC Signal Transmission In this disclosure, the proposed multiple hypothesis IBC can be used as an alternative to the current IBC mode, or the encoder can adaptively select either the IBC mode or the multiple hypothesis IBC mode.
[0392] In some embodiments, multiple hypothesis IBC can be used as an alternative to the current IBC model, i.e., always using multiple hypotheses for prediction.
[0393] In yet another embodiment, one of the multiple hypothesis IBC methods described above is used in conjunction with the current IBC mode. A flag is transmitted in the bitstream to indicate whether the multiple hypothesis IBC mode is applied to the CU.
[0394] In yet another embodiment, more than one of the multiple hypothesis IBC methods described above is used in conjunction with the current IBC mode. First, a signaling flag is transmitted in the bitstream to indicate whether a multiple hypothesis IBC mode is applied. Then, a signaling index is transmitted to indicate which of the multiple hypothesis IBC methods is applied to the CU.
[0395] In yet another embodiment, the multiple hypothesis IBC method described above is used in conjunction with the current IBC mode. The multiple hypothesis IBC can be used as an alternative to the current IBC mode based on certain coding information of the current block, such as SAD (Sum of Absolute Differences), SSE (Sum of Squared Errors), quantization parameters (QP) associated with TB / CB and / or slices, the nearest neighbor prediction mode of the CU (e.g., IBC mode or intra-frame or inter-frame mode) and / or slice type (e.g., I-slice, P-slice, or B-slice).
[0396] Combined Intra-Frame TMP - Intra / Inter-Frame Prediction According to one or more embodiments of this disclosure, coded blocks encoded in intra-frame TMP mode are combined with coded blocks encoded in intra-frame mode or inter-frame mode. Different methods can be used to achieve this objective.
[0397] In the first approach, the decoder / encoder can combine a coded block encoded using intra-TMP mode with a coded block encoded using intra-mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, a coded block encoded using intra-TMP mode is treated as a coded block encoded using inter-frame combining mode and combined with a coded block encoded using planar intra-prediction mode. In another example, similar to the combination technique of CIIP with TIMD and TM combining in ECM, a coded block encoded using intra-TMP mode is combined with a coded block encoded using intra-prediction mode derived from TIMD.
[0398] When combining coded blocks encoded using intra-TMP mode with coded blocks encoded using intra-mode mode, the weights can be determined by referring to the weight design of CIIP technology in VVC or ECM. Alternatively, weights can be designed based on template matching methods. The sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template can be used to calculate the weights of the intra-TMP coded block and the intra-coded block. For example, the weights of the intra-TMP coded block and the intra-coded block can be determined as follows: For the intra-coded block, such as... Figure 22 As shown, the SATD between the predicted samples and the reconstructed samples of the current block template is calculated as follows: The predicted samples of the current block template are obtained by intra-prediction using the reference samples of the template in the intra-frame mode of the current block. For intra-TMP coded blocks, the SATD between the predicted samples and the reconstructed samples of the current block template is calculated as follows: The predicted samples of the current block template are predicted using the reference samples pointed to by the block vector of the first block. Intra-frame TMP coded blocks. and intra-frame coded blocks The weights are determined as follows:
[0399]
[0400] When using the template of the current block to calculate the weights of intra-TMP coded blocks and intra-coded blocks, if both the left and top templates are available, both can be used; otherwise, if only the left or top template is available, only that template can be used. The decision to use both the left and top templates, only the top template, or only the left template can be made during the RDO process or based on predefined criteria. For example, if the intra-mode of the current block is planar or DC mode, both the left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0401] When combining coded blocks encoded in intra-TMP mode with coded blocks encoded in intra-mode mode, the weights derived from the template matching method can be compared with the weights derived from CIIP in the reference ECM during the RDO process. This means that a block-level flag needs to be transmitted in the bitstream to indicate which method is used; alternatively, weights derived from the template matching method can replace all or part of the weights derived from CIIP in the reference ECM. For example, if the intra-mode of the current block is planar or DC mode, the weights of the intra-coded block and the IBC coded block are determined based on the template matching method; otherwise, the weights are determined by referring to CIIP in the ECM.
[0402] In the second approach, the decoder / encoder can combine coded blocks encoded using intra-TMP mode with coded blocks encoded using inter-frame mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, a coded block encoded using intra-TMP mode is treated as a coded block encoded using planar intra-frame mode and combined with a coded block encoded using inter-frame combining mode. In another example, a coded block encoded using intra-TMP mode is treated as a coded block encoded using inter-frame combining mode and combined with a coded block encoded using inter-frame combining mode by equal averaging.
[0403] In the third approach, the decoder / encoder can combine coded blocks encoded in intra-TMP mode with coded blocks encoded in both intra-mode and inter-mode. Various methods can be employed in this combination. In one example, the coded blocks encoded in intra-TMP mode, intra-mode, and inter-mode are directly combined by equal averaging. In another example, the coded blocks encoded in intra-TMP mode are first combined separately with the coded blocks encoded in both intra-mode and inter-mode, as presented in the first and second approaches. Then, the individual combination results are combined by equal averaging.
[0404] GPM predicted using intra-frame TMP and intra-frame TMP According to one or more embodiments of this disclosure, an intra-frame TMP tool and a GPM tool are combined in the form of utilizing intra-frame TMP and GPM predicted by intra-frame TMP. Different methods can be used to achieve this objective.
[0405] In the first approach, both "inter-frame" components of the GPM (Gross Frame Prediction) using inter-frame and inter-frame prediction methods in VVC are replaced with intra-frame TMPs. This means that the prediction results of the two intra-frame TMPs are weighted and averaged according to the segmentation lines in the coded blocks. The weights can be obtained by referring to the GPM using inter-frame and inter-frame prediction methods in VVC.
[0406] In the second method, both “inter-frame” parts of the GPM in the ECM that utilize inter-frame and inter-frame prediction methods are replaced with intra-frame TMPs. Template matching tools can be used to further improve encoding and decoding performance.
[0407] Utilizing intra-frame TMP and intra-frame prediction GPM According to one or more embodiments of this disclosure, an intra-frame TMP tool and a GPM tool are combined in a manner that utilizes intra-frame TMP and intra-frame predicted GPM. Different methods can be used to achieve this objective.
[0408] In the first method, the “inter-frame” portion of the GPM in the ECM, which utilizes inter-frame and intra-frame prediction methods, is replaced with the intra-frame TMP, where the intra-frame TMP prediction results are weighted and averaged with the intra-frame prediction results to obtain the final prediction signal.
[0409] GPM using intra-frame TMP and inter-frame prediction According to one or more embodiments of this disclosure, the intra-frame TMP tool and the GPM tool are combined in a form that utilizes intra-frame TMP and inter-frame predicted GPM. Different methods can be used to achieve this objective.
[0410] In the first method, an “inter-frame” portion of the GPM in VVC, which utilizes inter-frame and inter-frame prediction methods, is replaced with an intra-frame TMP, where the intra-frame TMP prediction results are weighted and averaged with the inter-frame merged prediction results to obtain the final prediction signal.
[0411] In the second approach, an "inter-frame" portion of the GPM in the ECM, which utilizes inter-frame and inter-frame prediction methods, is replaced with an intra-frame TMP. Template matching tools can be used to further improve encoding and decoding performance.
[0412] Simplified GPM form of intra-TMP and intra-prediction combination According to one or more embodiments of this disclosure, the intra-frame TMP tool and the GPM tool are combined in a simplified form that utilizes intra-frame TMP and intra-frame prediction GPM, such as by combining intra-frame TMP and intra-frame prediction in a certain segmentation mode, which can save bit overhead in segmentation representation. Different methods can be used to achieve this goal.
[0413] In the first method, for a dividing line, such as 45 degrees, the upper left part of the coded block is encoded and decoded using intra-frame prediction mode, and the lower right part of the coded block is encoded and decoded using intra-frame TMP prediction mode. They are then averaged in GPM form to obtain the final prediction signal.
[0414] Combining IBC with TIMD mode According to one or more embodiments of this disclosure, the IBC tool is combined with the TIMD tool. Different methods can be used to achieve this objective.
[0415] In the first method, the IBC mode is treated as an intra-prediction mode added to the MPM list. Then, the IBC mode is compared with other intra-prediction modes in the MPM list using template matching cost. Finally, the TIMD method is used to fuse the two modes with minimum cost and the second minimum cost to obtain the final prediction result.
[0416] In the second method, the conventional TIMD prediction results are first obtained, then the template matching cost of the IBC model and the conventional TIMD prediction results are calculated, and finally the TIMD method is used to fuse the IBC model and the conventional TIMD prediction results to obtain the final prediction result.
[0417] Combining intra-frame TMP mode with TIMD mode According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a TIMD tool. Different methods can be used to achieve this goal.
[0418] In the first method, the intra-TMP mode is treated as an intra-prediction mode added to the MPM list. Then, the intra-TMP mode is compared with other intra-prediction modes in the MPM list using template matching cost. Finally, the TIMD method is used to fuse the two modes with the minimum cost and the second minimum cost to obtain the final prediction result.
[0419] In the second method, the regular TIMD prediction result is first obtained, then the template matching cost of the intra-frame TMP mode and the regular TIMD prediction result are calculated, and finally the TIMD method is used to fuse the intra-frame TMP mode and the regular TIMD prediction result to obtain the final prediction result.
[0420] Combining intra-frame TMP with LIC According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a LIC tool. Different methods can be used to achieve this objective.
[0421] In the first approach, the intra-frame TMP mode is treated as an inter-frame mode, and LIC is used to model the local illumination variation between the current block and its intra-frame TMP prediction block as a function of the local illumination variation between the current block template and the reference block template. This function is a linear equation as used in the conventional LIC method.
[0422] Combining IBC with OBMC According to one or more embodiments of this disclosure, the IBC tool is combined with the OBMC tool. Different methods can be used to achieve this objective.
[0423] In the first approach, the IBC mode is treated as an inter-frame mode, and the conventional OBMC method is applied, using block vector information from neighboring blocks to refine the top and left boundary pixels of the IBC-coded CU with weighted prediction.
[0424] In the second approach, the IBC mode is treated as an inter-frame mode, and a template-matching-based OBMC method is applied to refine the top and left boundary pixels of the IBC-coded CU.
[0425] It should be noted that when IBC is combined with OBMC, for a CU encoded in IBC mode, when using the regular OBMC method or the template matching-based OBMC method to refine the top and left boundary pixels of the current CU using the shift information of neighboring blocks, the neighboring blocks can be encoded in IBC mode or intra-frame TMP mode.
[0426] Combining intra-frame TMP with OBMC According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with an OBMC tool. Different methods can be used to achieve this objective.
[0427] In the first approach, the intra-frame TMP mode is treated as an inter-frame mode, and the conventional OBMC method is applied, using block vector information from neighboring blocks to refine the top and left boundary pixels of the intra-frame TMP coded CU using weighted prediction.
[0428] In the second method, the intra-frame TMP mode is treated as an inter-frame mode, and a template-matching-based OBMC method is applied to refine the top and left boundary pixels of the intra-frame TMP encoded CU.
[0429] It should be noted that when combining intra-frame TMP with OBMC, for CUs encoded in intra-frame TMP mode, when using the regular OBMC method or the template matching-based OBMC method to refine the top and left boundary pixels of the current CU using the shift information of neighboring blocks, the neighboring blocks can be encoded in intra-frame TMP mode or IBC mode.
[0430] Derivation of non-adjacent candidates in IBC AMVP mode or IBC merge mode According to one or more embodiments of this disclosure, the candidate derivation process for IBC merging mode or IBC AMVP mode is extended by using not only adjacent neighbor blocks but also non-adjacent neighbor blocks. The relevant content is summarized into sections such as "candidate scanning and candidate pruning," "candidate reordering," "motion information storage," "application scope," and "application fast size," and is presented as follows: Candidate scan and candidate pruning For candidate scans, non-adjacent neighboring blocks are scanned and selected using the following method: Scanning area and distance: In one or more embodiments, non-adjacent neighboring blocks can be scanned from the left and top regions of the current coded block. The scan distance can be defined as the number of coded blocks from the scan position to the left or top of the current coded block.
[0431] like Figure 24 As shown, multiple non-adjacent neighboring blocks can be scanned to the left or above the current encoded block. Figure 24 The distances shown represent the number of coded blocks to the left or top of the current block from each candidate location. For example, a region to the left of the current block with a "distance 2" indicates that the candidate neighboring block in that region is 2 blocks away from the current block. Similar indications can be applied to other scan regions with different distances.
[0432] In one or more embodiments, non-adjacent neighbor blocks at each distance may have the same block size as the current coded block, as shown in Figure 25(a). Note that when non-adjacent neighbor blocks at each distance have the same block size as the current coded block, the block size value is adaptively changed according to the partitioning granularity at each different region in the image.
[0433] In some embodiments, non-adjacent neighbor blocks at each distance may have a different block size than the current coded block, as shown in Figure 25(b). Note that when non-adjacent neighbor blocks at each distance have a different block size than the current coded block, the value of the block size can be predefined as a constant value, such as 4 × 4, 8 × 8, or 16 × 16.
[0434] Based on the defined scan distance, the total size of the scanned region to the left or above the current encoded block can be determined by a configurable distance value. In one or more embodiments, the maximum scan distance to the left and top can use the same or different values. For example, the maximum distance to the left and top shares the same value of 2. Multiple maximum scan distance values can be determined by the encoder side and transmitted in the bitstream via signaling. Alternatively, multiple maximum scan distance values can be predefined as multiple fixed values, such as values 2 or 4. When the maximum scan distance is predefined as a value of 4, this indicates that the scanning process terminates when the candidate list is full or all non-adjacent neighboring blocks with a maximum distance of 4 have been scanned (whichever arrives first).
[0435] In one or more embodiments, the start neighbor block and the end neighbor block may be location-dependent within each scan area at a specific distance.
[0436] In one or more embodiments, for the left-side scan region, the starting neighbor block can be the lower-left neighbor block of the starting neighbor block in an adjacent scan region with a small distance. For example, as... Figure 24 As shown, the starting neighboring block of the "distance 2" scan region to the left of the current block is the lower-left neighboring block of the starting neighboring block of the "distance 1" scan region. The ending neighboring block can be the block to the left of the ending neighboring block in the upper scan region with a smaller distance. For example, as... Figure 24 As shown, the end neighboring block of the "distance 2" scan area to the left of the current block is the left-adjacent neighboring block of the end neighboring block of the "distance 1" scan area above the current block.
[0437] Similarly, for the upper scan region, the starting neighbor block can be the upper-right neighbor of the starting neighbor block in an adjacent scan region with a smaller distance. The ending neighbor block can be the upper-left neighbor of the ending neighbor block in an adjacent scan region with a smaller distance.
[0438] In one or more embodiments, the sampling interval between the start neighbor block and the end neighbor block within each scan region at a specific distance can be location-dependent. In one or more embodiments, the sampling interval between the start neighbor block and the end neighbor block is smaller in scan regions with smaller distances. For example, as... Figure 24 As shown, each neighboring block between the start and end neighboring blocks is scanned in a scan region with a "distance 1"; and every two neighboring blocks between the start and end neighboring blocks are scanned in a scan region with a "distance 2". In one or more embodiments, the sampling interval between the start and end neighboring blocks may be the same or different for different side scan regions with a specific distance. For example, for a left-side scan region and an upper-side scan region with a specific distance, the sampling interval between the start and end neighboring blocks is the same.
[0439] Scanning order: When scanning neighboring blocks in a non-adjacent region, a certain order and / or rules can be followed to determine the selection of neighboring blocks to be scanned.
[0440] In one or more embodiments, the left-side region may be scanned first, followed by the upper region. For example... Figure 24 As shown, you can first scan the three non-adjacent rows on the left (e.g., from distance 1 to distance 3), and then scan the three non-adjacent rows above the current block.
[0441] In some embodiments, the left and upper regions can be scanned alternately. For example, as Figure 24 As shown, the left scanning area with a "distance of 1" is scanned first, and then the upper area with a "distance of 1" is scanned.
[0442] For scanning areas located on the same side (e.g., the left or upper region), the scanning order is from areas with smaller distances to areas with larger distances. This order can be flexibly combined with other embodiments of the scanning order. For example, the left and upper regions can be scanned alternately, and the regions on the same side can be arranged in order from small to large distances.
[0443] The scanning order within each scanned region at a specific distance can be defined. In one or more embodiments, for the left scanned region, scanning can begin from the bottom neighboring block to the top neighboring block. For the upper scanned region, scanning can begin from the right block to the left block.
[0444] In one or more embodiments, non-adjacent regions in one direction may be scanned first, followed by scanning non-adjacent regions in other directions. Within a direction, the scanning order can be defined. In one or more embodiments, within each direction, the scanning can start from a smaller distance and proceed to a larger distance.
[0445] In some embodiments, non-adjacent regions in different directions can be scanned alternately. For example, such as Figures 26 to 27 As shown, the system first scans non-adjacent regions with small distances in the direction with an angle value of 225 degrees. Then, it sequentially scans non-adjacent regions with small distances in the directions with angle values of 45, 90, 180, and 135 degrees. Next, it scans non-adjacent regions with larger distances in the direction with an angle value of 225 degrees, followed by non-adjacent regions with larger distances in the directions with angle values of 45, 90, 180, and 135 degrees.
[0446] In some examples, a total of 18 blocks are scanned, such as Figure 26 As shown, the scanned blocks are indicated by a boxed integer n, where n is in the range of 1 to 18 (inclusive), and where n represents the scan order.
[0447] In some examples, a total of 48 blocks are scanned, such as Figure 27 As shown, the scanned blocks are indicated by a boxed integer n, where n is in the range of 1 to 48 (inclusive), and where n represents the scan order; in these examples, degree values of 270, 0, 247.5, 22.5, 202.5, 67.5, 157.5 and 112.5 can also be used to determine the scan order.
[0448] Scan terminated: For non-adjacent candidates, neighboring blocks encoded using IBC mode or intra-frame TMP mode are defined as qualified candidates.
[0449] In one or more embodiments, the scanning process can be performed interactively. For example, a scan performed in a specific region at a specific distance can stop when the first X qualified candidates are identified, where X is a predefined positive value. For example, as Figure 24 As shown, scanning in the left scan region at a distance of 1 can stop when the first or more qualified candidates are identified. The next iteration of the scanning process then begins by targeting another scan region, which is regulated by a predefined scan order / rules.
[0450] In one or more embodiments, X can be defined for each distance. For example, X is set to 1 for each distance, meaning that for each distance, if the first qualified candidate is found, the scan is terminated and the scan process restarts from a different distance in the same area or from the same or different distances in different areas. Note that the value of X can be set to the same value or different values for different distances. If the maximum number of qualified candidates is found from all allowed distances in an area (e.g., defined by the maximum distance), the scan process for that area is terminated completely.
[0451] In another embodiment, X can be defined for a region. For example, X is set to 3, which means that for the entire region (e.g., the region to the left or above the current block), if the first 3 qualified candidates are found, the scan terminates and the scan process restarts from the same or different distance in another region. Note that the value of X can be set to the same value or different values for different regions. If the maximum number of qualified candidates is found from all regions, the entire scan process terminates completely.
[0452] The value of X can be defined for both distance and region. For example, X can be set to 3 for each region (e.g., the region to the left or above the current block), and X can be set to 1 for each distance. The value of X can be set to the same value or different values for different regions and distances.
[0453] In some embodiments, the scanning process can be performed continuously. For example, a scan performed in a specific region at a specific distance can stop when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or when the maximum allowed number of candidates has been reached. The maximum allowed number of candidates can be set in different ways. In one example, the maximum allowed number of candidates is set to a predefined value, which can be set to the maximum allowed size of the IBC merge candidate list (equal to 28 in ECM), or the maximum allowed size of the IBC merge candidate list minus 1, or other values. In another example, the maximum allowed number of candidates can be set to different values depending on the encoding conditions, and this value is transmitted in the bitstream.
[0454] During the candidate scanning process, each candidate non-adjacent neighbor block is identified and scanned according to the scanning method described above. For ease of implementation, each candidate non-adjacent neighbor block can be indicated or located by a specific scanning position. For example, the lower right position is used for both the upper and left non-adjacent neighbor blocks.
[0455] After identifying a qualified candidate following the above process, that candidate can undergo a similarity check against all existing candidates in the candidate list. Details of the similarity check can be found in the existing similarity check rules in the current IBC candidate derivation. If a new qualified candidate is found to be similar to any existing candidate in the candidate list, the new qualified candidate is removed / pruned.
[0456] It should be noted that the above candidate scanning and candidate pruning processes can be the same or different for IBC AMVP candidate derivation and IBC merge candidate derivation. For example, Figure 26 The candidate scanning and pruning process presented can be used for both IBCAMVP candidate derivation and IBC merge candidate derivation. In another example, Figure 26The candidate scanning and pruning process presented can be used for IBC AMVP candidate derivation. Figure 27 The candidate scanning and pruning process presented can be used for IBC merging candidate derivation. In the third example, Figure 26 The candidate scanning and pruning process presented can be used for IBC AMVP candidate derivation. Figure 41 The candidate scanning and pruning process presented can be used for IBC merging candidate derivation. In the three examples above, Figure 26 , Figure 27 and Figure 41 Examples of scan area, scan distance, and scan order information are presented. Regarding scan termination, in one example, a scan performed in a specific area at a specific distance stops when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or the maximum allowed number of IBC merge candidates has been reached. In another example, a scan performed in a specific area at a specific distance stops when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or the maximum allowed number of IBC merge candidates has been reached minus one. To illustrate this more clearly... Figure 26 and Figure 41 ,exist Figure 26 and Figure 41 Each non-adjacent neighboring block at a distance in the code has the same block size as the currently encoded block. Figure 26 The maximum scan distance on the left and top sides is set to 4 (excluding adjacent blocks), and... Figure 41 The maximum scan distance on the left and top sides is set to 7 (excluding adjacent blocks). Figure 26 and Figure 41 In the middle, the bottom right scan position is used for non-adjacent neighboring blocks along the bottom left, top right, and top left directions. Figure 26 and Figure 41 The middle and lower middle (horizontal index equal to half the block width, vertical index equal to the vertical index of the non-adjacent neighboring blocks along the upper right direction within the same scan distance) scan position is used for non-adjacent neighboring blocks along the upper direction. Figure 26 and Figure 41 The middle and middle-right (vertical index equal to half the block height, horizontal index equal to the horizontal index of the non-adjacent neighboring blocks along the lower left direction within the same scan distance) scan positions are used for non-adjacent neighboring blocks along the left direction.
[0457] Candidate Reordering When inserting spatially non-adjacent candidates into the IBC candidate list, all spatially non-adjacent candidates can be grouped and inserted as a whole into different positions in the IBC candidate list, or spatially non-adjacent candidates can be divided into several subgroups and each subgroup can be inserted into a different position in the IBC candidate list.
[0458] In one or more embodiments, spatially non-adjacent candidates may be inserted into the IBC candidate list in the following order: 1. Spatial BVP from spatially adjacent neighboring blocks 2. Spatial BVP from non-adjacent neighboring blocks 3. Historical BVP from FIFO table 4. Paired average BVP 5. For example Figure 13 As shown, BVP candidates are located in the IBC reference region. 6. Zero BVP In another embodiment, spatially non-adjacent candidates can be inserted into the IBC candidate list in the following order: 1. Spatial BVP from spatially adjacent neighboring blocks 2. First X-space BVP from non-adjacent neighboring blocks 3. Historical BVP from FIFO table 4. Other Y-space BVPs from non-adjacent neighboring blocks 5. Paired average BVP 6. For example Figure 13 As shown, BVP candidates are located in the IBC reference region. 7. Zero BVP The values of X and Y can be predefined fixed values (such as value 2), values transmitted via signaling received by the decoder (parameters transmitted via signaling at the sequence / strip / block / CTU level), values that can be configured at the encoder / decoder, values dynamically determined based on the number of available neighboring blocks to the left and above each individual coded block (e.g., X <= 3, Y <= 3), or any combination of methods for determining the values of X and Y. In one example, the value of X can be the same as the value of Y. In another example, the value of X can be different from the value of Y.
[0459] Since candidates placed towards the end of the IBC candidate list may incur higher signal transmission overhead when selected by the encoder and transmitted, the order of the above different categories of candidates can be designed using the following different methods: In one or more embodiments, the order of these candidates remains the same as the insertion order described above. An adaptive reordering method can then be applied to reorder these candidates; the adaptive reordering method can be a template matching (ARMC) based method.
[0460] In one or more embodiments, before inserting spatially non-adjacent candidates into the IBC candidate list, an adaptive reordering method (which may be a template matching-based method (ARMC)) may be applied to the derived spatially non-adjacent candidates, and then the first X candidates may be inserted into the IBC candidate list based on the above insertion method.
[0461] The value of X can be a predefined fixed value (such as value 2), a value transmitted by signaling received by the decoder (a parameter transmitted by signaling at the sequence / strip / block / CTU level), a value that can be configured at the encoder / decoder, a value that is dynamically determined based on the number of available neighboring blocks to the left and above each individual coded block (e.g., X <= 3), or any combination of methods for determining the value of X.
[0462] The above reordering methods can be selected and applied based on different factors: In one or more embodiments, the reordering method can be selected based on the type of video frame / strip. For example, for low-latency images or stripes, all spatially non-adjacent candidates can be placed after all spatially adjacent candidates. For non-low-latency images or stripes, the first X spatially non-adjacent candidates can be placed after the spatially adjacent candidates, and the remaining spatially non-adjacent candidates can be placed after the historical BVP candidates.
[0463] It is important to note that the above candidate reordering process can be the same or different for both the IBC AMVP candidate list derivation and the IBC merged candidate list derivation. For example, for both the IBC AMVP candidate list derivation and the IBC merged candidate list derivation, all spatially non-adjacent candidates are placed after all spatially adjacent candidates. In another example, for the IBC AMVP candidate list derivation, all spatially non-adjacent candidates are placed after all spatially adjacent candidates; for the IBC merged candidate list derivation, the first X spatially non-adjacent candidates can be placed after the spatially adjacent candidates, and the remaining spatially non-adjacent candidates can be placed after the historical BVP candidates.
[0464] Sports information storage When scanning spatially non-adjacent neighbor blocks based on the candidate derivation method provided above, the selected spatially non-adjacent neighbor block can be an IBC coded block or an intra-frame TMP coded block. In the case of both IBC coded blocks and intra-frame TMP coded blocks, motion information can include translational BV.
[0465] Whether it's an IBC coded block or an intra-frame TMP coded block, the motion information of these blocks may need to be stored in memory after they are encoded. To save memory usage, spatially non-adjacent blocks may be restricted to a certain region.
[0466] like Figure 28 As shown, the allowed non-adjacent regions used for scanning spatially non-adjacent blocks can be restricted to a finite region size.
[0467] In one or more embodiments, the restricted region may be applied to the IBC or intra-frame TMP spatial neighbor block.
[0468] The size of the allowed non-adjacent regions can be defined based on the size of the current coding tree unit (CTU), such as an integer (e.g., 1 or 2 or other integers) or a fraction (e.g., 0.5 or 0.25 or other fractions) of the current CTU size.
[0469] The size of the allowed non-adjacent regions can be defined based on a fixed number of pixels or samples, for example, 128 samples above and / or to the left of the current CTU.
[0470] The size (e.g., based on the CTU size or the number of samples) can be a pre-specified value or a value determined at the encoder and carried in the bitstream by a signal transmission.
[0471] The size of the restricted area can be defined separately for the top and left non-adjacent neighbor blocks. In some examples, a non-adjacent neighbor block can be a spatially non-adjacent neighbor block, a top non-adjacent neighbor block can be a top spatially non-adjacent neighbor block, and a left non-adjacent neighbor block can be a left spatially non-adjacent neighbor block. In one example, the top non-adjacent neighbor block can be restricted within the current CTU, or outside the current CTU but within a fixed number of samples / pixels from the top of the current CTU, so that no additional line buffer is needed to store the motion information of the aforementioned non-adjacent neighbor block. For example, if the existing line buffer already covers the neighboring area of 8 sample rows from the top of the current CTU, the fixed number can be defined as 8. In another example, the left non-adjacent neighbor block can be restricted within the current CTU, or outside the current CTU but within a predefined number or a number of samples / pixels from the left boundary of the current CTU.
[0472] As shown in Figure 29, if the allowed non-adjacent region extends beyond the current CTU, the allowed non-adjacent region above the current CU (non-adjacent IBC neighbor or intra-TMP neighbor) may have a significant memory cost. In this case, the actual memory cost increases proportionally with the image width and the maximum allowable scan distance in the vertical direction. To reduce the cost of the line buffer (i.e., the non-adjacent region above the current CTU), the height of the non-adjacent region above the current CTU can be limited to a value h (as shown in Figure 29(a)). Note that this value of h can be configurable or signaled to the decoder. In cases where IBC motion and intra-TMP motion are stored in separate buffers, such as the example shown in Figure 29(b), different methods exist for storing motion in the line buffer: In one method, the line buffer used to store IBC motion can indicate that the buffer region where CU B is located is set to invalid because CU B is not an IBC CU. In another method, the line buffer used to store IBC motion can indicate that the buffer region where CU B is located is set to valid and the IBC motion is copied from CU A because CU A is the adjacent IBC neighbor of CU B.
[0473] In the example of Figure 29, the height value h and the width value w can be set to multiples of 4 for easier implementation. In one embodiment, the values h and w can be set to the minimum value of the IBC CU (e.g., 4).
[0474] When the allowed non-adjacent regions used for scanning spatially non-adjacent blocks are limited to a finite region size, such as Figure 28 As in the example in Figure 29, the scanned non-adjacent neighboring locations may be outside the allowed non-adjacent area. In this case, different methods can be used to address this issue: In one approach, the scanning process can indicate that there is no valid neighboring information at the scanned location.
[0475] In another approach, the scanning process can project or crop this out-of-range location to another location within a permissible non-adjacent area. For example... Figure 30As shown in the example, there are two locations (i.e., two points 3001) outside the permissible non-adjacent area. These two locations are projected / truncated to two other locations, which are at the same vertical / horizontal coordinates but within the permissible non-adjacent area. The new projected / truncate location can be located on the boundary of the permissible non-adjacent area closest to the original location. In cases where the permissible non-adjacent area beyond the current CTU is set to values w and h that are set to the minimum (e.g., 4) of the IBC CU, the new projected / truncate location can be interchangeably set on one boundary or the other, because the buffer is so small that it can store motion information from only one CU, and in this case, there is no difference between truncating to one boundary side or the other.
[0476] In another approach, the allowed non-adjacent spatial regions can include three regions. For example... Figure 31 As shown, these three regions are the spatial regions located outside the current CTU and adjacent to it: the upper left, left, and upper regions. The height of the upper left and upper regions, which are allowed to be non-adjacent, is defined as h, while their width can depend on the image width. The width of the allowed non-adjacent region on the left is defined as w, and its height is equal to the height of the current CTU. (See diagram). Figure 31 As shown, if the scanned positions are not adjacent ( Figure 31 If one of the points (3001) exceeds the allowed spatial area, the new location of the projection / truncation can be defined in a different way (via...). Figure 31 The dashed arrow in the diagram corresponds to point 3002 (one of the points 3001). For non-adjacent areas to the top left, the new projection / truncation position is always the pixel position adjacent to the top-left position of the current CTU. For example, if the top-left position of the current CTU is (ctu_x, ctu_y), then the new projection / truncation position is (ctu_x - 1, ctu_y - 1). For the aforementioned non-adjacent areas, the new projection / truncation position has the same horizontal coordinate, but the vertical coordinate becomes (ctu_y - 1). For non-adjacent areas to the left, the new projection / truncation position has the same vertical coordinate, but the horizontal coordinate becomes (ctu_x - 1).
[0477] When motion information of an IBC coded block is stored in memory, the motion information can be stored at the granularity of the smallest IBC block size (e.g., a 4 × 4 block). If the current IBC coded block is a coding unit larger than the smallest IBC block, a different method can be used to store the motion information. In one or more embodiments, the motion information stored at each smallest IBC block (e.g., a 4 × 4 block) within the current block is simply a duplicate copy of the motion information for the current block.
[0478] Alternatively or additionally, motion information of IBC coded blocks can be stored at different granularities a × b (e.g., 8 × 8, 8 × 16, 16 × 8, or 16 × 16, etc.) instead of the minimum IBC block size (e.g., 4 × 4 granularity), where the granularity values of a and b can be configurable or determined by the encoder and then signaled to the decoder. Without loss of generality, an 8 × 8 granularity (e.g., a = b = 8) is used as an illustrative example. If we further assume the minimum IBC block size is 4 × 4, it means that each 8 × 8 block can only store one set of IBC motion information, representing a single IBC model, even if the four 4 × 4 sub-blocks within this 8 × 8 block may come from more than one IBC block, such as... Figure 32 As shown. Figure 32 In this example, the four 4 × 4 sub-blocks A, B, C, and D form an 8 × 8 block / region, storing only one IBC model information. However, these four 4 × 4 sub-blocks come from four different IBC blocks, representing four IBC models and including four sets of IBC motion information. In this case, there may be different ways to obtain and store a set of IBC information: In one or more embodiments, a set of multiple sets of available IBC motion information can be selected and stored. In one example, IBC motion information at a fixed or configurable location (e.g., the smallest IBC block in the upper left corner) is selected for motion storage. In another example, the average IBC motion information of multiple models can be calculated for motion storage.
[0479] IBC motion information at selected neighboring IBC blocks can be simplified / compressed before storage. In one embodiment, each saved BV can be compressed before storage to further reduce memory size. One example is data compression using general techniques. For instance, such techniques are provided to save a composite value consisting of an exponent and a mantissa to approximate each saved BV.
[0480] The methods described above can be applied in any combination for storing motion information. For example, the use of restricted regions defined for non-adjacent blocks can be combined with the use of compressed IBC motion information.
[0481] Application Scope The methods described above in the "Candidate Scanning and Candidate Pruning," "Candidate Reordering," and "Motion Information Storage" sections primarily aim to derive spatially non-adjacent candidates for IBC AMVP or IBC Merge modes. When temporal candidates are also used in IBC AMVP or IBC Merge modes, they can also utilize non-adjacent neighboring blocks in co-located images. Unlike spatially non-adjacent candidates primarily extracted from the left and upper regions of the current block, temporally non-adjacent candidates can be extracted from the left, upper, right, lower, and co-located regions of the current block. Furthermore, the methods described above in the "Candidate Scanning and Candidate Pruning," "Candidate Reordering," and "Motion Information Storage" sections can also be applied in a similar manner to temporally non-adjacent candidates in IBC AMVP or IBC Merge modes.
[0482] Application block size The methods for deriving spatial / temporally non-adjacent candidates for IBC AMVP or IBC merging modes, provided in the sections on "Candidate Scanning and Candidate Pruning," "Candidate Reordering," "Motion Information Storage," and "Application Scope," can be applied to different combinations of block sizes. The applied block size for spatial / temporally adjacent candidates and the applied block size for spatial / temporally non-adjacent candidates can be the same or different. In the first example, both spatially adjacent and spatially non-adjacent candidates are applied only to coding blocks with an area greater than 16. In the second example, both spatially adjacent and spatially non-adjacent candidates are applied to all coding block sizes. In the third example, spatially adjacent candidates are applied to coding blocks with an area greater than 16, and spatially non-adjacent candidates are applied to all coding block sizes. In the fourth example, spatially adjacent candidates are applied to all coding block sizes, and spatially non-adjacent candidates are applied to coding blocks with an area greater than 16. In these four examples of applied block sizes, for scan area, scan distance, and scan order, in one example, according to... Figure 26 The candidate scanning and pruning process presented in the example derives spatially non-adjacent candidates for the IBC AMVP and IBC merge modes. In another example, based on... Figure 26 The candidate scanning and pruning process presented in the paper is used to derive spatially non-adjacent candidates for the IBCAMVP pattern, and based on... Figure 41The candidate scanning and pruning process presented above is used to derive spatially non-adjacent candidates for the IBC merging pattern. In the four application block size examples above, regarding scan termination, in one example, a scan performed in a specific region at a specific distance stops when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or when the maximum allowed number of IBC merging candidates has been reached. In another example, a scan performed in a specific region at a specific distance stops when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or when the maximum allowed number of IBC merging candidates has been reduced by 1. The above examples of application block size, scan region, scan distance and scan order, and scan termination can be freely combined to constitute different examples of entire schemes utilizing non-adjacent candidates in IBC.
[0483] Sub-block-based IBC mode According to one or more embodiments of this disclosure, the IBC mode is extended to the sub-block level, wherein sub-blocks within the current block can have their own BV for motion compensation. Different example methods can be used to achieve this goal.
[0484] In the first example method, such as Figure 33 As shown, the BV of a sub-block in the current block is obtained by reusing the BV of a sub-block within the same block in the same frame. If the BV of a sub-block within the same block cannot be obtained, such as if the sub-block is intra-coded, the BV of the sub-block can be set to the BV of the same block. Besides the same block in the same frame, blocks at other locations can also be used to obtain sub-block-level BVs, where non-temporally adjacent candidates can be referenced. Additionally, template matching methods can be used to refine the BVs of the same block in the same frame or sub-blocks at other locations.
[0485] In the second example method, such as Figure 34 As shown, the BV of the left or top child block in the current block is obtained by refining the BV of the current block using template matching. The BV of the current block can be obtained through regular IBC mode, TM IBC mode, or other modes. For the left child block in the current block, only the left template can be used to refine its BV. For the top child block in the current block, only the top template can be used to refine its BV. For the top-left child block in the current block, both the left and top templates can be used to refine its BV.
[0486] Multiple template patterns for TM IBC mode and TM regular mode According to one or more embodiments of this disclosure, the template patterns of TM IBC mode and TM regular mode are expanded by importing more types of templates. Different methods can be used to achieve this goal.
[0487] In the first approach, in addition to the currently used template mode that uses both the top and left templates, other template modes (such as using only the left template, using only the top template, etc.) can also be used in TM IBC mode or TM regular inter-frame mode. When more than one template mode is used in TM IBC mode or TM regular inter-frame mode, the final template mode to be used for the current block can be determined based on predefined standards or flags transmitted in the bitstream.
[0488] IBC and Intra-Frame TMP for Decongestion of Deblocking Filter Boundary Strength According to one or more embodiments of this disclosure, when the boundary strength is obtained, the deblocking filter treats blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode equally. Different methods can be used to achieve this objective.
[0489] In the first method, when the boundary strength of the deblocking filter is obtained, blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode are both considered blocks encoded in intra-frame mode. The boundary strength criterion for blocks encoded in intra-frame mode can then be applied to both IBC and intra-frame TMP blocks. For example, if two neighboring blocks are both encoded in IBC or intra-frame TMP mode, or if one of the two neighboring blocks is encoded in either IBC or intra-frame TMP mode, the boundary strength is set to a predetermined positive integer, such as 2.
[0490] In the second method, when obtaining the boundary strength of the deblocking filter, blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode are both considered as blocks encoded in inter-frame mode. The boundary strength criterion for blocks encoded in inter-frame mode can then be applied to blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode. For example, if two neighboring blocks are both encoded in IBC mode, or both neighboring blocks are both encoded in intra-frame TMP mode, or one of two neighboring blocks is encoded in IBC mode and the other in intra-frame TMP mode, then if the two block vectors of these two neighboring blocks are different (or the absolute difference between the horizontal or vertical components of these two block vectors is greater than a threshold (e.g., half a pixel)), the boundary strength is set to a predetermined positive integer, such as 1. In another example, if the difference between the two block vectors of these two neighboring blocks is large, the boundary strength is set to a larger value.
[0491] In the third method, when a block encoded in IBC mode or a block encoded in intra-frame TMP mode is combined with an intra-frame tool (such as combining IBC with intra-frame prediction, i.e., using GPM with IBC and intra-frame prediction, or combining intra-frame TMP with intra-frame prediction, i.e., using GPM with intra-frame TMP and intra-frame prediction, etc.), the block is considered to be a block encoded in intra-frame mode when the boundary strength of the deblocking filter is obtained.
[0492] In the fourth method, when a block encoded in IBC mode or a block encoded in intra-frame TMP mode is combined with an inter-frame tool (such as combining IBC with inter-frame prediction, i.e., using GPM of IBC and inter-frame prediction, or combining intra-frame TMP with inter-frame prediction, i.e., using GPM of intra-frame TMP and inter-frame prediction, etc.), the block is considered to be a block encoded in inter-frame mode when the boundary strength of the deblocking filter is obtained.
[0493] Combination of IBC and intra-frame TMP According to one or more embodiments of this disclosure, the IBC tool is combined with the intra-frame TMP tool. Different methods can be used to achieve this goal.
[0494] In the first approach, the IBC tool is combined with the intra-TMP tool in the form of a GPM utilizing IBC and intra-TMP prediction. In one example, the two “inter-frame” portions of the GPM utilizing inter-frame and inter-prediction methods in VVC are replaced with IBC and intra-TMP predictions, respectively. This means that the IBC prediction results and the intra-TMP prediction results are weighted and averaged against each other according to the segmentation lines in the coded block. The weights can be obtained by referring to the GPM utilizing inter-frame and inter-prediction methods in VVC. In another example, the two “inter-frame” portions of the GPM utilizing inter-frame and inter-prediction methods in ECM are replaced with IBC and intra-TMP predictions, where template matching tools can be used to further improve encoding and decoding performance.
[0495] In the second approach, the IBC tool and the intra-TMP tool are combined in the form of CIIP, which utilizes predictions from both IBC and intra-TMP. In one example, fixed weights are used to weight the IBC and intra-TMP predictions, where the specific weight values depend on the coding patterns of neighboring blocks. In another example, fixed weight values are obtained based on a template matching method.
[0496] TM-based reordering of MMVD candidates for GPM using MMVD model According to one or more embodiments of this disclosure, a template matching (TM)-based method is used to reorder MMVD candidates for GPM utilizing MMVD patterns. Different methods may be used to achieve this objective.
[0497] In the first method, similar to GPM using template matching, different templates are used to reorder the MMVD candidates for different parts of a GPM partition with a predefined partitioning pattern. After reordering, a signal transmission index is used to represent the selected MMVD candidates for a portion of the GPM partition.
[0498] In the second method, a template is used to reorder the MMVD candidates for different parts of the GPM partition under different partitioning modes. After reordering, a signal transmission index is used to represent the selected MMVD candidates for a portion of the GPM partition.
[0499] IBC / Regular Inter-Frame / Affine Inter-Frame HMVP Candidate Utilization According to one or more embodiments of this disclosure, HMVP candidates are utilized based on their relative position to the current block in IBC / regular inter-frame / affine inter-frame formats. Different methods can be used to achieve this goal.
[0500] In the first method, when encoding the current block, all IBC / regular inter-frame / affine inter-frame HMVP candidates are first divided into several groups based on the distance between the center position of the HMVP candidate and the center position of the current block. For example, some predefined thresholds are used. , … , will have less than HMVP candidates with a distance equal to or greater than [a certain value] are divided into group 1, which includes those with a distance equal to or greater than [a certain value]. But smaller than HMVP candidates with a distance equal to or greater than 1 are divided into groups 2, ..., and those with a distance equal to or greater than 1 are divided into groups 2, ... But smaller than The HMVP candidates are divided into groups N based on their distance. Different methods can be used to set the specific value of the predefined threshold. For example, It is set to the maximum value of the current block's height and width. Set as Twice as much, ... Set as N times. Then, in each partitioned group, the first one is selected in a predefined scan order. Valid candidates. Predefined scan order can be defined in different ways. For example, HMVP candidates in each partition group can be scanned in the following order: left, top, bottom left, top right, top left. The value can be set to the same value in all partitioned groups, or it can be set to different values in different partitioned groups. Finally, the selected candidates from each partitioned group are added to the IBC / regular inter-frame / affine inter-frame candidate list in order from smaller distance to larger distance.
[0501] In the second method, after dividing all IBC / regular inter-frame / affine inter-frame HMVP candidates into several groups based on the distance between the center position of the HMVP candidate and the center position of the current block, the first candidate in each group is selected based on the template matching method. Valid candidates. Specifically, HMVP candidates in each partition group are reordered based on a template matching method, and then the first one with the minimum template matching cost is selected. Valid candidates. Finally, the selected candidates from each partition group are added to the IBC / regular inter-frame / affine inter-frame candidate list in order from smaller to larger distances.
[0502] When utilizing HMVP candidates based on the encoding order of the HMVP candidates (the original scheme in VVC or ECM) or based on the relative position of the HMVP candidates to the current block (the provided scheme), the number of HMVP candidates can be further increased for IBC / regular inter-frame / affine inter-frame HMVP candidates. This can be achieved using different methods. In one example, similar to the original design in VVC or ECM, only one HMVP table exists to store HMVP candidates, but the size of the HMVP table is further increased. In another example, several HMVP tables exist to store HMVP candidates; different HMVP tables can be assigned to different CTUs, or different HMVP tables can be assigned to different regions of the encoded image. For example, there are 5 HMVP tables, one assigned to the current CTU, one to the left neighboring CTU, one to the top neighboring CTU, one to the top right neighboring CTU, and one to the top left neighboring CTU. The HMVP table size can be set to the same number for different HMVP tables, or different numbers for different HMVP tables.
[0503] When managing IBC HMVP candidates, coded blocks encoded using either IBC mode or intra-TMP mode can be added to the IBC HMVP table, or only coded blocks encoded using IBC mode can be added. When adding a coded block to the IBC HMVP table, coded blocks with all block sizes can be added, or only coded blocks with an area greater than 16 can be added.
[0504] Construction of the intra-MPM list associated with IBC or intra-TMP coded blocks According to one or more embodiments of this disclosure, during the construction of the intra-MPM list for the current block, when a neighboring block is IBC-coded or intra-TMP-coded, the intra-prediction mode pointed to by the block vector of the IBC or intra-TMP is used to construct the intra-MPM list for the current block. Different methods can be used to achieve this objective.
[0505] In the first method, when the current block is encoded using a regular intra-frame mode, a multi-reference line intra-frame mode, a GPM using inter-frame and intra-frame modes, a spatial GPM mode, or a GPM using IBC and intra-frame modes to construct the intra-frame MPM list of the current block, when a neighboring block is IBC encoded or intra-frame TMP encoded, the intra-frame prediction mode pointed to by the block vector of the IBC or intra-frame TMP is used instead of setting the intra-frame prediction mode of the IBC or intra-frame TMP encoded block to a planar mode.
[0506] Figure 35 A computing environment (or computing device) 1610 coupled to a user interface 1650 is shown. The computing environment 1610 may be part of a data processing server. In some embodiments, the computing device 1610 may perform any of the various methods or processes (such as encoding / decoding methods or processes) described above according to various examples of this disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.
[0507] Processor 1620 typically controls the overall operation of computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1620 may include one or more processors for executing instructions to perform all or some of the steps described above. Furthermore, processor 1620 may include one or more modules that facilitate interaction between processor 1620 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0508] Memory 1630 is configured to store various types of data to support the operation of computing environment 1610. Memory 1630 may include predefined software 1632. Examples of such data include instructions for any application or method operating on computing environment 1610, video datasets, image data, etc. Memory 1630 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0509] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 can be coupled to encoders and decoders.
[0510] Figure 42 A flowchart of a video decoding method according to an example of this disclosure is shown.
[0511] In step 4210, the processor 1620 may obtain one or more IBC candidates from a plurality of non-adjacent neighboring blocks on the decoder side according to a scanning rule corresponding to the intra-block copy (IBC) mode, so as to apply the IBC mode to the current block, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block.
[0512] In step 4220, the processor 1620 can obtain the prediction of the current block on the decoder side by applying the IBC mode based on one or more IBC candidates.
[0513] In one or more examples, the scanning rules include: determining at least one scanning region based on the maximum scanning distance indicating the maximum number of blocks to one side of the current block; and determining the scanning position within the scanning region based on the orientation of the scanning region relative to the current block.
[0514] In one or more examples, the method further includes: determining the maximum scan distance by the decoder as 4 blocks to the left or top of the current block; or, determining the maximum scan distance by the decoder as 7 blocks to the left or top of the current block.
[0515] In one or more examples, determining the maximum scan distance as 7 blocks to the left or top of the current block includes: in response to determining that an IBC merge mode should be applied to the current block, determining the maximum scan distance as 7 blocks to the left or top of the current block.
[0516] In one or more examples, determining the maximum scan distance as four blocks to the left or top of the current block includes: in response to determining that the IBC Adaptive Motion Vector Prediction (AMVP) mode should be applied to the current block, determining the maximum scan distance as four blocks to the left or top of the current block.
[0517] In one or more examples, determining the scan position in the scan region based on the orientation of the scan region relative to the current block includes: determining the scan position as the lower right scan position in the scan region in response to determining that the scan region is in the lower left, upper right, or upper left direction relative to the current block; determining the scan position as the lower center scan position in the scan region in response to determining that the scan region is in the upper direction relative to the current block; or determining the scan position as the middle right scan position in the scan region in response to determining that the scan region is in the left direction relative to the current block.
[0518] In one or more examples, the scanning rule further includes terminating scanning in at least one scanning region in response to the following conditions: the size of the candidate list, which includes one or more IBC candidates, has reached a fixed size; or all non-adjacent neighboring blocks within the maximum scan distance have been scanned.
[0519] Figure 43 It shows the corresponding to, for example Figure 42 The flowchart of the video encoding method shown in the video decoding method is as follows.
[0520] In step 4310, the processor 1620 may obtain one or more IBC candidates from a plurality of non-adjacent neighboring blocks on the encoder side according to a scanning rule corresponding to the intra-block copy (IBC) mode, so as to apply the IBC mode to the current block, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block.
[0521] In step 4320, the processor 1620 can obtain the prediction of the current block on the encoder side by applying the IBC mode based on one or more IBC candidates.
[0522] In one or more examples, the scanning rules include: determining at least one scanning region based on the maximum scanning distance indicating the maximum number of blocks to one side of the current block; and determining the scanning position within the scanning region based on the orientation of the scanning region relative to the current block.
[0523] In one or more examples, the method further includes: determining the maximum scan distance by the encoder as 4 blocks to the left or above the current block; or, determining the maximum scan distance by the encoder as 7 blocks to the left or above the current block.
[0524] In one or more examples, determining the maximum scan distance as 7 blocks to the left or top of the current block includes: determining the maximum scan distance as 7 blocks to the left or top of the current block in response to determining that an IBC merge mode should be applied to the current block.
[0525] In one or more examples, determining the maximum scan distance as four blocks to the left or top of the current block includes: in response to determining that the IBC Adaptive Motion Vector Prediction (AMVP) mode should be applied to the current block, determining the maximum scan distance as four blocks to the left or top of the current block.
[0526] In one or more examples, determining the scan position in the scan region based on the orientation of the scan region relative to the current block includes: determining the scan position as the lower right scan position in the scan region in response to determining that the scan region is in the lower left, upper right, or upper left direction relative to the current block; determining the scan position as the lower center scan position in the scan region in response to determining that the scan region is in the upper direction relative to the current block; or determining the scan position as the middle right scan position in the scan region in response to determining that the scan region is in the left direction relative to the current block.
[0527] In one or more examples, the scanning rule further includes terminating scanning in at least one scanning region in response to the following conditions: the size of the candidate list, which includes one or more IBC candidates, has reached a fixed size; or all non-adjacent neighboring blocks within the maximum scan distance have been scanned.
[0528] Figure 44 A flowchart of a video decoding method according to an example of this disclosure is shown.
[0529] In step 4410, the processor 1620 may obtain one or more intra-block copy (IBC) candidates from multiple neighboring blocks on the decoder side according to the size of the current block, so as to apply the IBC mode to the current block.
[0530] In step 4420, the processor 1620 can obtain a prediction of the current block on the decoder side by applying an IBC mode based on one or more IBC candidates.
[0531] In one or more examples, obtaining one or more IBC candidates from multiple neighboring blocks based on the size of the current block includes: obtaining one or more IBC candidates by scanning multiple non-adjacent neighboring blocks and multiple neighboring blocks in response to determining that the size of the current block is greater than a predefined size; obtaining one or more IBC candidates by scanning at least multiple non-adjacent neighboring blocks or multiple neighboring blocks in response to determining that the size of the current block is arbitrary; obtaining one or more IBC candidates by scanning multiple neighboring blocks in response to determining that the size of the current block is greater than a predefined size; or obtaining one or more IBC candidates by scanning multiple non-adjacent blocks in response to determining that the size of the current block is greater than a predefined size, wherein the multiple non-adjacent neighboring blocks are at least one block away from the current block.
[0532] In one or more examples, multiple non-adjacent neighboring blocks are configured to be scanned by the decoder according to scanning rules corresponding to the IBC mode.
[0533] In one or more examples, the scanning rules include: determining at least one scanning region by the decoder based on the maximum scan distance indicating the maximum number of blocks to one side of the current block; and determining the scan position within the scanning region by the decoder based on the orientation of the scanning region relative to the current block.
[0534] In one or more examples, the method further includes: determining the maximum scan distance by the decoder as 4 blocks to the left or top of the current block; or, determining the maximum scan distance by the decoder as 7 blocks to the left or top of the current block.
[0535] In one or more examples, determining the maximum scan distance as 7 blocks to the left or top of the current block includes: in response to determining that an IBC merge mode should be applied to the current block, determining the maximum scan distance as 7 blocks to the left or top of the current block.
[0536] In one or more examples, determining the maximum scan distance as 4 blocks to the left or top of the current block includes: in response to determining that the IBC AMVP mode should be applied to the current block, determining the maximum scan distance as 4 blocks to the left or top of the current block.
[0537] In one or more examples, determining the scan position in the scan region based on the orientation of the scan region relative to the current block includes: determining the scan position as the lower right scan position in the scan region in response to determining that the scan region is in the lower left, upper right, or upper left direction relative to the current block; determining the scan position as the lower center scan position in the scan region in response to determining that the scan region is in the upper direction relative to the current block; or determining the scan position as the middle right scan position in the scan region in response to determining that the scan region is in the left direction relative to the current block.
[0538] In one or more examples, the scanning rule further includes terminating scanning in at least one scanning region in response to the following conditions: the size of the candidate list, which includes one or more IBC candidates, has reached a fixed size; or all non-adjacent neighboring blocks within the maximum scan distance have been scanned.
[0539] Figure 45 It shows the corresponding to, for example Figure 44 The flowchart of the video encoding method shown in the video decoding method is as follows.
[0540] In step 4510, the processor 1620 may obtain one or more intra-block copy (IBC) candidates from multiple neighboring blocks on the encoder side according to the size of the current block, so as to apply the IBC mode to the current block.
[0541] In step 4520, the processor 1620 can obtain the prediction of the current block on the encoder side by applying the IBC mode based on one or more IBC candidates.
[0542] In one or more examples, obtaining one or more IBC candidates from multiple neighboring blocks based on the size of the current block includes: obtaining one or more IBC candidates by scanning multiple non-adjacent neighboring blocks and multiple neighboring blocks in response to determining that the size of the current block is greater than a predefined size; obtaining one or more IBC candidates by scanning at least multiple non-adjacent neighboring blocks or multiple neighboring blocks in response to determining that the size of the current block is arbitrary; obtaining one or more IBC candidates by scanning multiple neighboring blocks in response to determining that the size of the current block is greater than a predefined size; or obtaining one or more IBC candidates by scanning multiple non-adjacent blocks in response to determining that the size of the current block is greater than a predefined size, wherein the multiple non-adjacent neighboring blocks are at least one block away from the current block.
[0543] In one or more examples, multiple non-adjacent neighboring blocks are configured to be scanned by the encoder according to scanning rules corresponding to the IBC mode.
[0544] In one or more examples, the scanning rules include: determining at least one scanning region by the encoder based on the maximum scanning distance indicating the maximum number of blocks to one side of the current block; and determining the scanning position in the scanning region by the encoder based on the orientation of the scanning region relative to the current block.
[0545] In one or more examples, the method further includes: determining the maximum scan distance by the encoder as 4 blocks to the left or above the current block; or, determining the maximum scan distance by the encoder as 7 blocks to the left or above the current block.
[0546] In one or more examples, determining the maximum scan distance as 7 blocks to the left or top of the current block includes: in response to determining that an IBC merge mode should be applied to the current block, determining the maximum scan distance as 7 blocks to the left or top of the current block.
[0547] In one or more examples, determining the maximum scan distance as 4 blocks to the left or top of the current block includes: in response to determining that the IBC AMVP mode should be applied to the current block, determining the maximum scan distance as 4 blocks to the left or top of the current block.
[0548] In one or more examples, determining the scan position in the scan region based on the orientation of the scan region relative to the current block includes: determining the scan position as the lower right scan position in the scan region in response to determining that the scan region is in the lower left, upper right, or upper left direction relative to the current block; determining the scan position as the lower center scan position in the scan region in response to determining that the scan region is in the upper direction relative to the current block; or determining the scan position as the middle right scan position in the scan region in response to determining that the scan region is in the left direction relative to the current block.
[0549] In one or more examples, the scanning rule further includes terminating scanning in at least one scanning region in response to the following conditions: the size of the candidate list, which includes one or more IBC candidates, has reached a fixed size; or all non-adjacent neighboring blocks within the maximum scan distance have been scanned.
[0550] Figure 46 A flowchart of a video decoding method according to an example of this disclosure is shown.
[0551] In step 4610, the processor 1620 can obtain a list of historical motion vector prediction values (HMVP) on the decoder side based on at least one block attribute of the current block.
[0552] In step 4620, the processor 1620 can obtain the prediction of the current block on the decoder side by applying the intra-block copy (IBC) mode based on the HMVP list.
[0553] In one or more examples, at least one block attribute includes any one or any combination of the following: the block position of the current block, the block size, or the image region.
[0554] In one or more examples, obtaining an HMVP list based on at least one block attribute of the current block includes obtaining different HMVP lists for different coding tree blocks that have different block positions or are located in different image regions.
[0555] In one or more examples, different HMVP lists may have the same list size or different list sizes.
[0556] In one or more examples, the HMVP list includes any one or any combination of the following: blocks decoded in IBC mode; blocks decoded in intra template matching prediction (TMP) mode; or blocks decoded with a block size larger than a predefined size.
[0557] Figure 47 It shows the corresponding to, for example Figure 46 The flowchart of the video encoding method shown in the video decoding method is as follows.
[0558] In step 4710, the processor 1620 can obtain a list of historical motion vector prediction values (HMVP) on the encoder side based on at least one block attribute of the current block.
[0559] In step 4720, the processor 1620 can obtain the prediction of the current block on the encoder side by applying the intra-block copy (IBC) mode based on the HMVP list.
[0560] In one or more examples, at least one block attribute includes any one or any combination of the following: the block position, block size, or image region of the current block.
[0561] In one or more examples, obtaining an HMVP list based on at least one block attribute of the current block includes obtaining different HMVP lists for different coding tree blocks that have different block positions or are located in different image regions.
[0562] In one or more examples, different HMVP lists may have the same list size or different list sizes.
[0563] In one or more examples, the HMVP list includes any one or any combination of the following: blocks encoded in IBC mode; blocks encoded in intra template matching prediction (TMP) mode; or blocks encoded in a block size larger than a predefined size.
[0564] Figure 48 A flowchart of a video decoding method according to an example of this disclosure is shown.
[0565] In step 4810, the processor 1620 may, on the decoder side, obtain an intra-prediction mode associated with the block vector of the neighboring block in response to determining that the neighboring blocks of the current block are encoded and decoded in either IBC mode or intra template matching prediction (TMP) mode.
[0566] In step 4820, the processor 1620 can construct a list of intra-most probable modes (MPMs) for the current block on the decoder side based on the intra-prediction mode.
[0567] Figure 49 It shows the corresponding to, for example Figure 49 The flowchart of the video encoding method shown in the video decoding method is as follows.
[0568] In step 4910, the processor 1620 may, on the encoder side, obtain an intra-prediction mode associated with the block vector of the neighboring block in response to determining that the neighboring blocks of the current block are encoded and decoded in either IBC mode or intra-template matching prediction (TMP) mode.
[0569] In step 4920, the processor 1620 can construct a list of intra-most probable modes (MPMs) for the current block on the encoder side based on the intra-prediction mode.
[0570] In some examples, an apparatus for video encoding and decoding is provided. The apparatus includes a processor 1620 and a memory 1640, the memory being configured to store instructions executable by the processor; wherein the processor, when executing the instructions, is configured to perform actions such as... Figures 42 to 49 Any of the methods shown.
[0571] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a memory 1630 containing multiple programs and / or storing a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The multiple programs can be executed by a processor 1620 in a computing environment 1610 to perform the above-described methods. In one example, the multiple programs can be executed by a processor 1620 in a computing environment 1610 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 1610 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-described decoding method based on the received bitstream or data stream. In another example, multiple programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bit stream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by an encoder (e.g., Figure 2 The video encoder 20 in the video encoder (e.g., the one described above) generates the video for the decoder (e.g., the one described above) using the encoding method described above. Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media may be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0572] In one embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method is provided.
[0573] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0574] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above, is also provided. In another embodiment, a computer program product including, for example, multiple programs in a memory 1630, which can be executed by a processor 1620 in a computing environment 1610 to perform the methods described above, is also provided. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0575] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0576] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.
[0577] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.
[0578] The description in this disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0579] Unless otherwise specified, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0580] The examples were chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A video decoding method, comprising: A decoder for applying an intra-block copy (IBC) mode to the current block obtains one or more IBC candidates from a plurality of non-adjacent neighboring blocks according to a scanning rule corresponding to the IBC mode, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block; and The decoder obtains the prediction of the current block by applying the IBC mode based on the one or more IBC candidates.
2. The method of claim 1, wherein, The scanning rules include: At least one scan region is determined based on the maximum scan distance indicating the maximum number of blocks to one side of the current block; and The scanning position in the scanning area is determined based on the orientation of the scanning area relative to the current block.
3. The method of claim 2, further comprising: The decoder determines the maximum scanning distance to be 4 blocks to the left or top of the current block; or The decoder determines the maximum scanning distance to be 7 blocks to the left or top of the current block.
4. The method of claim 3, wherein, Determining the maximum scanning distance as 7 blocks to the left or top of the current block includes: In response to determining that the IBC merge mode should be applied to the current block, the maximum scan distance is determined to be 7 blocks to the left or top of the current block.
5. The method of claim 3, wherein, Determining the maximum scanning distance as four blocks to the left or top of the current block includes: In response to determining that the IBC Adaptive Motion Vector Prediction (AMVP) mode should be applied to the current block, the maximum scan distance is determined to be 4 blocks to the left or top of the current block.
6. The method of claim 2, wherein, Determining the scan position in the scan area based on the orientation of the scan area relative to the current block includes: In response to determining that the scanned area is located in the lower left, upper right, or upper left direction relative to the current block, the scanned position is determined as the lower right scanned position in the scanned area; In response to determining that the scan region is located in the upward direction relative to the current block, the scan position is determined as the lower-middle scan position within the scan region; or In response to determining that the scanned region is located to the left relative to the current block, the scanned position is determined as the middle right scanned position within the scanned region.
7. The method of claim 2, wherein, The scanning rules further include: The scanning in the at least one scanning region is terminated in response to the following conditions being met: The size of the candidate list, including one or more IBC candidates, has reached a fixed size; or All non-adjacent neighboring blocks within the maximum scanning distance have been scanned.
8. A video decoding method, comprising: The decoder, which applies the intra-block copy (IBC) mode to the current block, obtains one or more IBC candidates from multiple neighboring blocks based on the size of the current block; as well as The decoder obtains the prediction of the current block by applying the IBC mode based on the one or more IBC candidates.
9. The method of claim 8, wherein, Obtaining one or more IBC candidates from multiple neighboring blocks based on the size of the current block includes: In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple non-adjacent neighboring blocks and multiple neighboring blocks; In response to determining that the size of the current block is arbitrary, the one or more IBC candidates are obtained by scanning at least a plurality of non-adjacent neighboring blocks or a plurality of neighboring blocks; In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple adjacent blocks; or In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple non-adjacent blocks. Among them, the plurality of non-adjacent neighboring blocks are at least one block away from the current block.
10. The method of claim 9, wherein, The plurality of non-adjacent neighboring blocks are configured to be scanned by the decoder according to scanning rules corresponding to the IBC mode.
11. The method of claim 10, wherein, The scanning rules include: The decoder determines at least one scan region based on the maximum scan distance indicating the maximum number of blocks to one side of the current block; and The decoder determines the scanning position in the scanning region based on the orientation of the scanning region relative to the current block.
12. The method of claim 11, further comprising: The decoder determines the maximum scanning distance to be 4 blocks to the left or top of the current block; or The decoder determines the maximum scanning distance to be 7 blocks to the left or top of the current block.
13. The method of claim 12, wherein, Determining the maximum scanning distance as 7 blocks to the left or top of the current block includes: In response to determining that the IBC merge mode should be applied to the current block, the maximum scan distance is determined to be 7 blocks to the left or top of the current block.
14. The method of claim 12, wherein, Determining the maximum scanning distance as four blocks to the left or top of the current block includes: In response to determining that the IBC AMVP mode should be applied to the current block, the maximum scan distance is determined to be 4 blocks to the left or top of the current block.
15. The method of claim 11, wherein, Determining the scan position in the scan area based on the orientation of the scan area relative to the current block includes: In response to determining that the scanned area is located in the lower left, upper right, or upper left direction relative to the current block, the scanned position is determined as the lower right scanned position in the scanned area; In response to determining that the scan region is located in the upward direction relative to the current block, the scan position is determined as the lower-middle scan position within the scan region; or In response to determining that the scanned region is located to the left relative to the current block, the scanned position is determined as the middle right scanned position within the scanned region.
16. The method of claim 11, wherein, The scanning rules further include: The scanning in the at least one scanning region is terminated in response to the following conditions being met: The size of the candidate list, including one or more IBC candidates, has reached a fixed size; or All non-adjacent neighboring blocks within the maximum scanning distance have been scanned.
17. A video coding method, comprising: An encoder for applying an intra-block copy (IBC) mode to the current block obtains one or more IBC candidates from a plurality of non-adjacent neighboring blocks according to a scanning rule corresponding to the IBC mode, wherein the plurality of non-adjacent neighboring blocks are at least one block away from the current block; and The encoder obtains the prediction of the current block by applying the IBC mode based on the one or more IBC candidates.
18. The method of claim 17, wherein, The scanning rules include: At least one scan region is determined based on the maximum scan distance indicating the maximum number of blocks to one side of the current block; and The scanning position in the scanning area is determined based on the orientation of the scanning area relative to the current block.
19. The method of claim 18, further comprising: The encoder determines the maximum scanning distance to be 4 blocks to the left or above the current block; or The encoder determines the maximum scanning distance to be 7 blocks to the left or above the current block.
20. The method of claim 19, wherein, Determining the maximum scanning distance as 7 blocks to the left or top of the current block includes: In response to determining that the IBC merge mode should be applied to the current block, the maximum scan distance is determined to be 7 blocks to the left or top of the current block.
21. The method of claim 19, wherein, Determining the maximum scanning distance as four blocks to the left or top of the current block includes: In response to determining that the IBC Adaptive Motion Vector Prediction (AMVP) mode should be applied to the current block, the maximum scan distance is determined to be 4 blocks to the left or top of the current block.
22. The method of claim 18, wherein, Determining the scan position in the scan area based on the orientation of the scan area relative to the current block includes: In response to determining that the scanned area is located in the lower left, upper right, or upper left direction relative to the current block, the scanned position is determined as the lower right scanned position in the scanned area; In response to determining that the scan region is located in the upward direction relative to the current block, the scan position is determined as the lower-middle scan position within the scan region; or In response to determining that the scanned region is located to the left relative to the current block, the scanned position is determined as the middle right scanned position within the scanned region.
23. The method of claim 18, wherein, The scanning rules further include: The scanning in the at least one scanning region is terminated in response to the following conditions being met: The size of the candidate list, including one or more IBC candidates, has reached a fixed size; or All non-adjacent neighboring blocks within the maximum scanning distance have been scanned.
24. A video coding method, comprising: The encoder, which applies the Intra-Block Copy (IBC) mode to the current block, obtains one or more IBC candidates from multiple neighboring blocks based on the size of the current block; as well as The encoder obtains the prediction of the current block by applying the IBC mode based on the one or more IBC candidates.
25. The method of claim 24, wherein, Obtaining one or more IBC candidates from multiple neighboring blocks based on the size of the current block includes: In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple non-adjacent neighboring blocks and multiple neighboring blocks; In response to determining that the size of the current block is arbitrary, the one or more IBC candidates are obtained by scanning at least a plurality of non-adjacent neighboring blocks or a plurality of neighboring blocks; In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple adjacent blocks; or In response to determining that the size of the current block is greater than a predefined size, the one or more IBC candidates are obtained by scanning multiple non-adjacent blocks. Among them, the plurality of non-adjacent neighboring blocks are at least one block away from the current block.
26. The method of claim 25, wherein, The plurality of non-adjacent neighboring blocks are configured to be scanned by the encoder according to scanning rules corresponding to the IBC mode.
27. The method of claim 26, wherein, The scanning rules include: The encoder determines at least one scan area based on the maximum scan distance indicating the maximum number of blocks to one side of the current block; and The encoder determines the scanning position in the scanning area based on the orientation of the scanning area relative to the current block.
28. The method of claim 27, further comprising: The encoder determines the maximum scanning distance to be 4 blocks to the left or above the current block; or The encoder determines the maximum scanning distance to be 7 blocks to the left or above the current block.
29. The method of claim 28, wherein, Determining the maximum scanning distance as 7 blocks to the left or top of the current block includes: In response to determining that the IBC merge mode should be applied to the current block, the maximum scan distance is determined to be 7 blocks to the left or top of the current block.
30. The method of claim 28, wherein, Determining the maximum scanning distance as four blocks to the left or top of the current block includes: In response to determining that the IBC AMVP mode should be applied to the current block, the maximum scan distance is determined to be 4 blocks to the left or top of the current block.
31. The method of claim 27, wherein, Determining the scan position in the scan area based on the orientation of the scan area relative to the current block includes: In response to determining that the scanned area is located in the lower left, upper right, or upper left direction relative to the current block, the scanned position is determined as the lower right scanned position in the scanned area; In response to determining that the scan region is located in the upward direction relative to the current block, the scan position is determined as the lower-middle scan position within the scan region; or In response to determining that the scanned region is located to the left relative to the current block, the scanned position is determined as the middle right scanned position within the scanned region.
32. The method of claim 27, wherein, The scanning rules further include: The scanning in the at least one scanning region is terminated in response to the following conditions being met: The size of the candidate list, including one or more IBC candidates, has reached a fixed size; or All non-adjacent neighboring blocks within the maximum scanning distance have been scanned.
33. A video decoding method, comprising: The decoder obtains a list of historical motion vector prediction values (HMVP) based on at least one block attribute of the current block; as well as The decoder obtains the prediction of the current block by applying the Intra-Block Copy (IBC) mode based on the HMVP list.
34. The method of claim 33, wherein, The at least one block attribute includes any one or any combination of the following: The current block's position, block size, or image area.
35. The method of claim 33, wherein, The HMVP list is obtained based on at least one block attribute of the current block, including: Different HMVP lists are obtained for different coding tree blocks with different block positions or located in different image regions.
36. The method of claim 35, wherein, The different HMVP lists may have the same list size or different list sizes.
37. The method of claim 33, wherein, The list of HMVPs includes any one or any combination of the following: Blocks decoded using IBC mode; Blocks decoded using intra-frame template matching prediction (TMP) mode; or Decode blocks using a fast size larger than the predefined size.
38. A video coding method, comprising: The encoder obtains a list of historical motion vector prediction values (HMVP) based on at least one block attribute of the current block; as well as The encoder obtains the prediction of the current block by applying the Intra-Block Copy (IBC) mode based on the HMVP list.
39. The method of claim 38, wherein, The at least one block attribute includes any one or any combination of the following: The current block's position, block size, or image area.
40. The method of claim 38, wherein, The HMVP list is obtained based on at least one block attribute of the current block, including: Different HMVP lists are obtained for different coding tree blocks with different block positions or located in different image regions.
41. The method of claim 40, wherein, The different HMVP lists may have the same list size or different list sizes.
42. The method of claim 38, wherein, The list of HMVPs includes any one or any combination of the following: Blocks encoded in IBC mode; Blocks encoded using intra-frame template matching prediction (TMP) mode; or Blocks encoded with a block size larger than the predefined size.
43. A video decoding method, comprising: In response to determining that the neighboring blocks of the current block are encoded and decoded in either IBC mode or Intra Template Matching Prediction (TMP) mode, the decoder obtains the intra prediction mode associated with the block vector of the neighboring blocks. as well as The decoder constructs a list of most probable intra modes (MPMs) for the current block based on the intra prediction mode.
44. A video coding method, comprising: In response to determining that the neighboring blocks of the current block are encoded and decoded in either IBC mode or Intra Template Matching Prediction (TMP) mode, the encoder obtains the intra prediction mode associated with the block vector of the neighboring blocks. as well as The encoder constructs a list of most probable intra modes (MPMs) for the current block based on the intra prediction mode.
45. A video decoding apparatus, comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 1-16, 33-37 or 43 when executing the instructions.
46. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 1-16, 33-37, or 43.
47. A video encoding apparatus, comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 17-32, 38-42 or 44 when executing the instructions.
48. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 17-32, 38-42, or 44.
49. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1-16, 33-37 or 43.
50. A non-transitory computer-readable storage medium for storing a bit stream generated by the method of any one of claims 17-32, 38-42 or 44.