Method and apparatus for intra block copy and intra template matching
By applying deblocking filters in intra-frame or inter-frame modes and processing video blocks based on boundary strength, the problems of low prediction efficiency in intra-frame block copying and intra-frame template matching are solved, achieving more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202480024737.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-10
- Filing Date
- 2024-04-10
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in terms of intra-frame block copying and intra-frame template matching prediction, making it difficult to effectively reduce video data redundancy and improve encoding and decoding efficiency.
A deblocking filter is used to apply predefined standards in intra-frame or inter-frame modes, and video blocks are processed based on boundary strength. Encoding and decoding are performed through intra-frame block copying and intra-frame template matching prediction modes.
It improves the efficiency of video encoding and decoding, reduces redundancy in video data, and enhances the quality of encoding and decoding.
Smart Images

Figure CN121002879A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to U.S. Provisional Application No. 63 / 458,433, filed April 10, 2023, entitled “Methods and Devices for Intra Block Copy and Intra Template Matching,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This disclosure relates to video encoding and decoding and compression, and more specifically, but not limited to, methods and apparatus for improving the encoding and decoding efficiency of intra block copy (IBC) and intra template matching prediction (intra TMP). Background Technology
[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited storage resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0005] This disclosure provides examples of techniques related to improving intra-frame block copying methods in video encoding or decoding processes.
[0006] According to a first aspect of this disclosure, a method for video decoding is provided. In this method, a decoder can obtain a first block and a second block, wherein the first block is encoded in either an intra-block copy (IBC) mode or an intra-template matching prediction (TMP) mode, and the second block is encoded in either an intra-TMP mode or an IBC mode. Furthermore, the decoder can obtain the boundary strength of a deblocking filter by applying a predefined criterion in one of the intra-mode or inter-mode. Moreover, the decoder can apply the deblocking filter to the first block and the second block based on the boundary strength.
[0007] According to a second aspect of this disclosure, a method for video encoding is provided. In this method, an encoder can obtain a first block and a second block, wherein the first block is encoded in either an intra-block copy (IBC) mode or an intra-template matching prediction (TMP) mode, and the second block is encoded in either an intra-TMP mode or an IBC mode. Furthermore, the encoder can obtain the boundary strength of a deblocking filter by applying a predefined criterion in one of the intra-mode or inter-mode. Additionally, the encoder can apply the deblocking filter to the first block and the second block based on the boundary strength. Furthermore, the encoder can generate a bitstream based on applying the deblocking filter to the first block and the second block.
[0008] According to a third aspect of this disclosure, an apparatus for video decoding is provided. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the first aspect when executing the instructions.
[0009] According to a fourth aspect of this disclosure, an apparatus for video encoding is provided. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the second aspect when executing the instructions.
[0010] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause one or more computer processors to perform the method according to the first aspect.
[0011] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause one or more computer processors to perform the method according to the second aspect.
[0012] According to a seventh aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first aspect.
[0013] According to the eighth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by the method according to the second aspect. Attached Figure Description
[0014] A more specific description of the examples of this disclosure will be presented with reference to the specific examples shown in the accompanying drawings. Given that these drawings depict only a few examples and are therefore not intended to limit the scope, the examples will be described and explained using additional features and details through the use of the drawings.
[0015] Figure 1 This is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some examples of this disclosure.
[0016] Figure 2 This is a block diagram illustrating an exemplary video encoder according to some examples of this disclosure.
[0017] Figure 3 This is a block diagram illustrating an exemplary video decoder according to some examples of this disclosure.
[0018] Figures 4A to 4E This is a block diagram illustrating, according to some examples of this disclosure, how a frame can be recursively divided into multiple video blocks of different sizes and shapes.
[0019] Figure 5 A diagram showing the locations of spatial candidates according to some examples of this disclosure.
[0020] Figure 6 A schematic diagram of candidate pairs considered in a redundancy check for spatial candidates, according to some examples of this disclosure, is shown.
[0021] Figure 7 A graph showing the scaling of motion vectors for time candidates according to some examples of this disclosure is shown.
[0022] Figure 8 A diagram showing candidate positions for time candidates according to some examples of this disclosure.
[0023] Figure 9 A diagram showing merged mode (MMVD) search points with motion vector difference according to some examples of this disclosure is provided.
[0024] Figure 10 The following are examples of unidirectional predictive motion vector selection for geometric partitioning mode (GPM) according to this disclosure.
[0025] Figure 11 The diagram shows the upper and left adjacent blocks in CIIP weight derivation according to some examples of this disclosure.
[0026] Figure 12 The current CTU processing order and its available reference points in the current CTU and the left CTU are shown as examples of some of the examples of this disclosure.
[0027] Figure 13 Examples of filling candidates for replacing zero vectors in the IBC list are shown according to this disclosure.
[0028] Figure 14 The diagram illustrates reference regions for IBC when a CTU(m,n) is encoded, according to some examples of this disclosure. A shaded block (m,n) with dots represents the current CTU; a block shaded with a " / " indicates a reference region; and an unshaded block indicates an invalid reference region.
[0029] Figure 15 IBC reference areas for camera-captured content are shown as some examples according to this disclosure.
[0030] Figures 16A-16B This illustrates a method for dividing angle patterns according to some examples of this disclosure.
[0031] Figures 17A-17D GPM with inter-frame prediction and intra-frame prediction is shown as some examples according to this disclosure. Figures 17A-17C The available IPM candidates are shown. Figure 17D An example of GPM with intra-frame and intra-frame prediction is shown.
[0032] Figure 18 The edges on the template are shown as some examples according to this disclosure.
[0033] Figure 19 The intra-frame template matching search region is shown in some examples according to this disclosure.
[0034] Figure 20 Templates for template-based OBMCs are shown as examples of some of the features provided in this disclosure.
[0035] Figure 21 The present disclosure illustrates some examples of intra-coded block partitioning methods and corresponding weights for angular and planar modes.
[0036] Figure 22 The template used for intra-frame coded blocks and its reference samples are shown.
[0037] Figure 23 The template for the IBC coded block and its reference sample points are shown.
[0038] Figure 24 An example non-adjacent neighbor block is shown for use in IBC AMVP or merge candidate.
[0039] Figure 25 shows different sizes of non-adjacent neighboring blocks: (a) neighboring blocks with the same size as the current block and (b) neighboring blocks with different sizes (e.g., 4×4 or 8×8) than the current block.
[0040] Figure 26 An example of a spatially adjacent block used to derive a non-adjacent spatial candidate for an IBC pattern is shown, where the numbers in the non-adjacent adjacent blocks indicate the scan order.
[0041] Figure 27 Another example of spatially adjacent blocks used to derive non-adjacent spatial candidates for the IBC pattern is shown, where the numbers in the non-adjacent blocks indicate the scan order, and the numbers after the arrows indicate the degree values of the angles.
[0042] Figure 28 This shows an example where non-adjacent spatial regions are restricted to half the size of the current CTU above and to the left.
[0043] Figure 29 illustrates motion storage in non-adjacent spatial neighborhoods (IBC neighborhood CUs or non-IBC neighborhood CUs): (a) Permissible non-adjacent spatial regions beyond the current CTU; (b) Motion storage in the row buffer (A is an IBC CU; B is a non-IBC CU).
[0044] Figure 30 This shows the projection / cropping of non-adjacent neighboring locations if the scanned non-adjacent neighboring locations exceed the permissible spatial area.
[0045] Figure 31 This illustrates another example of projecting / clipping non-adjacent neighboring locations if the scanned non-adjacent neighboring locations exceed the allowed spatial area (e.g., beyond the current CTU and available row buffers).
[0046] Figure 32 This shows that the granularity of IBC motion storage differs from the minimum IBC block size.
[0047] Figure 33 An example of a sub-block-based IBC mode is shown, where the BV of a sub-block in the current block is obtained by reusing the BV of a sub-block in a co-block in a co-picture.
[0048] Figure 34 An example of a sub-block-based IBC pattern is shown, where the BV of the left or top sub-block in the current block is obtained by refining the BV of the current block using a template matching method.
[0049] Figure 35This is a diagram illustrating a computing environment coupled with a user interface according to some examples of this disclosure.
[0050] Figure 36 This is a graph showing the ramp function of the GPM mixing weights based on the displacement (d) from the predicted sample location to the GPM partition boundary and the mixing region size (τ).
[0051] Figure 37 is a diagram showing the candidates for spatial GPM.
[0052] Figure 38 This is a diagram showing the GPM template.
[0053] Figure 39 This is a graph showing GPM mixing.
[0054] Figure 40 This is a flowchart illustrating some examples of methods for video decoding according to this disclosure.
[0055] Figure 41 This illustrates some examples of the use of, for example, according to this disclosure. Figure 40 The flowchart shown corresponds to the video encoding method of the video decoding method. Detailed Implementation
[0056] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. Numerous non-limiting details are set forth in the following detailed description to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0057] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. The singular forms “a,” “the,” and “the” in this disclosure and the appended claims are also intended to include the plural forms unless otherwise expressly indicated throughout the disclosure. It should also be understood that the term “and / or” as used in this disclosure refers to and includes one or any or all possible combinations of the plurality of related items listed.
[0058] Throughout this specification, references to "an embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise expressly stated, the features, structures, elements, or characteristics described in connection with one or more embodiments also apply to other embodiments.
[0059] Throughout this disclosure, unless otherwise expressly stated, the terms “first,” “second,” “third,” etc., are used only as names to refer to related elements (e.g., devices, components, compositions, steps, etc.) and do not imply any spatial or temporal order. For example, “first device” and “second device” can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be named arbitrarily.
[0060] The terms "module," "submodule," "circuit," "subcircuit," "circuit system," "subcircuit system," "unit," or "subunit" can include memory (shared, dedicated, or grouped) storing code or instructions executable by one or more processors. A module can include one or more circuits, with or without stored code or instructions. A module or circuit can include one or more components that are directly or indirectly connected. These components may or may not be physically attached to each other or located adjacent to each other.
[0061] As used herein, the terms “if” or “when” may be understood, depending on the context, to mean “at the time of” or “in response to”. If these terms appear in the claims, they may not indicate that the relevant limitation or feature is conditional or optional. For example, a method may include the steps of: i) performing a function or action X’ when or if condition X exists, and ii) performing a function or action Y’ when or if condition Y exists. The method may be implemented with the ability to perform both function or action X’ and function or action Y’. Thus, both function X’ and Y’ can be performed at different times in multiple executions of the method.
[0062] A unit or module can be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a purely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.
[0063] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.
[0064] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, switch, base station, or any other means that may facilitate communication from source device 12 to target device 14.
[0065] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.
[0066] like Figure 1As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.
[0067] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be sent directly to the target device 14 via the output interface 22 of the source device 12. Alternatively, the encoded video data can be stored on the storage device 32 for later access by the target device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or transmitter.
[0068] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0069] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0070] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0071] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0072] In some implementations, components of source device 12 (e.g., video source 18, video encoder 20, or the following references) Figure 1 The components included in the video encoder 20 and the output interface 22) at least a portion of the components and / or the components of the target device 14 (e.g., the input interface 28, the video decoder 30 or the following references) Figure 3At least a portion of the components included in the video decoder 30 and the display device 34 can operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), where the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network can be located in one or more client devices, and these client devices can communicate with server computers in the cloud computing service network via wireless communication networks (e.g., cellular communication networks, short-range wireless communication networks, or Global Navigation Satellite System (GNSS) communication networks) or wired communication networks (e.g., local area network (LAN) communication networks or power line communication (PLC) networks). In one embodiment, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” may be used interchangeably as appropriate. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service networks. Instead, this disclosure may also be implemented in any other type of computing environment currently known or developed in the future.
[0073] Figure 2 This is a block diagram illustrating another exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".
[0074] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copying (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that, regarding the CCSAO technology, this application is not limited to the embodiments described herein, but can also be applied to situations where an offset is selected for any other component among the luminance, Cb, and Cr chrominance components based on any one of the luminance, Cb, and Cr chrominance components to modify that other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein can be any one of the luminance, Cb, and Cr chrominance components, the second component mentioned herein can be any other one of the luminance, Cb, and Cr chrominance components, and the third component mentioned herein can be the remaining components among the luminance, Cb, and Cr chrominance components. In some examples, the loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or can be distributed among one or more of the described fixed or programmable hardware units.
[0075] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding modes) when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.
[0076] like Figure 2As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of sample points with sample values. Sample points in the array may also be referred to as pixels or image elements (pel). The number of sample points in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of sample points with sample values, but its dimension is smaller than that of the video frame. The number of sample points in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. A video block can be further divided into one or more block partitions or sub-blocks (which can then re-form blocks) by iteratively using, for example, QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring, for example, to HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.
[0077] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.
[0078] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.
[0079] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.
[0080] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0081] The motion estimation unit 42 calculates the motion vector for a video block in an inter-frame predictive coding frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1), each of the first and second reference frame lists identifying one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.
[0082] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation unit 42. Upon receiving motion vectors for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel differences forming the residual video block may include a luminance component difference or a chrominance component difference, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.
[0083] In some implementations, the intra-BC unit 48 can generate vectors and acquire prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-BC unit 48 can, for example, use various intra-prediction modes to encode the current block during individual encoding passes and test their performance through rate-distortion analysis. Next, the intra-BC unit 48 can select a suitable intra-prediction mode from the various tested intra-prediction modes to use and generate an intra-mode indicator accordingly. For example, the intra-BC unit 48 can use rate-distortion analysis to calculate rate-distortion values for the various tested intra-prediction modes and select the intra-prediction mode with the best rate-distortion characteristics from the tested modes as the suitable intra-prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-frame prediction mode exhibits the optimal rate-distortion value for the block.
[0084] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction according to the embodiments described herein. In any case, for intra-frame block copying, in terms of pixel differences, the predicted block may be a block considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values for sub-integer pixel positions.
[0085] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values of the predicted block from the pixel values of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.
[0086] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra-block copy prediction performed by the intra-BC unit 48, the intra-prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. The intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.
[0087] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.
[0088] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0089] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1 The video decoder 30 shown, or archived in, for example Figure 1 The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.
[0090] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values for use in motion estimation.
[0091] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.
[0092] Figure 3 This is a block diagram illustrating another exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 2 The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.
[0093] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).
[0094] Video data memory 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding modes) when decoding video data. Video data memory 79 and DPB 92 can be formed of any memory device from a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown in... Figure 3 The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip along with other components of the video decoder 30, or off-chip relative to those components.
[0095] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.
[0096] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-predictive unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-predictive mode transmitted by the signal and reference data from the previous decoded block of the current frame.
[0097] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.
[0098] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.
[0099] Motion compensation unit 82 and / or intra-frame prediction (BC) unit 85 determine prediction information for video blocks in the current video frame by parsing motion vectors and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) used to encode video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the reference frame list for the frame, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.
[0100] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.
[0101] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.
[0102] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0103] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1 On the display device 34).
[0104] In a typical video coding process, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of chrominance samples (Cb). SCr is a two-dimensional array of chrominance samples (Cr). In other instances, a frame may be monochrome and therefore consist of only a two-dimensional array of luma samples.
[0105] like Figure 4AAs shown, the video encoder 20 (or more specifically, the segmentation unit in the predictive processing unit of the video encoder 20) generates a coded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block may be an N×N sample block.
[0106] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or a combination thereof, and divide the CTU into smaller CUs. Figures 4B-4E This is a block diagram illustrating how, according to some embodiments of this disclosure, a frame is recursively divided into multiple video blocks of different sizes and shapes. For example... Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image can include, for example, two corresponding coding blocks (CBs) of luminance and chrominance samples of the same frame size, as well as syntax elements for encoding the samples of the coding blocks. In monochrome images or images with three separate color planes, a CU can include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that... Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / ternary / binary partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and ternary tree structures. Figure 4E As shown, a coded block with width W and height H has five possible segmentation types: quad segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation.
[0107] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.
[0108] Video encoder 20 can use intra-frame prediction or inter-frame prediction to generate prediction blocks for the PU. If video encoder 20 uses intra-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0109] After the video encoder 20 generates a predicted luminance block, a predicted Cb block, and a predicted Cr block for one or more PUs of the CU, the video encoder 20 can generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from the original luminance coding block of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0110] In addition, such as Figure 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.
[0111] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0112] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.
[0113] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding the samples of the prediction block of the PU for the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.
[0114] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.
[0115] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "motion vector predictor" (MVP) of the current CU) by exploring their spatial and temporal correlations.
[0116] Instead of the above combination Figure 2 The method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit 42, into the video bitstream, and subtracting the motion vector prediction factor of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0117] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current CU.
[0118] Generally, the basic inter-frame prediction scheme used in VVC is almost the same as that in HEVC, except for further extensions, additions and / or improvements to several prediction tools, such as extended merge prediction, MMVD and GPM.
[0119] Extended merge forecast
[0120] With the continuous improvement of video data acquisition technology and the finer size of video blocks used to preserve the details of video data, the amount of data required to represent the motion vector of the current image has also increased significantly. One way to overcome this challenge is to use the motion information (e.g., motion vectors) of the spatially adjacent CUs, temporally co-located CUs, etc. of the current CU as an approximation (e.g., prediction) of the motion information of the current CU, which is also known as the "Motion Vector Prediction (MVP)" of the current CU.
[0121] Similar to the process of selecting a prediction block from a reference picture during inter-frame prediction of a coded block, both video encoder 20 and video decoder 30 need to construct an MVP candidate list for the current CU using a set of rules, and then select an MVP candidate from the MVP candidate list as the MVP of the current CU. By doing so, it is not necessary to transmit the MVP candidate list itself between video encoder 20 and video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for video encoder 20 and video decoder 30 to use the same MVP candidate selected from the MVP candidate list to encode and decode the current CU.
[0122] In VVC, the MVP candidate list is constructed by sequentially including the following five types of MVPs: —The spatial MVP from the spatially adjacent CUs (i.e., spatial candidates); —The time MVP from the time-isolated CU (i.e., the time candidate); —History-based MVP (HMVP) from a First-In-First-Out (FIFO) table; —Paired average MVP; and —Zero MVP.
[0123] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded in merge mode, the index of the best MVP candidate is encoded using truncated unary binarization. The first bit of the index is encoded using the context, and the remaining bits of the index are encoded using bypass encoding.
[0124] The process for obtaining each type of MVP is described below. Similar to HEVC, VVC also supports obtaining a list of MVP candidates in parallel across all CUs within a given region.
[0125] MVP is obtained from spatial candidates.
[0126] In VVC, based on spatial candidates (e.g., with...) Figure 5 The MVP obtained from the current CU 101 (adjacent CUs) is the same as the MVP obtained from spatial candidates in HEVC, except that the positions of the first two spatial candidates have been swapped. From the location... Figure 5 Up to four spatial candidates are selected from the indicated spatial candidates (i.e., top position B0, left position A0, top right position B1, bottom left position A1, and top left position B2). The process is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because the one or more CUs belong to other stripes or tiles) or are intra-coded.
[0127] After adding the CU at position B0 as a candidate to the merged candidate list, redundancy checks are performed on the addition of the remaining candidates to the merged candidate list. This ensures that candidates with the same motion information are excluded from the merged candidate list, thereby improving encoding and decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only... Figure 6 Pairs linked by arrowed lines are considered, and a candidate is added to the merged candidate list only if the motion information of the candidate in the pair used for redundancy checking differs from the motion information of the candidate to be added. The spatial MVP obtained from the candidates in the merged candidate list is added to the MVP candidate list.
[0128] MVP is selected based on time candidate.
[0129] During the process of deriving the MVP based on the time candidate, only one time candidate is added to the merge candidate list. Specifically, when deriving the MVP based on this time candidate, for the current CU (e.g., ...), Figure 7 The curr_CU 303 in the image is based on the image belonging to the same position (e.g., Figure 7 The corresponding CU of col_pic 302 in (e.g., Figure 7 The scaled motion vector (col_CU 301) is used as a temporal candidate to obtain the MVP candidate list, and this scaled motion vector is added as a temporal MVP candidate. The list of reference images and their indices for obtaining the corresponding CU are explicitly signaled in the strip header. Figure 7 As shown, the scaled motion vector is obtained (i.e., scaled) based on the motion vector of the co-located CU using the image sequence count (POC) distance (i.e., tb and td), where tb is defined as the current image (e.g., ...). Figure 7 Reference image for curr_pic 304 (e.g., Figure 7 The difference between curr_ref 305 in the image and the current image is the POC difference, while td is defined as the reference image of the co-located image (e.g., Figure 7 The POC difference between col_ref 306 and the corresponding image. The reference image index for the time candidate is set to zero.
[0130] like Figure 8 As shown, the position of the time candidate (i.e., the co-occurring CU) in the current CU 401 is selected from positions C0 and C1. If the CU at position C0 in the co-occurring picture is unavailable, intra-coded, or located outside the current CTU line, the CU at position C1 is used as the co-occurring CU for obtaining the time MVP candidate. Otherwise, the CU at position C0 is used as the co-occurring CU for obtaining the time MVP candidate.
[0131] HMVP candidate
[0132] After the spatial MVP and temporal MVP, HMVP candidates are added to the MVP candidate list. Motion information from previously coded blocks is stored in the HMVP table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during encoding / decoding. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-coded CU exists, the associated motion information is added as the last entry in the HMVP table as a new HMVP candidate.
[0133] The size of the HMVP table is set to 6. When inserting a new HMVP candidate into the HMVP table, a constrained FIFO rule is used, where a redundancy check is first applied to check if a duplicate HMVP already exists in the HMVP table. If found, the duplicate HMVP is removed from the HMVP table, all subsequent HMVP candidates are shifted forward, and the duplicate HMVP is added to the last entry in the HMVP table.
[0134] HMVP candidates can be used during the MVP candidate list construction process. The latest few HMVP candidates in the HMVP table are checked sequentially and inserted into the MVP candidate list after the temporal MVP candidates. Redundancy checks are applied to the HMVP candidates relative to the spatial and / or temporal MVP candidates.
[0135] To reduce the number of redundant verification operations, the following simplifications are introduced: —Redundancy checks are performed on the last two entries in the HMVP table relative to the spatial MVP candidates obtained from the spatial candidates at positions A1 and B1, respectively; and —The process of building the MVP candidate list based on the HMVP candidates ends once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1.
[0136] Get Paired Average MVP Candidates
[0137] Pairwise average MVP candidates are generated by averaging the MVPs obtained using the first two predefined pairs of existing merge candidate lists. The first merge candidate in a predefined pair can be defined as p0Cand, and the second merge candidate in a predefined pair can be defined as p1Cand. For each reference image list individually, the average motion vector is calculated based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available for a reference image list, they are averaged even if they point to different reference images, and the reference image of the averaged motion vector is set as the reference image of p0Cand. If only one motion vector is available for a reference image list, that motion vector is used directly. If no motion vector is available for a reference image list, the motion vector and reference image index for that reference image list remain invalid.
[0138] Zero MVP
[0139] If the MVP candidate list is not full after adding pairwise average MVP candidates, insert zero MVPs at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0140] MMVD
[0141] As mentioned above, in merge mode, motion information (i.e., MVP candidates) is implicitly obtained from the MVP candidate list constructed for the current CU and directly used as the MV of the current CU to generate prediction samples for the current CU. This may lead to a certain error between the actual MV of the current CU and the implicitly obtained MVP. To improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, where the motion vector difference (MVD) of the current CU is added to the implicitly obtained MVP to obtain the MV of the current CU. After sending the regular merge flag, the MMVD flag is signaled to specify whether the MMVD mode is used for the current CU.
[0142] In MMVD mode, after selecting an MVP candidate from the first two MVP candidates in the MVP candidate list, MMVD information is signaled. The MMVD information includes an MMVD candidate flag to specify which of the first two MVP candidates is selected as the basis for the MV, a distance index to indicate the motion amplitude information of the MVD, and a direction index to indicate the motion direction information of the MVD.
[0143] The distance index indicating the motion amplitude information of the specified MVD is compared with the reference image of the current CU (e.g., Figure 9 In the L0 reference diagram 501 or L1 reference diagram 503, the MVP candidate is pointed to by the selected MVP candidate (e.g., Figure 9 The dashed circle in the diagram represents a predefined offset of the starting point. The MVD can be derived from this offset and then added to the selected MVP candidate. Table 1 below specifies the relationship between the distance index and the predefined offset.
[0144]
[0145] Table 1
[0146] The direction index specifies the sign of the MVD, which represents the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and the predefined signs. In some examples, the meaning of the MVD sign can vary depending on the information of the selected MVP candidate. When the selected MVP candidate is a one-way predictive MV or a two-way predictive MV (where both MVs point to the same side of the current image (i.e., the POCs of the two reference images of the current image (e.g., the reference images of List 0 and List 1, also referred to as the L0 reference image and L1 reference image, respectively)), the sign in Table 2 specifies the sign of the MVD added to the selected MVP candidate. When the selected MVP candidate is a bidirectional prediction MV (where the two MVs point to different sides of the current image (i.e., the POC of one reference image of the current image is greater than the POC of the current image, while the POC of the other reference image of the current image is less than the POC of the current image), if the POC distance for the L0 reference image (i.e., the POC distance between the L0 reference image and the current image) is greater than the POC distance for the L1 reference image (i.e., the POC distance between the L1 reference image and the current image), then the sign in Table 2 specifies the sign of the MVD for list 0 (MVD0) added to the MVP for list 0 (MVP0) of the selected MVP candidate, and the sign of the MVD for list 1 (MVD1) added to the MVP for list 1 of the selected MVP candidate is the opposite of the sign in Table 2; otherwise, if the POC distance for the L1 reference image is greater than the POC distance for the L0 reference image, then the sign in Table 2 specifies the sign of the MVD1 added to MVP1, and the sign of the MVD0 added to MVP0 is the opposite of the sign in Table 2. The symbols in the text are reversed.
[0147]
[0148] Table 2
[0149] MVD is scaled based on the POC distance. If the POC distances for the L0 and L1 reference images are the same, no scaling of the MVD is required. Otherwise, if the POC distance for the L0 reference image is greater than the POC distance for the L1 reference image, MVD1 is scaled. If the POC distance for the L1 reference image is greater than the POC distance for the L0 reference image, MVD0 is scaled.
[0150] GPM
[0151] In VVC, GPM is supported for inter-frame prediction. A CU-level flag is used to signal GPM as a merging mode; other merging modes include regular merging, MMVD, CIIP, and sub-block merging. For each possible CU size... ( in GPM supports a total of 64 partitions, with a possible CU size of [missing information]. Excluding 8 64 and 64 8.
[0152] When using GPM, the CU is divided into two parts by geometrically positioned straight lines. The position of the dividing line is mathematically determined based on the angle and offset parameters of the specific partition. Each part of the CU obtained through geometric partitioning is predicted inter-frame using its own motion; and only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, like traditional bidirectional prediction, only two motion-compensated predictions are needed for each CU.
[0153] If GPM is used for the current CU, then the geometric partition index (indicating the angle and offset of the geometric partition) and two merge indexes (one merge index for each partition) are further signaled.
[0154] The unidirectional prediction candidate list is directly obtained from the merge candidate list constructed through the extended merge prediction process described above. Let n represent the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector of the nth merge candidate in the merge candidate list (where X equals the parity of n) is used as the nth unidirectional prediction motion vector of GPM. Figure 10 In this context, these motion vectors are labeled with "x". If the corresponding LX motion vector of the nth merge candidate in the merge candidate list does not exist, the L(1-X) motion vector of the same merge candidate is used instead as the unidirectional predicted motion vector of GPM.
[0155] CIIP
[0156] In VVC, when encoding a CU in merge mode, if the CU contains at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if the width and height of the CU are less than 128 luma samples, an additional flag is signaled to indicate whether CIIP mode is applied to the current CU. In CIIP mode, the prediction signal is obtained by combining the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal in CIIP mode is obtained using the same inter-frame prediction process applied in regular merge mode; and the intra-frame prediction signal in CIIP mode is obtained according to the regular intra-frame prediction process using planar mode. Then, a weighted average is used to combine the intra-frame prediction signal and the inter-frame prediction signal, wherein the combination is based on the top neighbor block and the left neighbor block of the current CU 1601 (e.g., ...). Figure 11 The weight values are calculated for the encoding pattern shown below: —If the top adjacent block is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. —If the left adjacent block is available and is intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. —If (isIntraLeft + isIntraTop) equals 2, then set the weight value to 3; —Otherwise, if (isIntraLeft + isIntraTop) equals 1, then set the weight value to 2; —Otherwise, set the weight value to 1.
[0157] —The prediction signal in CIIP mode is obtained as follows. :
[0158] in It is the inter-frame prediction signal in CIIP mode. It is an intra-frame prediction signal in CIIP mode. It represents the weight value, and >> indicates a right shift operation.
[0159] Intra-block copying in Universal Video Codec (VVC)
[0160] Intra Block Copy (IBC) is a tool used in HEVC extensions to SCC. IBC significantly improves the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has been reconstructed within the current image. The luma block vector of the IBC-coded CU is integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-coded CUs are considered a third prediction mode in addition to intra-prediction mode or inter-prediction mode. IBC mode is suitable for CUs with both width and height less than or equal to 64 luma samples.
[0161] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-merging modes, a block vector search is first performed using a hash-based search. If the hash search does not return any valid candidates, a local search based on block matching is performed.
[0162] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4×4 sub-blocks. For larger current blocks, a hash key is determined to match the hash key of a reference block when all hash keys of all 4×4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the reference block with the lowest cost is selected.
[0163] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.
[0164] At the CU level, the IBC mode is transmitted using flag signals, and the IBC mode can be transmitted as IBCAMVP mode or IBC skip / merge mode as follows: IBC Skip / Merge Mode: Uses merge candidate indices to indicate which block vector from the list of neighboring candidate IBC blocks is used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and paired candidates.
[0165] IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighborhood and one from the upper neighborhood (in the case of IBC encoding). When either neighbor is unavailable, a default block vector is used as the predictor. A flag is sent to indicate the block vector predictor index.
[0166] IBC Reference Area
[0167] To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of predefined regions to include the current CTU region and specific regions of the left CTU. Figure 12 The reference area of the IBC mode is shown, where each block represents a 64×64 luminance sample unit.
[0168] Based on the current location of the encoding CU within the current CTU, apply the following operations: If the current block falls within the upper-left 64×64 block of the current CTU, then in addition to referencing the reconstructed samples in the current CTU, the current block can also use CPR mode to reference reference samples in the lower-right 64×64 block of the left CTU. The current block can also use CPR mode to reference reference samples in the lower-left 64×64 block of the left CTU and the upper-right 64×64 block of the left CTU.
[0169] If the current block falls within the upper right 64×64 block of the current CTU, then in addition to referencing the reconstructed samples in the current CTU, if the brightness position (0, 64) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to refer to the reference samples in the lower left and lower right 64×64 blocks of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64×64 block of the left CTU.
[0170] If the current block falls within the lower left 64×64 block of the current CTU, then in addition to referencing the already reconstructed samples in the current CTU, if the brightness position (64, 0) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference reference samples in the upper right and lower right 64×64 blocks of the left CTU. Otherwise, the current block can also use CPR mode to reference reference samples in the lower right 64×64 block of the left CTU.
[0171] If the current block falls within the lower right 64×64 block of the current CTU, the current block can use CPR mode to refer only to the samples already reconstructed in the current CTU.
[0172] This limitation allows the use of local on-chip memory to implement the IBC mode for hardware implementation.
[0173] Interaction between IBC and other coding tools
[0174] The interaction between the IBC mode and other inter-frame coding tools in VVC (such as Paired Merge Candidate, History-Based Motion Vector Predictor (HMVP), Combined Intra / Inter-Frame Prediction Mode (CIIP), Merge Mode with Motion Vector Difference (MMVD), and Geometric Partitioning Mode (GPM)) is as follows: IBC can be used with pairwise merge candidates and HMVP. New pairwise IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC motions are inserted into a history buffer for future reference.
[0175] IBC cannot be used in combination with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM.
[0176] When using a dual-tree partition, chroma-coded blocks are not allowed to use IBC.
[0177] Unlike HEVC screen content encoding extensions, the current image is no longer included as one of the reference images in reference image list 0 for IBC prediction. The process of deriving motion vectors for the IBC mode does not include all neighboring blocks in inter-frame modes, and vice versa. The following IBC design aspects are applied: IBC shares the same process as regular MV merging, including pairwise merging candidates and history-based motion predictors, but does not allow TMVP and zero vectors because they are invalid for IBC models.
[0178] Separate HMVP buffers (5 candidates each) are used for regular MV and IBC.
[0179] The block vector constraint is implemented in the form of a bitstream consistency constraint. The encoder needs to ensure that there are no invalid vectors in the bitstream, and that merging should not be used if a merging candidate is invalid (out of range or 0). This bitstream consistency constraint is expressed using a virtual buffer as described below.
[0180] For deblocking, IBC is treated as an inter-frame mode.
[0181] If the current block is encoded using IBC prediction mode, AMVR does not use quarter-pixel precision; instead, AMVR is signaled to indicate only whether the MV is in half-pixel precision (inter-pel) or in 4-integer-pel.
[0182] The number of IBC merge candidates can be signaled separately from the number of regular merge candidates, sub-block merge candidates, and geometric merge candidates in the strip header.
[0183] The concept of a virtual buffer is used to describe the permissible reference region and effective block vector for IBC prediction modes. The CTU size is represented as ctbSize, and the width of the virtual buffer ibcBuf is wIbcBuf = 128x128 / ctbSize, and the height is hIbcBuf = ctbSize. For example, for a 128×128 CTU size, the size of ibcBuf is also 128×128; for a 64×64 CTU size, the size of ibcBuf is 256×64; and for a 32×32 CTU size, the size of ibcBuf is 512×32.
[0184] The size of the VPDU is min(ctbSize, 64) in each dimension, W v = min(ctbSize, 64).
[0185] The virtual IBC buffer ibcBuf is maintained as follows.
[0186] When decoding each CTU line begins, refresh the entire ibcBuf with an invalid value of -1.
[0187] When starting decoding of the VPDU (xVPDU, yVPDU) relative to the top-left corner of the image, set ibcBuf[x][y] = 1, where x = xVPDU%wIbcBuf,…,xVPDU%wIbcBuf + W v – 1; y = yVPDU%ctbSize,…,yVPDU%ctbSize + W v 1.
[0188] After decoding, the CU contains (x, y) values relative to the top-left corner of the image. Settings: ibcBuf[ x % wIbcBuf ][ y % ctbSize ] = recSample[ x ][ y ] For a block covering coordinates (x, y), the block vector is valid if the following condition is true for the block vector bv = (bv[0], bv[1]); otherwise, the block vector is invalid: ibcBuf[ (x + bv[0])% wIbcBuf] [ (y + bv[1]) % ctbSize ] should not be equal to -1.
[0189] Intra-block copying in Enhanced Compression Model (ECM)
[0190] In ECM, IBC is improved in the following ways.
[0191] IBC Merge / AMVP List Construction
[0192] The following modifications are made to the IBC merge / AMVP list build: An IBC merge / AMVP candidate can only be inserted into the IBC merge candidate list / AMVP candidate list if it is valid.
[0193] Candidates in the upper right space, lower left space, and upper left space, as well as a pairwise average candidate, can be added to the IBC merge candidate list / AMVP candidate list.
[0194] Template-based adaptive reordering (ARMC-TM) was applied to the IBC merge list.
[0195] The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC merge candidates using full pruning, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0196] Candidate zero vectors used to populate the IBC merge list / AMVP list are replaced by the BVP candidate set located in the IBC reference region. Zero vectors are invalid as block vectors in IBC merge mode, and therefore, they are discarded as BVPs in the IBC candidate list.
[0197] Three candidates are located at the nearest corner of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as follows. Figure 13 As shown.
[0198] Block vector candidates derived from intra-frame TMP for IBC
[0199] In this method, the block vector (BV) derived from IntraTMP (Intra-Temporal Matching Prediction) is used for Intra-Block Copy (IBC). The stored IntraTMP BVs and IBC BVs of neighboring blocks are used as spatial BV candidates in the construction of the IBC candidate list.
[0200] The IntraTMP block vector is stored in the IBC block vector buffer, and the current IBC block can use the IBC BV and the IntraTMP BV of adjacent blocks as BV candidates in the IBC BV candidate list. The IntraTMP block vector is added to the IBC block vector candidate list as a spatial candidate.
[0201] IBC with template matching
[0202] Template matching is used in both the IBC merge mode and the IBC AMVP mode in IBC.
[0203] The IBC-TM merge list has been modified compared to the regular IBC merge mode, allowing candidates to be selected based on a pruning method that utilizes the motion distance between candidates, as in the regular TM merge mode. The ending zero motionfulfillment mechanism is replaced by motion vectors to the left (-W, 0), above (0, -H), and upper left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0204] In IBC-TM merging mode, a template matching method is used to refine the selected candidates before the RDO or decoding process. IBC-TM merging mode competes with regular IBC merging mode and sends a TM merging flag via signaling.
[0205] In IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. Each of these three selected candidates is refined using a template matching method and ranked according to their resulting template matching costs. Then, during motion estimation, only the top two candidates are considered as usual.
[0206] Template matching refinement for both the IBC-TM merging mode and the AMVP mode is quite simple because the IBC motion vector is constrained to be (i) an integer and (ii) within the reference region, such as Figure 12 As shown in the diagram. Therefore, in IBC-TM merge mode, all thinning is performed with integer precision, and in IBC-TM AMVP mode, they are performed with either integer precision or 4-pixel precision based on the AMVR value. Such thinning only accesses samples that have not been interpolated. In both cases, the motion vectors thinned and the template used in each thinning step must adhere to the constraints of the reference region.
[0207] IBC Reference Area
[0208] The reference area for IBC is extended to the two CTU rows above. Figure 14The reference region used for encoding CTU(m,n) is shown. Specifically, for the CTU(m,n) to be encoded, the reference region includes CTUs with indices (m-2,n-2)…(W,n-2), (0,n-1)…(W,n-1), (0,n)…(m,n), where W represents the maximum horizontal index within the current tile, strip, or image. This setup ensures that for a CTU of size 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is capped at [–(C << 1), C >> 2] in the horizontal direction and [–C, C >> 2] in the vertical direction to accommodate the reference region expansion, where C represents the CTU size.
[0209] IBC merging mode with block vector difference
[0210] In ECM, an IBC merging mode with block vector difference is used. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD direction is two horizontal directions and two vertical directions.
[0211] A basic candidate is selected from the top five candidates in the reordered IBC merge list. Then, all possible MBVD refinement locations (20×4) for each basic candidate are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement location. Finally, the top 8 refinement locations with the lowest template SAD cost are reserved as available locations for MBVD index encoding.
[0212] IBC Adaptation for Camera-Captured Content
[0213] When adapting IBC for camera-captured content, the IBC reference range is reduced from 2 CTU lines to 2 × 128 lines, such as Figure 15 As shown in the diagram. On the encoder side, to reduce complexity, the local search range is centered on the first block vector predictor of the current CU, set to [-8,8] in the horizontal direction and [-8,8] in the vertical direction. This encoder modification is not applicable to SCC sequences.
[0214] CIIP combined with TIMD and TM
[0215] In CIIP mode, prediction samples are generated by weighting the inter-prediction signals that combine candidate predictions using CIIP-TM and the intra-prediction signals that are predicted using intra-prediction modes derived using TIMD. This method is only applied to coded blocks with an area less than or equal to 1024.
[0216] The TIMD export method is used to export intra-prediction modes from CIIP. Specifically, it selects the intra-prediction mode with the smallest SATD value from the TIMD mode list and maps it to one of 67 regular intra-prediction modes.
[0217] Furthermore, it is proposed to modify the weights (wIntra, wInter) for the two tests when the derived intra-prediction mode is an angle mode. For near-horizontal mode (2 <= angle mode index < 34), such as Figure 16A The current block is divided vertically as shown; for near-vertical mode (34 <= angle mode index <= 66), as... Figure 16B The current block is divided horizontally as shown.
[0218] Table 3 shows (wIntra, wInter) for different sub-blocks.
[0219]
[0220] Table 3. Modified weights for angle mode.
[0221] Using CIIP-TM, a CIIP-TM merge candidate list is constructed for the CIIP-TM pattern. The merge candidates are refined through template matching. The CIIP-TM merge candidates are also reordered into regular merge candidates using the ARMC method. The maximum number of CIIP-TM merge candidates is two.
[0222] Multiple Hypothesis Prediction (MHP)
[0223] In multi-hypothesis inter-frame prediction mode, in addition to the traditional bidirectional prediction signal, one or more additional motion-compensated prediction signals are transmitted. The final overall prediction signal is obtained by weighted superposition of samples. The bidirectional prediction signal is utilized. and the first additional inter-frame prediction signal / hypothesis The final prediction signal is obtained as follows. : (2) Based on the mapping presented in Table 4, the weighting factors Specifyed by the new syntax element add_hyp_weight_idx:
[0224] Table 4. Mapping between add_hyp_weight_idx and α
[0225] Similar to the above, more than one additional prediction signal can be used. The final overall prediction signal is iteratively accumulated with each additional prediction signal.
[0226] (3)
[0227] The final overall prediction signal was obtained as the final (That is, the one with the largest index n) Within this mode, up to two additional prediction signals can be used (i.e., n is limited to 2).
[0228] Motion parameters for each additional prediction hypothesis can be explicitly signaled by specifying a reference index, motion vector predictor index, and motion vector difference, or implicitly signaled by specifying a merging index. A separate multi-hypothesis merging flag distinguishes between these two signaling modes.
[0229] For inter-frame AMVP mode, MHP is applied only when unequal weights are selected in BCW in bidirectional prediction mode.
[0230] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0231] Geometric Partitioning Pattern (GPM) in ECM
[0232] GPM with Combined Motion Vector Difference (MMVD)
[0233] GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. First, a signal flag is sent to the GPMCU to specify whether to use this mode. If this mode is used, each geometric partition of the GPM CU can further determine whether to send MVD using signals. If MVD is sent for a geometric partition, the motion of that partition is further refined using the signaled MVD information after selecting a GPM merging candidate. All other procedures remain the same as for GPM.
[0234] MVD is signaled as a pair of distance and direction, similar to MMVD. GPM with MMVD (GPM-MMVD) involves nine candidate distances (¼ pixel, ½ pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, MVD is shifted left by 2, just like MMVD.
[0235] GPM with Template Matching (TM)
[0236] Template matching is applied to GPM. When GPM mode is enabled for CU, a CU-level flag is signaled to indicate whether TM is applied to both geometric partitions. TM is used to refine motion information for each geometric partition. When TM is selected, a template is constructed based on the partition angle using left neighbor samples, top neighbor samples, or left and top neighbor samples, as shown in Table 5. Motion is then refined by minimizing the difference between the current template and the template in the reference image using the same search mode in the merging mode (where the half-pixel interpolation filter is disabled).
[0237]
[0238] Table 5. Templates for the first and second geometric partitions, where A indicates the use of the top sample point, L indicates the use of the left sample point, and L+A indicates the use of both the left and top sample points.
[0239] The GPM candidate list is constructed as follows: 1. Interleaved list-0 MV candidates and list-1 MV candidates are derived directly from the regular merge candidate list, where list-0 MV candidates have higher priority than list-1 MV candidates. A pruning method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates.
[0240] 2. Further, interleaved list-1 MV candidates and list-0 MV candidates are derived directly from the regular merge candidate list, where list-1 MV candidates have higher priority than list-0 MV candidates. The same pruning method with adaptive thresholds is also applied to remove redundant MV candidates.
[0241] 3. Fill the zero MV candidate list until the GPM candidate list is full.
[0242] GPM-MMVD and GPM-TM are enabled only for one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (if at least one GPM-MMVD flag is true), the value of the GPM-TM flag is inferred to be false.
[0243] GPM with inter-frame prediction and intra-frame prediction
[0244] In GPM with inter-frame prediction and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame prediction samples and intra-frame prediction samples for each GPM segmentation region. Inter-frame prediction samples are derived from the inter-frame GPM, while intra-frame prediction samples are derived from the intra-frame prediction mode (IPM) candidate list and the index sent by the encoder signal. The IPM candidate list size is predefined as 3. Available IPM candidates are the parallel angle mode (parallel mode) relative to the GPM block boundary, the vertical angle mode (vertical mode) relative to the GPM block boundary, and the planar mode, as shown below. Figures 17A-17D As shown. Furthermore, for example... Figure 17D The GPM with intra-frame prediction and intra-frame prediction shown is constrained to reduce signaling overhead for IPM and avoid increasing the size of intra-frame prediction circuitry on the hardware decoder. Furthermore, direct motion vectors and IPM storage on the GPM mixing region are introduced to further improve coding performance.
[0245] In DIMD and neighbor-based IPM export, parallel modes are registered first. Therefore, if no identical IPM candidate exists in the list, up to two IPM candidates can be registered from the decoder-side intra-mode export (DIMD) method and / or neighbor block export. For neighbor mode export, there are up to five available neighbor block locations, but they are limited by the angle of the GPM block boundary, as shown in Table 6. These locations have been used for GPM with template matching (GPM-TM).
[0246]
[0247] Table 6. Positions of available neighboring blocks for IPM candidate derivation based on the angle of the GPM block boundary. A and L represent the top and left sides of the predicted block.
[0248] GPM-intraframe can be combined with GPM with motion vector difference combining (GPM-MMVD). TIMD is used as an IPM candidate within GPM frames to further improve coding performance. Parallel modes can be registered first, followed by TIMD, DIMD, and IPM candidates from neighboring blocks.
[0249] Template matching-based reordering for GPM partitioning patterns
[0250] In template-match-based reordering of GPM partitioning patterns, given the motion information of the current GPM block, the corresponding TM cost value for the GPM partitioning pattern is calculated. Then, all GPM partitioning patterns are reordered in ascending order based on their TM cost values. Instead of sending the GPM partitioning pattern, an index using Golomb-Rice code is signaled to indicate the exact position of the GPM partitioning pattern in the reordering list.
[0251] The reordering method for GPM partitioning patterns is a two-step process performed after generating the corresponding reference templates for the two GPM partitions in the coding unit, as shown below: • Extend the GPM partition edge to the reference templates of the two GPM partitions to obtain 64 reference templates, and calculate the corresponding TM cost for each of the 64 reference templates; • The TM cost values based on the GPM partitioning pattern are reordered in ascending order, and the 32 best patterns are marked as available partitioning patterns.
[0252] The edges on the template extend from the edges of the current CU, such as Figure 18 As shown, however, the GPM blending process is not used in template areas that cross the edge.
[0253] After reordering in ascending order using TM cost, the index is signaled.
[0254] Geometric Partitioning Mode (GPM) with Adaptive Hybridization
[0255] In VVC, the final predicted samples are generated by weighted averaging the predictions of the two predicted signals. Two integer mixing matrices (W0 and W1) are used. The weights in the GPM mixing matrix are derived from the ramp function based on the displacement from the predicted sample location to the GPM partition boundary. The mixing region size is fixed at two (two samples on each side of the GPM partition partition boundary).
[0256] Improve the blending process in ECM by adding four additional blending region sizes (one-quarter, half, double, and four times the existing region size), such as Figure 36 As shown, the CU-level flags are encoded to signal the selected mixing region size. Furthermore, extended weighted precision is utilized, where the maximum value of the weights is changed from 8 (in VVC) to 32 to accommodate the extended mixing region size.
[0257] Spatial Geometric Partitioning Pattern (SGPM)
[0258] SGPM is an intra-frame mode of an inter-frame coding tool similar to GPM, where two prediction parts are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, as shown in Figure 37. 26 segmentation modes and 3 intra-frame prediction modes are used to form a combination. The length of the candidate list is set to 16. The selected candidate index is sent using a signal.
[0259] Reorder the list using a template ( Figure 38 The SAD (Search Allocation) between the template's prediction and reconstruction is used for sorting. In one example, the template size is fixed at 1. In some other examples, the template size can be set differently.
[0260] For each segmentation mode, the same intra-frame to inter-frame GPM list is used to export an IPM list for each segment. The IPM list size is set to 3. In the list, the TIMD export mode is replaced by two export modes with horizontal and vertical orientations.
[0261] SGPM mode is applied with restricted block size: 4 <= width <= 64, 4 <= height <= 64, width < height 8. Height < Width 8. Width Height >= 32.
[0262] Adaptive blending is also used in spatial GPM, where Figure 39 The mixing depth τ shown is derived as follows: If min(width, height) == 4, then choose 1 / 2τ; Otherwise, if min(width, height) == 8, then choose τ; Otherwise, if min(width, height) == 16, then choose 2τ; Otherwise, if min(width, height) == 32, then choose 4τ; Otherwise, choose 8τ.
[0263] Intra-frame template matching
[0264] Intra-frame template matching prediction (intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and performs the same prediction operation on the decoder side.
[0265] By comparing the L-shaped causal neighbors of the current block with Figure 19 The prediction signal is generated by matching another block within a predefined search region, which includes: R1: Current CTU R2: Top Left CTU R3: Up CTU R4: Left CTU The sum of absolute differences (SAD) is used as the cost function.
[0266] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0267] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a BlkW SearchRange_h = a BlkH in" " is a constant that controls the trade-off between gain and complexity. In practice, " "Equals 5."
[0268] For CUs with a width and height less than or equal to 64, enable the intra-frame template matching tool. The maximum CU size used for intra-frame template matching is configurable.
[0269] When DIMD is not used in the current CU, the intra-template matching prediction mode is signaled at the CU level using a dedicated flag.
[0270] Fusion for Template-Based Intra-Frame Mode Export (TIMD)
[0271] For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. First, the two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are then weighted and fused after applying the PDPC procedure, and this weighted intra-prediction is used to encode the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0272] The costs of the two selected modes are compared to a threshold. In the test, cost factor 2 is applied as follows: costMode2 < 2 costMode1.
[0273] If the condition is true, then fusion is applied; otherwise, only mode1 is used.
[0274] The weights of the patterns are calculated from their SATD costs as follows: Weight 1 = costMode2 / (costMode1 + costMode2) Weight 2 = 1 - Weight 1 Division is performed using the same lookup table (LUT)-based integerization scheme used by CCLM.
[0275] Localized lighting compensation (LIC)
[0276] LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its predicted block as a function of the local illumination variation between the current block template and the reference block template. The parameters of the function can be represented by the scaling factor α and the offset β, forming a linear equation, i.e., α p[x]+β is used to compensate for illumination variations, where p[x] is the reference sample point pointed to by the MV at position x on the reference image. When surround motion compensation is enabled, surround offset should be considered to clip the MV. Since α and β can be derived based on the current block template and the reference block template, they require no signaling overhead except for signaling the LIC flag for AMVP mode to indicate the use of LIC.
[0277] The local illumination compensation proposed in JVET-O0066 is used for unidirectional prediction inter-frame CU with the following modifications.
[0278] Intra-frame adjacent samples can be used for LIC parameter export; Disable LIC for blocks with fewer than 32 luminance samples; For both non-subblock and affine modes, LIC parameter export is performed based on the template block samples corresponding to the current CU, rather than the partial template block samples corresponding to the first top-left 16 × 16 element. Samples of the reference block template are generated by using a MC with block MV without rounding it to integer pixel precision.
[0279] OBMC
[0280] When applying OBMC, as described in JVET-L0101, motion information from neighboring blocks and weighted predictions are used to refine the upper and left boundary pixels of the CU.
[0281] The following conditions should not be used for OBMC: When OBMC is disabled at the SPS level When the current block has intra-frame mode or IBC mode When the current block applies a LIC When the current brightness block area is less than or equal to 32 Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using motion information from adjacent sub-blocks. Enabled for sub-block-based encoding tools: Affine AMVP mode; Affine merging mode and sub-block-based temporal motion vector prediction (SbTMVP). Bilateral matching based on sub-blocks.
[0282] When OBMC mode is used with CIIP mode with LMCS, inter-frame blending is performed before LMCS mapping of inter-frame samples. LMCS is applied to the blended inter-frame samples, which are then combined with intra-frame samples in CIIP mode where LMCS is applied.
[0283] in, This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by motion through adjacent blocks in the original domain, and and It's the weight.
[0284] OBMC based on template matching
[0285] In the template matching-based OBMC scheme, instead of directly using weighted prediction, the predicted value of the CU boundary sample derivation method is determined based on the template matching cost, including using only the motion information of the current block, or using the motion information of adjacent blocks, or a hybrid mode.
[0286] In this scheme, for each block with a size of 4×4 at the top CU boundary, the template size is equal to 4×1. If N adjacent blocks have the same motion information, the template size is enlarged to 4N×1 because the MC operation can process them all at once. For each left block with a size of 4×4 at the left CU boundary, the left template size is equal to 1×4 or 1×4N (…). Figure 20 ).
[0287] For each 4×4 top block (or N groups of 4×4 blocks), follow these steps to derive the predicted values for the boundary samples.
[0288] Taking block A as the current block, and its adjacent block AboveNeighbor_A as an example, the operations on the left block are performed in the same way.
[0289] First, based on the following three types of motion information, the three template matching costs (Cost1, Cost2, Cost3) are measured by the SAD between the reconstructed sample points of the template and their corresponding reference sample points derived through the MC process: Calculate Cost1 based on the motion information of A.
[0290] Cost2 is calculated based on the motion information of AboveNeighbor_A.
[0291] Cost3 is calculated based on the weighted predictions of motion information of A and AboveNeighbor_A, which have weighting factors of 3 / 4 and 1 / 4 respectively.
[0292] Secondly, by comparing Cost1, Cost2, and cost3, a method is selected to calculate the final prediction result of the boundary sample points.
[0293] The original MC result using the motion information of the current block is represented as Pixel1, and the MC result using the motion information of neighboring blocks is represented as Pixel2. The final prediction result is represented as NewPixel.
[0294] If Cost1 is the minimum, then NewPixel(i,j) = Pixel1(i,j).
[0295] If (Cost2 + (Cost2 >> 2) + (Cost2 >> 3)) <= Cost1, then use mixed mode 1.
[0296] For a luminance block, the number of mixed pixel rows is 4.
[0297] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16) 5
[0298] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4) 3
[0299] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8) 4
[0300] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16) 5
[0301] For chroma blocks, the number of mixed pixel rows is 1.
[0302] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16) 5
[0303] If Cost1 <= Cost2, then use Mixed Mode 2.
[0304] For a luminance block, the number of mixed pixel rows is 2.
[0305] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8) 4
[0306] NewPixel(i,1)=(31×Pixel1(i,1)+Pixel2(i,1)+16) 5
[0307] For chroma blocks, the number of mixed pixel rows / columns is 1.
[0308] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8) 4
[0309] Otherwise, use mixed mode 3.
[0310] For a luminance block, the number of mixed pixel rows is 4.
[0311] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4) 3
[0312] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8) 4
[0313] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16) 5
[0314] For chroma blocks, the number of mixed pixel rows is 1.
[0315] NewPixel(i,0)=(7×Pixel1(i,0)+Pixel2(i,0)+4) 3
[0316] Currently, the IBC tool is not combined with the GPM tool. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0317] Currently, coded blocks encoded in IBC mode are not combined with coded blocks encoded in intra-frame or inter-frame modes. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0318] Currently, the weights of intra-frame and inter-frame coded blocks in CIIP are predefined in a fixed manner. Therefore, this disclosure provides an example of adaptively determining weights based on a template matching method, which can improve prediction accuracy and coding performance.
[0319] Currently, the number of block vectors (BVs) in IBC tools is singular. Therefore, this disclosure provides examples of increasing the number of block vectors (BVs) and combining prediction results, which can improve prediction accuracy and coding performance.
[0320] Currently, coded blocks encoded in intra-frame TMP mode are not combined with coded blocks encoded in intra-frame or inter-frame modes. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0321] Currently, intra-frame TMP tools are not combined with GPM tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0322] Currently, the IBC tool is not combined with the TIMD tool. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0323] Currently, intra-frame TMP tools are not combined with TIMD tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0324] Currently, intra-frame TMP tools are not combined with LIC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0325] Currently, IBC tools are not combined with OBMC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0326] Currently, intra-frame TMP tools are not combined with OBMC tools. Therefore, this disclosure provides examples of combining them, which can improve prediction accuracy and coding performance.
[0327] Currently, the candidate derivation process of IBC merge mode and IBC AMVP mode only uses adjacent neighbor blocks and some non-adjacent neighbor blocks in the upper left region. Therefore, this disclosure provides a further expansion to include more non-adjacent neighbor blocks, which can improve prediction accuracy and coding performance.
[0328] Currently, IBC mode typically uses block-level BV for motion compensation. Therefore, this disclosure provides a further introduction of sub-block-based IBC mode, which can improve prediction accuracy and coding performance.
[0329] Currently, TM IBC mode and TM regular inter-frame mode use both left and top templates for motion refinement. Therefore, this disclosure provides a further extended template mode, which can improve prediction accuracy and coding performance.
[0330] Currently, deblocking filters treat blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode differently when obtaining boundary strength. Therefore, a unified approach is provided to unify these two modes, which can improve coding performance.
[0331] In order to address the aforementioned problems, this disclosure provides a method for further improving the existing design of IBC. Generally, the main features of the technology proposed in this disclosure are summarized below.
[0332] The IBC tool can be combined with the GPM tool, and the combination can be in the form of a GPM with IBC and IBC prediction, a GPM with IBC and intra-frame prediction, or a GPM with IBC and inter-frame prediction.
[0333] As a simplified version of the IBC tool combined with the GPM tool, for a predefined orientation (such as 45 degrees), the upper left part is predicted using intra-frame mode and the lower right part is predicted using IBC mode, and then they are averaged and weighted to obtain the final predicted signal.
[0334] The IBC tool is combined with the CIIP tool, where IBC prediction is combined with intra-frame prediction mode, or IBC prediction is combined with inter-frame prediction mode.
[0335] In CIIP, the weights of intra-coded blocks and inter-coded blocks are adaptively determined based on a template matching method.
[0336] The IBC tool is combined with the MHP tool, which obtains more than one BV prediction and then performs a weighted average of them to obtain the final prediction signal.
[0337] The intra-frame TMP tool is combined with the CIIP tool, where intra-frame TMP is combined with intra-frame prediction mode, or intra-frame TMP is combined with inter-frame prediction mode.
[0338] The intra-frame TMP tool can be combined with the GPM tool. The combination can be a combination of GPM with intra-frame TMP and intra-frame TMP prediction, a combination of GPM with intra-frame TMP and intra-frame prediction, or a combination of GPM with intra-frame TMP and inter-frame prediction.
[0339] As a simplified version of the intra-frame TMP tool combined with the GPM tool, for a predefined orientation (such as 45 degrees), the upper left part is predicted using the intra-frame mode and the lower right part is predicted using the intra-frame TMP mode, and then they are averaged and weighted to obtain the final predicted signal.
[0340] The IBC tool is combined with the TIMD tool, where the IBC mode is used in conjunction with the intra-prediction mode in MPM for TIMD fusion.
[0341] The intra-frame TMP tool is combined with the TIMD tool, where the intra-frame TMP mode is used in conjunction with the intra-frame prediction mode in MPM for TIMD fusion.
[0342] The intra-frame TMP tool is combined with the LIC tool, where the LIC tool is used to compensate for local illumination variations between the current block and its intra-frame TMP prediction block.
[0343] The IBC tool is combined with the OBMC tool, where the top and left boundary pixels of the current block predicted by IBC are refined by the OBMC tool.
[0344] The intra-frame TMP tool is combined with the OBMC tool, where the top and left boundary pixels of the current block predicted by the intra-frame TMP are refined by the OBMC tool.
[0345] The candidate derivation process for IBC merge mode or IBC AMVP mode is expanded by using not only adjacent neighbor blocks but also non-adjacent neighbor blocks.
[0346] The IBC mode is extended to the sub-block level, where sub-blocks in the current block have their own BV for motion compensation.
[0347] Template modes for extended TM IBC mode and TM regular inter-frame mode, where motion refinement is performed using only the left template, only the top template, etc.
[0348] The deblocking filter treats blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode equally when obtaining boundary strength.
[0349] In some examples, the disclosed methods can be applied independently or in combination.
[0350] GPM with IBC and IBC forecast
[0351] According to one or more embodiments of this disclosure, an IBC tool is combined with a GPM tool in the form of IBC and GPM with IBC prediction. Different methods can be used to achieve this objective.
[0352] In the first approach, the two "inter-frame" portions of the GPM (Genre Prediction Processing) with inter-frame and inter-prediction methods in the VVC are replaced with IBCs. This means that the prediction results of the two IBCs are merged and weighted averaged according to the dividing lines in the coded block. The weights can be obtained by referring to the GPM with inter-frame and inter-prediction methods in the VVC.
[0353] In the second approach, the two “inter-frame” parts of the GPM with inter-frame and inter-frame prediction methods in the ECM are replaced with IBC, where some template matching tools can be used to further improve coding performance.
[0354] When combining IBC tools with GPM tools in the form of a GPM with IBC and IBC predictions, the IBC prediction results can come from regular merge candidates, TM refined merge candidates, or merge candidates with block vector difference (MBVD). In some examples, regular merge candidates, TM refined merge candidates, or MBVD candidates can exist in the ECM. In the first approach, the two IBC prediction results come from the same type of merge candidate. For example, both IBC prediction results come from regular merge candidates, or both from TM refined merge candidates, or both from MBVD candidates, where the merge indices of the two IBC prediction results are different. In the second approach, the two IBC prediction results come from different kinds of merge candidates. For example, one IBC prediction result comes from a regular merge candidate, while the other IBC prediction result comes from a TM refined merge candidate. When TM refined merge candidates are used in a GPM with IBC and IBC predictions, the TM refined merge candidates can be directly reused in a GPM with IBC and IBC predictions, or similarly, different templates can be used for different parts of the GPM partitions of a predefined GPM partitioning pattern.
[0355] When combining an IBC tool with a GPM tool in the form of a GPM with IBC and IBC prediction, different methods can be used to encode GPM segmentation patterns. In the first method, similar to GPM in VVC, all allowed segmentation patterns are encoded with equal probability. In the second method, all allowed GPM segmentation patterns are divided into groups, and two indices are encoded to identify the transmitted GPM segmentation pattern, where the first index is used to determine which group is used for signal transmission, and the second index is used to determine a specific index within the selected group for signal transmission. For example, all allowed GPM segmentation patterns are divided into two groups, with GPM segmentation patterns along the horizontal or vertical direction forming one group and GPM segmentation patterns along other directions forming another. First, a flag indicating context encoding or bypass encoding is encoded to determine which group the transmitted GPM segmentation pattern belongs to for signal transmission, and then indices encoded with equal probability are encoded to determine which index within the selected group the transmitted GPM segmentation pattern belongs to for signal transmission. In the third method, a TM-based method is used to encode GPM segmentation patterns. In one example, similar to GPM in ECM, a TM-based method is used to reorder all allowed GPM partition patterns, and then a signaled index is sent, using Golomb-Rice code to indicate the exact position of the GPM partition pattern in the reordered list. In some examples, reordering may be based on comparison template matching costs. In another example, after dividing all allowed GPM partition patterns into groups, a TM-based method is used to reorder the GPM partition patterns in the selected groups, and then a signaled index is sent, using Golomb-Rice code to indicate the exact position of the GPM partition pattern in the reordered selected group. In a fourth method, since both the GPM partition patterns and the two merge indices of the two GPM partitions need to be sent to the bitstream, similar to spatial GPM, a TM-based method is used to reorder all combinations of the GPM partition patterns and the two merge indices of the two GPM partitions, and then a signaled index is sent, using Golomb-Rice code to indicate the exact position of the corresponding combination of the GPM partition patterns and the two merge indices of the two GPM partitions in the reordered list.
[0356] When combining IBC tools with GPM tools in the form of GPM with IBC and IBC prediction, different methods can be used to mix the two GPM partitions. In the first method, adaptive mixing is utilized. For CUs using GPM encoding with IBC and IBC prediction, comparisons are made during the RDO process as follows: Figure 36The first method shows different blending widths, and the index of the selected blending width is sent to the bitstream via a signal. The second method allows the blending width to be selected using predefined criteria, eliminating the need for an index to be sent to the bitstream. In one example, similar to spatial GPM, the blending width is selected based on the width and height of the current CU. In another example, similar to GPM in VVC, only one blending width is used for all CU sizes. The third method uses different blending methods for different types of content. For example, hard blending with a blending width of zero is used for screen content; adaptive blending is used for natural content.
[0357] When combining the IBC tool with the GPM tool in the form of a GPM with IBC and IBC prediction, different methods can be used to traverse motion information of a CU encoded with GPM using IBC and IBC prediction. In the first method, if the center position of a 4x4 block in the current CU is located within GPM partition A, it is filled with the block vector of the merge index of GPM partition A, and vice versa for GPM partition B. It should be noted that in this method, partitions A and B are adjusted by GPM partition lines, meaning that both partitions A and B can contain some mixed regions. In the second method, if the center position of a 4x4 block in the current CU is located within a non-mixed region of GPM partition A or B, the center position is filled only with the block vector of the merge index of the corresponding GPM partition. If the center position of a 4x4 block in the current CU is located within a mixed region of a GPM partition, the center position is filled with the weighted average of the block vectors of the merge index of the two GPM partitions. In the third method, regardless of where the GPM partition line is located, the motion information of the current CU is only filled with the block vector of either GPM partition A or GPM partition B.
[0358] GPM with IBC and Intra-frame Prediction
[0359] According to one or more embodiments of this disclosure, an IBC tool is combined with a GPM tool in the form of IBC and intra-frame prediction GPM. Different methods can be used to achieve this goal.
[0360] In the first method, the “inter-frame” portion of the GPM, which has both inter-frame and intra-frame prediction methods, in the ECM is replaced with IBC, wherein the IBC prediction results are combined and weighted averaged with the intra-frame prediction results to obtain the final prediction signal.
[0361] The unified GPM with intra-frame and intra-frame prediction, GPM with IBC and intra-frame prediction, and GPM with IBC and IBC prediction.
[0362] According to one or more embodiments of this disclosure, coding tools for GPM with intra-frame and intra-frame prediction, GPM with IBC and intra-frame prediction, and GPM with IBC and IBC prediction are unified into a single set of syntax elements.
[0363] Specifically, first, a flag indicating whether the CU is GPM-coded is encoded into the bitstream. If this flag is true, the GPM split mode is sent to the bitstream. For a GPM split partition, first, a flag indicating whether the GPM split partition is intra-coded is encoded into the bitstream. If this flag is true, the intra-mode index is sent to the bitstream; otherwise, the IBC merge index is sent to the bitstream. For another GPM split partition, the same syntax elements are sent. It should be noted that when both GPM split partitions are intra-coded or both are IBC-coded, the intra-mode index or IBC merge index of the two GPM split partitions will be different.
[0364] GPM with IBC and inter-frame prediction
[0365] According to one or more embodiments of this disclosure, an IBC tool is combined with a GPM tool in the form of a GPM with IBC and inter-frame prediction. Different methods can be used to achieve this objective.
[0366] In the first method, an “inter-frame” portion of the GPM with inter-frame and inter-frame prediction methods in the VVC is replaced with an IBC, wherein the IBC merged prediction results are weighted and averaged with the inter-frame merged prediction results to obtain the final prediction signal.
[0367] In the second approach, an "inter-frame" section of the GPM with inter-frame and inter-frame prediction methods in the ECM is replaced with an IBC, where template matching tools can be used to further improve coding performance.
[0368] Simplified IBC and intra-frame prediction combination in GPM form
[0369] According to one or more embodiments of this disclosure, combining IBC tools with GPM tools in a simplified form of GPM with IBC and intra-frame prediction, such as combining IBC and intra-frame prediction in a specific segmentation mode, can save bit overhead in segmentation representation. Different methods can be used to achieve this goal.
[0370] In the first method, for a dividing line, such as 45 degrees, the upper left portion of the coded block is encoded in intra-frame prediction mode and the lower right portion of the coded block is encoded in IBC prediction mode, and then they are averaged in GPM form to obtain the final predicted signal.
[0371] Combined IBC - Intra / Inter-frame Prediction
[0372] According to one or more embodiments of this disclosure, coded blocks encoded in IBC mode are combined with coded blocks encoded in intra-frame mode or inter-frame mode. Different methods can be used to achieve this objective.
[0373] In the first approach, the decoder / encoder can combine coded blocks encoded in IBC mode with coded blocks encoded in intra-frame mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, coded blocks encoded in IBC merge mode are treated as coded blocks encoded in inter-frame merge mode, and the coded blocks encoded in IBC merge mode are combined with coded blocks encoded in planar intra-frame prediction mode. In another example, similar to the combination of CIIP in ECM with TIMD and TM merging techniques, coded blocks encoded in IBC merge-TM mode are combined with coded blocks encoded in TIMD-derived intra-frame prediction mode.
[0374] When combining IBC-coded blocks with intra-frame-coded blocks, the weights can be designed similarly to the CIIP technique in VVC and the combination of CIIP with TIMD and TM merging techniques in ECM. That is, 1) the weights of both IBC-coded blocks and intra-frame-coded blocks are greater than zero and less than one, or the weight of the intra-frame-coded block gradually changes from one to zero from one region to another in the current block (the weight of the IBC-coded block is the opposite); 2) the weights of IBC-coded blocks and intra-frame-coded blocks can be determined based on the coding modes of adjacent blocks and the intra-frame mode of the current block; 3) the weights of IBC-coded blocks and intra-frame-coded blocks can be uniform throughout the current block or different in different positions within the current block.
[0375] For example, the weights of IBC-coded blocks and intra-coded blocks can be determined as follows: When the upper and left adjacent blocks of the current block are both intra-coded and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 1 / 4 and 3 / 4, respectively. When the upper and left adjacent blocks of the current block are both IBC-coded and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 3 / 4 and 1 / 4, respectively. When one upper or left adjacent block is IBC-coded, the other adjacent block is intra-coded, and the intra-mode of the current block is planar, the weights of the IBC-coded blocks and intra-coded blocks in the entire current block are 1 / 2 and 1 / 2, respectively.
[0376] When the intra-frame mode of the current block is close to the horizontal angle mode (2 <= angle mode index < 34), such as Figure 16A The current block is vertically divided as shown; when the intra-frame mode of the current block is close to the vertical angle mode (34 <= angle mode index <= 66), as shown... Figure 16B As shown, the current block is horizontally divided. Table 7 shows the weights of the IBC coded blocks (wIBC) and intra-coded blocks (wIntra) for different sub-blocks. Furthermore, the weights of the IBC and intra-coded blocks can be determined in the CIIP-PDPC version. In this version, the intra-mode of the current block is set to planar mode. As the combination position moves from the upper left to the lower right in the current block, the weight of the intra-coded block gradually decreases, and vice versa for the weight of the IBC coded block.
[0377]
[0378] Table 7. Modified weights for angle mode.
[0379] When combining IBC-encoded blocks with intra-frame-encoded blocks, the weights can also be masked, meaning the weights of the IBC and intra-frame-encoded blocks can be one or zero for different regions of the current block. Specific weights for IBC and intra-frame-encoded blocks can be determined based on the coding modes of adjacent blocks and the intra-frame mode of the current block. For example, the weights of IBC and intra-frame-encoded blocks can be determined as follows: when the intra-frame mode of the current block is close to the horizontal angle mode (2 <= angle mode index < 34), if the upper and left adjacent blocks of the current block are both intra-frame-encoded, then the weight of the intra-frame-encoded block is one in the left 3 / 4 region of the current block and zero in the right 1 / 4 region of the current block. Figure 21 As shown in (a), the weights of IBC coded blocks are conversely related; if only one adjacent block is intra-coded, the weight of the intra-coded block is one in the left half region of the current block and zero in the right half region of the current block, as shown in (a). Figure 21 As shown in (b), the weights of IBC coded blocks are conversely related; if neither the upper nor left adjacent block of the current block is intra-coded, the weight of the intra-coded block is one in the left 1 / 4 region of the current block and zero in the right 3 / 4 region of the current block, as shown in (b). Figure 21 As shown in (c), the weights for IBC coded blocks are also the opposite.
[0380] When the intra-frame mode of the current block is close to the vertical angle mode (34 <= angle mode index <= 66), if the upper and left adjacent blocks of the current block are both intra-coded, then the weight of the intra-coded block is one in the top 3 / 4 region of the current block and zero in the bottom 1 / 4 region of the current block. Figure 21 As shown in (d), the weights of IBC coded blocks are conversely related; if only one adjacent block is intra-coded, the weight of the intra-coded block is one in the top half region of the current block and zero in the bottom half region of the current block, as shown in (d). Figure 21As shown in (e), the weights of IBC coded blocks are conversely related; if neither the upper nor left adjacent block of the current block is intra-coded, then the weight of the intra-coded block is one in the top 1 / 4 region of the current block and zero in the bottom 3 / 4 region of the current block, as shown in (e). Figure 21 As shown in (f), the weights for IBC coded blocks are the same, and vice versa.
[0381] When the intra-frame mode of the current block is planar mode, if the upper and left adjacent blocks of the current block are both intra-coded, then the weight of the intra-coded block is one in the upper left 3 / 4 region of the current block (horizontal index less than 1 / 2 width of the current block or vertical index less than 1 / 2 height of the current block), and zero in the lower right 1 / 4 region of the current block (horizontal index equal to or greater than 1 / 2 width of the current block and vertical index equal to or greater than 1 / 2 height of the current block). Figure 21 As shown in (g), the weights of IBC coded blocks are conversely related; if only the upper adjacent blocks are intra-coded, the weight of the intra-coded block is one in the top half region of the current block and zero in the bottom half region of the current block, as shown in (g). Figure 21 As shown in (e), the weights of IBC coded blocks are conversely related; if only the left adjacent block is intra-coded, the weight of the intra-coded block is one in the left half region of the current block and zero in the right half region of the current block, as shown in (e). Figure 21 As shown in (b), the weights of IBC coded blocks are conversely related; if neither the upper nor left adjacent blocks of the current block are intra-coded, the weight of the intra-coded block is one in the upper left 1 / 4 region of the current block (horizontal index less than 1 / 2 the width of the current block and vertical index less than 1 / 2 the height of the current block), and zero in the lower right 3 / 4 region of the current block (horizontal index equal to or greater than 1 / 2 the width of the current block or vertical index equal to or greater than 1 / 2 the height of the current block), as shown in (b). Figure 21 As shown in (h), the weights for IBC coded blocks are the same, and vice versa.
[0382] The two weight design methods described above can be used independently or in combination. For example, when the intra-frame mode of the current block is planar mode, if neither the upper nor left adjacent block of the current block is intra-coded, the weights of the IBC coded block and the intra-coded block can be designed similarly to the CIIP technique in VVC. Under other conditions, the weights of the IBC coded block and the intra-coded block can be designed in the mask version.
[0383] When combining IBC-coded blocks with intra-frame-coded blocks, weights can be designed based on template matching methods. This involves using the sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template to calculate the weights of the IBC and intra-frame coded blocks. SATD is a widely used block matching criterion in scenarios such as fractional motion estimation in video compression. For example, the weights of IBC and intra-frame coded blocks can be determined as follows: For intra-frame coded blocks, such as... Figure 22 As shown, the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The prediction samples of the current block template are used for intra-frame prediction based on the reference samples of the template in the intra-frame mode of the current block. For IBC coded blocks, such as Figure 23 As shown, the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The prediction samples for the current block template are predicted using reference samples pointed to by the block vector of the current block. The weights of the IBC encoded blocks... Weights of intra-coded blocks The following decision can be made:
[0384] When using the template of the current block to calculate the weights of IBC and intra-coded blocks, if both left and top templates are available, both can be used; or if either the left or top template is available, only the left template or only the top template can be used. The use of both left and top templates, only the top template, or only the left template can be determined during Rate-Distortion Optimization (RDO) or within predefined criteria. RDO techniques typically minimize distortion (video quality loss) relative to the amount of data required to encode the video. For predefined criteria, for example, if the intra-mode of the current block is planar or DC, both left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0385] When combining a coded block encoded in IBC mode with a coded block encoded in intra-mode, the weights derived based on the template matching method can be compared with the weights derived from the CIIP in the reference ECM during RDO (or the weights designed in the mask version). This means that flags at the coded block level need to be sent to the bitstream to signal which method to use; alternatively, the weights derived based on the template matching method can be used to replace all or part of the weights derived from the CIIP in the reference ECM (or the weights designed in the mask version). For example, if the intra-mode of the current block is planar or DC mode, the weights of the intra-coded block and the IBC coded block are determined based on the template matching method; otherwise, the weights are determined by referring to the CIIP in the ECM (or the mask version).
[0386] In the second approach, the decoder / encoder can combine coded blocks encoded in IBC mode with coded blocks encoded in inter-frame mode. Various methods can be used in this combination. In one example, similar to CIIP in VVC, coded blocks encoded in IBC merge mode are treated as coded blocks encoded in planar intra-frame mode, and the IBC merge mode coded blocks are combined with the inter-frame merge mode coded blocks. In another example, coded blocks encoded in IBC merge mode are treated as coded blocks encoded in inter-frame merge mode, and the IBC merge mode coded blocks are combined with the inter-frame merge mode coded blocks by equal averaging.
[0387] In the third approach, the decoder / encoder can combine coded blocks encoded in IBC mode with coded blocks encoded in intra-frame mode and inter-frame mode. Various methods can be used in this combination. In one example, the coded blocks encoded in IBC mode, intra-frame mode, and inter-frame mode are directly combined by equal averaging. In another example, firstly, the coded blocks encoded in IBC mode are combined separately with the coded blocks encoded in intra-frame mode and inter-frame mode, as presented in the first and second approaches. Then, the individual combination results are combined by equal averaging.
[0388] CIIP Improvements
[0389] When designing the weights of intra-coded blocks and inter-coded blocks in CIIP, the weights can be designed based on template matching methods. This involves using the sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template to calculate the weights of the intra-coded blocks and inter-coded blocks. For example, the weights of intra-coded blocks and inter-coded blocks can be determined as follows: For intra-coded blocks, such as... Figure 22As shown, the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The prediction samples of the current block template are used for intra-frame prediction based on the reference samples of the template in the intra-frame mode of the current block. For inter-coded blocks, such as... Figure 23 As shown in the figure (where "BV of IBC merge candidate" is replaced by "MV of inter-frame merge candidate"), the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The prediction samples of the current block template are predicted using reference samples of the template with the motion vector of the current block. The weights of the inter-coded blocks... Weights of intra-coded blocks The following decision was made:
[0390] When using the template of the current block to calculate the weights of inter-coded blocks and intra-coded blocks, if both the left and top templates are available, both can be used; otherwise, if either the left or top template is available, only the left template or only the top template can be used. The use of both the left and top templates, only the top template, or only the left template can be determined during the RDO process or in a predefined standard. For example, if the intra-mode of the current block is planar or DC mode, both the left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0391] When deriving the weights of intra-coded blocks and inter-coded blocks in CIIP using a template matching-based method, the weights derived using the template matching method can be compared with those derived using the original method in the RDO process. This means that flags at the coded block level need to be sent to the bitstream to signal which method was used; alternatively, the weights derived using the template matching method can be used to replace all or part of the weights derived using the original method. For example, if the intra-frame mode of the current block is planar mode or DC mode, the weights of intra-coded blocks and inter-coded blocks in CIIP are determined using the template matching method; otherwise, the weights are determined using the original method.
[0392] Multiple Hypothesis IBC Prediction
[0393] According to one or more embodiments of this disclosure, the number of block vectors (BVs) in the IBC tool is increased to two or more, and two or more hypotheses are combined to obtain the final prediction result. Different methods can be used to achieve this goal.
[0394] In the first approach, the encoder / decoder can combine two hypotheses corresponding to two BVs to obtain the final prediction. Various methods can be used to achieve this. In one example, the two BVs corresponding to the minimum rate distortion metric and the second minimum rate distortion metric in the IBC AMVP mode are averaged equally to obtain the final prediction. In another example, the predictions corresponding to the IBC AMVP mode and the predictions corresponding to the IBC merging mode are averaged equally to obtain the final prediction.
[0395] In the second approach, the decoder / encoder can combine more hypotheses corresponding to more BVs to obtain the final prediction result. Various methods can be used to achieve this. In one example, the iterative accumulation method proposed in the Multiple Hypothesis Prediction (MHP) technique is used to obtain the final prediction result. In another example, all BVs corresponding to the minimum rate distortion metric, the second minimum rate distortion metric, the third minimum rate distortion metric, ..., in the IBC AMVP mode are averaged equally to obtain the final prediction result.
[0396] Predicted block candidate export
[0397] In some embodiments, candidate prediction blocks are searched and selected based on a criterion of minimizing template matching cost, i.e., the top N candidates that result in the minimum BV matching cost are selected. The BV matching cost may not be limited to SAD (sum of absolute differences) and SSE (sum of squared errors).
[0398] In some embodiments, candidate prediction blocks can be selected based on a predefined pattern (i.e., a planar pattern).
[0399] In some embodiments, prediction block candidates can be selected based on adjacent predefined patterns (i.e., top predefined pattern, left predefined pattern).
[0400] Fixed Multiple Hypothesis IBC
[0401] In this embodiment, the weighting factors used to generate the final prediction block are predefined and fixed on both the encoder and decoder sides. As an example, an equal weighting factor, i.e., 1 / N, can be used for all candidate blocks.
[0402] Adaptive Multiple Hypothesis IBC
[0403] To adapt to the diverse characteristics of video content, an adaptive multiple hypothesis IBC method is also proposed.
[0404] In some embodiments, a weighting factor can be derived based on the BV matching cost. The BV matching costs of N candidates are expressed as... …、 The weighting factors are calculated as follows.
[0405] (4)
[0406] It should be noted that BV matching costs can be measured using (but are not limited to) SAD and SSE.
[0407] In another embodiment, the weighting factor can be derived / switched based on the block size or syntax element signaled at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0408] In yet another embodiment, the weighting factors can be derived on the encoder side and then signaled to the decoder in the bitstream. The N candidate prediction blocks are represented as follows: , … The current block is represented as The weighting factor can be solved using the following mathematical formula: (5) The Wiener-Hopf equations can be used as an ALF to solve mathematical expression (5). The derived filter coefficients are then quantized to integer type and sent as signals at the block level.
[0409] In yet another embodiment, the weighting factors can be derived on the encoder side and then signaled to the decoder in the bitstream. The N candidate prediction blocks are represented as follows: , … The current block is represented as The weighting factor can be solved using the following mathematical formula: (6) The mathematical expression (6) can be solved using LDL decomposition or Gaussian elimination.
[0410] In another embodiment, the weighting factor is derived based on a template, and the derived weighting factor is applied to prediction block candidates to generate the final prediction block. The template for the prediction candidate is represented as... , …、 The current block is represented as The weighting factor can then be derived using the following mathematical formula: (7) The Wiener-Hopf equations can be used to find the mathematical expression (7). Then, the final predicted block can be calculated as follows: ,in This represents the i-th prediction block candidate.
[0411] The IBC model utilizes nonlocal correlation to improve prediction accuracy, where similar blocks are searched and used to generate the final prediction block. In this embodiment, a combination of nonlocal mean filtering and multiple hypothesis IBC is proposed, as described below. In the first step, N prediction block candidates are searched and identified, as performed in IBC. In the second step, weighting factors are calculated as follows.
[0412] (8)
[0413] in Used to measure the distance between the template of the i-th prediction block candidate and the template of the current block. Used as a weighting degree, and It is a normalization constant: (9) In order to calculate the weighting factor in mathematical formula (8), the strength of the weighting must first be determined. Several methods are proposed in this disclosure to determine the weighting strength.
[0414] In the first approach, a candidate list of weighted strength values, including some typical weighted strength values, is defined and fixed at both the encoder and decoder sides. At the encoder side, rate-distortion optimization is used to examine the weighted strength values, and the optimal weighted strength value is identified and sent as a signal in the bitstream to the decoder side.
[0415] In the second method, the templates of the predicted block candidates and the template of the current block are used to estimate the weighted intensity value. The template of the predicted candidate is represented as... , …、 The current block is represented as The weighted strength value can then be derived using the following mathematical formula: (10) In the third method, the weighted intensity value can be estimated using the variance and QP value of the template of the current block. That is, the relationship between the weighted intensity value, QP value and template variance can be fitted offline.
[0416] To better utilize the nonlocal correlations in IBC, this embodiment uses Singular Value Decomposition (SVD) to generate the final prediction block from the prediction block candidates. The width and height of the current block are represented by W and H, and the area of the current block is represented by... .
[0417] Step 1. As performed in FIBC, search and identify K candidate prediction blocks. .
[0418] Step 2. Current block K predicted block candidate building block group And it is arranged as a matrix: (11) in, It is by grouping Each candidate arrangement in the array is a column vector of size . The matrix.
[0419] Step 3. For the matrix Perform SVD decomposition.
[0420] (12)
[0421] Step 4. For the singular value matrix Apply soft threshold calculation.
[0422] (13)
[0423] in Using threshold shrink A function of the diagonal elements. For The k-th diagonal element in the layer, it is layered nonlinear function at the point shrink: (14) It is a contraction singularity at the diagonal position The matrix formed.
[0424] Step 5. Perform inverse SVD to obtain filtered block groups.
[0425] (15)
[0426] One of the key steps is determining the threshold for each diagonal element in step 4. In this disclosure, the threshold is calculated as follows. The threshold is estimated for each group of image patches using the following mathematical formula: (16) in It is the standard deviation of the noise, and The original block is used in the group The standard deviation of the original block in the k-th dimension of the SVD space. The deviation of the original block in the SVD space is estimated as follows.
[0427] (17)
[0428] in yes The k-th singular value. When When the value is zero, the soft thresholding operation is skipped. Additionally, using... and A parameterized power function is used to estimate the noise bias by utilizing the bias of the prediction block.
[0429] (18)
[0430] in The calculation is as follows: (19) here Represents the candidate vector of the prediction block The i-th pixel.
[0431] Multi-hypothesis IBC signaling
[0432] In this disclosure, the proposed multi-hypothesis IBC can be used as an alternative to the current IBC mode, or the encoder can adaptively select either the IBC mode or the multi-hypothesis IBC mode.
[0433] In some embodiments, multiple hypothesis IBC can be used as an alternative to the current IBC model, i.e., always using multiple hypotheses for prediction.
[0434] In another embodiment, one of the multi-hypothesis IBC methods described above is used in conjunction with the current IBC mode. A flag is signaled in the bitstream to indicate whether the multi-hypothesis IBC mode is applied to the CU.
[0435] In another embodiment, more than one of the multi-hypothesis IBC methods described above is used in conjunction with the current IBC mode. First, a signal flag is sent in the bitstream to indicate whether a multi-hypothesis IBC mode is applied. Then, a signal index is sent to indicate which of the multi-hypothesis IBC methods is applied to the CU.
[0436] In another embodiment, the multi-hypothesis IBC method described above is used in conjunction with the current IBC mode. Multi-hypothesis IBC can be used as an alternative to the current IBC mode based on certain coding information of the current block, such as SAD (Sum of Absolute Differences), SSE (Sum of Squared Errors), quantization parameters (QP) associated with TB / CB and / or slices, the neighboring prediction modes of the CU (e.g., IBC mode or intra-frame or inter-frame mode), and / or slice type (e.g., I slice, P slice, or B slice).
[0437] Combined intra-frame TMP - intra / inter-frame prediction
[0438] According to one or more embodiments of this disclosure, a coded block encoded in intra-frame TMP mode is combined with a coded block encoded in intra-frame mode or inter-frame mode. Different methods can be used to achieve this objective.
[0439] In the first approach, the decoder / encoder can combine coded blocks encoded in intra-frame TMP mode with coded blocks encoded in intra-frame mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, coded blocks encoded in intra-frame TMP mode are treated as coded blocks encoded in inter-frame combining mode, and the coded blocks encoded in intra-frame TMP mode are combined with coded blocks encoded in planar intra-frame prediction mode. In another example, similar to the combination of CIIP in ECM with TIMD and TM combining techniques, coded blocks encoded in intra-frame TMP mode are combined with coded blocks encoded in TIMD-derived intra-frame prediction mode.
[0440] When combining coded blocks encoded in intra-TMP mode with coded blocks encoded in intra-frame mode, the weights can be determined by referring to the weight design of CIIP technology in VVC or ECM, or the weights can be designed based on template matching methods. The weights of the intra-TMP coded block and the intra-frame coded block can be calculated using the sum of absolute differences (SAD), sum of squared differences (SSD), or sum of absolute transform differences (SATD) between the predicted and reconstructed samples of the current block template. For example, the weights of the intra-TMP coded block and the intra-frame coded block can be determined as follows: For the intra-frame coded block, such as... Figure 22 As shown, the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The predicted samples of the current block template are used for intra-frame prediction based on the reference samples of the template in the intra-frame mode of the current block. For intra-frame TMP coded blocks, the SATD between the predicted samples and reconstructed samples of the current block template is calculated as follows: The reference samples pointed to by the block vector of the first block are used to predict the predicted samples of the current block template. Weights of intra-frame TMP coded blocks. Weights of intra-coded blocks The following decision was made:
[0441] When using the template of the current block to calculate the weights of intra-TMP coded blocks and intra-coded blocks, if both the left and top templates are available, both can be used; otherwise, if either the left or top template is available, only the left template or only the top template can be used. The use of both the left and top templates, only the top template, or only the left template can be determined during the RDO process or in a predefined standard. For example, if the intra-mode of the current block is planar or DC mode, both the left and top templates are used; if the intra-mode of the current block is close to a horizontal angle mode (2 <= angle mode index < 34), only the left template is used; if the intra-mode of the current block is close to a vertical angle mode (34 <= angle mode index <= 66), only the top template is used.
[0442] When combining a coded block encoded in intra-TMP mode with a coded block encoded in intra-mode, the weights derived from the template matching method can be compared with the weights derived from the CIIP in the reference ECM during RDO. This means that flags at the coded block level need to be sent to the bitstream to signal which method to use; alternatively, the weights derived from the template matching method can be used to replace all or part of the weights derived from the CIIP in the reference ECM. For example, if the intra-mode of the current block is planar or DC mode, the weights of the intra-coded block and the IBC coded block are determined based on the template matching method; otherwise, the weights are determined by referring to the CIIP in the ECM.
[0443] In the second approach, the decoder / encoder can combine coded blocks encoded in intra-TMP mode with coded blocks encoded in inter-frame mode. Various methods can be employed in this combination. In one example, similar to CIIP in VVC, coded blocks encoded in intra-TMP mode are treated as coded blocks encoded in planar intra-frame mode, and these blocks are combined with coded blocks encoded in inter-frame merge mode. In another example, coded blocks encoded in intra-TMP mode are treated as coded blocks encoded in inter-frame merge mode, and these blocks are combined by equal averaging.
[0444] In the third approach, the decoder / encoder can combine coded blocks encoded in intra-frame TMP mode with coded blocks encoded in both intra-frame and inter-frame modes. Various methods can be employed in this combination. In one example, the coded blocks encoded in intra-frame TMP mode, intra-frame mode, and inter-frame mode are directly combined by equal averaging. In another example, firstly, the coded blocks encoded in intra-frame TMP mode are combined separately with the coded blocks encoded in both intra-frame and inter-frame modes, as presented in the first and second approaches. Then, the individual combination results are combined by equal averaging.
[0445] GPM with intra-frame TMP and intra-frame TMP prediction
[0446] According to one or more embodiments of this disclosure, an intra-TMP tool is combined with a GPM tool in the form of intra-TMP and intra-TMP prediction. Different methods can be used to achieve this objective.
[0447] In the first method, the two "inter-frame" portions of the GPM with inter-frame and inter-frame prediction methods in the VVC are replaced with intra-frame TMPs. This means that the prediction results of the two intra-frame TMPs are weighted and averaged against each other based on the dividing lines in the coded block. The weights can be obtained by referring to the GPM with inter-frame and inter-frame prediction methods in the VVC.
[0448] In the second approach, the two “inter-frame” parts of the GPM in the ECM, which have inter-frame and inter-frame prediction methods, are replaced with intra-frame TMPs, where some template matching tools can be used to further improve coding performance.
[0449] GPM with intra-frame TMP and intra-frame prediction
[0450] According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a GPM tool in the form of intra-frame TMP and intra-frame prediction GPM. Different methods can be used to achieve this objective.
[0451] In the first method, the “inter-frame” portion of the GPM, which has both inter-frame and intra-frame prediction methods, in the ECM is replaced with an intra-frame TMP, wherein the intra-frame TMP prediction results are weighted and averaged with the intra-frame prediction results to obtain the final prediction signal.
[0452] GPM with intra-frame TMP and inter-frame prediction
[0453] According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a GPM tool in the form of intra-frame TMP and inter-frame prediction GPM. Different methods can be used to achieve this objective.
[0454] In the first method, an “inter-frame” portion of the GPM in the VVC, which has inter-frame and inter-frame prediction methods, is replaced with an intra-frame TMP, wherein the intra-frame TMP prediction results are weighted and averaged with the inter-frame merged prediction results to obtain the final prediction signal.
[0455] In the second approach, an "inter-frame" portion of the GPM in the ECM, which has inter-frame and inter-frame prediction methods, is replaced with an intra-frame TMP, where template matching tools can be used to further improve coding performance.
[0456] Simplified combination of intra-frame TMP and intra-frame prediction in GPM form
[0457] According to one or more embodiments of this disclosure, combining intra-TMP tools with GPM tools in a simplified form of GPM with intra-TMP and intra-prediction, such as combining intra-TMP and intra-prediction in a specific segmentation mode, can save bit overhead in segmentation representation. Different methods can be used to achieve this goal.
[0458] In the first method, for a dividing line, such as 45 degrees, the upper left portion of the coded block is encoded in intra-frame prediction mode, and the lower right portion of the coded block is encoded in intra-frame TMP prediction mode. They are then averaged in GPM form to obtain the final predicted signal.
[0459] Combining IBC with TIMD mode
[0460] According to one or more embodiments of this disclosure, the IBC tool is combined with the TIMD tool. Different methods can be used to achieve this objective.
[0461] In the first method, the IBC mode is treated as an intra-prediction mode added to the MPM list. Then, the template matching cost is used to compare the IBC mode with other intra-prediction modes in the MPM list. Finally, the TIMD method is used to fuse the two modes with the minimum cost and the second minimum cost to obtain the final prediction result.
[0462] In the second method, the regular TIMD prediction results are first obtained, then the template matching cost of the IBC model and the regular TIMD prediction results is calculated, and finally the TIMD method is used to fuse the IBC model and the regular TIMD prediction results to obtain the final prediction result.
[0463] Combining intra-frame TMP with TIMD mode
[0464] According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a TIMD tool. Different methods can be used to achieve this goal.
[0465] In the first method, the intra-TMP mode is treated as an intra-prediction mode added to the MPM list. Then, the intra-TMP mode is compared with other intra-prediction modes in the MPM list using template matching cost. Finally, the TIMD method is used to fuse the two modes with the minimum cost and the second minimum cost to obtain the final prediction result.
[0466] In the second method, the regular TIMD prediction results are first obtained, then the template matching cost of the intra-frame TMP mode and the regular TIMD prediction results is calculated, and finally the TIMD method is used to fuse the intra-frame TMP mode and the regular TIMD prediction results to obtain the final prediction result.
[0467] Combine intra-frame TMP with LIC
[0468] According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with a LIC tool. Different methods can be used to achieve this goal.
[0469] In the first approach, the intra-frame TMP mode is treated as an inter-frame mode, and LIC is used to model the local illumination variation between the current block and its intra-frame TMP prediction block as a function of the local illumination variation between the current block template and the reference block template. This function is a linear equation as used in the conventional LIC method.
[0470] Combining IBC with OBMC
[0471] According to one or more embodiments of this disclosure, the IBC tool is combined with the OBMC tool. Different methods can be used to achieve this objective.
[0472] In the first approach, the IBC mode is treated as an inter-frame mode, and the conventional OBMC method is applied to refine the top and left boundary pixels of the IBC-encoded CU using block vector information from neighboring blocks with weighted prediction.
[0473] In the second approach, the IBC mode is treated as an inter-frame mode, and the template matching-based OBMC method is applied to refine the top and left boundary pixels of the IBC-encoded CU using a template matching-based approach.
[0474] It should be noted that when IBC is combined with OBMC, for a CU encoded in IBC mode, when using the regular OBMC method or the template matching-based OBMC method to refine the top and left boundary pixels of the current CU using the shift information of adjacent blocks, adjacent blocks can be encoded in IBC mode or intra-frame TMP mode.
[0475] Combining intra-frame TMP with OBMC
[0476] According to one or more embodiments of this disclosure, an intra-frame TMP tool is combined with an OBMC tool. Different methods can be used to achieve this objective.
[0477] In the first approach, the intra-frame TMP mode is treated as an inter-frame mode, and the conventional OBMC method is applied to refine the top and left boundary pixels of the intra-frame TMP coded CU using block vector information of neighboring blocks with weighted prediction.
[0478] In the second approach, the intra-frame TMP mode is treated as an inter-frame mode, and the template-matching-based OBMC method is applied to refine the top and left boundary pixels of the intra-frame TMP-coded CU using a template-matching-based approach.
[0479] It should be noted that when combining intra-frame TMP with OBMC, for a CU encoded in intra-frame TMP mode, when using the regular OBMC method or the template matching-based OBMC method to refine the top and left boundary pixels of the current CU using the shift information of adjacent blocks, adjacent blocks can be encoded in intra-frame TMP mode or IBC mode.
[0480] Non-adjacent candidate exports for IBC AMVP or merge mode
[0481] According to one or more embodiments of this disclosure, the candidate derivation process for IBC merge mode or IBC AMVP mode is extended by using not only adjacent neighbor blocks but also non-adjacent neighbor blocks. The relevant content is summarized in the "Candidate Scan and Candidate Pruning," "Candidate Reordering," "Motion Information Storage," and "Application Scope" sections, and is presented as follows: Candidate scan and candidate pruning For candidate scans, non-adjacent neighbor blocks are scanned and selected using the following method: Scan area and distance: In one or more embodiments, non-adjacent neighboring blocks can be scanned from the left and top regions of the current coded block. The scan distance can be defined as the number of coded blocks from the scan position to the left or top of the current coded block.
[0482] like Figure 24 As shown, multiple rows of non-adjacent neighboring blocks can be scanned to the left or above the current encoded block. Figure 24 The distances shown represent the number of coded blocks to the left or top of the current block from each candidate location. For example, a region to the left of the current block with a "distance 2" indicates that the candidate neighboring block in that region is 2 blocks away from the current block. Similar indications can be applied to other scan regions with different distances.
[0483] In one or more embodiments, non-adjacent neighbor blocks at each distance may have the same block size as the current coded block, as shown in Figure 25(a). Note that when non-adjacent neighbor blocks at each distance have the same block size as the current coded block, the value of the block size adaptively changes according to the partitioning granularity at each different region in the image.
[0484] In some embodiments, non-adjacent neighbor blocks at each distance may have a different block size than the current coded block, as shown in Figure 25(b). Note that when non-adjacent neighbor blocks at each distance have a different block size than the current coded block, the block size value can be predefined as a constant value, such as 4x4, 8x8, or 16x16.
[0485] Based on the defined scan distance, the total size of the scanned region to the left or above the current encoded block can be determined by a configurable distance value. In one or more embodiments, the maximum scan distance to the left and top sides can use the same value or different values. For example, the maximum distance to both the left and top sides shares the same value of 2. The maximum scan distance value can be determined by the encoder side and signaled in the bitstream. Alternatively, the maximum scan distance value can be predefined as a fixed value, such as 2 or 4. When the maximum scan distance is predefined as 4, it indicates that the scanning process should terminate when the candidate list is full or all non-adjacent neighboring blocks with a maximum distance of 4 have been scanned (whichever comes first).
[0486] In one or more embodiments, within each scanned region at a specific distance, the start and end adjacent blocks can be location-dependent.
[0487] In one or more embodiments, for the left-side scan region, the starting neighbor block can be the lower-left neighbor block of the starting neighbor block of an adjacent scan region with a small distance. For example, as... Figure 24 As shown, the starting neighboring block of the "distance 2" scan area to the left of the current block is the adjacent lower-left neighboring block of the starting neighboring block of the "distance 1" scan area. The ending neighboring block can be the adjacent left block of the ending neighboring block of the upper scan area with a smaller distance. For example, as... Figure 24 As shown, the end neighbor block of the "distance 2" scan area to the left of the current block is the adjacent left neighbor block of the end neighbor block of the "distance 1" scan area above the current block.
[0488] Similarly, for the upper scan region, the starting neighbor block can be the upper-right neighbor block of the starting neighbor block of an adjacent scan region with a smaller distance. The ending neighbor block can be the upper-left neighbor block of the ending neighbor block of an adjacent scan region with a smaller distance.
[0489] In one or more embodiments, within each scan region at a specific distance, the sampling interval between the start neighbor block and the end neighbor block can be location-dependent. In one or more embodiments, in scan regions with smaller distances, the sampling interval between the start neighbor block and the end neighbor block is smaller. For example, as... Figure 24 As shown, each neighboring block between the start and end neighboring blocks is scanned in a scan region with a "distance 1"; every two neighboring blocks between the start and end neighboring blocks are scanned in a scan region with a "distance 2". In one or more embodiments, the sampling interval between the start and end neighboring blocks may be the same or different for different side scan regions at a specific distance. For example, the sampling interval between the start and end neighboring blocks is the same for the left and top scan regions at a specific distance.
[0490] Scanning order: When scanning neighboring blocks in a non-adjacent region, certain order and / or rules can be followed to determine the selection of neighboring blocks to be scanned.
[0491] In one or more embodiments, the left-side region may be scanned first, followed by the upper region. For example... Figure 24 As shown, you can first scan the three rows of non-adjacent regions on the left (e.g., from distance 1 to distance 3), and then scan the three rows of non-adjacent regions above the current block.
[0492] In some embodiments, the left and upper regions can be scanned alternately. For example, as Figure 24 As shown, the left scanning area with a "distance of 1" is scanned first, and then the upper area with a "distance of 1" is scanned.
[0493] For scanning areas located on the same side (e.g., the left or upper region), the scanning order is from the smaller distance region to the larger distance region. This order can be flexibly combined with other embodiments of the scanning order. For example, the left and upper regions can be scanned alternately, and the order of regions on the same side is scheduled from the smaller distance to the larger distance.
[0494] Within each scanned region at a specific distance, a scanning order can be defined. In one or more embodiments, for the left scanned region, scanning can begin from the bottom neighboring block to the top neighboring block. For the upper scanned region, scanning can begin from the right block to the left block.
[0495] In one or more embodiments, non-adjacent regions along one direction may be scanned first, followed by scanning non-adjacent regions along other directions. Within a direction, a scanning order can be defined. In one or more embodiments, within each direction, scanning may begin at a smaller distance and progress to a larger distance.
[0496] In some embodiments, non-adjacent regions along different directions can be scanned alternately. For example, such as Figures 26-27 As shown, first, non-adjacent regions with small distances along a direction with an angle value of 225 degrees are scanned. Then, non-adjacent regions with small distances along directions with angle values of 45, 90, 180, and 135 degrees are scanned consecutively. Next, non-adjacent regions with larger distances along a direction with an angle value of 225 degrees are scanned, followed by non-adjacent regions with larger distances along directions with angle values of 45, 90, 180, and 135 degrees.
[0497] In some examples, a total of 18 blocks are scanned, such as Figure 26 As shown, the scan block is indicated by a boxed integer n, where n is in the range of 1 to 18 (inclusive), and n represents the scan order.
[0498] In some examples, a total of 48 blocks are scanned, such as Figure 27 As shown, the scan block is indicated by a boxed integer n, where n is in the range of 1 to 48, and n represents the scan order, including the end value; in these examples, degree values of 270, 0, 247.5, 22.5, 202.5, 67.5, 157.5 and 112.5 can be additionally used to determine the scan order.
[0499] Scan terminated: For non-adjacent candidates, adjacent blocks encoded in IBC mode or intra-frame TMP mode are defined as qualified candidates.
[0500] In one or more embodiments, the scanning process can be performed interactively. For example, a scan performed in a specific region at a specific distance can stop when the first X qualified candidates are identified, where X is a predefined positive value. For example, as Figure 24 As shown, when the first or more qualified candidates are identified, scanning in the left scan region with a distance of 1 can be stopped. Then, the next iteration of the scanning process begins by targeting another scan region, which is regulated by a predefined scan order / rules.
[0501] In one or more embodiments, X can be defined for each distance. For example, at each distance, X is set to 1, meaning that if a first qualified candidate is found, the scan is terminated for each distance, and the scan process restarts from a different distance in the same area or from the same or different distances in different areas. Note that the value of X can be set to the same value or different values for different distances. If the maximum number of qualified candidates are found from all allowed distances of the area (e.g., adjusted by the maximum distance), the scan process for a region is terminated completely.
[0502] In another embodiment, X can be defined for a region. For example, X is set to 3, which means that if the first 3 qualified candidates are found, the scan terminates for the entire region (e.g., the region to the left or above the current block), and the scan process restarts from the same or a different distance in another region. Note that the value of X can be set to the same value or different values for different regions. If the maximum number of qualified candidates are found from all regions, the entire scan process terminates completely.
[0503] The value of X can be defined for both distance and region. For example, for each region (e.g., the region to the left or above the current block), X is set to 3, and for each distance, X is set to 1. For different regions and distances, the value of X can be set to the same value or different values.
[0504] In some embodiments, the scanning process can be performed continuously. For example, scanning in a specific area at a specific distance can be stopped when all adjacent blocks covered have been scanned and no more qualified candidates are identified or the maximum allowed number of candidates has been reached.
[0505] During the candidate scanning process, each candidate non-adjacent neighbor block is determined and scanned by following the scanning method described above. For ease of implementation, each candidate non-adjacent neighbor block can be indicated or located by a specific scanning position. For example, the lower right position is used for both the upper and left non-adjacent neighbor blocks.
[0506] After identifying a qualified candidate by following the above process, that candidate can undergo a similarity check against all existing candidates already in the candidate list. Details of the similarity check can be found in the existing similarity check rules in the current IBC candidate export. If a newly qualified candidate is found to be similar to any existing candidate in the candidate list, the newly qualified candidate is removed / pruned.
[0507] It should be noted that for IBC AMVP and merged candidate export, the above candidate scanning and candidate pruning processes can be the same or different. For example, Figure 26 The candidate scanning and pruning process presented can be used for both IBC AMVP and merging candidate exports. In another example, Figure 26 The candidate scanning and pruning process presented in the document can be used for IBC AMVP candidate export. Figure 27 The candidate scanning and trimming process presented in the document can be used for IBC merge candidate export.
[0508] Candidate Reordering
[0509] When inserting non-adjacent spatial candidates into the IBC candidate list, all non-adjacent spatial candidates can be grouped as a whole and inserted into different positions in the IBC candidate list, or the non-adjacent spatial candidates can be divided into several subgroups and each subgroup can be inserted into a different position in the IBC candidate list.
[0510] In one or more embodiments, non-adjacent space candidates can be inserted into the IBC candidate list by following this order: 1. Spatial BVP from adjacent spatial neighborhoods 2. Spatial BVP from non-adjacent spatial neighborhoods 3. Historical BVP from FIFO table 4. Paired average BVP 5. For example Figure 13 BVP candidates located in the IBC reference region are shown. 6. Zero BVP In another embodiment, non-adjacent spatial candidates can be inserted into the IBC candidate list by following this order: 1. Spatial BVP from adjacent spatial neighborhoods 2. First X-space BVP from a non-adjacent spatial neighborhood 3. Historical BVP from FIFO table 4. Other Y-space BVPs from non-adjacent spatial neighborhoods 5. Paired average BVP 6. For example Figure 13 BVP candidates located in the IBC reference region are shown. 7. Zero BVP The values of X and Y can be predefined fixed values, such as the value 2, or values sent via signals received by the decoder (parameters sent via signals at the sequence / slice / block / CTU level), or configurable values at the encoder / decoder, or values dynamically determined based on the number of available neighbors to the left and above each individual coded block (e.g., X <= 3, Y <= 3), or any combination of methods for determining the values of X and Y. In one example, the value of X can be the same as the value of Y. In another example, the value of X can be different from the value of Y.
[0511] Since placing candidates at the end of the IBC candidate list could incur higher signaling overhead if selected by the encoder and signaled, the order of the above-mentioned different categories of candidates can be designed in the following different ways: In one or more embodiments, the order of the candidates remains the same as the insertion order described above. An adaptive reordering method can be applied to reorder the candidates subsequently; the adaptive reordering method can be a template matching (ARMC) based method.
[0512] In one or more embodiments, before inserting non-adjacent spatial candidates into the IBC candidate list, an adaptive reordering method can first be applied to the derived non-adjacent spatial candidates (the adaptive reordering method can be a template matching method (ARMC)), and then the first X candidates can be inserted into the IBC candidate list based on the above insertion method.
[0513] The value of X can be a predefined fixed value, such as the value 2, or a value sent by a signal received by the decoder (a parameter sent by a signal at the sequence / slice / block / CTU level), or a configurable value at the encoder / decoder, or a value dynamically determined based on the number of available neighbors to the left and above each individual coded block (e.g., X <= 3), or any combination of methods for determining the value of X.
[0514] The reordering method described above can be selected and applied based on various factors: In one or more embodiments, the reordering method can be selected based on the type of video frame / slice. For example, for low-latency images or slices, all non-adjacent spatial candidates can be placed after all neighboring spatial candidates. For non-low-latency images or slices, the first X non-adjacent spatial candidates can be placed after neighboring spatial candidates, and the remaining non-adjacent spatial candidates can be placed after historical BVP candidates.
[0515] It should be noted that the above candidate reordering process can be the same or different for both IBC AMVP and merge candidate list export. For example, for both IBC AMVP and merge candidate list export, all non-adjacent space candidates are placed after all adjacent space candidates. In another example, for IBC AMVP candidate list export, all non-adjacent space candidates are placed after all adjacent space candidates; for IBC merge candidate list export, the first X non-adjacent space candidates can be placed after neighboring space candidates, and the remaining non-adjacent space candidates can be placed after historical BVP candidates.
[0516] Sports information storage
[0517] When scanning non-adjacent spatial neighborhoods based on the candidate derivation methods proposed above, the selected non-adjacent spatial neighborhoods can be either IBC coded blocks or intra-frame TMP coded blocks. In the case of both IBC coded blocks and intra-frame TMP coded blocks, motion information can include translational BV.
[0518] For IBC coded blocks or intra-frame TMP coded blocks, once these blocks have been encoded, their motion information may need to be stored in memory. To save memory usage, non-adjacent spatial neighborhoods can be restricted to a certain region.
[0519] like Figure 28 As shown, the allowed non-adjacent regions used for scanning non-adjacent spatial neighbor blocks can be restricted to a finite region size.
[0520] In one or more embodiments, the restricted region may be applied to an IBC spatial neighbor block or an intra-frame TMP spatial neighbor block.
[0521] The size of the allowed non-adjacent region can be defined based on the size of the current coding tree unit (CTU), for example, an integer (e.g., 1 or 2 or other integers) or a fraction (e.g., 0.5 or 0.25 or other fractions) of the current CTU size.
[0522] The size of the allowed non-adjacent region can be defined based on a fixed number of pixels or samples (e.g., 128 samples above and / or to the left of the current CTU).
[0523] This size (e.g., depending on the CTU size or the number of samples) can be a prefix value determined at the encoder and carried in the bitstream or a value sent by a signal.
[0524] The size of the restricted region can be defined separately for the top and left non-adjacent neighbor blocks. In some examples, a non-adjacent neighbor block can be a non-adjacent spatial neighborhood, a top non-adjacent neighbor block can be a top non-adjacent spatial neighborhood, and a left non-adjacent neighbor block can be a left non-adjacent spatial neighborhood. In one example, the aforementioned non-adjacent neighbor blocks can be restricted to within the current CTU, or outside the current CTU but within a maximum fixed number of samples / pixels away from the top of the current CTU, so that no additional line buffer is needed to store the motion information of the aforementioned non-adjacent neighbor blocks. For example, if 8 sample rows of the adjacent region away from the top of the current CTU have already been covered by the existing line buffer, the fixed number can be defined as 8. In another example, the left non-adjacent neighbor block can be restricted to within the current CTU, or outside the current CTU but within a predefined or signaled number of samples / pixels away from the left boundary of the current CTU.
[0525] As shown in Figure 29, if the allowed non-adjacent regions extend beyond the current CTU, the allowed non-adjacent regions above the current CU (for non-adjacent IBC neighborhoods or intra-frame TMP neighborhoods) may have a large memory cost. In this case, the actual memory cost increases proportionally with the image width and the maximum allowable scan distance in the vertical direction. To reduce the cost of the line buffer (i.e., the aforementioned non-adjacent regions outside the current CTU), the height of the aforementioned non-adjacent regions outside the current CTU can be limited to a value h (as shown in Figure 29(a)). It should be noted that this value of h can be configurable or signaled to the decoder. In the case where IBC motion and intra-frame TMP motion are stored in separate buffers, for the example shown in Figure 29(b), different methods may exist to store motion in the line buffer: In one method, the line buffer used to store IBC motion may indicate that the buffer region where CU B is located is set to invalid because CU B is not an IBC CU. In another method, the line buffer used to store IBC motion may indicate that the buffer region where CU B is located is set to valid and that the IBC motion is copied from CUA because CUA is a neighboring IBC neighborhood of CU B.
[0526] In the example of Figure 29, for ease of implementation, the height value h and the width value w can be set to multiples of 4. In one embodiment, the values h and w can be set to the minimum value of the IBC CU (e.g., 4).
[0527] When the allowed non-adjacent regions used for scanning non-adjacent spatial neighbor blocks are limited to a finite region size, such as Figure 28 As in the example in Figure 29, the scanned non-neighboring locations can be outside the allowed non-neighboring areas. In this case, different methods can be used to address the problem: In one approach, the scanning process can indicate that the scanned location does not have valid neighborhood information.
[0528] In another method, the scanning process can project or crop the location outside the range to another location within a permitted non-adjacent area. For example... Figure 30As shown in the example, there are two locations (i.e., two spots 3001) outside the permissible non-adjacent area. These two locations are projected / cropped to two other locations, which are located at the same vertical / horizontal coordinates but within the permissible non-adjacent area. The new location of the projection / cropping can be located on the boundary of the permissible non-adjacent area closest to the original location. If the permissible non-adjacent area outside the current CTU is set to values w and h with a minimum value (e.g., 4) set to the IBC CU, the new location of the projection / cropping can be interchangeably set on one boundary or the other because the buffer is so small that it can only store motion information from one CU, and in this case, there is no difference between cropping to one boundary side or the other.
[0529] In another approach, the allowed non-adjacent spatial regions can include three regions. For example... Figure 31 As shown, these three regions are the spatial regions located outside and adjacent to the current CTU: the upper left, left side, and upper region. The height of the upper left and upper regions within the allowed non-adjacent spatial regions is defined as h, while their width can depend on the image width. The width of the allowed non-adjacent spatial region on the left is defined as w, and its height is equal to the height of the current CTU. (See diagram for details.) Figure 31 As shown, if the scanned non-adjacent positions ( Figure 31 If one of the spots in 3001 exceeds the allowed spatial area, the new location of the projection / cropping can be defined in a different way (with...). Figure 31 (One of the corresponding spots in spot 3002 is shown in spot 3001 indicated by the dashed arrow). For the top-left non-adjacent region, the new projection / cropping position is always the pixel position adjacent to the top-left position of the current CTU. For example, if the top-left position of the current CTU is (ctu_x, ctu_y), then the new projection / cropping position is (ctu_x - 1, ctu_y - 1). For the aforementioned non-adjacent regions, the new projection / cropping position has the same horizontal coordinate, but the vertical coordinate becomes (ctu_y - 1). For the left-side non-adjacent region, the new projection / cropping position has the same vertical coordinate, but the horizontal coordinate becomes (ctu_x - 1).
[0530] When motion information of an IBC coded block is stored in memory, it can be stored at a granularity of the smallest IBC block size (e.g., a 4×4 block). When the current IBC coded block is a coding unit with a size larger than the smallest IBC block, motion information can be stored in different ways. In one or more embodiments, the motion information stored at each smallest IBC block (e.g., a 4×4 block) within the current block is simply a duplicate copy of the motion information for the current block.
[0531] Alternatively or additionally, the motion information of the IBC coded block can be stored at different granularities a × b (e.g., 8×8, 8×16, 16×8, or 16×16 granularities, etc.) instead of the minimum IBC block size (e.g., 4×4 granularity), where the granularity values of a and b can be configurable or determined at the encoder and subsequently signaled to the decoder. Without loss of generality, it uses an 8×8 granularity (e.g., a = b = 8) as an illustrative example. If we further assume the minimum IBC block size is 4×4, it indicates that each 8×8 block can store only a set of IBC motion information, representing a single IBC model, even though the four 4×4 sub-blocks within that 8×8 block can come from more than one IBC block, as shown in... Figure 32 In. Figure 32 In this example, four 4×4 sub-blocks A, B, C, and D form an 8×8 block / region, and only one IBC model information is stored. However, these four 4×4 sub-blocks come from four different IBC blocks, representing four IBC models and including four sets of IBC motion information. In this case, there may be different ways to derive and store a single set of IBC information: In one or more embodiments, one of a plurality of sets of available IBC motion information can be selected and stored. In one example, IBC motion information at a fixed or configurable location (e.g., the top-left smallest IBC block) is selected for motion storage. In another example, the average IBC motion information of multiple models can be calculated for motion storage.
[0532] IBC motion information at selected adjacent IBC blocks can be simplified / compressed before storage. In one embodiment, each saved BV can be compressed before storage to further reduce memory size. One example is using general techniques for data compression. For instance, it is proposed to save a composite value from an exponent and a mantissa to approximate each saved BV.
[0533] The methods described above can be applied to motion information storage in any combination. For example, defined restricted regions of non-adjacent neighboring blocks can be combined with the use of compressed IBC motion information.
[0534] Application Scope
[0535] The methods proposed in the "Candidate Scanning and Pruning," "Candidate Reordering," and "Motion Information Storage" sections primarily aim to derive spatially non-adjacent candidates for IBC AMVP or merged modes. When temporal candidates are also used for IBC AMVP or merged modes, they can also utilize non-adjacent neighboring blocks in co-located images. Unlike spatially non-adjacent candidates, which are mainly extracted from the left and top regions of the current block, temporally non-adjacent candidates can be extracted from the left, top, right, bottom, and co-located regions of the current block. Furthermore, the methods proposed in the "Candidate Scanning and Pruning," "Candidate Reordering," and "Motion Information Storage" sections can also be applied to temporally non-adjacent candidates for IBC AMVP or merged modes in a similar manner.
[0536] Sub-block-based IBC mode
[0537] According to one or more embodiments of this disclosure, the IBC mode is extended to the sub-block level, where sub-blocks within the current block can have their own BV for motion compensation. Different example methods can be used to achieve this goal.
[0538] In the first example method, such as Figure 33 As shown, the BV of a sub-block in the current block is obtained by reusing the BV of a sub-block within a co-op block in the co-op image. If the BV of a sub-block within the co-op block cannot be obtained, such as when the sub-block is intra-coded, the BV of the sub-block can be set to the BV of the co-op block. Besides co-op blocks in the co-op image, sub-block-level BVs can also be obtained using blocks at other locations, where temporally non-adjacent candidates can be referenced. Furthermore, the BV of a sub-block within a co-op block or other locations within the co-op image can be refined using template matching methods.
[0539] In the second example method, such as Figure 34 As shown, the BV of the left or top child block in the current block is obtained by refining the BV of the current block using a template matching method. The BV of the current block can be obtained using the regular IBC mode, TM IBC mode, or other modes. For the left child block in the current block, only the left template can be used to refine the BV of the child block. For the aforementioned child blocks in the current block, only the aforementioned templates can be used to refine the BV of the child block. For the top-left child block in the current block, both the left template and the top template can be used to refine the BV of the child block.
[0540] Multiple template patterns for TM IBC mode and TM regular mode
[0541] According to one or more embodiments of this disclosure, the template patterns for TM IBC mode and TM regular mode are expanded by importing more types of templates. This objective can be achieved using different methods.
[0542] In the first method, in addition to the currently used template modes of using both the top and left templates, other template modes, such as using only the left template or only the top template, can also be used in TM IBC mode or TM regular inter-frame mode. When more than one template mode is used in TM IBC mode or TM regular inter-frame mode, the final template mode used for the current block can be determined based on predefined criteria or flags sent to the bitstream.
[0543] IBC and intra-frame TMP derived from deblocking filter boundary strength
[0544] According to one or more embodiments of this disclosure, the deblocking filter processes blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode equally when obtaining boundary strength. Different methods can be used to achieve this objective.
[0545] In the first method, when the boundary strength of the deblocking filter is obtained, blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode are both considered as blocks encoded in intra-frame mode. The boundary strength criterion for blocks encoded in intra-frame mode can then be applied to both blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode. For example, if two adjacent blocks are both encoded in IBC mode or intra-frame TMP mode, or if one of the two adjacent blocks is encoded in IBC mode or intra-frame TMP mode, the boundary strength is set to a predetermined positive integer, such as 2.
[0546] In the second method, when obtaining the boundary strength of the deblocking filter, blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode are both considered as blocks encoded in inter-frame mode. The boundary strength criterion for blocks encoded in inter-frame mode can then be applied to blocks encoded in IBC mode and blocks encoded in intra-frame TMP mode. For example, if two adjacent blocks are both encoded in IBC mode, or both adjacent blocks are encoded in intra-frame TMP mode, or one of two adjacent blocks is encoded in IBC mode and the other in intra-frame TMP mode, and if the two block vectors of the two adjacent blocks are different (or the absolute difference between the horizontal or vertical components of the two block vectors is greater than a threshold (e.g., half a pixel)), the boundary strength is set to a predetermined positive integer, such as 1. In another example, if the difference between the two block vectors of two adjacent blocks is large, the boundary strength is set to a larger value.
[0547] In the third method, when a block encoded in IBC mode or a block encoded in intra-frame TMP mode is combined with an intra-frame tool, such as combining IBC with intra-frame prediction, combining GPM with IBC and intra-frame prediction, combining intra-frame TMP with intra-frame prediction, combining GPM with intra-frame TMP and intra-frame prediction, etc., the block is considered as a block encoded in intra-frame mode when the boundary strength of the deblocking filter is obtained.
[0548] In the fourth method, when a block encoded in IBC mode or a block encoded in intra-frame TMP mode is combined with inter-frame tools, such as combining IBC with inter-frame prediction, combining GPM with IBC and inter-frame prediction, combining intra-frame TMP with inter-frame prediction, combining GPM with intra-frame TMP and inter-frame prediction, etc., the block is considered as a block encoded in inter-frame mode when the boundary strength of the deblocking filter is obtained.
[0549] Figure 35 A computing environment (or computing device) 1610 coupled to a user interface 1650 is shown. The computing environment 1610 may be part of a data processing server. In some embodiments, the computing device 1610 may perform any of the various methods or processes (e.g., encoding / decoding methods or processes) described above, according to various examples of this disclosure. The computing environment 1610 includes a processor 1620, a memory 1630, and an input / output (I / O) interface 1640.
[0550] Processor 1620 typically controls the overall operation of computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 1620 may include one or more modules that facilitate interaction between processor 1620 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0551] Memory 1630 is configured to store various types of data to support the operation of computing environment 1610. Memory 1630 may include predefined software 1632. Examples of such data include instructions for any application or method operating on computing environment 1610, video datasets, image data, etc. Memory 1630 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0552] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 can be coupled to encoders and decoders.
[0553] Figure 40 This is a flowchart illustrating some examples of methods for video decoding according to this disclosure.
[0554] In step 4010, the processor 1620 may obtain a first block and a second block on the decoder side, wherein the first block is encoded in either Intra-Block Copy (IBC) mode or Intra-Template Match Prediction (TMP) mode, and the second block is encoded in either Intra-TMP mode or IBC mode.
[0555] In step 4020, the processor 1620 can obtain the boundary strength of the deblocking filter on the decoder side by applying a predefined criterion in either intra-frame mode or inter-frame mode.
[0556] In step 4030, the processor 1620 may apply a deblocking filter to the first block and the second block on the decoder side and based on the boundary strength.
[0557] In some examples, the predefined standard is based on an intra-frame predefined standard. In one or more examples, in step 4020, the processor 1620 can obtain the boundary strength of the deblocking filter on the decoder side by applying an intra-frame predefined standard in the intra-frame mode.
[0558] In one or more examples, the first block and the second block are adjacent blocks of the current block, and in order to obtain the boundary strength of the deblocking filter by applying a predefined intra-frame criterion in intra-frame mode, the processor 1620 may set the boundary strength to a fixed value. In one or more examples, the fixed value is a first predetermined positive integer.
[0559] In some examples, the predefined standard is based on an inter-frame predefined standard. In one or more examples, in step 4020, the processor 1620 can obtain the boundary strength of the deblocking filter on the decoder side by applying an inter-frame predefined standard in the inter-frame mode.
[0560] In one or more examples, the first block and the second block are adjacent blocks of the current block, the first block's first block vector includes a first horizontal component and a first vertical component, and the second block's second block vector includes a second horizontal component and a second vertical component.
[0561] In one or more examples, to obtain the boundary strength of the deblocking filter by applying a predefined inter-frame criterion in inter-frame mode, the processor 1620 may perform one of the following actions on the decoder side: setting the boundary strength to a fixed value in response to determining that the first block vector is different from the second block vector; or setting the boundary strength to a fixed value in response to determining that the first absolute difference between the first horizontal component and the second horizontal component is greater than a first threshold; or setting the boundary strength to a fixed value in response to determining that the second absolute difference between the first vertical component and the second vertical component is greater than a second threshold. In one or more examples, the fixed value is a second predetermined positive integer.
[0562] In one or more examples, the first block and the second block are adjacent blocks of the current block, the first block's first block vector includes a first horizontal component and a first vertical component, and the second block's second block vector includes a second horizontal component and a second vertical component. In one or more examples, in order to obtain the boundary strength of the deblocking filter by applying a predefined inter-frame criterion in the inter-frame mode, the processor 1620 may perform one of the following actions on the decoder side: setting the boundary strength to a first value positively correlated with a first absolute difference between the first horizontal component and the second horizontal component; or setting the boundary strength to a second value positively correlated with a second absolute difference between the first vertical component and the second vertical component.
[0563] In some examples, processor 1620 may further combine the coding mode of the first or second block with the intra-frame mode on the decoder side, and obtain intra-frame prediction of one of the first or second blocks based on the boundary strength of the deblocking filter. In one or more examples, in order to combine the coding mode of the first or second block with the intra-frame mode, processor 1620 may perform one of the following actions as described in this disclosure: combining IBC with intra-frame prediction; combining Geometric Partitioning Mode (GPM) with IBC and intra-frame prediction; combining intra-frame TMP with intra-frame prediction; or combining GPM with intra-frame TMP and intra-frame prediction.
[0564] In some examples, processor 1620 may further combine the coding mode of the first or second block with the inter-frame mode on the decoder side, and obtain inter-frame prediction of one of the first or second blocks based on the boundary strength of the deblocking filter. In one or more examples, in order to combine the coding mode of the first or second block with the inter-frame mode, processor 1620 may perform one of the following actions as described in this disclosure: combining IBC with inter-frame prediction; combining Geometric Partition Mode (GPM) with IBC and inter-frame prediction; combining intra-frame TMP with inter-frame prediction; or combining GPM with intra-frame TMP and inter-frame prediction.
[0565] Figure 41It shows the relationship with, for example Figure 40 The flowchart shown corresponds to the video encoding method for the video decoding method.
[0566] In step 4110, the processor 1620 may obtain a first block and a second block on the encoder side, wherein the first block is encoded in either Intra-Block Copy (IBC) mode or Intra-Template Match Prediction (TMP) mode, and the second block is encoded in either Intra-TMP mode or IBC mode.
[0567] In step 4120, the processor 1620 can obtain the boundary strength of the deblocking filter on the encoder side by applying a predefined criterion in either intra-frame mode or inter-frame mode.
[0568] In step 4130, the processor 1620 may apply a deblocking filter to the first and second blocks on the encoder side and based on the boundary strength.
[0569] In step 4140, processor 1620 can generate a bitstream on the encoder side based on the deblocking filter applied to the first and second blocks in step 4130.
[0570] In some examples, the predefined standard is based on an intra-frame predefined standard. In one or more examples, in step 4120, the processor 1620 can obtain the boundary strength of the deblocking filter on the encoder side by applying an intra-frame predefined standard in the intra-frame mode.
[0571] In one or more examples, the first block and the second block are adjacent blocks of the current block, and in order to obtain the boundary strength of the deblocking filter by applying a predefined intra-frame criterion in intra-frame mode, the processor 1620 may set the boundary strength to a fixed value. In one or more examples, the fixed value is a first predetermined positive integer.
[0572] In some examples, the predefined criteria are based on inter-frame predefined criteria. In one or more examples, in step 4120, the processor 1620 can obtain the boundary strength of the deblocking filter on the encoder side by applying an inter-frame predefined criterion in the inter-frame mode.
[0573] In one or more examples, the first block and the second block are adjacent blocks of the current block, the first block's first block vector includes a first horizontal component and a first vertical component, and the second block's second block vector includes a second horizontal component and a second vertical component.
[0574] In one or more examples, to obtain the boundary strength of the deblocking filter by applying a predefined inter-frame criterion in inter-frame mode, the processor 1620 may perform one of the following actions on the encoder side: setting the boundary strength to a fixed value in response to determining that the first block vector is different from the second block vector; or setting the boundary strength to a fixed value in response to determining that the first absolute difference between the first horizontal component and the second horizontal component is greater than a first threshold; or setting the boundary strength to a fixed value in response to determining that the second absolute difference between the first vertical component and the second vertical component is greater than a second threshold. In one or more examples, the fixed value is a second predetermined positive integer.
[0575] In one or more examples, the first block and the second block are adjacent blocks of the current block, the first block's first block vector includes a first horizontal component and a first vertical component, and the second block's second block vector includes a second horizontal component and a second vertical component. In one or more examples, in order to obtain the boundary strength of the deblocking filter by applying a predefined inter-frame criterion in the inter-frame mode, the processor 1620 may perform one of the following actions on the encoder side: setting the boundary strength to a first value positively correlated with a first absolute difference between the first horizontal component and the second horizontal component; or setting the boundary strength to a second value positively correlated with a second absolute difference between the first vertical component and the second vertical component.
[0576] In some examples, processor 1620 may further combine the coding mode of the first or second block with the intra-frame mode on the encoder side, and obtain intra-frame prediction of one of the first or second blocks based on the boundary strength of the deblocking filter. In one or more examples, in order to combine the coding mode of the first or second block with the intra-frame mode, processor 1620 may perform one of the following actions as described in this disclosure: combining IBC with intra-frame prediction; combining Geometric Partitioning Mode (GPM) with IBC and intra-frame prediction; combining intra-frame TMP with intra-frame prediction; or combining GPM with intra-frame TMP and intra-frame prediction.
[0577] In some examples, processor 1620 may further combine the coding mode of the first or second block with the inter-frame mode on the encoder side, and obtain inter-frame prediction of one of the first or second blocks based on the boundary strength of the deblocking filter. In one or more examples, in order to combine the coding mode of the first or second block with the inter-frame mode, processor 1620 may perform one of the following actions as described in this disclosure: combining IBC with inter-frame prediction; combining Geometric Partition Mode (GPM) with IBC and inter-frame prediction; combining intra-frame TMP with inter-frame prediction; or combining GPM with intra-frame TMP and inter-frame prediction.
[0578] In some examples, an apparatus for video encoding is provided. The apparatus includes: a processor 1620; and a memory 1640 configured to store instructions executable by the processor; wherein the processor, when executing the instructions, is configured to perform actions such as... Figures 40-41 Any of the methods shown.
[0579] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a plurality of programs in memory 1630 and / or a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by processor 1620 in computing environment 1610 to perform the above-described methods. In one example, the plurality of programs can be executed by processor 1620 in computing environment 1610 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 1610 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.
[0580] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.
[0581] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0582] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 1630, the plurality of programs being executable by a processor 1620 in a computing environment 1610 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0583] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0584] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.
[0585] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.
[0586] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0587] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0588] The examples chosen and described are intended to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A method for video decoding, comprising: The decoder obtains a first block and a second block, wherein the first block is encoded in either an intra-block copy (IBC) mode or an intra-template match prediction (TMP) mode, and the second block is encoded in either the intra-TMP mode or the IBC mode. The decoder obtains the boundary strength of the deblocking filter by applying a predefined criterion in either intra-frame mode or inter-frame mode; and The deblocking filter is applied to the first block and the second block by the decoder and based on the boundary strength.
2. The method according to claim 1, wherein, The predefined standard is based on an intra-frame predefined standard; as well as The boundary intensity of the deblocking filter obtained by the decoder includes: The decoder obtains the boundary strength of the deblocking filter by applying the predefined intra-frame standard in the intra-frame mode.
3. The method according to claim 2, wherein, The first block and the second block are adjacent blocks of the current block; as well as The step of obtaining the boundary intensity of the deblocking filter by the decoder through applying the predefined intra-frame standard in the intra-frame mode includes: Set the boundary strength to a first predetermined positive integer.
4. The method according to claim 1, wherein, The predefined standard is based on inter-frame predefined standards; as well as The boundary intensity of the deblocking filter obtained by the decoder includes: The decoder obtains the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode.
5. The method according to claim 4, wherein, The first block and the second block are adjacent blocks of the current block; and The step of the decoder obtaining the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode includes one of the following actions: In response to determining that the first block vector of the first block is different from the second block vector of the second block, the boundary strength is set to a second predetermined positive integer; In response to determining that the first absolute difference between the first horizontal component of the first block vector and the second horizontal component of the second block vector is greater than a first threshold, the boundary strength is set to the second predetermined positive integer; or In response to determining that the second absolute difference between the first vertical component of the first block vector and the second vertical component of the second block vector is greater than a second threshold, the boundary strength is set to the second predetermined positive integer.
6. The method according to claim 4, wherein, The first block and the second block are adjacent blocks of the current block; and The step of the decoder obtaining the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode includes one of the following actions: The boundary strength is set to a first value that is positively correlated with the first absolute difference between the first horizontal component of the first block vector of the first block and the second horizontal component of the second block vector of the second block; or The boundary strength is set to a second value that is positively correlated with the second absolute difference between the first vertical component of the first block vector and the second vertical component of the second block vector.
7. The method according to claim 2, further comprising: The decoder combines the encoding mode of the first block or the second block with the intra-frame mode; as well as The decoder obtains an intra-frame prediction of either the first block or the second block based on the boundary strength of the deblocking filter.
8. The method according to claim 7, wherein, Combining the encoding mode of the first block or the second block with the intra-frame mode includes one of the following actions: Combine the IBC with the intra-frame prediction; Combine the geometric partitioning mode (GPM) with the IBC and the intra-frame prediction; Combine the intra-frame TMP with the intra-frame prediction; or The GPM is combined with the intra-frame TMP and the intra-frame prediction.
9. The method according to claim 4, further comprising: The decoder combines the encoding mode of the first block or the second block with the inter-frame mode; as well as The decoder obtains inter-frame prediction of either the first block or the second block based on the boundary strength of the deblocking filter.
10. The method according to claim 9, wherein, Combining the encoding mode of the first block or the second block with the inter-frame mode includes one of the following actions: Combine the IBC with the inter-frame prediction; Combine the geometric partitioning mode (GPM) with the IBC and the inter-frame prediction; Combine the intra-frame TMP with the inter-frame prediction; or The GPM is combined with the intra-frame TMP and the inter-frame prediction.
11. A method for video encoding, comprising: The encoder obtains a first block and a second block, wherein the first block is encoded in either an intra-block copy (IBC) mode or an intra-template match prediction (TMP) mode, and the second block is encoded in either the intra-TMP mode or the IBC mode. The encoder obtains the boundary strength of the deblocking filter by applying a predefined criterion in either intra-frame mode or inter-frame mode. The deblocking filter is applied to the first block and the second block by the encoder and based on the boundary strength; and The encoder generates a bitstream by applying the deblocking filter to the first block and the second block.
12. The method according to claim 11, wherein, The predefined standard is based on an intra-frame predefined standard; as well as The boundary strength of the deblocking filter obtained by the encoder includes: The encoder obtains the boundary strength of the deblocking filter by applying the predefined intra-frame standard in the intra-frame mode.
13. The method according to claim 12, wherein, The first block and the second block are adjacent blocks of the current block; as well as The step of obtaining the boundary intensity of the deblocking filter by the encoder through applying the predefined intra-frame standard in the intra-frame mode includes: Set the boundary strength to a first predetermined positive integer.
14. The method according to claim 11, wherein, The predefined standard is based on inter-frame predefined standards; as well as The boundary strength of the deblocking filter obtained by the encoder includes: The encoder obtains the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode.
15. The method according to claim 14, wherein, The first block and the second block are adjacent blocks of the current block; and The process by which the encoder obtains the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode includes one of the following actions: In response to determining that the first block vector of the first block is different from the second block vector of the second block, the boundary strength is set to a second predetermined positive integer; In response to determining that the first absolute difference between the first horizontal component of the first block vector and the second horizontal component of the second block vector is greater than a first threshold, the boundary strength is set to the second predetermined positive integer; or In response to determining that the second absolute difference between the first vertical component of the first block vector and the second vertical component of the second block vector is greater than a second threshold, the boundary strength is set to the second predetermined positive integer.
16. The method of claim 14, wherein, The first block and the second block are adjacent blocks of the current block; and The process by which the encoder obtains the boundary strength of the deblocking filter by applying the predefined inter-frame standard in the inter-frame mode includes one of the following actions: The boundary strength is set to a first value that is positively correlated with the first absolute difference between the first horizontal component of the first block vector of the first block and the second horizontal component of the second block vector of the second block; or The boundary strength is set to a second value that is positively correlated with the second absolute difference between the first vertical component of the first block vector and the second vertical component of the second block vector.
17. The method of claim 12, further comprising: The encoder combines the encoding mode of the first block or the second block with the intra-frame mode; as well as The encoder obtains an intra-frame prediction of either the first block or the second block based on the boundary strength of the deblocking filter.
18. The method according to claim 17, wherein, Combining the encoding mode of the first block or the second block with the intra-frame mode includes one of the following actions: Combine the IBC with the intra-frame prediction; Combine the geometric partitioning mode (GPM) with the IBC and the intra-frame prediction; Combine the intra-frame TMP with the intra-frame prediction; or The GPM is combined with the intra-frame TMP and the intra-frame prediction.
19. The method of claim 14, further comprising: The encoder combines the encoding mode of the first block or the second block with the inter-frame mode; as well as The encoder obtains inter-frame prediction of either the first block or the second block based on the boundary strength of the deblocking filter.
20. The method according to claim 19, wherein, Combining the encoding mode of the first block or the second block with the inter-frame mode includes one of the following actions: Combine the IBC with the inter-frame prediction; Combine the geometric partitioning mode (GPM) with the IBC and the inter-frame prediction; Combine the intra-frame TMP with the inter-frame prediction; or The GPM is combined with the intra-frame TMP and the inter-frame prediction.
21. An apparatus for video decoding, comprising: One or more processors; as well as One or more processors; as well as A memory, coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method according to any one of claims 1 to 10 when executing the instructions.
22. An apparatus for video encoding, comprising: One or more processors; as well as A memory, coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method according to any one of claims 11 to 20 when executing the instructions.
23. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1 to 10.
24. A non-transitory computer-readable storage medium for storing a bit stream generated by the method of any one of claims 11 to 20.