Method and apparatus for filtered intra block copy
By determining the template region sample values of the reference block and the current block during intra-block copying, a set of filter coefficients is obtained, which solves the problem of low efficiency in intra-block copying in the prior art and achieves more efficient video encoding and decoding and bit rate optimization.
Patent Information
- Application Number
- CN202480024155.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-19
- Filing Date
- 2024-04-18
- Publication Date
- 2025-11-04
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in intra-frame block copying (FIBC), making it difficult to effectively utilize redundant information in video data for efficient compression.
By determining the sample values of the template regions of the reference block and the current block, a set of filter coefficients corresponding to the filter shape is obtained. These coefficients and shapes are used to derive the predicted sample values of the current block, and the current block is reconstructed based on the predicted sample values, thus improving the encoding and decoding process of intra-frame block copying.
It improves the efficiency of video encoding and decoding, reduces bitrate usage, and maintains or improves video quality.
Smart Images

Figure CN120898424A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application is based on and claims priority to provisional application No. 63497157, filed on April 19, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to video encoding / decoding and compression. More specifically, this application relates to methods and apparatus for improving the encoding / decoding efficiency of filtered intra-block copy (FIBC). Background Technology
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, and video streaming devices. Electronic devices send and receive digital video data via communication networks or otherwise transmit it, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video codecs are used to compress video data according to one or more video codec standards before it is transmitted or stored. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data into a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0004] Embodiments of this disclosure provide methods and apparatus for improving the encoding and decoding efficiency of image / video blocks using FIBC technology.
[0005] According to one aspect of this disclosure, a method for video decoding is provided, comprising: determining a reference block in a video frame from a bitstream for predicting a current block in the video frame; obtaining a set of filter coefficients corresponding to a filter shape based at least on sample values from both a template region associated with the reference block and a template region associated with the current block; deriving predicted sample values of the current block using the set of filter coefficients and the filter shape, based at least on a plurality of predicted sample values of the current block; and reconstructing the current block based on the predicted sample values.
[0006] According to one aspect of this disclosure, a method for video coding is provided, comprising: determining a reference block in a video frame for predicting a current block in the video frame; obtaining a set of filter coefficients corresponding to a filter shape based at least on sample values from both a template region associated with the reference block and a template region associated with the current block; deriving predicted sample values of the current block using the set of filter coefficients and the filter shape based at least on a plurality of predicted sample values of the current block; and generating a bitstream based on the predicted sample values.
[0007] According to one aspect of this disclosure, a method for video decoding is provided, comprising: determining a reference block in a video frame from a bitstream used for predicting a current block in a video frame; obtaining a set of filter coefficients corresponding to a filter shape based at least on sample values from both a template region associated with the reference block and a template region associated with the current block, wherein the filter shape is a rectangle having a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted; deriving predicted sample values for the current block based on at least one of a plurality of predicted sample values for the current block and a plurality of corresponding reconstructed sample values associated with the reference block using the set of filter coefficients and the filter shape; and reconstructing the current block based on the predicted sample values.
[0008] According to one aspect of this disclosure, a method for video coding is provided, comprising: determining a reference block in a video frame for predicting a current block in the video frame; obtaining a set of filter coefficients corresponding to a filter shape, based at least on sample values from both a template region associated with the reference block and a template region associated with the current block, wherein the set of filter coefficients comprises a rectangle having a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted; deriving predicted sample values for the current block based on at least one of a plurality of predicted sample values for the current block and a plurality of corresponding reconstructed sample values associated with the reference block, using the set of filter coefficients and the filter shape; and generating a bitstream based on the predicted sample values.
[0009] According to one aspect of this disclosure, a method for video decoding is provided, comprising: obtaining a merge candidate list for intra-block copy (IBC) prediction of a current block, the merge candidate list including a plurality of candidates already encoded using IBC; reordering the plurality of candidates in the merge candidate list based on a template matching score for each of the plurality of candidates, the template matching score being obtained based on the difference between a sample value of a template of a candidate and a corresponding reference sample value of a reference template of a reference block of the candidate, wherein, in response to determining that a candidate is encoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC; and reconstructing the current block based on the reordered merge candidate list.
[0010] According to one aspect of this disclosure, a method for video coding is provided, comprising: obtaining a merge candidate list for intra-block copy (IBC) prediction of a current block, the merge candidate list including a plurality of candidates already encoded using IBC; reordering the plurality of candidates in the merge candidate list based on a template matching score for each of the plurality of candidates, the template matching score being obtained based on the difference between a sample value of a candidate's template and a corresponding reference sample value of a reference template of a reference block of the candidate, wherein, in response to determining that a candidate is encoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC; and encoding the current block based on the reordered merge candidate list to generate a bitstream.
[0011] According to one aspect of this disclosure, an apparatus is provided, comprising: one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform operations of the methods of this disclosure.
[0012] According to one aspect of this disclosure, a computer program product is provided that stores computer-executable instructions, which, when executed, cause the one or more processors to perform the operations of the methods of this disclosure.
[0013] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores instructions, which, when executed by a computing device having one or more processors, cause the one or more processors to perform a decoding method of this disclosure and store a bitstream to be decoded by the decoding method of this disclosure, or to perform an encoding method of this disclosure and store a bitstream generated by the encoding method of this disclosure.
[0014] According to one aspect of this disclosure, a computer-readable medium for storing bitstreams is provided, wherein the bitstreams are decoded by performing the operations of the methods of this disclosure, or the bitstreams are obtained by performing the operations of the methods of this disclosure.
[0015] According to one aspect of this disclosure, a method for receiving a bitstream to be decoded by the decoding method of this disclosure is provided.
[0016] According to one aspect of this disclosure, a method for transmitting a bit stream generated by the encoding method of this disclosure is provided.
[0017] It should be understood that the foregoing general description and the following detailed description are merely examples and are not intended to limit this disclosure. Attached Figure Description
[0018] Examples consistent with this disclosure are illustrated in conjunction with the accompanying drawings, which are incorporated in and form part of this specification, and are used together with the specification to explain the principles of this disclosure.
[0019] Figure 1 This is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.
[0020] Figure 2 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.
[0021] Figure 3 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.
[0022] Figures 4A to 4E This is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.
[0023] Figure 5 A schematic diagram of the spatial candidate locations is shown.
[0024] Figure 6 A schematic diagram of candidate pairs considering redundancy checks for spatial candidates is shown.
[0025] Figure 7 A schematic diagram of scaling the motion vector of the time candidate is shown.
[0026] Figure 8 A schematic diagram of candidate positions for time-based candidates is shown.
[0027] Figure 9 A schematic diagram of a merging pattern with motion vector difference (MMVD) search points is shown.
[0028] Figure 10 The selection of unidirectional predicted motion vectors in the Geometric Partitioning Mode (GPM) is shown.
[0029] Figure 11 The top and left adjacent blocks used in the CIIP weight derivation are shown.
[0030] Figure 12 The current CTU processing order and its available reference points in the current CTU and the left CTU are shown.
[0031] Figure 13 The filling candidates for replacing the zero vector in the IBC list are shown.
[0032] Figure 14 The reference region of IBC is shown when CTU(m,n) is encoded and decoded.
[0033] Figure 15 The IBC reference area for the content captured by the camera is shown.
[0034] Figures 16A to 16B The method for dividing the angle pattern is shown.
[0035] Figure 17A , Figure 17B and Figure 17C The available IPM candidates are shown, and Figure 17D An example of GPM with intra-frame prediction and intra-frame prediction is shown.
[0036] Figure 18 The edges on the template are shown.
[0037] Figure 19 The intra-frame template matching search area used is shown.
[0038] Figure 20 A template for OBMC based on template matching is shown.
[0039] Figure 21 The template in the reference image and the reference sample points of the template are shown.
[0040] Figure 22 This shows a template and reference points for a block that has sub-block motion, using motion information of the current block's sub-blocks.
[0041] Figure 23 This shows the luminance block used to derive the direct block vector.
[0042] Figure 24 The method for dividing intra-coded blocks and the corresponding weights for angular and planar modes are shown.
[0043] Figure 25A schematic diagram of the filter shape of the reference block and the training region is shown.
[0044] Figure 26 This shows examples of predictions for different positions within the current block.
[0045] Figure 27 A schematic diagram is shown corresponding to the spatial terms of adjacent brightness samples.
[0046] Figure 28 A schematic diagram showing examples of different shapes / numbers of filter taps.
[0047] Figure 29 A schematic diagram showing examples of different shapes / numbers of filter taps.
[0048] Figure 30 A schematic diagram showing examples of different shapes / numbers of filter taps.
[0049] Figure 31 A schematic diagram showing the possible locations of the candidate regions is provided.
[0050] Figure 32 A schematic diagram of the possible candidate locations is shown.
[0051] Figure 33 The workflow of a method for decoding video data according to one or more aspects of this disclosure is shown.
[0052] Figure 34 The workflow of a method for encoding video data according to one or more aspects of this disclosure is shown.
[0053] Figure 35 The workflow of a method for decoding video data according to one or more aspects of this disclosure is shown.
[0054] Figure 36 The workflow of a method for encoding video data according to one or more aspects of this disclosure is shown.
[0055] Figure 37 A schematic diagram of the filter shape of the reference block and the training region is shown.
[0056] Figure 38 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0057] Figure 39 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0058] Figure 40 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0059] Figure 41 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0060] Figure 42 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0061] Figure 43 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0062] Figure 44 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0063] Figure 45 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0064] Figure 46 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0065] Figure 47 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0066] Figure 48 The workflow of a method for video decoding according to one or more aspects of this disclosure is shown.
[0067] Figure 49 The workflow of a method for video encoding according to one or more aspects of this disclosure is shown.
[0068] Figure 50 This is a schematic diagram illustrating a computing environment coupled to a user interface according to some embodiments of the present disclosure. Detailed Implementation
[0069] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting details are set forth to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0070] It should be noted that the terms "first," "second," etc., used in the description, claims, and drawings of this disclosure are used to distinguish objects and are not intended to describe any particular order or sequence. It should be understood that data used in this manner can be interchanged under appropriate conditions, such that embodiments of this disclosure described herein can be implemented in orders other than those shown in the drawings or described in this disclosure.
[0071] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data for later decoding by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.
[0072] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to communication standards such as wireless communication protocols and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 12 to target device 14.
[0073] In some other implementations, encoded video data may be sent from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by target device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital versatile discs (DVDs), compressed optical disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 may access the stored video data from storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing the encoded video data stored on the file server, including wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or combinations thereof. Transmission of the encoded video data from storage device 32 can be streaming, downloading, or a combination of both.
[0074] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources such as or combinations of such sources: a video capture device (e.g., a camera), a video archive including previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or combinations of these sources. As an example, if video source 18 is a camera in a security surveillance system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and can be applied to wireless and / or wired applications.
[0075] The captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be sent directly to target device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may also include a modem and / or transmitter.
[0076] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. Encoded video data transmitted on link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0077] In some implementations, target device 14 may include display device 34, which may be an integrated display device as well as an external display device configured to communicate with target device 14. Display device 34 displays decoded video data to the user and may include any of a variety of display devices, such as liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.
[0078] Video encoder 20 and video decoder 30 may operate according to proprietary or industry standards (such as VVC, HEVC, MPEG-4, Part 10, AVC) or extensions to such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally contemplated that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.
[0079] The video encoder 20 and video decoder 30 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, either of which may be integrated into the respective device as part of a combined encoder / decoder (CODEC).
[0080] In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or as referred to below) Figure 2 The components described include those in the video encoder 20 and output interface 22) and / or at least a portion of the components of the target device 14 (e.g., input interface 28, video decoder 30, or as referred to below). Figure 3 The components described herein, including those in the video decoder 30 and the display device 34, can operate within a cloud computing service network that provides software, platforms, and / or infrastructure, such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network may be provided in one or more client devices, and the one or more client devices may communicate with server computers in the cloud computing service network via wireless communication networks (e.g., cellular communication networks, short-range wireless communication networks, or Global Navigation Satellite System (GNSS) communication networks) or wired communication networks (e.g., local area network (LAN) communication networks or power line communication (PLC) networks). In one embodiment, at least a portion of the operations described herein may be implemented as a cloud-based service provided by one or more server computers, which are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” may be used interchangeably where appropriate. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service networks. Rather, this disclosure may be implemented in any other type of computing environment currently known or to be developed in the future.
[0081] Figure 2 This is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and decoding and inter-frame predictive coding and decoding on video blocks within a video frame. Intra-frame predictive coding and decoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame predictive coding and decoding relies on temporal prediction to reduce or remove temporal redundancy in video data within neighboring video frames or pictures of a video sequence. It should be noted that the term "frame" can be used in the field of video coding and decoding as a synonym for the terms "image" or "picture".
[0082] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (IBC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that this application is not limited to the embodiments described herein regarding CCSAO technology. Rather, this application can be applied to situations where an offset is selected for the luminance, Cb, and Cr chrominance components based on any other component among the luminance, Cb, and Cr chrominance components, to modify any of the components based on the selected offset. The first component mentioned herein can be any one of the luminance, Cb, and Cr chrominance components; the second component mentioned herein can be any other one of the luminance, Cb, and Cr chrominance components; and the third component mentioned herein can be the remaining components among the luminance, Cb, and Cr chrominance components. In some examples, the loop filter can be omitted, and the decoded video block can be directly provided to the DPB 64 by the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or it can be distributed among one or more of the described fixed or programmable hardware units.
[0083] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.
[0084] like Figure 2As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of sample points with sample values. Sample points in the array may also be referred to as pixels or image elements (pel). The number of sample points in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of sample points with sample values, but its dimension is smaller than that of the video frame. The number of sample points in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. Video blocks can be further divided into one or more block partitions or sub-blocks (which can then re-form blocks) by iteratively using, for example, QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the term "block" or "video block" as used herein can be a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring, for example, to HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.
[0085] Prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of several intra-frame or inter-frame predictive coding modes. Prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coded block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coded block for subsequent use as part of a reference frame. Prediction processing unit 41 also provides at least one of the syntax elements (e.g., motion vectors, intra-frame or inter-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.
[0086] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 can perform intra-predictive coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding / decoding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.
[0087] In some implementations, motion estimation unit 42 determines the inter-frame prediction mode of the current video frame by generating motion vectors based on a predetermined mode within the video frame sequence. These motion vectors indicate the displacement of a video block within the current video frame relative to a predicted block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or image relative to a predicted block within a reference frame and to the current block being encoded / decoded within the current frame. The predetermined mode may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors for intra-frame BC encoding, such as block vectors, in a manner similar to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 41.
[0088] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0089] The motion estimation unit 42 calculates motion vector information for video blocks in inter-frame predictive coding frames by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1). Each of the first and second reference frame lists identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the determined motion vector information to the motion compensation unit 44, and then to the entropy coding unit 56.
[0090] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vector information determined by motion estimation unit 42. Upon receiving motion vector information for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel differences forming the residual video block may include luminance component differences or chrominance component differences, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.
[0091] In some implementations, the intra-BC unit 48 may generate vectors and extract prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 may determine the intra-prediction mode to be used to encode the current block. In some examples, the intra-BC unit 48 may, for example, encode the current block using various intra-prediction modes during individual encoding passes and test its performance through rate-distortion analysis. Next, the intra-BC unit 48 may select an appropriate intra-prediction mode from among the various tested intra-prediction modes for use and generate an intra-prediction mode indicator accordingly. For example, the intra-BC unit 48 may use rate-distortion analysis for various tested intra-prediction modes to calculate rate-distortion values and select the intra-prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra-prediction mode to be used. Rate-distortion analysis typically determines the amount of distortion (or error) between a coded block and the original uncoded block used to generate the coded block, as well as the bit rate (i.e., number of bits) used to generate the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate of various coded blocks to determine which intra-frame prediction mode exhibits the best rate-distortion value for that block.
[0092] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction as described herein. In any case, for intra-frame block copying, the predicted block may be a block that is considered to closely match the block to be encoded / decoded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and the identification of the predicted block may include the calculation of values for sub-integer pixel positions.
[0093] Regardless of whether the predicted block comes from the same frame predicted intra-frame or a different frame predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel value of the predicted block from the pixel value of the current video block being encoded, thereby forming a pixel difference. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.
[0094] Intra-prediction processing unit 46 can perform intra-prediction on the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, or intra-block copy prediction performed by intra-BC unit 48, as described above. Specifically, intra-prediction processing unit 46 can determine the intra-prediction mode to be used to encode the current block. To this end, intra-prediction processing unit 46 can encode the current block using various intra-prediction modes, for example, during individual encoding passes, and intra-prediction processing unit 46 (or, in some examples, mode selection unit) can select an appropriate intra-prediction mode to be used from the tested intra-prediction modes. Intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to entropy coding unit 56. Entropy coding unit 56 can encode information indicating the selected intra-prediction mode in the bitstream.
[0095] After prediction processing unit 41 determines the prediction block of the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.
[0096] The transform processing unit 52 can send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can then perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0097] After quantization, the entropy coding unit 56 entropies the quantized transform coefficients into the video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be used as follows: Figure 1 As shown, it is sent to the video decoder 30, or as... Figure 1 The data shown is archived in storage device 32 for later transmission to video decoder 30 or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements of the current video frame being encoded.
[0098] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate a reference block for prediction of other video blocks. As mentioned above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of a frame stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values for motion estimation.
[0099] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block can then be used by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter-frame prediction of another video block in subsequent video frames.
[0100] Figure 3 This is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform generally the same functions as described above. Figure 2 The decoding process is the inverse of the encoding process described in the video encoder 20. For example, the motion compensation unit 82 can generate prediction data based on motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on intra-frame prediction mode indicators received from the entropy decoding unit 80.
[0101] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be divided among one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application alone or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).
[0102] Video data memory 79 may store video data, such as encoded video bitstreams, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or through access to physical data storage media (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) storing encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for decoding video data by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding / decoding modes). Video data memory 79 and DPB 92 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, Figure 3 The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, those skilled in the art will understand that the video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30, or off-chip relative to those components.
[0103] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or video block level. Entropy decoding unit 80 of video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.
[0104] When a video frame is encoded as an intra-predictive coded (I) frame or for intra-coded prediction blocks in other types of frames, the intra-prediction unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-prediction mode transmitted by the signal and reference data from the previous decoded block of the current frame.
[0105] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vector information and other syntax elements received from the entropy decoding unit 80. Each prediction block can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.
[0106] In some examples, when a video block is encoded according to the intra-BC mode described herein, the intra-BC unit 85 generates a predicted block for the current video block based on block vector information and other syntax elements received from the entropy decoding unit 80. The predicted block can be located within a reconstructed region of the same image as the current video block, defined by the video encoder 20.
[0107] Motion compensation unit 82 and / or intra-frame BC unit 85 determine prediction information for video blocks in the current video frame by parsing vector information and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode for encoding video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the frame's reference frame list, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.
[0108] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.
[0109] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.
[0110] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0111] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores a reference frame for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later display on a display device (e.g., Figure 1 Presented on the display device 34).
[0112] In typical video encoding and decoding processes, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of chrominance samples (Cb). SCr is a two-dimensional array of chrominance samples (Cr). In other cases, a frame may be monochromatic and therefore consist of only a two-dimensional array of luma samples.
[0113] like Figure 4AAs shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and / or other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block may be an N×N sample block.
[0114] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or combinations thereof, and divide the CTU into smaller CUs. Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image may include, for example, two corresponding coding blocks (CB) for luma samples and chroma samples, as well as syntax elements for encoding the samples of the coding blocks. In monochrome images or images with three separate color planes, a CU may include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that... Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree partitioning / ternary tree partitioning / binary tree partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to a binary and / or ternary tree structure. Figure 4E As shown, a CB with width W and height H has five possible segmentation types: quadrilateral segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal ternary segmentation, and vertical ternary segmentation.
[0115] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more M×N PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. A PU of the CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.
[0116] Video encoder 20 can use intra-frame prediction or inter-frame prediction to generate prediction blocks for the PU. If video encoder 20 uses intra-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of one or more frames different from the frame associated with the PU.
[0117] After the video encoder 20 generates one or more PU predictive luminance blocks, Cb blocks, and Cr blocks of the CU, the video encoder 20 can generate luminance residual blocks of the CU by subtracting the predictive luminance blocks of the CU from its original luminance decomposition blocks, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in a predictive luminance block of the CU and a corresponding sample in the original luminance decomposition block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks of the CU, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in a Cb block of the CU's predictive Cb block and a corresponding sample in the original Cb decomposition block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in a Cr block of the CU's predictive Cr block and a corresponding sample in the original Cr encoding block of the CU.
[0118] In addition, such as Figure 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.
[0119] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block of the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.
[0120] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 may quantize the coefficient blocks. Quantization generally refers to the process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream comprising bit sequences forming representations of encoded frames and associated data, which may be stored in storage device 32 or transmitted to target device 14.
[0121] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding samples of the prediction blocks of the PU of the current CU to the corresponding samples of the transform blocks of the TU of the current CU. After reconstructing the coded blocks of each CU of the frame, the video decoder 30 can reconstruct the frame.
[0122] As mentioned above, video coding primarily uses two modes for video compression: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block from a reference video block.
[0123] However, with the continuous improvement of video data capture technologies for preserving details in video data and the increasing finer video tile sizes, the amount of data required to represent the motion vector of the current frame has also increased significantly. One way to overcome this challenge is to take advantage of the fact that not only do groups of adjacent CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these adjacent CUs are also similar. Therefore, it is possible to use the motion information of spatially adjacent CUs and / or temporally co-located CUs as an approximation of the motion information (e.g., motion vector) of the current CU by exploring the spatial and temporal correlations of the current CU (also known as the "motion vector predictor (MVP)" of the current CU).
[0124] Instead of encoding the actual motion vectors of the current CU determined by the motion estimation unit 42 into the video bitstream, as described above... Figure 2 As described, the motion vector difference (MVD) of the current CU is generated by subtracting the motion vector predictor of the current CU from the actual motion vector of the current CU. By doing so, it is not necessary to encode the motion vectors determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0125] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a code block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using potential candidate motion vectors associated with spatially adjacent CUs and / or temporally co-located CUs. They then select a member from this motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current CU.
[0126] Typically, the basic inter-frame prediction scheme used in VVC remains almost identical to that in HEVC, except that several prediction tools are further extended, added, and / or improved, such as extended merge prediction, MMVD, and GPM.
[0127] Extended merge forecast
[0128] With the continuous improvement of video data capture technologies for preserving details in video data and the increasing finer video block sizes, the amount of data required to represent the motion vector of the current image has also increased significantly. One way to overcome this challenge is to use motion information (e.g., motion vectors) of spatially adjacent CUs, temporally co-located CUs, etc., of the current CU as an approximation (e.g., prediction) of the motion information of the current CU, which is also called the "motion vector predictor (MVP)" of the current CU. The "motion vector" used throughout this disclosure includes not only motion vectors between CUs from different frames (e.g., between temporally co-located CUs in inter-frame prediction) but also block vectors between CUs within the same frame (e.g., between spatially adjacent CUs in intra-frame prediction).
[0129] Similar to the process of selecting a prediction block from a reference picture during inter-frame prediction of a coded block, both video encoder 20 and video decoder 30 need to construct an MVP candidate list for the current CU using a set of rules, and then select an MVP candidate from the MVP candidate list as the MVP of the current CU. By doing so, it is not necessary to transmit the MVP candidate list itself between video encoder 20 and video decoder 30, and the index of the MVP candidate selected from the MVP candidate list is sufficient for video encoder 20 and video decoder 30 to use the same MVP candidate selected from the MVP candidate list to encode and decode the current CU.
[0130] In VVC, the MVP candidate list is constructed by including the following five types of MVPs in sequence:
[0131] —The spatial MVP of spatially adjacent CUs (i.e., spatial candidates);
[0132] —The temporal MVP of the CU that is temporally co-located (i.e., the temporal candidate);
[0133] —History-based MVP (HMVP) from a First-In-First-Out (FIFO) table;
[0134] —Paired average MVP; and
[0135] —Zero MVP.
[0136] The size of the MVP candidate list is signaled in the sequence parameter set header, and the maximum allowed size of the MVP candidate list is 6. For each CU encoded and decoded in merge mode, a truncated unary binary representation is used to encode the index of the best MVP candidate. The first binary number of the index is encoded using the context, and the other binary numbers used for the index are bypassed.
[0137] The export process for each type of MVP is shown below. For example, in HEVC, VVC also supports exporting the MVP candidate list of all CUs in parallel within a certain size region.
[0138] Deriving MVP from Spatial Candidates
[0139] In VVC, from spatial candidates (e.g., Figure 5 The MVP derived from the CU adjacent to the current CU 101 is the same as that in HEVC, except that the positions of the first two space candidates are swapped. Figure 5 Up to four spatial candidates are selected from the spatial candidates at the locations depicted: top position B0, left position A0, upper right position B1, lower left position A1, and upper left position B2. Derivation is performed in the order of the CUs at positions B0, A0, B1, A1, and B2. The CU at position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because said one or more CUs belong to other slices or tiles) or if intra-frame encoding / decoding is performed.
[0140] After adding the CU at position B0 as a candidate to the merged candidate list, the remaining candidates are added to the merged candidate list for redundancy checking. This ensures that candidates with the same motion information are excluded from the merged candidate list, thus improving encoding and decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only those using the same motion information are considered. Figure 6 The arrows in the diagram link the pairs of candidates, and a candidate is added to the merged candidate list only if the motion information of the candidate in the pair used for redundancy check is different from that of the candidate to be added. The spatial MVP derived from the candidates in the merged candidate list is added to the MVP candidate list.
[0141] Derive MVP from time candidates
[0142] During the MVP derivation from the time candidate, only one time candidate is added to the merge candidate list. Specifically, when deriving the MVP from that time candidate, it is based on the co-occurring CU (e.g., Figure 7 col_CU301) is used as the current CU (e.g., Figure 7 The corresponding image of curr_CU303 in the image (e.g., Figure 7 The time candidate (col_pic302) is used to derive the scaling motion vector, and the scaling motion vector is added as a time MVP candidate to the MVP candidate list. The list of reference images and their indices for deriving the co-located CU are explicitly signaled in the slice header. Figure 7 As shown, the scaled motion vector is obtained from the motion vector of the co-located CU using the image sequence count (POC) distance (i.e., tb and td), where tb is defined as the current image (e.g., ...). Figure 7 Reference image for curr_pic304 (e.g., Figure 7 The difference between curr_ref305 in the image and the current image in the POC, and td is defined as the reference image of the co-located image (e.g., Figure 7 The POC difference between col_ref306 and the corresponding image. The reference image index for the time candidate is set to zero.
[0143] like Figure 8 As depicted, the position of the temporal candidate (i.e., the co-occurring CU) in the current CU401 is selected between positions C0 and C1. If the CU at position C0 in the co-occurring picture is unavailable, intra-coded, or outside the current CTU line, then the CU at position C1 is used as the co-occurring CU for deriving the temporal MVP candidate. Otherwise, the CU at position C0 is used as the co-occurring CU for deriving the temporal MVP candidate.
[0144] Exporting HMVP candidates
[0145] After the spatial MVP and temporal MVP, HMVP candidates are added to the MVP candidate list. Motion information of previously encoded / decoded blocks is stored in the HMVP table and used as the MVP of the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-subblock is encoded / decoded via inter-frame encoding / decoding, the associated motion information is added to the last entry of the HMVP table as a new HMVP candidate.
[0146] The size of the HMVP table is set to 6. When a new HMVP candidate is inserted into the HMVP table, a constrained FIFO rule is used, where a redundancy check is first applied to find if a duplicate HMVP already exists in the HMVP table. If found, the duplicate HMVP is removed from the HMVP table, and then all HMVP candidates are shifted forward, with the duplicate HMVP added to the last entry in the HMVP table.
[0147] HMVP candidates can be used during the MVP candidate list construction process. The latest few HMVP candidates in the HMVP table are checked sequentially, and then inserted into the MVP candidate list after the temporal MVP candidates. Redundancy checks are applied to the HMVP candidates relative to the spatial and / or temporal MVP candidates.
[0148] To reduce the number of redundant check operations, the following simplification is introduced:
[0149] —The last two entries in the HMVP table are redundantly checked compared to the spatial MVP candidates derived from the spatial candidates at positions A1 and B1, respectively; and
[0150] —The process of building the MVP candidate list from the HMVP candidate is terminated once the total number of available MVP candidates reaches the maximum allowed size of the MVP candidate list minus 1.
[0151] Derivation of Paired Average MVP Candidates
[0152] Pairwise average MVP candidates are generated by averaging the MVPs derived from the first two predefined merge candidate pairs in an existing merge candidate list. The first merge candidate in a predefined pair can be defined as p0Cand, and the second merge candidate in a predefined pair can be defined as p1Cand. The average motion vector is calculated separately for each list of reference images based on the availability of motion vectors for p0Cand and p1Cand. If two motion vectors are available for a list of reference images, then these two motion vectors are averaged even if they point to different reference images, and the reference image for the average motion vector is set to the reference image of p0Cand; if only one motion vector is available for a list of reference images, then the motion vector is used directly; if no motion vector is available for a list of reference images, then the motion vectors and reference image indices for this list remain invalid.
[0153] Zero MVP
[0154] If the MVP candidate list is not full after adding paired average MVP candidates, insert zero MVPs at the end of the MVP candidate list until the maximum allowed size of the MVP candidate list is reached.
[0155] MMVD
[0156] As described above, in the merge mode, motion information (i.e., MVP candidates) is implicitly derived from the MVP candidate list constructed for the current CU and directly used as the MV of the current CU to generate prediction samples for the current CU. This may result in some error between the actual MV of the current CU and the implicitly derived MVP. To improve the accuracy of the MV of the current CU, MMVD is introduced in VVC, where the motion vector difference (MVD) of the current CU is added to the implicitly derived MVP to obtain the MV of the current CU. The MMVD flag is signaled after the regular merge flag is sent to indicate whether the MMVD mode is used for the current CU.
[0157] In MMVD mode, after an MVP candidate is selected from the first two MVP candidates in the MVP candidate list, MMVD information is signaled. The MMVD information includes an MMVD candidate flag, which specifies which of the first two MVP candidates is selected as the MV base, a distance index for indicating the motion amplitude information of the MVD, and a direction index for indicating the motion direction information of the MVD.
[0158] The distance index for specifying the motion magnitude information of the MVD indicates the reference image of the current CU pointed to from the selected MVP candidate point (e.g., Figure 9 The starting point in the L0 reference image 501 or L1 reference image 503 (e.g., by) Figure 9 The dashed circle in the table represents the predefined offset, and MVD can be derived from the offset and added to the selected MVP candidate. The relationship between the distance index and the predefined offset is specified in Table 1 below. Table 1
[0159] The direction index specifies the sign of the MVD, indicating the direction of the MVD relative to the starting point. Table 2 specifies the relationship between the direction index and predefined symbols. It should be noted that the meaning of the sign of the MVD can vary depending on the information of the selected MVP candidate. When the selected MVP candidate is a non-predictive MV or a bidirectional predictive MV with two MVs pointing to the same side of the current image (i.e., the POCs of both reference images of the current image (e.g., the reference image of List 0 and the reference image of List 1, which are also referred to as the L0 reference image and the L1 reference image, respectively) that are both greater than the POC of the current image, or both are less than the POC of the current image), the sign in Table 2 specifies the sign of the MVD added to the selected MVP candidate. When the selected MVP candidate is a bidirectional predicted MV with two MVs pointing to different sides of the current image (i.e., the POC of one reference image of the current image is greater than the POC of the current image, and the POC of the other reference image of the current image is less than the POC of the current image), if the POC distance of the L0 reference image (i.e., the POC distance between the L0 reference image and the current image) is greater than the POC distance of the L1 reference image (i.e., the POC distance between the L1 reference image and the current image), then the plus or minus sign in Table 2 specifies the plus or minus sign of the MVD added to the list of MVPs in list 0 of MVP0, and the plus or minus sign of the MVD added to the list of MVPs in list 1 of MVP1, and the MVD added to the list 1 of MVD1 of MVP1 is opposite to the plus or minus sign in Table 2; otherwise, if the POC distance of the L1 reference image is greater than the POC distance of the L0 reference image, then the plus or minus sign in Table 2 specifies the plus or minus sign of the MVD1 added to MVP1, and the plus or minus sign of the MVD0 added to MVP0 is opposite to the plus or minus sign in Table 2. Table 2
[0160] MVD is scaled based on the POC distance. If the POC distances of the L0 and L1 reference images are the same, then MVD does not need to be scaled. Otherwise, if the POC distance of the L0 reference image is greater than that of the L1 reference image, then MVD1 is scaled. If the POC distance of the L1 reference image is greater than that of the L0 reference image, then MVD0 is scaled.
[0161] GPM
[0162] In VVC, GPM is supported for inter-frame prediction. GPM uses CU-level flags as a merging mode for signal values. Other merging modes include regular merging mode, MMVD mode, CIIP mode, and sub-block merging mode. For all other 8... 64 and 64 Every possible CU size other than 8 ( GPM supports a total of 64 partitions.
[0163] When using GPM, the CU is divided into two parts by a geometrically positioned straight line. The position of the segmentation line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the CU obtained through geometric segmentation uses its own motion for inter-frame prediction; and only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, similar to regular bidirectional prediction, only two motion-compensated predictions are needed for each CU.
[0164] If GPM is used for the current CU, it further signals the partitioning mode (indicating the angle and offset of the geometric partition) and the geometric partition indexes of two merged indexes (one for each partition).
[0165] A unidirectional prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process described above. Let n denote the index of the unidirectional prediction motion vector in the unidirectional prediction candidate list. The LX motion vector of the nth merge candidate in the merge candidate list (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for GPM. These motion vectors are... Figure 10 The value is marked with "x". If the corresponding LX motion vector of the nth merge candidate in the merge candidate list does not exist, the L(1-X) motion vector of the same merge candidate is used to replace the unidirectional predicted motion vector of GPM.
[0166] CIIP
[0167] In VVC, when a CU is encoded in merge mode, if the CU includes at least 64 luma samples (i.e., the width of the CU multiplied by the height of the CU is equal to or greater than 64), and if both the width and height of the CU are less than 128 luma samples, an additional flag is signaled to indicate whether CIIP mode should be applied to the current CU. In CIIP mode, the prediction signal is obtained by combining the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal in CIIP mode is derived using the same inter-frame prediction process applied in regular merge mode; and the intra-frame prediction signal in CIIP mode is derived following the regular intra-frame prediction process with planar mode. Then, a weighted average is used to combine the intra-frame prediction signal and the inter-frame prediction signal, where the encoding patterns of the top and left adjacent blocks of the current CU 1601 (e.g., ...) are considered. Figure 11 The weight values are calculated as follows (as shown):
[0168] - If the top adjacent block is available and is intra-coded, isIntraTop is set to 1; otherwise, isIntraTop is set to 0.
[0169] - If the left adjacent block is available and is intra-coded, then isIntraLeft is set to 1; otherwise, isIntraLeft is set to 0.
[0170] - If (isIntraLeft+isIntraTop) equals 2, then the weight value is set to 3;
[0171] Otherwise, if (isIntraLeft+isIntraTop) equals 1, the weight value is set to 2;
[0172] Otherwise, the weight value is set to 1.
[0173] Predicted signal under CIIP mode Export as follows:
[0174] in, It is the inter-frame prediction signal in CIIP mode. It is an intra-frame prediction signal in CIIP mode. It is a weight value, and This indicates a right shift operation.
[0175] Intra-block copying in Universal Video Coding (VVC)
[0176] Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has already been reconstructed within the current frame. The luma block vector of an IBC-coded CU is integer-precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-frame-element and 4-frame-element motion vector precision. IBC-coded CUs are considered a third prediction mode in addition to intra-frame prediction mode or inter-frame prediction mode. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.
[0177] On the encoder side, hash-based motion estimation is performed on the IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-merging modes, a block vector search is first performed using a hash-based search. If the hash search does not return any valid candidates, a local search based on block matching is performed.
[0178] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. The hash key calculation for each location in the current image is based on 4x4 sub-blocks. For larger current blocks, a hash key matching the reference block is determined when all hash keys of all 4x4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost for each matching reference is calculated, and the block vector cost with the lowest cost is selected.
[0179] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.
[0180] At the CU level, IBC mode is signaled using flags, and it can signal either IBCAMVP mode or IBC skip / merge mode, as shown below:
[0181] IBC Skip / Merge Mode: The merge candidate index is used to indicate which block vectors from the list of neighboring candidate IBC coding blocks are used to predict the current block. The merge list includes spatial, HMVP, and paired candidates.
[0182] IBC AMVP mode: Encodes block vector differences in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the top neighbor (if IBC encoding is used). When either neighbor is unavailable, the default block vector is used as the predictor. A signaling flag indicates the index of the block vector prediction value.
[0183] IBC Reference Area
[0184] To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of predefined regions, including the current CTU region and some regions of the left CTU. Figure 12 The reference area for the IBC mode is shown, where each block represents a 64x64 luminance sample unit.
[0185] Depending on the location of the current encoding CU within the current CTU, the following applies:
[0186] If the current block falls within the top-left 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, the current block can also use CPR mode to reference reference samples in the bottom-right 64x64 block of the left CTU. The current block can also use CPR mode to reference reference samples in the bottom-left 64x64 block of the left CTU and the top-right 64x64 block of the left CTU.
[0187] If the current block falls within the upper right 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, if the brightness position (0, 64) has not yet been reconstructed relative to the current CTU, the current block can also use CPR mode to reference the reference samples in the lower left and lower right 64x64 blocks of the left CTU; otherwise, the current block can also reference the reference samples in the lower right 64x64 block of the left CTU.
[0188] If the current block falls within the lower left 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, if the brightness position (64,0) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference the reference samples in the upper right and lower right 64x64 blocks of the left CTU. Otherwise, the current block can also use CPR mode to reference the reference samples in the lower right 64x64 block of the left CTU.
[0189] If the current block falls within the bottom right 64x64 block of the current CTU, it can only use CPR mode to reference samples that have already been reconstructed in the current CTU.
[0190] This limitation allows the use of local on-chip memory for hardware implementation to implement the IBC mode.
[0191] IBC Interaction with Other Encoding Tools
[0192] The interaction between IBC mode and other inter-frame coding tools in VVC (such as Paired Merge Candidate, History-Based Motion Vector Predictor (HMVP), Combined Intra / Inter-Frame Prediction Mode (CIIP), Merge Mode with Motion Vector Difference (MMVD), and Geometric Partitioning Mode (GPM)) is as follows:
[0193] IBC can be used with pairwise merge candidates and HMVP. New pairwise IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC movements are inserted into a history buffer for future reference.
[0194] IBC cannot be used in combination with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM.
[0195] When using DUAL_TREE partitioning, IBC is not allowed for chroma-coded blocks.
[0196] Unlike in HEVC screen content encoding / decoding extensions, the current image is no longer included as one of the reference images in reference image list 0 for IBC prediction. The derivation process of motion vectors for IBC modes excludes all adjacent blocks in inter-frame modes, and vice versa. The following IBC design aspects are applied:
[0197] IBC shares the same process as regular MV merging, including pairwise merging candidates and history-based motion predictors, but does not allow TMVP and zero vectors because they are invalid for IBC mode.
[0198] Separate HMVP buffers (each HMVP buffer has 5 candidates) are used for regular MV and IBC.
[0199] Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream, and that merging should not be used if a merging candidate is invalid (out of range or zero). Such bitstream consistency constraints are expressed using a virtual buffer as described below.
[0200] For deblocking, IBC is processed as an inter-frame mode.
[0201] If the current block is encoded and decoded using IBC prediction mode, then AMVR does not use quarter-image elements; instead, AMVR is signaled to indicate only whether the MV is between pixels or 4 integer pixels.
[0202] The number of IBC merge candidates can be signaled separately in the slice header from the number of regular, sub-block, and geometric merge candidates.
[0203] The concept of a virtual buffer is used to describe the permissible reference region and effective block vector of an IBC prediction mode. Representing the CTU size as ctbSize, the virtual buffer ibcBuf has a width wIbcBuf = 128x128 / ctbSize and a height hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32.
[0204] The size of the VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize,64).
[0205] The virtual IBC buffer ibcBuf is maintained as follows.
[0206] At the beginning of decoding each CTU line, refresh the entire ibcBuf with an invalid value of -1.
[0207] When decoding VPDU(xVPDU, yVPDU) relative to the top-left corner of the image, set ibcBuf[x][y] = -1, where x = xVPDU%wIbcBuf, ..., xVPDU%wIbcBuf + W v -1;y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1.
[0208] After decoding, the CU includes (x, y) relative to the top-left corner of the image, set.
[0209] ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]
[0210] For a block covering coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid:
[0211] ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1.
[0212] Intra-block copying in Enhanced Compression Model (ECM)
[0213] In ECM, IBC is improved in the following ways.
[0214] IBC Merge / AMVP List Construction
[0215] The IBC merge / AMVP list construction has been modified as follows:
[0216] Only valid IBC merge / AMVP candidates can be added to the IBC merge / AMVP candidate list.
[0217] Candidates in the top right, bottom left, and top left spaces, as well as a pairwise average candidate, can be added to the IBC merge / AMVP candidate list.
[0218] Template-based adaptive reordering (ARMC-TM) is applied to the IBC merge list.
[0219] The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC merge candidates with full pruning, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0220] The zero vector candidate filling the IBC merge / AMVP list is replaced with a set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in the IBC merge mode, and therefore, it is discarded as a BVP in the IBC candidate list.
[0221] Three candidates are located at the nearest corner of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as follows. Figure 13 As shown.
[0222] IBC with template matching
[0223] Template matching is used in both IBC merge mode and IBC AMVP mode in IBC.
[0224] Compared to the merge list used in the regular IBC merge mode, the IBC-TM merge list is modified to select candidates based on a pruning method that uses the motion distance between candidates as in the regular TM merge mode. The zero motion at the end is replaced by motion vectors at the left (-W, 0), top (0, -H), and top-left corner (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0225] In IBC-TM merging mode, the selected candidates are refined using a template matching method prior to the RDO or decoding process. IBC-TM merging mode competes with the regular IBC merging mode and signals the TM merging flag.
[0226] In the IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merging list. Each of these three selected candidates is refined using a template matching method and ranked according to the template matching cost they produce. Then, during motion estimation, typically only the top two candidates are considered.
[0227] Template matching refinement is very simple for both the IBC-TM merging mode and the AMVP mode because the IBC motion vector is constrained to (i) integers and (ii) within the reference region, such as Figure 12 As shown. Therefore, in IBC-TM merge mode, all thinning is performed with integer precision, and in IBC-TM AMVP mode, the thinning is performed with integer or 4 image element precision depending on the AMVR value. Such thinning only accesses samples that are not interpolated. In both cases, the motion vectors of the thinning and the template used in each thinning step must adhere to the constraints of the reference region.
[0228] IBC Reference Area
[0229] The IBC reference area extends to the two CTU rows above. Figure 14 The reference region used to encode CTU(m, n) is shown. Specifically, for the CTU(m, n) to be encoded, the reference region includes CTUs with indices (m-2, n-2)…(W, n-2), (0, n-1)…(W, n-1), (0, n)…(m, n), where W represents the maximum horizontal index within the current tile, slice, or image. This setup ensures that for a CTU of size 128, IBC does not require additional memory in the current ETM platform. The sample-by-sample block vector search (or local search) range is limited to [–(C<<1), C>>2] in the horizontal direction and [–C, C>>2] in the vertical direction to accommodate the expansion of the reference region, where C represents the CTU size.
[0230] IBC merging mode with block vector difference
[0231] In ECM, an IBC merging mode with block vector difference is used. The distance set is {1-image element, 2-image element, 4-image element, 8-image element, 12-image element, 16-image element, 24-image element, 32-image element, 40-image element, 48-image element, 56-image element, 64-image element, 72-image element, 80-image element, 88-image element, 96-image element, 104-image element, 112-image element, 120-image element, 128-image element}, and the BVD directions are two horizontal directions and two vertical directions.
[0232] A basic candidate is selected from the top five candidates in the reordered IBC merge list. Then, all possible MBVD refinement positions (20×4) for each basic candidate are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and the reference at each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are retained as available positions and thus used for MBVD index encoding and decoding.
[0233] IBC adaptation for content captured by the camera
[0234] When adapting IBC for content captured by the camera, the IBC reference range is reduced from 2 CTU lines to 2x128 lines, such as Figure 15 As shown. On the encoder side, to reduce complexity, the local search range is set to [-8,8] in the horizontal direction and [-8,8] in the vertical direction, centered on the first block vector predictor of the current CU. This encoder modification is not applied to the SCC sequence.
[0235] CIIP combined with TIMD and TM
[0236] In CIIP mode, prediction samples are generated by weighting the inter-frame prediction signals that combine candidate predictions using CIIP-TM and the intra-frame prediction signals that are predicted using TIMD-derived intra-frame prediction modes. This method is only applicable to coded blocks with an area of 1024 or less.
[0237] The TIMD export method is used to export intra-prediction modes from CIIP. Specifically, it selects the intra-prediction mode with the smallest SATD value from the TIMD mode list and maps it to one of 67 regular intra-prediction modes.
[0238] Additionally, if the exported intra-prediction mode is an angle mode, it is also recommended to modify the weights of the two tests (wIntra, wInter). For near-horizontal mode (2 <= angle mode index < 34), such as Figure 16A As shown, the current block is vertically divided; for near-vertical mode (34 <= angle mode index < 66), the current block is horizontally divided, as shown. Figure 16B As shown.
[0239] Table 3 shows the different sub-blocks (wIntra, wInter). Table 3 shows the modified weights used in angle mode.
[0240] Using CIIP-TM, a CIIP-TM merge candidate list is constructed for the CIIP-TM pattern. The merge candidates are refined through template matching. The ARMC method also reorders the CIIP-TM merge candidates into regular merge candidates. The maximum number of CIIP-TM merge candidates is two.
[0241] Multiple Hypothesis Prediction (MHP)
[0242] In multi-hypothesis inter-frame prediction mode, in addition to the traditional bidirectional prediction signal, one or more additional motion-compensated prediction signals are signaled. The overall prediction signal is obtained by weighted superposition of samples. This utilizes dual prediction signals. and the first additional inter-frame prediction signal / hypothesis The obtained prediction signal It was obtained in the following way: (2)
[0243] According to the mapping given in Table 4, the weighting factor Specifyed by the new syntax element add_hyp_weight_idx: Table 4 add_hyp_weight_idx and The mapping between them.
[0244] Similar to the above, more than one additional prediction signal can be used. The overall prediction signal is obtained by iteratively accumulating each additional prediction signal. (3)
[0245] The obtained total predicted signal is used as the final prediction signal. (That is, the one with the largest index n) Within this mode, up to two additional prediction signals can be used (i.e., n is limited to 2).
[0246] The motion parameters for each additional prediction hypothesis can be explicitly signaled by specifying a reference index, motion vector predictor index, and motion vector difference, or implicitly signaled by specifying a merging index. A separate multi-hypothesis merging flag distinguishes these two signaling modes.
[0247] For inter-frame AMVP mode, MHP is only applied when unequal weights are selected in BCW in bidirectional prediction mode.
[0248] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0249] Geometric Partitioning Pattern (GPM) in ECM
[0250] GPM with Combined Motion Vector Difference (MMVD)
[0251] GPM in VVC extends the existing GPM unidirectional MV by applying motion vector refinement on top of it. First, a signal flag is sent to the GPMCU to indicate whether this mode is used. If this mode is used, each geometric partition of the GPM CU can further determine whether to signal the MVD. If the MVD is signaled for a geometric partition, the motion of the partition is further refined using the signaled MVD information after selecting a GPM merge candidate. All other processes remain the same as in GPM.
[0252] Similar to MMVD, MVD signals a pair of distances and directions. GPM with MMVD (GPM-MMVD) involves nine candidate distances (¼-picture element, ½-picture element, 1-picture element, 2-picture element, 3-picture element, 4-picture element, 6-picture element, 8-picture element, 16-picture element) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, MVD shifts left by 2, just like MMVD.
[0253] GPM with Template Matching (TM)
[0254] Template matching is applied to GPM. When GPM mode is enabled for CU, a signal is sent to the CU level flag to indicate whether TM is applied to both geometric partitions. TM is used to refine motion information for each geometric partition. When TM is selected, a template is constructed based on the segmentation angle using adjacent samples from the left and top, or adjacent samples from the left and top, as shown in Table 5. Motion is then refined by using the same search mode of the merging mode of the disabled half-image element interpolation filter to minimize the difference between the current template and the template in the reference image. Table 5
[0255] Table 5 shows the templates for the first and second geometric partitions, where A indicates the use of the upper sample points, L indicates the use of the left sample points, and L+A indicates the use of both the left and upper sample points.
[0256] The GPM candidate list is constructed as follows:
[0257] 1. Derive interleaved list-0 and list-1 MV candidates directly from the regular merge candidate list, where list-0 MV candidates have higher priority than list-1 MV candidates. Apply a pruning method with an adaptive threshold based on the current CU size to remove redundant MV candidates.
[0258] 2. Derive interleaved list-1 MV candidates and list-0 MV candidates directly from the regular merged candidate list, where list-1 MV candidates have higher priority than list-0 MV candidates. The same pruning method with an adaptive threshold is also applied to remove redundant MV candidates.
[0259] 3. Zero MV candidates are filled until the GPM candidate list is full.
[0260] GPM-MMVD and GPM-TM are dedicated to a single GPM CU. This is accomplished by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (if at least one GPM-MMVD flag is true), the GPM-TM flag is inferred to be false.
[0261] GPM with inter-frame prediction and intra-frame prediction
[0262] In GPM with inter-frame prediction and intra-frame prediction, the final prediction samples are generated by weighting the inter-frame prediction samples and intra-frame prediction samples for each GPM-separated region. Inter-frame prediction samples are derived from the inter-frame GPM, while intra-frame prediction samples are derived from the intra-prediction mode (IPM) candidate list and the index signaled from the encoder. The IPM candidate list size is predefined as 3. Available IPM candidates are the parallel angle mode (parallel mode) for GPM block boundaries, the vertical angle mode (vertical mode) for GPM block boundaries, and the planar mode, respectively, such as... Figures 17A to 17C As shown. Furthermore, as... Figure 17D As shown, GPM with intra-frame prediction is limited to reduce the signaling overhead of IPM and avoid increasing the size of the intra-frame prediction circuitry on the hardware decoder. Furthermore, direct motion vectors and IPM storage on the GPM-mixed region are introduced to further improve coding performance.
[0263] In DIMD and neighbor-mode-based IPM derivation, parallel modes are registered first. Therefore, if no identical IPM candidate exists in the list, the maximum two IPM candidates derived from the decoder-side intra-frame mode derivation (DIMD) method and / or neighbor block derivation can be registered. For neighbor-mode derivation, there are at most five available neighbor block locations, but they are limited by the angle of the GPM block boundaries as shown in Table 6, which has been used for GPM with template matching (GPM-TM). Table 6
[0264] Table 6 shows the positions of available neighboring blocks derived from the IPM candidate derivation based on the angles of the GPM block boundaries. A and L represent the top and left sides of the predicted block.
[0265] GPM-intraframe can be combined with GPM-MMVD (GPM-MMVD). TIMD is used as an IPM candidate for GPM-intraframe to further improve coding performance. Parallel mode can be registered first, followed by TIMD, DIMD, and IPM candidates for adjacent blocks.
[0266] Template-match-based reordering of GPM splitting patterns
[0267] In template-match-based reordering of GPM split patterns, given the motion information of the current GPM block, the corresponding TM cost value for the GPM split pattern is calculated. Then, all GPM split patterns are reordered in ascending order based on their TM cost values. Instead of sending GPM split patterns, a signal is sent using Golomb-Rice codes to indicate the exact index of the GPM split pattern within the reordering list.
[0268] The GPM partition reordering method is a two-step process performed after generating the corresponding reference templates for the two GPM partitions in the coding unit, as shown below:
[0269] • Extend the GPM partition edge to the reference template of two GPM partitions to generate 64 reference templates and calculate the corresponding TM cost for each of the 64 reference templates;
[0270] • Reorder the GPM split patterns in ascending order based on their TM cost values and mark the best 32 as available split patterns.
[0271] like Figure 18 As shown, the edges on the template extend from the edges of the current CU, but the GPM blending process is not used in the template area that crosses the edges.
[0272] After using incremental reordering with TM cost, signal the index.
[0273] Intra-frame template matching
[0274] Intra-template matching prediction (intra-TMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, with its L-shaped template matching the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and performs the same prediction operation on the decoder side.
[0275] The prediction signal is obtained by comparing the L-shaped causal neighbors of the current block with... Figure 19 It is generated by matching another block in a predefined search region, which consists of the following parts:
[0276] R1: Current CTU
[0277] R2: Top Left CTU
[0278] R3: Above CTU
[0279] R4: Left CTU
[0280] The sum of absolute differences (SAD) is used as the cost function.
[0281] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0282] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:
[0283] SearchRange_w=a BlkW
[0284] SearchRange_h=a BlkH
[0285] in' ' is a constant that controls the trade-off between gain and complexity. In fact, ' 'Equals 5.'
[0286] For CUs with a width and height less than or equal to 64, the internal template matching tool is enabled. The maximum CU size for intra-frame template matching is configurable.
[0287] When the current CU does not use DIMD, the intra-template matching prediction mode is signaled at the CU level through a dedicated flag.
[0288] Fusion for Template-Based Intra-Frame Mode Export (TIMD)
[0289] For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. First, the two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are weighted and fused after applying the PDPC procedure, and this weighted intra-prediction is used to encode the current CU. Location-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0290] The costs of the two chosen modes are compared to a threshold. In the test, cost factor 2 is applied as follows:
[0291] costMode2<2 costMode1
[0292] If the condition is true, then apply fusion; otherwise, use only mode 1.
[0293] The weights of the patterns are calculated based on their SATD costs as follows:
[0294] Weight 1 = costMode2 / (costMode1 + costMode2)
[0295] Weight 2 = 1 - Weight 1
[0296] Division is performed using the same lookup table (LUT)-based integration scheme used by CCLM.
[0297] Localized lighting compensation (LIC)
[0298] LIC is an inter-frame prediction technique used to model the local illumination changes between the current block and its predicted block based on the local illumination changes between the current block template and the reference block template. The parameters of the function can be represented by scaling α and offset β, forming a linear equation, i.e., α p[x]+β is used to compensate for illumination variations, where p[x] is a reference sample point at position x on the reference image, pointed to by the MV. When surround motion compensation is enabled, surround offset should be considered to clip the MV. Since α and β can be derived based on the current block template and the reference block template, they require no signaling overhead other than signaling the LIC flag for AMVP mode to indicate the use of the LIC.
[0299] The local illumination compensation proposed in JVET-O0066 is used for unidirectional prediction of inter-frame CUs with the following modifications:
[0300] Intra-frame adjacent samples can be used for LIC parameter derivation;
[0301] Disable LIC for blocks with fewer than 32 luminance samples;
[0302] For both non-sub-block mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than the partial template block samples corresponding to the first upper left 16×16 cell; and
[0303] Samples of the reference block template are generated by using a MC with block MV without rounding it to integer pixel precision.
[0304] OBMC
[0305] When applying OBMC, motion information of neighboring blocks with weighted predictions as described in JVET-L0101 is used to refine the top and left boundary pixels of the CU.
[0306] The following conditions should not be used for OBMC:
[0307] When OBMC is disabled at the SPS level;
[0308] When the current block has intra-frame mode or IBC mode;
[0309] When applying a LIC to the current block; and
[0310] When the current luminance block area is less than or equal to 32.
[0311] Sub-block boundary OBMC is performed by applying the same blending to the boundary pixels of the top, left, bottom, and right sub-blocks using motion information from adjacent sub-blocks. This enables sub-block-based coding tools.
[0312] Affine AMVP mode;
[0313] Affine merging mode and sub-block-based temporal motion vector prediction (SbTMVP); and
[0314] Bilateral matching based on sub-blocks.
[0315] When using OBMC mode in CIIP mode with LMCS, inter-frame blending is performed before LMCS mapping of inter-frame samples. LMCS is applied to the blended inter-frame samples that are combined with intra-frame samples applied in CIIP mode.
[0316] in This represents the sample points predicted by the motion of the current block in the original domain. This represents the sample points predicted in the mapping domain. This represents the sample points predicted by the motion of neighboring blocks in the original domain, and and It's the weight.
[0317] OBMC based on template matching
[0318] In the template matching-based OBMC scheme, instead of directly using weighted prediction, the predicted value of the CU boundary sample derivation method is determined based on the template matching cost, including using only the motion information of the current block, or using the motion information of adjacent blocks, or a hybrid mode.
[0319] In this scheme for each block with a size of 4×4 at the top CU boundary, the template size is equal to 4×1. If N adjacent blocks have the same motion information, the template size is increased to 4N×1 because MC operations can be processed at once. For each left block with a size of 4×4 at the left CU boundary, the left template size is equal to 1×4 or 1×4N (…). Figure 20 ).
[0320] For each 4x4 top block (or N (4x4 block group), and then derive the predicted values of the boundary samples after the following steps.
[0321] Taking block A as the current block and its adjacent block AboveNeighbor_A as an example, the operations on the left block are performed in the same way.
[0322] First, the matching costs (Cost1, Cost2, Cost3) of three templates are measured by the SAD between the reconstructed sample points of the template and their corresponding reference sample points. The corresponding reference sample points are derived by the MC process based on the following three types of motion information:
[0323] Calculate Cost1 based on the motion information of A.
[0324] Cost2 is calculated based on the motion information of AboveNeighbor_A.
[0325] Cost3 is calculated based on the weighted prediction of the motion information of A and AboveNeighbor_A, where the weighting factors are 3 / 4 and 1 / 4, respectively.
[0326] Secondly, by comparing Cost1, Cost2, and Cost3, a method is selected to calculate the final prediction result of the boundary sample points.
[0327] The original MC result using the motion information of the current block is represented as Pixel1, and the MC result using the motion information of neighboring blocks is represented as Pixel2. The final prediction result is represented as NewPixel.
[0328] If Cost1 is the minimum, then NewPixel(i,j) = Pixel1(i,j).
[0329] If (Cost2 + (Cost2 >> 2)) + (Cost2 >> 3))) <= Cost1, then use mixed mode 1.
[0330] For a luminance block, the number of mixed pixel rows is 4.
[0331] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5
[0332] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3
[0333] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4
[0334] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5
[0335] For a chroma block, the number of mixed pixel rows is 1.
[0336] NewPixel(i,0)=(26×Pixel1(i,0)+6×Pixel2(i,0)+16)>>5
[0337] If Cost1 <= Cost2, then use Mixed Mode 2.
[0338] For a luminance block, the number of mixed pixel rows is 2.
[0339] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4
[0340] NewPixel(i,1)=(31×Pixel1(i,1)+Pixel2(i,1)+16)>>5
[0341] For chroma blocks, the number of mixed pixel rows / columns is 1.
[0342] NewPixel(i,0)=(15×Pixel1(i,0)+Pixel2(i,0)+8)>>4
[0343] Otherwise, use mixed mode 3.
[0344] For a luminance block, the number of mixed pixel rows is 4.
[0345] NewPixel(i,1)=(7×Pixel1(i,1)+Pixel2(i,1)+4)>>3
[0346] NewPixel(i,2)=(15×Pixel1(i,2)+Pixel2(i,2)+8)>>4
[0347] NewPixel(i,3)=(31×Pixel1(i,3)+Pixel2(i,3)+16)>>5
[0348] For a chroma block, the number of mixed pixel rows is 1.
[0349] NewPixel(i,0)=(7×Pixel1(i,0)+Pixel2(i,0)+4)>>3
[0350] Template Matching-Based Adaptive Reordering of Merging Candidates (ARMC-TM)
[0351] The reordering method is applied to the regular merge pattern, template matching (TM) merge pattern, and affine merge pattern (excluding SbTMVP candidates). For the TM merge pattern, merge candidates are reordered before the refinement process.
[0352] After constructing the merge candidate list, the merge candidates are divided into several subgroups. For the regular merge mode and the TM merge mode, the subgroup size is set to 5. For the affine merge mode, the subgroup size is set to 3. The merge candidates in each subgroup are reordered incrementally based on the cost value based on template matching. For simplicity, the merge candidates in the last subgroup, but not the first subgroup, are not reordered.
[0353] The template matching cost of merging candidates is measured by the sum of absolute differences (SAD) between the samples of the current block's template and their corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located using the motion information of the merging candidates.
[0354] When merging candidates utilize bidirectional prediction, the reference samples for the template of the merging candidates are also generated through bidirectional prediction, such as... Figure 21 As shown.
[0355] For sub-block-based merge candidates with a sub-block size equal to Wsub × Hsub, the above template includes several sub-templates of size Wsub × 1, and the left template includes several sub-templates of size 1 × Hsub. For example... Figure 22 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0356] Direct block vectors for chroma blocks
[0357] Direct block vectors are used for chroma blocks in dual-tree slices. When a chroma dual-tree is activated, a signal flag is sent to indicate whether to use IBC mode for encoding and decoding of the chroma blocks. Figure 23 If one of the luma blocks in the five locations shown is encoded / decoded in IBC or intra-frame TMP mode, then its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform block vector scaling.
[0358] While existing IBC schemes offer significant improvements to intra-frame coding in ECM, there is room for further performance enhancements. Simultaneously, some aspects of existing Convolutional Cross-Component Model (CCCM) patterns need simplification for efficient codec hardware implementation or improvement for better coding efficiency. Furthermore, a better balance needs to be struck between implementation complexity and its coding / decoding efficiency benefits.
[0359] In order to address the aforementioned problems, this disclosure provides a method for further improving the existing design of IBC. Generally, the main features of the technology proposed in this disclosure are summarized below.
[0360] The CCCM tool is used to filter IBC predictions. Filtered Intra-Block Copying (FIBC) is a special intra-prediction mode that applies filters to IBC-based prediction blocks to increase prediction accuracy and adapt the characteristics of the copied blocks to the local neighborhood.
[0361] In FIBC, training samples can be adjacent to the current block. Knowing references from local regions can improve the accuracy of predictions.
[0362] In FIBC, training samples do not need to be adjacent to the current block. Knowing references from non-local regions can also improve the accuracy of predictions.
[0363] In FIBC, only one assumption can be used: that is, the best matching block that results in the minimum matching cost is selected as the final prediction.
[0364] In FIBC, multiple hypotheses can also be used.
[0365] It should be understood that the accompanying drawings in this disclosure can be combined with all the examples mentioned in this disclosure, and the disclosed methods can be applied independently or in combination.
[0366] Filtered intra-block copy (FIBC)
[0367] According to one or more embodiments of this disclosure, the CCCM tool is used to filter IBC predictions. Different methods can be used to achieve this goal. Existing CCCM modes apply various filters to predict chroma sample values based on corresponding luma sample values. Unlike CCCM, Filtered Intra-Block Copying (FIBC) is a special intra-prediction mode that applies filters to the IBC-based prediction block to predict the target luma or chroma sample of the current block based on the corresponding luma or chroma sample of the reference block, respectively, in order to increase prediction accuracy and adapt the characteristics of the copied block to the local neighborhood.
[0368] According to one or more embodiments of this disclosure, the IBC prediction is further filtered. Different methods can be used to achieve this objective. Filtered Intra-Block Copying (FIBC) is a special intra-prediction mode that applies filters to the predicted blocks based on intra-block copying to increase prediction accuracy and adapt the characteristics of the copied blocks to the local neighborhood.
[0369] According to one or more embodiments of this disclosure, reconstructed luminance / chrominance samples on the template region of a reference block are used as input to a filter during the training phase, and corresponding reconstructed luminance / chrominance samples in the template region of the current block are the target. In one example, Figure 25 A filter shape (cross-shaped) and training region for a reference block are shown. It should be understood that for this filter shape, both the template region and the boundary regions of the template region can be part of the training region of the reference block. Reconstructed samples in the boundary regions can be used for training when available, and they are filled with the nearest available samples when unavailable. Conversely, the training region of the current block can be determined as the template region of the current block. In the filtering phase, after the filter coefficients of the filter have been trained / determined through the training phase, the filter can be applied to the corresponding sample values of the reference block and the boundary regions of the reference block to predict each of the sample values of the current block.
[0370] According to one or more embodiments of this disclosure, predicted samples can be used as input to a filter during the prediction process. In one example, Figure 26 The diagram illustrates the prediction samples used in the prediction process, where gray areas (as shown by patterns with sparse dots) are prediction samples, small patches without patterns represent reconstructed samples, and patterns with dense dots represent locations to be predicted, which will become prediction samples after being predicted. The following section, in conjunction with the appendix... Figure 44 and 45 Describes a method for generating predictions for video decoding / encoding.
[0371] According to one or more embodiments of this disclosure, the filter coefficients (i.e., parameters) are derived using a regression-based MSE minimization technique (i.e., LDL decomposition) that exists in the ECM and is utilized by other tools (such as CCCM).
[0372] According to one or more embodiments of this disclosure, a convolutional N-tap (N is an integer and greater than 1) filter may include an (N-1-M) tap space term (M is an integer), M nonlinear terms, and a bias term. The (N-1-M) tap space term corresponds to adjacent sample values, such as those from... Figure 27 The brightness samples of the reconstructed reference block shown are (i.e., L0, L1, ..., L8). In this example, the formula for each new predicted brightness sample is as follows:
[0373] in, Is with The associated coefficients, and This is the offset (i.e., 1 << (bitDepth - 1)). The reference brightness sample value of the upper-left sample adjacent to the current block can be used as the offsetLuma value. The location and number of spatial and nonlinear terms can differ. Examples of filter taps with different shapes / numbers are as follows... Figure 28 As shown. For another example, use the different positions and quantities shown in the table below.
[0374] According to one or more embodiments of this disclosure, the filter shape may be rectangular, N M (N and M are integers and greater than 1). Figure 29 Examples of filter taps of different shapes / numbers are shown. The corresponding center point (C) position may differ according to one or more embodiments of the invention. The corresponding center point (C) may also be referred to as the position to be predicted. Examples of different positions of the center point (C) are shown in... Figure 29 As shown in the image.
[0375] like Figure 29As shown, the position to be predicted in each of the filter shapes 0-11 is identified by the letter "C". For example, shapes 0 and 1 are both rectangles with a width of 3 rows and a height of 3 rows. Shape 0 is identified using the position to be predicted at the bottom right of the rectangle (i.e., the third row and third column), while shape 1 is identified using different positions to be predicted at the center of the rectangle (i.e., the second row and second column). Similarly, shape 2 is a rectangle with a width of 3 rows and a height of 4 rows, and is identified using the position to be predicted at the third row and second column. Shapes 3 and 4 are both rectangles with a width of 4 rows and a height of 4 rows. Shape 3 is identified using the position to be predicted at the third row and third column, while shape 4 is identified using different positions to be predicted at the bottom right of the rectangle (i.e., the fourth row and fourth column). Shapes 5 and 6 are both rectangles with a width of 2 rows and a height of 8 rows. Shape 5 is identified using the predicted position at the bottom right of the rectangle (i.e., the eighth row and the second column), while shape 6 is identified using different predicted positions at the fifth row and the second column. Shapes 7 and 8 are both rectangles with a width of 8 rows and a height of 2 rows. Shape 7 is identified using the predicted position at the bottom right of the rectangle (i.e., the second row and the eighth column), while shape 8 is identified using different predicted positions at the second row and the fifth column. Shape 9 is a rectangle with a width of 6 rows and a height of 2 rows, and is identified using the predicted position at the bottom right of the rectangle (i.e., the second row and the sixth column). Shapes 10 and 11 are both rectangles with a width of 2 rows and a height of 6 rows. Shape 10 is identified using the predicted position at the fourth row and the second column, while shape 11 is identified using different predicted positions at the bottom right of the rectangle (i.e., the sixth row and the second column).
[0376] The width / height of the filter shape and the location to be predicted in this disclosure are not limited to... Figure 29 The shape shown. In this disclosure, filtered IBC can be performed using appropriate filter shapes from among different filter shapes to further increase prediction accuracy.
[0377] According to one or more embodiments of this disclosure, the filter shape may be rectangular and exclude the lower right sample point, and the number used is N. M-1 (N and M are integers and greater than 1). Figure 30 Examples of different shapes / numbers of filter taps are shown, where C is the corresponding center point location and / or the location to be predicted.
[0378] like Figure 30As shown, filter shapes 0-12 are rectangles with different widths and / or different heights and no bottom-right sample / position, which is the location to be predicted and identified by the letter "C" (see the unpatterned block). For example, shape 0 has a width and height of 3 rows; shape 1 has a width and height of 3 rows and 2 rows; shape 2 has a width and height of 3 rows and 4 rows; shape 3 has a width and height of 4 rows and 3 rows; shape 4 has a width and height of 4 rows; shape 5 has a width and height of 2 rows and 8 rows; shape 6 has a width and height of 2 rows and 3 rows; shape 7 has a width and height of 2 rows and 4 rows; shape 8 has a width and height of 8 rows and 2 rows; shape 9 has a width and height of 6 rows and 2 rows; shape 10 has a width and height of 4 rows and 2 rows; shape 11 has a width and height of 2 rows and 5 rows; and shape 12 has a width and height of 2 rows and 6 rows. The width and height of the rectangle without a bottom-right sample in this disclosure are not limited to... Figure 30 The shape shown.
[0379] The following text combines Figure 46 and Figure 47 Description of usage such as Figures 29-30 The filter shape shown is used for a filtered IBC method of video decoding / encoding.
[0380] According to one or more embodiments of this disclosure, the number of filter taps can be predefined or signaled at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0381] According to one or more embodiments of this disclosure, the template size and shape may be the same as in the intra-frame TMP, and the template size used for training depends on their availability in the four rows above and to the left of the current block.
[0382] According to one or more embodiments of this disclosure, depending on the availability of the current block, the template size for training can be up to 5 rows above and to the left of the current block.
[0383] According to one or more embodiments of this disclosure, the template size and shape may be the same as those in CCCM, and the template size used for training is 6 rows above and to the left of the current block, depending on their availability.
[0384] According to one or more embodiments of this disclosure, the template size for training can be N rows above and to the left of the current block, depending on their availability, where N is an integer.
[0385] According to one or more embodiments of this disclosure, the template size for training can be N rows above the current block, depending on their availability, where N is an integer.
[0386] According to one or more embodiments of this disclosure, the template size for training can be N rows to the left of the current block, depending on their availability, where N is an integer.
[0387] According to one or more embodiments of this disclosure, the size of the template used for training can depend on the filter shape. In one example, if the height of the filter shape is greater than its width, then the size of the template used for training can be N rows above the current block, depending on its availability, where N is an integer. Similarly, in another example, if the width of the filter shape is greater than its height, then the size of the template used for training can be N rows to the left of the current block, depending on its availability, where N is an integer. That is, instead of an L-shape, the template used for training the filter can be a rectangular shape, and the width of the template (in the case above the current block) or the height of the template (in the case to the left of the current block) can depend on the width or height of the current block. Therefore, the corresponding template associated with the reference block will be above or to the left of the reference block, which has the same size and shape. Such a template for training can be used in conjunction with the following... Figure 46 and Figure 47 The method described is a filtered IBC.
[0388] According to one or more embodiments of this disclosure, reference samples / template regions of the reference block / template region of the current block can be predefined or signaled / switched at different codec levels (such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level).
[0389] According to one or more embodiments of this disclosure, location information can be used to calculate model parameters, including utilizing horizontal / vertical / diagonal distances and their nonlinear terms. One or more pieces of location information can be used for this purpose. In one example, the location-based parameters are related to the vertical and horizontal coordinates (Xc, Yc) of the center brightness sample point and are calculated relative to the upper-left coordinates (Xt1, Ytl) of the block, for example, Xc-Xtl+Yc-Ytl. In another example, the location-based parameters are related to the vertical and horizontal coordinates (Xc, Yc) of the center brightness sample point and are calculated relative to the upper-left coordinates (Xtl, Ytl) of the block, for example, Xc-Xtl+Yc-Ytl, Xc-Xtl, Yc-Ytl. In yet another example, the position-based parameters are related to the vertical and horizontal coordinates (Xc, Yc) of the center brightness sample point, and are calculated relative to the top-left coordinates (Xt1, Ytl) of the block, for example, (Xc-Xt1+Yc-Ytl) / N, where N is a predefined number, such as 2. In yet another example, the position-based parameters are related to the vertical and horizontal coordinates (Xc, Yc) of the center brightness sample point, and they are calculated relative to the top-left coordinates (Xt1, Ytl) of the block, for example, (Xc-Xt1+Yc-Yt1) / N1, (Xc-Xt1) / N2, (Yc-Yt1) / N3, where N1~N3 are predefined numbers, such as 2, 3, and 4. In yet another example, the position-based nonlinear term is expressed as a power of two of the horizontal / vertical / diagonal distances, for example, (Xc-Xtl+Yc-Ytl). (Xc-Xtl+Yc-Ytl), (Xc-Xtl) (Xc-Xtl), (Yc-Ytl) (Yc-Ytl), where (Xc, Yc) are the vertical and horizontal coordinates of the center brightness sample point, and (Xtl, Ytl) is the top-left coordinate.
[0390] According to one or more embodiments of this disclosure, an enable flag can be signaled in the bitstream to indicate the FIBC mode being used. The enable flag can be signaled at different codec levels (such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level).
[0391] According to one or more embodiments of this disclosure, instead of explicitly signaling the selected mode flag, the mode flag can be derived at the decoder to save bit overhead.
[0392] According to one or more embodiments of this disclosure, no additional control flags are required, and the FIBC mode will be derived under certain predefined conditions (e.g., a specific mode, a specific block size, a specific partition). When the predefined conditions are matched, the FIBC mode will be derived based on previously decoded information.
[0393] According to one or more embodiments of this disclosure, samples in regions not adjacent to the current block can be used to derive the model of the current block. In one embodiment, a candidate region list with N candidates can be constructed by sequentially examining potential M×M regions. If an examined region is available, it is added to the candidate region list. For example, a candidate region list with 6 candidates is constructed by sequentially examining potential 8×8 regions. The top-left position of the potential 8×8 regions is predetermined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (2, 0 ... xStep,0), (0, -2) yStep), (-2 xStep,2 yStep), (2) xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep), (2) xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 yStep), (-xStep / 2,0), (0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep=Max(width,16), yStep=Max(height,16). Figure 31 Some possible locations of the candidate regions are shown.
[0394] According to one or more embodiments of this disclosure, a non-nearest neighbor candidate with N candidates can be constructed by the positions and inclusion order of spatial non-nearest neighbor candidates from two sets of spatial non-nearest neighbor candidates in an inter-frame merging mode. If an area is available for inspection, it is added to the candidate area list. Figure 32 Some possible candidate locations are shown.
[0395] According to one or more embodiments of this disclosure, inherited parameters of FIBC from previously decoded TB / CB / slice / picture / sequence levels can be used in the current block. According to one or more embodiments of this disclosure, a control flag is signaled at the TB / CB / slice / picture / sequence level to indicate whether the inherited FIBC signaling is enabled or disabled. When the control flag is signaled as enabled, the inherited FIBC flag is further signaled to the decoder to indicate whether the inherited FIBC is used at the signaling level.
[0396] According to one or more embodiments of this disclosure, derived parameters from previously decoded TB / CB / slice / picture / sequence level FIBCs can be stored and used as the current FIBC (referred to as the inherited FIBC). In one embodiment, a history-based FIBC (H-FIBC) table can be maintained, similar to an HMVP table. In one embodiment, an index value can be signaled in the bitstream to indicate which candidate model in the H-FIBC table is selected. In one embodiment, the corresponding table can be updated after decoding an FIBC encoded block. In one embodiment, the size of the H-FIBC table is N. N is an integer (e.g., 4, 5, 6, 7).
[0397] According to one or more embodiments of this disclosure, the FIBC flag can be inherited from an IBC HMVP candidate.
[0398] According to one or more embodiments of this disclosure, the FIBC flag can be inherited from the IBC space MVP from the spatially adjacent CU.
[0399] According to one or more embodiments of this disclosure, the FIBC flag can be inherited from the IBC time MVP from a time-coordinated CU.
[0400] Multiple Hypothesis FIBC
[0401] According to one or more embodiments of this disclosure, more than one prediction block candidate is used and weighted to generate the final prediction for the current block. Assume N prediction block candidates are used.
[0402] Predicted block candidate export
[0403] In one embodiment, candidate prediction blocks are searched and selected based on a criterion of minimizing template matching cost; that is, the top N candidates that result in the minimum template matching cost are selected. Template matching cost may not be limited to SAD (sum of absolute differences) and SSE (sum of squared errors).
[0404] In one embodiment, prediction block candidates can be selected based on a predefined pattern (i.e., a planar pattern).
[0405] In one embodiment, prediction block candidates can be selected based on adjacent predefined patterns (i.e., top predefined pattern, left predefined pattern).
[0406] Fixed Multiple Hypothesis FIBC
[0407] In this embodiment, the weighting factor for generating the final prediction block is predefined and fixed on both the encoder and decoder sides. As an example, equal weighting factors can be used, i.e., 1 / N for all candidate blocks.
[0408] Adaptive Multihypothesis Intra-FIBC
[0409] To adapt to the different characteristics of video content, an adaptive multi-hypothesis intra-FIBC method is also proposed.
[0410] In one embodiment, the weighting factor can be derived based on the template matching cost. The template matching costs of N candidates are represented as... , , ..., The weighting factors are calculated as follows. (4)
[0411] It should be noted that template matching costs can be measured using (but are not limited to) SAD and SSE.
[0412] In another embodiment, the weighting factor can be derived / switched based on the block size or syntax element signaled at different codec levels (such as SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level).
[0413] In another embodiment, the weighting factor can be derived on the encoder side and then signaled to the decoder in the bitstream. The N prediction block candidates are represented as... , … And represent the current block as The weighting factor can then be solved using the following equation: (5)
[0414] Equation (5) can be solved using the Wiener-Hopf equation as an ALF. The derived filter coefficients are then quantized to integer type, and the derived filter coefficients are signaled at the block level.
[0415] In yet another embodiment, weighting factors are derived based on a template, and these derived weighting factors are applied to prediction block candidates to generate the final prediction block. The template for the prediction candidate is represented as... , … And represent the current block as The weighting factor can then be derived using the following equation: (6)
[0416] Equation (6) can be solved using the Wiener-Hopf equation. Then, the final predicted block can be computed as... ,in This represents the i-th prediction block candidate.
[0417] The FIBC model leverages nonlocal correlation to improve prediction accuracy, where similar blocks are searched and used to generate the final prediction block. In this embodiment, a combination of nonlocal mean filtering and multi-hypothesis FIBC is proposed, as described below. In the first step, N prediction block candidates are searched and identified for inclusion in FIBC. In the second step, weighting factors are calculated as follows. (7)
[0418] in Used to measure the distance between the template of the i-th prediction block candidate and the template of the current block. Used as a weighting degree, and Normalization constant: (8)
[0419] To calculate the weighting factors in equation (7), the strength of the weights must first be determined. Several methods are proposed in this disclosure to determine the weighting strength.
[0420] In the first approach, a candidate list of weighted strength values, including some typical weighted strength values, is defined and fixed on both the encoder and decoder sides. On the encoder side, rate-distortion optimization is used to examine the weighted strength values, and the optimal weighted strength value is identified in the bitstream and signaled to the decoder side.
[0421] In the second method, the templates of the predicted block candidates and the template of the current block are used to estimate the weighted intensity value. The template of the predicted candidate is represented as... , … And represent the current block as Then, the weighted strength value can be solved using the following equation: (9)
[0422] In the third method, the QP value and variance of the template of the current block can be used to estimate the weighted intensity value, that is, the relationship between the weighted intensity value, QP value and template variance can be fitted offline.
[0423] To better utilize the nonlocal correlations in FIBC, in this embodiment, singular value decomposition (SVD) is used to generate the final predicted block from the prediction block candidates. The width and height of the current block are represented by W and H, and the area of the current block is represented by... .
[0424] Step 1. As performed in FIBC, search and identify K candidate prediction blocks. .
[0425] Step 2. Current block K predicted block candidate building block group And it is arranged as a matrix: (10)
[0426] in, It is the size of The matrix, by grouping Each candidate arrangement in the array is a column vector.
[0427] Step 3. For the matrix Perform SVD decomposition. (11)
[0428] Step 4. For the singular value matrix Apply soft thresholding. (12)
[0429] in It is contraction With threshold A function of the diagonal elements. For The k-th diagonal element in the array is level nonlinear function at the point Shrink: (13)
[0430] It is composed of singular values that contract at the diagonal position. The matrix formed.
[0431] Step 5. Perform inverse SVD to obtain the filtered patch group. (14)
[0432] One of the key steps is determining a threshold for each diagonal element in step 4. In this disclosure, the threshold is calculated as follows. The threshold is estimated for each group of image patches using the following equation: (15)
[0433] in It is the standard deviation of the noise, and It is a group The standard deviation of the original block in the k-th dimension of the SVD space. The deviation estimate of the original block in the SVD space is as follows. (16)
[0434] in yes The k-th singular value. When When the value is zero, the soft thresholding operation is skipped. Additionally, using... and The parameterized power function uses the bias of the prediction block to estimate the bias of the noise. (17)
[0435] in The calculation is as follows: (18)
[0436] Here, Represents the candidate vector of the prediction block The i-th pixel.
[0437] Multiple Assumptions FIBC Signaling
[0438] In this disclosure, the proposed multi-hypothesis FIBC can be used as a replacement for the current FIBC mode, or the encoder can adaptively select either the FIBC mode or the multi-hypothesis FIBC mode.
[0439] In one embodiment, the proposed multiple hypothesis FIBC is used as a replacement for the current FIBC pattern, i.e., multiple hypotheses are always used for prediction.
[0440] In another embodiment, one of the multi-hypothesis FIBC methods described above is used in conjunction with the current FIBC mode. A signal flag is sent in the bitstream to indicate whether the multi-hypothesis FIBC mode is applied to the CU.
[0441] In another embodiment, more than one of the multi-hypothesis FIBC methods described in the preceding sections is used in conjunction with the current FIBC mode. First, a signal flag is sent in the bitstream to indicate whether the multi-hypothesis FIBC mode is applied. Then, an index is signaled to indicate which of the multi-hypothesis FIBC methods is applied to the CU.
[0442] Coordination of filters for FIBC and FTMP modes
[0443] TMP prediction can also be filtered using CCCM tools, a process known as Filtered Template Matching Prediction (FTMP). The FTMP process is identical to the FIBC process, except that FTMP does not require a signaled block vector from the encoder to find the reference block. Instead, in FTMP, the reference block is determined on the decoder side by searching for the most similar L-shaped template in the reconstructed portion of the current frame, and this corresponding block is used as the reference block for the current block to be predicted. In other words, the L-shaped template associated with the reference block is the template most similar to the L-shaped template associated with the current block in the reconstructed portion of the frame. This operation for determining the reference block is the same as in intra-frame TMP. After determining the reference block, the same filtering process as in FIBC is applied to predict the target luma or chroma samples of the current block based on the corresponding luma or chroma samples of the reference block, respectively. For example, as... Figure 25 As shown, a cross-shaped filter can be applied to the corresponding sample values (sample values of the reference block and the boundary region of the reference block) to predict each of the sample values of the current block.
[0444] According to one or more embodiments of this disclosure, the same filter shape and / or template region can be applied to both prediction in FIBC mode and prediction in FTMP mode. For example, before deciding which mode to apply, the decoder and / or encoder can try both FIBC mode and FTMP mode with the same filter shape and / or template region to achieve better performance or reduced cost. Different methods can be used to achieve this goal.
[0445] In the first example, the filter operations proposed to be used in FTMP mode are also applied to FIBC mode. In one example, the 6-tap filter (a cross with 5 spatial components and a bias term) used in FTMP mode and the template region used for training (4 rows wide above and to the left of the current block in terms of samples, depending on their availability) can also be applied to FIBC mode in the same CU.
[0446] In the second example, the filter operations proposed to be used in FIBC mode are also applied to FTMP mode. In one example, the 2-tap filter (a single-sample filter with one spatial component and a bias term) used in FIBC mode and the template region used for training (one row above and to the left of the current block, depending on their availability) can also be applied to FTMP mode in the same CU.
[0447] Coordination of filters for FIBC, FTMP and CCCM modes
[0448] In Convolutional Cross-Component Model (CCCM) mode, filters are used to predict chroma sample values based on corresponding luma sample values. In CCCM mode, the set of chroma sample values in the reconstructed region of the current block to be predicted, along with the corresponding luma sample values, are used to determine the filter coefficients of the CCCM filter. In one example, the chroma sample values to be predicted and their corresponding luma sample values are isotopic sample values. Although the training results for the filter coefficients may differ, the same filter shape and / or template region can be reused in CCCM, FIBC, and FTMP for better performance / reduced cost. In one example, the filter shape and / or template region can be signaled to the decoder by the encoder. In another example, the filter shape and / or template region can be derived by the encoder based on predetermined rules (e.g., from a predefined candidate set).
[0449] According to one or more embodiments of this disclosure, the same filter shape and / or template area can be applied to at least two of the FIBC, FTMP, and CCCM modes. Different methods can be used to achieve this goal.
[0450] In the first example, the filter operations proposed to be used in CCCM mode are also applied to FIBC mode. In one example, the 7-tap filter (a cross with 5 spatial components, a nonlinear term, and a bias term) used in CCCM mode and the template region used for training (6 rows above and to the left of the current block, depending on their availability) can also be applied to FIBC mode in the same CU.
[0451] In the second example, it is proposed that one of the filter operations used in CCCM mode be applied to both FIBC and FTMP modes. In one example, an 11-tap filter (with 9 spatial components, a nonlinear term, and a bias term) is used for CCCM mode. The three squares) and the template region for training (six rows above and to the left of the current block, depending on their availability) can also be applied to FIBC mode and FTMP mode.
[0452] Figure 33 The workflow of a method 3300 for video decoding according to one or more aspects of this disclosure is shown.
[0453] At step 3310, method 3300 includes determining a reference block in a reconstructed portion of a video frame for predicting a current block in the video frame, wherein the L-shaped template associated with the reference block is the template most similar to the L-shaped template associated with the current block in the reconstructed portion of the video frame.
[0454] At step 3320, method 3300 includes obtaining a set of filter coefficients corresponding to the filter shape based on sample values from both the training region associated with the reference block and the training region associated with the current block.
[0455] At step 3330, method 3300 includes deriving predicted sample values for the current block based on a plurality of corresponding sample values associated with the reference block, using a set of filter coefficients and a filter shape.
[0456] At step 3340, method 3300 includes reconstructing the current block based on predicted sample values.
[0457] Figure 34 The workflow of a method 3400 for video encoding according to one or more aspects of this disclosure is shown.
[0458] At step 3410, method 3400 includes dividing a video frame into multiple blocks.
[0459] At step 3420, method 3400 includes determining a reference block in a reconstructed portion of a video frame for predicting a current block in the video frame, wherein the L-shaped template associated with the reference block is the template most similar to the L-shaped template associated with the current block in the reconstructed portion of the video frame.
[0460] At step 3430, method 3400 includes obtaining filter coefficients corresponding to the filter shape based on sample values from both the training region associated with the reference block and the training region associated with the current block.
[0461] At step 3440, method 3400 includes deriving predicted sample values for the current block based on multiple corresponding sample values associated with the reference block, using a set of filter coefficients and a filter shape.
[0462] At step 3450, method 3400 includes generating a bitstream based on predicted sample values.
[0463] Figure 35 The workflow of a method 3500 for video decoding according to one or more aspects of this disclosure is shown.
[0464] At step 3510, method 3500 includes determining at least one of a filter shape and a template region, wherein at least one of the filter shape and template region will be used in at least two of a filtered intra-block copy (FIBC) mode, a filtered template matching prediction (FTMP) mode, and a convolutional cross-component model (CCCM) mode to predict sample values of the current block in a video frame.
[0465] At step 3520, method 3500 includes using a template region to train each filter coefficient in a set of filter coefficients corresponding to the filter shape for at least two of the FIBC, FTMP, and CCCM modes.
[0466] At step 3530, method 3500 includes using a set of filter coefficients to derive the predicted sample values for the current block.
[0467] At step 3540, method 3500 includes reconstructing the current block based on predicted sample values.
[0468] In one example, training the filter coefficient set for FIBC mode includes: determining a reference block in the reconstructed portion of a video frame for predicting the current block, wherein the reference block is determined based on a block vector received from the bitstream; and obtaining the filter coefficient set for FIBC mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training regions associated with the reference block and the training regions associated with the current block are determined at least in part based on a template region for FIBC mode.
[0469] In one example, training the filter coefficient set for FTMP mode includes: determining a reference block in the reconstructed portion of a video frame for predicting the current block, wherein the L-shaped template associated with the reference block is the template most similar to the L-shaped template associated with the current block in the reconstructed portion of the video frame; and obtaining the filter coefficient set for FTMP mode based on sample values from both the training region associated with the reference block and the training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined at least in part based on the template region for FTMP mode.
[0470] In one example, training the filter coefficient set for CCCM mode includes: determining the set of chroma sample values in the template region; and obtaining the filter coefficient set for CCCM mode based on the set of chroma sample values and the corresponding luminance sample values for the set of chroma sample values.
[0471] In one example, the filter shape includes at least one of the following: a cross-shaped filter shape corresponding to 5 spatial terms, nonlinear terms, and bias terms; a single-sample filter shape corresponding to 1 spatial term and bias term; or a 3-sampling filter shape corresponding to 9 spatial terms, nonlinear terms, and bias terms. 3. Square filter shape.
[0472] In one example, the template area includes at least one of the following: four lines above and to the left of the current block; one line above and to the left of the current block; or six lines above and to the left of the current block.
[0473] Figure 36 The workflow of a method 3600 for video encoding according to one or more aspects of this disclosure is shown.
[0474] At step 3610, method 3600 includes dividing a video frame into multiple blocks.
[0475] At step 3620, method 3600 includes determining at least one of a filter shape and a template region, wherein at least one of the filter shape and template region will be used in at least two of a filtered intra-block copy (FIBC) mode, a filtered template matching prediction (FTMP) mode, and a convolutional cross-component model (CCCM) mode to predict sample values of the current block in a video frame.
[0476] At step 3630, method 3600 includes using a template region to train each filter coefficient in a set of filter coefficients corresponding to the filter shape for at least two of the FIBC, FTMP and CCCM modes respectively.
[0477] At step 3640, method 3600 includes deriving the predicted sample values of the current block using the filter coefficient set.
[0478] At step 3650, method 3600 includes generating a bitstream based on predicted sample values.
[0479] In one example, training a set of filter coefficients for FIBC mode includes: determining a reference block in the reconstructed portion of a video frame for predicting the current block, wherein the reference block is determined based on a block vector to be transmitted via a bitstream; and obtaining a set of filter coefficients for FIBC mode based on sample values from both a training region associated with the reference block and a training region associated with the current block, wherein the training regions associated with the reference block and the training regions associated with the current block are determined at least in part based on a template region for FIBC mode.
[0480] In one example, training the filter coefficient set for FTMP mode includes: determining a reference block in the reconstructed portion of a video frame for predicting the current block, wherein the L-shaped template associated with the reference block is the template most similar to the L-shaped template associated with the current block in the reconstructed portion of the video frame; and obtaining the filter coefficient set for FTMP mode based on sample values from both the training region associated with the reference block and the training region associated with the current block, wherein the training region associated with the reference block and the training region associated with the current block are determined at least in part based on the template region for FTMP mode.
[0481] In one example, training the filter coefficient set for CCCM mode includes: determining the set of chroma sample values in the template region; and obtaining the filter coefficient set for CCCM mode based on the set of chroma sample values and the corresponding luminance sample values for the set of chroma sample values.
[0482] In one example, the filter shape includes at least one of the following: a cross-shaped filter shape corresponding to 5 spatial terms, nonlinear terms, and bias terms; a single-sample filter shape corresponding to 1 spatial term and bias term; or a 3-sampling filter shape corresponding to 9 spatial terms, nonlinear terms, and bias terms. 3. Square filter shape.
[0483] In one example, the template area includes at least one of the following: four lines above and to the left of the current block; one line above and to the left of the current block; or six lines above and to the left of the current block.
[0484] Reference region and padding process in filtered intra-block copy (FIBC)
[0485] According to one or more embodiments of this disclosure, filter coefficients are calculated by minimizing the MSE between predicted and reconstructed luminance and / or chrominance samples in a reference region. In one example, Figure 25 The reference area includes the luminance / chromaticity samples above and to the left of the CU. An extension of the area shown in blue (slashed portion) is needed to support the "side samples" of the plus spatial filter, and this can be achieved using different methods.
[0486] In the first method, it is proposed to fill in the missing sample by using the nearest available sample when the sample is unavailable.
[0487] The second method proposes to fill with the nearest available sample point regardless of whether the nearest available sample point is unavailable.
[0488] According to one or more embodiments of this disclosure, the reference region may extend to the right by a CU width and extend below the CU boundary by a CU height.
[0489] According to one or more embodiments of this disclosure, the reference region can be adjusted to include only available samples. In one example, the reference region includes N rows of luminance / chrominance samples above and to the left of the CU. N is an integer and / or has a maximum upper limit (e.g., 4, 5, 6, 7).
[0490] Adaptive reordering of merging candidates using filtered intra-block copying (FIBC)
[0491] According to one or more embodiments of this disclosure, during the process of extending the ARMC-TM to the IBC merge list, when the merge candidates are predicted using FIBC, the reference samples of the templates for the merge candidates are also generated by FIBC.
[0492] According to one or more embodiments of this disclosure, filter coefficients are calculated by minimizing the MSE between predicted and reconstructed luminance samples and / or chrominance samples in a reference region.
[0493] According to one or more embodiments of this disclosure, the template size and shape may not be included in the reference samples of the merged candidate templates. In one example, Figure 37 The reference area includes three rows of brightness samples above and to the left of the CU.
[0494] According to one or more embodiments of this disclosure, during the process of extending the ARMC-TM to the IBC merge list, when the merge candidates are predicted using FIBC, reference samples of the templates for the merge candidates are generated by IBC-LIC to reduce complexity.
[0495] According to one or more embodiments of this disclosure, during the process of extending the ARMC-TM to the IBC merge list, when the merge candidates are predicted using FIBC, the reference samples of the template of the merge candidates are generated by IBC without filtering (i.e., unfiltered IBC) to reduce complexity.
[0496] Merging candidates using filtered intra-block copying
[0497] According to one or more embodiments of this disclosure, the IBC predictions from the merged candidates are further filtered. Different methods can be used to achieve this objective.
[0498] According to one or more embodiments of this disclosure, an enable flag can be signaled in the bitstream to indicate the FIBC merging mode used. The enable flag can be signaled at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0499] According to one or more embodiments of this disclosure, a mode flag can be derived at the decoder to save bit overhead, rather than explicitly signaling the selected mode flag.
[0500] Direct block vector of chroma block copied using filtered intra-block format
[0501] According to one or more embodiments of this disclosure, the direct block vector of the chroma block is further filtered. Different methods can be used to achieve this objective.
[0502] According to one or more embodiments of this disclosure, an enable flag can be signaled in the bitstream to indicate the FIBC merging mode used. The enable flag can be signaled at the SPS / DPS / VPS / SEI / APS / PPS / PH / SH / region / CTU / CU / subblock / sample level.
[0503] According to one or more embodiments of this disclosure, the mode flag can be inherited from the luma block at the decoder to save bit overhead, rather than explicitly signaling the selected mode flag.
[0504] Figure 38 The flowchart of a method 3800 for video decoding according to one or more aspects of this disclosure is shown. Method 3800 may be provided by a decoder (e.g., Figure 3 The video decoder 30) is executed.
[0505] At step 3810, a bitstream including video frames may be received, and in order to predict the current block in the video frame, a reference block in the same video frame may be determined. For example, the reference block may be determined using a block vector indicated in the bitstream used for the current block. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has already been reconstructed within the current video frame.
[0506] At step 3820, a set of filter coefficients corresponding to the filter shape can be obtained by training sample values of the template region associated with the reference block and the template region associated with the current block. The template region associated with the reference block can be expanded based on the filter shape, for example, to further include sample values of edge points of the filter shape.
[0507] In one example, the template region associated with the reference block, which includes samples to the left and above the reference block, can be extended to further include one or more lines to the right of the reference block and one or more lines below the reference block.
[0508] In another example, the template region associated with the reference block may extend from the reference block in four directions, namely, to the left and right, up and down, to further include one or more additional lines from the four sides surrounding the reference block.
[0509] In one or more examples, the extended row or rows may correspond to a CU width or a CU height. The sample values in the extended rows can be filled with the nearest available sample value. In another example, the sample filled with the nearest available sample value is an unavailable sample, such as a sample that has not yet been reconstructed to the right or below the reference block.
[0510] In another example, the template region associated with the reference block can be expanded to include only available samples. For example, the template region associated with the reference block is expanded to include up to N rows to the left of the reference block and up to N rows above the reference block, where N is an integer (e.g., 4, 5, 6, 7).
[0511] At step 3830, each of the predicted sample values for the current block can be derived based on a set of filter coefficients and a filter shape, using multiple corresponding sample values associated with the reference block. For example, the multiple corresponding sample values associated with the reference block can be used as input to the obtained filter to output each of the predicted sample values for the current block. The input sample values associated with the reference block may include sample values in the extended template region.
[0512] At step 3840, the current block can be reconstructed based on the predicted sample values. For example, the current block can be reconstructed by combining the predicted sample values with the residuals carried in the bitstream.
[0513] Figure 39 The flowchart of method 3900 for video encoding according to one or more aspects of this disclosure is shown. Method 3900 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 3900 can be the corresponding steps of method 3800.
[0514] At step 3910, a reference block in the video frame captured by the camera can be determined for predicting the current block in the same video frame. For example, this can be determined based on block matching (BM) performed at the encoder.
[0515] At step 3920, a filter with a set of filter coefficients and a filter shape can be obtained by training the sample values of the template region associated with the reference block and the template region associated with the current block, wherein the template region associated with the reference block is extended based on the filter shape. For example, the template region associated with the reference block is extended in the same manner as described in reference method 3800.
[0516] At step 3930, each of the predicted sample values of the current block can be derived using the obtained filter based on multiple corresponding sample values associated with the reference block.
[0517] At step 3940, a bitstream can be generated based on the predicted sample values.
[0518] Figure 40 The workflow of a method 4000 for video decoding according to one or more aspects of this disclosure is illustrated. Method 4000 may be provided by a decoder (e.g., Figure 3 The video decoder 30) is executed.
[0519] At step 4010, a merge candidate list for intra-block copy (IBC) prediction of the current block can be obtained. For example, the merge candidate list for IBC prediction can be constructed according to the above description or other criteria. The merge candidate list may include multiple candidates that have already been encoded using IBC, each candidate having, for example, a block vector.
[0520] At step 4020, the plurality of candidates may be reordered based on the template matching score of each of the plurality of candidates in the merged candidate list. The template matching score is calculated based on the difference between the sample value of the candidate's template and the corresponding reference sample value of the reference template of the reference block pointed to by the block vector of the candidate. In the calculation of the template matching score, in response to determining that the candidate is encoded and decoded using filtered IBC (i.e., FIBC), the corresponding reference sample value of the reference template is also obtained using filtered IBC (i.e., FIBC).
[0521] At step 4030, the current block can be reconstructed based on the reordered list of merge candidates.
[0522] In one example, the IBC prediction for the current block can be obtained from the candidates in the merged candidate list, for example, by using the block vector of that candidate. A determination is then made regarding whether the IBC prediction for the current block has been filtered.
[0523] In one example, the determination is based on signaling transmitted in the bitstream or inference derived at the decoder. If a determination is made regarding filtering the IBC prediction for the current block, then FIBC is performed on the IBC prediction.
[0524] In one example, at least a portion of the template and at least a portion of the reference template are not used to obtain a template matching score. For example, such as Figure 37 As shown in the grid, samples adjacent to the reference block and candidates are not used to calculate the template matching score.
[0525] Figure 41 The workflow of a method 4100 for video encoding according to one or more aspects of this disclosure is illustrated. Method 4100 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 4100 can be the corresponding steps of method 4000.
[0526] At step 4110, a merge candidate list for intra-block copy (IBC) prediction of the current block can be obtained. For example, the merge candidate list for IBC prediction can be constructed according to the above description or other criteria. The merge candidate list may include multiple candidates already encoded with IBC, each candidate having, for example, a block vector.
[0527] At step 4120, the plurality of candidates can be reordered based on the template matching score of each of the plurality of candidates in the merged candidate list. The template matching score is calculated based on the difference between the sample value of the candidate's template and the corresponding reference sample value of the reference template of the reference block pointed to by the block vector of the candidate. In the calculation of the template matching score, in response to determining that the candidate is encoded and decoded using filtered IBC (i.e., FIBC), the corresponding reference sample value of the reference template is also obtained using filtered IBC (i.e., FIBC).
[0528] At step 4130, a bitstream can be generated by encoding the current block based on a reordered list of merge candidates.
[0529] In one example, the IBC prediction for the current block can be obtained from the candidates in the merge candidate list, for example, by using the block vector of that candidate. A determination is made regarding whether the IBC prediction for the current block should be filtered. If a determination is made regarding filtering the IBC prediction for the current block, FIBC is performed on the IBC prediction.
[0530] In one example, signaling indicating the determination can be transmitted in the bit stream.
[0531] Figure 42 The workflow of a method 4200 for video decoding according to one or more aspects of this disclosure is illustrated. Method 4200 can be provided by a decoder (e.g., Figure 3 The video decoder 30) is executed.
[0532] At step 4210, the block vector of the chroma block from the bitstream can be determined by using the block vector of the luma block, which is IBC-encoded and associated with the chroma block. For example, the block vector of the chroma block can be directly inherited from the luma block.
[0533] At step 4220, the IBC prediction of the chroma block can be obtained based on the inherited block vector.
[0534] At step 4230, a determination is made regarding the filtering of the IBC prediction for the chroma block, i.e., using FIBC.
[0535] In one example, the determination can be made based on signaling in the bitstream.
[0536] In another example, the mode to be used by a chroma block (e.g., FIBC or IBC) can be directly inherited from the mode used by an associated luma block (e.g., FIBC or IBC), making explicit signaling possible.
[0537] At step 4240, the chroma block can be reconstructed based on the filtered IBC prediction.
[0538] In one example, the filter shape and / or filter coefficients of the chroma block can be computed by using a template of the chroma block at the decoder (e.g., in a similar manner to the luma block).
[0539] In another example, the filter shape and / or filter coefficients of the chroma block can be directly inherited from the filter shape and / or filter coefficients of the associated luma block.
[0540] Figure 43 The workflow of a method 4300 for video encoding according to one or more aspects of this disclosure is illustrated. Method 4300 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 4300 can be the corresponding steps of method 4200.
[0541] At step 4310, the block vector of the chroma block can be determined based on the luma block that is encoded and decoded using intra-block copy (IBC) and associated with the chroma block.
[0542] At step 4320, IBC prediction of the chroma block can be obtained based on the block vector.
[0543] At step 4330, the IBC prediction of the chroma block can be filtered.
[0544] In one example, signaling can be generated that indicates the IBC prediction of the chroma block will be filtered.
[0545] In another example, an IBC prediction that indicates the chroma block will be filtered by the FIBC mode used by the associated luma block can be generated.
[0546] At step 4340, a bitstream can be generated based on filtered IBC prediction.
[0547] In one example, signaling indicating that the IBC prediction of the chroma block will be filtered can be transmitted in the bitstream.
[0548] In another example, if the signaling indicating that the IBC prediction of the chroma block will be filtered can be directly inherited from the mode used by the associated luma block, then the signaling can be omitted from the bitstream.
[0549] Figure 44 The workflow of a method 4400 for video decoding according to one or more aspects of this disclosure is illustrated. Method 4400 may be provided by a decoder (e.g., Figure 3The video decoder 30) is executed. Method 4400 can be combined with methods 3300, 3500, 3800, 4000, and 4200, and the above description of methods 3300, 3500, 3800, 4000, and 4200 can be applied at least in part to some steps of the methods.
[0550] At step 4410, method 4400 may determine a reference block in a video frame from the bitstream for use in predicting the current block in the video frame. In one example, the reference block may be determined based on merge candidates.
[0551] At step 4420, method 4400 may obtain a set of filter coefficients corresponding to the filter shape based at least on sample values from both the template region associated with the reference block and the template region associated with the current block. In one example, the filter shape may be a rectangle without a bottom-right sample, as described above. Figure 30 The template region associated with the reference block and the template region associated with the current block are sample regions used to train filter coefficient sets in the vicinity of the reference block and the current block (e.g., adjacent or not adjacent to the current block), respectively. The template regions associated with the reference block and the template regions associated with the current block can be the same as the template regions used for template matching. In one example, the template region associated with the current block and the corresponding template region associated with the reference block can be L-shaped or rectangular with the size described above.
[0552] At step 4430, method 4400 can derive the predicted sample values of the current block using the filter coefficient set and filter shape, at least based on multiple predicted sample values of the current block. For example, as Figure 26 As shown in Example 2, the filter shape is Figure 30 In shape 4, the predicted sample points in the region with sparse points can be multiple predicted sample point values used to derive the predicted sample point values (as shown by dense points) for the location to be predicted, and the location to be predicted can be the predicted sample point values to be derived in step 4430.
[0553] In one example, such as Figure 26 As shown in Example 1, based on the position to be predicted in the current block and the filter shape, further prediction can be made based on the current block (e.g., Figure 26 The method 4400 derives the predicted sample value of the current block from multiple reconstructed sample values of adjacent reconstructed regions (small blocks without patterns in the reconstructed region). That is, when the location to be predicted is close to the edge of the current block and the filter used to predict this location covers samples outside the current block due to the filter shape, the method 4400 can derive the predicted sample value of the location to be predicted in the current block based on both multiple predicted sample values of the current block and multiple reconstructed sample values of the reconstructed regions adjacent to the current block.
[0554] In another example, method 4400 may further derive the predicted sample values of the current block based on multiple corresponding reconstructed sample values associated with the reference block, as described in steps 3300 and 3800. The multiple predicted sample values of the current block and the multiple corresponding reconstructed sample values associated with the reference block may be weighted based on a weighting factor to derive the predicted sample values of the current block. For example, method 4400 may derive the predicted sample values of the current block based on the weighted meaning of each of the multiple reconstructed sample values associated with the reference block and each corresponding predicted sample value associated with the current block and / or the corresponding reconstructed sample value.
[0555] At step 4440, method 4400 may reconstruct the current block based on the predicted sample values.
[0556] Figure 45 The workflow of a method 4500 for video encoding according to one or more aspects of this disclosure is illustrated. Method 4500 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 4500 may be corresponding steps of method 4400, therefore the above description of method 4400 and methods 3400, 3600, 3900, and 4100 may be applied at least in part to some steps of the method.
[0557] At step 4510, method 4500 may determine a reference block in the video frame for use in predicting the current block in the video frame.
[0558] At step 4520, method 4500 may obtain a set of filter coefficients corresponding to the filter shape based at least on sample values from both the template region associated with the reference block and the template region associated with the current block.
[0559] At step 4530, method 4500 can derive the predicted sample values of the current block using the filter coefficient set and filter shape based at least on multiple predicted sample values of the current block.
[0560] In one example, method 4500 can also derive the predicted sample values of the current block based on multiple reconstructed sample values of the reconstructed regions adjacent to the current block, according to the location to be predicted in the current block and the filter shape.
[0561] In another example, method 4500 may further derive the predicted sample values of the current block based on multiple corresponding reconstructed sample values associated with the reference block. The multiple predicted sample values of the current block and the multiple corresponding reconstructed sample values associated with the reference block may be weighted based on a weighting factor to derive the predicted sample values of the current block.
[0562] At step 4540, method 4500 can generate a bitstream based on the predicted sample values.
[0563] Figure 46 The flowchart of a method 4600 for video decoding according to one or more aspects of this disclosure is shown. Method 4600 may be provided by a decoder (e.g., Figure 3 The video decoder 30) is executed. Method 4600 can be combined with methods 3300, 3500, 3800, 4000, 4200, and 4400, and the above description of methods 3300, 3500, 3800, 4000, 4200, and 4400 can be applied at least in part to some steps of the methods.
[0564] At step 4610, method 4600 may determine a reference block in a video frame from the bitstream for use in predicting the current block in the video frame. In one example, the reference block may be determined based on merge candidates.
[0565] At step 4620, method 4600 may obtain a set of filter coefficients corresponding to a filter shape based at least on sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the filter shape is a rectangle with a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted. The filter shape may be a rectangle with a width greater than its height, a rectangle with a height greater than its width, or a rectangle with a width equal to its height. Figure 29 As shown, the location to be predicted within the rectangle differs for different filter shapes. Figure 30 As shown, the filter shape can be a rectangle without a bottom right sample point, and the bottom right sample point can be the location to be predicted.
[0566] In one example, the template region associated with the current block may depend on the filter shape. For instance, if the height of the filter shape is greater than its width, for example... Figure 29 If shapes 2, 5, 6, 10, and 11 are present, then the template region associated with the current block includes several rows above the current block, for example, N rows (where N is an integer); or if the width of the filter shape is greater than the height of the filter shape, for example... Figure 29If shapes 7, 8, and 9 are given, then the template region associated with the current block includes several rows to the left of the current block, for example, N rows (where N is an integer). Similarly, the template region associated with the reference block, corresponding to the template region associated with the current block, can also depend on the filter shape. For example, if the height of the filter shape is greater than its width, then the template region associated with the reference block includes several rows above the reference block, for example, N rows (where N is an integer); or if the width of the filter shape is greater than its height, then the template region associated with the reference block includes several rows to the left of the current block, for example, N rows (where N is an integer).
[0567] At step 4630, method 4600 may derive the predicted sample values of the current block based on at least one of multiple predicted sample values of the current block and multiple corresponding reconstructed sample values associated with the reference block, utilizing a set of filter coefficients and a filter shape. In one example, method 4600 may derive the predicted sample values of the current block based on multiple corresponding reconstructed sample values associated with the reference block, as described above in steps 3330 or 3830. In another example, method 4600 may derive the predicted sample values of the current block based on multiple predicted sample values of the current block, or multiple predicted sample values of the current block and multiple corresponding reconstructed sample values associated with the reference block, as described above in step 4430.
[0568] At step 4640, method 4600 may reconstruct the current block based on the predicted sample values.
[0569] Figure 47 The workflow of a method 4700 for video encoding according to one or more aspects of this disclosure is illustrated. Method 4700 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 4700 may be corresponding steps of method 4600, and therefore the above description of method 4600 and methods 3400, 3600, 3900, and 4100 may be applied at least in part to some steps of method 4700.
[0570] At step 4710, method 4700 may determine a reference block in the video frame for use in predicting the current block in the video frame.
[0571] At step 4720, method 4700 may obtain a set of filter coefficients corresponding to a filter shape based at least on sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the filter shape is a rectangle with a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted. The filter shape may be a rectangle with a width greater than its height, a rectangle with a height greater than its width, or a rectangle with a width equal to its height. The position to be predicted within the rectangle may differ for different filter shapes. The filter shape may also be a rectangle without a bottom-right sample, and the bottom-right sample may be the position to be predicted.
[0572] In one example, the template region associated with the reference block and the template region associated with the current block depend on the filter shape. The template region associated with the current block includes: multiple rows above the current block if the height of the filter shape is greater than the width of the filter shape; or multiple rows to the left of the current block if the width of the filter shape is greater than the height of the filter shape. The template region associated with the reference block includes: multiple rows above the reference block if the height of the filter shape is greater than the width of the filter shape; or multiple rows to the left of the reference block if the width of the filter shape is greater than the height of the filter shape.
[0573] At step 4730, method 4700 may derive the predicted sample values of the current block based on at least one of multiple predicted sample values of the current block and multiple corresponding reconstructed sample values associated with the reference block, using a set of filter coefficients and a filter shape.
[0574] At step 4740, method 4700 can generate a bitstream based on the predicted sample values.
[0575] Figure 48 The workflow of a method 4800 for video decoding according to one or more aspects of this disclosure is illustrated. Method 4800 may be provided by a decoder (e.g., Figure 3 The video decoder 30) is executed. Method 4800 can be combined with methods 3300, 3500, 3800, 4200, 4400, and 4600, and the above description of methods 3300, 3500, 3800, 4200, 4400, and 4600 can be applied at least in part to some steps of the methods. Step 4820 of method 4800 can be different from step 4020 of method 4000, and the above description of steps 4010 and 4030 of method 4000 can be applied to the methods.
[0576] At step 4810, method 4800 may obtain a merge candidate list for intra-block copy (IBC) prediction of the current block, the merge candidate list including multiple candidates that have been encoded with IBC.
[0577] At step 4820, method 4800 may reorder the plurality of candidates based on the template matching score of each of the plurality of candidates in the merged candidate list, the template matching score being obtained based on the difference between the sample value of a candidate's template and the corresponding reference sample value of a reference template of the candidate's reference block, wherein, in response to determining that a candidate is encoded and decoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC (i.e., unfiltered IBC). In one example, unfiltered IBC may include IBC Local Illumination Compensation (LIC). That is, in response to determining that a candidate is encoded and decoded using filtered IBC, the corresponding reference sample value of the reference template can be obtained using IBC-LIC.
[0578] At step 4830, method 4800 may reconstruct the current block based on a reordered list of merge candidates. In one example, FIBC may be used to decode the current block, for example, according to methods 3300, 3500, 3800, 4200, 4400, and 4600 described above. FIBC flags indicating the use of FIBC to encode the current block may be inherited from the merge candidates. A merge candidate may be selected from the reordered list of merge candidates. The merge candidate may be one of the following: a motion vector predictor (HMVP) candidate based on IBC history, an IBC spatial MVP candidate from spatially adjacent coding units, and an IBC temporal MVP candidate from temporally co-located coding units.
[0579] Figure 49 The workflow of a method 4900 for video encoding according to one or more aspects of this disclosure is illustrated. Method 4900 may be generated by an encoder (e.g., Figure 2 The video decoder 20) is executed. The steps of method 4900 can be corresponding steps of method 4800; therefore, the above descriptions of method 4800 and methods 3400, 3600, 3900, 4300, and 4500 can be applied at least partially to some steps of method 4900. Step 4920 of method 4900 can differ from step 4120 of method 4100, and the above descriptions of steps 4110 and 4130 of method 4100 can be applied to method 4900.
[0580] At step 4910, method 4900 may obtain a merge candidate list for intra-block copy (IBC) prediction of the current block, the merge candidate list including multiple candidates that have been encoded with IBC.
[0581] At step 4920, method 4900 may reorder the plurality of candidates based on the template matching score of each of the plurality of candidates in the merged candidate list, the template matching score being obtained based on the difference between the sample value of a candidate's template and the corresponding reference sample value of a reference template of the candidate's reference block, wherein, in response to determining that a candidate is encoded and decoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC (i.e., unfiltered IBC). In one example, unfiltered IBC may include IBC Local Illumination Compensation (LIC). That is, in response to determining that a candidate is encoded and decoded using filtered IBC, the corresponding reference sample value of the reference template can be obtained using IBC-LIC.
[0582] At step 4930, method 4900 can generate a bitstream by encoding the current block based on a reordered list of merge candidates. In one example, FIBC can be used to encode the current block, for example, according to methods 3400, 3600, 3900, 4300, 4500, and 4700 described above. FIBC flags indicating the use of FIBC for encoding the current block can be inherited from the merge candidates. Merge candidates can be selected from the reordered list of merge candidates. Merge candidates can be one of the following: IBC history-based motion vector predictor (HMVP) candidates, IBC spatial MVP candidates from spatially adjacent coding units, and IBC temporal MVP candidates from temporally co-located coding units.
[0583] Figure 50 A computing environment 5010 coupled to a user interface 5050 is shown. The computing environment 5010 may be part of a data processing server. The computing environment 5010 includes a processor 5020, memory 5030, and input / output (I / O) interface 5040.
[0584] Processor 5020 typically controls the overall operation of computing environment 5010, such as operations associated with display, data acquisition, data communication, and image processing. Processor 5020 may include one or more processors to execute instructions to perform all or some of the steps described above. Furthermore, processor 5020 may include one or more modules that facilitate interaction between processor 5020 and other components. The processor may be a central processing unit (CPU), microprocessor, single-chip machine, graphics processing unit (GPU), etc.
[0585] Memory 5030 is configured to store various types of data to support the operation of computing environment 5010. Memory 5030 may include predefined software 5032. Examples of such data include any application or method instructions for operation on computing environment 5010, video datasets, image data, etc. Memory 5030 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0586] I / O interface 5040 provides an interface between processor 5020 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 5040 can be coupled to encoders and decoders.
[0587] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a memory 5030 containing a plurality of programs and / or storing a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by a processor 5020 in a computing environment 5010 to perform the above-described methods. In one example, the plurality of programs can be executed by a processor 5020 in a computing environment 5010 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 5010 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 5020 in the computing environment 5010 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 5020 in the computing environment 5010 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 5020 in the computing environment 5010 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.
[0588] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.
[0589] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 5020); and a non-transitory computer-readable storage medium or memory 5030 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0590] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 5030 is also provided, the plurality of programs being executable by a processor 5020 in a computing environment 5010 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0591] In an embodiment, the computing environment 5010 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0592] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.
[0593] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.
[0594] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0595] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0596] The examples chosen and described are intended to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A method for video decoding, comprising: A reference block is determined in a video frame from a bitstream to be used to predict the current block in the video frame; A set of filter coefficients corresponding to the filter shape is obtained based at least on sample values from both the template region associated with the reference block and the template region associated with the current block. Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on at least a plurality of predicted sample values of the current block; as well as The current block is reconstructed based on the predicted sample values.
2. The method according to claim 1, wherein, The process of deriving the predicted sample values for the current block includes: Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on the multiple predicted sample values of the current block and the multiple reconstructed sample values of the reconstructed region adjacent to the current block, according to the position to be predicted in the current block and the filter shape.
3. The method according to claim 1, wherein, The process of deriving the predicted sample values for the current block includes: Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on the plurality of predicted sample values of the current block and the plurality of corresponding reconstructed sample values associated with the reference block.
4. The method according to claim 3, wherein, The plurality of predicted sample values of the current block and the plurality of corresponding reconstructed sample values associated with the reference block are weighted based on a weighting factor to derive the predicted sample values of the current block.
5. A method for video encoding, comprising: A reference block in the video frame is determined for use in predicting the current block in the video frame; A set of filter coefficients corresponding to the filter shape is obtained based at least on sample values from both the template region associated with the reference block and the template region associated with the current block. Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on at least a plurality of predicted sample values of the current block; as well as A bitstream is generated based on the predicted sample values.
6. The method according to claim 5, wherein, The process of deriving the predicted sample values for the current block includes: Using the filter coefficient set and the filter shape, based on the position to be predicted in the current block and the filter shape, and based on the multiple predicted sample values of the current block and the multiple reconstructed sample values of the reconstructed region adjacent to the current block, the predicted sample values of the current block are derived.
7. The method according to claim 5, wherein, The process of deriving the predicted sample values for the current block includes: Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on the plurality of predicted sample values of the current block and the plurality of corresponding reconstructed sample values associated with the reference block.
8. The method according to claim 7, wherein, The plurality of predicted sample values of the current block and the plurality of corresponding reconstructed sample values associated with the reference block are weighted based on a weighting factor to derive the predicted sample values of the current block.
9. A method for video decoding, comprising: A reference block is determined in a video frame from a bitstream to be used to predict the current block in the video frame; A set of filter coefficients corresponding to a filter shape is obtained based at least on the sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the filter shape is a rectangle with a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted. Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on at least one of multiple predicted sample values of the current block and multiple corresponding reconstructed sample values associated with the reference block; and The current block is reconstructed based on the predicted sample values.
10. The method according to claim 9, wherein, The filter shape is a rectangle with a width greater than its height, a rectangle with a height greater than its width, or a rectangle with a width equal to its height.
11. The method according to claim 9, wherein, The position to be predicted within the rectangle is different for different filter shapes.
12. The method according to claim 9, wherein, The filter shape is a rectangle without a bottom right sample point, and the bottom right sample point is the location to be predicted.
13. The method according to claim 9, wherein, The template region associated with the reference block and the template region associated with the current block depend on the filter shape.
14. The method according to claim 13, wherein, The template region associated with the current block includes: If the height of the filter shape is greater than the width of the filter shape, then it consists of multiple rows above the current block; or If the width of the filter shape is greater than the height of the filter shape, then it is multiple rows of the row to the left of the current block.
15. The method according to claim 14, wherein, The template region associated with the reference block includes: If the height of the filter shape is greater than the width of the filter shape, then it consists of multiple rows above the reference block; or If the width of the filter shape is greater than the height of the filter shape, then it is multiple rows on the left side of the reference block.
16. A method for video encoding, comprising: A reference block in the video frame is determined for use in predicting the current block in the video frame; A set of filter coefficients corresponding to a filter shape is obtained based at least on the sample values from both the template region associated with the reference block and the template region associated with the current block, wherein the filter shape is a rectangle with a width of N rows and a height of M rows, where N and M are integers greater than 1, and the filter shape is identified using the position to be predicted. Using the filter coefficient set and the filter shape, the predicted sample values of the current block are derived based on at least one of multiple predicted sample values of the current block and multiple corresponding reconstructed sample values associated with the reference block; and A bitstream is generated based on the predicted sample values.
17. The method according to claim 16, wherein, The filter shape is a rectangle with a width greater than its height, a rectangle with a height greater than its width, or a rectangle with a width equal to its height.
18. The method according to claim 16, wherein, The position to be predicted within the rectangle is different for different filter shapes.
19. The method of claim 16, wherein, The filter shape is a rectangle without a bottom right sample point, and the bottom right sample point is the location to be predicted.
20. The method of claim 16, wherein, The template region associated with the reference block and the template region associated with the current block depend on the filter shape.
21. The method according to claim 20, wherein, The template region associated with the current block includes: If the height of the filter shape is greater than the width of the filter shape, then it consists of multiple rows above the current block; or If the width of the filter shape is greater than the height of the filter shape, then it is multiple rows to the left of the current block.
22. The method according to claim 21, wherein, The template region associated with the reference block includes: If the height of the filter shape is greater than the width of the filter shape, then it consists of multiple rows above the reference block; or If the width of the filter shape is greater than the height of the filter shape, then it is multiple rows on the left side of the reference block.
23. A method for video decoding, comprising: Obtain a merge candidate list of intra-block copy (IBC) predictions for the current block, the merge candidate list including multiple candidates encoded using IBC; Based on the template matching score of each of the plurality of candidates in the merged candidate list, the plurality of candidates are reordered, wherein the template matching score is obtained based on the difference between the sample value of the candidate's template and the corresponding reference sample value of the reference template of the candidate's reference block, wherein, in response to determining that the candidate is encoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC; as well as The current block is reconstructed based on the reordered list of merge candidates.
24. The method according to claim 23, wherein, The unfiltered IBC includes IBC Local Illumination Compensation (LIC).
25. The method according to claim 23, wherein, The current block is decoded using FIBC, and the FIBC flag is inherited from the merge candidate.
26. The method of claim 25, wherein, The merged candidate is one of the following: a motion vector predictor (HMVP) candidate based on IBC history, an IBC spatial MVP candidate from spatially adjacent coding units, and an IBC temporal MVP candidate from temporally co-located coding units.
27. A method for video encoding, comprising: Obtain a merge candidate list of intra-block copy (IBC) predictions for the current block, the merge candidate list including multiple candidates encoded using IBC; Based on the template matching score of each of the plurality of candidates in the merged candidate list, the plurality of candidates are reordered, wherein the template matching score is obtained based on the difference between the sample value of the candidate's template and the corresponding reference sample value of the reference template of the candidate's reference block, wherein, in response to determining that the candidate is encoded using filtered IBC (FIBC), the corresponding reference sample value of the reference template is obtained using unfiltered IBC; as well as A bitstream is generated by encoding the current block based on a reordered list of merge candidates.
28. The method according to claim 27, wherein, The unfiltered IBC includes IBC Local Illumination Compensation (LIC).
29. The method according to claim 27, wherein, The current block is encoded using FIBC, and the FIBC flag is inherited from the merge candidate.
30. The method according to claim 29, wherein, The merged candidate is one of the following: a motion vector predictor (HMVP) candidate based on IBC history, an IBC spatial MVP candidate from spatially adjacent coding units, and an IBC temporal MVP candidate from temporally co-located coding units.
31. An apparatus comprising: One or more processors; as well as One or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operation of the method described in any one of claims 1-30.
32. A computer program product storing computer-executable instructions, which, when executed, cause one or more processors to perform the operations of the method described in any one of claims 1-30.
33. A computer-readable storage medium storing instructions that, when executed by a computing device having one or more processors, cause the one or more processors to perform the following operations: Perform the decoding method according to any one of claims 1-4, 9-15, 23-26, and store the bitstream to be decoded by the decoding method according to any one of claims 1-4, 9-15, 23-26, or Perform the encoding method according to any one of claims 5-8, 16-22, 27-30, and store the bit stream generated by the encoding method according to any one of claims 5-8, 16-22, 27-30.
34. A computer-readable medium storing a bitstream, wherein The bitstream will be decoded by performing the operation described in any one of claims 1-4, 9-15, 23-26, or The bit stream is obtained by performing the operations described in any one of claims 5-8, 16-22, 27-30.
35. A method for receiving a bitstream to be decoded by the decoding method according to any one of claims 1-4, 9-15, 23-26.
36. A method for transmitting a bit stream generated by the encoding method according to any one of claims 5-8, 16-22, 27-30.