Method and apparatus for intra block copy
By employing context-adaptive binary arithmetic coding and fractional motion information processing in video coding, the intra-frame block copying process is optimized, solving the problem of low efficiency in existing technologies and achieving more efficient video coding performance.
Patent Information
- Application Number
- CN202480023410.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-29
- Filing Date
- 2024-03-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video coding technologies suffer from inefficiency and poor coding performance during intra-frame block copying, especially when dealing with different stripe types, frame types, and video resolutions, lacking effective context-adaptive binary arithmetic coding and motion information processing methods.
Context-adaptive binary arithmetic coding (CABAC) is used in intra-block copy (IBC) mode to adaptively code for different stripe types, frame types and video resolutions. Fractional motion information processing is used to optimize the determination of block vectors (BV) and the generation of predicted blocks, including the selection of interpolation filters and skipping process, as well as the processing of various merging modes and refinement positions.
It improves the efficiency and performance of video encoding, especially when dealing with different video resolutions and frame types, reducing encoding complexity and improving encoding quality.
Smart Images

Figure CN120982102A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to U.S. Provisional Application No. 63 / 455,544, filed on March 29, 2023, entitled “Methods and Devices for IntraBlock Copy,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This disclosure relates to video encoding, decoding, and compression, and specifically, but not limited to, methods and apparatus for improving intra-frame block copying methods in the video encoding or decoding process. Background Technology
[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices transmit and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0005] This disclosure provides examples of techniques related to improving intra-frame block copying methods in video encoding or decoding processes.
[0006] According to a first aspect of this disclosure, a method for video decoding is provided. In this method, the decoder can obtain a syntax element associated with an intra-block copy (IBC) mode and transmit the syntax element jointly or separately using context-adaptive binary arithmetic coding (CABAC), which uses different context bits for different stripe types, different frame types, different video components, or different video resolutions.
[0007] According to a second aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; obtain a prediction block of the current block based on the BV; and in response to a condition being satisfied, perform a padding process or partially or completely skip the padding process by the decoder on one or more samples associated with an interpolation filter.
[0008] According to a third aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; obtain a prediction block of the current block based on the BV; and in response to a condition being satisfied, determine an interpolation filter or skip an interpolation process, wherein the condition is configured to indicate whether to switch the interpolation filter or skip the interpolation process, and whether the interpolation filter is to be applied to the prediction block to obtain a filtered prediction for the current block.
[0009] According to a fourth aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain a plurality of base candidates from an intra block copy (IBC) merge list; obtain a plurality of merge mode with block vector difference (MBVD) refinement positions for each base candidate based on at least one candidate distance set; reorder the plurality of MBVD refinement positions; select a first number of MBVD refinement positions with lowest template SAD cost; and obtain a block vector difference based on the first number of MBVD refinement positions, wherein the at least one candidate distance set comprises any one or any combination of the following sets: {1 / 4-pixel, 1 / 2-pixel, 1-pixel, 2-pixel, 3-pixel, 4-pixel, 6-pixel, 8-pixel, 10-pixel, 12-pixel, 14-pixel, 16-pixel, 18-pixel, 20-pixel, 22-pixel, 24-pixel, 26-pixel, 28-pixel, 30-pixel, 32-pixel}; {1 / 2-pixel, 1-pixel, 2-pixel, 3-pixel, 4-pixel, 6-pixel, 8-pixel, 10-pixel, 12-pixel, 16-pixel, 20-pixel, 24-pixel, 28-pixel, 32-pixel, 36-pixel, 40-pixel, 44-pixel, 48-pixel, 52-pixel, 56-pixel}; {1 / 2-pixel, 1-pixel, 2-pixel, 3-pixel, 4-pixel, 6-pixel, 8-pixel, 10-pixel, 12-pixel, 14-pixel, 16-pixel, 18-pixel, 20-pixel, 22-pixel, 24-pixel, 26-pixel, 28-pixel, 30-pixel, 32-pixel, 34-pixel}; or {1-pixel, 2-pixel, 4-pixel, 6-pixel, 8-pixel, 12-pixel, 16-pixel, 20-pixel, 24-pixel, 32-pixel, 40-pixel, 48-pixel, 56-pixel, 64-pixel, 72-pixel, 80-pixel, 88-pixel, 96-pixel, 104-pixel, 112-pixel}.
[0010] According to a fifth aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; obtain a prediction block of the current block based on the BV; obtain a reference template based on the BV and a template of the current block; obtain parameters of a linear filter based on the reference template; and apply the linear filter to the prediction block to obtain a filtered prediction for the current block.
[0011] According to a sixth aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a plurality of luma blocks co-located at a plurality of predefined positions relative to a position of a current chroma block in an intra block copy (IBC) mode; obtain a plurality of luma block vectors (BVs) of the plurality of luma blocks based on the fractional motion information; determine a chroma BV based on the plurality of luma BVs of the plurality of luma blocks; and obtain a predicted chroma block based on the chroma BV.
[0012] According to a seventh aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; in response to applying a reconstruction reordering IBC (RRIBC) mode, set a horizontal component or a vertical component of the BV to zero by the decoder according to a flipping type of the RRIBC mode; and obtain a prediction block of the current block based on the BV.
[0013] According to an eighth aspect of the disclosure, a method for video decoding is provided. In the method, a decoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; in response to applying a block vector predictor candidate clustering and block vector difference sign derivation (IBC-BVPC) mode and an IBC advanced motion vector prediction (AMVP) mode for reconstruction reordering IBC, skip a motion refinement process for the BV, or skip a motion refinement process for a null component of the BV, the null component of the BV including a horizontal component or a vertical component; and obtain a prediction block of the current block based on the BV.
[0014] According to a ninth aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder can obtain magnitudes of block vector difference (BVD) components, each of the BVD components comprising an integer BVD component part and a fractional BVD component part; obtain a context coded BVD sign prediction index; obtain a plurality of BV candidates by creating a combination between a BVD sign and an absolute BVD value and adding the combination to a BV prediction value, each of the BV candidates comprising an integer part and a fractional part; derive a BVD sign prediction cost for each of the BV candidates based on a template matching cost; rank the plurality of BV candidates based on the BVD sign prediction cost; and derive the BVD sign prediction index corresponding to a true BVD sign.
[0015] According to a tenth aspect of the present disclosure, a method for video decoding is provided. In the method, a decoder can obtain a pairwise intra block copy (IBC) candidate based on at least two precious IBC candidates in a candidate list of an IBC merge mode or an IBC advanced motion vector prediction (AMVP) mode; and determine corresponding properties of the pairwise IBC candidate associated with properties of the two precious IBC candidates.
[0016] According to an eleventh aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder can obtain syntax elements related to an intra block copy (IBC) mode, and signal the syntax elements jointly or individually with context adaptive binary arithmetic coding (CABAC) using different context bins for different slice types, different frame types, different video components, or different video resolutions.
[0017] According to a twelfth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; obtain a prediction block of the current block based on the BV; and in response to a condition being satisfied, perform, by the encoder, a padding process on one or more samples associated with an interpolation filter, or partially or completely skip the padding process.
[0018] According to a thirteenth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; obtain a prediction block of the current block based on the BV; and in response to a condition being satisfied, determine an interpolation filter or skip an interpolation process, wherein the condition is configured to indicate whether to switch the interpolation filter or skip the interpolation process, and whether the interpolation filter is to be applied to the prediction block to obtain a filtered prediction for the current block.
[0019] According to a fourteenth aspect of the disclosure, a method for video coding is provided. In the method, an encoder can obtain a plurality of base candidates from an intra block copy (IBC) merge list; obtain a plurality of merge by disparity vector difference (MBVD) refinement positions for each base candidate based on at least one candidate distance set; reorder the plurality of MBVD refinement positions; select a first number of MBVD refinement positions with lowest template SAD cost; and obtain a block vector difference based on the first number of MBVD refinement positions, wherein the at least one candidate distance set comprises any one or any combination of the following sets: {1 / 4-pel, 1 / 2-pel, 1-pel, 2-pel, 3-pel, 4-pel, 6-pel, 8-pel, 10-pel, 12-pel, 14-pel, 16-pel, 18-pel, 20-pel, 22-pel, 24-pel, 26-pel, 28-pel, 30-pel, 32-pel}; {1 / 2-pel, 1-pel, 2-pel, 3-pel, 4-pel, 6-pel, 8-pel, 10-pel, 12-pel, 16-pel, 20-pel, 24-pel, 28-pel, 32-pel, 36-pel, 40-pel, 44-pel, 48-pel, 52-pel, 56-pel}; {1 / 2-pel, 1-pel, 2-pel, 3-pel, 4-pel, 6-pel, 8-pel, 10-pel, 12-pel, 14-pel, 16-pel, 18-pel, 20-pel, 22-pel, 24-pel, 26-pel, 28-pel, 30-pel, 32-pel, 34-pel}; or {1-pel, 2-pel, 4-pel, 6-pel, 8-pel, 12-pel, 16-pel, 20-pel, 24-pel, 32-pel, 40-pel, 48-pel, 56-pel, 64-pel, 72-pel, 80-pel, 88-pel, 96-pel, 104-pel, 112-pel}.
[0020] According to a fifteenth aspect of the disclosure, a method for video coding is provided. In the method, an encoder can obtain fractional motion information for a current block in an intra block copy (IBC) mode; obtain a block vector (BV) for the current block based on the fractional motion information; obtain a prediction block for the current block based on the BV; obtain a reference template based on the BV and a template of the current block; obtain parameters of a linear filter based on the reference template; and apply the linear filter to the prediction block to obtain a filtered prediction for the current block.
[0021] According to a sixteenth aspect of the present disclosure, a method for video coding is provided. In the method, an encoder can obtain fractional motion information of a plurality of luma blocks co-located at a plurality of predefined positions relative to a position of a current chroma block in an intra block copy (IBC) mode; obtain a plurality of luma block vectors (BVs) of the plurality of luma blocks based on the fractional motion information; determine a chroma BV based on the plurality of luma BVs of the plurality of luma blocks; and obtain a predicted chroma block based on the chroma BV.
[0022] According to a seventeenth aspect of the present disclosure, a method for video coding is provided. In the method, an encoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; in response to applying a reconstruction reordering IBC (RRIBC) mode, set a horizontal component or a vertical component of the BV to zero by the encoder according to a flipping type of the RRIBC mode; and obtain a predicted block of the current block based on the BV.
[0023] According to an eighteenth aspect of the present disclosure, a method for video coding is provided. In the method, an encoder can obtain fractional motion information of a current block in an intra block copy (IBC) mode; obtain a block vector (BV) of the current block based on the fractional motion information; in response to applying a block vector predictor candidate clustering and block vector difference sign derivation (IBC-BVPC) mode and an IBC advanced motion vector prediction (AMVP) mode for reconstruction reordering IBC, skip a fractional motion estimation process for the BV or skip a fractional motion estimation process for null components of the BV, the null components of the BV including a horizontal component or a vertical component; and obtain a predicted block of the current block based on the BV.
[0024] According to a nineteenth aspect of the present disclosure, a method for video coding is provided. In the method, an encoder can obtain magnitudes of block vector difference (BVD) components, each of the BVD components including an integer BVD component part and a fractional BVD component part; obtain a context coded BVD sign predictor index; obtain a plurality of BV candidates by creating a combination between a BVD sign and an absolute BVD value and adding the combination to a BV predictor, each of the BV candidates including an integer part and a fractional part; derive a BVD sign predictor cost of each of the BV candidates based on a template matching cost; rank the plurality of BV candidates based on the BVD sign predictor cost; and derive the BVD sign predictor index corresponding to a true BVD sign.
[0025] According to a twentieth aspect of the present disclosure, a method for video encoding is provided. In the method, an encoder can obtain a pairwise intra block copy (IBC) candidate based on at least two precious IBC candidates in a candidate list of an IBC merge mode or an IBC advanced motion vector prediction (AMVP) mode; and determine corresponding properties of the pairwise IBC candidate associated with properties of the two precious IBC candidates.
[0026] In a twenty-first aspect of the present disclosure, some embodiments of the present disclosure provide an apparatus for video decoding. The apparatus can include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Moreover, the one or more processors, upon execution of the instructions, are configured to perform the method according to the first aspect.
[0027] In a twenty-second aspect of the present disclosure, some embodiments of the present disclosure provide an apparatus for video encoding. The apparatus can include one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Moreover, the one or more processors, upon execution of the instructions, are configured to perform the method according to the second aspect.
[0028] In a twenty-third aspect of the present disclosure, some embodiments of the present disclosure provide a non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect.
[0029] In a twenty-fourth aspect of the present disclosure, some embodiments of the present disclosure provide a non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the second aspect.
[0030] In a twenty-fifth aspect of the present disclosure, some embodiments of the present disclosure provide a non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to the first aspect.
[0031] In a twenty-sixth aspect of the present disclosure, some embodiments of the present disclosure provide a non-transitory computer-readable storage medium for storing a bitstream generated by the method according to the second aspect.
[0032] It should be understood that both the foregoing general description and the following detailed description are merely examples, but are not limiting on the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0034] Figure 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks in accordance with some embodiments of the present disclosure.
[0035] Figure 2 is a block diagram illustrating an exemplary video encoder in accordance with some embodiments of the present disclosure.
[0036] Figure 3 is a block diagram illustrating an exemplary video decoder in accordance with some embodiments of the present disclosure.
[0037] Figures 4A-4E is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes in accordance with some embodiments of the present disclosure.
[0038] Figures 5A-5B illustrates an example of a 4-parameter affine model in accordance with some examples of the present disclosure.
[0039] Figure 5C illustrates an example of a 6-parameter affine model in accordance with some examples of the present disclosure.
[0040] Figure 6 illustrates an example of neighboring neighboring blocks for inherited affine merge candidates in accordance with some examples of the present disclosure.
[0041] Figure 7 illustrates an example of neighboring neighboring blocks for constructed affine merge candidates in accordance with some examples of the present disclosure.
[0042] Figure 8 illustrates a current CTU processing order and its available reference samples in the current CTU and the left CTU in accordance with some examples of the present disclosure.
[0043] Figure 9 illustrates a padding candidate for replacing zero vectors in an IBC list in accordance with some examples of the present disclosure.
[0044] Figure 10 illustrates a reference region for IBC when encoding a CTU (m, n) in accordance with some examples of the present disclosure.
[0045] Figure 11 illustrates an IBC reference region for camera-captured content in accordance with some examples of the present disclosure.
[0046] Figures 12A-12B illustrates a partitioning method for angular modes in accordance with some examples of the present disclosure.
[0047] Figure 13A Spatial neighboring blocks used by an ATVMP are shown in accordance with some examples of the present disclosure.
[0048] Figure 13B An example of deriving a sub-CU motion field by applying motion displacement and scaling from a spatial neighbor to motion information from a corresponding collocated CU is shown in accordance with some examples of the present disclosure.
[0049] Figure 14 is a flowchart of decoding a bin in accordance with some examples of the present disclosure.
[0050] Figure 15 An intra template matching search region used in some examples of the present disclosure is shown.
[0051] Figure 16A An example of BV adjustment for horizontal flipping in accordance with some examples of the present disclosure is shown.
[0052] Figure 16B An example of BV adjustment for vertical flipping in accordance with some examples of the present disclosure is shown.
[0053] Figure 17 An example of five positions in a reconstructed luma sample in accordance with some examples of the present disclosure is shown.
[0054] Figure 18 An example of a prediction process of a DBV method in accordance with some examples of the present disclosure is shown.
[0055] Figure 19 An example of AMVP IBC candidate clustering based on L2 distance and TM cost in accordance with some examples of the present disclosure is shown.
[0056] Figure 20 is a diagram illustrating a computing environment coupled with a user interface in accordance with some examples of the present disclosure.
[0057] Figure 21 is a flowchart illustrating a method for video decoding in accordance with some examples of the present disclosure.
[0058] Figure 22 is a flowchart illustrating a method for video decoding in accordance with some examples of the present disclosure.
[0059] Figure 23 is a flowchart illustrating a method for video decoding in accordance with some examples of the present disclosure.
[0060] Figure 24 is a flowchart illustrating a method for video decoding in accordance with some examples of the present disclosure.
[0061] Figure 25 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0062] Figure 26 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0063] Figure 27 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0064] Figure 28 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0065] Figure 29 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0066] Figure 30 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0067] Figure 31 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0068] Figure 32 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0069] Figure 33 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0070] Figure 34 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0071] Figure 35 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0072] Figure 36 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0073] Figure 37 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0074] Figure 38 FIG. 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.
[0075] Figure 39is a flowchart illustrating a method for video encoding, in accordance with some examples of the present disclosure.
[0076] Figure 40 is a flowchart illustrating a method for video encoding, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION
[0077] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to help provide an understanding of the subject matter presented herein. But the subject matter presented can be practiced without the exact details, and can
[0078] It should be noted that the terms "first," "second," and the like in the description and in the claims of the present disclosure, as well as in the drawings, are used to distinguish between similar objects and are not necessarily used to describe a particular sequential or chronological order. It is understood that the use of such terms in the description is not meant to limit the scope of the present disclosure to a particular order or sequence, unless otherwise specifically stated.
[0079] Figure 1 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel, in accordance with some implementations of the present disclosure. As shown in Figure 1 As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will be later decoded by a destination device 14. Source device 12 and destination device 14 can comprise any of a variety of electronic devices, including a cloud server, a server computer, a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0080] In some embodiments, destination device 14 can receive, via link 16, encoded video data to be decoded. Link 16 can comprise any type of communication medium or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, link 16 can comprise a communication medium to enable source device 12 to transmit encoded video data directly to destination device 14 in real-time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from source device 12 to destination device 14.
[0081] In other embodiments, encoded video data can be transmitted from output interface 22 to storage device 32. Subsequently, encoded video data in storage device 32 can be accessed by destination device 14 via input interface 28. Storage device 32 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Video Disk (DVD), Compact Disk-Read Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In a further example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Destination device 14 can access stored video data from storage device 32 via streaming or download. The file server can be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 14 can access the encoded video data from the file server through any standard data connection, including a wireless channel (e.g., a wireless fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server. The transmission of encoded video data from storage device 32 can be a streaming transmission, a download transmission, or a combination of both streaming and download transmissions.
[0082] As Figure 1As shown in FIG. 1, source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 can include a source such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface to receive video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As one example, if video source 18 is a video camera of a security surveillance system, source device 12 and destination device 14 can form a camera phone or video phone. However, the implementations described in this application can be applied to video coding in general, and can have application to wireless and / or wired applications.
[0083] The captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored onto storage device 32 for later access by destination device 14 or other devices, for decoding and / or playback. Output interface 22 can further include a modem and / or a transmitter.
[0084] Destination device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 can include a receiver and / or modem and receives encoded video data over link 16. The encoded video data communicated over link 16, or provided on storage device 32, can include a variety of syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements can be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0085] In some implementations, destination device 14 can include display device 34, which can be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and can include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0086] Video encoder 20 and video decoder 30 can operate according to a proprietary standard or industry standard, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the application is not limited to a specific video coding / decoding standard and can be applicable to other video coding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that video decoder 30 of destination device 14 can be configured to decode video data according to any of these current or future standards.
[0087] Video encoder 20 and video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. When implemented partially in software, an electronic device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video coding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device.
[0088] In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or the components of video encoder 20 described below with reference to FIG. 2, and / or output interface 22), and / or the components of destination device 14 (e.g., input interface 28, video decoder 30, or the components of video decoder 30 described below with reference to FIG. 3, and / or memory 32) are integrated in a multimedia computing device. Figure 2 In some implementations, at least a portion of the components of source device 12 (e.g., video source 18, video encoder 20, or the components of video encoder 20 described below with reference to FIG. 2, and / or output interface 22), and / or the components of destination device 14 (e.g., input interface 28, video decoder 30, or the components of video decoder 30 described below with reference to FIG. 3, and / or memory 32) are integrated in a multimedia computing device. Figure 3At least a portion of the components described as included in video decoder 30, and display device 34, can operate in a cloud computing services network, such as a software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), which can provide software, platforms, and / or infrastructure. In some implementations, one or more components of source device 12 and / or destination device 14 that are not included in the cloud computing services network can be disposed in one or more client devices, and the one or more client devices can communicate with server computers in the cloud computing services network through a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In embodiments, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers implemented by at least a portion of the components of source device 12 and / or at least a portion of the components of destination device 14 in the cloud computing services network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing services network can be a private cloud, a public cloud, or a hybrid cloud. The terms such as “cloud,” “cloud computing,” “cloud-based,” and the like can be used interchangeably herein as appropriate without departing from the scope of the disclosure. It should be understood that the disclosure is not limited to implementation in the above-described cloud computing services network. Rather, the disclosure can also be implemented in any other type of computing environment currently known or developed in the future.
[0089] Figure 2 FIG. 1 illustrates a block diagram of an example video encoder 20 according to some embodiments described in this application. Video encoder 20 can perform intra-prediction encoding and inter-prediction encoding on video blocks within a video frame. Intra-prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video coding, the term “frame” can be used as a synonym for the term “image” or “picture.”
[0090] As Figure 2As shown in FIG. 1, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, summer 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58, inverse transform processing unit 60, and summer 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between summer 62 and DPB 64 to filter block boundaries to remove blockiness artifacts from reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of summer 62. It should be noted that for the CCSAO technique, the present application is not limited to the embodiments described herein, but can also be applied to a case where an offset is selected for any one of a luma component, a Cb chroma component, and a Cr chroma component to any other one of the luma component, the Cb chroma component, and the Cr chroma component to modify the any other one based on the selected offset. Furthermore, it should also be noted that the first component referred to herein can be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component referred to herein can be any other one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component referred to herein can be the remaining one of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filters can be omitted, and the decoded video blocks can be provided directly by summer 62 to DPB 64. Video encoder 20 can take the form of a fixed or programmable hardware encoder, or can be dispersed in one or more of the illustrated fixed or programmable hardware encoders.
[0091] Video data memory 40 can store video data to be encoded by the components of video encoder 20. The video data in video data memory 40 can be obtained, for example, from video source 18, as shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20 (e.g., in intra- or inter-coding modes). Video data memory 40 and DPB 64 can be formed by a variety of memory devices, including a Figure 1 suggested in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20 (e.g., in intra- or inter-coding modes). Video data memory 40 and DPB 64 can be formed by a variety of memory devices, including a
[0092] As Figure 2As shown in FIG. 1, after receiving video data, partitioning unit 45 within prediction processing unit 41 partitions the video data into video blocks. This partitioning can also include partitioning of a video frame into slices, tiles, or other larger coding units (CUs) according to a predefined splitting structure (such as a quadtree (QT) structure) associated with the video data, for example. A video frame is or can be considered as a two-dimensional array or matrix of sample values. A sample in the array can also be referred to as a pixel or a pel. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. A video frame can be divided into a plurality of video blocks, for example, by using QT partitioning. A video block is or can be considered as a two-dimensional array or matrix of sample values again, but with smaller dimensions than the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. A video block can be further partitioned into one or more block partitions or sub-blocks (which can form blocks again) by using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof, for example, iteratively. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to HEVC and VVC, for example, a block or video block can be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or can be or correspond to a respective block (such as a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.
[0093] Prediction processing unit 41 can select one of a plurality of possible predictive encoding modes, such as one of a plurality of intra-predictive encoding modes or one of a plurality of inter-predictive encoding modes, for the current video block based on error results (e.g., rate and distortion levels). Prediction processing unit 41 can provide the resulting intra- or inter-predicted block to summer 50 to generate a residual block, and to summer 62 to reconstruct the encoded block for use as part of a reference frame at a later time. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partitioning information, and other such syntax information, to entropy encoding unit 56.
[0094] To select a suitable intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 can perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block being coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 can perform multiple coding passes, e.g., to select a suitable coding mode for each block of video data.
[0095] In some implementations, motion estimation unit 42 determines an inter-prediction mode for a current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, the motion vectors indicating displacement of video blocks within the current video frame relative to prediction blocks within a reference video frame. Motion estimation performed by motion estimation unit 42 is a process of generating motion vectors that estimate motion for video blocks. For example, a motion vector can indicate displacement of a video block within a current video frame or picture relative to a prediction block within a reference frame that is related to a current block being coded within the current frame. The predetermined pattern can designate video frames in the sequence as P-frames or B-frames. Intra-BC unit 48 can determine vectors for intra-BC coding (e.g., block vectors) in a similar manner as motion vectors determined by motion estimation unit 42 for inter-prediction, or can utilize block vectors determined by motion estimation unit 42.
[0096] In terms of pixel difference, a prediction block for a video block can be or can correspond to a block or reference block of a reference frame that is deemed to closely match the video block being coded, the pixel difference can be determined by a sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some implementations, video encoder 20 can calculate values for sub-integer pixel positions of reference frames stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference frame. Thus, motion estimation unit 42 can perform a motion search relative to full-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.
[0097] Motion estimation unit 42 calculates motion vectors for video blocks in an inter-prediction coded frame by comparing locations of the video blocks to locations of prediction blocks of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of the first and second reference frame lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vectors to motion compensation unit 44, which then sends to entropy encoding unit 56.
[0098] Motion compensation performed by motion compensation unit 44 can involve fetching or generating a prediction block based on a motion vector determined by motion estimation unit 42. After receiving a motion vector for a current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block can include luma component differences or chroma component differences or both. Motion compensation unit 44 can also generate syntax elements associated with a video block of a video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements can include, for example, syntax elements defining motion vectors used to identify a prediction block, any flags indicating a prediction mode, or any other syntax information described herein. It is noted that motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are illustrated separately for conceptual purposes.
[0099] In some implementations, intra BC unit 48 can generate vectors and fetch prediction blocks in a manner similar to that described above in connection with motion estimation unit 42 and motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are referred to as block vectors rather than motion vectors. Specifically, intra BC unit 48 can determine an intra prediction mode to use for encoding the current block. In some examples, intra BC unit 48 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and test their performance through rate-distortion analysis. Next, intra BC unit 48 can select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, intra BC unit 48 can use rate-distortion analysis to compute rate-distortion values for the various tested intra prediction modes, and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block that was encoded to produce the encoded block, and the bit rate (i.e., the number of bits) used to produce the encoded block. Intra BC unit 48 can compute a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion values for the block.
[0100] In other examples, intra BC unit 48 can perform such functions for intra BC prediction according to the implementations described herein using motion estimation unit 42 and motion compensation unit 44 in whole or in part. In either case, for intra block copy, in terms of pixel difference, the prediction block can be a block that is deemed to closely match the block to be encoded, the pixel difference can be determined by SAD, SSD, or other difference metric, and identifying the prediction block can include calculating values for sub-integer pixel positions.
[0101] Whether the prediction block is from the same frame according to intra prediction or from a different frame according to inter prediction, video encoder 20 can form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values forming the residual video block can include both luma component differences and chroma component differences.
[0102] As an alternative to inter prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra block copy prediction performed by intra BC unit 48 as described above, intra prediction processing unit 46 can intra predict the current video block. In particular, intra prediction processing unit 46 can determine an intra prediction mode to use for encoding the current block. To this end, intra prediction processing unit 46 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and intra prediction processing unit 46 (or, in some examples, a mode selection unit) can select an appropriate intra prediction mode to use from among the tested intra prediction modes. Intra prediction processing unit 46 can provide information indicative of the selected intra prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 can encode the information indicative of the selected intra prediction mode into the bitstream.
[0103] After prediction processing unit 41 determines a prediction block for the current video block via inter prediction or intra prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, e.g., a discrete cosine transform (DCT) or a conceptually similar transform.
[0104] Transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, quantization unit 54 can then perform a scan on the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 can perform the scan.
[0105] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy encoding methodology or technique. The encoded bitstream can then be transmitted to video decoder 30 as shown in FIG. 3 or archived, such as in storage device 32 as shown in FIG. 3 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video frame that is being encoded. Figure 1 Figure 1
[0106] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transforms, respectively, to reconstruct the residual video block in the pixel domain for use in generating reference blocks for predicting other video blocks. As noted above, motion compensation unit 44 can generate a motion compensated prediction block from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0107] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block can then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block to inter predict another video block in a subsequent video frame.
[0108] Figure 3 FIG. 3 is a block diagram illustrating an example video decoder 30, in accordance with some embodiments of the present application. Video decoder 30 includes video data memory 79, entropy decoding unit 80, prediction processing unit 81, inverse quantization unit 86, inverse transform processing unit 88, adder 90, and DPB 92. Prediction processing unit 81 further includes motion compensation unit 82, intra prediction unit 84, and intra BC unit 85. Video decoder 30 can perform a decoding process generally reciprocal to the encoding process described above in connection with video encoder 20. Figure 2 The decoding process described in connection with video encoder 20 is substantially reciprocal. For example, motion compensation unit 82 can generate prediction data based on motion vectors received from entropy decoding unit 80, while intra prediction unit 84 can generate prediction data based on intra prediction mode indicators received from entropy decoding unit 80.
[0109] In some examples, the components of video decoder 30 can be tasked to perform the implementations of the present application. Moreover, in some examples, the implementations of the present disclosure can be distributed among one or more of the components of video decoder 30. For example, intra BC unit 85 can perform the implementations of the present application, alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 can not include intra BC unit 85, and the functionality of intra BC unit 85 can be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0110] Video data memory 79 can store video data, such as an encoded video bitstream, to be decoded by the other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source, such as a camera, via wired or wireless network communication of video data, or by accessing physical data storage media (e.g., a flash drive or hard disk). Video data memory 79 can include an encoded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra- or inter-coding modes). Video data memory 79 and DPB 92 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For Figure 3 illustrative purposes, video data memory 79 and DPB 92 are depicted as two distinct components of video decoder 30. But it will be readily apparent to one of ordinary skill in the art that video data memory 79 and DPB 92 can be provided by same memory device or separate memory devices. In some examples, video data memory 79 can be on-chip with other components of video decoder 30, or off-chip relative to those components.
[0111] During the decoding process, video decoder 30 receives an encoded video bitstream that represents encoded video frames and associated syntax elements. Video decoder 30 can receive the syntax elements at the video frame level and / or video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra-prediction mode indicators, among other syntax elements. Entropy decoding unit 80 then forwards the motion vectors, or intra-prediction mode indicators, among other syntax elements, to prediction processing unit 81.
[0112] When a video frame is coded as an intra-predicted (I) frame or an intra-coded prediction block in another type of frame, intra-prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-prediction mode signaled and reference data from previously decoded blocks of the current frame.
[0113] When a video frame is coded as an inter-predicted (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 produces one or more prediction blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks can be produced from a reference frame within one of the reference frame lists. Video decoder 30 can construct the reference frame lists, i.e., List 0 and List 1, using default construction techniques based on reference frames stored in DPB 92.
[0114] In some examples, when a video block is coded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 produces a prediction block for the current video block based on block vectors and other syntax elements received from entropy decoding unit 80. The prediction block can be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.
[0115] Motion compensation unit 82 and / or intra BC unit 85 determine the prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to produce the prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode used to code the video block of the video frame (e.g., intra-prediction or inter-prediction), the inter-prediction frame type (e.g., B or P), the construction information for one or more of the reference frame lists for the frame, the motion vectors for each inter-predicted coded video block of the frame, the inter-prediction status for each inter-predicted coded video block of the frame, and other information used to decode the video block in the current video frame.
[0116] Similarly, intra BC unit 85 can use some of the received syntax elements, such as flags, to determine that the current video block is predicted using the intra BC mode, the construction information for which video blocks of the frame are within the reconstructed region and should be stored in DPB 92, the block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information used to decode the video block in the current video frame.
[0117] Motion compensation unit 82 can also perform interpolation using interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks. In this case, motion compensation unit 82 can determine the interpolation filters used by video encoder 20 from the syntax elements received and use these interpolation filters when generating the prediction blocks.
[0118] Inverse quantization unit 86 inverse quantizes quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 using the same quantization parameter calculated by video encoder 20 for each video block in the video frame to determine a degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual blocks in the pixel domain.
[0119] After motion compensation unit 82 or intra BC unit 85 generates the prediction block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs the decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 to the corresponding prediction block generated by motion compensation unit 82 and intra BC unit 85. In-loop filter 91, e.g., a de-blocking filter, a SAO filter, a CCSAO filter, and / or an ALF, can be located between adder 90 and DPB 92 to further process the decoded video block. In some examples, in-loop filter 91 can be omitted, and the decoded video block can be directly provided by adder 90 to DPB 92. The decoded video blocks in a given frame are then stored in DPB 92, which stores reference frames for subsequent motion compensation of video blocks. DPB 92 or a separate memory device from DPB 92 can also store decoded video for later presentation on a display device (e.g., display device 34 of FIG. 1). Figure 1
[0120] In a typical video coding process, a video sequence generally includes an ordered set of frames or pictures. Each frame can include three arrays of samples, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame can be monochrome, and thus include only one two-dimensional array of luma samples.
[0121] As Figure 4A As shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, such as 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block may be an N×N sample block.
[0122] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or combinations thereof, and divide the CTU into smaller CUs. Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image can include, for example, two corresponding coding blocks (CBs) of luminance and chrominance samples of the same frame size, as well as syntax elements for encoding the samples of the coding blocks. In monochrome images or images with three separate color planes, a CU can include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that... Figure 4C and Figure 4DThe quad-tree partitioning depicted in the middle is for illustrative purposes only, and one CTU can be split into multiple CUs based on quad-tree partitioning / triple-tree partitioning / binary-tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to a quad-tree structure, and each quad-tree leaf CU can be further partitioned according to binary and triple tree structures. As shown in Figure 4E The coding block having width W and height H has five possible partitioning types, i.e., quad partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0123] In some implementations, video encoder 20 can further partition the coding block of a CU into one or more (M x N) PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. A PU of a CU can include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In a monochrome picture or a picture having three separate color planes, a PU can include a single PB and syntax structures used to predict the PB. Video encoder 20 can generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, the Cb PB, and the Cr PB of each PU of a CU.
[0124] Video encoder 20 can use intra prediction or inter prediction to generate the predicted blocks for a PU. If video encoder 20 uses intra prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0125] After video encoder 20 generates the predicted luma blocks, the predicted Cb blocks, and the predicted Cr blocks for one or more PUs of a CU, video encoder 20 can generate a luma residual block for the CU by subtracting the predicted luma blocks of the CU from the original luma coding block of the CU such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 can generate a Cb residual block and a Cr residual block for the CU such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU can indicate a difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0126] Furthermore, as Figure 4CAs shown in FIG. 1, video encoder 20 can partition a picture into one or more slices, each including one or more CUs. Video encoder 20 can partition a picture into one or more tiles, each including one or more CUs. Video encoder 20 can partition a picture into one or more warps, each including one or more CUs. Video encoder 20 can partition a picture into one or more tiles and one or more warps. Video encoder 20 can partition a picture into one or more slices, one or more tiles, one or more warps, or some combination of these partition types.
[0127] Video encoder 20 can apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block can be a two- dimensional array of transform coefficients. A transform coefficient can be a scalar. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.
[0128] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to a process that reduces the bit depth of transform coefficients to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After video encoder 20 quantizes a coefficient block, video encoder 20 can entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream that includes a sequence of bits, which forms a representation of encoded frames and associated data, which is maintained in storage device 32 or sent to destination device 14.
[0129] After receiving the bitstream generated by video encoder 20, video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 can reconstruct the frames of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 can perform inverse transforms on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs the coding blocks of the current CU by adding the samples of the prediction blocks for the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of a frame, video decoder 30 can reconstruct the frame.
[0130] As noted above, video coding primarily uses two modes (i.e., intra prediction (or intra-frame prediction) and inter prediction (or inter-frame prediction)) to achieve video compression. It should be noted that IBC can be considered as a third mode of intra prediction. Between the two modes, inter prediction contributes more to coding efficiency than intra prediction due to the use of motion vectors to predict the current video block from a reference video block.
[0131] However, as video data capture technology is constantly improving and finer video block sizes are used to preserve details in the video data, the amount of data required to represent the motion vectors for the current frame also increases substantially. One way to overcome this challenge benefits from the fact that not only a group of neighboring CUs in both spatial and temporal domains have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, the motion information (e.g., motion vectors) of the spatial neighboring CUs and / or temporal collocated CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (which is also referred to as the “motion vector predictor” (MVP) of the current CU) by exploring their spatial and temporal correlations.
[0132] Instead of encoding the actual motion vector of the current CU determined by the motion estimation unit 42 into the video bitstream as described above in connection with Figure 2 the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. By doing so, the motion vectors determined by the motion estimation unit 42 for each CU of the frame do not need to be encoded into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be reduced significantly.
[0133] Similar to the process of selecting a prediction block in a reference frame during inter prediction of a coding block, both video encoder 20 and video decoder 30 need to employ a set of rules for constructing a list of motion vector candidates (also referred to as a "merge list") for a current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally collocated CUs of the current CU, and then selecting one member from the list of motion vector candidates as the motion vector predictor for the current CU. By doing so, there is no need to send the list of motion vector candidates itself from video encoder 20 to video decoder 30, and the index of the selected motion vector predictor within the list of motion vector candidates is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor within the list of motion vector candidates to encode and decode the current CU.
[0134] The main focus of the present disclosure is to further enhance the intra block copy method by improving the coding efficiency and / or reducing its coding complexity.
[0135] Affine model
[0136] In HEVC, only translational motion model is applied to motion compensated prediction. However, in real world, there are many kinds of motions, e.g., zoom-in / zoom-out, rotation, perspective motion and other irregular motions. In VVC, affine motion compensated prediction is applied by signaling one flag for each inter coded block to indicate whether translational or affine motion model is applied to inter prediction. In current VVC, one affine coded block supports two affine modes, including 4-parameter affine mode and 6-parameter affine mode.
[0137] 4-parameter affine model has the following parameters: two parameters for translational motion in horizontal and vertical directions respectively, one parameter for scaling motion and one parameter for rotational motion in both directions. In this model, the horizontal scaling parameter is equal to the vertical scaling parameter, while the horizontal rotation parameter is equal to the vertical rotation parameter. To better adapt to the motion vector and affine parameters, these affine parameters will be derived from two MVs (also known as control point motion vectors (CPMV)) located at the top-left and top-right corners of the current block. As shown in Figures 5A-5B The affine motion field of a block is described by two CPMVs (V0, V1). Based on the control point motion, the motion field (v x , v y ) of one affine coded block is described as:
[0138]
[0139] 6-parameter affine mode has the following parameters: two parameters for translational motion in horizontal and vertical directions, respectively, two parameters for scaling motion and rotational motion in horizontal direction, respectively, two parameters for scaling motion and rotational motion in vertical direction, respectively. The 6-parameter affine motion model is coded with three CPMVs. As shown in FIG. 5, the three control points of a 6-parameter affine block are located at the top-left corner, top-right corner and bottom-left corner of the block. The motion at the top-left control point is related to translational motion, and the motion at the top-right control point is related to horizontal rotational and scaling motion, and the motion at the bottom-left control point is related to vertical rotational and scaling motion. Compared with the 4-parameter affine motion model, the 6-parameter affine motion model can have different rotational and scaling motion in horizontal direction than those in vertical direction. Assuming (V0, V1, V2) are the MVs of the top-left corner, top-right corner and bottom-left corner of the current block in FIG. 5, the motion vectors (v x , v y ) for each sub-block can be obtained using the three MVs at the control points as follows:
[0140]
[0141] Affine merge mode
[0142] In affine merge mode, the CPMVs of the current block are not explicitly signaled, but derived from neighboring blocks. Specifically, in this mode, the motion information of spatial neighboring blocks is used to generate the CPMVs of the current block. The size of the affine merge mode candidate list is limited. For example, in the current VVC design, there can be at most five candidates. The encoder can evaluate and select the best candidate index based on rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. The affine merge candidates can be decided in three ways:
[0143] • Inherited from neighboring affine coded blocks
[0144] • Constructed from translational MVs from neighboring blocks
[0145] • Zero MV
[0146] For the inherited approach, there can be at most two candidates. These candidates are obtained from the neighboring blocks located at the left-bottom of the current block (e.g., the scan order is from A0 to A1 as shown in FIG. 6A) and from the neighboring blocks located at the right-top of the current block (e.g., the scan order is from B0 to B2 as shown in FIG. 6B) if available. Figure 6 Figure 6
[0147] For the constructed approach, the candidate is a combination of translational MVs of neighboring blocks, which is generated in two steps.
[0148] • Step 1: Obtain four translational MVs from available neighbors.
[0149] o MV1: MV from one of the three neighboring blocks close to the top-left corner of the current block. As shown, the scan order is B2, B3 and A2. Figure 7
[0150] o MV2: MV from one of the two neighboring blocks close to the top-right corner of the current block. As shown, the scan order is Bl and B0. Figure 7
[0151] o MV3: MV from one of the two neighboring blocks close to the bottom-left corner of the current block. As shown, the scan order is Al and A0. Figure 7
[0152] o MV4: MV from the temporal collocated block of the neighboring block close to the bottom-right corner of the current block. As shown, the neighboring block is T. Figure 7
[0153] Step 2: Derive combinations based on the four translation MVs from Step 1.
[0154] o Combination 1: MV1, MV2, MV3
[0155] o Combination 2: MV1, MV2, MV4
[0156] o Combination 3: MV1, MV3, MV4
[0157] o Combination 4: MV2, MV3, MV4
[0158] o Combination 5: MV1, MV2
[0159] o Combination 6: MV1, MV3
[0160] When the merge candidate list is not full after being populated with inherited candidates and constructed candidates, insert a zero MV at the end of the list.
[0161] Affine AMVP mode
[0162] The affine AMVP (Advanced Motion Vector Prediction) mode can be applied to a CU whose width and height are both greater than or equal to 16. An affine flag at CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used or not, and then another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMVs of the current CU and their prediction values CPMVPs are signaled in the bitstream. The size of the affine AMVP candidate list is 2, and the affine AMVP candidate list is generated by using the following four types of CPMV candidates in the following order:
[0163] - Inherited affine AMVP candidate extrapolated from CPMVs of neighboring CUs
[0164] - Constructed affine AMVP candidate CPMVP derived using translational MVs of neighboring CUs
[0165] - Translational MVs from neighboring CUs
[0166] - Temporal MVs from collocated CUs
[0167] - Zero MV
[0168] The checking order of inherited affine AMVP candidates is the same as that of inherited affine merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting inherited affine motion predictors into the candidate list, no pruning process is applied.
[0169] Constructed AMVP candidates are derived from the same spatial neighbors as affine merge mode. The same checking order as in affine merge candidate construction is used. In addition, the reference picture index of the neighboring blocks is also checked. The first block that is inter coded in the checking order and has the same reference picture as in the current CU is used. When the current CU is coded using 4-parameter affine mode and both mv0 and mv1 are available, mv0 and mv1 are added as one candidate to the affine AMVP candidate list. When the current CU is coded using 6-parameter affine mode and all three CPMVs are available, they are added as one candidate to the affine AMVP candidate list. Otherwise, the constructed AMVP candidate will be set as unavailable.
[0170] If the number of affine AMVP list candidates is still less than 2 after inserting valid inherited affine AMVP candidates and constructed AMVP candidates, mv0, mv1 and mv2 are added in turn as translational MVs to predict all control point MVs of the current CU if available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the list.
[0171] Intra block copy in Versatile Video Coding (VVC)
[0172] Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency for screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has already been reconstructed within the current picture. The luma block vector of an IBC coded CU has integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pel motion vector precision and 4-pel motion vector precision. An IBC coded CU is considered as a third prediction mode different from the intra prediction mode or the inter prediction mode. IBC mode is applicable to a CU whose width and height are both less than or equal to 64 luma samples.
[0173] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check for blocks whose width or height is not larger than 16 luma samples. For non-merge mode, block vector search is first performed using hash-based search. If hash search does not return a valid candidate, a local search based on block matching will be performed.
[0174] In hash-based search, the hash key (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key computation for each location in the current picture is based on 4x4 sub-blocks. For larger size of the current block, the hash key is determined to match the hash key of the reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference locations. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matched reference is computed and the reference with the smallest cost is selected.
[0175] In block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0176] At CU level, the IBC mode is signaled using a flag and it can be signaled as IBC AMVP mode or IBC skip / merge mode as follows:
[0177] - IBC skip / merge mode: a merge candidate index is used to indicate which block vectors from the list of neighboring candidate IBC coded blocks are used to predict the current block. The merge list is composed of spatial candidates, HMVP candidates and pairwise candidates.
[0178] - IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. The block vector prediction method uses two candidates as the predictor, one from the left neighbor and one from the above neighbor (if IBC coded). When either neighbor is not available, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0179] IBC reference region
[0180] To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstructed part of a pre-defined region that includes the region of the current CTU and a certain region of the left CTU. Figure 8 The reference region of IBC mode is illustrated, where each block represents a 64x64 luma sample unit.
[0181] Depending on the location of the current coding CU within the current CTU, the following applies:
[0182] - If the current block falls into the top-left 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, the current block can also refer to the reference samples in the bottom-right 64x64 block of the left CTU using CPR mode. Using CPR mode, the current block can also refer to the reference samples in the bottom-left 64x64 block and the top-right 64x64 block of the left CTU.
[0183] - If the current block falls into the top-right 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, in the case that the luma position (0, 64) is not yet reconstructed with respect to the current CTU, the current block can also refer to the reference samples in the bottom-left 64x64 block and the bottom-right 64x64 block of the left CTU using CPR mode; otherwise, the current block can also refer to the reference samples in the bottom-right 64x64 block of the left CTU.
[0184] - If the current block falls into the bottom-left 64x64 block of the current CTU, in addition to the already reconstructed samples in the current CTU, in the case that the luma position (64, 0) is not yet reconstructed with respect to the current CTU, the current block can also refer to the reference samples in the top-right 64x64 block and the bottom-right 64x64 block of the left CTU using CPR mode. Otherwise, the current block can also refer to the reference samples in the bottom-right 64x64 block of the left CTU using CPR mode.
[0185] - If the current block falls into the bottom-right 64x64 block of the current CTU, the current block can only refer to the already reconstructed samples in the current CTU using CPR mode.
[0186] This restriction allows IBC mode to use the local on-chip memory of hardware implementations for its implementation.
[0187] Interaction of IBC with other coding tools
[0188] Interaction of IBC mode with other inter coding tools in VVC (such as paired merge candidate, history-based motion vector predictor (HMVP), combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partition mode (GPM)) is as follows:
[0189] - IBC can be used together with paired merge candidate and HMVP. A new paired IBC merge candidate can be generated by averaging the two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference.
[0190] - IBC cannot be used in combination with the following inter tools: affine motion, CIIP, MMVD, and GPM.
[0191] - When using DUAL_TREE partition, IBC is not allowed for chroma coding blocks.
[0192] Unlike in HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in the reference picture list 0 for IBC prediction. The derivation process of the motion vector for IBC mode excludes all neighboring blocks in inter mode and vice versa. The following IBC design aspects are applied:
[0193] - IBC shares the same process as in regular MV merge, including paired merge candidate and history-based motion predictor, but TMVP and zero vector are not allowed as they are not valid for IBC mode.
[0194] - Separate HMVP buffers (5 candidates each) are used for regular MV and IBC.
[0195] - The block vector constraints are implemented in the form of bitstream conformance constraints, the encoder needs to make sure that there are no invalid vectors in the bitstream and merge should not be used if the merge candidate is invalid (out of range or 0). This bitstream conformance constraint is expressed in terms of virtual buffers as follows.
[0196] - For deblocking, IBC is treated as inter mode.
[0197] - If the current block is coded using IBC prediction mode, AMVR does not use quarter-pel; instead, AMVR is signaled to indicate only if the MV is inter-pel or 4 integer-pel.
[0198] The number of IBC merge candidates can be signaled in the slice header separately from the number of regular merge candidates, subblock merge candidates, and geometric merge candidates.
[0199] A virtual buffer concept is used to describe the allowable reference region and the effective block vector for the IBC prediction mode. Let the CTU size be denoted as ctbSize, the virtual buffer ibcBuf has a width of wIbcBuf = 128 x 128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128 x 128, the size of ibcBuf is also 128 x 128; for a CTU size of 64 x 64, the size of ibcBuf is 256 x 64; and for a CTU size of 32 x 32, the size of ibcBuf is 512 x 32.
[0200] The size of a VPDU is min(ctbSize, 64) in each dimension, W v = min(ctbSize, 64).
[0201] The virtual IBC buffer ibcBuf is maintained as follows.
[0202] - When starting the decoding of each CTU row, flush the entire ibcBuf with invalid values -1.
[0203] - When starting the decoding of a VPDU (xVPDU, yVPDU) relative to the top-left corner of the picture, set ibcBuf[x][y] = -1 for x = xVPDU % wIbcBuf,..., xVPDU % wIbcBuf + W v -1; y = yVPDU % ctbSize,..., yVPDU % ctbSize + W v -1.
[0204] - After decoding, a CU contains (x, y) relative to the top-left corner of the picture, set ibcBuf[x % wIbcBuf][y % ctbSize] = recSample[x][y]
[0205] For a block covering coordinates (x, y), the block is valid if the following holds for a block vector bv = (bv[0], bv[1]):
[0206] ibcBuf[(x + bv[0]) % wIbcBuf][(y + bv[1]) % ctbSize] should not be equal to -1.
[0207] Intra block copy in enhanced compression model (ECM)
[0208] In ECM, IBC is improved from the following aspects.
[0209] IBC merge / AMVP list construction
[0210] IBC merge / AMVP list construction is modified as follows:
[0211] • An IBC merge / AMVP candidate can only be inserted into the IBC merge / AMVP candidate list if it is valid.
[0212] • The top-right, bottom-left and top-left spatial candidates and one pair-wise average candidate can be added to the IBC merge / AMVP candidate list.
[0213] • Template based adaptive reordering (ARMC-TM) is applied to the IBC merge list.
[0214] The HMVP table size of IBC is increased to 25. After deriving up to 20 IBC merge candidates by full pruning, they are reordered together. After reordering, the first 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC merge list.
[0215] The zero vector candidate used to fill the IBC merge / AMVP list is replaced by a set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in IBC merge mode and thus, it is discarded as a BVP in the IBC candidate list.
[0216] Three candidates are located at the nearest corners of the reference region and the other three candidates are determined to be in the middle of the three sub-regions (A, B and C) whose coordinates are determined by the width, height of the current block and the ΔX and ΔY parameters as depicted in Figure 9
[0217] IBC with template matching
[0218] Template matching is used in IBC for both IBC merge mode and IBC AMVP mode.
[0219] The IBC-TM merge list is modified compared to the list used by regular IBC merge mode such that the candidates are selected according to a pruning method and the motion distance between the candidates is the same as for regular TM merge mode. The zero motion padding at the end is replaced by the motion vector of the left (-W, 0), top (0, -H) and top-left (-W, -H) where W is the width and H is the height of the current CU.
[0220] In IBC-TM merging mode, the selected candidates are refined using a template matching method before the RDO or decoding process. IBC-TM merging mode competes with the regular IBC merging mode and uses a signal transmission TM-merging flag.
[0221] In IBC-TM AMVP mode, a maximum of three candidates are selected from the IBC-TM merge list. Each of these three selected candidates is refined using a template matching method and ranked according to its resulting template matching cost. Then, during motion estimation, only the first two are considered as usual.
[0222] Template matching refinement for IBC-TM merging mode and AMVP mode is very simple because the IBC motion vector is constrained to (i) integers and (ii) within the reference region, such as Figure 8 As shown. Therefore, in IBC-TM merge mode, all thinning is performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision depending on the AMVR value. This thinning only accesses samples that are not interpolated. In both cases, the thinning motion vector and the template used in each thinning step must adhere to the constraints of the reference region.
[0223] IBC Reference Area
[0224] The IBC reference area extends to the two CTU rows above. Figure 10 The diagram illustrates the reference region used to encode a CTU(m,n). Specifically, for a CTU(m,n) to be encoded, the reference region includes CTUs with indices (m–2,n–2)…(W,n–2), (0,n–1)…(W,n–1), (0,n)…(m,n), where W represents the maximum horizontal index within the current tile, strip, or image. This setting ensures that IBC does not require additional memory on the current ETM platform when the CTU size is 128. The per-sample block vector search (or local search) range is limited to: horizontal [–(C<<1),C>>2], vertical [–C,C>>2], to accommodate the expansion of the reference region, where C represents the CTU size.
[0225] IBC merging mode using block vector difference
[0226] The ECM employs an IBC merging mode utilizing block vector differences. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.
[0227] Basic candidates are selected from the top five candidates in the reordered IBC merge list. Then, all possible MBVD refinement positions (20×4) for each basic candidate are reordered based on the SAD cost between the template (the row above the current block and the column to the left of the current block) and its reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are reserved as available positions and thus used for MBVD index encoding.
[0228] IBC adaptation for camera-captured content
[0229] When adapting to IBC for camera-captured content, the IBC reference range decreases from 2 CTU lines to 2 × 128 lines, such as Figure 11 As shown. On the encoder side, to reduce complexity, the local search range is set to a horizontal [-8,8] and a vertical [-8,8] range centered on the first block vector prediction value of the current CU. This encoder modification is not applicable to SCC sequences.
[0230] Sub-block-based temporal motion vector prediction (SbTMVP)
[0231] VVC supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP uses motion fields from co-located images to improve motion vector prediction and merging patterns for CUs in the current image. The same co-located images used by TMVP are used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects:
[0232] –TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level;
[0233] –TMVP obtains the temporal motion vector from the co-op block in the co-op image (the co-op block is the lower right or center block relative to the current CU), while SbTMVP applies motion shift before obtaining the temporal motion information from the co-op image, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0234] Figures 13A-13B The diagram illustrates the SbTVMP process. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, it checks... Figure 13A In the spatial neighbor A1, if A1 has a motion vector that uses a co-located image as its reference image, then that motion vector is selected as the motion shift to apply. If no such motion is identified, then the motion shift is set to (0,0).
[0235] In the second step, such as Figure 13BAs shown, the motion shift identified in step 1 (i.e., the coordinates added to the current block) is applied to obtain sub-CU level motion information (motion vectors and reference indices) from the co-location image. Figure 13B The example assumes that motion shift is set to prevent movement of A1. Then, for each sub-CU, motion information of the sub-CU is derived using motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-location image. After identifying the motion information of the co-location sub-CU, the motion information is converted into a motion vector and reference index for the current sub-CU in a manner similar to the TMVP process of HEVC, wherein temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.
[0236] In VVC, a sub-block-based merge list containing a combination of SbTVMP candidates and affine merge candidates is used for signal transmission in sub-block-based merge mode. The SbTVMP mode is enabled / disabled by the Sequence Parameter Set (SPS) flag. If the SbTVMP mode is enabled, the SbTVMP prediction is added as the first entry in the sub-block-based merge candidate list, followed by the affine merge candidate. The size of the sub-block-based merge list is transmitted in SPS, and the maximum allowed size of the sub-block-based merge list in VVC is 5.
[0237] The sub-CU size used in SbTMVP is fixed at 8×8, and similar to the affine merging mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.
[0238] The encoding logic for the additional SbTMVP merge candidate is the same as that for other merge candidates, that is, for each CU in the P or B stripe, an additional RD check is performed to determine whether to use the SbTMVP candidate.
[0239] Intra-frame template matching
[0240] Intra-frame template matching prediction (intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame that matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and performs the same prediction operation on the decoder side.
[0241] A prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in a predefined search region in Figure 4, which consists of the following:
[0242] R1: Current CTU
[0243] R2: Top Left CTU
[0244] R3: Above CTU
[0245] R4: Left CTU
[0246] The sum of absolute differences (SAD) is used as the cost function.
[0247] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0248] The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportionally to the block dimensions (BlkW, BlkH) to ensure a fixed number of SAD comparisons per pixel. That is:
[0249] SearchRange_w=a*BlkW
[0250] SearchRange_h = a * BlkH
[0251] Here, 'a' is a constant that controls the trade-off between gain and complexity. In fact, 'a' equals 5.
[0252] For CUs with width and height dimensions less than or equal to 64, enable the intra-frame template matching tool. This maximum CU size used for intra-frame template matching is configurable.
[0253] When DIMD is not used for the current CU, the intra-frame template matching prediction mode is transmitted at the CU level using a dedicated flag.
[0254] Probability estimation techniques for CABAC in AVC and HEVC
[0255] CABAC (Context-Based Adaptive Binary Arithmetic Coding) was initially introduced in the H.264 / AVC standard as one of two supported entropy coding schemes. In CABAC, arithmetic coding consists of two modules: codeword mapping (also known as binarization) and probability estimation. During codeword mapping, syntax elements are mapped to a string of binary bits. This mapping is achieved through a binarizer, which converts syntax elements into several sets of binary bits based on different binarization schemes. In practice, various binarization schemes can be applied for this conversion, such as fixed-length codes, unary codes, truncated unary codes, and k-order exponential Golomb codes. The purpose of the probability estimation module is to determine the probability that a binary bit is either 1 or 0. In AVC, the probability of a binary bit is calculated based on an exponential aging model, where the probability that a current binary bit is equal to 1 or 0 depends on the value of previously encoded binary bits. Furthermore, according to common statistics, the influence of bits immediately preceding a current binary bit is generally greater than that of bits encoded much earlier. With this in mind, a parameter α is introduced in CABAC that controls the number N of previously encoded bits used to estimate the probability of the current bit, i.e., N = 1 / α. This parameter translates to an adaptive rate at which the probability is updated as the number of encoded bits increases. Specifically, using the adaptive parameter α, the probability that a bit is a minimum probability symbol (LPS) is recursively calculated as follows:
[0256] p(t+1)=p(t)·(1-α)+x(t)·α (3)
[0257] Where p(t) is the probability of the LPS symbol at time t; where p(t+1) is the updated probability of the LPS symbol at time t+1; x(t) is equal to 1 when the current bit is an LPS symbol and equal to 0 when the current bit is a maximum probability symbol (MPS). In the CABAC engines of AVC and HEVC, the probability is updated independently according to (3) for each syntax element with a fixed value of α≈1 / 19.69 (i.e., when estimating the probability of a current bit, approximately 19.69 previously encoded bits are considered). Furthermore, to avoid the use of multiplication during probability estimation, the probability p(t) in equation (3) (which is a real number ranging from 0 to 1) is quantized into a fixed set of probability states. For example, in AVC and HEVC, the probability has 7 bits of precision, corresponding to 128 probability states.
[0258] In AVC and HEVC, a video bitstream typically consists of one or more independently decodeable stripes. At the beginning of each stripe, the probabilities of all contexts are initialized to some predefined values. Theoretically, given the statistical properties of a given context, a uniform distribution (i.e., p) should be used.init The context probabilities are initialized with a value of 0.5. However, to allow the probabilities of a context to reach their corresponding statistical distribution more quickly, it is found beneficial to provide each context with some appropriate initial probability values (which may not be equally probable). Specifically, in AVC and HEVC, given a sliceQP... Y The initial QP, the initial probability state InitProbState of a context, is calculated as follows:
[0259]
[0260] Here, SlopeIdx and OffsetIdx (both in the range of 0 to 15) are two initialization parameters that are predefined and stored as lookup tables (LUTs) for calculating the initial probability of a context. As shown in equation (4), the initial probability state is modeled by a linear function of the striped QP, where the slope is equal to (m >> 4) and the offset is equal to n.
[0261] Probability estimation techniques for CABAC in VVC
[0262] Aside from the following key differences, the probability estimation module used in VVC is almost identical to the probability estimation modules used in AVC and HEVC:
[0263] 1. VVC maintains two probability estimates for each context, where each context has its own probability adaptive rate α in (3). The final probability actually used for arithmetic coding is the average of these two estimates.
[0264] 2. In VVC, multiple probabilistic LUTs are predefined and used to initialize the probabilities of different contexts for a stripe. Similar to AVC and HEVC, the initial probability estimates are based on a linear model with the stripe QP as input. However, in VVC, the obtained values represent the actual probability values; while in AVC / HEVC, the obtained values represent the indices of the probability states.
[0265] Multiple Hypothesis Probability Estimation
[0266] Clearly, since the statistical properties of all grammatical elements differ, using a fixed adaptive parameter for all grammatical elements may not be optimal. On the other hand, numerous scientific studies have demonstrated that using multiple probability estimators can achieve better estimation accuracy compared to a single estimator. Therefore, a multi-hypothesis probability estimation scheme is applied in the CABAC design of VVC, where two different adaptive parameters α0 and α1 are utilized, corresponding to a slow and a fast probability adaptation speed, respectively. In this way, two different probabilities can be calculated for each binary bit using the two adaptive parameters, and then averaged to generate the final probability of the binary bit, i.e.,
[0267]
[0268] Here, α0 and α1 are two adaptive parameters associated with the two probabilistic assumptions. In VVC, a training algorithm designed to jointly optimize the adaptive parameters and the initial probabilities is used to independently select the values of α0 and α1 for each context. Specifically, according to the current design, for each context, α0 is allowed to be selected from a set of predefined values {1 / 4, 1 / 8, 1 / 16, 1 / 32}, and α1 is allowed to be selected from another set of predefined values {1 / 32, 1 / 64, 1 / 128, 1 / 256, 1 / 512}.
[0269] Initial probability calculation
[0270] Similar to AVC / HEVC, VVC's CABCA procedure also calls a QP-related probability initialization procedure at the beginning of each stripe. However, unlike AVC / HEVC, which initializes the state of a probabilistic state machine, VVC directly obtains the actual values of the initial probabilities, as shown below.
[0271]
[0272] SlopeIdx and OffsetIdx are two initialization parameters used to calculate the slope and offset of the linear model, and they are represented with 3 digits of precision. and These are the two initial probabilities calculated for the two probability estimators.
[0273] Entropy coding in ECM
[0274] Extended precision
[0275] The increase in intermediate precision used in the arithmetic coding engine includes three elements. First, the precision for both probability states is increased to 15 bits, compared to 10 bits and 14 bits in VVC. Second, the LPS range update process is modified as follows:
[0276] if q>=16384
[0277] q=2 15 –1–q
[0278] R LPS =((range*(q>>6))>>9)+1,
[0279] Where range is a 9-bit variable representing the width of the current interval, q is a 15-bit variable representing the probability state of the current context model, and R LPS This is the updated range of LPS. This operation can also be achieved by looking up 512×256 entries in a 9-bit lookup table. Third, on the encoder side, the 256-entry lookup table used for bit estimation in the VTM is expanded to 512 entries.
[0280] Window size based on strip type
[0281] Since the statistics differ for different stripe types, it is beneficial to update the context's probability state at a rate that provides a more accurate probability estimate for a given stripe type (e.g., a more accurate prediction of the probability that a bit will be 1 or 0). Therefore, for each context model, three window sizes are predefined for I-strip, B-strip, and P-strip, just like the initialization parameters.
[0282] The context initialization parameters and window size were retrained.
[0283] Improved probability estimation of CABAC
[0284] Multi-hypothesis probability estimation with adaptive weights
[0285] Multiple hypothesis-based (MHP-AW) probabilities are estimated using adaptive weights. Specifically, two separate probability estimates, p0 and p1, are maintained for each context and updated based on their own fitness. However, instead of using a simple average, multiple weights are introduced to obtain the resulting probability p for binary arithmetic encoding, as follows:
[0286] p = (ω0·p0 + ω1·p1) >> s
[0287] Here, ω0 and ω1 are weights selected from the predefined set {10,12,16,20,22}; s is a bitwise right shift value, which equals 5 when (ω0+ω1)≤32, and 6 otherwise. For each context model, three different sets of weights are predefined under I-slice, B-slice, and P-slice types. Weights for I-slice type are only allowed for intra-slice use, while weights for B-slice and P-slice types can be switched for inter-slice use at the slice level.
[0288] CABAC initialization based on previous inter-frame striping and windowing adjustments
[0289] After encoding the last CTU, the context initialization stored at the previously encoded image can be used to initialize inter-frame slices with the same slice type, QP, and time ID. For each slice type, the buffer size used to store the previously initialized slices is set to equal to 5, and when the buffer is full, the entry with the smallest QP and time ID is removed first before storing the initialization.
[0290] CABAC employs two probabilistic states updated with short and long window sizes, respectively. Since the predefined window size for each context model is not optimal for different statistics in different regions, the window size is adjusted based on the previously encoded bits for each context.
[0291] The short and long window sizes used in CABAC updates are adjusted using two incremental parameters stored in a lookup table for each context and obtained using previously encoded bits as indexes. The previously encoded bits are used as indices to retrieve the adjustment parameters from the lookup table: delta0 for the short window and delta1 for the long window. The original short and long window sizes, stored in the existing initialization table and defined for this context model, are denoted as shift0 and shift1, respectively. The actual window sizes used to encode the current bits after adjustment are (shift0 + delta0) and (shift1 + delta1), where shift0 and shift1 are predefined window sizes stored in the context initialization table.
[0292] Reconstruction Reordering IBC (RRIBC) for Screen Content Encoding
[0293] Symmetry is frequently observed in video content, especially in text character regions within screen content sequences and in computer-generated graphics. To further improve the coding efficiency of IBC in ECM, a Reconstruction Reordering IBC (RRIBC) mode for screen content video coding is proposed.
[0294] When RRIBC is applied, the samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the predicted block is obtained without flipping. On the decoder side, the reconstructed block is flipped back to recover the original block.
[0295] Specifically, two flipping methods are supported for RRIBC coded blocks: horizontal flipping and vertical flipping. First, the syntax flags of the IBC AMVP coded block are signaled to indicate whether the reconstruction is flipped. If the reconstruction is flipped, another flag specifying the flipping type is further signaled. For IBC merging, the flipping type is inherited from neighboring blocks, and no syntax signaling is used. Considering horizontal or vertical symmetry, the current block and the reference block are typically horizontally or vertically aligned. Therefore, when applying a horizontal flip, the vertical component of the BV is not signaled and is inferred to be equal to 0. Similarly, when applying a vertical flip, the horizontal component of the BV is not signaled and is inferred to be equal to 0.
[0296] To better utilize symmetry properties, a flip-aware BV adjustment method is applied to refine the block vector candidates. For example, as... Figure 16A and Figure 16B As shown, (x nbr ,y nbr ) and (x cur ,y cur ) represent the coordinates of the center sample points of the neighboring blocks and the current block, respectively, BV nbr and BV cur These represent the block values (BV) of the neighboring block and the current block, respectively. Instead of directly inheriting the BV from the neighboring block, when encoding the neighboring block with horizontal flipping, motion shifts are added to the BV. nbr The horizontal component (represented as BV) nbr h To calculate BV cur The horizontal component, i.e., BV cur h =2(x nbr -x cur )+BV nbr h Similarly, in the case of encoding neighboring blocks using vertical flipping, motion shifting is added to the BV. nbr The vertical component (denoted as BV) nbr v To calculate BV cur The vertical component, i.e., BV cur v =2(y nbr -y cur )+BV nbr v .
[0297] Direct Block Vector (DBV) Mode for Chromatography Prediction
[0298] Direct Block Vector (DBV) is provided to improve the encoding / decoding efficiency of chroma components when dual-tree is activated in an intra-strip. When chroma dual-tree is activated in an intra-strip, for a chroma CU encoded in DBV mode, if... Figure 17 If the center block is encoded using IBC or IntraTmp mode, its block vector bvL is used to obtain the chroma block vector bvC. The bv scaling process is determined based on template matching. If the luma block is encoded using RRIBC, bvL undergoes the same flip-sensory BV adjustment as RRIBC. Then, as... Figure 18 The depiction involves determining the corresponding offset position (xCb+bvC_hor, yCb+bvC_ver) by using the position of the current chroma block (xCb, yCb) and its bvC, and then performing block copy prediction.
[0299] Block Vector Difference Prediction for IBC Blocks (IBC-BVDP)
[0300] IBC-BVDP is a technique for predicting the sign and magnitude of the x and y components of the block vector difference (BVD) of an IBC block. Specifically, the BVD sign and suffix bits of the exponential Golomb code representing the BVD magnitude are predicted by estimating the template matching cost of candidate blocks to utilize the regular CABAC mode at the entropy coding level instead of its bypass mode. Specifically, the most significant bits of the magnitude suffixes for the horizontal and vertical BVD components are predicted, and the predicted match is encoded in the bitstream using the CABAC context mode. The less significant bits of the magnitude suffixes for the horizontal and vertical BVD components are encoded in bypass mode. The maximum number of bits to predict for the PU is controlled by a macro and is currently set to 4 in ECM8.0.
[0301] BVP candidate clustering and BVD symbol derivation for reconstructing reordered IBC patterns (IBC-BVPC)
[0302] In this IBC-BVPC mode, the IBC AMVP list construction is modified based on the clustering of BVP candidates according to the distance between BVP candidates and the sign prediction of BVD (if BV has an empty component).
[0303] For blocks whose BV has two non-empty components, clustering of the BVP candidates is applied before selecting two AMVP candidates. Clustering is used when the number of valid BVP candidates exceeds two, and up to six BVP candidates are clustered based on the L2 Euclidean distance between these candidates. The radius (R) determines the vector group (e.g., ... Figure 19 As shown), the following is a logarithmic function of the current block width (cbWidth) and height (cbHeight):
[0304] R=log2((cbWidth·cbHeight)>>MIN_PU_SIZE) (7)
[0305] The clustering method is applied in the order of the candidate list, and candidates assigned to a group are removed from the list of subsequent clusters. Within each group, the BVP with the lowest TM cost is selected as the representative candidate for that group. The motion estimation process is performed on the representative candidates from the first two groups, just as in the regular IBC AMVP list.
[0306] Conversely, BVs with a null component (including RRIBC blocks) are signaled to the decoder via the bvOneNullComp flag. Instead of invoking the AMVP IBC list construction, two new BVP candidates are determined, which are adjusted to the boundaries of the valid IBC search region according to the horizontal or vertical direction indicated by the bvNullCompDir flag.
[0307] AMVP BVP0 is set to the nearest valid location to the current block (-cbWidth or -cbHeight), so BVD (if not empty) is always negative, pointing to the left for a BV with an empty vertical component, or pointing upwards for a BV with an empty horizontal component. Similarly, AMVP BVP1 is set to the furthest location from the current block in the valid reference region, i.e., the left or top boundary of the IBC search region. Therefore, if BVP1 is selected, BVD is always positive, pointing to the right for a BV with an empty vertical component and to the bottom for a BV with an empty horizontal component.
[0308] The optimal IBC AMVP index is transmitted via signaling, allowing the symbols of the non-empty BVD components to be derived on the decoder side. Therefore, the absolute values of the non-empty components of the BVD are transmitted to the decoder via signaling, thereby improving encoding / decoding efficiency. The RRIBC mode is transmitted via signaling using existing syntax flags, and the direction of the flipped mode is obtained from the bvNullCompDir flag.
[0309] IBC with local illumination compensation (IBC-LIC)
[0310] Intra-Block Copy with Local Illumination Compensation (IBC-LIC) is an encoding / decoding tool that uses linear equations to compensate for local illumination variations within an image between an IBC-encoded CU and its predicted block. The parameters of the linear equations are derived in the same way as those for LIC used for inter-frame prediction, except that the reference template is generated using block vectors from the IBC-LIC. IBC-LIC can be applied to both IBC AMVP and IBC merging modes. In IBC AMVP mode, an IBC-LIC flag is signaled to indicate the use of IBC-LIC. In IBC merging mode, the IBC-LIC flag is inferred from merging candidates.
[0311] In video codecs, intra-block copying is well-known for accurately predicting both screen content and artificially generated content, where patterns and edges can repeat within a frame. Intra-block copying can also be beneficial for predicting natural content with repeating textures in the current frame. For codec scenarios with limited repeating content, intra-block copying can be omitted while still transmitting its minimum signaling bits. In such cases, to further improve the codec efficiency of intra-block copying, a more flexible on / off control mechanism with varying granularity is desired.
[0312] In inter-frame prediction encoding / decoding modes, fractional motion vectors are used to improve prediction accuracy. However, in the current intra-block copy mode, only integer motion vectors are used. This study aims to explore the encoding / decoding benefits of using fractional motion vectors for intra-block copy. When using fractional motion in intra-block copy, several subsequent issues need to be addressed: fractional motion derivation, signal transmission, interpolation padding, interpolation filter selection, and interaction with other encoding / decoding tools.
[0313] In this disclosure, the encoding and decoding tools for intra-frame block copying are improved in the following ways:
[0314] Flexible on / off control mechanism
[0315] • CABAC context window
[0316] • Interpolation-based fractional intra-frame block copying
[0317] ○Score Movement Search
[0318] ○Fractional motion refinement
[0319] ○Conditional sample / pixel fill for fractional interpolation
[0320] ○ Interpolation filter switching
[0321] ○Multiple Hypothesis Fractional Intra-Block Copying
[0322] • Signal transmission of motion information
[0323] • IBC Merge / AMVP Campaign Candidate List Construction
[0324] Combinations with intra-frame template matching
[0325] • Combination with IBC merging mode utilizing block vector differences
[0326] • Combination of IBC-LIC and IBC using fractional motion vectors
[0327] • Combination of DBV with IBC using fractional or integer motion vectors
[0328] • Combination of RRIBC and IBC using fractional motion vectors
[0329] • Combination of IBC-BVPC and IBC using fractional motion vectors
[0330] • Combination of IBC-BVDP and IBC using fractional motion vectors
[0331] • Combination of paired candidates with other IBC coding modes
[0332] Flexible on / off control mechanism
[0333] This section presents several methods for controlling the on / off state of IBC mode application. On / off control indicates whether IBC mode can be enabled for the current sequence, frame, stripe, CTU, or block (at different granularities). If IBC mode is enabled, additional flags (e.g., whether IBC mode is enabled or disabled for a specific block) and / or information (e.g., block vectors) can be transmitted via signaling. If IBC mode is disabled, flags or information are no longer transmitted via signaling.
[0334] In some embodiments, the on / off control of intra-frame block copying can be based on an explicit signaling method.
[0335] In one embodiment, on / off control is based on one or more sequence-level, frame-level, stripe-level, coding tree unit (CTU)-level, or block-level flags, or any combination of different level flags. When any combination of different level flags is used, the transmission of lower-level flags depends on the on / off state of higher-level flags. In one example, if a frame-level flag indicates that IBC mode is off, no further stripe-level or block-level flags are transmitted. Otherwise, lower-level flags(s) are transmitted further.
[0336] In another embodiment, on / off control is based on different regions. The purpose of the region concept is to provide more flexible granularity for IBC on / off control.
[0337] In one embodiment, the region here can be defined as a non-overlapping area within a frame, strip, or CTU. For all blocks located within a specific region, a single on / off control flag can be signed to indicate whether IBC mode is disabled for all of these blocks. The size of the region can be predefined as a set of fixed values, such as M×N, or a set of values transmitted via signaling.
[0338] In some other embodiments, the on / off control of intra-block copying can be based on local information and does not require explicit signal transmission.
[0339] In some embodiments, the on / off control is based on prediction information. In one embodiment, IBC mode is always off for inter-frame prediction blocks. In another embodiment, IBC mode is always off for inter-frame prediction blocks with unidirectional and / or bidirectional prediction.
[0340] In another embodiment, IBC mode is always disabled for blocks encoded in sub-block mode. Sub-block mode is a mode that divides the current block into sub-blocks, and each sub-block can have its own motion information. Examples include affine mode and SbTMVP mode. In yet another embodiment, IBC mode is always disabled for blocks not encoded in sub-block mode.
[0341] In some other embodiments, the on / off control is based on other encoding information. In one embodiment, the IBC model is always off when one or more other encoding modes are applied to the current block. For example, the IBC mode is always off when affine mode is enabled.
[0342] In some other embodiments, the on / off control is based on frame type. In one embodiment, IBC mode is always off for B-frames and / or P-frames.
[0343] In some other embodiments, the on / off control is based on block information. In one embodiment, IBC mode is always off for coded blocks smaller than a certain size (e.g., 8×8 blocks) or larger than a certain size (e.g., 64×64). In one embodiment, IBC mode is always off for wide blocks (e.g., blocks whose width is M times their height) or long blocks (e.g., blocks whose height is N times their width), where the values of M and N can be fixed (e.g., M=2, N=3) or transmitted as signals at the sequence or frame level.
[0344] CABAC context window
[0345] In current IBC designs, one or more IBC mode-related flags can exist with CABAC context coding. For example, block-level IBC enable flags are context-coded. Since statistics can differ for different slices or frame types, it is desirable for the context probability state to be updated at a rate that provides a more accurate probability estimate for a given slice / frame type (e.g., a more accurate prediction of the probability that a bit is 1 or 0).
[0346] In some embodiments, for each context model associated with the IBC mode, three windows can be predefined for three different stripes (including the I strip, B strip, and P strip).
[0347] In some embodiments, for each context model associated with the IBC mode, two windows can be predefined for different stripes with two different prediction modes (including intra-frame prediction stripes (I stripes) and inter-frame prediction stripes (B stripes and P stripes)).
[0348] Based on multiple context windows updated under a given stripe / frame type, one or more IBC mode-related flags (e.g., motion accuracy, interpolation filter selection for motion-compensated block prediction or / and template prediction, etc.) can be further signaled by using different context binary bits.
[0349] In one example, motion precision flags can be signaled individually or in combination for different strip / frame types. The motion precision flags signaled can depend on the currently supported precision type for a given strip / frame type (e.g., 1 pixel, 4 pixels, or fractional pixels, such as 1 / 2, or / and 1 / 4, or / and 1 / 8 pixels, or / and 1 / 16 pixels).
[0350] In another example, signal transmission interpolation filters can be further selected individually or jointly for different stripe / frame types (e.g., 12-tap, 6-tap, 4-tap, 2-tap, or / and 0-tap filters).
[0351] In some embodiments, one or more IBC mode-related flags (e.g., motion accuracy, interpolation filter selection for motion-compensated block prediction or / and template prediction, etc.) may be transmitted separately or jointly for different video component types.
[0352] In one example, motion precision flags can be supported, and then the luminance and chrominance components can be signaled in different ways. In one embodiment, 4-pixel, 1-pixel, 1 / 2-pixel, and 1 / 4-pixel precisions can be supported and signaled for the luminance component, while only 4-pixel, 1-pixel, and 1 / 2-pixel precisions are supported for the chrominance component.
[0353] When supporting different motion precisions and transmitting signals with different motion precisions for different video components, corresponding interpolation filters for each motion precision can be defined in similar or different ways for different video components. For example, the 2-tap filters for the luma and chroma components can be the same or different. Furthermore, the interpolation filters used at each motion precision can be the same or different, and the application of interpolation filters at each motion precision and for each video component can be defined individually or jointly.
[0354] In some embodiments, one or more IBC mode-related flags (e.g., motion accuracy, interpolation filter selection for motion-compensated block prediction or / and template prediction, etc.) may be further transmitted individually or jointly by taking into account different strip / frame types, different video components and / or different video resolutions.
[0355] When multiple windows are defined for different stripes or frames, the context window size and initialization parameters can also be retrained individually or jointly.
[0356] Fractional intra-frame copying based on interpolation
[0357] Score motion search
[0358] In one embodiment, fractional motion search can be performed on the encoder side, and the final motion signal can be transmitted to the decoder side. After subtracting the motion predictions known to both the encoder and decoder, the transmitted motion signal can be in the form of motion difference. The motion search can be performed in three steps:
[0359] • In step 1, the best N integer motion vectors with the minimum distortion cost (e.g., sum of absolute differences (SAD)) can be searched first.
[0360] • In step 2, half-pixel thinning is applied around each of the N integer motion vectors. In this step 2, M optimal half-pixel positions can be obtained (the optimal M positions indicate the M half-pixel motion differences with the lowest rate distortion cost). For example, an encoder or decoder can obtain the M optimal half-pixel positions with the lowest rate distortion cost. If K of the N integer motion vectors are selected, the output can be a total of K*M half-pixel positions.
[0361] In step 3, quarter-pixel thinning is applied around the optimal half-pixel position for each of the N integer motion vectors. In this step 3, a set of Q quarter-pixel positions can be obtained for each of the K*M half-pixel positions obtained in step 2. The optimal R quarter-pixel positions from all K*M*Q candidate positions can then be generated. The optimal position among the R positions (e.g., the position with the minimum rate distortion) can be determined by full rate distortion calculation and transmitted to the decoder. For example, the encoder can generate the optimal R quarter-pixel positions, then select the best position with the minimum rate distortion among the R positions, and transmit the best position among the R positions to the decoder. The values of N, M, K, Q, and R are position integers.
[0362] Following these three steps, (e.g., in the format of motion vector difference) the optimally thinned motion vector is transmitted via signal transmission (after half-pixel and / or quarter-pixel thinning). In this disclosure, and in this and the following sections, the motion vector can be used interchangeably with a block vector that identifies a reference / predicted block in the same picture / frame.
[0363] In another embodiment, fractional motion search can be performed at both the encoder and decoder sides, eliminating the need to transmit the final fractional motion via signaling. In this approach, a template-matching-based method can be used to find the optimal fractional motion.
[0364] In one or more embodiments, an inverse L-shaped sample / pixel region adjacent to the coded block can be used as a matching template, and the pixel / sample width can be prefixed, configurable, or transmitted as a signal in a sequence or / and picture, or / and strip, or / and CTU level.
[0365] Within a constrained search region (defined by the number of configurable or signal-transmitted CTUs, or the number of CTU lines, or the number of samples, prefixed from the upper, left, and / or upper left spatial regions), the template similarity between any adjacent / non-adjacent reference blocks and the current coded block is calculated, and the best N reference blocks with the closest similarity are selected as candidates in the template list.
[0366] Use a signal transmission to indicate whether to use the template matching method. If the flag is true, then a further signal transmission should be used to indicate which candidate in the template list should be used, and its index.
[0367] In another embodiment, both the encoder search method and the template matching method are used in combination. For example, an integer motion and fractional thinning method is first applied at the encoder, and then another template thinning is further applied at both the encoder and decoder sides. Since the encoder search method is already accurate enough, template thinning can be performed with high precision and in a smaller area. For example, motion thinning on the encoder side is performed with a maximum precision of half a pixel or a quarter pixel, while template thinning can be further performed with a precision of a quarter pixel, an eighth pixel, or a sixteen-pixel.
[0368] Fractional motion refinement
[0369] With or without a fractional motion search process, the initial motion vector (MV) can be identified. The initial MV can be adjusted for two reasons:
[0370] • For smaller signaling overhead, the initial Mv can be rounded to a specific precision or value to minimize the difference between the initial Mv and the selected predicted Mv value.
[0371] • For lower signaling overhead, a few of the least important bits in the initial MV can be discarded.
[0372] With or without making the above adjustments, it may be necessary to refine the initial Mv on the decoder side.
[0373] In one or more embodiments, method-based template matching can be used. In one example, an inverse L-shaped sample / pixel region adjacent to the coded block can be used as a matching template. The initial Mv can be refined at the integer pixel and / or fractional pixel level. The potential refinement set can be {1 / 4 pixel, 2 / 4 pixel, 3 / 4 pixel} or / and {1 / 8 pixel, 3 / 8 pixel, 5 / 8 pixel, 7 / 8 pixel}, and the refinement directions are two horizontal directions and two vertical directions (positive and negative values). The refined Mv that generates the prediction block with the most similar template is selected as the final Mv. Note that if the most similar template is selected, the selected refinement can be implicitly obtained by the decoder, or if multiple refined Mvs with N most similar templates are obtained, the selected refinement can be explicitly obtained by the encoder.
[0374] In one or more embodiments, additional flags can be transmitted via signaling to indicate whether this fractional motion refinement is applied. These additional flags can be transmitted as sequences, images, strips, or CTU levels.
[0375] Conditional sample / pixel padding for fractional interpolation
[0376] When using fractional Mv, interpolation operations may require a greater number of pixels / samples than the current block. The actual difference depends on the interpolation filter tap length. In cases where some pixels / samples are unavailable, a pixel / sample padding process may be necessary. Different padding schemes can be used.
[0377] In one or more embodiments, a repeating fill type can be used. Unavailable pixel / sample locations can be filled with the same value of the nearest available pixel / sample in the same row or column. This repeating fill can be performed first in the horizontal direction (left and right border fills) and then in the vertical direction (top and bottom border fills). Alternatively, this repeating fill can be performed first in the vertical direction (top and bottom border fills) and then in the horizontal direction (left and right border fills).
[0378] In one or more embodiments, a symmetrical fill type can be used. Unavailable pixel / sample locations can be filled with pixels at positions symmetrical to the fill boundary. This fill can be performed first in the horizontal direction (left or right boundary fill) and then in the vertical direction (top or bottom boundary fill). Alternatively, this repeating fill can be performed first in the vertical direction (top or bottom boundary fill) and then in the horizontal direction (left or right boundary fill).
[0379] In one or more embodiments, the filling process may be conditionally skipped or simplified.
[0380] In one embodiment, the filling process may be partially or completely skipped depending on the value of the fractional portion of the motion vector. In one example, if the horizontal or vertical portion of the motion vector is equal to zero, the corresponding vertical or horizontal filling may be skipped or not. If both directions of the motion vector are equal to zero, the filling process may be completely skipped or the entire filling process may still be performed.
[0381] In another embodiment, if all the required samples involved in the interpolation process related to a particular motion vector are checked to be valid (e.g., all samples involved in the interpolation process are located inside the valid reference region, such as...), then... Figure 8 , 10 If (as shown in Figure 11), then the process can be skipped entirely or the entire filling process can still be performed.
[0382] In another embodiment, regarding the interpolation process, the number of desired left or top samples located outside the reference block pointed to by the integer MV or the integer part of MV is one less than the number of desired right or bottom samples, and the padding size of the left and top samples can be reduced by one. For example, if a 12-tap interpolation filter is used for an IBC coded block with non-zero fractional parts in the horizontal and vertical directions of the motion vector, the padding size of the top and left samples is 5, while the padding size of the bottom and right samples is 6.
[0383] In another embodiment, in some cases, the interpolation process is only used for motion search (e.g., encoder fractional motion estimation) or motion reordering (e.g., ARMC), motion refinement (e.g., IBC-DBV, IBC template matching), and parameter generation for prediction refinement (e.g., template prediction generation for IBC-LIC parameter derivation), but not for the final prediction generation of the current coding block, and the padding process can be skipped (e.g., by forcing the fractional part to zero).
[0384] Interpolation filter switching
[0385] For various reasons, the interpolation filter may need to be switched. For example, if the image / video content is noisy and a smooth filter effect is desired, a longer tap length may be preferred. Conversely, if reduced padding complexity is required or the image / video content has rich textured edges, a shorter tap length may be preferred.
[0386] In one or more embodiments, filter switching can be determined on the decoder side by analyzing image / video content (such as gradient histograms), which does not require signaling bits.
[0387] In one or more other embodiments, filter switching can be evaluated on the encoder side and transmitted as a signal at different granularities (based on sequence, image, strip, CTU level, or region).
[0388] In one or more other embodiments, filter switching can be evaluated on the encoder side and transmitted by signal for different frame / strip types (e.g., I-strip, B-strip, and P-strip).
[0389] In one or more other embodiments, filter switching can be evaluated on the encoder side and transmitted as a signal for different video components (e.g., luminance and chrominance components, or Y, Cb and Cr components).
[0390] In one or more other embodiments, filter switching can be evaluated on the encoder side and transmitted as a signal for motion accuracy (e.g., 1 pixel, 4 pixels, or fractional pixels, such as 1 / 2, or / and 1 / 4, or / and 1 / 8 pixels, or / and 1 / 16 pixels).
[0391] In one or more other embodiments, filter switching can be evaluated on the encoder side and transmitted by signal for different interpolation scenarios (e.g., regular compensation prediction for the current block, or compensation prediction for templates such as the top template and / or left template of the current block).
[0392] In another embodiment, the filter switching provided in any combination of the scenarios specified above can be predefined and does not require signal transmission. For example, a 2-tap interpolation filter can always be used if the following three conditions are met: chroma component, template prediction generation, and natural content.
[0393] In another embodiment, the filter switching described above may depend on other flags and may not require signal transmission. For example, if a video block is marked as encoded in a specific mode (intra-frame TMP mode) or with a specific motion precision, the filter switching may be determined accordingly.
[0394] In one embodiment, filter switching is predefined according to a specific mode. For example, when using ARMC, IBC template matching mode, IBC-BVDP, and IBC-BVPC modes, the reference template generation for template matching cost calculation can use interpolation filters with shorter lengths (e.g., 2-tap, 4-tap, 6-tap) compared to default filters (e.g., default 12-tap or 8-tap filters for luma and default 6-tap filters for chroma). For IBC-LIC and processing IBC chroma samples, the default luma interpolation filter (e.g., 12-tap filter) and / or chroma interpolation filter (e.g., 6-tap filter) can be used respectively. In another example, for template matching-related tools such as ARMC, IBC template matching mode, IBC-BVDP, and IBC-BVPC modes, the reference template can be generated by ignoring fractional parts (e.g., truncating or rounding to integer values), and the interpolation and padding processes can be skipped entirely.
[0395] The filter switching methods provided above can be applied in any combination.
[0396] Multiple assumptions fractional intra-block copy
[0397] When multiple motion vectors (from motion search and / or motion refinement) are available, multiple prediction blocks can be generated. Multi-hypothesis intra-block replication can be used when averaging multiple similar blocks can produce better block predictions.
[0398] In one or more embodiments, the number of hypotheses may be predefined, configured, or signaled. Additionally, the weights used to average the multiple prediction hypotheses may also be predefined, configured, or signaled.
[0399] In one or more other embodiments, the number of multiple hypotheses can be implicitly determined on the decoder side. For example, if N prediction blocks can be generated, and the value transmitted by signal is N, which is outside the range (0 to N-1) of a valid single prediction block, then multiple hypotheses are enabled, and the average of all N prediction blocks can be used.
[0400] Multiple hypotheses can be generated from N motion / block vectors, where each motion / block vector can generate a specific motion-compensated prediction block, and N is a positive integer. The N motion / block vectors can be obtained from the same candidate list or different candidate lists. In one example, the N motion / block vectors can be obtained from the same IBC merge candidate list or AMVP list, or partially from both the IBC merge candidate list and the IBC AMVP candidate list. In another example, the N motion / block vectors can be obtained entirely or partially based on an intra-template matching method.
[0401] Interpolation process for template matching
[0402] When applying template-based adaptive reordering (ARMC-TM) to IBC merging and / or AMVP mode, or / and applying template-based motion refinement to IBC merging and / or AMVP mode, it may be necessary to adaptively use fractional motion / block vectors.
[0403] For Template-Based Adaptive Reordering (ARMC-TM), it may be necessary to compute a template-based distortion cost for each motion / block vector candidate. Since this template-based distortion cost is only used for candidate reordering and not for the final compensation prediction, the fractional part of each motion / block vector candidate (if a non-zero fractional part exists) may or may not be considered for template-based distortion computation. Specifically, in one or more examples, fractional motion-based interpolation may or may not be performed on each motion / block vector candidate without a non-zero fractional part.
[0404] Similarly, when calculating the template distortion cost for each motion refinement location, it may or may not be necessary to consider the fractional part for each location.
[0405] Signal transmission of motion information
[0406] When motion vectors in intra-block copying support multiple levels of precision, permissible signal transmission methods can be defined accordingly.
[0407] In one or more embodiments, only one level of precision with zero motion vector difference is allowed. This precision may be predefined, configurable, or transmitted via signaling. For example, this precision may be predefined as the highest precision supported by the motion vector, such as 1 / 4 pixel or 1 / 8 pixel.
[0408] In one or more embodiments, multiple precisions with zero motion vector difference are allowed. These multiple precisions can be predefined, configurable, or transmitted via signals. For example, such precision can be predefined as the highest or second highest precision supported by the motion vector, such as 1 / 4 pixel and 1 pixel. When multiple precisions with zero motion vector difference are allowed, after the zero motion vector difference indication transmitted via signals (1 flag or 1 binary bit), one or more additional flags are transmitted to indicate which precision is used.
[0409] When multiple precisions are supported, the current precision flag can be signed in different ways. In one example, a flag indicating whether the current precision is greater than 0 is first signaled. If so, another flag indicating whether the current precision is greater than 1 is further signaled. Alternatively, a second flag indicating whether the current precision is greater than 1 can be implicitly obtained at the decoder without explicit signaling. In one example, the value of the motion / block vector difference can be used to achieve this (e.g., even or odd motion / block vector differences can indicate a specific motion precision value). In this document, values of 0, 1, or other values greater than 1 can be predefined or configured to represent different motion vector (or motion vector difference) precisions (e.g., 0 for 1 pixel precision, 1 for 1 / 2 pixel precision, 2 for 1 / 4 pixel precision, and 3 for 1 / 8 pixel precision).
[0410] In one or more embodiments, multiple MV candidates in the IBC merge / AMVP motion candidate list are divided into different groups. In one example, the grouping criterion could be MV precision, where MV candidates in the same group have the same actual MV precision. Actual MV precision is defined as the MV precision after right-shifting all the least significant zeros of MV.
[0411] IBC Merge / AMVP Campaign Candidate List Construction
[0412] In one or more embodiments, multiple MV candidates in the IBC merge / AMVP motion candidate list are grouped into different groups. In one example, the grouping criterion could be Mv precision, where Mv candidates in the same group have the same actual MV precision. Actual MV precision is defined as the Mv precision after right-shifting all the least significant zeros of the MV.
[0413] In one or more other embodiments, multiple lists of IBC merge / AMVP motion candidates are created, in addition to the exit list. For each list, only Mv candidates with the same actual Mv precision are added. Similarly, the actual Mv precision is defined as the Mv precision after right-shifting all the least significant zeros of Mv.
[0414] In cases where multiple candidate list groups and / or multiple candidate lists are generated, the group index and / or candidate list index need to be determined before the actual Mv candidate index can be decided. In one or more other examples, the group index and / or candidate list index can be evaluated on the encoder side and then signaled to the decoder. In one or more other still examples, the group index and / or candidate list index can be inherited from a specific neighboring block without explicit signaling.
[0415] Combination with intra-frame template matching
[0416] When the motion vector for intra-block copying is determined, a prediction block can be generated based on that motion vector. When combined with intra-template matching, the prediction block generated by intra-block copying is further refined through intra-template matching. Specifically, for a predefined search range around the prediction block generated by intra-block copying, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. In this method, the prediction block generated by intra-block copying is considered the starting block position, which is used to guide the subsequent block search process in the intra-template matching method.
[0417] The combination of intra-block copying and intra-template matching prediction modes can be signaled at the CU level using a dedicated flag. Alternatively, the original intra-block copying flag above the original intra-template matching mode flag can be used to indicate the combination of intra-block copying and intra-template matching prediction modes.
[0418] In one or more other examples, the combination of intra-block copying and intra-template matching prediction can generate improved motion vectors.
[0419] In one example, the intra-block copy (IBC) mode provides an initial motion vector, which can be further refined using an intra-template matching method.
[0420] In another example, motion / block vectors obtained from the intra-template matching method can be reused to generate an IBC merge candidate list or an IBC AMVP candidate list. In this case, motion / block vectors generated for spatially adjacent or non-adjacent neighboring blocks can be cached or saved. In one example, the motion / block vectors of neighboring blocks encoded in the intra-template matching method can be saved in a history motion vector table. Alternatively or additionally, the motion / block vectors of neighboring blocks encoded in the intra-template matching method can be saved in the encoder's local cache and then reused by the encoder during the motion search process.
[0421] In other examples, combining intra-block duplication and intra-template matching prediction can generate improved prediction blocks. In one example, two prediction blocks can be generated separately using intra-block duplication and intra-template matching prediction methods, and a weighted average of these two prediction blocks can be generated to represent the final prediction block of the current coded block. In another example, multiple prediction blocks (e.g., N>1) can be generated separately, where M out of the N prediction blocks (e.g., M<=N) can be generated by intra-block duplication, and S out of the N prediction blocks (e.g., S<=N) can be generated by intra-template matching prediction. When combining N prediction blocks (e.g., N>1), the weight values can be obtained using matching cost (e.g., an example of matching cost calculation can be based on an L-shaped template, where a higher matching cost indicates a lower weight value, and a lower matching cost indicates a higher weight value) or a least-squares flavor method.
[0422] Combination with IBC merging mode utilizing block vector difference
[0423] With fractional Mv supported, the IBC merging mode utilizing block vector differences can be extended by employing more candidate distance values. In one or more other embodiments, two distance sets may exist. The first set is the exit integer distance set, while the second set is added separately for fractional distances. In one example, the fractional distance set could be {1 / 8 pixel, 2 / 8 pixel, 3 / 8 pixel, 4 / 8 pixel, 5 / 8 pixel, 6 / 8 pixel, 7 / 8 pixel}. In another example, the fractional distance set could be {1 / 8 pixel, 2 / 8 pixel, 4 / 8 pixel}. The BVD directions of the second set are also two horizontal directions and two vertical directions.
[0424] In another example, the new distance set can be defined as {1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 14 pixel, 16 pixel, 18 pixel, 20 pixel, 22 pixel, 24 pixel, 26 pixel, 28 pixel, 30 pixel, 32 pixel}.
[0425] In another example, the new distance set can be defined as {1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 16 pixel, 20 pixel, 24 pixel, 28 pixel, 32 pixel, 36 pixel, 40 pixel, 44 pixel, 48 pixel, 52 pixel, 56 pixel}.
[0426] In another example, the new distance set can be defined as {1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 14 pixel, 16 pixel, 18 pixel, 20 pixel, 22 pixel, 24 pixel, 26 pixel, 28 pixel, 30 pixel, 32 pixel, 34 pixel}.
[0427] In another example, the new distance set can be defined as {1 pixel, 2 pixels, 4 pixels, 6 pixels, 8 pixels, 12 pixels, 16 pixels, 20 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels}.
[0428] The distance set proposed above can be selected or adaptively determined using predefined rules:
[0429] In one embodiment, only one of the above sets is used.
[0430] In another embodiment, multiple sets can be used, and signal transmission-based selection can be used to switch between different sets.
[0431] In another embodiment, multiple sets can be used, and conditional selection can be used to switch between different sets without the need for signal transmission.
[0432] In one example, a distance set can be selected when one or another encoding mode (e.g., RRIBC mode, palette mode) is allowed at certain levels (e.g., sequence, image, stripe), while another distance set can be selected when one or another encoding mode is not allowed at certain levels.
[0433] Combination of IBC-LIC and IBC utilizing fractional motion vectors
[0434] In IBC-LIC, after generating prediction blocks from reference blocks using integer copying (using integer motion vectors or the integer part of fractional motion vectors) or interpolation filtering (using fractional motion vectors), a linear filtering process is further applied to refine the prediction blocks. Note that the parameters of the linear filtering process are obtained using a reference template generated from the same motion vectors used to generate the current block. When IBC blocks encoded with fractional motion vectors are combined with IBC-LIC modes, the reference templates for LIC parameter derivation can be generated in different ways:
[0435] In one example, the reference template used for IBC-LIC parameter derivation is generated by modifying an integer motion vector from the fractional motion vector used in the current IBC block. This modification can be based on truncation operations (e.g., directly discarding the fractional portion of the fractional motion vector) or rounding operations (e.g., rounding to the nearest integer value).
[0436] In another example, the reference template used for IBC-LIC parameter derivation is generated using the same fractional motion vector used in the current IBC block. However, the interpolation filter used to generate the reference template can be the same as or different from the interpolation filter used to generate the prediction block for the current block. For example, given a 1 / 4-pixel motion vector of the current block, a 12-tap filter can be used for prediction generation, but when generating the reference template for IBC-LIC, the same 12-tap filter or a different filter (e.g., a 4-tap, a 2-tap, or a 0-tap filter) can be used. Note that if some samples are outside the effective IBC reference area, a shorter tap filter may result in less computation and / or less padding. If a different filter is used to generate the reference template, the filter can be predefined or transmitted via signal transmission at different levels (based on sequence, picture, strip, CTU level, or region).
[0437] DBV combined with IBC utilizing fractional or integer motion vectors
[0438] Several methods have been proposed to improve the DBV mode when it is used for IBC coded blocks with integer or fractional motion vectors:
[0439] In the current DBV mode, luminance motion vectors are selected from a co-located luminance block at five locations according to a fixed order of the inspection process (e.g., center block, top left block, top right block, bottom left block, and bottom right block). Improvements to luminance block selection are proposed, either through signal transmission by the encoder or determination by the decoder. One example of the decoder-determined method is to use an L-shaped chromaticity template matching cost to sort N (N<=5) luminance block vectors. Another example is based on block content analysis (e.g., gradient-based edge detection, prioritizing luminance blocks on the same edge).
[0440] In the current DBV mode, template-based MV thinning is performed. However, the current thinning only allows positive thinning (e.g., +1 or +2 on the current MV component). It is proposed that negative thinning (e.g., -1 or -2) should also be allowed. When fractional motion vectors are supported in the selected luma block, it is proposed that fractional thinning (e.g., 1 / 2 pixel, 1 / 4 pixel, -1 / 2 pixel, 3 / 4 pixel thinning) be allowed on the current motion vector.
[0441] In the current DBV mode, template-based MV thinning is performed separately for the Cb and Cr components. If the thinned MV differs for Cb and Cr, potential artifacts (e.g., chromaticity component misalignment) may be observed in the predicted samples. Improvements to MV thinning in different ways are proposed:
[0442] In one embodiment, MV refinement is allowed only for one chromaticity component (Cb or Cr), and the refined MV is reused when predicting for another component (Cr or Cb).
[0443] In another embodiment, MV refinement is performed jointly for Cb and Cr. In this method, the reference template cost is calculated by combining the matching errors of the Cb and Cr templates.
[0444] Different methods can be proposed when using the selected luminance motion vector to generate the chrominance motion vector:
[0445] In one embodiment, the fractional portion of the luminance motion vector (if available) can be truncated or rounded to an integer value. In this way, subsequent prediction generation and MV refinement are performed using integer motion vectors, and no interpolation or padding processes are required.
[0446] In another embodiment, the fractional portion of the luminance vector (if available) can be retained to obtain the chroma motion vector, and chroma prediction generation requires an interpolation process; however, the fractional portion of the chroma motion vector can be truncated or rounded to an integer value only during the template-based MV thinning process. After thinning, the fractional portion is added back to the thinned chroma motion vector. In this case, the template-based MV thinning process does not require interpolation and padding, but chroma prediction generation still requires interpolation and padding.
[0447] In another embodiment, the fractional part of the luminance vector can be considered for both prediction generation and MV refinement.
[0448] When the fractional part is available for the resulting chroma motion vector, the prediction generation process and / or the mv refinement process may require interpolation filtering. Different or the same interpolation filters can be applied to both processes. In one example, a default interpolation filter (such as the default 6-tap chroma filter) can be applied to both processes. Alternatively, shorter filters (such as 2-tap or 4-tap filters) can be applied to the chroma prediction generation process and / or the mv refinement process.
[0449] Given a specific interpolation filter, computational complexity and bandwidth consumption are more expensive for smaller block sizes. For some video formats such as YUV420, the minimum size of the chroma block (e.g., a 2×2 block) can be smaller than the minimum size of the corresponding luma block (e.g., a 4×4 block), so the worst-case complexity of the chroma component (e.g., here the worst-case refers to the complexity of the interpolation process performed on the coded block with the smallest size) is higher than the worst-case complexity of the luma component. It may be desirable to constrain the worst-case interpolation complexity of the chroma component. Under this consideration, it may be desirable to completely avoid the interpolation filtering process (e.g., without a fractional part or ignoring the fractional part of the chroma motion vector) or to simplify the interpolation filtering process by using shorter tap filters (e.g., 2-tap filters). Alternatively, block size-dependent filter switching can be proposed. In one example, for chroma blocks with smaller block sizes (e.g., 2×2, 2×4, 4×2, etc.), the fractional part of the luma motion vector (if available) can be truncated or rounded to an integer value. In another example, for chroma blocks with smaller block sizes (e.g., 2×2, 2×4, 4×2, etc.), interpolation filters with fewer taps than the default chroma interpolation filter can be used (e.g., 2-tap filters, 4-tap filters). The selected filters can be predefined, or explicitly (e.g., by a dedicated flag indicating a specific filter index) or implicitly (e.g., depending on other existing flags, such as another coding mode) signaled at different granularities (e.g., based on sequence, picture, stripe, CTU level, or region).
[0450] Combination of RRIBC and IBC using fractional motion vectors
[0451] When RIBC is applied to an IBC-coded block, the block's motion vector can indicate one of three flip types: no flip, horizontal flip, and vertical flip. For horizontal or vertical flip types, the vertical or horizontal component of the IBC motion vector is set to zero.
[0452] When the RRIBC mode is combined with the regular IBC AMVP mode, encoder-side fractional motion estimation and / or decoder-side motion refinement may or may not be allowed for motion vectors indicating horizontal or vertical flips.
[0453] When fractional motion vectors are not allowed for either of the two RIBC flip modes (horizontal flip and vertical flip types), motion precision is limited to integer pixels (e.g., 1 pixel or 4 pixels). Therefore, for IBC AMVP mode, motion precision signal transmission is also limited to integer pixels. Alternatively, when zero mvd is determined on the decoder side, the corresponding motion precision (e.g., the highest supported integer pixel, such as 1 pixel) can be obtained directly without further signal transmission.
[0454] Combination of IBC-BVPC and IBC using fractional motion vectors
[0455] When IBC-BVPC is applied to an IBC coded block, the motion vector of the block may have two non-empty components or one non-empty component (e.g., only the horizontal component or the vertical component is not zero).
[0456] When the IBC-BVPC mode is combined with the regular IBC AMVP mode, for motion vectors with only one non-empty component (e.g., only the horizontal component or the vertical component is non-zero), encoder-side fractional motion estimation and / or decoder-side motion refinement may or may not be allowed.
[0457] When the IBC-BVPC mode, which has only one non-empty component, does not allow fractional motion vectors, the motion precision is limited to integer pixels (e.g., 1 pixel or 4 pixels). Therefore, for the IBC AMVP mode, the motion precision signal transmission is also limited to integer pixels. Alternatively, when zero mvd is determined on the decoder side, the corresponding motion precision (e.g., the highest supported integer pixel, such as 1 pixel) can be obtained directly without further signal transmission.
[0458] Combination of IBC-BVDP and IBC utilizing fractional motion vectors
[0459] When the IBC-BVDP mode is applied to an IBC block with fractional motion vectors, the most significant bits of the magnitude suffixes for the horizontal and vertical components of the BVD can be predicted in different ways:
[0460] If the fractional part of the MVD prior to prediction is non-zero (e.g., the horizontal component and / or the vertical component has a non-zero fractional part), the most significant binary bit selected for the magnitude suffix to be used for prediction may or may not include the fractional binary bit.
[0461] In one embodiment, when fractional bits are included for prediction, the generated latent motion vectors are sorted together without grouping. For example, if a total of 4 bits are predicted, with 2 bits from the integer part of the MVD and the other 2 bits from the fractional part of the MVD, the total size of the latent motion vectors is 16 before considering sign prediction. When sorted together, these 16 vectors are compared together with their corresponding template matching costs.
[0462] In another embodiment, when fractional bits are included for prediction, the generated latent motion vectors are sorted separately by grouping. For example, if a total of 4 bits are predicted, with 2 bits from the integer part of the MVD and the other 2 bits from the fractional part of the MVD, the total size of the latent motion vectors is 16 before considering sign prediction. When sorting separately, the prediction of fractional bits is performed separately from the prediction of integer bits by fixing one combination of two integer bits and then predicting two fractional bits, or by fixing one combination of fractional bits and then predicting two integer bits.
[0463] In another embodiment, when the fractional bits are not included for prediction, the fractional bits are always signaled as the current bypass mode at the entropy coding level. Once the fractional bits have been signaled, the remaining integer bits can be predicted as the original IBC-BVDP mode. In this case, the template matching cost calculation may or may not take into account the bypassed fractional bits signaled.
[0464] In one embodiment, fractional bits are considered for integer bit prediction, such that a reference template is generated by interpolation filtering. The interpolation filter used in the reference template generation can be predefined (e.g., always a 2-tap filter) or switched at different granularities (e.g., based on sequence, image, strip, CTU level, or region).
[0465] In another embodiment, fractional bits are not considered for integer bit prediction, such that the reference template is generated without interpolation filtering (e.g., based on truncation operations (e.g., directly discarding the fractional portion of the fractional motion vector) or rounding operations (e.g., rounding to the nearest integer value)).
[0466] Combination of paired candidates with other IBC coding modes
[0467] In the current IBC model, new paired IBC candidates can be generated by averaging two previous IBC candidates in the IBC merge / AMVP list.
[0468] When a pair of candidates is generated by averaging two fractional motion candidates, the new candidate can have higher accuracy than the currently supported highest accuracy. For example, averaging a 1 / 2 pixel mv and a 1 / 4 pixel mv can result in a new candidate with 1 / 8 pixel accuracy.
[0469] In one embodiment, the new pair of candidates can be rounded to or truncated to a supported motion precision. For example, a new 1 / 8 pixel pair of candidates can be truncated to 1 / 4 pixel (discarding fractions higher than 1 / 4 pixel) or rounded to 1 / 4 pixel.
[0470] In another embodiment, the new paired candidate can remain at the current high precision. For the IBC-AMVP list, a subsequent rounding process will be performed on the encoder side. For the IBC merge list, the current precision will always be retained.
[0471] When generating paired candidates by averaging two predicted candidates, these two candidates can indicate the same or different RRIBC flipping patterns. For example, one candidate indicates a horizontal flip, while the other indicates no flip.
[0472] In one embodiment, if two candidates have the same RRIBC flip mode (e.g., no flip mode, horizontal flip mode, vertical flip mode), then a new pair of candidates can be set to have the same RRIBC flip mode. Otherwise, a new pair of candidates can be set to have a no flip mode.
[0473] In another embodiment, new paired candidates can always be set to have a non-flipped mode, regardless of whether the two candidates have the same or different RRIBC flip modes.
[0474] When generating paired candidates by averaging two predicted candidates, the two candidates can indicate the same or different IBC-LIC flags.
[0475] In one embodiment, if two candidates have the same IBC-LIC flag (e.g., both are IBC-LIC enabled or disabled), a new pair of candidates can be set to have the same IBC-LIC flag. Otherwise, a new pair of candidates can be set to have IBC-LIC disabled.
[0476] In another embodiment, if one of the two candidates has IBC-LIC enabled, the new paired candidate can always be set to have IBC-LIC enabled.
[0477] In another embodiment, new paired candidates can always be set to have IBC-LIC off, regardless of whether the two candidates have the same or different IBC-LIC flags.
[0478] Figure 20 A computing environment 2010 coupled to a user interface 2050 is shown. The computing environment 2010 may be part of a data processing server. The computing environment 2010 includes a processor 2020, memory 2030, and input / output (I / O) interfaces 2040.
[0479] Processor 2020 typically controls the overall operation of computing environment 2010, such as operations associated with display, data acquisition, data communication, and image processing. Processor 2020 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 2020 may include one or more modules that facilitate interaction between processor 2020 and other components. Processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0480] Memory 2030 is configured to store various types of data to support the operation of computing environment 2010. Memory 2030 may include predefined software 2032. Examples of such data include instructions for any application or method operating on computing environment 2010, video datasets, image data, etc. Memory 2030 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0481] I / O interface 2040 provides an interface between processor 2020 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, home button, start scan button, and stop scan button. I / O interface 2040 can be coupled to encoder and decoder.
[0482] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a memory 2030 containing multiple programs and / or storing a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The multiple programs can be executed by a processor 2020 in a computing environment 2010 to perform the above-described methods. In an embodiment, the multiple programs can be executed by a processor 2020 in a computing environment 2010 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 2010 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 2020 in the computing environment 2010 to perform the above-described decoding method based on the received bitstream or data stream. In another example, multiple programs can be executed by the processor 2020 in the computing environment 2010 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 2020 in the computing environment 2010 to (e.g., to...) Figure 3 The video decoder 30 in the medium transmits the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by an encoder (e.g., Figure 2 The video encoder 20 in the video is generated using, for example, the encoding method described above, for use by the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media may be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0483] In one embodiment, a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method is provided.
[0484] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 2020); and a non-transitory computer-readable storage medium or memory 2030 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.
[0485] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above, is also provided. In another embodiment, a computer program product including, for example, multiple programs in a memory 2030, which can be executed by a processor 2020 in a computing environment 2010 to perform the methods described above, is also provided. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0486] In an embodiment, the computing environment 2010 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0487] Figure 21 This is a flowchart illustrating a method for video decoding according to an example of the present disclosure. In step 2101, the method includes: obtaining a syntax element associated with an intra-block copy (IBC) mode by a decoder, wherein the syntax element is signaled jointly or separately by an encoder using context-adaptive binary arithmetic coding (CABAC), which uses different context bits for different stripe types, different frame types, different video components, or different video resolutions.
[0488] In some embodiments of this disclosure, multiple CABAC context windows are predefined for different stripe types or different frame types.
[0489] In some embodiments of this disclosure, syntax elements are used to indicate motion information accuracy, interpolation filter selection for motion compensation block prediction, or interpolation filter selection for template prediction.
[0490] In some embodiments of this disclosure, different video components include luminance components and chrominance components.
[0491] Figure 22 This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 2201, the method includes: obtaining fractional motion information of the current block by a decoder in intra-block copy (IBC) mode. In step 2202, the method includes: obtaining a block vector (BV) of the current block by the decoder based on the fractional motion information. In step 2203, the method includes: obtaining a predicted block of the current block by the decoder based on the BV. In step 2204, the method includes: in response to a satisfied condition, performing a padding process on one or more samples associated with an interpolation filter by the decoder, or partially or completely skipping the padding process.
[0492] In some embodiments of this disclosure, the condition includes: determining that one or more samples associated with the interpolation filter are unavailable and that the fractional portion of the BV has a value of zero, the fractional portion including a horizontal portion or a vertical portion; and that the decoder partially or completely skips the padding process for one or more samples, including: skipping the padding process associated with the fractional portion of the BV, the value of which is zero.
[0493] In some embodiments of this disclosure, the condition includes: all samples associated with the interpolation filter are located within the valid reference region, and the decoder partially or completely skips the padding process for one or more samples, including: completely skipping the padding process.
[0494] In some embodiments of this disclosure, the process of filling one or more samples by the decoder includes determining the fill size of a left or top sample located outside a reference block pointed to by an integer BV or an integer portion of BV to be one sample size smaller than the fill size of a right or bottom sample located outside the reference block.
[0495] In some embodiments of this disclosure, the method further includes: the decoder performing a padding process for any one or any combination of the following processes: motion search, motion refinement, or parameter generation for prediction refinement; or skipping the padding process when an interpolation filter is applied to the prediction block to obtain a filtered prediction for the current block.
[0496] Figure 23 This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 2301, the method includes: obtaining fractional motion information of the current block by the decoder in intra-block copy (IBC) mode. In step 2302, obtaining the block vector (BV) of the current block by the decoder based on the fractional motion information. In step 2303, the method includes: obtaining a predicted block of the current block by the decoder based on the BV. In step 2304, the method includes: determining an interpolation filter or skipping the interpolation process by the decoder in response to a satisfied condition, wherein the condition is configured to indicate whether to switch the interpolation filter or skip the interpolation process, and whether to apply the interpolation filter to the predicted block to obtain a filtered prediction for the current block.
[0497] In some embodiments of this disclosure, the condition includes: receiving, by the decoder, a syntax element transmitted by the encoder in the form of a signal, the syntax element being used to indicate switching the interpolation filter for any one or any combination of a particular stripe type, a particular frame type, a particular video component, a particular IBC mode, or a particular interpolation scenario.
[0498] In some embodiments of this disclosure, the condition includes any one or any combination of the following: processing a specific strip type, processing a specific frame type, processing a specific video component, applying a specific motion precision, applying a specific IBC mode, or applying a specific interpolation scenario.
[0499] In some embodiments of this disclosure, the specific video component includes IBC luma samples or IBC chroma samples; or, the specific IBC mode includes: adaptive reordering with merged candidates (ARMC) mode, template matching mode, block vector difference prediction for IBC blocks (IBC-BVDP) mode, block vector prediction candidate clustering and block vector difference sign derivation for reconstructing reordered IBC (IBC-BVPC) mode, or IBC with local illumination compensation (IBC-LIC) mode.
[0500] In some embodiments of this disclosure, the condition includes applying ARMC mode, template matching mode, IBC-BVDP mode, or IBC-BVPC mode, and determining the interpolation filter by the decoder or skipping the interpolation process by the decoder includes: determining the interpolation filter as an interpolation filter having a smaller length compared to a predefined length; or skipping the interpolation process.
[0501] In some embodiments of this disclosure, the condition includes applying the IBC-LIC mode or processing IBC chroma samples, and the decoder determining the interpolation filter includes: determining the interpolation filter as a preset luminance interpolation filter or a preset chroma interpolation filter.
[0502] Figure 24This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 2401, the method includes: obtaining a plurality of basic candidates from an intra-block copy (IBC) merging list by a decoder. In step 2402, obtaining a plurality of merged pattern (MBVD) refinement positions with block vector differences for each basic candidate by the decoder based on at least one candidate distance set. In step 2403, the method includes: reordering the plurality of MBVD refinement positions by the decoder. In step 2404, the method includes: selecting a first number of MBVD refinement positions with the lowest template SAD cost by the decoder. In step 2405, the method includes: obtaining block vector differences by a decoder based on a first number of MBVD refinement locations, wherein the at least one candidate distance set includes any one or any combination of the following sets: {1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 14 pixel, 16 pixel, 18 pixel, 20 pixel, 22 pixel, 24 pixel, 26 pixel, 28 pixel, 30 pixel, 32 pixel}; {1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 16 pixel, 20 pixel, 24 pixel, 28 pixel, 32 pixel}. {1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 40 pixels, 44 pixels, 48 pixels, 52 pixels, 56 pixels}; {1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 10 pixels, 12 pixels, 14 pixels, 16 pixels, 18 pixels, 20 pixels, 22 pixels, 24 pixels, 26 pixels, 28 pixels, 30 pixels, 32 pixels, 34 pixels}; or, {1 pixel, 2 pixels, 4 pixels, 6 pixels, 8 pixels, 12 pixels, 16 pixels, 20 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels}.
[0503] In some embodiments of this disclosure, obtaining multiple MBVD refinement positions for each basic candidate by the decoder based on at least one candidate distance set includes: selecting at least one distance set from at least one candidate distance set based on syntax elements transmitted by the encoder using signals; or selecting at least one distance set corresponding to a specific coding mode from at least one candidate distance set in response to satisfying a condition associated with a specific coding mode.
[0504] In some embodiments of this disclosure, the condition includes applying a specific encoding pattern at a specific level.
[0505] Figure 25This is a flowchart illustrating a method for video decoding according to an example of the present disclosure. In step 2501, the method includes: obtaining fractional motion information of the current block by the decoder in intra-block copy (IBC) mode. In step 2502, the method includes: obtaining a block vector (BV) of the current block by the decoder based on the fractional motion information. In step 2503, the method includes: obtaining a predicted block of the current block by the decoder based on the BV. In step 2504, the method includes: obtaining a reference template by the decoder based on the BV and a template of the current block. In step 2505, the method includes: obtaining parameters of a linear filter by the decoder based on the reference template. In step 2506, the method includes: applying a linear filter to the predicted block by the decoder to obtain a filtered prediction for the current block.
[0506] In some embodiments of this disclosure, obtaining a reference template by the decoder based on the BV and the template of the current block includes: obtaining an integer motion vector by modifying the BV; and obtaining a reference template based on the integer motion vector and the template of the current block.
[0507] In some embodiments of this disclosure, obtaining an integer motion vector by modifying BV includes: discarding the fractional portion of BV; or, rounding the value of BV to the nearest integer value.
[0508] In some embodiments of this disclosure, obtaining a reference template by the decoder based on the BV and the template of the current block includes: in response to determining that one or more samples associated with a first interpolation filter are unavailable, performing a filling process on one or more samples to obtain a reference template, wherein the first interpolation filter is the same as or different from a second interpolation filter to be applied to the prediction block.
[0509] In some embodiments of this disclosure, the first interpolation filter is different from the second interpolation filter to be applied to the prediction block, and the first interpolation filter is predefined or transmitted by signaling at a specific level.
[0510] Figure 26 This is a flowchart illustrating a method for video decoding according to an example of the present disclosure. In step 2601, the method includes: obtaining fractional motion information of multiple luma blocks located at multiple predefined positions relative to the position of the current chroma block by a decoder in intra-block copy (IBC) mode. In step 2602, the method includes: obtaining multiple luma block vectors (BVs) of the multiple luma blocks by the decoder based on the fractional motion information. In step 2603, the method includes: determining a chroma BV by the decoder based on the multiple luma BVs of the multiple luma blocks. In step 2604, the method includes: obtaining a predicted chroma block by the decoder based on the chroma BV.
[0511] In some embodiments of this disclosure, determining the chromaticity BV by the decoder based on multiple luminance BVs of multiple luminance blocks includes: calculating the chromaticity template matching cost by the decoder based on the luminance BV of each of the multiple luminance blocks and the current chromaticity block; and determining the chromaticity BV as the luminance BV with the lowest chromaticity template matching cost by the decoder.
[0512] In some embodiments of this disclosure, determining the chromaticity BV by the decoder based on multiple luminance BVs of multiple luminance blocks includes: performing gradient-based edge detection on the multiple luminance blocks; and determining the chromaticity BV by the decoder as the luminance BV of a luminance block having the same edge as the chromaticity block.
[0513] In some embodiments of this disclosure, the method further includes: obtaining a template-based BV refinement by a decoder based on intra-frame template matching (ITM); and obtaining a refined chroma BV by a decoder based on the template-based BV refinement and the chroma BV.
[0514] In some embodiments of this disclosure, prediction refinement includes positive refinement or negative refinement.
[0515] In some embodiments of this disclosure, obtaining a refined chroma BV by the decoder based on the template-based BV refinement and the chroma BV includes: applying the template-based BV refinement to a first color component of the chroma BV; and reusing the predicted refinement to a second color component of the chroma BV.
[0516] In some embodiments of this disclosure, obtaining template-based BV refinement by the decoder includes: obtaining template-based BV refinement based on a reference template cost calculated by combining the matching error of the first color template component and the second color template component; and obtaining refined chroma BV by the decoder based on the template-based BV refinement and the chroma BV includes: applying the template-based BV to the first color component and the second color component of the chroma BV.
[0517] In some embodiments of this disclosure, the method further includes: truncating or rounding the value of chroma BV to an integer value.
[0518] In some embodiments of this disclosure, determining the chroma block vector (BV) by the decoder based on multiple luminance BVs of multiple luminance blocks includes:
[0519] Multiple values of luminance BV are truncated or rounded into multiple integer values; and the chrominance BV is determined by the decoder based on the multiple integer values.
[0520] In some embodiments of this disclosure, the method further includes: applying a first interpolation filter to the predicted chroma block to obtain a filtered prediction for the current chroma block, wherein the first interpolation filter is the same as or different from a second interpolation filter used to obtain template-based BV refinement, and the first interpolation filter or the second interpolation filter is predefined.
[0521] In some embodiments of this disclosure, the method further includes: in response to determining that the size of the current chroma block is smaller than a predefined size, truncating or rounding the value of the chroma BV to an integer value, or using a filter with a smaller size compared to the redefined size when performing a first interpolation process on the predicted chroma block or a second interpolation process on the chroma BV.
[0522] In some embodiments of this disclosure, filters with smaller dimensions are predefined or used for signal transmission at a specific level.
[0523] Figure 27 This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 2701, the method includes: obtaining fractional motion information of the current block by the decoder in Intra-Block Copy (IBC) mode. In step 2702, the method includes: obtaining the block vector (BV) of the current block by the decoder based on the fractional motion information. In step 2703, the method includes: in response to applying a Reconstruction Reordering IBC (RRIBC) mode, setting either the horizontal or vertical component of the BV to zero by the decoder according to the flip type of the RRIBC mode. In step 2704, the method includes: obtaining the predicted block of the current block by the decoder based on the BV.
[0524] In some embodiments of this disclosure, the method further includes: performing a motion refinement process on only the horizontal or vertical component of the BV in response to applying the RRIBC mode and the IBC Advanced Motion Vector Prediction (AMVP) mode.
[0525] In some embodiments of this disclosure, the method further includes: limiting the motion precision to integer pixels in response to the application of an RRIBC mode that disallows fractional motion vectors, or receiving a motion vector difference (MVD) equal to zero.
[0526] Figure 28This is a flowchart illustrating a method for video decoding according to an example of the present disclosure. In step 2801, the method includes: obtaining fractional motion information of the current block by the decoder in Intra-Block Copy (IBC) mode. In step 2802, the method includes: obtaining the block vector (BV) of the current block by the decoder based on the fractional motion information. In step 2803, the method includes: skipping motion refinement of the BV, or skipping motion refinement of the empty components of the BV, including horizontal or vertical components, in response to applying the Block Vector Prediction Candidate Clustering and Block Vector Difference Sign Derivation (IBC-BVPC) mode and the IBC Advanced Motion Vector Prediction (AMVP) mode for reconstructing reordered IBC. In step 2804, the method includes: obtaining the predicted block of the current block by the decoder based on the BV.
[0527] In some embodiments of this disclosure, the method further includes: limiting the motion precision to integer pixels in response to the application of the IBC-BVPC mode disallowing fractional motion vectors, or receiving a motion vector difference (MVD) equal to zero.
[0528] Figure 29 This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 2901, the method includes: obtaining amplitudes of block vector difference (BVD) components by a decoder, each of these BVD components including an integer BVD component portion and a fractional BVD component portion. In step 2902, the method includes: obtaining a context-coded BVD symbol prediction index by a decoder. In step 2903, the method includes: obtaining a plurality of BV candidates by a decoder by creating a combination between BVD symbols and absolute BVD values and adding that combination to a BV prediction value, each BV candidate including an integer portion and a fractional portion. In step 2904, the method includes: obtaining a BVD symbol prediction cost for each BV candidate by a decoder based on a template matching cost. In step 2905, the method includes: sorting the plurality of BV candidates by a decoder based on the BVD symbol prediction costs. In step 2906, the method includes: obtaining a BVD symbol prediction index corresponding to a true BVD symbol by a decoder.
[0529] In some embodiments of this disclosure, the BVD symbol prediction index includes integer bits and fractional bits.
[0530] In some embodiments of this disclosure, sorting multiple BV candidates by the decoder based on BVD symbol prediction costs includes: obtaining a first sorting result by sorting the multiple BV candidates based on a first BVD symbol prediction cost, the first BVD symbol prediction cost being obtained based on the integer portion of each BV candidate; and obtaining a second sorting result by sorting the multiple BV candidates based on a second BVD symbol prediction cost, the second BVD symbol prediction cost being obtained based on the fractional portion of each BV candidate, in response to a first combination of integer bits of the BVD symbol prediction index already obtained based on the first sorting result. The decoder obtaining the BVD symbol prediction index corresponding to the true BVD symbol includes: obtaining a first combination of integer bits of the BVD symbol prediction index based on the first sorting result; and obtaining a second combination of fractional bits of the BVD symbol prediction index based on the second sorting result.
[0531] In some embodiments of this disclosure, sorting multiple BV candidates by the decoder based on BVD symbol prediction cost includes: obtaining a first sorting result by sorting the multiple BV candidates based on a first BVD symbol prediction cost, the first BVD symbol prediction cost being obtained based on the fractional portion of each BV candidate; and obtaining a second sorting result by sorting the multiple BV candidates based on a second BVD symbol prediction cost, the second BVD symbol prediction cost being obtained based on the integer portion of each BV candidate, in response to a first combination of fractional bits of the BVD symbol prediction index already obtained based on the first sorting result. The decoder obtaining the BVD symbol prediction index corresponding to the true BVD symbol includes: obtaining a first combination of fractional bits of the BVD symbol prediction index based on the first sorting result; and obtaining a second combination of integer bits of the BVD symbol prediction index based on the second sorting result.
[0532] In some embodiments of this disclosure, the fractional BVD component portion is obtained by the decoder receiving fractional binary bits representing the fractional BVD component portion from the encoder, and the decoder obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost including: performing an interpolation filtering process for each BV candidate, wherein the interpolation filter associated with the interpolation filtering process is predefined or switched at different levels; and obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost calculated by taking into account both the integer and fractional portions of each BV candidate.
[0533] In some embodiments of this disclosure, the fractional BVD component portion is obtained by the decoder receiving fractional binary bits representing the fractional BVD component portion from the encoder, and the decoder obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost including: skipping the interpolation filtering process for each BV candidate; and obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost calculated only considering the integer portion of each BV candidate.
[0534] Figure 30 This is a flowchart illustrating a method for video decoding according to an example of this disclosure. In step 3001, the method includes: obtaining pairwise intra-block copy (IBC) candidates by a decoder based on at least two valuable IBC candidates from a candidate list of IBC merging mode or IBC advanced motion vector prediction (AMVP) mode. In step 3002, the method includes: determining corresponding attributes of the pairwise IBC candidates associated with attributes of the two valuable IBC candidates by the decoder.
[0535] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates associated with the attributes of two precious IBC candidates includes: determining the corresponding attribute of the pair of IBC candidates based on the different attributes of the two precious IBC candidates in response to the two precious IBC candidates having different attributes; determining the corresponding attribute of the pair of IBC candidates to be the same attribute in response to the two precious IBC candidates having the same attribute; or determining the corresponding attribute of the pair of IBC candidates to a predefined value.
[0536] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to two precious IBC candidates having different motion precisions, rounding or truncating the pair of IBC candidates to a predefined motion precision, or determining the motion precision of the pair of IBC candidates to be the higher of the different motion precisions.
[0537] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to two precious IBC candidates having different Reconstruction Reordered IBC (RRIBC) flipping modes, setting the pair of IBC candidates to have a non-flipping mode, wherein the different RRIBC flipping modes include any combination of a non-flipping mode, a horizontal flipping mode, and a vertical flipping mode.
[0538] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates to be the same in response to two precious IBC candidates having the same attribute includes: setting the pair of IBC candidates to have the same RRIBC flip mode in response to two precious IBC candidates having the same RRIBC flip mode, wherein the same RRIBC flip mode includes any one of a non-flip mode, a horizontal flip mode, and a vertical flip mode.
[0539] In some embodiments of this disclosure, determining the corresponding attributes of paired IBC candidates as predefined values includes setting the paired IBC candidates to have an RRIBC non-flipped mode.
[0540] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to the two precious IBC candidates having different IBC (IBC-LIC) flags with local illumination compensation (these flags include an IBC-LIC on flag and an IBC-LIC off flag), setting the pair of IBC candidates to have an IBC-LIC on flag or an IBC-LIC off flag.
[0541] In some embodiments of this disclosure, determining the corresponding attributes of a pair of IBC candidates to be the same in response to two valuable IBC candidates having the same attribute includes: setting the pair of IBC candidates to have the same IBC-LIC flag in response to two valuable IBC candidates having the same IBC-LIC flag, wherein the same IBC-LIC flag includes either an IBC-LIC on flag or an IBC-LIC off flag.
[0542] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates as a predefined value includes setting the pair of IBC candidates to have an IBC-LIC off flag.
[0543] Figure 31 This is a flowchart illustrating a method for video encoding according to an example of this disclosure. In step 3101, the method includes: obtaining syntax elements associated with an intra-block copy (IBC) mode by an encoder. In step 3102, the method includes: signaling syntax elements jointly or separately by the encoder using context-adaptive binary arithmetic coding (CABAC), which uses different context bits for different stripe types, different frame types, different video components, or different video resolutions.
[0544] In some embodiments of this disclosure, multiple CABAC context windows are predefined for different stripe types or different frame types.
[0545] In some embodiments of this disclosure, syntax elements are used to indicate motion information accuracy, interpolation filter selection for motion compensation block prediction, or interpolation filter selection for template prediction.
[0546] In some embodiments of this disclosure, different video components include luminance components and chrominance components.
[0547] Figure 32 This is a flowchart illustrating a method for video coding according to an example of this disclosure. In step 3201, the method includes: obtaining fractional motion information of the current block by an encoder in intra-block copy (IBC) mode. In step 3202, the method includes: obtaining a block vector (BV) of the current block by the encoder based on the fractional motion information. In step 3203, the method includes: obtaining a predicted block of the current block by the encoder based on the BV. In step 3204, the method includes: in response to a satisfied condition, performing a padding process on one or more samples associated with an interpolation filter by the encoder, or partially or completely skipping the padding process.
[0548] In some embodiments of this disclosure, the condition includes: determining that one or more samples associated with the interpolation filter are unavailable and that the fractional portion of the BV has a value of zero, the fractional portion including a horizontal portion or a vertical portion; and that the encoder partially or completely skips the padding process for one or more samples, including: skipping the padding process associated with the fractional portion of the BV, the fractional portion having a value of zero.
[0549] In some embodiments of this disclosure, the condition includes: all samples associated with the interpolation filter are located within the valid reference region, and the encoder partially or completely skips the filling process for one or more samples, including: completely skipping the filling process.
[0550] In some embodiments of this disclosure, the process of filling one or more samples by the encoder includes determining the fill size of a left or top sample located outside a reference block pointed to by an integer BV or an integer portion of BV to be one sample size smaller than the fill size of a right or bottom sample located outside the reference block.
[0551] In some embodiments of this disclosure, the method further includes: the encoder performing a filling process for any one or any combination of the following processes: motion search, motion refinement, or parameter generation for prediction refinement; or skipping the filling process when an interpolation filter is applied to the prediction block to obtain a filtered prediction for the current block.
[0552] Figure 33This is a flowchart illustrating a method for video coding according to an example of this disclosure. In step 3301, the method includes: obtaining fractional motion information of the current block by an encoder in intra-block copy (IBC) mode. In step 3302, obtaining a block vector (BV) of the current block by the encoder based on the fractional motion information. In step 3303, the method includes: obtaining a prediction block of the current block by the encoder based on the BV. In step 3304, the method includes: determining an interpolation filter or skipping the interpolation process by the encoder in response to a satisfied condition, wherein the condition is configured to indicate whether to switch the interpolation filter or skip the interpolation process, and whether to apply the interpolation filter to the prediction block to obtain a filtered prediction for the current block.
[0553] In some embodiments of this disclosure, the condition includes: receiving a syntax element transmitted by the encoder from the encoder, the syntax element being used to indicate switching the interpolation filter for any one or any combination of a particular stripe type, a particular frame type, a particular video component, a particular IBC mode, or a particular interpolation scenario.
[0554] In some embodiments of this disclosure, the condition includes any one or any combination of the following: processing a specific strip type, processing a specific frame type, processing a specific video component, applying a specific motion precision, applying a specific IBC mode, or applying a specific interpolation scenario.
[0555] In some embodiments of this disclosure, the specific video component includes IBC luma samples or IBC chroma samples; or, the specific IBC mode includes: adaptive reordering with merged candidates (ARMC) mode, template matching mode, block vector difference prediction for IBC blocks (IBC-BVDP) mode, block vector prediction candidate clustering and block vector difference sign derivation for reconstructing reordered IBC (IBC-BVPC) mode, or IBC with local illumination compensation (IBC-LIC) mode.
[0556] In some embodiments of this disclosure, the condition includes applying ARMC mode, template matching mode, IBC-BVDP mode, or IBC-BVPC mode, and determining the interpolation filter by the encoder or skipping the interpolation process by the encoder includes: determining the interpolation filter as an interpolation filter having a smaller length compared to a predefined length; or skipping the interpolation process.
[0557] In some embodiments of this disclosure, the condition includes applying the IBC-LIC mode or processing IBC chroma samples, and the encoder determining the interpolation filter includes: determining the interpolation filter as a preset luminance interpolation filter or a preset chroma interpolation filter.
[0558] Figure 34This is a flowchart illustrating a method for video encoding according to an example of the present disclosure. In step 3401, the method includes: obtaining a plurality of basic candidates from an intra-block copy (IBC) merge list by an encoder. In step 3402, obtaining a plurality of merge mode (MBVD) refinement positions with block vector differences for each basic candidate by the encoder based on at least one candidate distance set. In step 3403, the method includes: reordering the plurality of MBVD refinement positions by the encoder. In step 3404, the method includes: selecting a first number of MBVD refinement positions with the lowest template SAD cost by the encoder. In step 3405, the method includes: obtaining block vector differences by a decoder based on a first number of MBVD refinement locations, wherein the at least one candidate distance set includes any one or any combination of the following sets: {1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 14 pixel, 16 pixel, 18 pixel, 20 pixel, 22 pixel, 24 pixel, 26 pixel, 28 pixel, 30 pixel, 32 pixel}; {1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 10 pixel, 12 pixel, 16 pixel, 20 pixel, 24 pixel, 28 pixel, 32 pixel}. {1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 40 pixels, 44 pixels, 48 pixels, 52 pixels, 56 pixels}; {1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 10 pixels, 12 pixels, 14 pixels, 16 pixels, 18 pixels, 20 pixels, 22 pixels, 24 pixels, 26 pixels, 28 pixels, 30 pixels, 32 pixels, 34 pixels}; or, {1 pixel, 2 pixels, 4 pixels, 6 pixels, 8 pixels, 12 pixels, 16 pixels, 20 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels}.
[0559] In some embodiments of this disclosure, obtaining multiple MBVD refinement positions for each basic candidate by the encoder based on at least one candidate distance set includes: signaling syntax elements to indicate at least one distance set from at least one candidate distance set; or selecting at least one distance set from at least one candidate distance set in response to satisfying a condition associated with a particular coding pattern.
[0560] In some embodiments of this disclosure, the condition includes applying a specific encoding pattern at a specific level.
[0561] Figure 35This is a flowchart illustrating a method for video coding according to an example of the present disclosure. In step 3501, the method includes: obtaining fractional motion information of the current block by an encoder in intra-block copy (IBC) mode. In step 3502, the method includes: obtaining a block vector (BV) of the current block by the encoder based on the fractional motion information. In step 3503, the method includes: obtaining a prediction block of the current block by the encoder based on the BV. In step 3504, the method includes: obtaining a reference template by the encoder based on the BV and a template of the current block. In step 3505, the method includes: obtaining parameters of a linear filter by the encoder based on the reference template. In step 3506, the method includes: applying a linear filter to the prediction block by the encoder to obtain a filtered prediction for the current block.
[0562] In some embodiments of this disclosure, obtaining a reference template by the encoder based on BV and the template of the current block includes: obtaining an integer motion vector by modifying BV; and obtaining a reference template based on the integer motion vector and the template of the current block.
[0563] In some embodiments of this disclosure, obtaining an integer motion vector by modifying BV includes: discarding the fractional portion of BV; or, rounding the value of BV to the nearest integer value.
[0564] In some embodiments of this disclosure, obtaining a reference template by the encoder based on the BV and the template of the current block includes: in response to determining that one or more samples associated with a first interpolation filter are unavailable, performing a filling process on one or more samples to obtain a reference template, wherein the first interpolation filter is the same as or different from a second interpolation filter to be applied to the prediction block.
[0565] In some embodiments of this disclosure, the first interpolation filter is different from the second interpolation filter to be applied to the prediction block, and the first interpolation filter is predefined or transmitted by signaling at a specific level.
[0566] Figure 36 This is a flowchart illustrating a method for video encoding according to an example of the present disclosure. In step 3601, the method includes: obtaining fractional motion information of multiple luma blocks located at multiple predefined positions relative to the position of the current chroma block by an encoder in intra-block copy (IBC) mode. In step 3602, the method includes: obtaining multiple luma block vectors (BVs) of the multiple luma blocks by the encoder based on the fractional motion information. In step 3603, the method includes: determining a chroma BV by the encoder based on the multiple luma BVs of the multiple luma blocks. In step 3604, the method includes: obtaining a predicted chroma block by the encoder based on the chroma BVs.
[0567] In some embodiments of this disclosure, determining the chromaticity BV by the encoder based on multiple luminance BVs of multiple luminance blocks includes: calculating the chromaticity template matching cost by the encoder based on the luminance BV of each of the multiple luminance blocks and the current chromaticity block; and determining the chromaticity BV by the encoder as the luminance BV with the lowest chromaticity template matching cost.
[0568] In some embodiments of this disclosure, determining the chromaticity BV by the encoder based on multiple luminance BVs of multiple luminance blocks includes: performing gradient-based edge detection on the multiple luminance blocks; and determining the chromaticity BV by the encoder as the luminance BV of a luminance block that has the same edge as the chromaticity block.
[0569] In some embodiments of this disclosure, the method further includes: obtaining template-based BV thinning by an encoder based on intra-frame template matching (ITM); and obtaining a thinned chroma BV by an encoder based on the template-based BV thinning and the chroma BV.
[0570] In some embodiments of this disclosure, prediction refinement includes positive refinement or negative refinement.
[0571] In some embodiments of this disclosure, obtaining a refined chroma BV by an encoder based on the template-based BV refinement and the chroma BV includes: applying the template-based BV refinement to a first color component of the chroma BV; and reusing the predicted refinement to a second color component of the chroma BV.
[0572] In some embodiments of this disclosure, obtaining template-based BV refinement by the encoder includes: obtaining template-based BV refinement based on a reference template cost calculated by combining the matching error of the first color template component and the second color template component; and obtaining refined chroma BV by the encoder based on the template-based BV refinement and the chroma BV includes: applying the template-based BV to the first color component and the second color component of the chroma BV.
[0573] In some embodiments of this disclosure, the method further includes: truncating or rounding the value of chroma BV to an integer value.
[0574] In some embodiments of this disclosure, the encoder determines the chromaticity block vector (BV) based on multiple luminance BVs in multiple luminance blocks, including:
[0575] Multiple values of luminance BV are truncated or rounded into multiple integer values; and the encoder determines the chrominance BV based on the multiple integer values.
[0576] In some embodiments of this disclosure, the method further includes: applying a first interpolation filter to the predicted chroma block to obtain a filtered prediction for the current chroma block, wherein the first interpolation filter is the same as or different from a second interpolation filter used to obtain template-based BV refinement, and the first interpolation filter or the second interpolation filter is predefined.
[0577] In some embodiments of this disclosure, the method further includes: in response to determining that the size of the current chroma block is smaller than a predefined size, truncating or rounding the value of the chroma BV to an integer value, or using a filter with a smaller size compared to the redefined size when performing a first interpolation process on the predicted chroma block or a second interpolation process on the chroma BV.
[0578] In some embodiments of this disclosure, filters with smaller dimensions are predefined or used for signal transmission at a specific level.
[0579] Figure 37 This is a flowchart illustrating a method for video encoding according to an example of this disclosure. In step 3701, the method includes: obtaining fractional motion information of the current block by an encoder in Intra-Block Copy (IBC) mode. In step 3702, the method includes: obtaining a block vector (BV) of the current block by the encoder based on the fractional motion information. In step 3703, the method includes: setting the horizontal or vertical component of the BV to zero by the encoder according to the flip type of the RRIBC mode in response to applying a Reconstruction Reordering IBC (RRIBC) mode. In step 3704, the method includes: obtaining a predicted block of the current block by the encoder based on the BV.
[0580] In some embodiments of this disclosure, the method further includes: limiting the motion precision to integer pixels in response to the application of an RRIBC mode that disallows fractional motion vectors, or receiving a motion vector difference (MVD) equal to zero.
[0581] Figure 38This is a flowchart illustrating a method for video coding according to an example of the present disclosure. In step 3801, the method includes: obtaining fractional motion information of the current block by an encoder in Intra-Block Copy (IBC) mode. In step 3802, the method includes: obtaining a block vector (BV) of the current block by the encoder based on the fractional motion information. In step 3803, the method includes: skipping the fractional motion estimation process for the BV, or skipping the fractional motion estimation process for the empty components of the BV, which include horizontal or vertical components, in response to applying the Block Vector Prediction Candidate Clustering and Block Vector Difference Sign Derivation (IBC-BVPC) mode for reconstructing reordered IBC and the IBC Advanced Motion Vector Prediction (AMVP) mode. In step 3804, the method includes: obtaining a predicted block of the current block by the encoder based on the BV.
[0582] In some embodiments of this disclosure, the method further includes: limiting the motion precision to integer pixels in response to the application of the IBC-BVPC mode disallowing fractional motion vectors, or receiving a motion vector difference (MVD) equal to zero.
[0583] Figure 39 This is a flowchart illustrating a method for video coding according to an example of this disclosure. In step 3901, the method includes: obtaining amplitudes of block vector difference (BVD) components by an encoder, each of these BVD components including an integer BVD component portion and a fractional BVD component portion. In step 3902, the method includes: obtaining a context-coded BVD symbol prediction index by an encoder. In step 3903, the method includes: obtaining a plurality of BV candidates by an encoder by creating a combination between BVD symbols and absolute BVD values and adding that combination to a BV prediction value, each BV candidate including an integer portion and a fractional portion. In step 3904, the method includes: obtaining a BVD symbol prediction cost for each BV candidate by an encoder based on a template matching cost. In step 3905, the method includes: sorting the plurality of BV candidates by an encoder based on the BVD symbol prediction costs. In step 3906, the method includes: obtaining a BVD symbol prediction index corresponding to a true BVD symbol by an encoder.
[0584] In some embodiments of this disclosure, the BVD symbol prediction index includes integer bits and fractional bits.
[0585] In some embodiments of this disclosure, the encoder sorting of multiple BV candidates based on BVD symbol prediction costs includes: obtaining a first sorting result by sorting the multiple BV candidates based on a first BVD symbol prediction cost, which is obtained based on the integer portion of each BV candidate; and obtaining a second sorting result by sorting the multiple BV candidates based on a second BVD symbol prediction cost, which is obtained based on the fractional portion of each BV candidate, in response to a first combination of integer bits of the BVD symbol prediction index already obtained based on the first sorting result. The encoder obtaining the BVD symbol prediction index corresponding to the true BVD symbol includes: obtaining a first combination of integer bits of the BVD symbol prediction index based on the first sorting result; and obtaining a second combination of fractional bits of the BVD symbol prediction index based on the second sorting result.
[0586] In some embodiments of this disclosure, the encoder sorting of multiple BV candidates based on BVD symbol prediction costs includes: obtaining a first sorting result by sorting the multiple BV candidates based on a first BVD symbol prediction cost, which is obtained based on the fractional portion of each BV candidate; and obtaining a second sorting result by sorting the multiple BV candidates based on a second BVD symbol prediction cost, which is obtained based on the integer portion of each BV candidate, in response to a first combination of fractional bits of the BVD symbol prediction index already obtained based on the first sorting result. The encoder obtaining the BVD symbol prediction index corresponding to the true BVD symbol includes: obtaining a first combination of fractional bits of the BVD symbol prediction index based on the first sorting result; and obtaining a second combination of integer bits of the BVD symbol prediction index based on the second sorting result.
[0587] In some embodiments of this disclosure, the fractional BVD component portion is obtained by the encoder receiving fractional binary bits representing the fractional BVD component portion from the encoder, and the encoder obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost including: performing an interpolation filtering process for each BV candidate, wherein the interpolation filter associated with the interpolation filtering process is predefined or switched at different levels; and obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost calculated by taking into account both the integer and fractional portions of each BV candidate.
[0588] In some embodiments of this disclosure, the fractional BVD component portion is obtained by the encoder receiving fractional binary bits representing the fractional BVD component portion from the encoder, and the encoder obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost includes: skipping the interpolation filtering process for each BV candidate; and obtaining the BVD symbol prediction cost for each BV candidate based on the template matching cost calculated only considering the integer portion of each BV candidate.
[0589] Figure 40 This is a flowchart illustrating a method for video coding according to an example of this disclosure. In step 4001, the method includes: obtaining pairwise intra-block copy (IBC) candidates by an encoder based on at least two precious IBC candidates from a candidate list of IBC merging mode or IBC advanced motion vector prediction (AMVP) mode. In step 4002, the method includes: determining corresponding attributes of the pairwise IBC candidates associated with attributes of the two precious IBC candidates by the encoder.
[0590] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates associated with the attributes of two precious IBC candidates includes: determining the corresponding attribute of the pair of IBC candidates based on the different attributes of the two precious IBC candidates in response to the two precious IBC candidates having different attributes; determining the corresponding attribute of the pair of IBC candidates to be the same attribute in response to the two precious IBC candidates having the same attribute; or determining the corresponding attribute of the pair of IBC candidates to a predefined value.
[0591] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to two precious IBC candidates having different motion precisions, rounding or truncating the pair of IBC candidates to a predefined motion precision, or determining the motion precision of the pair of IBC candidates to be the higher of the different motion precisions.
[0592] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to two precious IBC candidates having different Reconstruction Reordered IBC (RRIBC) flipping modes, setting the pair of IBC candidates to have a non-flipping mode, wherein the different RRIBC flipping modes include any combination of a non-flipping mode, a horizontal flipping mode, and a vertical flipping mode.
[0593] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates to be the same in response to two precious IBC candidates having the same attribute includes: setting the pair of IBC candidates to have the same RRIBC flip mode in response to two precious IBC candidates having the same RRIBC flip mode, wherein the same RRIBC flip mode includes any one of a non-flip mode, a horizontal flip mode, and a vertical flip mode.
[0594] In some embodiments of this disclosure, determining the corresponding attributes of paired IBC candidates as predefined values includes setting the paired IBC candidates to have an RRIBC non-flipped mode.
[0595] In some embodiments of this disclosure, in response to two precious IBC candidates having different attributes, determining the corresponding attributes of a pair of IBC candidates based on the different attributes of the two precious IBC candidates includes: in response to the two precious IBC candidates having different IBC (IBC-LIC) flags with local illumination compensation (these flags include an IBC-LIC on flag and an IBC-LIC off flag), setting the pair of IBC candidates to have an IBC-LIC on flag or an IBC-LIC off flag.
[0596] In some embodiments of this disclosure, determining the corresponding attributes of a pair of IBC candidates to be the same in response to two valuable IBC candidates having the same attribute includes: setting the pair of IBC candidates to have the same IBC-LIC flag in response to two valuable IBC candidates having the same IBC-LIC flag, wherein the same IBC-LIC flag includes either an IBC-LIC on flag or an IBC-LIC off flag.
[0597] In some embodiments of this disclosure, determining the corresponding attribute of a pair of IBC candidates as a predefined value includes setting the pair of IBC candidates to have an IBC-LIC off flag.
[0598] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.
[0599] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.
[0600] The description in this disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.
[0601] Unless otherwise specified, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0602] The examples were chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.
Claims
1. A method for video decoding, the method comprising: The decoder obtains the fractional motion information of the current block in Intra-Block Copy (IBC) mode; The decoder obtains the block vector (BV) of the current block based on the fractional motion information; The decoder obtains the predicted block of the current block based on the BV; as well as In response to the condition being met, the decoder performs a filling process on one or more samples associated with the interpolation filter, or partially or completely skips the filling process.
2. The method as described in claim 1, wherein, The conditions include: determining that one or more samples associated with the interpolation filter are unavailable and that the fractional portion of the BV is zero, the fractional portion comprising a horizontal or vertical portion, and The decoder partially or completely skips the filling process for the one or more samples, including: Skip the filling process associated with the fractional portion of the BV, where the value of the fractional portion is zero.
3. The method as described in claim 1, wherein, The conditions include: all samples associated with the interpolation filter are located within the valid reference region, and The decoder partially or completely skips the filling process for the one or more samples, including: The filling process is completely skipped.
4. The method of claim 1, wherein, The process of filling the one or more sample points by the decoder includes: The fill size of the left or top sample located outside the reference block pointed to by the integer BV or the integer part of BV is determined to be 1 sample size smaller than the fill size of the right or bottom sample located outside the reference block.
5. The method of claim 1, further comprising: The decoder performs the filling process for any one or any combination of the following processes: motion search, motion refinement, or parameter generation for predictive refinement.
6. The method of claim 1, further comprising: The filling process is skipped when the interpolation filter is applied to the prediction block to obtain a filtered prediction for the current block.
7. A method for video decoding, the method comprising: The decoder obtains fractional motion information of multiple luma blocks located at multiple predefined positions relative to the current chroma block in intra-block copy (IBC) mode; The decoder obtains multiple luminance block vectors (BVs) of the multiple luminance blocks based on the fractional motion information; The decoder determines the chromaticity BV based on the multiple luminance BVs of the multiple luminance blocks; as well as The decoder obtains the predicted chroma block based on the chroma BV.
8. The method of claim 7, wherein, The decoder determines the chromaticity BV based on the plurality of luminance BVs of the plurality of luminance blocks, including: The decoder calculates the chroma template matching cost based on the luminance BV of each of the plurality of luminance blocks and the current chroma block; and The decoder determines the chroma BV as the luminance BV with the lowest chroma template matching cost.
9. The method of claim 7, wherein, The decoder determines the chromaticity BV based on the plurality of luminance BVs of the plurality of luminance blocks, including: Gradient-based edge detection is performed on the plurality of brightness blocks; and The decoder determines the chroma BV as the luminance BV of the luminance block that has the same edge as the chroma block.
10. The method of claim 7, further comprising: The decoder obtains template-based BV refinement based on intra-frame template matching (ITM); as well as The decoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV.
11. The method of claim 10, wherein, The prediction refinement includes positive refinement or negative refinement.
12. The method of claim 10, wherein, The decoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV, including: The template-based BV refinement is applied to the first color component of the chromaticity BV; and The refined prediction is then reused for the second color component of the chromaticity BV.
13. The method of claim 10, wherein, The template-based BV refinement obtained by the decoder includes: The template-based BV refinement is obtained based on a reference template cost calculated by combining the matching error between the first and second color template components. The decoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV, including: The template-based BV is applied to the first and second color components of the chromaticity BV.
14. The method of claim 7, further comprising: The value of the chromaticity BV is truncated or rounded to an integer value.
15. The method of claim 7, wherein, The decoder determines the chroma block vector (BV) based on multiple luminance BVs of the plurality of luminance blocks, including: The multiple values of the multiple luminance BV are truncated or rounded to multiple integer values; and The decoder determines the chroma BV based on the plurality of integer values.
16. The method of claim 10, further comprising: A first interpolation filter is applied to the predicted chroma block to obtain a filtered prediction for the current chroma block. in, The first interpolation filter may be the same as or different from the second interpolation filter used to obtain the template-based BV refinement, and either the first interpolation filter or the second interpolation filter is predefined.
17. The method of claim 10, further comprising: In response to determining that the size of the current chroma block is smaller than a predefined size, the value of the chroma BV is truncated or rounded to an integer value, or when performing a first interpolation process on the predicted chroma block or a second interpolation process on the chroma BV, a filter with a smaller size compared to the redefined size is used.
18. The method of claim 10, wherein, The filter with the smaller size is predefined or used for signal transmission at a specific level.
19. A method for video decoding, the method comprising: The block vector difference (BVD) component amplitude is obtained by the decoder, each of the BVD components including an integer BVD component portion and a fractional BVD component portion; The decoder obtains the context-encoded BVD symbol prediction index; The decoder obtains multiple BV candidates by creating a combination between the BVD symbol and the absolute BVD value and adding the combination to the BV prediction value. Each BV candidate includes an integer part and a fractional part. The decoder obtains the BVD symbol prediction cost for each BV candidate based on the template matching cost; The decoder sorts the plurality of BV candidates based on the BVD symbol prediction cost; as well as The decoder obtains the BVD symbol prediction index corresponding to the real BVD symbol.
20. The method of claim 19, wherein, The BVD symbol prediction index includes integer binary bits and fractional binary bits.
21. The method of claim 20, wherein, The decoder sorts the plurality of BV candidates based on the BVD symbol prediction cost, including: A first ranking result is obtained by ranking the plurality of BV candidates based on a first BVD symbol prediction cost, wherein the first BVD symbol prediction cost is obtained based on the integer portion of each BV candidate; and In response to the first combination of integer bits of the BVD symbol prediction index already obtained based on the first ranking result, a second ranking result is obtained by ranking the plurality of BV candidates based on a second BVD symbol prediction cost, which is obtained based on the score portion of each BV candidate. The BVD symbol prediction index obtained by the decoder corresponding to the real BVD symbol includes: The first combination of the integer binary bits of the BVD symbol prediction index is obtained based on the first sorting result; and Based on the second sorting result, a second combination of the fractional binary bits of the BVD symbol prediction index is obtained.
22. The method of claim 20, wherein, The decoder sorts the plurality of BV candidates based on the BVD symbol prediction cost, including: A first ranking result is obtained by ranking the plurality of BV candidates based on a first BVD symbol prediction cost, wherein the first BVD symbol prediction cost is obtained based on the score portion of each BV candidate; and In response to the first combination of the fractional bits of the BVD symbol prediction index already obtained based on the first ranking result, a second ranking result is obtained by ranking the plurality of BV candidates based on a second BVD symbol prediction cost, which is obtained based on the integer portion of each BV candidate. The BVD symbol prediction index obtained by the decoder corresponding to the real BVD symbol includes: The first combination of the fractional binary bits of the BVD symbol prediction index is obtained based on the first sorting result; and Based on the second sorting result, a second combination of the integer binary bits of the BVD symbol prediction index is obtained.
23. The method of claim 19, wherein, The fractional BVD component is obtained by the decoder receiving fractional binary bits representing the fractional BVD component from the encoder, and The BVD symbol prediction cost obtained by the decoder based on the template matching cost includes: An interpolation filtering process is performed for each BV candidate, wherein the interpolation filter associated with the interpolation filtering process is predefined or switched at different levels; and The BVD symbol prediction cost for each BV candidate is obtained based on the template matching cost calculated by taking into account both the integer part and the fractional part of each BV candidate.
24. The method of claim 19, wherein, The fractional BVD component is obtained by the decoder receiving fractional binary bits representing the fractional BVD component from the encoder, and The BVD symbol prediction cost obtained by the decoder based on the template matching cost includes: Skip the interpolation filtering process for each BV candidate; and The BVD symbol prediction cost for each BV candidate is obtained based on the template matching cost calculated by considering only the integer part of each BV candidate.
25. A method for video encoding, the method comprising: The encoder obtains the fractional motion information of the current block in Intra-Block Copy (IBC) mode; The encoder obtains the block vector (BV) of the current block based on the fractional motion information; The encoder obtains the predicted block of the current block based on the BV; as well as In response to the condition being met, the encoder performs a filling process on one or more samples associated with the interpolation filter, or partially or completely skips the filling process.
26. The method of claim 25, wherein, The conditions include: determining that one or more samples associated with the interpolation filter are unavailable and that the fractional portion of the BV is zero, the fractional portion comprising a horizontal or vertical portion, and The encoder partially or completely skips the filling process for the one or more samples, including: Skip the filling process associated with the fractional portion of the BV, where the value of the fractional portion is zero.
27. The method of claim 25, wherein, The conditions include: all samples associated with the interpolation filter are located within the valid reference region, and The encoder partially or completely skips the filling process for the one or more samples, including: The filling process is completely skipped.
28. The method of claim 25, wherein, The process of filling the one or more sample points by the encoder includes: The fill size of the left or top sample located outside the reference block pointed to by the integer BV or the integer part of BV is determined to be 1 sample size smaller than the fill size of the right or bottom sample located outside the reference block.
29. The method of claim 25, further comprising: The encoder performs the filling process for any one or any combination of the following processes: motion search, motion refinement, or parameter generation for predictive refinement.
30. The method of claim 25, further comprising: The filling process is skipped when the interpolation filter is applied to the prediction block to obtain a filtered prediction for the current block.
31. A method for video encoding, the method comprising: The encoder obtains fractional motion information of multiple luma blocks located at multiple predefined positions relative to the current chroma block in intra-block copy (IBC) mode; The encoder obtains multiple luminance block vectors (BVs) of the multiple luminance blocks based on the fractional motion information; The encoder determines the chromaticity BV based on the plurality of luminance BVs of the plurality of luminance blocks; as well as The encoder obtains the predicted chromaticity block based on the chromaticity BV.
32. The method of claim 31, wherein, The encoder determines the chromaticity BV based on the plurality of luminance BVs of the plurality of luminance blocks, including: The encoder calculates the chroma template matching cost based on the luminance BV of each of the plurality of luminance blocks and the current chroma block; and The encoder determines the chromaticity BV as the luminance BV with the lowest chromaticity template matching cost.
33. The method of claim 31, wherein, The encoder determines the chromaticity BV based on the plurality of luminance BVs of the plurality of luminance blocks, including: Gradient-based edge detection is performed on the plurality of brightness blocks; and The encoder determines the chroma BV as the luminance BV of the luminance block that has the same edge as the chroma block.
34. The method of claim 31, further comprising: The encoder obtains template-based BV refinement based on intra-frame template matching (ITM); as well as The encoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV.
35. The method of claim 34, wherein, The prediction refinement includes positive refinement or negative refinement.
36. The method of claim 34, wherein, The encoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV, including: The template-based BV refinement is applied to the first color component of the chromaticity BV; and The refined prediction is then reused for the second color component of the chromaticity BV.
37. The method of claim 34, wherein, The template-based BV refinement obtained by the encoder includes: The template-based BV refinement is obtained based on a reference template cost calculated by combining the matching error between the first and second color template components. The encoder obtains the refined chroma BV based on the template-based BV refinement and the chroma BV, including: The template-based BV is applied to the first and second color components of the chromaticity BV.
38. The method of claim 31, further comprising: The value of the chromaticity BV is truncated or rounded to an integer value.
39. The method of claim 31, wherein, The encoder determines the chroma block vector (BV) based on multiple luminance BVs of the plurality of luminance blocks, including: The multiple values of the multiple luminance BV are truncated or rounded to multiple integer values; and The encoder determines the chromaticity BV based on the plurality of integer values.
40. The method of claim 34, further comprising: A first interpolation filter is applied to the predicted chroma block to obtain a filtered prediction for the current chroma block. in, The first interpolation filter may be the same as or different from the second interpolation filter used to obtain the template-based BV refinement, and either the first interpolation filter or the second interpolation filter is predefined.
41. The method of claim 34, further comprising: In response to determining that the size of the current chroma block is smaller than a predefined size, the value of the chroma BV is truncated or rounded to an integer value, or when performing a first interpolation process on the predicted chroma block or a second interpolation process on the chroma BV, a filter with a smaller size compared to the redefined size is used.
42. The method of claim 34, wherein, The filter with the smaller size is predefined or used for signal transmission at a specific level.
43. A method for video encoding, the method comprising: The encoder obtains the magnitude of the block vector difference (BVD) components, each of which includes an integer BVD component portion and a fractional BVD component portion; The encoder obtains the context-coded BVD symbol prediction index; The encoder obtains multiple BV candidates by creating a combination between BVD symbols and absolute BVD values and adding the combination to the BV prediction value. Each BV candidate includes an integer part and a fractional part. The encoder obtains the BVD symbol prediction cost for each BV candidate based on the template matching cost; The encoder sorts the plurality of BV candidates based on the BVD symbol prediction cost; as well as The encoder obtains the BVD symbol prediction index corresponding to the real BVD symbol.
44. The method of claim 43, wherein, The BVD symbol prediction index includes integer binary bits and fractional binary bits.
45. The method of claim 44, wherein, The encoder sorts the plurality of BV candidates based on the BVD symbol prediction cost, including: A first ranking result is obtained by ranking the plurality of BV candidates based on a first BVD symbol prediction cost, wherein the first BVD symbol prediction cost is obtained based on the integer portion of each BV candidate; and In response to the first combination of integer bits of the BVD symbol prediction index already obtained based on the first ranking result, a second ranking result is obtained by ranking the plurality of BV candidates based on a second BVD symbol prediction cost, which is obtained based on the score portion of each BV candidate. The BVD symbol prediction index obtained by the encoder corresponding to the true BVD symbol includes: The first combination of the integer binary bits of the BVD symbol prediction index is obtained based on the first sorting result; and Based on the second sorting result, a second combination of the fractional binary bits of the BVD symbol prediction index is obtained.
46. The method of claim 44, wherein, The encoder sorts the plurality of BV candidates based on the BVD symbol prediction cost, including: A first ranking result is obtained by ranking the plurality of BV candidates based on a first BVD symbol prediction cost, wherein the first BVD symbol prediction cost is obtained based on the score portion of each BV candidate; and In response to the first combination of the fractional bits of the BVD symbol prediction index already obtained based on the first ranking result, a second ranking result is obtained by ranking the plurality of BV candidates based on a second BVD symbol prediction cost, which is obtained based on the integer portion of each BV candidate. The BVD symbol prediction index obtained by the encoder corresponding to the true BVD symbol includes: The first combination of the fractional binary bits of the BVD symbol prediction index is obtained based on the first sorting result; and Based on the second sorting result, a second combination of the integer binary bits of the BVD symbol prediction index is obtained.
47. The method of claim 43, further comprising: The encoder transmits the fractional BVD component via signals, where the fractional binary bits represent the fractional BVD component. The BVD symbol prediction cost obtained by the encoder based on the template matching cost for each BV candidate includes: An interpolation filtering process is performed for each BV candidate, wherein the interpolation filter associated with the interpolation filtering process is predefined or switched at different levels; and The BVD symbol prediction cost for each BV candidate is obtained based on the template matching cost calculated by taking into account both the integer part and the fractional part of each BV candidate.
48. The method of claim 43, further comprising: The encoder transmits the fractional BVD component via signals, where the fractional binary bits represent the fractional BVD component. The BVD symbol prediction cost obtained by the encoder based on the template matching cost for each BV candidate includes: Skip the interpolation filtering process for each BV candidate; and The BVD symbol prediction cost for each BV candidate is obtained based on the template matching cost calculated by considering only the integer part of each BV candidate.
49. An apparatus for video decoding, the apparatus comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 1 to 24 when executing the instructions.
50. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 1 to 24.
51. An apparatus for video encoding, the apparatus comprising: One or more processors; as well as A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method as described in any one of claims 25 to 48 when executing the instructions.
52. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 25 to 48.
53. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method of any one of claims 1 to 24.
54. A non-transitory computer-readable storage medium for storing a bit stream generated by the method of any one of claims 25 to 48.