Method and device for video processing and medium

Through the method of time-domain block vector prediction and time-domain block vector candidates, the problem of insufficient encoding and decoding efficiency in existing video encoding and decoding technologies is solved, and more efficient video encoding and decoding is achieved.

CN120435864APending Publication Date: 2025-08-05DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089830.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-28
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in the video block vector prediction, which makes it difficult for the existing technology to effectively improve the encoding and decoding efficiency.

Method used

By using the method of time domain block vector prediction and time domain block vector candidates, the encoding and decoding efficiency is improved by determining the conversion between the current video block and the bit stream.

Benefits of technology

It improves the encoding and codec efficiency and effectiveness of video encoding and codec and improves the performance of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120435864A_ABST
    Figure CN120435864A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. In the method, for a conversion between a current video block of a video and a bitstream of the video, at least one of a time domain block vector (BV) prediction or a time domain BV candidate for the current video block is determined. A target candidate for the current video block is determined based on the base candidate. The conversion is performed based on at least one of the time-domain BV prediction or the time-domain BV candidate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to temporal block vector (BV) prediction or temporal BV candidates. Background Art

[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided. The method includes: determining, for conversion between a current video block of a video and a bitstream of the video, at least one of a temporal block vector (BV) prediction or a temporal BV candidate for the current video block; and performing the conversion based on the temporal BV prediction or the temporal BV candidate. The method according to the first aspect of the present disclosure utilizes the temporal BV prediction or the temporal BV candidate. This improves the efficiency of the BV prediction, thereby improving both codec effectiveness and codec efficiency.

[0005] In a second aspect, another method for video processing is provided. The method includes: determining a block vector prediction (BVP) for a sub-block of a current video block for conversion between a current video block and a bitstream of the video, where the current video block is encoded or decoded using a sub-block-based temporal motion vector prediction (SbTMVP) mode; and performing conversion based on the BVP. The method according to the second aspect of the present disclosure determines the BVP for the sub-block of the current video block encoded or decoded using SbTMVP. This improves codec effectiveness and efficiency.

[0006] In a third aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. These instructions, when executed by the processor, cause the processor to perform the method according to the first aspect or the second aspect of the present disclosure.

[0007] In a fourth aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect or the second aspect of the present disclosure.

[0008] In a fifth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining at least one of a temporal block vector (BV) prediction or a temporal BV candidate for a current video block of the video; and generating a bitstream based on at least one of the temporal BV prediction or the temporal BV candidate.

[0009] In a sixth aspect, a method for storing a bitstream of a video is provided. The method includes: determining at least one of a temporal block vector (BV) prediction or a temporal BV candidate for a current video block of the video; generating a bitstream based on the at least one of the temporal BV prediction or the temporal BV candidate; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] In a seventh aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a block vector prediction (BVP) of a sub-block of a current video block of the video, the current video block being encoded or decoded using a sub-block-based temporal motion vector prediction (SbTMVP) mode; and generating a bitstream based on the BVP.

[0011] In an eighth aspect, a method for storing a bitstream of a video is provided. The method includes determining a block vector prediction (BVP) of a sub-block of a current video block of the video, the current video block being encoded or decoded using a sub-block-based temporal motion vector prediction (SbTMVP) mode; generating a bitstream based on the BVP; and storing the bitstream in a non-transitory computer-readable recording medium.

[0012] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0014] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0015] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;

[0016] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0017] Figure 4 shows the spatial neighborhood used in IBC vector prediction;

[0018] Figure 5 The current CTU processing order and its available reference samples in the current and left CTUs are shown;

[0019] Figure 6 Shows the spatial proximity used in IBC Merge / AMVP list construction;

[0020] Figure 7 shows the fill candidates for replacing the zero vector in the IBC list;

[0021] Figure 8 shows the IBC reference area depending on the current CU position;

[0022] Figure 9 The reference area for IBC when CTU(m,n) is encoded and decoded is shown. The blue block represents the current CTU; the green block represents the reference area; and the white block represents the invalid reference area;

[0023] Figure 10A A diagram showing BV adjustment for horizontal flipping is shown;

[0024] Figure 10B A diagram showing BV adjustment for vertical flipping is shown;

[0025] Figure 11 The intra-frame template matching search area used is shown;

[0026] Figure 12 shows the use of IntraTMP block vectors for IBC blocks;

[0027] Figure 13A An example of an IBC block vector candidate list in which only IBC block vectors are present is shown;

[0028] Figure 13B An example of an IBC block vector candidate list in which an IBC block vector and an IntraTMP block vector exist is shown;

[0029] Figure 14 The template and the reference points of the template in the reference picture are shown;

[0030] Figure 15 A template of a block with sub-block motion using motion information of a sub-block of a current block and a reference sample of the template are shown;

[0031] Figure 16 Shows the location of the spatial merge candidate;

[0032] Figure 17 shows candidate pairs considered for redundancy check of spatial merge candidates;

[0033] Figure 18 A diagram showing motion vector scaling for temporal Merge candidates is shown;

[0034] Figure 19 The candidate positions of the time domain Merge candidates C0 and C1 are shown;

[0035] Figure 20 The spatial neighboring blocks used to derive spatial Merge candidates are shown;

[0036] Figure 21A shows the spatial neighborhood blocks used by ATVMP;

[0037] Figure 21B Shows the derivation of sub-CU motion fields by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs;

[0038] Figure 22A Candidate positions for spatial candidates are shown;

[0039] Figure 22B Candidate positions for time domain candidates are shown;

[0040] Figure 23 The candidate positions for the temporal BV candidate are shown, and the spatial domain can be left, above, upper right, lower left, or upper left;

[0041] Figure 24 Candidate positions for the time domain BV candidate are shown;

[0042] Figure 25 A first pattern of candidate positions for a time-domain BV candidate is shown;

[0043] Figure 26 A second pattern of candidate positions for time-domain BV candidates is shown;

[0044] Figure 27 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;

[0045] Figure 28 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and

[0046] Figure 29A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0047] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION

[0048] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0049] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0050] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.

[0051] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0052] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment

[0053] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0054] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0055] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0056] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.

[0057] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0058] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.

[0059] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0060] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0061] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0062] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.

[0063] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0064] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0065] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0066] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0067] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0068] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0069] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0070] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0071] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0072] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0073] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0074] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0075] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0076] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0077] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0078] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0079] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0080] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0081] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.

[0082] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0083] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0084] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0085] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0086] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.

[0087] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.

[0088] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0089] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0090] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to image / video coding and decoding, and more particularly to time-domain block vector prediction. It can be applied to existing video coding standards such as HEVC or standard VVC (Versatile Video Codec). It can also be adapted to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 Visual and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. The H.262 standard is based on a hybrid video codec architecture that uses temporal prediction and transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly, and at the April 2018 JVET meeting, the new video codec standard was officially named the Versatile Video Codec (VVC). The first version of the VVC Test Model (VTM) was also released. The VVC working draft and the VTM test model were subsequently updated after each meeting. VVC achieved technical completion (FDIS) at the July 2020 meeting. In January 2021, JVET established Exploratory Experiments (EEs) to leverage novel traditional algorithms to achieve enhanced compression efficiency beyond VVC capabilities. Soon after, ECM was established as a common software foundation for long-term exploratory work toward the next generation of video codec standards. 2.1. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is in integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra prediction mode or the inter prediction mode. The IBC mode is suitable for CUs with a width and height that are less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In the hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4×4 sub-block. For larger current block sizes, a hash key is determined to match the hash key of the reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matching reference is calculated, and the block vector cost with the minimum cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows: – IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vectors from a list of neighboring candidate IBC coded blocks are used to predict the current block. The Merge list includes spatial, HMVP, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same manner as motion vector differences. The block vector prediction method uses two candidates as predictors: one from the left neighbor and one from the top neighbor (if IBC is encoded). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.1.1. Simplification of IBC vector prediction The BV prediction values of the Merge mode and AMVP mode in IBC will share a common prediction value list including the following elements: ● Two adjacent airspace locations (such as Figure 4 A1, B1 in FIG, which show the spatial neighborhood used in IBC vector prediction), 5 HMVP entries, ●The default is zero vector. For Merge mode, at most the first 6 entries of the list will be used; for AMVP mode, the first 2 entries of the list will be used. The list also meets the shared Merge list region requirement (sharing the same list within SMR). 2.1.2.IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined area including the area of the current CTU and some areas of the left CTU. Figure 5 The reference area of the IBC mode is shown, where each block represents a 64x64 luma sample unit. Figure 5 The current CTU processing order and its available reference samples in the current CTU and the left CTU are shown. Depending on the location of the current codec CU position within the current CTU, the following applies: – If the current block falls within the upper left 64x64 block of the current CTU, in addition to the reconstructed samples in the current CTU, it can also refer to the reference samples in the lower right 64x64 block of the left CTU, using CPR mode. The current block can also refer to the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU, using CPR mode. – If the current block falls into the upper right 64x64 block of the current CTU, in addition to the samples that have been reconstructed in the current CTU, if the luma position (0,64) relative to the current CTU has not been reconstructed, the current block can also refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU, using the CPR mode; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls into the lower left 64x64 block of the current CTU, in addition to the samples that have been reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU, using CPR mode. Otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU, using CPR mode. – If the current block falls into the lower right 64x64 block of the current CTU, it can only refer to the reconstructed samples in the current CTU, using CPR mode. This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementation. 2.1.3. IBC Interaction with Other Encoding Tools The interaction between IBC mode and other inter-frame codec tools in VVC (such as paired merge candidates, history-based motion vector predictor (HMVP), intra-frame inter-frame joint prediction mode (CIIP), merge mode with motion vector difference (MMVD) and geometric partition mode (GPM)) is as follows: – IBC can be used with pairwise merge candidates and HMVP. A new pairwise IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, the IBC motion is inserted into the history cache for future reference. – IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM. – When using DUAL_TREE partitioning, IBC is not allowed for chroma codec blocks. Unlike the HEVC screen content codec extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects apply: – IBC shares the same process as in regular MV Merge, including the use of paired merge candidates and history-based motion prediction values, but TMVP and zero vectors are not allowed because they are invalid for IBC mode. – Separate HMVP caches (5 candidates each) for regular MV and IBC. – Block vector constraints are implemented as bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream and that if the merge candidate is invalid (out of range or 0), then the merge is not used. This bitstream consistency constraint is expressed in terms of a virtual buffer as described below. – For deblocking, IBC is handled as inter mode. – If the current block is encoded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to only indicate whether the MV is inter-pixel or 4-integer pixels. - The number of IBC Merge candidates may be signaled in the slice header separately from the number of normal, sub-block and geometric Merge candidates. The concept of a virtual buffer is used to describe the allowable reference area for IBC prediction modes and valid block vectors. Denoting the CTU size as ctbSize, the virtual buffer, ibcBuf, has a width of wIbcBuf = 128x128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows. – At the beginning of decoding each CTU line, the entire ibcBuf is flushed with an invalid value of -1. – When decoding VPDU (xVPDU, yVPDU) starts relative to the upper left corner of the picture, set ibcBuf[x][y]=0, where x=xVPDU%wIbcBuf,...,xVPDU%wIbcBuf+W v -1;y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1. – After decoding, the CU contains the (x, y) relative to the top left corner of the picture, set ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]. For a block covering coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1. 2.1.4.IBC Virtual Cache Test The luminance block vector bvL (luminance block vector with 1 / 16 fractional sample accuracy) shall obey the following constraints: –CtbSizeY is greater than or equal to ((yCb+(bvL[1]>>4))&(CtbSizeY-1))+cbHeight. – IbcVirBuf[0][(x+(bvL[0]>>4))&(IbcBufWidthY-1)][(y+(bvL[1]>>4))&(CtbSizeY-1)] shall not be equal to -1 for x=xCb..xCb+cbWidth-1 and y=yCb..yCb+cbHeight-1. Otherwise, bvL is considered to be an invalid bv. Samples are processed in CTB units. The array size of each luma CTB in both width and height is CtbSizeY in samples. – (xCb, yCb) is the luminance position of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture, –cbWidth specifies the width of the current codec block in luminance samples. –cbHeight specifies the height of the current codec block in luma samples. 2.2.IBC Merge / AMVP List Construction The IBC Merge / AMVP list construction is modified as follows: ●An IBC Merge / AMVP candidate can be inserted into the IBC Merge / AMVP candidate list only if it is valid. ● Upper right airspace candidate, lower left airspace candidate and upper left airspace candidate (B0, A0 and B2, such as Figure 6 As shown, it shows the spatial neighboring positions used in IBC Merge / AMVP list construction), and a pairwise average candidate can be added to the IBC Merge / AMVP candidate list. ● Template-based adaptive reordering (ARMC-TM) is applied to the IBC Merge list. The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC Merge candidates with full deduplication, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBCMerge list. The candidates for the zero vector used to populate the IBC Merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. The zero vector is invalid as a block vector in IBC Merge mode and, therefore, is discarded as a BVP in the IBC candidate list. Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX parameter and the ΔY parameter, as shown in Figure 7 , which shows the padding candidates for replacing the zero vector in the IBC list. 2.3. IBC with Template Matching Template matching is used in IBC for both IBC Merge mode and IBC AMVP mode. Compared to the list used by the conventional IBC Merge mode, the IBC-TM Merge list is modified so that candidates are selected according to the deduplication method using the motion distance between candidates as in the conventional TM Merge mode. The ending zero motion fulfillment is replaced by the motion vectors of the left (-W, 0), above (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU. In IBC-TM Merge mode, the selected candidates are refined with a template matching method before the RDO or decoding process.The IBC-TM Merge mode has been made competitive with the regular IBC Merge mode and the TM-Merge flag is signaled. In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM Merge list. Each of those 3 selected candidates is refined using a template matching method and ranked according to their resulting template matching cost. Only the top 2 candidates are then typically considered in the motion estimation process. Template matching refinement for both IBC-TM Merge mode and AMVP mode is very simple because IBC motion vectors are constrained (i) to be integers and (ii) to be Figure 8 The reference region shown in Figure 1 shows the IBC reference region depending on the current CU position. Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision depending on the AMVR value. This refinement only accesses samples that are not interpolated. In both cases, the refined motion vectors and the template used in each refinement step must adhere to the constraints of the reference region. 2.4.IBC Reference Area The reference area of the IBC extends to the upper two CTU rows. Figure 9 The reference region used for encoding and decoding CTUs (m, n) is shown. Specifically, for a CTU (m, n) to be encoded or decoded, the reference region includes CTUs with indices (m-2, n-2) ... (W, n-2), (0, n-1) ... (W, n-1), (0, n) ... (m, n), where W represents the maximum horizontal index within the current slice, slice, or picture. When the CTU size is 256, the reference region is limited to one CTU row above. This setting ensures that IBC does not require additional memory in current ETM platforms for CTU sizes of 128 or 256. The per-sample block vector search (or local search) range is limited horizontally to [-(C<<1), C>>2] and vertically to [-C, C>>2] to accommodate reference region expansion, where C represents the CTU size. 2.5. Reconstruction Reordering IBC (RR-IBC) Allows the reconstruction-reordering IBC (RR-IBC) mode for blocks encoded and decoded with IBC. When RR-IBC is applied, the samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the predicted block is derived without flipping. On the decoder side, the reconstructed block is flipped back to restore the original block. For blocks encoded by RR-IBC, two flipping methods are supported, horizontal flipping and vertical flipping. First, for blocks encoded by IBCAVP, a syntax flag is signaled to indicate whether the reconstruction is flipped, and if flipped, another flag specifying the flipping type is further signaled. For IBC Merge, the flipping type is inherited from the adjacent block without syntax signaling. Taking into account horizontal symmetry or vertical symmetry, the current block and the reference block are usually aligned horizontally or vertically. Therefore, when horizontal flipping is applied, the vertical component of BV is not transmitted through the signal and is presumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of BV is not transmitted through the signal and is presumed to be equal to 0. Figure 10A A diagram showing BV adjustment for horizontal flipping is shown. Figure 10B A diagram showing BV adjustment for vertical flipping is shown. In order to better utilize the symmetry property, a flip-aware BV adjustment scheme is applied to refine the block vector candidates. Figure 10A and Figure 10B As shown, (x nbr ,y nbr ) and (x cur ,y cur ) represent the coordinates of the center sample points of the neighboring blocks and the current block, BV nbr and BV cur Represents the BV of the neighboring block and the current block respectively. Instead of inheriting the BV directly from the neighboring block, the motion displacement is added to the BV when the neighboring block is coded with horizontal flipping. nbr The horizontal component (expressed as BV nbr h ) to calculate BV cur The horizontal component of BV cur h =2(x nbr -x cur )+BV nbr h Similarly, in the case where the neighboring blocks are coded with vertical flipping, the motion displacement is added to BV nbr The vertical component (expressed as BV nbr v ) to calculate BV cur The vertical component of BV cur v =2(y nbr -y cur )+BV nbr v . 2.6. IBC Merge Mode with Block Vector Difference (IBC-MBVD) Affine MMVD and GPM-MMVD have been adopted as extensions of the conventional MMVD mode. It is natural to extend the MMVD mode to the IBCMercure mode. In IBC-MBVD, the distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions. A base candidate is selected from the first five candidates in the reordered IBC Merge list. And all possible MBVD refinement positions (20×4) for each base candidate are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. Finally, the first 8 refinement positions with the lowest template SAD cost remain as available positions and are therefore used for MBVD index encoding and decoding. The MBVD index is binarized by the Rice codec with a parameter equal to 1. Blocks encoded with IBC-MBVD do not inherit the flip type from neighboring blocks encoded with RR-IBC. 2.7. Intra-frame template matching Intra Template Matching (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. Figure 11 The intra-frame template matching search area used is shown. By combining the L-shaped causal neighbors of the current block with the following Figure 11 Generate a prediction signal by matching it with another block in a predefined search area: R1: current CTU, R2: Upper left CTU, R3: Upper CTU, R4: left CTU. The sum of absolute differences (SAD) is used as the cost function. In each region, the decoder searches for the template with the minimum SAD relative to the current one and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) with a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW, SearchRange_h=a*BlkH. Where 'a' is a constant that controls the gain / complexity tradeoff. In practice, 'a' is equal to 5. The intra template matching tool is enabled for CUs with width and height both less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.8. Using Block Vectors Derived from IntraTMP for IBC It is proposed to use the block vector derived from IntraTMP for IBC. The proposed method is to store the IntraTMP block vector in the IBC block vector cache, and the current IBC block can use both the IBC BV and IntraTMP BV of the neighboring blocks as BV candidates for the IBC BV candidate list, such as Figure 12 As shown, the Figure 12 The use of the IntraTMP block vector for an IBC block is shown. Figure 13A and Figure 13B An example of comparing block vector candidates from neighboring blocks of IBC-only codecs in an IBC block vector candidate list with block vector candidates from neighboring blocks of IBC and IntraTMP codecs in a proposed IBC block vector candidate list is shown. The IntraTMP block vector is added to the IBC block vector candidate list as a spatial candidate. Figure 13A An example of an IBC block vector candidate list in which only IBC block vectors exist is shown. Figure 13B An example of an IBC block vector candidate list in which an IBC block vector and an IntraTMP block vector exist is shown. It should be noted that the proposed method makes IBC block vector prediction more efficient by using different block vectors without additional memory for storing block vectors. 2.9. Adaptive Reordering of Merge Candidates with Template Matching (ARMC-TM) Merge candidates are adaptively reordered using Template Matching (TM). The reordering method is applied to the normal Merge mode, TM Merge mode and Affine Merge mode (excluding SbTMVP candidates). For TM Merge mode, Merge candidates are reordered before the refinement process. First, an initial Merge candidate list is constructed according to a given checking order, such as spatial, TMVP, non-adjacent, HMVP, paired, and virtual Merge candidates. The candidates in the initial list are then divided into several subgroups. For Template Matching (TM) Merge mode, Adaptive DMVR mode, each Merge candidate in the initial list is first refined by using TM / multi-pass DMVR. The Merge candidates in each subgroup are reordered to generate a reordered Merge candidate list, and the reordering is based on the cost value based on template matching. The index of the selected Merge candidate in the reordered Merge candidate list is transmitted to the decoder via a signal. For simplicity, the Merge candidates in the last rather than the first subgroup are not reordered. During the Merge motion vector candidate list, all zero candidates from the ARMC reordering process are excluded. For the normal Merge mode and the TM Merge mode, the subgroup size is set to 5. For the affine Merge mode, the subgroup size is set to 3. ●Cost calculation During the reordering process, the template matching cost of the Merge candidate is measured by the SAD between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate. When the Merge candidate uses bidirectional prediction, the reference samples of the template of the Merge candidate are also located by Figure 14 The bidirectional prediction shown in Figure 14 The template and the reference points of the template in the reference picture are shown. ●Refinement of the initial Merge candidate list When multi-pass DMVR is used to derive refined motion to the initial Merge candidate list, only the first pass of multi-pass DMVR (i.e., PU level) is applied in the reordering. When template matching is used to derive refined motion, the template size is set equal to 1. When the block is flat (where the block width is greater than 2 times the height) or the block is narrow and long (where the height is greater than 2 times the width), only the upper template or the left template is used during the motion refinement of the TM. The TM is expanded to perform 1 / 16 pixel MVD accuracy. The first four Merge candidates utilize refined motion reordering in TM Merge mode. For a sub-block-based Merge candidate with a sub-block size equal to Wsub×Hsub, the upper template includes several sub-templates of size Wsub×1, and the left template includes several sub-templates of size 1×Hsub. Figure 15 As shown, the Figure 15Templates of blocks with sub-block motion using motion information of sub-blocks of a current block and reference samples of the templates are shown, and reference samples of each sub-template are derived using motion information of sub-blocks in the first row and first column of the current block. Re-ranking criteria During the re-ranking process, a candidate is considered redundant if the cost difference between it and its predecessor is lower than the lambda value, e.g., |D1-D2| < λ, where D1 and D2 are the costs obtained during the first ARMC ranking and λ is the Lagrangian parameter used in the RD criterion on the encoder side. The proposed algorithm is defined as follows: - Determine the minimum cost difference between a candidate and its predecessor among all candidates in the list. ■ If the minimum cost difference is greater than or equal to λ, the list is considered large enough and reordering stops. ■ If the minimum cost difference is below λ, the candidate is considered redundant and it is moved at another position in the list. The other position is the first position where the candidate is sufficiently numerous compared to its predecessor. - The algorithm stops after a finite number of iterations (if the minimum cost difference is not lower than λ). This algorithm is applied to the normal, TM, BM and affine Merge modes. Similar algorithms are applied to the Merge MMVD prediction method and the symbolic MVD prediction method, which also use ARMC for reordering. The value of λ is set equal to the rate-distortion criterion λ used to select the best merge candidate for the low-delay configuration at the encoder side and the value λ corresponding to another QP for the random access configuration. A set of λ values corresponding to each signaled QP offset is provided in the SPS or in the slice header of the QP offset if it is not present in the SPS. ●Extension of AMVP model The ARMC design also applies to AMVP mode, where AMVP candidates are reordered based on their TM cost. For template matching for advanced motion vector prediction (TM-AMVP) mode, an initial AMVP candidate list is constructed, which is then refined from the TM to construct a refined AMVP candidate list. In addition, MVP candidates with a TM cost greater than a threshold, which is equal to five times the cost of the first MVP candidate, are skipped. Note that when surround motion compensation is enabled, MV candidates should consider surround offsets for cropping. 2.10. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Airspace MVP from airspace adjacent CU, 2) Temporal MVP from the same CU, 3) History-based MVP from FIFO table, 4) Pairwise average MVP, 5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU codec in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded using the context, and bypass coding is used for the other binary bits. The derivation process of Merge candidates for each category is provided in this session. As done in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a certain size. 2.10.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Figure 16 Select the maximum four Merge candidates from the candidates in the position depicted in Figure 16 The positions of the spatial Merge candidates are shown. The order of derivation is B1, A1, B0, A0 and B2. When one or more CUs at positions B0, A0, B1, A1 are not available (for example, because they belong to another strip or slice) or are intra-coded, only position B2 is considered. After adding the candidate for position B1, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, so that the coding efficiency is improved. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Figure 17 shows the candidate pairs that are considered for redundancy check of spatial Merge candidates. In contrast, only the pairs with Figure 17 The arrows in are used to link pairs, and a candidate is only added to the list if the corresponding candidate for redundancy check does not have the same motion information. 2.10.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, when deriving this time domain Merge candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list and the reference index used to derive the co-located CU are explicitly signaled in the slice header. Figure 18As shown by the dotted line in , a scaled motion vector for the temporal merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set to zero. The position of the time domain candidate is selected between candidate C0 and candidate C1, such as Figure 19 If the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used to derive the temporal merge candidate. 2.10.3. History-based Merge Candidate Derivation After spatial MVP and TMVP, history-based MVP (HMVP) merge candidates are added to the merge list. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-subblock inter-coded CU exists, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that up to 5 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find whether the same HMVP exists in the table. If found, the same HMVP is removed from the table, and all HMVP candidates are moved forward and the same HMVP is inserted into the last entry of the table. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the check list are inserted into the candidate list after the TMVP candidate. Redundancy check is applied to the HMVP candidates and to the spatial or temporal Merge candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: 1. The last two entries in the table are redundantly checked as A1 airspace candidates and B1 airspace candidates, respectively. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process from HMVP is terminated. 2.10.4. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list using the first two Merge candidates. The first Merge candidate is defined as p0Cand, and the second Merge candidate can be defined as p1Cand. The average motion vector is calculated for each reference list separately according to the availability of the motion vectors of p0Cand and p1Cand. If both motion vectors are available in one list, the two motion vectors are averaged even when they point to different reference images, and their reference images are set to one of p0Cand; if only one motion vector is available, the one motion vector is used directly; if no motion vector is available, the list is kept invalid. In addition, if the half-pixel interpolation filter index of p0Cand and p1Cand is different, it is set to 0. When the merge list is not full after adding pairwise average merge candidates, zero MVPs are inserted at the end until the maximum number of merge candidates is encountered. 2.10.5.Merge Estimation Region The Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). The generation of the Merge candidate list for the current CU does not include candidate blocks in the same MER as the current CU. In addition, the update process for the history-based motion vector prediction candidate list is only updated when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel) and where (xCb, yCb) is the top left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.10.6. Non-adjacent airspace candidates In ECM, non-adjacent spatial merge candidates as in JVET-L0399 are inserted after the TMVP in the regular merge candidate list. Figure 20 The pattern of spatial Merge candidates is shown in Figure 20 The spatial neighboring blocks used to derive spatial merge candidates are shown. The distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. No line buffer limit is applied. 2.10.7. ARMC based on MV candidate type Merge candidates of a single candidate type (e.g., TMVP or non-adjacent MVP (NA-MVP)) are reordered based on the ARMC TM cost value. The reordered candidates are then added to the Merge candidate list. The TMVP candidate type adds more TMVP candidates with more temporal locations and different inter-frame prediction directions to perform reordering and selection. In addition, the NA-MVP candidate type is further extended with more spatial non-adjacent locations. The target reference picture of the TMVP candidate can be selected from any reference picture in the list according to the scaling factor. The selected reference picture is the reference picture whose scaling factor is closest to 1. 2.11. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports sub-block-based temporal motion vector prediction (SbTMVP). Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the CU in the current picture. TMVP uses the same co-located pictures as SbTVMP. SbTMVP differs from TMVP in the following two main aspects: –TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; While TMVP extracts the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the center block relative to the current CU), SbTMVP applies a motion displacement before extracting the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks from the current CU. The SbTVMP process Figure 21A and Figure 21B Shown in. Figure 21A The spatial neighboring blocks used by ATVMP are shown. Figure 21B The sub-CU motion field is derived by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, check Figure 21A If A1 has a motion vector that uses the co-located picture as its reference picture, then that motion vector is selected as the motion offset to be applied. If no such motion is identified, then the motion offset is set to (0,0). In the second step, the motion offset identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located picture, as Figure 21B As shown in . Figure 21BThe example in assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the minimum motion grid covering the center sample) in the co-located picture is used to derive the motion information of the sub-CU. After identifying the motion information of the co-located sub-CU, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block-based Merge list containing SbTVMP candidates and affine Merge candidates is used for signaling of the sub-block-based Merge mode. The SbTVMP mode is enabled / disabled by the sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block-based Merge candidates, followed by the affine Merge candidates. The size of the sub-block-based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8x8, and as for the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8. The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, that is, for each CU in a P slice or a B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 3. Question In current BV prediction (eg, for both IBC Merge and IBC AMVP modes), time-domain BV prediction is not utilized. In order to further improve the efficiency of BV prediction, time-domain BV prediction is introduced. 4. Detailed solution The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. The term "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB, or a video processing unit including multiple samples / pixels. A block may be rectangular or non-rectangular. W and H are the width and height of the current block (eg, luma block). For blocks in IBC and intra-frame TMP codecs, the block vector (BV) is used to indicate the displacement from the current block to a reference block that has been or partially reconstructed within the current picture. In the following, a BV candidate is a BV prediction value or a search point. A block has BV information if it is IBC coded or intra TMP coded. Time domain BV prediction 1. In one example, time-domain BV prediction can be introduced into BV prediction. a. In one example, the BV prediction may be at least one of the following. (a) In one example, the BV prediction may be a regular IBC Merge prediction. (b) In one example, the BV prediction may be a conventional IBC AMVP prediction. (c) In one example, the BV prediction may be an IBC-TM Merge prediction. (d) In one example, the BV prediction may be an IBC-TM AMVP prediction. (e) In one example, the BV prediction may be a RR-IBC Merge prediction. (f) In one example, the BV prediction may be a RR-IBC AMVP prediction. (g) In one example, the BV prediction may be an IBC-MBVD prediction. (h) In one example, the BV prediction may be a string copy vector prediction. (i) In one example, the BV prediction may be any other BV prediction. 2. In one example, time-domain BV candidates may be introduced into the BV candidate list. a. In one example, the BV candidate list may be at least one of the following. (a) In one example, the BV candidate list may be a regular IBC Merge list. (b) In one example, the BV candidate list may be a regular IBC AMVP list. (c) In one example, the BV candidate list may be an IBC-TM Merge list. (d) In one example, the BV candidate list may be an IBC-TM AMVP list. (e) In one example, the BV candidate list may be an RR-IBC Merge list. (f) In one example, the BV candidate list may be a RR-IBC AMVP list. (g) In one example, the BV candidate list may be an IBC-MBVD basic candidate list. (h) In one example, the BV candidate list may be any other BV candidate list. 3. In one example, the temporal BV prediction or candidate may be derived in at least one of the following methods. a. In one example, if a motion grid (such as a 4x4 grid) covering a temporal position is available, has BV information, and its BV is valid for the current block, then the temporal position can be used for temporal BV candidate derivation. b. In one example, if a motion grid (such as a 4x4 grid) covering a temporal position is not available, or does not have BV information, or its BV is invalid for the current block, then the temporal position may not be used for temporal BV candidate derivation. c. In one example, if the motion grid (such as a 4x4 grid) covering a temporal position is outside the CTU row of the current block, the temporal position can be clipped to inside the CTU row of the current block and then used for temporal BV candidate derivation. (a) Alternatively, if a moving grid (such as a 4x4 grid) covering one temporal location Outside the CTU row of the current block, the temporal position may not be used for temporal BV candidate derivation. d. In one example, the position for the temporal BV candidate may be selected among several positions in the co-located picture. (a) These locations include: Figure 22B C0 and C1 in the co-located picture depicted in Figure 22B Candidate positions for time domain candidates are shown. (b) C0 may be checked first. If BV is not available in C0, then C1 is checked. (c) C1 may be checked first. If BV is not available in C1, C0 is checked. (d) If the CU at position C0 is not available, has no BV information, is outside the CTU row of the current block, or its BV is invalid for the current block, then position C1 is used. Otherwise, position C0 is used to derive the temporal BV candidate. This means that the priority order is C0->C1. (e) Alternatively, the position for the temporal BV candidate can be selected between positions C0 and C1 in the co-located picture, as Figure 22B Depicted. If the CU at position C1 is not available, has no BV information, is outside the CTU row of the current block, or its BV is invalid for the current block, then position C0 is used. Otherwise, position C1 is used to derive the temporal BV candidate. This means the priority order is C1->C0. e. Alternatively, when deriving the temporal BV candidate, the candidates corresponding to positions C0 and C1 in the co-located picture can be used, such as Figure 22B Described in. (a) For example, the derivation order is C0, C1. (b) Alternatively, the derivation order is C1, C0. f. In one example, the width and height of the co-located block in the co-located picture may be the same as the width and height of the current block in the current picture. g. In one example, the position of the co-located block in the co-located picture may be the same as the position of the current block in the current picture. h. In one example, the position of the co-located block in the co-located picture may be determined by adding a motion offset to the position of the current block in the current picture. (a) In one example, the motion displacement can be a motion vector of a spatial neighbor. 1) In one example, the spatial neighbors can be Figure 22A The left (A1), top (B1), top right (B0), bottom left (A0), or top left (B2) neighbor in Figure 22A Candidate positions for spatial candidates are shown. 2) In one example, if the spatial neighbor has a motion vector that uses the co-located picture as its reference picture, this motion vector can be selected as the motion displacement; if no such motion is identified, the spatial neighbor may not provide a motion displacement or the motion displacement may be set to (0,0). 3) In one example, if a spatial neighbor has a motion vector that uses the co-located picture as its reference picture, then this motion vector can be selected as the motion displacement; if no such motion is identified, then a motion vector of reference list 0 or reference list 1 can be scaled to point to the co-located picture, and the scaled motion vector can be used as the motion displacement. 4) In one example, the motion displacement may be derived in a predefined priority order, and the first N valid motion vector(s) may be used as the motion displacement(s). i. In one example, N can be 1, 2, 3, 4, or 5. ii. In one example, the priority order may be A1->B1->B0->A0->B2. iii. In one example, the priority order may be B1->A1->B0->A0->B2. iv. In one example, the priority order may be A0->A1->B0->B1->B2. (b) In one example, the motion displacement(s) with the top M minimum template matching costs may be used to derive temporal BV candidates. 1) In one example, M can be 1, 2, 3, 4, or 5. i. In one example, when deriving the temporal BV candidate, at least one of the following may be used: a candidate selected from position C0 or C1, a candidate selected from position C0Left or C1Left, a candidate selected from position C0Above or C1Above, a candidate selected from position C0AboveRight or C1AboveRight, a candidate selected from position C0BottomLeft or C1BottomLeft, a candidate selected from position C0AboveLeft or C1AboveLeft, and these positions are in the same picture, such as Figure 23 As shown, where CXSpatial is the motion displacement derived from the spatial neighbors added to CX (X is 0 or 1, and spatial is left, above, right above, left below, or left above). Figure 23 The candidate positions for the temporal BV candidates are shown, and the spatial domain can be left, above, above right, below left, or above left. (a) In one example, up to six time-domain BV candidates can be derived. (b) In one example, when deriving the temporal BV candidate, a candidate selected from position C0 or C1, a candidate selected from position C0Left or C1Left, which are in the same picture, may be used. Figure 24 In this example, at most two time-domain BV candidates can be derived. Figure 24 Candidate positions of the time-domain BV candidates are shown. (c) In one example, the priority order of C0 and C1 is C0->C1. (d) In one example, the priority order of C0 and C1 is C1->C0. (e) In one example, the priority order of C0Spatial and C1Spatial may be the same as the priority order of C0 and C1. 1) In one example, the airspace can be left, above, above right, below left, or above left. (f) In one example, the priority order of C0Spatial and C1Spatial may be opposite to the priority order of C0 and C1. 1) In one example, the airspace can be left, above, above right, below left, or above left. j. In one example, when deriving the temporal BV candidate, at least one of the following may be used: a candidate derived from positions C0 and C1, a candidate derived from positions C0Left and C1Left, a candidate derived from positions C0Above and C1Above, a candidate derived from positions C0AboveRight and C1AboveRight, a candidate derived from positions C0BottomLeft and C1BottomLeft, a candidate derived from positions C0AboveLeft and C1AboveLeft, and these positions are in the same picture, such as Figure 23 where CXSpatial is the motion displacement derived from the spatial neighbors added to CX (X is 0 or 1, and spatial is left, above, above right, below left, or above left). (a) In one example, up to 12 time-domain BV candidates can be derived. (b) In one example, when deriving temporal BV candidates, candidates derived from positions C0 and C1, and candidates derived from positions C0Left and C1Left, which are in the same picture, may be used. Figure 24 In this example, up to four time-domain BV candidates can be derived. (c) In one example, the derivation order of C0 and C1 is C0, C1. (d) In one example, the order of derivation of C0 and C1 is C1, C0. (e) In one example, the order of derivation of C0Spatial and C1Spatial can be the same as that of C0 The derivation order is the same as C1. 1) In one example, the airspace can be left, above, above right, below left, or above left. (f) In one example, the order of derivation of C0Spatial and C1Spatial may be opposite to the order of derivation of C0 and C1. 1) In one example, the airspace can be left, above, above right, below left, or above left. k. In one example, the temporal BV candidates can be derived from some specific temporal locations. (a) In one example, the time domain location may be predefined. (b) In one example, the time domain position can be derived based on some codec information. (c) In one example, the temporal position may be derived based on at least one of the position, width, or height of the current block. (d) In one example, the distance between the temporal BV candidate and the current coding block may be based on the width and height of the current coding block. 1) In one example, the mode of the time domain BV candidate is Figure 25 Shown in. Figure 25 The first pattern of candidate positions for time domain BV candidates is shown. For each search round, four time domain positions are checked. For each search round i (i>=0), the four time domain positions are {(x+W+i*W), (y+H+i*H)} (RB i ),,{(x+W / 2+i*W),(y+H / 2+i*H)}(Ctr i )、{(x+W+i*W),(y+H / 2)}(R i ), and {(x+W / 2),(y+H+i*H)}(B i ). i. In one example, if five search rounds are used, the 20 time domain positions are {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H / 2)}, {(x+W / 2), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+W / 2+2*W), (y+H / 2+2*H)}, {(x+W+2*W),(y+H / 2)},{(x+W / 2),(y+H+2*H)},{(x+W+3*W),(y+H+3*H)},{(x+W / 2+3*W),(y+H / 2+3*H)},{(x+W+3*W),(y+H / 2)},{ (x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}. i. In one example, for each search round i, RB i ->Ctr i Derived a time domain BV candidate in the order of priority, with R i ->B i A time-domain BV candidate is derived in the order of priority, and at most two time-domain BV candidates can be derived. ii. In one example, for each search round i, the derivation order is RB i 、Ctr i 、R i 、B i , up to four time-domain BV candidates can be derived. 2) In one example, the mode of the time domain BV candidate is Figure 26 Shown in. Figure 26 The second mode of candidate positions for time domain BV candidates is shown. For each search round, four time domain positions are checked. For each search round i (i>=1), the four time domain positions are {(x+W+i*W), (y+H+i*H)} (RB i ), {(x+W / 2+i*W),(y+H / 2+i*H)}(Ctr i )、{(x+W+i*W),(y+H / 2)}(R i ), and {(x+W / 2),(y+H+i*H)}(B i For each search round 0, the four time domain positions are {(x+W), (y+H)} (RB0), {(x+W / 2), (y+H / 2)} (Ctr0), {(x+W), (y+H-4)} (R0), {(x+W-4), (y+H)} (B0). i. In one example, if five search rounds are used, the 20 time domain positions are {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H-4))}, {(x+W-4), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+W / 2+2*W), (y+H / 2+2*H)}, {(x+W+2*W),(y+H / 2)},{(x+W / 2),(y+H+2*H)},{(x+W+3*W),(y+H+3*H)},{(x+W / 2+3*W),(y+H / 2+3*H)},{(x+W+3*W),(y+H / 2)},{ (x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}. ii. In one example, for each search round i, RB i ->Ctr i Derived a time domain BV candidate in the order of priority, with R i ->B i A time-domain BV candidate is derived in the order of priority, and at most two time-domain BV candidates can be derived. iii. In one example, for each search round i, the derivation order is RB i 、Ctr i 、R i 、B i , up to four time-domain BV candidates can be derived. 3) In one example, any other pattern of time-domain BV candidates may be used. 1. In one example, all the above time-domain BV candidates can be combined in any way. m. In one example, there may be a constraint on the maximum number of time-domain BV candidates (eg, N). (a) In one example, the number of time-domain BV candidates may be no greater than 5. (b) In one example, the number of time-domain BV candidates may be no greater than 4. (c) Alternatively, there may be a constraint on the maximum number (eg, M) of temporal BV candidates to be derived that can be unique (eg, after full deduplication). 1) In one example, M can be 5. 2) In one example, M can be 4. 3) In one example, M may vary according to the coding mode of the current block. i. In one example, for IBC-TM AMVP and / or IBC-TM Merge modes, M may be 1 or 2; for other IBC modes, M may be 4 or 5. n. In one example, redundancy checking or deduplication can be performed when deriving time-domain BV candidates. (a) In one example, full deduplication can be performed when deriving temporal BV candidates to ensure that candidates with the same or similar motion information are excluded from the BV candidate list. (b) In one example, partial deduplication can be performed when deriving time-domain BV candidates. o. In one example, the position of the time-domain BV candidate in the BV candidate list can be one of the following. (a) In one example, all time-domain BV candidates may be inserted before the HMVP candidate. (b) In one example, part of the time-domain BV candidates may be inserted before the HMVP candidate, and the remaining time-domain BV candidates may be inserted after the HMVP candidate. (c) In one example, all time-domain BV candidates may be inserted after the HMVP candidate. 4. In one example, the number of co-located pictures used to derive temporal BV / MV candidates may be N (eg, N is a positive integer). a. In one example, N may be greater than or equal to 1. b. In one example, the indication of the co-located pictures used to derive the temporal BV candidate may be signaled at the sequence level / group of picture level / picture level / slice level / slice group level, such as in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. c. In one example, N reference pictures with the top N smallest POC distances relative to the current picture may be selected as co-located pictures. d. In one example, N reference pictures with the top N smallest QP differences relative to the current picture may be selected as co-located pictures. e. In one example, N reference pictures with the first N smallest QPs may be selected as co-located pictures. 5. In one example, whether to use time-domain BV prediction (TBVP) and whether to use time-domain MV prediction (TMVP) can use the same indication. a. In one example, whether to use time domain BV prediction (TBVP) and whether to use time domain MV prediction (TMVP) can use different indicators. b. In one example, whether to use time domain BV prediction (TBVP) can be transmitted by signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. BV candidate re-ranking 6. In one example, a re-ranking / refinement process may be performed when deriving the BV candidate list. a. In one example, the re-ranking / refinement process can be based on template matching cost(s). b. In one example, when constructing the BV candidate list, N1 adjacent spatial candidates and / or N2 temporal candidates and / or N3 HMVP candidates and / or N4 pairwise average candidates and / or N5 predefined BV candidates may be partially or fully derived using full deduplication to ensure that there are no duplicate or similar candidates in the list, and then reordered together. After reordering, the top N candidates (such as those with the lowest cost) may be selected as the final candidates in the BV candidate list. i. In one example, N may be 6 and / or N1 may be 5 and / or N2 may be 10 and / or N3 may be 25 and / or N4 may be 1 and / or N5 may be 6. ii. In one example, there may be a constraint on the maximum number (eg, M) of BV candidates to be derived that can be unique (eg, after full deduplication). (i) In one example, M may be 20. iii. In one example, the adjacent spatial domain BV candidate may be composed of the left and / or above and / or above right and / or below left and / or above left spatial domain candidates (example in Figure 22A shown in ). iv. In one example, the time-domain BV candidates may consist of those specified in list item 3. v. In one example, the number of HMVP candidates and / or the HMVP table size may be increased to N2 (eg, 25). vi. In one example, for paired BV candidates, they may be generated by averaging predefined pairs of existing candidates in the motion candidate list. (ii) In one example, the predefined pairs may be defined as pairs in a set such as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the motion candidate indices in the motion candidate list. vii. In one example, the predefined BV candidates may be located in the IBC reference region. c. In one example, BV candidate type-based ARMC can be used to reorder BV candidates of a specific candidate type or multiple specific candidate types according to one or more criteria. i. In one example, when building a BV candidate list, M candidates with a particular candidate type (such as with the lowest cost) may be selected from the N re-ranked candidates with that candidate type. (i) In one example, M may vary according to the candidate type and / or coding mode of the current block. (ii) In one example, the candidate type may be a neighboring spatial domain BV candidate. For example, M is 4 and N is 5. (iii) In one example, the candidate type may be a time-domain BV candidate. For example, M is 4 and N is 10. (iv) In one example, the candidate type may be an HMVP candidate. For example, M is 10 and N is 25. (v) In one example, the candidate type may be a pairwise average BV candidate. For example, M is 1 and N is 6. (vi) In one example, the candidate type may be a predefined BV candidate, for example, M is 1 and N is 6. ii. In one example, multiple BV candidate types (i.e., candidate type combinations) can be Reorder together. (i) In one example, when constructing a BV candidate list, M candidates having any one of a specific BV candidate type (such as having the lowest cost) can be selected from N reordered candidates in the candidate type combination, where M can vary depending on the candidate type combination and / or the encoding mode of the current block. (ii) In one example, adjacent spatial candidates and / or temporal candidates and / or HMVP candidates and / or pairwise average candidates and / or predefined BV candidates may be reordered together. For example, M is 6 and N is 20. (iii) In one example, at least one BV candidate type of the BV candidates may first be reordered using BV candidate type-based ARMC. (iv) In one example, N1 HMVP candidates (such as those with the lowest cost) may be selected from the reordered candidates having the HMVP candidate type, and the selected N1 HMVP candidates may be reordered together with adjacent spatial domain candidates and / or temporal domain candidates and / or pairwise average candidates and / or predefined BV candidates. Finally, M candidates (such as those with the lowest cost) may be selected. (v) In one example, N2 time domain candidates (such as those with the lowest cost) may be selected from the reordered candidates having the time domain candidate type, and the selected N2 time domain candidates may be reordered together with adjacent spatial domain candidates and / or HMVP candidates and / or pairwise average candidates and / or predefined BV candidates. Finally, M candidates (such as those with the lowest cost) may be selected. (vi) In one example, if a candidate is re-ranked more than once, its re-ranking criterion (eg, template matching cost) may be reused. BVP against SbTMVP 7. In one example, BVP may be obtained for sub-blocks (such as 4x4 or 8x8) of a block coded using SbTMVP. a. In one example, the BVP can be obtained from the temporal location in the co-located block located by the SbTMVP. General information 8. The syntax elements disclosed above can be binarized as flags, fixed length codes, EG(x) codes, unary codes, truncated unary codes, truncated binary codes, etc. It can be signed or unsigned. 9. The syntax elements disclosed above can be encoded or decoded using at least one context model, or they can be bypassed. 10. The syntax elements (SEs) disclosed above may be conditionally signaled. a. The SE is signaled only when the corresponding function is applicable. 11. The syntax elements disclosed above can be transmitted by signal at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 12. In the above examples, a block may refer to a color component / sub-picture / slice / slice / codec tree unit (CTU) / CTU row / CTU group / codec unit (CU) / prediction unit (PU) / transform unit (TU) / codec tree block (CTB) / codec block (CB) / prediction block (PB) / transform block (TB) / block / subblock of a block / subregion within a block / any other region containing more than one sample or pixel. 13. Whether and / or how to apply the method disclosed above can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 14. Whether and / or how the above disclosed methods can be applied to be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. 15. Whether and / or how to apply the above disclosed method may depend on coded information such as block size, color format, single-tree partitioning / dual-tree partitioning, color component, slice / picture type.

[0091] Figure 27 27. A flow chart of a method 2700 for video processing according to an embodiment of the present disclosure is shown. The method 2700 is implemented during conversion between a current video block of a video and a bitstream of the video.

[0092] At block 2710, at least one of a temporal block vector (BV) prediction or a temporal BV candidate for the current video block is determined. For example, the temporal BV prediction may be introduced into the BV prediction. For another example, the temporal BV candidate may be introduced into the BV candidate list.

[0093] At block 2720, conversion is performed based on at least one of the temporal BV prediction or the temporal BV candidate. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.

[0094] Method 2700 enables the use of time-domain BV prediction or time-domain BV candidates. In this way, the efficiency of BV prediction can be improved, thereby improving codec efficiency and codec effectiveness.

[0095] In some embodiments, temporal BV prediction is introduced in at least one of the following: conventional intra block copy (IBC) Merge prediction, conventional IBC advanced motion vector prediction (AMVP) prediction, IBC template matching (IBC-TM) Merge prediction, IBC-TM AMVP prediction, reconstruction-reordering IBC (RR-IBC) Merge prediction, RR-IBC AMVP prediction, IBC Merge mode with block vector difference (IBC-MBVD) prediction, string copy vector prediction, or another BV prediction.

[0096] In some embodiments, the temporal BV candidate is included in a BV candidate list. In some embodiments, the BV candidate list includes at least one of the following: a conventional intra block copy (IBC) Merge candidate list, a conventional IBC advanced motion vector prediction (AMVP) candidate list, an IBC template matching (IBC-TM) Merge candidate list, an IBC-TM AMVP candidate list, a reconstruction-reordering IBC (RR-IBC) Merge candidate list, an RR-IBC AMVP candidate list, an IBC Merge mode with block vector difference (IBC-MBVD) basic candidate list, or another BV candidate list.

[0097] In some embodiments, determining at least one of a temporal BV prediction or a temporal BV candidate comprises: determining whether a set of conditions are satisfied, the set of conditions comprising: a first condition that a motion grid of a co-located block covering the temporal position of the current video block is available; a second condition that the motion grid has BV information; and a third condition that the BV associated with the motion grid is valid for the current video block; and if it is determined that the set of conditions are satisfied, determining at least one of a temporal BV prediction or a temporal BV candidate based on the temporal position. For example, if a motion grid (such as a 4x4 grid) covering a temporal position is available, has BV information, and its BV is valid for the current block, then the temporal position can be used for temporal BV candidate derivation.

[0098] In some embodiments, if at least one condition in the set of conditions is not met, the temporal position is not used to determine at least one of a temporal BV prediction or a temporal BV candidate. For example, if a motion grid (such as a 4x4 grid) covering a temporal position is not available, or does not have BV information, or its BV is invalid for the current block, then the temporal position may not be used for temporal BV candidate derivation.

[0099] In some embodiments, determining at least one of a temporal BV prediction or a temporal BV candidate includes: if it is determined that a motion grid of a co-located block covering the temporal position of the current video block is outside a codec tree unit (CTU) row of the current video block, performing a clipping operation on the temporal position to obtain a clipped temporal position inside the CTU row; and determining at least one of a temporal BV prediction or a temporal BV candidate based on the clipped temporal position. For example, if a motion grid (such as a 4x4 grid) covering a temporal position is outside a CTU row of the current block, the temporal position can be clipped to be inside the CTU row of the current block and then used for temporal BV candidate derivation.

[0100] In some embodiments, if the motion grid of the co-located block covering the temporal position of the current video block is outside the codec tree unit (CTU) row of the current video block, the temporal position is not used to determine at least one of the temporal BV prediction or the temporal BV candidate. That is, if the motion grid (such as a 4x4 grid) covering a temporal position is outside the CTU row of the current block, then the temporal position may not be used for temporal BV candidate derivation.

[0101] As used herein, the term "motion grid" may refer to a unit such as a minimum unit that stores motion information. In some embodiments, the motion grid comprises a 4×4 grid.

[0102] In some embodiments, determining at least one of a temporal BV prediction or a temporal BV candidate includes: determining a temporal position from a plurality of positions in a co-located picture of the current video block; and determining at least one of a temporal BV prediction or a temporal BV candidate based on the temporal position. For example, the position for the temporal BV candidate may be selected from among a plurality of positions in the co-located picture.

[0103] In some embodiments, the plurality of positions include a first position at the lower right of the co-located block of the current video block in the co-located block and a second position at the center of the co-located block. For example, the first position may be Figure 22B C0 in, and the second position can be Figure 22B C1 in.

[0104] In some embodiments, determining the time domain position includes: determining whether the BV is available in the first position; if it is determined that the BV is not obtained in the first position, determining whether the BV is available in the second position; and if it is determined that the BV is obtained in the second position, determining the second position as the time domain position.

[0105] In some embodiments, determining the time domain position includes: determining whether the BV is available in the second position; if it is determined that the BV is not obtained in the second position, determining whether the BV is available in the first position; and if it is determined that the BV is obtained in the first position, determining the first position as the time domain position.

[0106] In some embodiments, determining the time domain position includes determining the time domain position based on a priority order of the first position and the second position.

[0107] In some embodiments, the priority order includes an order in which the first position takes precedence over the second position. Determining the temporal position includes: determining whether at least one of the following conditions is satisfied: a first condition, the first condition being that a codec unit (CU) at the first position is unavailable; a second condition, the second condition being that the CU at the first position has no BV information; a third condition, the third condition being that the CU at the first position is outside a codec tree unit (CTU) row of the current video block; or a fourth condition, the fourth condition being that the BV of the CU at the first position is invalid for the current video block; if it is determined that at least one condition is satisfied, determining the second position as the temporal position; and if it is determined that none of the conditions are satisfied, determining the first position as the temporal position.

[0108] In some embodiments, the priority order includes an order in which the second position takes precedence over the first position. Determining the temporal position includes: determining whether at least one of the following conditions is satisfied: a first condition, the first condition being that a codec unit (CU) at the second position is unavailable; a second condition, the second condition being that the CU at the second position has no BV information; a third condition, the third condition being that the CU at the second position is outside a codec tree unit (CTU) row of the current video block; or a fourth condition, the fourth condition being that the BV of the CU at the second position is invalid for the current video block; if it is determined that at least one condition is satisfied, determining the first position as the temporal position; and if it is determined that none of the conditions are satisfied, determining the second position as the temporal position.

[0109] In some embodiments, a plurality of BV candidates for the current video block are determined based on a plurality of positions in a co-located block of the current video block.

[0110] In some embodiments, the plurality of positions include a first position at the lower right of the co-located block of the current video block in the co-located picture, and a second position at the center of the co-located block. For example, the first position may be Figure 22B In C0, the second position can be Figure 22B C1 in . A plurality of BV candidates are determined based on the plurality of positions and the order of the plurality of positions.

[0111] In some embodiments, the order includes one of: a first order in which the first position precedes the second position, or a second order in which the first position follows the second position.

[0112] In some embodiments, the width and height of the co-located block in the co-located picture are the same as the width and height of the current video block in the current picture.

[0113] In some embodiments, the position of the co-located block in the co-located picture is the same as the position of the current video block in the current picture.

[0114] In some embodiments, the position of the co-located block in the co-located picture is determined based on the motion displacement and the position of the current video block in the current picture.

[0115] In some embodiments, the motion displacement comprises motion vectors of spatial neighbors of the current video block.

[0116] In some embodiments, the spatial neighbor comprises one of a plurality of spatial neighbors. The plurality of spatial neighbors comprises: a first spatial neighbor to the left of the current video block, such as Figure 22A A1 shown in ; the second spatial neighbor above the current video block, such as Figure 22A B1 shown in ; the third spatial neighbor to the upper right of the current video block, such as Figure 22A B0 shown in ; the fourth spatial neighbor to the lower left of the current video block, such as Figure 22A A0 shown in ; and the fifth spatial neighbor to the upper left of the current video block, such as Figure 22A B2 shown in .

[0117] In some embodiments, determining the motion displacement includes determining at least one valid motion vector of at least one spatial neighbor of the current video block as the at least one motion displacement, the at least one motion displacement being determined according to a predetermined priority order of the plurality of spatial neighbors.

[0118] In some embodiments, the at least one valid motion vector comprises a number of valid motion vectors, the number being one of: 1, 2, 3, 4, or 5.

[0119] In some embodiments, the predetermined priority order includes one of the following: a first priority order of the first airspace neighbor, the second airspace neighbor, the third airspace neighbor, the fourth airspace neighbor and the fifth airspace neighbor; a second priority order of the second airspace neighbor, the first airspace neighbor, the third airspace neighbor, the fourth airspace neighbor and the fifth airspace neighbor; and a third priority order of the fourth airspace neighbor, the first airspace neighbor, the third airspace neighbor, the second airspace neighbor and the fifth airspace neighbor.

[0120] In some embodiments, if the candidate motion vector of the candidate spatial neighbor uses a co-located picture as a reference picture of the candidate spatial neighbor, the candidate motion vector is determined to be a motion displacement.

[0121] In some embodiments, no candidate spatial neighbor's candidate motion vector uses the co-located picture as its reference picture, the motion displacement includes a zero vector, or the candidate spatial neighbor does not have a motion displacement. For example, if the spatial neighbor has a motion vector that uses the co-located picture as its reference picture, then this motion vector may be selected as the motion displacement; if no such motion is identified, then the spatial neighbor may not provide a motion displacement or the motion displacement may be set to (0,0).

[0122] In some embodiments, a candidate motion vector without a candidate spatial neighbor uses a co-located picture as a reference picture of the candidate spatial neighbor, another motion vector of one of the first reference picture list or the second reference picture list is scaled to point to the co-located picture, and the scaled another motion vector is determined as a motion displacement.

[0123] In some embodiments, determining at least one of the BV predictions or the BV candidate comprises: determining a set of template matching costs for a set of motion displacements associated with the current video block; determining at least one motion displacement from the set of motion displacements based on an order of the set of template matching costs; and determining at least one of the BV predictions or the BV candidate based on the at least one motion displacement.

[0124] In some embodiments, the number of at least one motion displacement comprises one of: 1, 2, 3, 4, or 5.

[0125] In some embodiments, the temporal BV candidates include at least one temporal BV candidate selected from the following: a candidate determined based on a first position of a co-located block of the current video block in the co-located picture, or a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, and a set of candidates determined based on a set of shifted first positions or a set of shifted second positions, wherein the set of shifted first positions are shifted from the first position based on a set of motion displacements associated with a set of spatial neighbors of the current video block, and the set of shifted second positions are shifted from the second position based on the set of motion displacements.

[0126] In some embodiments, a set of spatial neighbors includes at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor above the right of the current video block, a fourth spatial neighbor below the left of the current video block, and a fifth spatial neighbor above the left of the current video block.

[0127] In some embodiments, the plurality of positions include a first position at the lower right of the co-located block of the current video block in the co-located block and a second position at the center of the co-located block. For example, the first position may be Figure 22B C0 in, and the second position can be Figure 22B C1 in.

[0128] In some embodiments, the number of the at least one time-domain BV candidates is less than or equal to six.

[0129] In some embodiments, the set of spatial neighbors includes a first spatial neighbor to the left of the current video block.

[0130] In some embodiments, the number of at least one temporal BV candidate is less than or equal to two.

[0131] In some embodiments, the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

[0132] In some embodiments, the priority order of the shifted first position and the shifted second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

[0133] In some embodiments, the displaced first position and the displaced second position are motion-displaced based on spatial neighbors, wherein the spatial neighbors include at least one of: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor above the right of the current video block, a fourth spatial neighbor below the left of the current video block, and a fifth spatial neighbor above the left of the current video block.

[0134] In some embodiments, the temporal BV candidates include at least one temporal BV candidate selected from the following: a candidate determined based on a first position of a co-located block of the current video block in the co-located picture, a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, a set of candidates determined based on a set of shifted first positions, the set of shifted first positions being shifted from the first position based on a set of motion displacements associated with a set of spatial neighbors of the current video block, and a set of candidates determined based on a set of shifted second positions, the set of shifted second positions being shifted from the second position based on a set of motion displacements.

[0135] In some embodiments, a set of spatial neighbors includes at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor above the right of the current video block, a fourth spatial neighbor below the left of the current video block, and a fifth spatial neighbor above the left of the current video block.

[0136] In some embodiments, the plurality of positions include a first position at the lower right of the co-located block of the current video block in the co-located picture and a second position at the center of the co-located block. For example, the first position may be Figure 22B C0 in, and the second position can be Figure 22B C1 in.

[0137] In some embodiments, the number of the at least one time-domain BV candidates is less than or equal to twelve.

[0138] In some embodiments, the set of spatial neighbors includes a first spatial neighbor to the left of the current video block.

[0139] In some embodiments, the number of the at least one time-domain BV candidates is less than or equal to four.

[0140] In some embodiments, the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

[0141] In some embodiments, the priority order of the shifted first position and the shifted second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

[0142] In some embodiments, the displaced first position and the displaced second position are motion-displaced based on spatial neighbors, wherein the spatial neighbors include at least one of: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor above the right of the current video block, a fourth spatial neighbor below the left of the current video block, and a fifth spatial neighbor above the left of the current video block.

[0143] In some embodiments, at least one temporal BV candidate is determined based on a set of temporal locations.

[0144] In some embodiments, a set of time domain locations is predefined.

[0145] In some embodiments, a set of time domain locations is determined based on codec information.

[0146] In some embodiments, the set of temporal positions is determined based on at least one of: a position of the current video block, a width of the current video block, or a height of the current video block.

[0147] In some embodiments, at least one distance between at least one temporal BV candidate and the current video block is based on a width and a height of the current video block.

[0148] In some embodiments, at least one time-domain BV candidate in the first mode is determined by multiple search rounds, wherein in a search round in the multiple search rounds, multiple time-domain positions are checked, wherein the multiple time-domain positions include: a position {(x+W+i*W), (y+H+i*H)} denoted as RBi, a position {(x+W / 2+i*W), (y+H / 2+i*H)} denoted as Ctri, a position {(x+W+i*W), (y+H / 2)} denoted as Ri, and a position {(x+W / 2), (y+H+i*H)} denoted as Bi, and wherein (x, y) represents the position of the current video block, W represents the width of the current video block, H represents the height of the current video block, i represents the index of the search round, and i is greater than or equal to 0.

[0149] In some embodiments, the plurality of search rounds includes 5 search rounds, and 20 time domain positions are examined during the 5 search rounds, the 20 time domain positions including: {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H / 2)}, {(x+W / 2), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+W / 2+2*W), (y+H / 2+2*H)}, {(x+W+2*W),(y+H / 2)}, {(x+W / 2),(y+H+2*H)}, {(x+W+3*W),(y+H+3*H)}, {(x+W / 2+3*W),(y+H / 2+3*H)}, {(x+W+3*W),(y+ H / 2)}, {(x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}. For example, these time domain locations are in Figure 25 The time domain BV candidate pattern is shown in Figure 25 Shown in.

[0150] In some embodiments, for a search round with index i, a first time-domain BV candidate is determined based on a priority order of RBi over Ctri, and a second time-domain BV candidate is determined based on a priority order of Ri over Bi, and at least one time-domain BV candidate includes at most two time-domain BV candidates.

[0151] In some embodiments, for a search round with index i, a first time-domain BV candidate is determined based on a priority order of RBi over Ctri, Ctri over Ri, and Ri over Bi, and the at least one time-domain BV candidate includes at most four time-domain BV candidates.

[0152] In some embodiments, at least one time-domain BV candidate in the second mode is determined by multiple search rounds, wherein in a search round in the multiple search rounds, multiple time-domain positions are checked, wherein for a search round with an index i greater than or equal to 1, the multiple time-domain positions include: a position denoted as RBi {(x+W+i*W), (y+H+i*H)}, a position denoted as Ctri {(x+W / 2+i*W), (y+H / 2+i*H)}, a position denoted as Ri {(x+W+i*W), (y+H / 2)}, and and a position {(x+W / 2), (y+H+i*H)} denoted as Bi, wherein (x, y) represents the position of the current video block, W represents the width of the current video block, H represents the height of the current video block, and wherein for the search round with index 0, the multiple time domain positions include {(x+W), (y+H)} denoted as RB0, {(x+W / 2), (y+H / 2)} denoted as Ctr0, {(x+W), (y+H-4)} denoted as R0, and {(x+W-4), (y+H)} denoted as B0.

[0153] In some embodiments, the plurality of search rounds includes 5 search rounds, and 20 time domain positions are examined during the 5 search rounds, the 20 time domain positions including: {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H-4))}, {(x+W-4), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+W / 2+2*W)} ,(y+H / 2+2*H)}, {(x+W+2*W),(y+H / 2)}, {(x+W / 2),(y+H+2*H)}, {(x+ W+3*W),(y+H+3*H)},{(x+W / 2+3*W),(y+H / 2+3*H)},{(x+W+3*W),(y+ H / 2)}, {(x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}. Patterns for time domain BV candidates can be found in Figure 26 Shown in.

[0154] In some embodiments, for a search round with index i, a first time-domain BV candidate is determined based on a priority order of RBi over Ctri, and a second time-domain BV candidate is determined based on a priority order of Ri over Bi, and at least one time-domain BV candidate includes at most two time-domain BV candidates.

[0155] In some embodiments, for a search round with index i, a first time-domain BV candidate is determined based on a priority order of RBi over Ctri, Ctri over Ri, and Ri over Bi, and the at least one time-domain BV candidate includes at most four time-domain BV candidates.

[0156] In some embodiments, at least one pattern of time-domain BV candidates is used. For example, the at least one pattern may be Figure 25 The first pattern shown, or Figure 26 The second pattern is shown. Alternatively, any other pattern of time-domain BV candidates may be used.

[0157] In some embodiments, the at least one time-domain BV candidate includes a first time-domain BV candidate determined in a first manner and a second time-domain BV candidate determined in a second manner. For example, all of the above time-domain BV candidates can be combined in any manner.

[0158] In some embodiments, the number of temporal BV candidates for the current video block is less than or equal to a threshold number.

[0159] In some embodiments, the number of temporal BV candidates after the full deduplication process is less than or equal to a threshold number.

[0160] In some embodiments, the threshold number is 5 or 4.

[0161] In some embodiments, the threshold number is based on the codec mode of the current video block.

[0162] In some embodiments, the codec mode includes at least one of IBC-TM AMVP mode or IBC-TM Merge mode and the threshold number is 1 or 2, and / or wherein the codec mode includes another IBC mode and the threshold number is 4 or 5.

[0163] In some embodiments, the method 2700 further includes: performing at least one of a redundancy check or a deduplication process on the at least one time-domain BV candidate.

[0164] In some embodiments, a full deduplication process is performed on multiple time domain BV candidates, and if the difference between the first motion information of the first time domain BV candidate and the second motion information of the second time domain BV candidate is less than or equal to a threshold, at least one of the first time domain BV candidate or the second time domain BV candidate is excluded from the time domain BV candidate list.

[0165] In some embodiments, the deduplication process includes a partial deduplication process.

[0166] In some embodiments, the method 2700 further includes: adding a plurality of temporal BV candidates to the BV candidate list of the current video block.

[0167] In some embodiments, multiple temporal BV candidates are added to the BV candidate list before history-based motion vector prediction (HMVP) candidates.

[0168] In some embodiments, a portion of the plurality of temporal BV candidates are added to the BV candidate list before a history-based motion vector prediction (HMVP) candidate, and the remaining portion of the plurality of temporal BV candidates are added to the BV candidate list after the HMVP candidate.

[0169] In some embodiments, multiple temporal BV candidates are added in the BV candidate list after the history-based motion vector prediction (HMVP) candidates.

[0170] In some embodiments, at least one temporal BV prediction or at least one temporal BV candidate for the current video block is determined based on a set of co-located pictures for the current video block.

[0171] In some embodiments, the number of co-located pictures in a group is greater than or equal to a first value. For example, the first value may be 1.

[0172] In some embodiments, the indication of a group of co-located pictures is included at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0173] In some embodiments, an indication of a group of co-located pictures is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0174] In some embodiments, a group of co-located pictures is selected from a plurality of co-located pictures based on at least one of: a plurality of picture count distances (POCs) of the plurality of co-located pictures relative to a current picture including the current video block, a plurality of quantization parameter (QP) differences of the plurality of co-located pictures relative to the current picture, or a plurality of QPs of the plurality of co-located pictures.

[0175] In some embodiments, a set of co-located pictures includes the first N co-located pictures with the smallest POC distance, N being a positive integer.

[0176] In some embodiments, a set of co-located pictures includes the first N co-located pictures with the smallest QP difference, N being a positive integer.

[0177] In some embodiments, a set of co-located pictures includes the first N co-located pictures with the smallest QP, where N is a positive integer.

[0178] In some embodiments, the indication in the bitstream indicates at least one of: whether temporal BV prediction (TBVP) is used for the conversion, or whether temporal motion vector prediction (TMVP) is used for the conversion.

[0179] In some embodiments, an indication in the bitstream indicates whether temporal BV prediction (TBVP) is used for the conversion, and another indication in the bitstream indicates whether temporal motion vector prediction (TMVP) is used for the conversion.

[0180] In some embodiments, the indication indicating whether temporal BV prediction (TBVP) is used for the conversion is included at at least one of: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.

[0181] In some embodiments, the indication is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0182] In some embodiments, a process is applied to determine the BV candidate list, the process including at least one of a re-ranking process or a refinement process. For example, the re-ranking / refinement process may be performed when deriving the BV candidate list.

[0183] In some embodiments, the process is based on the template matching cost of the BV candidate.

[0184] In some embodiments, determining the BV candidate list includes: determining a group of candidates, the group of candidates including at least one of the following: a first number of adjacent spatial candidates, a second number of temporal candidates, a third number of history-based motion vector prediction (HMVP) candidates, a fourth number of pair-wise average candidates, or a fifth number of predefined BV candidates; updating the group of candidates by performing a full deduplication process on the group of candidates to remove duplicate candidates; reordering the updated group of candidates; and determining the BV candidate list based on the reordering of the updated group of candidates.

[0185] In some embodiments, the BV candidate list includes the top N candidates with the lowest costs in the updated set of candidates, where N is a positive integer.

[0186] In some embodiments, N is 6, the first number is 5, the second number is 10, the third number is 25, the fourth number is 1, or the fifth number is 6.

[0187] In some embodiments, the number of candidates in the updated set of candidates is less than or equal to a threshold number.

[0188] In some embodiments, the threshold number is 20.

[0189] In some embodiments, the first number of adjacent spatial candidates includes at least one of the following: a spatial BV candidate to the left of the current video block, a spatial BV candidate above the current video block, a spatial BV candidate to the upper right of the current video block, a spatial BV candidate to the lower left of the current video block, or a spatial BV candidate to the upper left of the current video block.

[0190] In some embodiments, the third number of HMVP candidates or the size of the HMVP table is 25.

[0191] In some embodiments, the pair-wise average candidate is determined by averaging at least one predefined candidate pair in the motion candidate list.

[0192] In some embodiments, at least one predefined candidate pair includes {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers 0, 1, 2 and 3 represent the indexes of the motion candidates in the motion candidate list.

[0193] In some embodiments, the predefined BV candidates are located in the IBC reference region.

[0194] In some embodiments, BV candidate type-based Merge candidate adaptive reordering (ARMC) is applied to reorder BV candidates having at least one candidate type based on at least one criterion.

[0195] In some embodiments, a first number of candidates having the lowest costs of a first candidate type are selected from a second number of reordered candidates having the first candidate type, the first number of candidates being added to the BV candidate list.

[0196] In some embodiments, the first number is based on at least one of the first candidate type or a codec mode of the current video block.

[0197] In some embodiments, the first candidate type includes adjacent spatial domain BV candidates, the first number is 4, and the second number is 5.

[0198] In some embodiments, the first candidate type includes time-domain BV candidates, the first number is 4, and the second number is 10.

[0199] In some embodiments, the first candidate type includes history-based motion vector prediction (HMVP) BV candidates, the first number is 10, and the second number is 25.

[0200] In some embodiments, the first candidate type includes pair-wise average BV candidates, the first number is 1, and the second number is 6.

[0201] In some embodiments, the first candidate type comprises types of predefined BV candidates, the first number is 1, and the second number is 6.

[0202] In some embodiments, BV candidates of multiple BV candidate types are reordered together.

[0203] In some embodiments, a first number of candidates having lowest costs are selected from a second number of reordered candidates having at least one BV candidate type of the plurality of BV candidate types, the first number of candidates being added to the BV candidate list.

[0204] In some embodiments, the plurality of candidate types include adjacent spatial candidate types, temporal candidate types, history-based motion vector prediction (HMVP) candidate types, pairwise average candidate types, and predefined BV candidate types, the first number is 6, and the second number is 20.

[0205] In some embodiments, BV candidates of at least one candidate type are reordered based on Merge Candidate Adaptive Reordering (ARMC) based on the BV candidate type.

[0206] In some embodiments, the first number of candidates is determined by: selecting a third number of HMVP candidates from the reordered candidates having the HMVP candidate type; reordering the third number of HMVP candidates together with at least one of the following: adjacent spatial candidates, time domain candidates, pairwise average candidates, or predefined BV candidates; and selecting the first number of candidates based on the reordered candidates.

[0207] In some embodiments, the first number of candidates is determined by: selecting a fourth number of time domain candidates from the reordered candidates having a time domain candidate type; reordering the fourth number of time domain candidates together with at least one of the following: adjacent spatial domain candidates, HMVP candidates, pairwise average candidates, or predefined BV candidates; and selecting the first number of candidates based on the reordered candidates.

[0208] In some embodiments, if the candidates for the current video block are reordered more than once, the reordering criteria for the candidates used in the first reordering are reused in the second reordering. For example, the reordering criteria include the template matching cost of the candidates. In other words, if a candidate is reordered more than once, its reordering criteria (e.g., template matching cost) can be reused.

[0209] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, at least one of a temporal BV prediction or a temporal BV candidate for a current video block of the video is determined. A bitstream is generated based on at least one of the temporal BV prediction or the temporal BV candidate.

[0210] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. In this method, at least one of a temporal BV prediction or a temporal BV candidate for a current video block of the video is determined. A bitstream is generated based on the temporal BV prediction or the temporal BV candidate. The bitstream is stored in a non-transitory computer-readable recording medium.

[0211] Figure 28 28. A flow chart of a method 2800 for video processing according to an embodiment of the present disclosure is shown. The method 2800 is implemented during conversion between a current video block of a video and a bitstream of the video.

[0212] At block 2810, a block vector prediction (BVP) is determined for a sub-block of a current video block. The current video block is coded using a sub-block based temporal motion vector prediction (SbTMVP) mode. For example, a BVP may be obtained for a sub-block such as a 4x4 or 8x8 sub-block of the current video block coded using SbTMVP.

[0213] At block 2820, a conversion is performed based on the BVP. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.

[0214] Method 2800 enables determination of the BVP of a sub-block of a block coded using SbTMVP. In this way, codec efficiency and codec effectiveness can be improved.

[0215] In some embodiments, determining the BVP includes: determining a co-located block of the current video block based on the SbTMVP of the current video block; and determining the BVP based on a temporal position in the co-located block. For example, the BVP can be obtained from the temporal position in the co-located block located by the SbTMVP.

[0216] In some embodiments, the indication or syntax element in the bitstream is binarized into at least one of the following: a flag, a fixed-length code, a Euclidean geometry (x) (EG(x)) code, a unary code, a truncated unary code, or a truncated binary code. For example, the indication or syntax element is signed or unsigned.

[0217] In some embodiments, indications or syntax elements in the bitstream are encoded or decoded using at least one context model, or are bypassed.

[0218] In some embodiments, an indication or syntax element is included in the bitstream based on a condition.

[0219] In some embodiments, the condition includes that a functionality associated with the indication or syntax element is applicable.

[0220] In some embodiments, the indication or syntax element is at at least one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

[0221] In some embodiments, the indication or syntax element is in a codec structure, which codec structure includes at least one of the following: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header or slice group header.

[0222] In some embodiments, the current video block includes one of the following: a color component, a sub-picture, a slice, a codec tree unit (CTU), a CTU row, a group of CTUs, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, a sub-region within a block, or a region containing more than one sample or pixel.

[0223] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a video bitstream generated by a method performed by an apparatus for video processing. In the method, a base bit rate (BVP) of a sub-block of a current video block of the video is determined. The current video block is encoded and decoded using the SbTMVP mode. A bitstream is generated based on the BVP.

[0224] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. In this method, a base bit value (BVP) of a sub-block of a current video block of the video is determined. The current video block is encoded and decoded using the SbTMVP mode. A bitstream is generated based on the BVP. The bitstream is stored in a non-transitory computer-readable recording medium.

[0225] In some embodiments, information about whether and / or how to apply method 2700 and / or method 2800 is included in the bitstream.

[0226] In some embodiments, the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0227] In some embodiments, the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0228] In some embodiments, the information is indicated in a region comprising more than one sample or pixel.

[0229] In some embodiments, the region includes one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, slice, sub-picture.

[0230] In some embodiments, the information is based on encoded information.

[0231] In some embodiments, the coded information includes at least one of: a codec mode, a block size, a color format, a single-tree partition or a dual-tree partition, a color component, a slice type, or a picture type.

[0232] It should be understood that the method 2700 and / or the method 2800 can be applied individually or in any combination. Using the method 2700 and / or the method 2800, the coding effectiveness and / or coding efficiency can be improved.

[0233] Implementations of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0234] Item 1. A method for video processing, comprising: determining, for conversion between a current video block of a video and a bitstream of the video, at least one of a time-domain block vector (BV) prediction or a time-domain BV candidate for the current video block; and performing the conversion based on the time-domain BV prediction or the at least one of the time-domain BV candidates.

[0235] Item 2. The method according to Item 1, wherein the temporal BV prediction is introduced in at least one of the following: conventional intra block copy (IBC) Merge prediction, conventional IBC advanced motion vector prediction (AMVP) prediction, IBC template matching (IBC-TM) Merge prediction, IBC-TM AMVP prediction, reconstruction-reordering IBC (RR-IBC) Merge prediction, RR-IBC AMVP prediction, IBC Merge mode with block vector difference (IBC-MBVD) prediction, string copy vector prediction, or another BV prediction.

[0236] Clause 3. The method of clause 1, wherein the time-domain BV candidate is included in a BV candidate list.

[0237] Item 4. The method according to Item 3, wherein the BV candidate list includes at least one of the following: a conventional intra block copy (IBC) Merge candidate list, a conventional IBC advanced motion vector prediction (AMVP) candidate list, an IBC template matching (IBC-TM) Merge candidate list, an IBC-TM AMVP candidate list, a reconstruction-reordering IBC (RR-IBC) Merge candidate list, an RR-IBC AMVP candidate list, an IBC Merge mode with block vector difference (IBC-MBVD) basic candidate list, or another BV candidate list.

[0238] Item 5. A method according to any one of Items 1 to 4, wherein determining at least one of the time domain BV prediction or the time domain BV candidate includes: determining whether a set of conditions are satisfied, the set of conditions including: a first condition, the first condition being that a motion grid of a co-located block covering the time domain position of the current video block is available, a second condition, the second condition being that the motion grid has BV information, and a third condition, the third condition being that the BV associated with the motion grid is valid for the current video block; and if it is determined that the set of conditions is satisfied, determining at least one of the time domain BV prediction or the time domain BV candidate based on the time domain position.

[0239] Clause 6. The method of clause 5, wherein if at least one condition in the set of conditions is not satisfied, the temporal position is not used to determine at least one of the temporal BV prediction or the temporal BV candidate.

[0240] Item 7. A method according to any one of Items 1 to 6, wherein determining at least one of the time domain BV prediction or the time domain BV candidate includes: if it is determined that the motion grid of the co-located block covering the time domain position of the current video block is outside the codec tree unit (CTU) row of the current video block, performing a limiting operation on the time domain position to obtain a limited time domain position within the CTU row; and determining the time domain BV prediction or at least one of the time domain BV candidates based on the limited time domain position.

[0241] Item 8. A method according to any one of Items 1 to 6, wherein if the motion grid of the co-located block covering the temporal position of the current video block is outside the codec tree unit (CTU) row of the current video block, then the temporal position is not used to determine the temporal BV prediction or at least one of the temporal BV candidates.

[0242] Clause 9. A method according to any one of clauses 5 to 8, wherein the motion grid comprises a 4x4 grid.

[0243] Item 10. A method according to any one of Items 1 to 9, wherein determining the time domain BV prediction or at least one of the time domain BV candidates includes: determining a time domain position from multiple positions in a co-located picture of the current video block; and determining the time domain BV prediction or at least one of the time domain BV candidates based on the time domain position.

[0244] Item 11. The method of Item 10, wherein the plurality of positions comprises a first position at a lower right position of a co-located block of the current video block in the co-located picture, and a second position at a center position of the co-located block.

[0245] Item 12. A method according to Item 11, wherein determining the time domain position includes: determining whether BV is available in the first position; if it is determined that BV is not obtained in the first position, determining whether BV is available in the second position; and if it is determined that BV is obtained in the second position, determining the second position as the time domain position.

[0246] Item 13. A method according to Item 11, wherein determining the time domain position includes: determining whether BV is available in the second position; if it is determined that BV is not obtained in the second position, determining whether BV is available in the first position; and if it is determined that BV is obtained in the first position, determining the first position as the time domain position.

[0247] Item 14. The method according to Item 11, wherein determining the time domain position comprises: determining the time domain position based on a priority order of the first position and the second position.

[0248] Item 15. A method according to Item 14, wherein the priority order includes an order in which the first position takes precedence over the second position, and wherein determining the time domain position includes: determining whether at least one of the following conditions is met: a first condition, the first condition is that the codec unit (CU) at the first position is not available, a second condition, the second condition is that the CU at the first position has no BV information, a third condition, the third condition is that the CU at the first position is outside the codec tree unit (CTU) row of the current video block, or a fourth condition, the fourth condition is that the BV of the CU at the first position is invalid for the current video block; if it is determined that at least one condition is met, determining the second position as the time domain position; and if it is determined that no condition is met, determining the first position as the time domain position.

[0249] Item 16. A method according to Item 14, wherein the priority order includes an order in which the second position takes precedence over the first position, and wherein determining the time domain position includes: judging whether at least one of the following conditions is satisfied: a first condition, the first condition is that the codec unit (CU) at the second position is not available, a second condition, the second condition is that the CU at the second position has no BV information, a third condition, the third condition is that the CU at the second position is outside the codec tree unit (CTU) row of the current video block, or, a fourth condition, the fourth condition is that the BV of the CU at the second position is invalid for the current video block; if it is determined that at least one condition is satisfied, determining the first position as the time domain position; and if it is determined that no condition is satisfied, determining the second position as the time domain position.

[0250] Item 17. The method of any one of Items 1 to 16, wherein a plurality of BV candidates for the current video block are determined based on a plurality of positions in a co-located block of the current video block.

[0251] Item 18. A method according to Item 17, wherein the multiple positions include a first position at the lower right of the co-located block of the current block in the co-located picture, and a second position at the center position of the co-located block, and the multiple BV candidates are determined based on the multiple positions and the order of the multiple positions.

[0252] Item 19. The method according to Item 17, wherein the order comprises one of: a first order in which the first position precedes the second position, or a second order in which the first position follows the second position.

[0253] Item 20. The method of any one of Items 10 to 19, wherein a width and a height of the co-located block in the co-located picture are the same as a width and a height of the current video block in the current picture.

[0254] Item 21. The method of Item 20, wherein the position of the co-located block in the co-located picture is the same as the position of the current video block in the current picture.

[0255] Item 22. The method of Item 20, wherein the position of the co-located block in the co-located picture is determined based on a motion displacement and a position of the current video block in the current picture.

[0256] Item 23. The method of Item 22, wherein the motion displacement comprises a motion vector of a spatial neighbor of the current video block.

[0257] Item 24. A method according to Item 23, wherein the spatial neighbor includes a spatial neighbor among a plurality of spatial neighbors, the plurality of spatial neighbors including: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor to the upper right of the current video block, a fourth spatial neighbor to the lower left of the current video block, and a fifth spatial neighbor to the upper left of the current video block.

[0258] Item 25. A method according to Item 24, wherein determining the motion displacement includes: determining at least one valid motion vector of at least one spatial neighbor of the current video block as at least one motion displacement, and the at least one motion displacement is determined according to a predetermined priority order of multiple spatial neighbors.

[0259] Item 26. The method of Item 25, wherein the at least one valid motion vector comprises a number of valid motion vectors, the number being one of: 1, 2, 3, 4, or 5.

[0260] Item 27. A method according to Item 25 or Item 26, wherein the predetermined priority order includes one of the following: a first priority order of the first airspace neighbor, the second airspace neighbor, the third airspace neighbor, the fourth airspace neighbor and the fifth airspace neighbor, a second priority order of the second airspace neighbor, the first airspace neighbor, the third airspace neighbor, the fourth airspace neighbor and the fifth airspace neighbor, and a third priority order of the fourth airspace neighbor, the first airspace neighbor, the third airspace neighbor, the second airspace neighbor and the fifth airspace neighbor.

[0261] Item 28. A method according to Item 23 or Item 24, wherein if a candidate motion vector of a candidate spatial neighbor uses the co-located picture as a reference picture of the candidate spatial neighbor, the candidate motion vector is determined to be the motion displacement.

[0262] Item 29. A method according to Item 23 or Item 24, wherein a candidate motion vector without a candidate spatial neighbor uses the co-located picture as a reference picture for the candidate spatial neighbor, the motion displacement includes a zero vector, or the candidate spatial neighbor has no motion displacement.

[0263] Item 30. A method according to Item 23 or Item 24, wherein a candidate motion vector without a candidate spatial neighbor uses the co-located picture as a reference picture of the candidate spatial neighbor, another motion vector of one of the first reference picture list or the second reference picture list is scaled to point to the co-located picture, and the scaled another motion vector is determined as the motion displacement.

[0264] Item 31. A method according to any one of Items 1 to 30, wherein determining the BV prediction or at least one of the BV candidates includes: determining a set of template matching costs for a set of motion displacements associated with the current video block; determining at least one motion displacement from the set of motion displacements based on an order of the set of template matching costs; and determining the BV prediction or at least one of the BV candidates based on the at least one motion displacement.

[0265] Item 32. The method of Item 31, wherein the number of the at least one motion displacement comprises one of: 1, 2, 3, 4 or 5.

[0266] Item 33. A method according to any one of Items 1 to 32, wherein the temporal BV candidates include at least one temporal BV candidate selected from the following items: a candidate determined based on a first position of a co-located block of the current video block in a co-located picture, or a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, and a group of candidates determined based on a set of displaced first positions or a set of displaced second positions, wherein the set of displaced first positions is displaced from the first position based on a set of motion displacements associated with a set of spatial neighbors of the current video block, and the set of displaced second positions is displaced from the second position based on the set of motion displacements.

[0267] Item 34. A method according to Item 33, wherein the set of spatial neighbors includes at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor to the upper right of the current video block, a fourth spatial neighbor to the lower left of the current video block, and a fifth spatial neighbor to the upper left of the current video block.

[0268] Item 35. The method of Item 33 or Item 34, wherein the first position comprises a position to the lower right of the co-located block, and the second position comprises a center position of the co-located block.

[0269] Item 36. A method according to any one of Items 33 to 35, wherein the number of the at least one time-domain BV candidates is less than or equal to 6.

[0270] Item 37. The method of Item 33, wherein the set of spatial neighbors comprises a first spatial neighbor to the left of the current video block.

[0271] Item 38. The method of Item 37, wherein the number of the at least one time-domain BV candidates is less than or equal to 2.

[0272] Item 39. A method according to any one of items 33 to 38, wherein the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

[0273] Item 40. A method according to Item 39, wherein the priority order of the displaced first position and the displaced second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

[0274] Item 41. A method according to Item 40, wherein the first displaced position and the second displaced position are based on motion displacement of spatial neighbors, and the spatial neighbors include at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor to the upper right of the current video block, a fourth spatial neighbor to the lower left of the current video block, and a fifth spatial neighbor to the upper left of the current video block.

[0275] Item 42. A method according to any one of Items 1 to 32, wherein the temporal BV candidates include at least one temporal BV candidate selected from the following items: a candidate determined based on a first position of a co-located block of the current video block in a co-located picture, a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, a set of candidates determined based on a set of displaced first positions, the set of displaced first positions being displaced from the first positions based on a set of motion displacements associated with a set of spatial neighbors of the current video block, and a set of candidates determined based on a set of displaced second positions, the set of displaced second positions being displaced from the second positions based on the set of motion displacements.

[0276] Item 43. A method according to Item 42, wherein the set of spatial neighbors includes at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor to the upper right of the current video block, a fourth spatial neighbor to the lower left of the current video block, and a fifth spatial neighbor to the upper left of the current video block.

[0277] Item 44. The method of Item 42 or Item 43, wherein the first position comprises a lower right position of the co-located block, and the second position comprises a center position of the co-located block.

[0278] Item 45. The method of any one of Items 42 to 44, wherein the number of the at least one time-domain BV candidates is less than or equal to 12.

[0279] Item 46. The method of Item 42, wherein the set of spatial neighbors comprises a first spatial neighbor to the left of the current video block.

[0280] Item 47. The method of Item 46, wherein the number of the at least one time-domain BV candidates is less than or equal to 4.

[0281] Item 48. A method according to any one of items 42 to 47, wherein the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

[0282] Item 49. A method according to Item 48, wherein the priority order of the displaced first position and the displaced second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

[0283] Item 50. A method according to Item 49, wherein the first displaced position and the second displaced position are based on motion displacement of spatial neighbors, and the spatial neighbors include at least one of the following: a first spatial neighbor to the left of the current video block, a second spatial neighbor above the current video block, a third spatial neighbor to the upper right of the current video block, a fourth spatial neighbor to the lower left of the current video block, and a fifth spatial neighbor to the upper left of the current video block.

[0284] Item 51. A method according to any one of Items 1 to 50, wherein at least one time-domain BV candidate is determined based on a set of time-domain positions.

[0285] Item 52. The method of Item 51, wherein the set of time domain positions is predefined.

[0286] Item 53. The method of Item 51, wherein the set of time domain positions is determined based on codec information.

[0287] Item 54. The method of Item 51, wherein the set of temporal positions is determined based on at least one of: a position of the current video block, a width of the current video block, or a height of the current video block.

[0288] Item 55. The method of any one of Items 51 to 54, wherein at least one distance between the at least one temporal BV candidate and the current video block is based on a width and a height of the current video block.

[0289] Item 56. A method according to any one of Items 1 to 54, wherein at least one time-domain BV candidate in the first mode is determined by multiple search rounds, wherein in a search round in the multiple search rounds, multiple time-domain positions are checked, wherein the multiple time-domain positions include: a position {(x+W+i*W), (y+H+i*H)} denoted as RBi, a position {(x+W / 2+i*W), (y+H / 2+i*H)} denoted as Ctri, a position {(x+W+i*W), (y+H / 2)} denoted as Ri, and a position {(x+W / 2), (y+H+i*H)} denoted as Bi, and wherein (x, y) represents the position of the current video block, W represents the width of the current video block, H represents the height of the current video block, i represents the index of the search round, and i is greater than or equal to 0.

[0290] Item 57. A method according to Item 56, wherein the multiple search rounds include 5 search rounds, and 20 time domain positions are checked during the 5 search rounds, and the 20 time domain positions include: {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H / 2)}, {(x+W / 2), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+W / 2+2*W),(y+H / 2+2*H)},{(x+W+2*W),(y+H / 2)},{(x+W / 2),(y+H+2*H)}, {(x+W+3*W),(y+H+3*H)}, {(x+W / 2+3*W),(y+H / 2+3*H)}, {(x+W+3*W) ,(y+H / 2)}, {(x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+ 4*W), (y+H / 2+4*H)}, {(x+W+4*W), (y+H / 2)}, and {(x+W / 2), (y+H+4*H)}.

[0291] Item 58. A method according to Item 56 or Item 57, wherein for a search round with index i, a first time domain BV candidate is determined based on a priority order of RBi over Ctri, and a second time domain BV candidate is determined based on a priority order of Ri over Bi, and the at least one time domain BV candidate includes at most two time domain BV candidates.

[0292] Item 59. A method according to Item 56 or Item 57, wherein for a search round with index i, a first time domain BV candidate is determined based on a priority order of RBi over Ctri, Ctri over Ri and Ri over Bi, and the at least one time domain BV candidate includes at most four time domain BV candidates.

[0293] Item 60. A method according to any one of Items 1 to 54, wherein at least one time-domain BV candidate in the second mode is determined by a plurality of search rounds, wherein in a search round of the plurality of search rounds, a plurality of time-domain positions are checked, wherein for the search round with an index i greater than or equal to 1, the plurality of time-domain positions include: a position denoted as RBi {(x+W+i*W), (y+H+i*H)}, a position denoted as Ctri {(x+W / 2+i*W), (y+H / 2+i*H)}, a position denoted as Ri {(x+W+i*W), (y+H / 2)}, and a position {(x+W / 2), (y+H+i*H)} denoted as Bi, wherein (x, y) represents the position of the current video block, W represents the width of the current video block, and H represents the height of the current video block, and wherein for the search round with index 0, the multiple time domain positions include {(x+W), (y+H)} denoted as RB0, {(x+W / 2), (y+H / 2)} denoted as Ctr0, {(x+W), (y+H-4)} denoted as R0, and {(x+W-4), (y+H)} denoted as B0.

[0294] Item 61. A method according to Item 60, wherein the multiple search rounds include 5 search rounds, and 20 time domain positions are checked during the 5 search rounds, and the 20 time domain positions include: {(x+W), (y+H)}, {(x+W / 2), (y+H / 2)}, {(x+W), (y+H-4))}, {(x+W-4), (y+H)}, {(x+W+W), (y+H+H)}, {(x+W / 2+W), (y+H / 2+H)}, {(x+W+W), (y+H / 2)}, {(x+W / 2), (y+H+H)}, {(x+W+2*W), (y+H+2*H)}, {(x+ W / 2+2*W),(y+H / 2+2*H)},{(x+W+2*W),(y+H / 2)},{(x+W / 2),(y+H+2*H)}, {(x+W+3*W),(y+H+3*H)}, {(x+W / 2+3*W),(y+H / 2+3*H)}, {(x+W+3*W) ,(y+H / 2)}, {(x+W / 2),(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+ 4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}.

[0295] Item 62. A method according to Item 60 or Item 61, wherein for a search round with index i, a first time domain BV candidate is determined based on a priority order of RBi over Ctri, and a second time domain BV candidate is determined based on a priority order of Ri over Bi, and the at least one time domain BV candidate includes at most two time domain BV candidates.

[0296] Item 63. A method according to Item 60 or Item 61, wherein for a search round with index i, a first time domain BV candidate is determined based on a priority order of RBi over Ctri, Ctri over Ri and Ri over Bi, and the at least one time domain BV candidate includes at most four time domain BV candidates.

[0297] Item 64. A method according to any of Items 1 to 63, wherein at least one pattern of time-domain BV candidates is used.

[0298] Item 65. The method of any one of Items 1 to 64, wherein the at least one time-domain BV candidate comprises a first time-domain BV candidate determined in a first manner and a second time-domain BV candidate determined in a second manner.

[0299] Item 66. The method of any one of Items 1 to 65, wherein the number of temporal BV candidates for the current video block is less than or equal to a threshold number.

[0300] Item 67. The method of Item 66, wherein the number of time-domain BV candidates after a full deduplication process is less than or equal to the threshold number.

[0301] Item 68. The method of Item 66 or Item 67, wherein the threshold number is 5 or 4.

[0302] Item 69. The method of Item 66 or Item 67, wherein the threshold number is based on a codec mode of the current video block.

[0303] Item 70. A method according to item 69, wherein the coding mode includes at least one of the following: IBC-TMAMVP mode or IBC-TM Merge mode, and the threshold number is 1 or 2, and / or wherein the coding mode includes another IBC mode, and the threshold number is 4 or 5.

[0304] Item 71. The method according to any one of items 1 to 70 further includes: performing at least one of a redundancy check or a deduplication process on at least one time-domain BV candidate.

[0305] Item 72. A method according to Item 71, wherein a plurality of time domain BV candidates are subjected to a full deduplication process, and if the difference between the first motion information of a first time domain BV candidate and the second motion information of a second time domain BV candidate is less than or equal to a threshold, at least one of the first time domain BV candidate or the second time domain BV candidate is excluded from the time domain BV candidate list.

[0306] Item 73. A method according to Item 71, wherein the deduplication process includes a partial deduplication process.

[0307] Item 74. The method of any one of Items 1 to 73, further comprising: adding a plurality of temporal BV candidates to the BV candidate list of the current video block.

[0308] Item 75. The method of Item 74, wherein the plurality of temporal BV candidates are added to the BV candidate list before history-based motion vector prediction (HMVP) candidates.

[0309] Item 76. A method according to item 74, wherein a portion of the multiple time-domain BV candidates are added to the BV candidate list before the history-based motion vector prediction (HMVP) candidate, and the remaining portion of the multiple time-domain BV candidates are added to the BV candidate list after the HMVP candidate.

[0310] Item 77. The method of Item 74, wherein the plurality of temporal BV candidates are added to the BV candidate list after history-based motion vector prediction (HMVP) candidates.

[0311] Item 78. The method of any one of Items 1 to 77, wherein at least one temporal BV prediction or at least one temporal BV candidate for the current video block is determined based on a set of co-located pictures for the current video block.

[0312] Item 79. The method of Item 78, wherein the number of the set of co-located pictures is greater than or equal to a first value.

[0313] Item 80. The method of Item 78 or Item 79, wherein the indication of the set of co-located pictures is included at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0314] Item 81. A method according to item 80, wherein the indication of the set of co-located pictures is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.

[0315] Item 82. A method according to any one of items 78 to 81, wherein the set of co-located pictures is selected from a plurality of co-located pictures based on at least one of: a plurality of picture count distances (POCs) of the plurality of co-located pictures relative to a current picture including the current video block, a plurality of quantization parameter (QP) differences of the plurality of co-located pictures relative to the current picture, or a plurality of QPs of the plurality of co-located pictures.

[0316] Item 83. The method of Item 82, wherein the set of co-located pictures comprises the first N co-located pictures with the smallest POC distance, N being a positive integer.

[0317] Item 84. The method of Item 82, wherein the set of co-located pictures comprises the first N co-located pictures with the smallest QP differences, N being a positive integer.

[0318] Item 85. The method of Item 82, wherein the set of co-located pictures comprises the first N co-located pictures with the smallest QP, N being a positive integer.

[0319] Item 86. A method according to any one of items 1 to 85, wherein the indication in the bitstream indicates at least one of: whether temporal BV prediction (TBVP) is used for the conversion, or whether temporal motion vector prediction (TMVP) is used for the conversion.

[0320] Item 87. A method according to any one of Items 1 to 85, wherein an indication in the bitstream indicates whether temporal BV prediction (TBVP) is used for the conversion, and another indication in the bitstream indicates whether temporal motion vector prediction (TMVP) is used for the conversion.

[0321] Item 88. A method according to any one of items 1 to 87, wherein an indication of whether time-domain BV prediction (TBVP) is used for the conversion is included at at least one of the following: sequence level, picture group level, picture level, slice level, or slice group level.

[0322] Item 89. A method according to item 88, wherein the indication is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.

[0323] Item 90. The method according to any one of items 1 to 89 further comprises: determining a BV candidate list for the current video block, wherein a processing process is applied for determining the BV candidate list, the processing process comprising at least one of the following: a reordering process or a refinement process.

[0324] Item 91. The method of Item 90, wherein the processing is based on a template matching cost of a BV candidate.

[0325] Item 92. A method according to item 90 or item 91, wherein determining the BV candidate list includes: determining a group of candidates, the group of candidates including at least one of the following: a first number of adjacent spatial candidates, a second number of temporal candidates, a third number of history-based motion vector prediction (HMVP) candidates, a fourth number of pair-wise average candidates, or a fifth number of predefined BV candidates; updating the group of candidates by performing a full deduplication process on the group of candidates to remove duplicate candidates; reordering the updated group of candidates; and determining the BV candidate list based on the reordering of the updated group of candidates.

[0326] Item 93. The method of Item 92, wherein the BV candidate list comprises top N candidates with lowest costs in the updated set of candidates, N being a positive integer.

[0327] Item 94. The method of Item 93, wherein N is 6, the first number is 5, the second number is 10, the third number is 25, the fourth number is 1, or the fifth number is 6.

[0328] Item 95. The method of any one of Items 92 to 94, wherein the number of candidates in the updated set of candidates is less than or equal to a threshold number.

[0329] Item 96. The method of Item 95, wherein the threshold number is 20.

[0330] Item 97. A method according to any one of items 92 to 96, wherein the first number of adjacent spatial candidates includes at least one of the following: a spatial BV candidate to the left of the current video block, a spatial BV candidate above the current video block, a spatial BV candidate to the upper right of the current video block, a spatial BV candidate to the lower left of the current video block, or a spatial BV candidate to the upper left of the current video block.

[0331] Item 98. The method of any one of Items 92 to 97, wherein the third number of HMVP candidates or the size of the HMVP table is 25.

[0332] Item 99. The method of any one of Items 92 to 98, wherein the pair-wise average candidate is determined by averaging at least one predefined candidate pair in the motion candidate list.

[0333] Item 100. A method according to item 99, wherein the at least one predefined candidate pair includes {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers 0, 1, 2 and 3 represent the index of the motion candidates in the motion candidate list.

[0334] Item 101. The method of Item 92, wherein the predefined BV candidate is located in an IBC reference region.

[0335] Item 102. A method according to any one of items 92 to 101, wherein Merge candidate adaptive reordering (ARMC) based on BV candidate type is applied to reorder BV candidates having at least one candidate type based on at least one criterion.

[0336] Item 103. The method of item 102, wherein a first number of candidates having a first candidate type with lowest costs are selected from a second number of reordered candidates having the first candidate type, the first number of candidates being added to the BV candidate list.

[0337] Item 104. The method of Item 103, wherein the first number is based on at least one of: the first candidate type, or a codec mode of the current video block.

[0338] Item 105. The method of item 103 or item 104, wherein the first candidate type comprises adjacent spatial domain BV candidates, the first number is 4, and the second number is 5.

[0339] Item 106. The method of Item 103 or Item 104, wherein the first candidate type comprises time-domain BV candidates, the first number is 4, and the second number is 10.

[0340] Item 107. The method of Item 103 or Item 104, wherein the first candidate type comprises history-based motion vector prediction (HMVP) BV candidates, the first number is 10, and the second number is 25.

[0341] Item 108. The method of Item 103 or Item 104, wherein the first candidate type comprises pair-averaged BV candidates, the first number is 1, and the second number is 6.

[0342] Item 109. The method of Item 103 or Item 104, wherein the first candidate type comprises a type of predefined BV candidates, the first number is 1, and the second number is 6.

[0343] Item 110. The method of any one of Items 1 to 109, wherein BV candidates of a plurality of BV candidate types are reordered together.

[0344] Item 111. The method of item 110, wherein a first number of candidates having lowest costs are selected from a second number of reordered candidates having at least one BV candidate type of the plurality of BV candidate types, the first number of candidates being added to the BV candidate list.

[0345] Item 112. A method according to item 111, wherein the multiple candidate types include adjacent spatial candidate types, temporal candidate types, history-based motion vector prediction (HMVP) candidate types, pairwise average candidate types and predefined BV candidate types, the first number is 6, and the second number is 20.

[0346] Item 113. The method of Item 111 or Item 112, wherein the BV candidates of at least one candidate type are reordered based on BV candidate type-based Merge Candidate Adaptive Reordering (ARMC).

[0347] Item 114. A method according to item 112, wherein the first number of candidates is determined by: selecting a third number of HMVP candidates from the reordered candidates having the HMVP candidate type; reordering the third number of HMVP candidates together with at least one of the following: adjacent spatial candidates, time domain candidates, pairwise average candidates, or predefined BV candidates; and selecting the first number of candidates based on the reordered candidates.

[0348] Item 115. A method according to item 112, wherein the first number of candidates is determined by: selecting a fourth number of time domain candidates from the reordered candidates having the time domain candidate type; reordering the fourth number of time domain candidates together with at least one of the following: adjacent spatial domain candidates, HMVP candidates, pairwise average candidates, or predefined BV candidates; and selecting the first number of candidates based on the reordered candidates.

[0349] Item 116. The method of any one of Items 110 to 115, wherein if the candidates for the current video block are reordered more than once, a reordering criterion for the candidates used in a first reordering is reused in a second reordering.

[0350] Item 117. The method of Item 116, wherein the reordering criteria comprises template matching costs of the candidates.

[0351] Item 118. A method for video processing, comprising: determining, for conversion between a current video block of a video and a bitstream of the video, a block vector prediction (BVP) for a sub-block of the current video block, the current video block being encoded and decoded using a sub-block based temporal motion vector prediction (SbTMVP) mode; and performing the conversion based on the BVP.

[0352] Item 119. The method of Item 118, wherein determining the BVP comprises: determining a co-located block of the current video block based on the SbTMVP of the current video block; and determining the BVP based on a temporal position in the co-located block.

[0353] Item 120. A method according to any one of items 1 to 119, wherein the indication or syntax element in the bitstream is binarized into at least one of the following: a flag, a fixed-length code, a Euclidean geometry (x) (EG(x)) code, a unary code, a truncated unary code, or a truncated binary code.

[0354] Item 121. The method of Item 120, wherein the indication or the syntax element is signed or unsigned.

[0355] Item 122. A method according to any one of items 1 to 119, wherein the indication or syntax element in the bitstream is encoded or decoded using at least one context model, or is bypassed.

[0356] Item 123. A method according to any one of items 120 to 122, wherein the indication or the syntax element is included in the bitstream based on a condition.

[0357] Item 124. A method according to item 123, wherein the condition includes: the functionality associated with the indication or the syntax element is applicable.

[0358] Item 125. A method according to any one of items 122 to 124, wherein the indication or the syntax element is at at least one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

[0359] Item 126. A method according to any one of items 122 to 125, wherein the indication or the syntax element is in a codec structure, and the codec structure includes at least one of the following: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.

[0360] Item 127. A method according to any one of Items 1 to 126, wherein the current video block comprises one of the following: a color component, a sub-picture, a slice, a codec tree unit (CTU), a CTU row, a group of CTUs, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), a block, a subblock of a block, a subregion within a block, or a region containing more than one sample or pixel.

[0361] Item 128. A method according to any of items 1 to 127, wherein information on whether and / or how to apply the method is included in the bitstream.

[0362] Item 129. The method of Item 128, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0363] Item 130. A method according to item 128 or item 129, wherein the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.

[0364] Item 131. A method according to any one of Items 128 to 130, wherein the information is indicated in an area comprising more than one sample or pixel.

[0365] Item 132. The method of Item 131, wherein the region comprises one of: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

[0366] Item 133. A method according to any one of Items 128 to 132, wherein the information is based on encoded information.

[0367] Item 134. The method of Item 133, wherein the coded information comprises at least one of: a codec mode, a block size, a color format, a single tree partition or a dual tree partition, a color component, a slice type, or a picture type.

[0368] Item 135. The method of any one of Items 1 to 134, wherein the converting comprises encoding the current video block into the bitstream.

[0369] Item 136. The method of any one of Items 1 to 134, wherein the converting comprises decoding the current video block from the bitstream.

[0370] Item 137. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 136.

[0371] Item 138. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 136.

[0372] Item 139. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining at least one of a time-domain block vector (BV) prediction or a time-domain BV candidate for a current video block of the video; and generating the bitstream based on the time-domain BV prediction or the at least one of the time-domain BV candidates.

[0373] Item 140. A method for storing a bitstream of a video, comprising: determining at least one of a time-domain block vector (BV) prediction or a time-domain BV candidate for a current video block of the video; generating the bitstream based on the time-domain BV prediction or at least one of the time-domain BV candidates; and storing the bitstream in a non-transitory computer-readable recording medium.

[0374] Item 141. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a block vector prediction (BVP) of a sub-block of a current video block of the video, the current video block being encoded and decoded using a sub-block based temporal motion vector prediction (SbTMVP) mode; and generating the bitstream based on the BVP.

[0375] Item 142. A method for storing a bitstream of a video, comprising: determining a block vector prediction (BVP) for a subblock of a current video block of the video, the current video block being encoded and decoded using a subblock-based temporal motion vector prediction (SbTMVP) mode; generating the bitstream based on the BVP; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0376] Figure 29 A block diagram of a computing device 2900 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2900 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0377] It should be understood that Figure 29 The computing device 2900 shown in FIG. 2 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.

[0378] like Figure 29As shown, computing device 2900 comprises a general computing device 2900. Computing device 2900 may include at least one or more processors or processing units 2910, memory 2920, storage unit 2930, one or more communication units 2940, one or more input devices 2950, and one or more output devices 2960.

[0379] In some embodiments, computing device 2900 can be implemented as any user terminal or server terminal with computing capability. A server terminal can be a server, a large computing device, etc. provided by a service provider. A user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 2900 can support any type of interface to a user (such as a "wearable" circuit device, etc.).

[0380] The processing unit 2910 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 2920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 2900. The processing unit 2910 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0381] The computing device 2900 typically includes various computer storage media. Such media can be any media accessible by the computing device 2900, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 2920 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 2930 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 2900.

[0382] The computing device 2900 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 29 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0383] The communication unit 2940 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 2900 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 2900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0384] Input device 2950 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 2960 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 2940, computing device 2900 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 2900, or, if desired, any device that enables computing device 2900 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).

[0385] In some embodiments, some or all components of the computing device 2900 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0386] In an embodiment of the present disclosure, the computing device 2900 may be used to implement video encoding / decoding. The memory 2920 may include one or more video encoding / decoding modules 2925 having one or more program instructions. These modules are accessible and executable by the processing unit 2910 to perform the functions of the various embodiments described herein.

[0387] In an example embodiment performing video encoding, an input device 2950 may receive video data as input 2970 to be encoded. The video data may be processed, for example, by a video codec module 2925 to generate an encoded bitstream. The encoded bitstream may be provided as output 2980 via an output device 2960.

[0388] In an example embodiment performing video decoding, an input device 2950 may receive an encoded bitstream as input 2970. The encoded bitstream may be processed, for example, by a video codec module 2925 to generate decoded video data. The decoded video data may be provided as output 2980 via an output device 2960.

[0389] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For conversion between a current video block of a video and a bitstream of the video, determining at least one of a temporal block vector (BV) prediction or a temporal BV candidate for the current video block; as well as The converting is performed based on the at least one of the temporal BV prediction or the temporal BV candidate.

2. The method according to claim 1, wherein the temporal BV prediction is introduced in at least one of the following: Conventional Intra Block Copy (IBC) Merge prediction, Conventional IBC Advanced Motion Vector Prediction (AMVP) prediction, IBC Template Matching (IBC-TM) Merge prediction, IBC-TM AMVP forecast, Reconstruction-Reordering IBC (RR-IBC) Merge prediction, RR-IBC AMVP prediction, IBC Merge mode with Block Vector Difference (IBC-MBVD) prediction, the string to copy the vector predictions to, or Additional BV predictions. The method according to claim 1 , wherein the time-domain BV candidate is included in a BV candidate list.

4. The method according to claim 3, wherein the BV candidate list includes at least one of the following: Regular Intra Block Copy (IBC) Merge candidate list, Conventional IBC Advanced Motion Vector Prediction (AMVP) candidate list, IBC Template Matching (IBC-TM) Merge candidate list, IBC-TM AMVP candidate list, Reconstruction-Reordering IBC (RR-IBC) Merge candidate list, RR-IBC AMVP candidate list, IBC Merge Mode with Block Vector Difference (IBC-MBVD) base candidate list, or Additional BV candidate list.

5. The method according to any one of claims 1 to 4, wherein determining at least one of the temporal BV prediction or the temporal BV candidate comprises: Determine whether a set of conditions are met, the set of conditions comprising: A first condition is that a motion grid of a co-located block covering the temporal position of the current video block is available, The second condition is that the motion mesh has BV information, and a third condition, the third condition being that the BV associated with the motion grid is valid for the current video block; and If it is determined that the set of conditions is satisfied, at least one of the temporal BV prediction or the temporal BV candidate is determined based on the temporal position. 6 . The method of claim 5 , wherein if at least one condition in the set of conditions is not satisfied, the temporal position is not used to determine at least one of the temporal BV prediction or the temporal BV candidate.

7. The method according to any one of claims 1 to 6, wherein determining at least one of the temporal BV prediction or the temporal BV candidate comprises: If it is determined that a motion grid of a co-located block covering the temporal position of the current video block is outside a codec tree unit (CTU) row of the current video block, performing a clipping operation on the temporal position to obtain a clipped temporal position inside the CTU row; as well as At least one of the temporal BV prediction or the temporal BV candidate is determined based on the clipped temporal position.

8. The method of any one of claims 1 to 6, wherein if a motion grid of a co-located block covering a temporal position of the current video block is outside a codec tree unit (CTU) row of the current video block, then the temporal position is not used to determine at least one of the temporal BV prediction or the temporal BV candidate.

9. The method of any one of claims 5 to 8, wherein the motion grid comprises a 4x4 grid.

10. The method according to any one of claims 1 to 9, wherein determining at least one of the temporal BV prediction or the temporal BV candidate comprises: determining a temporal position from a plurality of positions in a co-located picture of the current video block; as well as Based on the temporal position, at least one of the temporal BV prediction or the temporal BV candidate is determined. 11 . The method of claim 10 , wherein the plurality of positions comprises a first position at a lower right position of a co-located block of the current video block in the co-located picture, and a second position at a center position of the co-located block.

12. The method of claim 11, wherein determining the time domain position comprises: determining whether a BV is available in the first location; If it is determined that the BV is not obtained in the first location, determining whether the BV is available in the second location; as well as If it is determined that the BV is obtained in the second position, the second position is determined as the time domain position.

13. The method of claim 11 , wherein determining the temporal location comprises: determining whether a BV is available in the second location; If it is determined that the BV is not obtained in the second location, determining whether the BV is available in the first location; as well as If it is determined that the BV is obtained in the first position, the first position is determined as the time domain position.

14. The method of claim 11 , wherein determining the temporal location comprises: The time domain position is determined based on a priority order of the first position and the second position.

15. The method of claim 14, wherein the priority order comprises an order in which the first position takes precedence over the second position, and Wherein determining the time domain position includes: Determine whether at least one of the following conditions is met: a first condition, wherein the first condition is that the codec unit (CU) at the first position is unavailable, The second condition is that the CU at the first position has no BV information, A third condition is that the CU at the first position is outside a codec tree unit (CTU) row of the current video block, or a fourth condition, the fourth condition being that the BV of the CU at the first position is invalid for the current video block; If it is determined that the at least one condition is satisfied, determining the second position as the time domain position; as well as If it is determined that no condition is met, the first position is determined as the time domain position.

16. The method of claim 14, wherein the priority order includes an order in which the second position takes precedence over the first position, and Wherein determining the time domain position includes: Determine whether at least one of the following conditions is met: a first condition, wherein the first condition is that the codec unit (CU) at the second position is unavailable, The second condition is that the CU at the second position has no BV information, a third condition, wherein the third condition is that the CU at the second position is outside a codec tree unit (CTU) row of the current video block, or a fourth condition, the fourth condition being that the BV of the CU at the second position is invalid for the current video block; If it is determined that the at least one condition is satisfied, determining the first position as the time domain position; as well as If it is determined that no condition is met, the second position is determined as the time domain position. 17 . The method of claim 1 , wherein a plurality of BV candidates for the current video block are determined based on a plurality of positions in a co-located block of the current video block.

18. The method according to claim 17, wherein the multiple positions include a first position at the lower right of the co-located block of the current block in the co-located picture, and a second position at the center position of the co-located block, and the multiple BV candidates are determined based on the multiple positions and the order of the multiple positions.

19. The method of claim 17, wherein the sequence comprises one of the following: the first position precedes the second position in a first order, or The first position is in a second order after the second position.

20. The method of any one of claims 10 to 19, wherein a width and a height of the co-located block in the co-located picture are the same as a width and a height of the current video block in the current picture.

21. The method of claim 20, wherein a position of the co-located block in the co-located picture is the same as a position of the current video block in the current picture.

22. The method of claim 20, wherein a position of the co-located block in the co-located picture is determined based on a motion displacement and a position of the current video block in the current picture.

23. The method of claim 22, wherein the motion displacement comprises motion vectors of spatial neighbors of the current video block.

24. The method of claim 23, wherein the spatial neighbor comprises one of a plurality of spatial neighbors, the plurality of spatial neighbors comprising: The first spatial neighbor on the left side of the current video block, A second spatial neighbor above the current video block, The third spatial neighbor to the upper right of the current video block, The fourth spatial neighbor to the lower left of the current video block, and The fifth spatial neighbor above and to the left of the current video block.

25. The method of claim 24, wherein determining the motion displacement comprises: At least one valid motion vector of at least one spatial neighbor of the current video block is determined as at least one motion displacement, wherein the at least one motion displacement is determined according to a predetermined priority order of a plurality of spatial neighbors.

26. The method of claim 25, wherein the at least one valid motion vector comprises a number of valid motion vectors, the number being one of: 1, 2, 3, 4, or 5.

27. The method of claim 25 or claim 26, wherein the predetermined priority order comprises one of the following: a first priority order of the first spatial neighbor, the second spatial neighbor, the third spatial neighbor, the fourth spatial neighbor, and the fifth spatial neighbor; a second priority order of the second spatial neighbor, the first spatial neighbor, the third spatial neighbor, the fourth spatial neighbor, and the fifth spatial neighbor, A third priority order of the fourth spatial neighbor, the first spatial neighbor, the third spatial neighbor, the second spatial neighbor, and the fifth spatial neighbor.

28. The method of claim 23 or claim 24, wherein if a candidate motion vector of a candidate spatial neighbor uses the co-located picture as a reference picture of the candidate spatial neighbor, the candidate motion vector is determined as the motion displacement.

29. The method of claim 23 or claim 24, wherein no candidate motion vector of a candidate spatial neighbor uses the co-located picture as a reference picture for the candidate spatial neighbor, the motion displacement comprises a zero vector, or the candidate spatial neighbor has no motion displacement.

30. A method according to claim 23 or claim 24, wherein a candidate motion vector without a candidate spatial neighbor uses the co-located picture as a reference picture of the candidate spatial neighbor, another motion vector of one of the first reference picture list or the second reference picture list is scaled to point to the co-located picture, and the scaled another motion vector is determined as the motion displacement.

31. The method of any one of claims 1 to 30, wherein determining at least one of the BV prediction or the BV candidate comprises: determining a set of template matching costs for a set of motion displacements associated with the current video block; determining at least one motion displacement from the set of motion displacements based on an order of the set of template matching costs; as well as At least one of the BV prediction or the BV candidate is determined based on the at least one motion displacement.

32. The method of claim 31 , wherein the number of the at least one motion displacement comprises one of: 1, 2, 3, 4, or 5.

33. The method according to any one of claims 1 to 32, wherein the time-domain BV candidate comprises at least one time-domain BV candidate selected from the following: a candidate determined based on a first position of a co-located block of the current video block in a co-located picture, or a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, and A set of candidates is determined based on a set of displaced first positions that are displaced from the first positions based on a set of motion displacements associated with a set of spatial neighbors of the current video block, or a set of displaced second positions that are displaced from the second positions based on the set of motion displacements.

34. The method of claim 33, wherein the set of spatial neighbors comprises at least one of: The first spatial neighbor on the left side of the current video block, A second spatial neighbor above the current video block, The third spatial neighbor to the upper right of the current video block, The fourth spatial neighbor to the lower left of the current video block, and The fifth spatial neighbor above and to the left of the current video block.

35. The method of claim 33 or claim 34, wherein the first position comprises a position to the lower right of the co-located block, and the second position comprises a position at the center of the co-located block.

36. The method according to any one of claims 33 to 35, wherein the number of the at least one time-domain BV candidate is less than or equal to 6.

37. The method of claim 33, wherein the set of spatial neighbors comprises a first spatial neighbor to the left of the current video block. The method of claim 37 , wherein the number of the at least one time-domain BV candidate is less than or equal to 2.

39. The method according to any one of claims 33 to 38, wherein the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

40. The method of claim 39, wherein the priority order of the shifted first position and the shifted second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

41. The method of claim 40, wherein the shifted first position and the shifted second position are based on motion displacements of spatial neighbors, the spatial neighbors comprising at least one of: The first spatial neighbor on the left side of the current video block, A second spatial neighbor above the current video block, The third spatial neighbor to the upper right of the current video block, The fourth spatial neighbor to the lower left of the current video block, and The fifth spatial neighbor above and to the left of the current video block.

42. The method according to any one of claims 1 to 32, wherein the time-domain BV candidates include at least one time-domain BV candidate selected from the following: a candidate determined based on a first position of a co-located block of the current video block in a co-located picture, a candidate determined based on a second position of the co-located block of the current video block in the co-located picture, a set of candidates determined based on a set of displaced first positions that are displaced from the first positions based on a set of motion displacements associated with a set of spatial neighbors of the current video block, and A set of candidates is determined based on a set of shifted second positions that are displaced from the second position based on the set of motion displacements.

43. The method according to claim 42, wherein The set of spatial neighbors includes at least one of the following: The first spatial neighbor on the left side of the current video block, A second spatial neighbor above the current video block, The third spatial neighbor to the upper right of the current video block, The fourth spatial neighbor to the lower left of the current video block, and The fifth spatial neighbor above and to the left of the current video block.

44. The method of claim 42 or claim 43, wherein the first position comprises a lower right position of the co-located block, and the second position comprises a center position of the co-located block.

45. The method according to any one of claims 42 to 44, wherein the number of the at least one time-domain BV candidate is less than or equal to 12.

46. The method of claim 42, wherein the set of spatial neighbors comprises a first spatial neighbor to the left of the current video block. The method of claim 46 , wherein the number of the at least one time-domain BV candidate is less than or equal to 4.

48. The method according to any one of claims 42 to 47, wherein the priority order of the first position and the second position is that the first position takes precedence over the second position, or the second position takes precedence over the first position.

49. The method of claim 48, wherein the priority order of the shifted first position and the shifted second position is the same as the priority order of the first position and the second position, or is opposite to the priority order of the first position and the second position.

50. The method of claim 49, wherein the shifted first position and the shifted second position are based on motion displacements of spatial neighbors, the spatial neighbors comprising at least one of: The first spatial neighbor on the left side of the current video block, A second spatial neighbor above the current video block, The third spatial neighbor to the upper right of the current video block, The fourth spatial neighbor to the lower left of the current video block, and The fifth spatial neighbor above and to the left of the current video block.

51. The method of any one of claims 1 to 50, wherein at least one temporal BV candidate is determined based on a set of temporal locations.

52. The method of claim 51, wherein the set of time domain locations is predefined.

53. The method of claim 51, wherein the set of time domain positions is determined based on codec information.

54. The method of claim 51, wherein the set of temporal positions is determined based on at least one of: a position of the current video block, a width of the current video block, or a height of the current video block.

55. The method of any one of claims 51 to 54, wherein at least one distance between the at least one temporal BV candidate and the current video block is based on a width and a height of the current video block.

56. The method according to any one of claims 1 to 54, wherein at least one temporal BV candidate in the first mode is determined by a plurality of search rounds, wherein in a search round of the plurality of search rounds a plurality of temporal positions are checked, The multiple time domain positions include: Represented as RB i The position {(x+W+i*W),(y+H+i*H)}, denoted as Ctr i The position {(x+W / 2+i*W),(y+H / 2+i*H)}, expressed as R i The position of {(x+W+i*W), (y+H / 2)}, and represented by B i The position of {(x+W / 2),(y+H+i*H)}, and Wherein (x, y) represents the position of the current video block, W represents the width of the current video block, H represents the height of the current video block, and i represents the index of the search round, where i is greater than or equal to 0.

57. The method of claim 56, wherein the plurality of search rounds comprises five search rounds, and 20 time-domain positions are examined during the five search rounds, the 20 time-domain positions comprising: {(x+W),(y+H)}, {(x+W / 2),(y+H / 2)}, {(x+W),(y+H / 2)}, {(x+W / 2),(y+H)}, {(x+W+W),(y+H+H)}, {(x+W / 2+W),(y+H / 2+H )}, {(x+W+W),(y+H / 2)}, {(x+W / 2),(y+H+H)}, {(x+W+2*W),(y+H+2*H)}, {(x+W / 2+2*W),(y+H / 2+2*H)}, {(x+W+2*W),(y+H / 2)}, {(x+W / 2),(y+H+2*H)}, {(x+W+3*W),(y+H+3*H)}, {(x+W / 2+3*W),(y+H / 2+3*H)}, {(x+W+3*W),(y+H / 2)}, {(x+W / 2) ,(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}.

58. The method of claim 56 or claim 57, wherein for a search round with index i, the first time domain BV candidate is based on RB i Takes precedence over Ctr i The priority order of R is determined, and the second time domain BV candidate is based on R i Prioritize B i and the at least one time-domain BV candidate includes at most two time-domain BV candidates.

59. The method of claim 56 or claim 57, wherein for a search round with index i, the first time domain BV candidate is based on RB i Takes precedence over Ctr i 、Ctr i Prioritizes R i And R i Prioritize B i and the at least one time-domain BV candidate includes at most four time-domain BV candidates.

60. The method according to any one of claims 1 to 54, wherein at least one temporal BV candidate in the second mode is determined by a plurality of search rounds, wherein in a search round of the plurality of search rounds, a plurality of temporal positions are checked, Wherein, for the search round with an index i greater than or equal to 1, the multiple time domain positions include: Represented as RB i The position {(x+W+i*W),(y+H+i*H)}, denoted as Ctr i The position {(x+W / 2+i*W),(y+H / 2+i*H)}, expressed as R i The position of {(x+W+i*W), (y+H / 2)}, and represented by B i The position of {(x+W / 2),(y+H+i*H)}, Where (x, y) represents the position of the current video block, W represents the width of the current video block, H represents the height of the current video block, and Wherein, for the search round with index 0, the multiple time domain positions include {(x+W), (y+H)} represented as RB0, {(x+W / 2), (y+H / 2)} represented as Ctr0, {(x+W), (y+H-4)} represented as R0, and {(x+W-4), (y+H)} represented as B0.

61. The method of claim 60, wherein the plurality of search rounds comprises five search rounds, and 20 time-domain positions are examined during the five search rounds, the 20 time-domain positions comprising: {(x+W),(y+H)}, {(x+W / 2),(y+H / 2)}, {(x+W),(y+H-4))}, {(x+W-4),(y+H)}, {(x+W+W),(y+H+H)}, {(x+W / 2+W),(y+H / 2+ H)}, {(x+W+W),(y+H / 2)}, {(x+W / 2),(y+H+H)}, {(x+W+2*W),(y+H+2*H)}, {(x+W / 2+2*W),(y+H / 2+2*H)}, {(x+W+2*W),(y+ H / 2)}, {(x+W / 2),(y+H+2*H)}, {(x+W+3*W),(y+H+3*H)}, {(x+W / 2+3*W),(y+H / 2+3*H)}, {(x+W+3*W),(y+H / 2)}, {(x+W / 2) ,(y+H+3*H)}, {(x+W+4*W),(y+H+4*H)}, {(x+W / 2+4*W),(y+H / 2+4*H)}, {(x+W+4*W),(y+H / 2)}, and {(x+W / 2),(y+H+4*H)}.

62. The method according to claim 60 or claim 61, wherein for the search round with index i, the first time domain BV candidate is based on RB i Takes precedence over Ctr i The priority order of R is determined, and the second time domain BV candidate is based on R i Prioritize B i and the at least one time-domain BV candidate includes at most two time-domain BV candidates.

63. The method of claim 60 or claim 61, wherein for a search round with index i, the first time domain BV candidate is based on RB i Takes precedence over Ctr i 、Ctr i Prioritizes R i And R i Prioritize B i and the at least one time-domain BV candidate includes at most four time-domain BV candidates.

64. A method according to any one of claims 1 to 63, wherein at least one pattern of time-domain BV candidates is used.

65. The method according to any one of claims 1 to 64, wherein the at least one time-domain BV candidate comprises a first time-domain BV candidate determined in a first manner and a second time-domain BV candidate determined in a second manner.

66. The method of any one of claims 1 to 65, wherein a number of temporal BV candidates for the current video block is less than or equal to a threshold number.

67. The method of claim 66, wherein the number of time-domain BV candidates after a full deduplication process is less than or equal to the threshold number.

68. A method according to claim 66 or claim 67, wherein the threshold number is 5 or 4.

69. The method of claim 66 or claim 67, wherein the threshold number is based on a codec mode of the current video block.

70. The method according to claim 69, wherein the codec mode comprises at least one of the following: IBC-TMAMVP mode or IBC-TM Merge mode, and the threshold number is 1 or 2, and / or The coding / decoding mode includes another IBC mode, and the threshold number is 4 or 5.

71. The method of any one of claims 1 to 70, further comprising: At least one of a redundancy check and a deduplication process is performed on at least one time-domain BV candidate.

72. The method according to claim 71, wherein a full deduplication process is performed on multiple time domain BV candidates, and if the difference between the first motion information of the first time domain BV candidate and the second motion information of the second time domain BV candidate is less than or equal to a threshold, at least one of the first time domain BV candidate or the second time domain BV candidate is excluded from the time domain BV candidate list.

73. The method of claim 71, wherein the deduplication process comprises a partial deduplication process.

74. The method of any one of claims 1 to 73, further comprising: Add multiple temporal BV candidates to the BV candidate list of the current video block.

75. The method of claim 74, wherein the plurality of temporal BV candidates are added to the BV candidate list before history-based motion vector prediction (HMVP) candidates.

76. The method of claim 74, wherein a portion of the plurality of temporal BV candidates are added to the BV candidate list before a history-based motion vector prediction (HMVP) candidate, and a remaining portion of the plurality of temporal BV candidates are added to the BV candidate list after the HMVP candidate.

77. The method of claim 74, wherein the plurality of temporal BV candidates are added to the BV candidate list after history-based motion vector prediction (HMVP) candidates.

78. The method of any one of claims 1 to 77, wherein at least one temporal BV prediction or at least one temporal BV candidate for the current video block is determined based on a set of co-located pictures for the current video block.

79. The method of claim 78, wherein the number of the set of co-located pictures is greater than or equal to a first value.

80. The method of claim 78 or claim 79, wherein the indication of the set of co-located pictures is included at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

81. The method of claim 80, wherein the indication of the set of co-located pictures is included in at least one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

82. The method of any one of claims 78 to 81, wherein the set of co-located pictures is selected from a plurality of co-located pictures based on at least one of: a plurality of picture count distances (POCs) of the plurality of co-located pictures relative to a current picture including the current video block, a plurality of quantization parameter (QP) differences of the plurality of co-located pictures relative to the current picture, or a plurality of QPs of the plurality of co-located pictures.

83. The method of claim 82, wherein the set of co-located pictures comprises first N co-located pictures with minimum POC distances, N being a positive integer.

84. The method of claim 82, wherein the set of co-located pictures comprises the first N co-located pictures with the smallest QP difference, N being a positive integer.

85. The method of claim 82, wherein the set of co-located pictures comprises the first N co-located pictures with the smallest QP, N being a positive integer.

86. The method of any one of claims 1 to 85, wherein the indication in the bitstream indicates at least one of: Whether to use time-domain BV prediction (TBVP) for the conversion, or Whether to use temporal motion vector prediction (TMVP) for the conversion.

87. A method according to any one of claims 1 to 85, wherein an indication in the bitstream indicates whether time-domain BV prediction (TBVP) is used for the conversion, and Another indication in the bitstream indicates whether temporal motion vector prediction (TMVP) is used for the conversion.

88. The method of any one of claims 1 to 87, wherein an indication of whether temporal BV prediction (TBVP) is used for the conversion is included at at least one of: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.

89. The method of claim 88, wherein the indication is included in at least one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

90. The method according to any one of claims 1 to 89, further comprising: A BV candidate list for the current video block is determined, wherein a processing process is applied for determining the BV candidate list, the processing process comprising at least one of: a reordering process or a refinement process.

91. The method of claim 90, wherein the processing is based on template matching costs of BV candidates.

92. The method of claim 90 or claim 91, wherein determining the BV candidate list comprises: determining a set of candidates, the set of candidates comprising at least one of: a first number of adjacent spatial candidates, a second number of temporal candidates, a third number of history-based motion vector prediction (HMVP) candidates, a fourth number of pairwise average candidates, or a fifth number of predefined BV candidates; updating the set of candidates by performing a full deduplication process on the set of candidates to remove duplicate candidates; reordering the updated set of candidates; as well as The BV candidate list is determined based on a reordering of the updated set of candidates.

93. The method of claim 92, wherein the BV candidate list comprises top N candidates with lowest costs in the updated set of candidates, N being a positive integer.

94. The method of claim 93, wherein N is 6, the first number is 5, the second number is 10, the third number is 25, the fourth number is 1, or the fifth number is 6.

95. The method of any one of claims 92 to 94, wherein the number of candidates in the updated set of candidates is less than or equal to a threshold number.

96. The method of claim 95, wherein the threshold number is 20.

97. The method according to any one of claims 92 to 96, wherein the first number of adjacent spatial domain candidates comprises at least one of the following: The spatial domain BV candidate to the left of the current video block, The spatial BV candidate above the current video block, The spatial domain BV candidate in the upper right corner of the current video block, The spatial domain BV candidate at the bottom left of the current video block, or The spatial domain BV candidate at the upper left of the current video block.

98. The method according to any one of claims 92 to 97, wherein the third number of HMVP candidates or the size of the HMVP table is 25.

99. The method of any one of claims 92 to 98, wherein the pair-wise average candidate is determined by averaging at least one predefined candidate pair in a motion candidate list.

100. The method of claim 99, wherein the at least one predefined candidate pair comprises {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, wherein the numbers 0, 1, 2, and 3 represent the indices of the motion candidates in the motion candidate list.

101. The method of claim 92, wherein the predefined BV candidate is located in an IBC reference region.

102. The method according to any one of claims 92 to 101, wherein Merge Candidate Adaptive Reordering (ARMC) based on BV Candidate Type is applied to reorder BV candidates having at least one candidate type based on at least one criterion.

103. The method of claim 102, wherein a first number of candidates having a first candidate type with lowest costs are selected from a second number of reordered candidates having the first candidate type, the first number of candidates being added to the BV candidate list.

104. The method of claim 103, wherein the first number is based on at least one of: the first candidate type, or a codec mode of the current video block.

105. The method of claim 103 or claim 104, wherein the first candidate type comprises adjacent spatial domain BV candidates, the first number is 4, and the second number is 5.

106. The method of claim 103 or claim 104, wherein the first candidate type comprises time-domain BV candidates, the first number is 4, and the second number is 10.

107. The method of claim 103 or claim 104, wherein the first candidate type comprises history-based motion vector prediction (HMVP) BV candidates, the first number is 10, and the second number is 25.

108. The method of claim 103 or claim 104, wherein the first candidate type comprises pair-averaged BV candidates, the first number is 1, and the second number is 6.

109. The method of claim 103 or claim 104, wherein the first candidate type comprises a type of predefined BV candidates, the first number is 1, and the second number is 6.

110. The method of any one of claims 1 to 109, wherein BV candidates of a plurality of BV candidate types are reordered together.

111. The method of claim 110, wherein a first number of candidates having lowest costs are selected from a second number of reordered candidates having at least one BV candidate type of the plurality of BV candidate types, the first number of candidates being added to the BV candidate list.

112. A method according to claim 111, wherein the multiple candidate types include adjacent spatial candidate types, temporal candidate types, history-based motion vector prediction (HMVP) candidate types, pairwise average candidate types and predefined BV candidate types, the first number is 6, and the second number is 20.

113. The method of claim 111 or claim 112, wherein BV candidates of at least one candidate type are reordered based on BV candidate type-based Merge Candidate Adaptive Reordering (ARMC).

114. The method of claim 112, wherein the first number of candidates is determined by: selecting a third number of HMVP candidates from the reordered candidates having the HMVP candidate type; reordering the third number of HMVP candidates together with at least one of: adjacent spatial candidates, temporal candidates, pairwise average candidates, or predefined BV candidates; and The first number of candidates is selected based on the reordered candidates.

115. The method of claim 112, wherein the first number of candidates is determined by: selecting a fourth number of time-domain candidates from the reordered candidates having the time-domain candidate type; reordering the fourth number of temporal candidates together with at least one of: adjacent spatial candidates, HMVP candidates, pairwise average candidates, or predefined BV candidates; and The first number of candidates is selected based on the reordered candidates.

116. The method according to any one of claims 110 to 115, wherein if the candidates of the current video block are reordered more than once, the reordering criteria of the candidates used in the first reordering are reused in the second reordering.

117. The method of claim 116, wherein the reordering criteria comprises template matching costs of the candidates.

118. A method for video processing, comprising: determining, for conversion between a current video block of a video and a bitstream of the video, a block vector prediction (BVP) for a sub-block of the current video block, the current video block being coded or decoded using a sub-block based temporal motion vector prediction (SbTMVP) mode; as well as The conversion is performed based on the BVP.

119. The method of claim 118, wherein determining the BVP comprises: Determining a co-located block of the current video block based on the SbTMVP of the current video block; as well as The BVP is determined based on a temporal position in the co-located block.

120. A method according to any one of claims 1 to 119, wherein the indication or syntax element in the bitstream is binarized into at least one of the following: a flag, a fixed-length code, a Euclidean geometry (x) (EG(x)) code, a unary code, a truncated unary code, or a truncated binary code.

121. The method of claim 120, wherein the indication or the syntax element is signed or unsigned.

122. The method according to any one of claims 1 to 119, wherein indications or syntax elements in the bitstream are encoded or decoded using at least one context model, or are bypassed.

123. The method according to any one of claims 120 to 122, wherein the indication or the syntax element is included in the bitstream based on a condition.

124. The method of claim 123, wherein the conditions include: The functionality associated with the indication or the syntax element is applicable.

125. The method of any one of claims 122 to 124, wherein the indication or the syntax element is at at least one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

126. The method of any one of claims 122 to 125, wherein the indication or the syntax element is in a codec structure, the codec structure comprising at least one of: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

127. The method of any one of claims 1 to 126, wherein the current video block comprises one of: Color component, sub-images, strips, piece, Codec Tree Unit (CTU), CTU line, CTU group, Codec Unit (CU), Prediction Unit (PU), Transformation Unit (TU), Codec Tree Block (CTB), Codec Block (CB), Prediction Block (PB), Transform Block (TB), piece, Sub-blocks of a block, a subregion within a block, or An area containing more than one sample or pixel.

128. The method according to any one of claims 1 to 127, wherein information on whether and / or how to apply the method is included in the bitstream.

129. The method of claim 128, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.

130. The method of claim 128 or claim 129, wherein the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

131. A method according to any one of claims 128 to 130, wherein the information is indicated in an area comprising more than one sample or pixel.

132. The method of claim 131 , wherein the region comprises one of a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

133. A method according to any one of claims 128 to 132, wherein the information is based on encoded information.

134. The method of claim 133, wherein the coded information comprises at least one of: a codec mode, a block size, a color format, a single tree partition or a dual tree partition, a color component, a slice type, or a picture type.

135. The method of any one of claims 1 to 134, wherein the converting comprises encoding the current video block into the bitstream.

136. The method of any one of claims 1 to 134, wherein the converting comprises decoding the current video block from the bitstream.

137. An apparatus for video processing comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 136.

138. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to perform the method according to any one of claims 1 to 136.

139. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining at least one of a temporal block vector (BV) prediction or a temporal BV candidate for a current video block of the video; as well as The bitstream is generated based on the time-domain BV prediction or the at least one of the time-domain BV candidates.

140. A method for storing a bitstream of a video, comprising: determining at least one of a temporal block vector (BV) prediction or a temporal BV candidate for a current video block of the video; Generate the bitstream based on at least one of the temporal BV prediction or the temporal BV candidate; as well as The bitstream is stored in a non-transitory computer-readable recording medium.

141. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a block vector prediction (BVP) for a sub-block of a current video block of the video, the current video block being coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; as well as The bitstream is generated based on the BVP.

142. A method for storing a bitstream of a video, comprising: determining a block vector prediction (BVP) for a sub-block of a current video block of the video, the current video block being coded using a sub-block based temporal motion vector prediction (SbTMVP) mode; generating the bitstream based on the BVP; as well as The bitstream is stored in a non-transitory computer-readable recording medium.