Method and device for video processing and medium
By employing prediction methods guided by block vectors and motion vectors, the problem of insufficient encoding and decoding efficiency in existing video encoding and decoding technologies is solved, achieving more efficient video processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-24
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved, especially in block vector prediction and motion vector prediction.
The block vector prediction method guided by block vectors and the motion vector prediction method guided by motion vectors are used to perform vector prediction by determining the guide vector and reference vector associated with the current video block, so as to improve the encoding and decoding efficiency.
It improves the efficiency and effectiveness of video encoding and decoding, and enhances the performance of video processing.
Smart Images

Figure CN121925852A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to block vector guided block vector prediction and motion vector guided motion vector prediction. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between a current video block and the bitstream of the video, determining a first guiding vector associated with the current video block, the first guiding vector including a first guiding block vector (BV) or a first guiding motion vector (MV); determining a first reference vector for a first reference block located based on the first guiding vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction for the current video block based at least on the first guiding vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction; and performing a conversion based on the vector prediction. The method according to the first aspect of this disclosure enables block vector prediction guided by block vectors and / or motion vector prediction guided by motion vectors, thereby improving encoding / decoding efficiency and / or encoding / decoding effectiveness.
[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a first guide vector associated with a current video block of the video, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV); determining a first reference vector for a first reference block located based on the first guide vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction of the current video block based at least on the first guide vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction; and generating a bitstream based on the vector prediction.
[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a first guiding vector associated with a current video block of the video, the first guiding vector including a first guiding block vector (BV) or a first guiding motion vector (MV); determining a first reference vector for a first reference block located based on the first guiding vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction of the current video block based at least on the first guiding vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction; generating a bitstream based on the vector prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other objects, features and advantages of the exemplary embodiments of this disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0011] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 The spatial proximity locations used in IBC vector prediction are shown; Figure 5The current CTU processing order and its available reference points in the current CTU and the left CTU are shown; Figure 6 This shows the airspace proximity locations used in the construction of the IBC Merge / AMVP list; Figure 7 The filling candidates for replacing the zero vector in the IBC list are shown; Figure 8 The IBC reference area is shown, depending on the current CU location; Figure 9 This shows the reference region for IBC when encoding and decoding CTU(m,n). Dark gray blocks represent the current CTU; light gray blocks represent the reference region; and white blocks represent invalid reference regions. Figure 10A A schematic diagram of BV adjustment for horizontal flipping is shown; Figure 10B A schematic diagram of BV adjustment for vertical flipping is shown; Figure 11 The intra-frame template matching search area used is shown; Figure 12 A schematic diagram illustrating the use of IntraTMP block vectors for IBC blocks is shown; Figure 13 Examples are shown of candidate lists of IBC block vectors that exist for (a) only IBC block vectors and (b) both IBC block vectors and IntraTMP block vectors; Figure 14 Reference sample points of the template are shown in the template and reference image; Figure 15 The template and reference sample points of the template are shown for a block with sub-block motion information using the current block; Figure 16 The locations of the spatial merge candidates are shown; Figure 17 The candidate pairs considered for redundancy checks of spatial merge candidates are shown. Figure 18 A schematic diagram of motion vector scaling for temporal Merge candidates is shown; Figure 19 The candidate positions, C0 and C1, for the time-domain Merge candidate are shown; Figure 20 The spatial neighbor block used to derive spatial merge candidates is shown; Figure 21A This shows the spatial neighbor blocks used by ATVMP; Figure 21BA schematic diagram is shown to derive the motion field of a sub-CU by applying motion displacements from spatial neighbors and scaling motion information from the corresponding co-located sub-CUs. Figure 22 The non-adjacent spatial domain candidates for the IBC mode are shown; Figure 23 This illustrates BV prediction guided by block vectors; Figure 24 The adjacent airspace locations are shown; Figure 25 B is shown n The five positions in the middle; Figure 26 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 27 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0012] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0018] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded images and associated data. A coded image is a coded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0023] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture containing the current video block.
[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0031] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1 and a motion vector indicating the spatial displacement between the current video block and the reference video block, the reference image containing the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction for the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0035] In one example, motion estimation unit 204 may indicate a value to video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0036] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block, which is stored in the buffer 213.
[0044] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is the overall inverse of the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge pattern. Using AMVP, several most likely candidates are derived based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0050] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies the inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video codec steps in detail, it should be understood that the corresponding decoding steps of the inverse codec will be implemented by the decoder. Furthermore, the term "video processing" includes video codec or compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compression format to another or at different compression bitrates.
[0056] 1. Brief Overview This disclosure relates to image / video codecs, and in particular to block vector guided block vector prediction. It can be applied to existing video codec standards such as HEVC or standard VVC (Video Codec for Multipurpose). It is also applicable to future video codec standards or video codecs.
[0057] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding.
[0058] To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. JVET meetings are held quarterly. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the April 2018 JVET meeting, and the first version of the VVC Test Model (VTM) was also released at that time. The VVC working draft and the VTM test model are updated after each meeting. The VVC project was technically completed (FDIS) at the July 2020 meeting.
[0059] In January 2021, JVET established an Exploratory Experiment (EE) with the goal of improving compression efficiency beyond the capabilities of VVC using novel traditional algorithms. Shortly thereafter, ECM was built as a general-purpose software foundation for long-term exploratory work toward next-generation video codec standards.
[0060] 2.1. Intra-Block Copy (IBC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precise. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.
[0061] On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-Merge mode, block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed.
[0062] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4×4 sub-blocks. For the larger current block, a hash key match with a reference block is determined when all hash keys of all 4×4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected.
[0063] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.
[0064] At the CU level, the IBC mode is transmitted via signaling using a flag, and it can be transmitted via signaling as IBCAMVP mode or IBC skip / Merge mode, as follows: – IBC Skip / Merge Mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and paired candidates.
[0065] – IBC AMVP Mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left nearest neighbor and one from the top nearest neighbor (if IBC encoding / decoding is used). When either nearest neighbor is unavailable, the default block vector is used as the prediction value. A flag is transmitted via signaling to indicate the index of the block vector prediction value.
[0066] 2.1.1. Simplification of IBC Vector Prediction The BV predictions of the Merge and AMVP modes in IBC will share a common list of predictions consisting of the following elements: • Two adjacent airspace locations (in Figure 4 (A1, B1 in the text) • 5 HMVP entries; • The default value is zero vector.
[0067] For Merge mode, up to the first 6 entries of the list will be used; for AMVP mode, the first 2 entries of the list will be used. The list must also conform to the shared Merge list region requirements (the same list within the shared SMR).
[0068] 2.1.2. IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined region, which includes the current CTU region and some regions of the left CTU. Figure 5 The reference area of the IBC mode is shown, where each block represents a 64x64 lumen sample unit.
[0069] Depending on the location of the current codec CU within the current CTU, the following applies: – If the current block falls within the top-left 64x64 block of the current CTU, then in addition to the already reconstructed samples in the current CTU, the CPR mode can be used to reference the reference samples in the bottom-right 64x64 block of the left CTU. The current block can also use the CPR mode to reference the reference samples in the bottom-left 64x64 block of the left CTU and the reference samples in the top-right 64x64 block of the left CTU.
[0070] – If the current block falls within the upper right 64x64 block of the current CTU, then in addition to the already reconstructed samples in the current CTU, if the brightness position (0,64) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference the reference samples in the lower left and lower right 64x64 blocks of the left CTU; otherwise, the current block can also reference the reference samples in the lower right 64x64 block of the left CTU.
[0071] – If the current block falls within the lower-left 64x64 block of the current CTU, then in addition to the already reconstructed samples in the current CTU, if the brightness position (64,0) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference reference samples in the upper-right and lower-right 64x64 blocks of the left CTU. Otherwise, the current block can also use CPR mode to reference reference samples in the lower-right 64x64 block of the left CTU.
[0072] – If the current block falls within the bottom right 64x64 block of the current CTU, it can only use CPR mode to refer to the reconstructed sample points in the current CTU.
[0073] This restriction allows for the use of local on-chip memory to implement IBC mode for hardware implementations.
[0074] 2.1.3. Interaction between IBC and other codec tools The interaction between IBC mode and other inter-frame coding / decoding tools in VVC, such as Paired Merge Candidate, History-Based Motion Vector Prediction (HMVP), Intra / Inter Joint Prediction Mode (CIIP), Merge Mode with Motion Vector Difference (MMVD), and Geometric Partitioning Mode (GPM), is as follows: – IBC can be used with pairwise merge candidates and HMVP. New pairwise IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC movements are inserted into the history cache for future reference.
[0075] – IBC cannot be used in combination with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM.
[0076] – When using DUAL_TREE segmentation, IBC is not allowed for chroma codec blocks.
[0077] Unlike HEVC screen content encoding / decoding extensions, the current image is no longer included as one of the reference images in reference image list 0 for IBC prediction. The derivation process of motion vectors for the IBC mode excludes all neighboring blocks in inter-frame modes and vice versa. The following IBC design aspects are applied: – IBC shares the same process as regular MV Merge, including pairwise Merge candidates and historical motion predictions, but does not allow TMVP and zero vectors because they are invalid for IBC mode.
[0078] – Separate HMVP caches (5 candidates each) are used for traditional MV and IBC.
[0079] – Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that no invalid vectors exist in the bitstream, and if a merge candidate is invalid (out of range or 0), then the merge should not be used. As described below, such bitstream consistency constraints are expressed based on virtual buffers.
[0080] – For deblocking, IBC is processed as an inter-frame mode.
[0081] – If the current block is encoded or decoded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is transmitted via signaling to indicate only whether the MV is an inter-frame pixel or 4 integer pixels.
[0082] – The number of IBC Merge candidates can be transmitted separately in the strip header from the number of regular candidates, sub-block candidates, and geometric Merge candidates.
[0083] The concept of a virtual buffer is used to describe the permissible reference region and valid block vector for IBC prediction modes. Representing the CTU size as ctbSize, the virtual buffer ibcBuf has a width wIbcBuf = 128x128 / ctbSize and a height hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32.
[0084] The size of the VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64).
[0085] The virtual IBC cache ibcBuf is maintained as follows.
[0086] – At the beginning of decoding each CTU line, refresh the entire ibcBuf with an invalid value of -1.
[0087] – At the beginning of decoding the VPDU(xVPDU, yVPDU) relative to the top left corner of the image, set ibcBuf[x][y] = 1, where x = xVPDU%wIbcBuf, …, xVPDU% wIbcBuf + W v 1;y = yVPDU%ctbSize,…, yVPDU%ctbSize + W v 1.
[0088] – After decoding the CU containing (x, y) relative to the top left corner of the image, set ibcBuf[ x % wIbcBuf ][ y % ctbSize ] = recSample[ x ][ y ].
[0089] For a block covering coordinates (x, y), if for the block vector bv = (bv[0], bv[1]) If the following are true, then it is valid; otherwise, it is invalid: ibcBuf[ (x + bv[0])% wIbcBuf][ (y + bv[1]) % ctbSize ] should not be equal to 1 .
[0090] 2.1.4. IBC Virtual Cache Test The luminance block vector bvL (luminance block vector with 1 / 16 fractional sample precision) must comply with the following constraints: - CtbSizeY is greater than or equal to ( ( yCb + ( bvL[ 1 ]>>4 ) )&( CtbSizeY 1 ) ) +cbHeight.
[0091] - For x = xCb..xCb + cbWidth 1 and y = yCb..yCb + cbHeight 1, IbcVirBuf[ 0 ][ ( x + (bvL[ 0 ]>>4 ) )&( IbcBufWidthY 1 ) ][ ( y + (bvL[ 1 ]>>4 ) )&(CtbSizeY 1) should not be equal to -1.
[0092] Otherwise, bvL is considered invalid bv.
[0093] Samples are processed in units of CTB. Each luminance CTB has an array size of CtbSizeY in both width and height, where CtbSizeY is in units of samples.
[0094] - (xCb, yCb) is the brightness position of the top-left sample of the current luma codec block relative to the top-left luma sample of the current image. - cbWidth specifies the width of the current codec block in the luma sample. - cbHeight specifies the height of the current codec block in the luminance sample.
[0095] 2.2. IBC Merge / AMVP List Construction The IBC Merge / AMVP list construction has been modified as follows: • It can be inserted into the IBC Merge / AMVP candidate list only if the IBC Merge / AMVP candidate is valid.
[0096] • Candidate airspace for the upper right, lower left, and upper left (e.g.) Figure 6 The candidates shown (B0, A0, and B2) and a pairwise average candidate can be added to the IBC Merge / AMVP candidate list.
[0097] • Template-based adaptive reordering (ARMC-TM) is applied to the IBC Merge list.
[0098] The HMVP table size for IBC is increased to 25. After deriving a maximum of 20 IBC Merge candidates using full deduplication, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list.
[0099] Zero vector candidates used to populate the IBC Merge / AMVP list are replaced by a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in the IBC Merge mode, and therefore are discarded as BVPs in the IBC candidate list.
[0100] Three candidates are located at the nearest corner of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), their coordinates determined by the width and height of the current block and the ΔX and ΔY parameters, as shown below. Figure 7 As shown.
[0101] 2.3. IBC with Template Matching Template matching is used in both the IBC Merge mode and the IBC AMVP mode in IBC.
[0102] Compared to the list used in the regular IBC Merge mode, the IBC-TM Merge list is modified to select candidates based on a deduplication method that has a motion distance between the candidates and those in the regular TM Merge mode. The final zero motion is replaced by motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0103] In IBC-TM Merge mode, the selected candidate is refined using a template matching method before the RDO or decoding process. IBC-TM Merge mode competes with the regular IBC Merge mode, and the TM-Merge flag is transmitted via signaling.
[0104] In the IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM Merge list. Each of these three selected candidates is refined using a template matching method and ranked according to its resulting template matching cost. Then, typically only the top two are considered in the motion estimation process.
[0105] Since the IBC motion vector is constrained to be (i) an integer and (ii) as... Figure 8 Within the reference area shown, template matching and refinement for both IBC-TM Merge and AMVP modes are relatively straightforward. Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with either integer or 4-pixel precision, depending on the AMVR value. Such refinement only accesses samples that are not interpolated. In both cases, the motion vectors for refinement and the template used in each refinement step must respect the constraints of the reference area.
[0106] 2.4. IBC Reference Area The IBC reference area is extended to the two CTU rows above. Figure 9 The reference region used for encoding and decoding CTUs (m, n) is shown. Specifically, for a CTU (m, n) to be encoded and decoded, the reference region includes CTUs indexed as (m-2, n-2)…(W, n-2), (0, n-1)…(W, n-1), (0, n)…(m, n), where W represents the maximum horizontal index within the current slice, strip, or picture. When the CTU size is 256, the reference region is limited to the uppermost CTU row. This setting ensures that for CTU sizes of 128 or 256, the IBC does not require additional memory in the current ETM platform. The block vector search (or local search) range per sample is limited to horizontal [-(C<<1), C>>2] and vertical [-C, C>>2] to accommodate the reference region expansion, where C represents the CTU size.
[0107] 2.5. Reconstructing the Reordered IBC (RR-IBC) The IBC (RR-IBC) mode allows for reconstruction and reordering of IBC blocks after IBC encoding and decoding. When RR-IBC is applied, samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the predicted block is derived without flipping. On the decoder side, the reconstructed block is flipped back to recover the original block.
[0108] Two flipping methods are supported for blocks encoded with RR-IBC: horizontal flipping and vertical flipping. First, for blocks encoded with IBC AMVP, a syntax flag is signaled indicating whether the reconstruction has been flipped. If it has been flipped, another flag is further signaled specifying the flipping type. For IBC Merge, the flipping type is inherited from the neighboring block, with no syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block are typically horizontally or vertically aligned. Therefore, when a horizontal flip is applied, the vertical component of the BV is not signaled and is presumed to be 0. Similarly, when a vertical flip is applied, the horizontal component of the BV is not signaled and is presumed to be 0.
[0109] To better utilize symmetry, a flip-aware BV adjustment method is applied to refine block vector candidates. For example, as Figure 7 As shown in A and 7B, ( x nbr , y nbr )and( x cur , y cur () represents the coordinates of the center sample points of the neighboring blocks and the current block, respectively. BV nbr and BV cur These represent the BV of the neighboring block and the current block, respectively. This is relevant when the neighboring block is encoded and decoded using a horizontal flipping method. BV cur The horizontal component does not inherit BV directly from neighboring blocks, but rather through... BV nbr The horizontal component (represented as) BV nbr h Add motion displacement to the calculation, that is, BV cur h =2( x nbr - x cur ) + BV nbr h Similarly, in the case where adjacent blocks are encoded and decoded using vertical flipping, BV cur The vertical component is through the direction BV nbr The vertical component (represented as) BV nbr v Add motion displacement to the calculation, that is, BV cur v =2( y nbr - y cur ) + BV nbr v . Figure 10A A schematic diagram of BV adjustment for horizontal flipping is shown; and Figure 10B A schematic diagram of BV adjustment for vertical flipping is shown.
[0110] 2.6. IBC Merge Mode with Block Vector Difference (IBC-MBVD) Affine-MMVD and GPM-MMVD have been adopted by ECM as extensions to the regular MMVD mode. Extending the MMVD mode to the IBC Merge mode is a natural progression.
[0111] In IBC-MBVD, the distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD direction is two horizontal directions and two vertical directions.
[0112] The basic candidate is selected from the top five candidates in the reordered IBC Merge list. All possible MBVD refinement positions (20×4) for each basic candidate are reordered based on the SAD cost between the template (the row above and the column to the left of the current block) and its reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are reserved as available positions and thus used for MBVD index encoding and decoding. The MBVD index is binarized using rice codes with a parameter equal to 1.
[0113] Blocks encoded and decoded by IBC-MBVD do not inherit the flip type from neighboring blocks encoded and decoded by RR-IBC.
[0114] 2.7. Intra-frame template matching Intra-frame template matching prediction (intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signaling, and the same prediction operation is performed on the decoder side.
[0115] The prediction signal is obtained by comparing the L-shaped causal nearest neighbors of the current block with... Figure 11 Another block of match in the predefined search area consisting of the following is generated: R1: Current CTU; R2: Top left CTU; R3: Above CTU; R4: Left CTU.
[0116] The sum of absolute differences (SAD) is used as the cost function.
[0117] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0118] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW, SearchRange_h = a * BlkH.
[0119] in" "" is a constant that controls the trade-off between gain and complexity. In practice, " "Equals 5."
[0120] Enable intra-frame template matching for CUs with width and height dimensions less than or equal to 64. The maximum CU size for intra-frame template matching is configurable.
[0121] When DIMD is not used in the current CU, the intra-template matching prediction mode is transmitted at the CU level via a dedicated flag.
[0122] 2.8. Using block vectors derived from IntraTMP for IBC A method is proposed that uses block vectors derived from IntraTMP for IBC. The proposed method stores the IntraTMP block vector in an IBC block vector cache, and the current IBC block can use both the IBC BV and the IntraTMP BV of neighboring blocks as BV candidates for the IBC BV candidate list, such as... Figure 12 As shown.
[0123] Figure 13 An example is shown comparing block vector candidates in the IBC block vector candidate list that come only from neighboring blocks of the IBC codec, and block vector candidates in the proposed IBC block vector candidate lists that come from both neighboring blocks of the IBC codec and neighboring blocks of the IntraTMP codec. IntraTMP block vectors are added to the IBC block vector candidate list as spatial domain candidates.
[0124] It should be noted that the proposed method makes IBC block vector prediction more efficient by using different block vectors without requiring additional memory for storing block vectors.
[0125] 2.9. Adaptive Reordering of Merge Candidates Using Template Matching (ARMC-TM) Merge candidates are adaptively reordered using template matching (TM). The reordering method is applied to the regular Merge pattern, TM Merge pattern, and affine Merge pattern (excluding SbTMVP candidates). For the TM Merge pattern, Merge candidates are reordered before the refinement process.
[0126] The initial Merge candidate list is first constructed according to a given inspection order, such as spatial candidates, TMVP candidates, non-adjacent candidates, HMVP candidates, paired candidates, and virtual Merge candidates. Then, the candidates in the initial list are divided into several subgroups. For Template Matching (TM) Merge mode and Adaptive DMVR mode, each Merge candidate in the initial list is first refined using TM / multi-pass DMVR. The Merge candidates in each subgroup are reordered to generate a reordered Merge candidate list, and the reordering is based on the cost value of template matching. The index of the selected Merge candidate in the reordered Merge candidate list is signaled to the decoder. For simplicity, the Merge candidates in the last subgroup (not the first subgroup) are not reordered. During the construction of the Merge motion vector candidate list, all zero candidates from the ARMC reordering process are excluded. For the regular Merge mode and TM Merge mode, the subgroup size is set to 5. For the affine Merge mode, the subgroup size is set to 3.
[0127] • Cost Calculation The template matching cost of the merge candidate during the reordering process is measured by the SAD between the samples of the current block's template and their corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located using the motion information of the merge candidate. When the merge candidate utilizes bidirectional prediction, the reference samples of the merge candidate's template are also generated through bidirectional prediction, such as... Figure 14 As shown.
[0128] • Refinement of the initial Merge candidate list When multi-pass DMVR is used to derive refined motion for the initial Merge candidate list, only the first pass of the multi-pass DMVR (i.e., PU level) is applied to the reordering. When template matching is used to derive refined motion, the template size is set to 1. When the block is a flat block with a width greater than twice its height or a narrow block with a height greater than twice its width, only the top or left template is used during TM motion refinement. TM is expanded to perform 1 / 16 pixel MVD precision. The first four Merge candidates are reordered in TM Merge mode using refined motion.
[0129] For sub-block-based merge candidates with a sub-block size equal to Wsub × Hsub, the upper template includes several sub-templates of size Wsub × 1, and the left template includes several sub-templates of size 1 × Hsub. For example... Figure 15 As shown, the motion information of the sub-blocks in the first row and first column of the current block is used to derive the reference sample points of each sub-template.
[0130] • Reordering criteria During the reordering process, a candidate is considered redundant if the cost difference between a candidate and its predecessor is less than the lambda value, for example, |D1-D2|<λ (where D1 and D2 are the costs obtained during the first ARMC sorting, and λ is the Lagrange parameter used in the RD criterion on the encoder side).
[0131] The proposed algorithm is defined as follows: - Determine the minimum cost difference between a candidate in the list and its predecessor.
[0132] If the minimum cost difference is greater than or equal to λ, the list is considered sufficiently different, and reordering stops.
[0133] If the minimum cost difference is less than λ, the candidate is considered redundant and is moved to another position in the list. This other position is the first position that is sufficiently different from the candidate's predecessor.
[0134] - The algorithm stops after a finite number of iterations (if the minimum cost difference is not less than λ).
[0135] This algorithm has been applied to the regular Merge pattern, TM Merge pattern, BM Merge pattern, and affine Merge pattern. Similar algorithms have been applied to Merge MMVD and Symbolic MVD prediction methods that also use ARMC for reordering.
[0136] The value of λ is set to be equal to the rate-distortion criterion used to select the best Merge candidate on the encoder side for the low-latency configuration, and is set to be equal to the value λ corresponding to another QP for the random access configuration. For QP offsets that do not exist in the SPS, a set of λ values corresponding to each QP offset transmitted through the signal is provided in the SPS or strip header.
[0137] • Extensions to the AMVP pattern The ARMC design also applies to the AMVP mode, where AMVP candidates are reordered based on TM costs. For the Template Matching (TM-AMVP) mode for advanced motion vector prediction, an initial AMVP candidate list is constructed, followed by refinement from TM to construct a refined AMVP candidate list. Additionally, MVP candidates with TM costs greater than a threshold (which is equal to five times the cost of the first MVP candidate) are skipped.
[0138] Note that when surround motion compensation is enabled, the MV candidate must be clipped to account for surround offset.
[0139] 2.10. Extended Merge Forecast In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: 1) Airspace MVP from adjacent CUs in airspace, 2) Temporal MVP from the same CU, 3) Historical MVP from FIFO table 4) Paired average MVP 5) Zero MV.
[0140] The size of the merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU encoding / decoding in merge mode, the index of the best merge candidate is encoded using rounded unary binarization (TU). The first bit of the merge index is encoded / decoded using the context, and bypass encoding / decoding is used for the other bits.
[0141] This section provides the derivation process for various merge candidates. Similar to HEVC, VVC also supports parallel derivation of the merge candidate list for all CUs within a region of a specific size.
[0142] 2.10.1. Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Figure 16 Of the candidates at the indicated positions, a maximum of four Merge candidates are selected. The derivation order is B. 1, A 1, B 0, A0 and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or when it is intra-frame encoded / decoded. After adding candidates at position B1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only... Figure 17 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information.
[0143] 2.10.2. Time-domain candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU and the reference indices are explicitly transmitted via signal transmission in the strip header. Figure 18 As shown by the dashed lines, the scaled motion vectors of the temporal merge candidate are obtained by scaling the motion vectors of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index of the temporal merge candidate is set to 0.
[0144] like Figure 19 As shown, the position for the temporal candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded or decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate.
[0145] 2.10.3. Historical Merge Candidate Derivation Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0146] HMVP Table Size S Setting it to 6 indicates that a maximum of 5 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, the first-in-first-out (FIFO) rule of the constraint is utilized, where a redundancy check is first applied to find if the same HMVP already exists in the table. If found, the same HMVP is removed from the table, and all subsequent HMVP candidates are shifted forward, with the same HMVP inserted as the last entry in the table.
[0147] HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked sequentially and inserted into the candidate list, following the TMVP candidates. Redundancy checks for spatial or temporal Merge candidates are applied to HMVP candidates.
[0148] To reduce the number of redundant check operations, the following simplifications are introduced: 1. The last two entries in the table perform redundancy checks on the A1 and B1 airspace candidates, respectively.
[0149] 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list building process from HMVP is terminated.
[0150] 2.10.4. Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list using the first two Merge candidates. The first Merge candidate is defined as p0Cand, and the second Merge candidate can be defined as p1Cand. For each reference list, the averaged motion vector is calculated separately based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available in a list, they are averaged even if they point to different reference images, and their reference image is set to the reference image of p0Cand; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, the list remains invalid. Furthermore, if the half-pixel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.
[0151] If the Merge list is not full after adding pairwise average Merge candidates, insert zero MVP at the end until the maximum number of Merge candidates is reached.
[0152] 2.10.5. Merge estimation region The Merge Estimation Region (MER) allows for the independent derivation of Merge candidate lists for CUs within the same Merge Estimation Region (MER). Candidate blocks located within the same MER as the current CU are not included in the generation of the current CU's Merge candidate list. Furthermore, the update process for the historical motion vector prediction candidate list is only updated when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left brightness sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and is transmitted via signaling as log2_parallel_merge_level_minus2 in the sequence parameter set.
[0153] 2.10.6. Non-adjacent airspace candidates In ECM, non-adjacent airspace merge candidates are inserted after the TMVP in the regular merge candidate list. The style of airspace merge candidates is as follows: Figure 20 As shown. The distance between non-adjacent spatial domain candidates and the current codec block is based on the width and height of the current codec block. Line buffer limits are not applied.
[0154] 2.10.7. ARMC Based on MV Candidate Types Merge candidates for a single candidate type (e.g., TMVP or Non-Adjacent MVP (NA-MVP)) are reordered based on ARMC™ arithmetic values. The reordered candidates are then added to the Merge candidate list. For the TMVP candidate type, more TMVP candidates with more temporal locations and different inter-frame prediction orientations are added to perform reordering and selection. Furthermore, the NA-MVP candidate type is further expanded using more spatially non-adjacent locations. The target reference image for a TMVP candidate can be selected from any reference image in the list based on a scaling factor. The selected reference image is the one with a scaling factor closest to 1.
[0155] 2.11. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-image to improve motion vector prediction and merge patterns for CUs in the current image. The same co-image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in two main aspects: – TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level; – TMVP obtains temporal motion vectors from co-op blocks in the co-op image (the co-op block is the lower right or center block relative to the current CU), and SbTMVP applies motion displacement before obtaining temporal motion information from the co-op image, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks from the current CU.
[0156] The SbTVMP process is as follows: Figure 18 A and Figure 18 As shown in B, SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, it checks... Figure 21A The spatial nearest neighbor A1 is selected. If A1 has a motion vector that uses a co-located image as its reference image, then that motion vector is chosen as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).
[0157] In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to the position of the current block. Figure 21B The corresponding image shown obtains motion information (motion vectors and reference indices) at the sub-CU level. Figure 21BThe example assumes the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-location image is used to derive the motion information for the sub-CU. After the motion information of the co-location sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.
[0158] In VVC, a sub-block-based combined Merge list containing both SbTVMP candidates and affine Merge candidates is used for signaling in sub-block-based Merge mode. SbTVMP mode is enabled / disabled via the Sequence Parameter Set (SPS) flag. If SbTVMP mode is enabled, the SbTVMP prediction is added as the first entry in the sub-block-based Merge candidate list, followed by the affine Merge candidate. The size of the sub-block-based Merge list is transmitted via signaling in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5.
[0159] The sub-CU size used in SbTMVP is fixed at 8x8, and like the affine Merge pattern, the SbTMVP pattern only applies to CUs with a width and height greater than or equal to 8.
[0160] The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, that is, for each CU in the P-strip or B-strip, an additional RD check is performed to determine whether to use the SbTMVP candidate.
[0161] 2.12. IBC-CIIP, IBC-GPM and IBC-LIC 2.12.1. IBC-CIIP First Option Intra-Block Copy and Intra-Prediction Combination (IBC-CIIP) is an encoding / decoding tool for CUs that uses IBC in Merge mode and intra-prediction to obtain two prediction signals, which are then weighted and summed to generate the final prediction. Specifically, if intra-prediction is in planar or DC mode, the final prediction is obtained as follows:
[0162] in These represent the IBC prediction signal and the intra-frame prediction signal, respectively. If both the upper CU and the left CU are intra-frame encoded and decoded, then... It is set to (1,2) if either the top CU or the left CU is intra-coded. It is set to (2,2) if both the top CU and the left CU are encoded and decoded by IBC. It is set to equal to (3,2). Otherwise (i.e., if intra-frame prediction is in directional mode), the final prediction is obtained by adaptively switching the prediction samples and IBC of the intra-frame mode. For illustrative purposes, assume the current CU size is... Furthermore, the intra-frame mode is either horizontal or vertical. If both the upper neighboring CU and the left neighboring CU are intra-coded, then the final predicted left... Partial (horizontal mode) or above Partial (vertical mode) is set as the intra-predictive signal; and if only one of the upper CU and the left CU is intra-coded, the final predicted left CU will be... Partial (horizontal mode) or above Partial (vertical mode) is set to intra-frame prediction signal; and if both the upper CU and the left CU are IBC or inter-frame encoded / decoded, the final predicted left CU will be... Partial (horizontal mode) or above A portion (vertical mode) is set as the intra-predicted signal. In the above, apart from the intra-predicted portion, the other portions of the final prediction are set as IBC prediction samples.
[0163] 2.12.2. IBC-CIIP Second Option Intra-Block Copy and Intra-Prediction Combination (IBC-CIIP) is an encoding / decoding tool for CUs that uses IBC and intra-prediction to obtain two prediction signals, which are then weighted and summed to generate the final prediction, as shown below:
[0164] in This represents the IBC prediction signal and the intra-frame prediction signal. For IBC Merge mode and IBCAMVP mode, It was set to equal to (13,4) and (1,1).
[0165] An intra-prediction mode (IPM) candidate list is used to generate the intra-prediction signal, and the IPM candidate list size is predefined to 2. The IPM index is transmitted via signaling to indicate which IPM to use.
[0166] 2.12.3. IBC with Geometric Segmentation (IBC-GPM) Intra-Block Copy with Geometric Partitioning (IBC-GPM) is an encoding / decoding tool that geometrically divides the CU into two sub-segments. IBC and intra-prediction are used to generate the prediction signals for the two sub-segments. IBC-GPM can be applied to either the regular IBC Merge mode or the IBC-TM Merge mode. The Intra-Prediction Mode (IPM) candidate list is constructed using the same method as GPM with Inter-Frame and Intra-Frame Prediction for intra-frame prediction, and the IPM candidate list size is predefined to 3. A total of 48 geometric partitioning modes are available, divided into two sets of geometric partitioning modes as follows: Table 1: Geometric segmentation patterns in the first geometric segmentation pattern set
[0167] Table 2: Geometric segmentation patterns in the second geometric segmentation pattern set
[0168] When using IBC-GPM, the IBC-GPM geometric segmentation mode set flag is transmitted via signaling to indicate whether a first or second geometric segmentation mode set is selected, followed by the geometric segmentation mode index. The IBC-GPM intra-frame flag is transmitted via signaling to indicate whether intra-frame prediction is used for the first sub-segmentation. When intra-frame prediction is used for sub-segmentation, the intra-frame prediction mode index is transmitted via signaling. When IBC is used for sub-segmentation, the merge index is transmitted via signaling.
[0169] 2.12.4. IBC with local illumination compensation (IBC-LIC) Intra-Block Copy with Local Illumination Compensation (IBC-LIC) is an encoder-decoder tool that uses linear equations to compensate for local illumination variations within an image between an IBC-encoded CU and its predicted blocks. Except for the use of block vectors to generate the reference template in IBC-LIC, the derivation of the parameters of the linear equations is the same as that for LIC used for inter-frame prediction. IBC-LIC can be applied to both IBC AMVP and IBC Merge modes. For IBC AMVP mode, an IBC-LIC flag is transmitted via signaling to indicate the use of IBC-LIC. For IBC Merge mode, the IBC-LIC flag is inferred from the Merge candidate.
[0170] 2.13. IBCs with non-adjacent airspace candidates Several new candidate blocks obtained from the BV of non-adjacent spatial neighboring blocks are proposed for insertion into the candidate lists of IBC Merge mode and IBC AMVP. For example... Figure 22As shown, a grid pattern similar to that used for non-adjacent spatial domain candidates in regular inter-frame connections is used to obtain non-adjacent BVs. For both IBC Merge and IBC AMVP, the obtained BVs are inserted between adjacent spatial domain candidates and HBVP candidates. Additionally, the restriction that spatial domain candidates are not used to construct the IBC Merge candidate list for 4x4 CUs has been removed.
[0171] 2.14. Bidirectional Prediction of IBC GPM Bidirectional prediction IBC GPM uses different IBCs to generate prediction samples for two GPM segmentations. In bidirectional prediction IBC GPM, the core IBC GPM design (e.g., 48 GPM modes, IBC Merge candidate list) remains the same as the core IBC GPM design in unidirectional prediction IBC GPM in ECM-9.0. Here, unidirectional prediction IBC GPM uses both IBC and intra-frame modes to generate prediction samples for each GPM segment.
[0172] Only the signaling and hybrid scheme are slightly modified compared to the existing unidirectional prediction IBC GPM. For signaling, two flags are transmitted via signaling to indicate the two segmented prediction modes. For hybridization, the GPM adaptive hybridization used for inter-frame GPM is also utilized in bidirectional prediction IBC GPM.
[0173] 2.15. IBC BVP-Merge and Bidirectional Prediction IBC Merge Inspired by AMVP-Merge, IBC BVP-Merge derives two required BVs from IBC Block Vector Prediction (BVP) and IBC Merge. Unlike the AMVP-Merge model, two distinct indices for IBC BVP and IBC Merge candidates are transmitted from the encoder to the decoder via signals.
[0174] Bidirectional predictive IBC Merge derives two BVs from an existing list of IBC Merge candidates using two different indices. These two indices are transmitted from the encoder to the decoder via a signal. The goal of bidirectional predictive IBC Merge is both IBC MBVD and IBC regular merge.
[0175] 2.16. Derivation of the IBC MBVD list The IBC MBVD list derivation allows for adaptive BVD offsets along the MBVD direction. Specifically, the MBVD candidate search is a two-step algorithm that begins by checking the template SAD cost of the offsets to the BVD along each direction at M-pixel intervals. The second step of the search checks the template SAD cost for the intermediate candidates around the selected candidate. The candidate with the lowest TM cost is included in the final MBVD list.
[0176] 2.17. Multiple Hypothesis Prediction ( MHP) In multi-hypothesis inter-frame prediction mode, in addition to the regular bidirectional prediction signal, one or more additional motion-compensated prediction signals are also transmitted via signal transmission. The resulting overall prediction signal is obtained by sample-by-sample weighted superposition. Utilizing the bidirectional prediction signal... and the first additional inter-frame prediction signal / hypothesis The obtained prediction signal The following was obtained: .
[0177] Weighting factors The new syntax element add_hyp_weight_idx is specified according to the following mapping.
[0178]
[0179] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0180] .
[0181] The resulting overall prediction signal is obtained as the last one. (i.e., has the largest index) Within this EE, a maximum of two additional prediction signals can be used (i.e., Limited to 2).
[0182] The motion parameters for each additional prediction hypothesis can be transmitted via signaling by explicitly specifying the reference index, motion vector prediction index, and motion vector difference, or by implicitly specifying the merge index. A separate multi-hypothesis merge flag distinguishes between these two signaling modes.
[0183] For inter-frame AMVP mode, MHP is applied only when unequal weights are selected in BCW in bidirectional prediction mode.
[0184] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0185] 3. Problem In current BV prediction (e.g., for both IBC Merge and IBC AMVP modes), block vector-guided BV prediction is not utilized. To further improve the efficiency of BV prediction, block vector-guided BV prediction has been introduced.
[0186] 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0187] The term "block" can refer to a codec tree block (CTB), codec tree unit (CTU), codec block (CB), CU, PU, TU, PB, TB, or a video processing unit comprising multiple samples / pixels. Blocks can be rectangular or non-rectangular.
[0188] W and H are the width and height of the current block (e.g., the luminance block).
[0189] For blocks encoded and decoded by IBC and intra-frame TMP, the block vector (BV) is used to indicate the displacement from the current block to a reference block that has been or partially reconstructed within the current frame.
[0190] In the following text, BV candidates are BV predictions or search points. A block has BV information if it is encoded using IBC or intra-frame TMP.
[0191] It should be noted that the proposed method can also be applied to other encoding / decoding methods that may require derivation of MV / BV.
[0192] In one example, assume the block vector BV associated with the current block B0. 0,1 Points to reference block B1, which has been or partially reconstructed within the current image. If B1 has a corresponding BV, then BV is selected. 1,2 The BV information, then when BV 0,2 When the current block is valid, it can obtain a value for B0 with BV. 0,2 = BV 0,1 +BV 1,2 The given BV 0,2 The new reference block B2. And BV 0,2 It is a block vector-guided BV prediction or candidate, which is called a first-order block vector-guided cascaded BVP. If B2 has a corresponding BV of BV... 2,3 The BV information, then when BV 0,3 When the current block is valid, it can obtain a value for B0 with BV. 0,3 = BV 0,2 +BV 2,3 The given BV 0,3 The new reference block B3. And BV 0,3 It is a block vector-guided BV prediction or candidate, which is called a second-order block vector-guided cascaded BVP. If B3 has a corresponding BV of BV... 3,4 The BV information, then when BV0,4 When the current block is valid, it can obtain a value for B0 with BV. 0,4 = BV 0,3 +BV 3,4 The given BV 0,4 The new reference block B4. And BV 0,4 It is a block vector-guided BV prediction or candidate, which is called a third-order block vector-guided cascaded BVP. Similarly, if B n With corresponding BV is BV n,n+1 The BV information, then when BV 0,n+1 When the current block is valid, it can obtain a value for B0 with BV. 0,n+1 = BV 0,n +BV n,n+1 The given BV 0,n+1 New reference block B n+1 And BV 0,n+1 It is a block vector-guided BV prediction or candidate, which is called n Cascaded BVP guided by block vector.
[0193] In one example, assume the block vector BV associated with the current block B0. 0,1 Points to reference block B1, which has been or partially reconstructed within the current image. If B1 has a corresponding BV, then BV is selected. 1,2 The BV information, then when BV 1,2 When the current block is valid, BV 1,2 (This is referred to as a first-order block vector guided direct BVP) is a block vector guided BV prediction or candidate. If B2 has a corresponding BV of BV, then BV is... 2,3 The BV information, then when BV 2,3 When the current block is valid, BV 2,3 (This is called a second-order block vector guided direct BVP) is a block vector guided BV prediction or candidate. If B3 has a corresponding BV of BV, then... 3,4 The BV information, then when BV 3,4 When the current block is valid, BV 3,4 (This is referred to as the third-order block vector guided direct BVP) is a block vector guided BV prediction or candidate. Similarly, if B n With corresponding BV is BV n,n+1 The BV information, then when BV n,n+1 When the current block is valid, BV n,n+1 (It is called) n The block vector-guided direct BVP is a block vector-guided BV prediction or candidate.
[0194] A schematic diagram of block vector-guided BV prediction is shown below. Figure 23 As shown.
[0195] Block Vector Guided BV / MV Prediction 1. In one example, block vector-guided BV and / or MV predictions can be incorporated into the BV and / or MV predictions.
[0196] a. In one example, BV and / or MV prediction can be at least one of the following.
[0197] (a) In one example, BV and / or MV predictions can be regular IBC and / or inter-frame merge predictions.
[0198] (b) In one example, BV and / or MV predictions can be regular IBC and / or inter-frame AMVP predictions.
[0199] (c) In one example, BV and / or MV predictions can be IBC and / or inter-frame-TM merge predictions.
[0200] (d) In one example, BV and / or MV predictions can be IBC and / or inter-frame-TM AMVP predictions.
[0201] (e) In one example, BV prediction can be RR-IBC Merge prediction.
[0202] (f) In one example, BV prediction can be RR-IBC AMVP prediction.
[0203] (g) In one example, BV and / or MV prediction can be IBC-MBVD and / or inter-frame MMVD prediction.
[0204] (h) In one example, BV and / or MV predictions can be IBC-CIIP and / or CIIP.
[0205] (i) In one example, BV and / or MV predictions can be IBC-GPM and / or GPM predictions.
[0206] (j) In one example, BV and / or MV predictions can be BV predictions only from the encoder.
[0207] (k) In one example, BV prediction can be a string copy vector prediction.
[0208] (l) In one example, the BV and / or MV prediction can be any other BV and / or MV prediction.
[0209] 2. In one example, block vector-guided BV and / or MV candidates can be added to the BV and / or MV candidate list.
[0210] a. In one example, the BV and / or MV candidate list can be at least one of the following.
[0211] (a) In one example, the BV and / or MV candidate list can be a regular IBC and / or inter-frame merge list.
[0212] (b) In one example, the BV and / or MV candidate list can be a regular IBC and / or inter-frame AMVP list.
[0213] (c) In one example, the BV and / or MV candidate list can be the IBC and / or inter-frame-TM merge list.
[0214] (d) In one example, the BV and / or MV candidate list can be the IBC and / or inter-frame-TM AMVP list.
[0215] (e) In one example, the BV candidate list can be the RR-IBC Merge list.
[0216] (f) In one example, the BV candidate list can be the RR-IBC AMVP list.
[0217] (g) In one example, the BV and / or MV candidate list can be the IBC-MBVD and / or inter-frame MMVD basic candidate list.
[0218] (h) In one example, the BV and / or MV candidate list can be the IBC-CIIP and / or CIIP Merge list.
[0219] (i) In one example, the BV and / or MV candidate list can be the IBC-CIIP AMVP and / or CIIP AMVP list.
[0220] (j) In one example, the BV and / or MV candidate list can be the IBC-GPM Merge and / or GPM Merge list.
[0221] (k) In one example, the BV and / or MV candidate list can be the IBC-GPM AMVP and / or GPM AMVP list.
[0222] (l) In one example, the BV and / or MV candidate list can be the BV candidate list of the encoder only.
[0223] (m) In one example, the BV and / or MV candidate list can be any other BV candidate and / or MV candidate list.
[0224] 3. In one example, the block vector-guided BV prediction or candidate can be derived using at least one of the following methods.
[0225] a. In one example, a block vector-guided BV prediction or candidate can be a block vector-guided cascaded BV prediction or candidate.
[0226] b. In one example, the block vector BV associated with the current block B0. 0,1 Point to reference block B1. If B1 has a value represented as BV 1,2 If the pointer to reference block B2 is BV, then BV 0,2 By BV 0,2 = BV 0,1 +BV 1,2 Given. And BV 0,2 It is a block vector-guided BV prediction or candidate, which is called a first-order block vector-guided cascaded BVP or candidate.
[0227] (a) In one example, it can be required that B1 should have been reconstructed, or partially reconstructed, within the current image.
[0228] (b) In one example, BV can be required. 0,2 The current block must be valid.
[0229] c. In one example, a block vector-guided BV prediction or candidate could be... n A cascaded BVP or candidate guided by an order (e.g., n≥1) block vector is denoted as BV. 0,n+1 .
[0230] (a) BV 0,n+1 = BV 0,1 +BV 1,2 +BV 2,3 +…+BV n-1,n +BV n,n+1. d. In one example, a block vector-guided BV prediction or candidate can be a block vector-guided direct BV prediction or candidate.
[0231] e. In one example, the block vector BV associated with the current block B0. 0,1 Point to reference block B1. If B1 has a value represented as BV 1,2 BV, then BV 1,2 (It is called a first-order block vector guided direct BVP or candidate) is a block vector guided BV prediction or candidate.
[0232] (a) In one example, it can be required that B1 should have been reconstructed, or partially reconstructed, within the current image.
[0233] (b) In one example, BV can be required. 1,2 The current block must be valid.
[0234] f. In one example, the block vector-guided BV prediction or candidate could be n A direct BVP or candidate guided by a block vector of order (e.g., n≥1) is denoted as BV. n,n+1 .
[0235] g. In one example, the block vector BV associated with the current block B0. 0,1 It can be derived from at least one of the following candidates.
[0236] (a) In one example, the block vector BV associated with the current block B0. 0,1 BV candidates can be derived from any existing BV candidate list before the BV guided by the derivation block vector.
[0237] 1) In one example, an existing BV candidate can be a block vector-guided BV candidate.
[0238] 2) In one example, if BV 0,i+1 yes i For cascaded BV candidates guided by block vectors, then correspondingly, BV 0,i+2 yes (i+1) Cascaded BV candidates guided by block vectors.
[0239] 3) In one example, if BV 0,i+1 yes i For cascaded BV candidates guided by block vectors, then correspondingly, BV i+1,i+2 yes (i+1) Direct BV candidates guided by block vectors.
[0240] (b) In one example, the block vector BV associated with the current block B0. 0,1 Some predefined BV candidates can be derived from the BV candidate list before the BVP guided by the derivation block vector.
[0241] 1) In one example, the predefined BV candidate can be the top N BV candidates.
[0242] 2) In one example, the predefined BV candidates can be the selected N BV candidates.
[0243] 3) In one example, a predefined BV candidate can be a BV candidate of a specified candidate type.
[0244] i. In one example, the specified candidate type can involve adjacent (e.g., Figure 6 or Figure 24 Airspace BV candidate.
[0245] ii. In one example, the specified candidate type can involve non-adjacent (e.g., Figure 22 Airspace BV candidate.
[0246] iii. In one example, the specified candidate type can involve HMVP BV candidates.
[0247] iv. In one example, the specified candidate type may involve adjacent time-domain BV candidates.
[0248] v. In one example, the specified candidate type can involve non-adjacent time-domain BV candidates.
[0249] vi. In one example, the specified candidate type can involve pairwise average BV candidates.
[0250] vii. In one example, the specified candidate type can involve block vector-guided BV candidates.
[0251] h. In one example, only BV candidates guided by first-order block vectors can be used.
[0252] i. In one example, only first-order and second-order block vector-guided BV candidates can be used.
[0253] j. In one example, only the first... n BV candidates guided by block vectors can be used.
[0254] 4. In one example, when from B n Derivation of BV n,n+1 At that time, (multiple) specific locations can be examined to derive BV.
[0255] a. In one example, to derive BV, only B n The center (e.g., Figure 25 The position of Ctr in the block can be checked. Assume the width and height of the block are represented as W and H, respectively.
[0256] (a) Assuming the top left position within the block is (0, 0), the center position of the block is defined as (W>>1, H>>1).
[0257] (b) Assuming the top-left position within the block is (0, 0), then the center position of the block is defined as ((W>>1)-1, H>>1).
[0258] (c) Assuming the top-left position within the block is (0, 0), then the center position of the block is defined as (W>>1, (H>>1)-1).
[0259] (d) Assuming the top-left position within the block is (0, 0), then the center position of the block is defined as ((W>>1)-1, (H>>1)-1).
[0260] b. In one example, to derive BV, only B n The top left (e.g., Figure 25 The LT position in the text can be checked.
[0261] c. In one example, to derive BV, B n The top left (e.g., Figure 25 LT in the middle), top right (e.g., Figure 25 RT in the middle), center (e.g., Figure 25 (Ctr in the middle), bottom left (e.g., Figure 25 (LB in the middle) and bottom right (e.g., Figure 25 At least one of the RB positions in the data can be checked.
[0262] (a) In one example, all available BVs can be used as BV candidates for block vector guidance (e.g., BV... n,n+1 or BV 0,n+1 Derivation.
[0263] (b) In one example, only the first available BV can be used as a block vector-guided BV candidate (e.g., BV...). n,n+1 or BV 0,n+1 Derivation.
[0264] d. In one example, to derive BV, B n The top left (e.g., Figure 25 LT in the middle), top right (e.g., Figure 25 RT in the middle), center (e.g., Figure 25 (Ctr in the middle), bottom left (e.g., Figure 25 (LB in the middle) and bottom right (e.g., Figure 25 All five positions of the RB position can be checked.
[0265] (a) In one example, the inspection order is center -> top left -> top right -> bottom left -> bottom right.
[0266] (b) In one example, all available BVs can be used as BV candidates for block vector guidance (e.g., BV... n,n+1 or BV 0,n+1 Derivation.
[0267] (c) In one example, only the first available BV can be used as a block vector-guided BV candidate (e.g., BV...).n,n+1 or BV 0,n+1 Derivation.
[0268] e. In one example, to derive BV, B n Some predefined locations can be checked.
[0269] (a) In one example, all available BVs can be used as BV candidates for block vector guidance (e.g., BV... n,n+1 or BV 0,n+1 Derivation.
[0270] (b) In one example, only the first available BV can be used as a block vector-guided BV candidate (e.g., BV...). n,n+1 or BV 0,n+1 Derivation.
[0271] f. In one example, if B is covered n If a motion mesh (such as a 4x4 mesh) is available at a specific location and has BV information, it can be used for block vector-guided BV candidates (e.g., BV...). n,n+1 or BV 0,n+1 Derivation.
[0272] 5. In one example, there may be a constraint on the maximum number (e.g., N) of BV candidates guided by the block vector.
[0273] a. Alternatively, there may be no constraint on the maximum number (e.g., N) of BV candidates guided by the block vector.
[0274] b. In one example, the number of BV candidates guided by the block vector may not exceed 5.
[0275] c. In one example, the number of BV candidates guided by the block vector may not exceed 4.
[0276] d. Alternatively, there may be a constraint on the maximum number (e.g., M) of unique (e.g., after complete deduplication) block vector-guided BV candidates to be derived.
[0277] (a) In one example, M can be 5.
[0278] (b) In one example, M can be 4.
[0279] (c) In one example, M can be 3.
[0280] e. In one example, when deriving BV candidates guided by block vectors, redundancy checks or deduplication can be performed.
[0281] (a) In one example, when deriving block vector-guided BV candidates, full deduplication can be performed to ensure that candidates with the same or similar motion information are excluded from the BV candidate list.
[0282] (b) In one example, partial deduplication can be performed when deducing BV candidates guided by block vectors.
[0283] 6. In one example, the position of the block vector-guided BV candidate in the BV candidate list can be one of the following.
[0284] a. In one example, all block vector-guided BV candidates can be inserted before the HMVP BV candidates.
[0285] b. In one example, a portion of the block vector-guided BV candidate can be inserted before the HMVP BV candidate, and the remaining block vector-guided BV candidate can be inserted after the HMVP BV candidate.
[0286] c. In one example, all block vector-guided BV candidates can be inserted after the HMVP BV candidates.
[0287] d. In one example, all block vector-guided BV candidates can be inserted before non-adjacent spatial BV candidates.
[0288] e. In one example, a partial block vector-guided BV candidate can be inserted before a non-adjacent spatial BV candidate, and the remaining block vector-guided BV candidates can be inserted after a non-adjacent spatial BV candidate.
[0289] f. In one example, all block vector-guided BV candidates can be inserted after non-adjacent spatial BV candidates.
[0290] 7. In one example, whether to use block vector-guided BV prediction can be signaled at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0291] BV Candidate Reordering 8. In one example, a reordering / refinement process can be performed when deriving the BV candidate list.
[0292] a. In one example, the reordering / refinement process can be based on template matching costs (multiple).
[0293] b. In one example, when constructing the BV candidate list, N1 adjacent spatial candidates and / or N2 temporal candidates and / or N3 non-adjacent spatial candidates and / or N4 HMVP candidates and / or N5 block vector-guided BV candidates and / or N6 pairwise average candidates and / or N7 predefined BV candidates can be derived partially or entirely using complete deduplication to ensure that there are no duplicate or similar candidates in the list, and then they are reordered together. After reordering, the top N candidates (such as those with the lowest cost) can be selected as the final candidates in the BV candidate list.
[0294] i. In one example, N can be 6 and / or N1 can be 5 and / or N2 can be 10 and / or N3 can be 35 and / or N4 can be 25 and / or N5 can be 10 and / or N6 can be 1 and / or N7 can be 6.
[0295] ii. In one example, there may be a constraint on the maximum number (e.g., M) of unique (e.g., after full deduplication) BV candidates to be derived.
[0296] (i) In one example, M can be 20.
[0297] (ii) In one example, M can be 28.
[0298] iii. In one example, adjacent airspace BV candidates can consist of airspace candidates to the left and / or above and / or upper right and / or lower left and / or upper left (example in...). Figure 6 (As shown in the image).
[0299] iv. In one example, the time-domain BV candidate can consist of those specified in item 3.
[0300] v. In one example, non-adjacent spatial domain BV candidates can be determined by... Figure 22 The components specified in the document.
[0301] vi. In one example, the number of HMVP BV candidates and / or the HMVP table size can be increased to N2 (e.g., 25).
[0302] vii. In one example, such as Figure 23 As shown, block vector-guided BV candidates can be block vector-guided cascaded BV candidates.
[0303] viii. In one example, such as Figure 23 As shown, block vector-guided BV candidates can be block vector-guided direct BV candidates.
[0304] ix. In one example, a pair of BV candidates can be generated by averaging predefined pairs of existing candidates in the motion candidate list.
[0305] (iii) In one example, a predefined pair can be defined as a pair in a set such as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the motion candidate indices in the motion candidate list.
[0306] x. In one example, a predefined BV candidate can be located in the IBC reference region.
[0307] c. In one example, ARMC based on BV candidate types can be used to reorder BV candidates with one or more specific candidate types, according to one or more criteria.
[0308] i. In one example, when constructing the BV candidate list, M candidates with a specific candidate type (such as having the lowest cost) can be selected from N reordered candidates with candidate types.
[0309] (i) In one example, M can vary depending on the candidate type and / or encoding / decoding mode of the current block.
[0310] (ii) In one example, the candidate type can be adjacent spatial BV candidates. For example, M is 4 and N is 5.
[0311] (iii) In one example, the candidate type can be a non-adjacent spatial BV candidate. For example, M is 6 and N is 35.
[0312] (iv) In one example, the candidate type can be a time-domain BV candidate. For example, M is 4 and N is 10.
[0313] (v) In one example, the candidate type can be an HMVP BV candidate. For example, M is 10 and N is 25.
[0314] (i) In one example, the candidate type can be a block vector-guided BV candidate. For example, M is 4 and N is 10.
[0315] (ii) In one example, the candidate type can be a pairwise average BV candidate. For example, M is 1 and N is 6.
[0316] (iii) In one example, the candidate type can be a predefined BV candidate. For example, M is 1 and N is 6.
[0317] ii. In one example, multiple BV candidate types (i.e., combinations of candidate types) can be reordered together.
[0318] (i) In one example, when constructing the BV candidate list, M candidates with any particular BV candidate type (such as having the lowest cost) can be selected from N reordered candidates that conform to the combination of candidate types, where M can vary depending on the combination of candidate types and / or the encoding / decoding mode of the current block.
[0319] (ii) In one example, adjacent spatial candidates and / or temporal candidates and / or non-adjacent spatial candidates and / or HMVP candidates and / or block vector-guided candidates and / or pairwise averaged candidates and / or predefined BV candidates can be reordered together. For example, M is 6 and N is 20.
[0320] (iii) In one example, at least one BV candidate type can be reordered using ARMC based on the BV candidate type.
[0321] (iv) In one example, N1 HMVP candidates (such as those with the lowest cost) can be selected from candidates that have been reordered with HMVP candidate type, and the selected N1 HMVP candidates can be reordered together with adjacent spatial candidates and / or temporal candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates. Finally, M candidates (such as those with the lowest cost) can be selected.
[0322] (v) In one example, N2 temporal candidates (such as those with the lowest cost) can be selected from candidates that have been reordered with temporal candidate types, and the selected N2 temporal candidates can be reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates. Finally, M candidates (such as those with the lowest cost) can be selected.
[0323] (vi) In one example, N3 block vector-guided candidates (such as those with the lowest cost) can be selected from candidates reordered with block vector-guided candidate types, and the selected N3 block vector-guided candidates can be reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or temporal candidates and / or pairwise averaged candidates and / or predefined BV candidates. Finally, M candidates (such as those with the lowest cost) can be selected.
[0324] (vii) In one example, if a candidate is reordered more than once, its reordering criteria (e.g., template matching cost) can be reused.
[0325] Two-way forecasting / multiple hypothesis IBC forecasting 9. In one example, block vector-guided BVP can be used for bidirectional forecasting / multiple hypothesis IBC forecasting.
[0326] a. In one example, at least one prediction can be derived from a common IBC prediction, and at least one prediction can be derived from a block vector-guided BVP.
[0327] (a) In one example, only one index can be transmitted from the encoder to the decoder via signaling to indicate IBC Merge / AMVP BV information for common IBC prediction.
[0328] (b) In one example, such as Figure 23 As shown, the block vector-guided BVP can be derived from the indicated IBC Merge / AMVP BV information.
[0329] (c) In one example, common IBC predictions can be derived from at least one of the following.
[0330] 1) In one example, common IBC predictions can be derived from regular IBC Merge predictions.
[0331] 2) In one example, common IBC predictions can be derived from regular IBC AMVP predictions.
[0332] 3) In one example, common IBC predictions can be derived from IBC-TM Merge predictions.
[0333] 4) In one example, common IBC predictions can be derived from IBC-TM AMVP predictions.
[0334] 5) In one example, common IBC predictions can be derived from RR-IBC Merge predictions.
[0335] 6) In one example, common IBC predictions can be derived from RR-IBC AMVP predictions.
[0336] 7) In one example, common IBC predictions can be derived from IBC-MBVD predictions.
[0337] 8) In one example, common IBC predictions can be derived from IBC-CIIP predictions.
[0338] 9) In one example, common IBC predictions can be derived from IBC-GPM predictions.
[0339] 10) In one example, common IBC predictions can be derived from BV predictions of encoder only.
[0340] 11) In one example, common IBC predictions can be derived from string copy vector predictions.
[0341] 12) In one example, the common IBC prediction can be derived from any other BV prediction.
[0342] 10. In one example, whether bidirectional prediction / multiple hypothesis IBC prediction is used can be signaled at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0343] General information 11. In all publicly available methods, the term “block vector (BV)” can be replaced with “motion vector (MV)”.
[0344] a. The BVP list can be replaced by the MVP list.
[0345] b. The term IBC prediction can be replaced by inter-frame prediction.
[0346] c. The term IBC Merge list can be replaced by Merge list.
[0347] d. The term IBC AMVP list can be replaced by the AMVP list.
[0348] e. “Block vector guided” can be replaced by “MV guided”.
[0349] f. In one example, the MV-guided MV prediction or candidate could be... n A cascaded MVP or candidate guided by an order (e.g., n≥1) motion vector is denoted as MV. 0,n+1 .
[0350] (a) MV 0,n+1 =MV 0,1 +MV 1,2 +MV 2,3 +…+MV n-1,n +MV n,n+1 .
[0351] (b) For example, MV 0,n+1 and MV n,n+1It can refer to the same reference image.
[0352] (c) For example, MV i,i+1 and MV j,j+1 It can refer to different reference images.
[0353] g. In one example, the MV-guided MV prediction or candidate could be... n A direct MVP or candidate guided by an order (e.g., n≥1) motion vector is denoted as MV. n,n+1 .
[0354] 12. The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. These can be signed or unsigned.
[0355] 13. The grammatical elements disclosed above can be encoded and decoded using at least one context model. Alternatively, they can be encoded and decoded using a bypass method.
[0356] 14. The syntax elements (SE) disclosed above can be transmitted conditionally via signals.
[0357] a. SE is transmitted via signal only if the corresponding function applies.
[0358] 15. The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0359] 16. In the above examples, a block can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU row / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample point or pixel.
[0360] 17. Whether and / or how the methods disclosed above can be signaled at the sequence level / picture group level / picture level / strip level / film group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / film group header.
[0361] 18. Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0362] 19. Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0363] Reference Figure 26 Another embodiment is described below. Figure 26 A flowchart of a method 2600 for video processing according to an embodiment of the present disclosure is shown. Method 2600 is implemented during the conversion between video units or video units of a video and a bitstream of a video.
[0364] At box 2610, for the conversion between the current video block and the video bitstream, a first guide vector associated with the current video block is determined. The first guide vector includes a first guide block vector (BV) or a first guide motion vector (MV).
[0365] At box 2620, a first reference vector for a first reference block located based on a first guide vector is determined. The first reference vector includes either a first reference block vector or a first reference motion vector. For example, as... Figure 23 The first reference block of B0 in the B0 is based on, for example, BV. 0,1 The first guide block (BV) is located. The first reference block vector (BV) is... 1,2 It is the BV of the first reference block. As used in this paper, the first reference vector is referred to as the direct BV guided by the first-order BV.
[0366] At box 2630, the vector prediction of the current video block is determined based on at least the first guide vector and the first reference vector. Vector prediction includes block vector prediction or motion vector prediction.
[0367] At box 2640, the transformation based on vector prediction is performed.
[0368] Method 2600 enables the application of block vector guided BVP or motion vector guided MVP for conversion. In this way, encoding / decoding efficiency and / or encoding / decoding effectiveness can be improved.
[0369] In some embodiments, block vector prediction includes at least one of the following: regular intra-block copy (IBC) merge prediction, regular IBC advanced motion vector prediction (AMVP) prediction, IBC with template matching (IBC-TM) merge prediction, IBC-TM AMVP prediction, reconstruction reordering IBC (RR-IBC) merge prediction, RR-IBC AMVP prediction, IBC Merge mode with block vector difference (IBC-MBVD) prediction, combination of intra-block copy and intra-prediction (IBC-CIIP), IBC with geometric segmentation (IBC-GPM) prediction, encoder-only BV prediction, string copy vector prediction, intra-template matching prediction (IntraTMP) merge prediction, or further BV prediction. Motion vector prediction includes at least one of the following: regular inter-frame merge prediction, regular inter-frame advanced motion vector prediction (AMVP) prediction, inter-frame template matching (TM) merge prediction, or inter-frame-TM AMVP prediction, inter-frame merge mode with motion vector difference (MMVD), intra-frame inter-frame joint prediction (CIIP), geometric segmentation (GPM) prediction, encoder-only MV prediction, or further MV prediction.
[0370] In some embodiments, determining the vector prediction of the current video block includes: determining the sum of a first guide block vector and a first reference block vector as the block vector prediction. For example, the block vector prediction may be... Figure 23 BV in 0,2 It is BV 0,1 and BV 1,2 The sum. As used in this article, BV 0,2 This is known as a first-order BV-guided cascaded BVP.
[0371] In some embodiments, vector predictions are candidates in a list of vector candidates, which may include a list of block vector candidates or a list of motion vector candidates.
[0372] In some embodiments, the vector candidate list includes at least one of the following: a regular intra-block copy (IBC) merge list, a regular inter-frame merge list, a regular IBC advanced motion vector prediction (AMVP) list, a regular inter-frame AMVP list, an IBC with template matching (IBC-TM) merge list, an inter-frame template matching (TM) merge list, an IBC-TM AMVP list, an inter-frame-TM AMVP list, a reconstruction reordering IBC (RR-IBC) merge list, an RR-IBC AMVP list, a basic candidate list for IBC merge mode with block vector difference (IBC-MBVD), a basic candidate list for inter-frame merge mode with motion vector difference (MMVD), a combined intra-block copy and intra-prediction (IBC-CIIP) merge list, an intra-inter-frame joint prediction (CIIP) merge list, an IBC-CIIP AMVP list, a CIIP AMVP list, an IBC with geometric segmentation (IBC-GPM) merge list, a GPM merge list, and an IBC-GPM merge list. AMVP list, GPM AMVP list, IntraTMP Merge list, encoder-only BV candidate list, encoder-only MV candidate list, further BV candidate list, or further MV candidate list.
[0373] In some embodiments, vector prediction includes at least one vector-guided cascaded prediction or candidate, and the at least one vector-guided cascaded prediction or candidate includes a first-order vector-guided cascaded prediction or candidate. The first-order vector-guided cascaded prediction or candidate may be the sum of a first guiding vector and a first reference vector of a first reference block, the first guiding vector pointing to the first reference block, and the first reference vector pointing to a second reference block.
[0374] In some embodiments, the first reference block has been or partially reconstructed within the current image.
[0375] In some embodiments, the sum of the first guiding vector and the first reference vector is valid for the current video block.
[0376] In some embodiments, such as Figure 23 As shown, at least one vector-guided cascaded prediction or candidate includes an n-order vector-guided cascaded prediction or candidate BV. 0,n+1 N is an integer greater than or equal to 1. BV 0,n+1 It can be determined by the following formula: BV 0,n+1 = BV 0,1 +BV 1,2 +…+BV n-1,n +BV n,n+1 BV 0,1BV represents the first guiding vector. 1,2 Let BV represent the first reference vector (also known as the direct BV candidate guided by first-order BV), ..., BV n-1,n Let (n-1)th reference vector (also known as the (n-1)th order BV-guided direct BV candidate), and BV n,n+1 Let BV represent the nth reference vector (also known as the nth-order BV-guided direct BV candidate), where BV 1,2 From the first reference block to the second reference block, BV n-1,n Pointing from the (n-1)th reference block to the nth reference block, and BV n,n+1 Point from the nth reference block to the (n+1)th reference block.
[0377] In some embodiments, vector prediction includes at least one vector-guided direct prediction or candidate, the at least one vector-guided direct prediction or candidate including a first-order vector-guided direct prediction or candidate. A first guiding vector points to a first reference block, and the first-order vector-guided direct prediction or candidate is a first reference vector pointing from the first reference block to a second reference block.
[0378] In some embodiments, the first reference block has been or partially reconstructed within the current image.
[0379] In some embodiments, first-order vector-guided direct predictions or candidates are valid for the current video block.
[0380] In some embodiments, at least one vector-guided direct prediction or candidate includes an n-order vector-guided direct prediction or candidate BV. n,n+1 n is an integer greater than or equal to 1. BV n,n+1 Point from the nth reference block to the (n+1)th reference block.
[0381] In some embodiments, BV is determined based on the nth reference block. n,n+1 This includes: examining at least one position associated with the nth reference block to derive a vector.
[0382] In some embodiments, at least one position includes a center position, wherein assuming the top-left position within a block of width W and height H is (0, 0), the center position of the block is represented by one of the following: (W>>1, H>>1), ((W>>1), H>>1), ((W>>1), (H>>1)-1), or ((W>>1)-1, (H>>1)-1).
[0383] In some embodiments, at least one position includes the top-left position of the nth reference block.
[0384] In some embodiments, at least one position includes at least one of the following: upper left position, upper right position, center position, lower left position, lower right position, or some predefined positions.
[0385] In some embodiments, the inspection order is center position, top left position, top right position, bottom left position, bottom right position.
[0386] In some embodiments, all available vectors are used to derive vector-guided vector candidates, or a first available vector is used to derive vector-guided vector candidates.
[0387] In some embodiments, checking at least one location includes: checking a motion mesh covering the location of the nth reference block; and if it is determined that the motion mesh is available and has BV information, the BV information is used for block vector-guided BV candidate derivation.
[0388] In some embodiments, the first guide vector associated with the current video block is determined according to at least one of the following: an existing vector candidate in a vector candidate list prior to determining vector prediction, or at least one predefined vector candidate in a vector candidate list prior to determining vector prediction.
[0389] In some embodiments, existing vector candidates include vector-guided vector candidates, and vector-guided vector candidates include at least one of the following: vector-guided cascaded candidates or vector-guided direct candidates.
[0390] In some embodiments, at least one predefined vector candidate includes the first N vector candidates or the selected N vector candidates, where N is a positive integer, or at least one predefined vector candidate includes vector candidates of at least one candidate type.
[0391] In some embodiments, at least one candidate type includes at least one of the following: adjacent spatial domain BV candidate, non-adjacent spatial domain BV candidate, history-based motion vector prediction (HMVP) BV candidate, adjacent temporal domain BV candidate, non-adjacent temporal domain BV candidate, pairwise averaged BV candidate, or block vector guided BV candidate.
[0392] In some embodiments, first-order vector-guided cascaded vector candidates or first-order vector-guided direct vector candidates are used for transformation.
[0393] In some embodiments, first-order vector-guided cascaded or direct vector candidates and second-order vector-guided cascaded or direct vector candidates are used for transformation.
[0394] In some embodiments, cascaded or direct vector candidates guided by the first N order vectors are used for transformation, where N is a positive integer.
[0395] In some embodiments, there is no limit to the number of vector-guided vector candidates.
[0396] In some embodiments, the number of vector-guided vector candidates is less than or equal to the maximum number. For example, the maximum number may be 5 or 4.
[0397] In some embodiments, the number of unique vector candidates after complete deduplication is less than or equal to a maximum number, which is 5, 4 or 3.
[0398] In some embodiments, when deriving vector-guided vector candidates, redundancy checks or deduplication are performed, and the vector candidates include block vector candidates or motion vector candidates.
[0399] In some embodiments, when deriving vector-guided vector candidates, full deduplication is performed to ensure that candidates with the same or similar motion information are excluded from the vector candidate list.
[0400] In some embodiments, partial deduplication is performed when deriving vector-guided vector candidates.
[0401] In some embodiments, the position of block vector-guided BV candidates in the BV candidate list is based on inserting block vector-guided BV candidates after historical motion vector prediction (HMVP) BV candidates.
[0402] In some embodiments, the position of a block vector-guided BV candidate in the BV candidate list is based on inserting a block vector-guided BV candidate after a non-adjacent spatial BV candidate.
[0403] In some embodiments, the position of block vector-guided BV candidates in the BV candidate list is based on one of the following: inserting a block vector-guided BV candidate before a history-based motion vector prediction (HMVP) BV candidate; inserting a portion of a block vector-guided BV candidate before an HMVP BV candidate and inserting the remaining block vector-guided BV candidate after an HMVP BV candidate; inserting a block vector-guided BV candidate before a non-adjacent spatial BV candidate; or inserting a portion of a block vector-guided BV candidate before a non-adjacent spatial BV candidate and inserting the remaining block vector-guided BV candidate after a non-adjacent spatial BV candidate.
[0404] In some embodiments, whether to use vector-guided vector prediction is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0405] In some embodiments, method 2600 further includes: applying a reordering or refinement process to derive a vector candidate list, the vector candidate list including a BV candidate list or an MV candidate list.
[0406] In some embodiments, the reordering or refinement process is based on template matching cost.
[0407] In some embodiments, the BV candidate list is constructed by: at least partially deriving, through complete deduplication of at least one of the following candidates: N1 adjacent spatial candidates, N2 temporal candidates, N3 non-adjacent spatial candidates, N4 HMVP candidates, N5 block vector-guided BV candidates, N6 pairwise averaged candidates, or N7 predefined BV candidates, where N1, N2, N3, N4, N5, N6, and N7 are positive integers; reordering the derived candidates; and selecting the top N candidates with the lowest cost as the final candidates in the BV candidate list, where N is a positive integer.
[0408] In some embodiments, N is 6, N1 is 5, N2 is 10, N3 is 35, N4 is 25, N5 is 10, N6 is 1, and N7 is 6.
[0409] In some embodiments, the number of unique BV candidates after complete deduplication is less than or equal to the maximum number.
[0410] In some embodiments, the maximum number is 20 or 28.
[0411] In some embodiments, adjacent airspace candidates consist of at least one of the following: left airspace candidate, upper airspace candidate, upper right airspace candidate, lower left airspace candidate, or upper left airspace candidate.
[0412] In some embodiments, time-domain BV candidates consist of adjacent time-domain BV candidates or non-adjacent time-domain BV candidates.
[0413] In some embodiments, non-adjacent spatial domain candidates consist of non-adjacent spatial domain candidates between regular frames.
[0414] In some embodiments, the number of HMVP BV candidates and / or the HMVP table size is increased to 25.
[0415] In some embodiments, block vector guided BV candidates include at least one block vector guided cascade or block vector guided direct BV candidate.
[0416] In some embodiments, paired BV candidates are generated by averaging the candidates of predefined pairs in a motion candidate list, the predefined pairs being in a set denoted as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where 0, 1, 2, and 3 represent motion candidate indices in the motion candidate list.
[0417] In some embodiments, predefined BV candidates are located in the intra-block copy reference region.
[0418] In some embodiments, the derived candidate reordering includes applying adaptive reordering of Merge candidates based on BV candidate type (ARMC) to reorder BV candidates with at least one candidate type based on a standard.
[0419] In some embodiments, M candidates with the lowest cost and candidate type are selected from N reordered candidates with that candidate type, where M and N are positive integers.
[0420] In some embodiments, M is based on the candidate type and / or encoding / decoding mode of the current video block.
[0421] In some embodiments, the candidate type includes adjacent spatial BV candidates, M is 4, and N is 5.
[0422] In some embodiments, the candidate type includes non-adjacent spatial BV candidates, where M is 6 and N is 35.
[0423] In some embodiments, the candidate type includes time-domain BV candidates, where M is 4 and N is 10.
[0424] In some embodiments, the candidate type includes HMVP BV candidates, M is 10, and N is 25.
[0425] In some embodiments, the candidate type includes block vector guided BV candidates, where M is 4 and N is 10.
[0426] In some embodiments, the candidate type includes pairwise averaged BV candidates, where M is 1 and N is 6.
[0427] In some embodiments, the candidate type includes predefined BV candidates, where M is 1 and N is 6.
[0428] In some embodiments, the derived candidate reordering includes reordering BV candidates of multiple BV candidate types together.
[0429] In some embodiments, when constructing the BV candidate list, M candidates with the lowest cost having at least one of a plurality of BV candidate types are selected from N reordered candidates of the plurality of BV candidate types, where M is a combination of the plurality of BV candidate types and / or the encoding / decoding mode of the current video block.
[0430] In some embodiments, the candidates for multiple BV candidate types include at least one of the following: adjacent spatial candidates, temporal candidates, non-adjacent spatial candidates, HMVP BV candidates, block vector guided candidates, pairwise averaged candidates, or predefined BV candidates, where M is 6 and N is 20.
[0431] In some embodiments, BV candidates of at least one BV candidate type are first reordered using ARMC based on the BV candidate type.
[0432] In some embodiments, N1 HMVP candidates are selected from candidates that have been reordered with HMVP candidate type, and the selected N1 HMVP candidates are reordered together with adjacent spatial candidates and / or temporal candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N1 is a positive integer.
[0433] In some embodiments, N2 temporal candidates are selected from candidates that have been reordered with temporal candidate types, and the selected N2 temporal candidates are reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N2 is a positive integer.
[0434] In some embodiments, N3 block vector-guided candidates are selected from candidates that have been reordered with block vector-guided candidate types, and the selected N3 block vector-guided candidates are reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or temporal candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N3 is a positive integer.
[0435] In some embodiments, if a candidate is reordered more than once, the reordering criteria for the candidate are reused, and the reordering criteria include template matching cost.
[0436] In some embodiments, block vector guided block vector prediction is used for at least one of the following: bidirectional prediction intra-block copy (IBC) prediction or multi-hypothesis IBC prediction.
[0437] In some embodiments, at least one prediction is derived from a common IBC prediction, and at least one prediction is derived from a block vector prediction guided by a block vector.
[0438] In some embodiments, an index is directed from the encoder to the decoder to indicate IBC Merge or AMVP BV information for common IBC predictions.
[0439] In some embodiments, block vector-guided block vector predictions are derived from indicated IBC Merge or AMVP BV information.
[0440] In some embodiments, the common IBC prediction is derived from at least one of the following: regular IBC Merge prediction, regular IBC AMVP prediction, IBC-TM Merge prediction, IBC-TM AMVP prediction, RR-IBC Merge prediction, RR-IBC AMVP prediction, IBC-MBVD prediction, IBC-CIIP prediction, IBC-GPM prediction, encoder-only BV prediction, string copy vector prediction, intra-template matching prediction (IntraTMP) Merge prediction, or further BV prediction.
[0441] In some embodiments, whether to use at least one of bidirectional prediction IBC prediction or multi-hypothesis IBC prediction is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0442] In some embodiments, BV is replaced by MV, BVP list is replaced by MVP list, IBC prediction is replaced by inter-frame prediction, IBC Merge list is replaced by Merge list, IBC AMVP list is replaced by AMVP list, and block vector-guided vector is replaced by MV-guided vector.
[0443] In some embodiments, MV-guided MV predictions or candidates include those represented as MV. 0,n+1 The cascaded MVP or candidate guided by the nth-order motion vector, and the MV 0,n+1 This is derived from the following: MV 0,n+1 =MV 0,1 +MV 1,2 +…+MV n-1,n +MV n,n+1 MV 0,1 MV represents the first guiding motion vector. 1,2 MV represents the first reference motion vector. n-1,n Let represent the (n-1)th reference motion vector, which points to the nth reference block in the first reference image, and MVn,n+1 This represents the nth reference motion vector, which points to the (n+1)th reference block in the second reference image.
[0444] In another example embodiment, BV is also used to derive MVP or MV candidates. For example, MV can be derived by adding MV. 0,1 , by MV 0,1 BV and MV of the positioning reference block 1,2 …、MV n-1,n and MV n,n+1 It has been confirmed.
[0445] In some embodiments, MV 0,n+1 and MV n,n+1 Refers to the same reference image. MV i,i+1 and MV j,j+1 Refers to different reference images.
[0446] In some embodiments, MV-guided MV predictions or candidates include those represented as MV. n,n+1 The nth-order motion vector guides the direct MVP or candidate. n,n+1 Point from the nth reference block in the first reference image to the (n+1)th reference block in the second reference image.
[0447] In some embodiments, syntax elements in the bitstream are binarized into one of the following: flags, fixed-length codes, exponential Golomb (x) (EG(x)) codes, unary codes, rounded unary codes, or rounded binary codes, and the syntax elements are signed or unsigned.
[0448] In some embodiments, syntax elements in the bitstream are encoded or decoded using at least one context model or are bypassed.
[0449] In some embodiments, a syntax element is included in the bitstream based on at least one condition, the at least one condition including a condition for a function associated with the syntax element to be applied to the transformation.
[0450] In some embodiments, syntax elements are included at one of the following: block level, sequence level, picture group level, picture level, strip level, slice group level, or codec structure, wherein the codec structure includes one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0451] In some embodiments, the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of a block, sub-region within a block, or region containing more than one sample point or pixel.
[0452] In some embodiments, whether and / or how method 2600 is applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0453] In some embodiments, whether and / or how method 2600 is applied includes one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.
[0454] In some embodiments, whether and / or how method 2600 is applied is based on encoded information, which includes at least one of the following: block size, color format, single-tree or dual-tree segmentation, color components, stripe type or picture type.
[0455] In some embodiments, the conversion includes encoding the current video block into a bitstream.
[0456] In some embodiments, the conversion includes decoding the current video block from the bitstream.
[0457] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores video data generated as a bitstream by a method performed by an apparatus for video processing. In this method, a first guide vector associated with a current video block is determined, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV). A first reference vector for a first reference block located based on the first guide vector is determined, the first reference vector including a first reference block vector or a first reference motion vector. Vector prediction of the current video block is determined at least based on the first guide vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction. The bitstream is generated based on the vector prediction.
[0458] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, a first guide vector associated with a current video block of the video is determined, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV). A first reference vector for a first reference block located based on the first guide vector is determined, the first reference vector including a first reference block vector or a first reference motion vector. Vector prediction of the current video block is determined based at least on the first guide vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction. A bitstream is generated based on the vector prediction. The bitstream is stored in a non-transitory computer-readable recording medium.
[0459] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0460] Item 1. A method for video processing, comprising: for a conversion between a current video block and a bitstream of the video, determining a first guiding vector associated with the current video block, the first guiding vector including a first guiding block vector (BV) or a first guiding motion vector (MV); determining a first reference vector for a first reference block located based on the first guiding vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction of the current video block based at least on the first guiding vector and the first reference vector, the vector prediction including a block vector prediction or a motion vector prediction; and performing the conversion based on the vector prediction.
[0461] Item 2. The method according to Item 1, wherein the block vector prediction includes at least one of the following: regular intra-frame block copy (IBC) merge prediction, regular IBC advanced motion vector prediction (AMVP) prediction, IBC with template matching (IBC-TM) merge prediction, IBC-TM AMVP prediction, reconstruction reordering IBC (RR-IBC) merge prediction, RR-IBCAMVP prediction, IBC merge mode with block vector difference (IBC-MBVD) prediction, combination of intra-frame block copy and intra-frame prediction (IBC-CIIP), IBC with geometric segmentation (IBC-GPM) prediction, encoder-only BV prediction, string copy vector prediction, intra-frame template matching prediction (IntraTMP) merge prediction, or further BV prediction, and wherein the motion vector prediction includes at least one of the following: regular inter-frame merge prediction, regular inter-frame advanced motion vector prediction (AMVP) prediction, inter-frame template matching (TM) merge prediction, or inter-frame-TM AMVP prediction, inter-frame merge mode with motion vector difference (MMVD), intra-frame joint prediction (CIIP), geometric segmentation (GPM) prediction, encoder-only MV prediction, or further MV prediction.
[0462] Item 3. The method according to Item 1 or 2, wherein determining the vector prediction of the current video block comprises: determining the sum of the first guide block vector and the first reference block vector as the block vector prediction.
[0463] Item 4. The method according to any one of items 1 to 3, wherein the vector prediction is a candidate in a vector candidate list, the vector candidate list including a block vector candidate list or a motion vector candidate list.
[0464] Item 5. The method according to Item 4, wherein the vector candidate list includes at least one of the following: a regular intra-block copy (IBC) merge list, a regular inter-frame merge list, a regular IBC advanced motion vector prediction (AMVP) list, a regular inter-frame AMVP list, an IBC with template matching (IBC-TM) merge list, an inter-frame template matching (TM) merge list, an IBC-TM AMVP list, an inter-frame-TM AMVP list, a reconstruction reordering IBC (RR-IBC) merge list, an RR-IBCAMVP list, a basic candidate list of IBC merge mode with block vector difference (IBC-MBVD), a basic candidate list of inter-frame merge mode with motion vector difference (MMVD), a combined intra-block copy and intra-prediction (IBC-CIIP) merge list, an intra-inter-frame joint prediction (CIIP) merge list, an IBC-CIIP AMVP list, a CIIP AMVP list, an IBC with geometric segmentation (IBC-GPM) merge list, and a GPM. Merge list, IBC-GPM AMVP list, GPM AMVP list, IntraTMP Merge list, encoder-only BV candidate list, encoder-only MV candidate list, further BV candidate list, or further MV candidate list.
[0465] Item 6. The method according to any one of items 1 to 5, wherein the vector prediction includes at least one vector-guided cascaded prediction or candidate, the at least one vector-guided cascaded prediction or candidate including a first-order vector-guided cascaded prediction or candidate, wherein the first-order vector-guided cascaded prediction or candidate is the sum of the first guide vector and the first reference vector of the first reference block, the first guide vector pointing to the first reference block, and the first reference vector pointing to the second reference block.
[0466] Item 7. The method described in Item 6, wherein the first reference block has been or partially reconstructed within the current image.
[0467] Item 8. The method according to Item 6, wherein the sum of the first guiding vector and the first reference vector is valid for the current video block.
[0468] Item 9. The method according to Item 6, wherein the at least one vector-guided cascaded prediction or candidate includes an n-order vector-guided cascaded prediction or candidate BV. 0,n+1 n is an integer greater than or equal to 1, where BV 0,n+1 It was determined by the following: BV 0,n+1 = BV 0,1 +BV1,2 +…+BV n-1,n +BV n,n+1 BV 0,1 BV represents the first guiding vector. 1,2 Let BV represent the first reference vector, ..., BV. n-1,n Let BV represent the (n-1)th reference vector. n,n+1 Let BV represent the nth reference vector. 1,2 From the first reference block to the second reference block, BV n-1,n Pointing from the (n-1)th reference block to the nth reference block, and BV n,n+1 Point from the nth reference block to the (n+1)th reference block.
[0469] Item 10. The method according to any one of items 1 to 5, wherein the vector prediction includes at least one vector-guided direct prediction or candidate, the at least one vector-guided direct prediction or candidate including a first-order vector-guided direct prediction or candidate, wherein the first guiding vector points to the first reference block, and the first-order vector-guided direct prediction or candidate is the first reference vector pointing from the first reference block to the second reference block.
[0470] Item 11. The method according to Item 10, wherein the first reference block has been or partially reconstructed within the current image.
[0471] Item 12. The method according to Item 10, wherein the first-order vector-guided direct prediction or candidate is valid for the current video block.
[0472] Item 13. The method according to Item 10, wherein the at least one vector-guided direct prediction or candidate includes an n-order vector-guided direct prediction or candidate BV. n,n+1 n is an integer greater than or equal to 1, where BV n,n+1 Point from the nth reference block to the (n+1)th reference block.
[0473] Item 14. The method according to Item 13, wherein BV is determined based on the nth reference block. n,n+1 This includes: examining at least one position associated with the nth reference block to derive a vector.
[0474] Item 15. The method according to Item 14, wherein the at least one position includes a center position, wherein assuming the top-left position within a block of width W and height H is (0, 0), the center position of the block is represented by one of the following: (W>>1, H>>1), ((W>>1)-1, H>>1), ((W>>1), (H>>1)-1), or ((W>>1)-1, (H>>1)-1).
[0475] Item 16. The method according to Item 14, wherein the at least one position includes the top-left position of the nth reference block.
[0476] Item 17. The method according to Item 14, wherein the at least one position includes at least one of the following: upper left position, upper right position, center position, lower left position, lower right position, or some predefined position.
[0477] Item 18. The method according to Item 17, wherein the inspection order is the center position, the top left position, the top right position, the bottom left position, and the bottom right position.
[0478] Item 19. The method according to Item 15, wherein all available vectors are used to derive vector-guided vector candidates, or wherein a first available vector is used to derive the vector-guided vector candidate.
[0479] Item 20. The method according to any one of items 14 to 19, wherein checking at least one location comprises: checking a motion mesh covering the location of the nth reference block; and if it is determined that the motion mesh is available and has BV information, the BV information is used for block vector-guided BV candidate derivation.
[0480] Item 21. The method according to any one of items 1 to 20, wherein the first guiding vector associated with the current video block is determined according to at least one of the following: an existing vector candidate in a vector candidate list prior to determining the vector prediction, or at least one predefined vector candidate in a vector candidate list prior to determining the vector prediction.
[0481] Item 22. The method according to Item 21, wherein the existing vector candidate includes vector-guided vector candidates, and the vector-guided vector candidate includes at least one of the following: vector-guided cascaded candidates or vector-guided direct candidates.
[0482] Item 23. The method according to Item 21, wherein the at least one predefined vector candidate includes the first N vector candidates or the selected N vector candidates, where N is a positive integer, or wherein the at least one predefined vector candidate includes vector candidates of at least one candidate type.
[0483] Item 24. The method according to Item 23, wherein the at least one candidate type includes at least one of the following: adjacent spatial domain BV candidate, non-adjacent spatial domain BV candidate, history-based motion vector prediction (HMVP) BV candidate, adjacent temporal domain BV candidate, non-adjacent temporal domain BV candidate, pairwise averaged BV candidate, or block vector guided BV candidate.
[0484] Item 25. The method according to any one of items 6 to 24, wherein a first-order vector-guided cascaded vector candidate or a first-order vector-guided direct vector candidate is used for the transformation.
[0485] Item 26. The method according to any one of items 6 to 24, wherein a first-order vector-guided cascaded or direct vector candidate and a second-order vector-guided cascaded or direct vector candidate are used for the transformation.
[0486] Item 27. The method according to any one of items 6 to 24, wherein cascaded or direct vector candidates guided by the first N order vectors are used for the transformation, where N is a positive integer.
[0487] Item 28. The method according to any one of items 1 to 27, wherein there is no limitation on the number of vector candidates guided by the vector.
[0488] Item 29. The method according to any one of items 1 to 27, wherein the number of vector candidates guided by the vector is less than or equal to the maximum number.
[0489] Item 30. The method according to Item 29, wherein the maximum number is 5 or 4.
[0490] Item 31. The method according to Item 29, wherein the number of vector candidates that are unique after complete deduplication is less than or equal to the maximum number, which is 5, 4 or 3.
[0491] Item 32. The method according to any one of items 1 to 31, wherein when deriving vector candidates guided by vectors, redundancy checks or deduplication are performed, said vector candidates including block vector candidates or motion vector candidates.
[0492] Item 33. The method according to Item 32, wherein when deriving the vector-guided vector candidates, complete deduplication is performed to ensure that candidates with the same or similar motion information are excluded from the vector candidate list.
[0493] Item 34. The method according to Item 32, wherein partial deduplication is performed when deriving the vector-guided vector candidate.
[0494] Item 35. The method according to any one of items 1 to 34, wherein the position of the block vector-guided BV candidate in the BV candidate list is based on: inserting the block vector-guided BV candidate after the history-based motion vector prediction (HMVP) BV candidate.
[0495] Item 36. The method according to any one of items 1 to 34, wherein the position of the block vector-guided BV candidate in the BV candidate list is based on: inserting the block vector-guided BV candidate after a non-adjacent spatial BV candidate.
[0496] Item 37. The method according to any one of items 1 to 35, wherein the position of the block vector-guided BV candidates in the BV candidate list is based on one of the following: inserting a block vector-guided BV candidate before a history-based motion vector prediction (HMVP) BV candidate; inserting a portion of a block vector-guided BV candidate before an HMVP BV candidate and inserting the remaining block vector-guided BV candidate after an HMVP BV candidate; inserting a block vector-guided BV candidate before a non-adjacent spatial BV candidate; or inserting a portion of a block vector-guided BV candidate before a non-adjacent spatial BV candidate and inserting the remaining block vector-guided BV candidate after a non-adjacent spatial BV candidate.
[0497] Item 38. The method according to any one of Items 1 to 37, wherein whether vector-guided vector prediction is used is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
[0498] Item 39. The method according to any one of items 1 to 38 further comprises: applying a reordering or refinement process to derive a vector candidate list, said vector candidate list including a BV candidate list or an MV candidate list.
[0499] Item 40. The method according to Item 39, wherein the reordering or refinement process is based on template matching cost.
[0500] Item 41. The method according to Item 39 or 40, wherein the BV candidate list is constructed by: at least partially deriving, through complete deduplication of at least one of the following candidates: N1 adjacent spatial candidates, N2 temporal candidates, N3 non-adjacent spatial candidates, N4 HMVP candidates, N5 block vector-guided BV candidates, N6 pairwise averaged candidates, or N7 predefined BV candidates, where N1, N2, N3, N4, N5, N6, and N7 are positive integers; reordering the derived candidates; and selecting the top N candidates with the lowest cost as the final candidates in the BV candidate list, where N is a positive integer.
[0501] Item 42. The method according to Item 41, wherein N is 6, N1 is 5, N2 is 10, N3 is 35, N4 is 25, N5 is 10, N6 is 1, and N7 is 6.
[0502] Item 43. The method according to Item 41, wherein the number of unique BV candidates after complete deduplication is less than or equal to the maximum number.
[0503] Item 44. The method according to Item 43, wherein the maximum number is 20 or 28.
[0504] Item 45. The method according to Item 41, wherein the adjacent airspace candidates consist of at least one of the following: left airspace candidate, upper airspace candidate, upper right airspace candidate, lower left airspace candidate, or upper left airspace candidate.
[0505] Item 46. The method according to Item 41, wherein the time-domain BV candidate is composed of adjacent time-domain BV candidates or non-adjacent time-domain BV candidates.
[0506] Item 47. The method according to Item 41, wherein the non-adjacent spatial domain candidates consist of non-adjacent spatial domain candidates between regular frames.
[0507] Item 48. The method according to Item 41, wherein the number of HMVP BV candidates and / or the HMVP table size is increased to 25.
[0508] Item 49. The method according to Item 41, wherein the block vector-guided BV candidate includes at least one block vector-guided cascade or block vector-guided direct BV candidate.
[0509] Item 50. The method according to Item 41, wherein paired BV candidates are generated by averaging the candidates of predefined pairs in a motion candidate list, the predefined pairs being in a set denoted as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where 0, 1, 2, and 3 represent motion candidate indices in the motion candidate list.
[0510] Item 51. The method according to Item 41, wherein the predefined BV candidate is located in the intra-block copy reference region.
[0511] Item 52. The method according to any one of items 41 to 51, wherein the derived candidate reordering comprises: applying adaptive reordering of Merge candidates based on BV candidate type (ARMC) to reorder BV candidates having at least one candidate type based on criteria.
[0512] Item 53. The method according to Item 52, wherein M candidates with the lowest cost having a candidate type are selected from N reordered candidates having said candidate type, where M and N are positive integers.
[0513] Item 54. The method according to Item 53, wherein M is based on the candidate type and / or encoding / decoding mode of the current video block.
[0514] Item 55. The method according to Item 53, wherein the candidate type includes adjacent spatial BV candidates, M is 4, and N is 5.
[0515] Item 56. The method according to Item 53, wherein the candidate type includes non-adjacent spatial BV candidates, M is 6, and N is 35.
[0516] Item 57. The method according to Item 53, wherein the candidate type includes time-domain BV candidates, M is 4, and N is 10.
[0517] Item 58. The method according to Item 53, wherein the candidate type includes HMVP BV candidates, M is 10, and N is 25.
[0518] Item 59. The method according to Item 53, wherein the candidate type includes block vector guided BV candidates, M is 4, and N is 10.
[0519] Item 60. The method according to Item 53, wherein the candidate type comprises pairwise averaged BV candidates, M is 1, and N is 6.
[0520] Item 61. The method according to Item 53, wherein the candidate type includes predefined BV candidates, M is 1, and N is 6.
[0521] Item 62. The method according to any one of items 41 to 61, wherein the derived candidate reordering comprises: reordering BV candidates of multiple BV candidate types together.
[0522] Item 63. The method according to Item 62, wherein when constructing the BV candidate list, M candidates with the lowest cost having at least one of the plurality of BV candidate types are selected from N reordered candidates of the plurality of BV candidate types, wherein M is a combination based on the plurality of BV candidate types and / or the encoding / decoding mode of the current video block.
[0523] Item 64. The method according to Item 63, wherein the candidates of the plurality of BV candidate types include at least one of the following: adjacent spatial candidates, temporal candidates, non-adjacent spatial candidates, HMVP BV candidates, block vector guided candidates, pairwise averaged candidates, or predefined BV candidates, wherein M is 6 and N is 20.
[0524] Item 65. The method according to any one of items 62 to 64, wherein at least one BV candidate type of BV candidate is first reordered using ARMC based on the BV candidate type.
[0525] Item 66. The method according to any one of items 62 to 65, wherein N1 HMVP candidates are selected from candidates that have been reordered with HMVP candidate type, and the selected N1 HMVP candidates are reordered together with adjacent spatial candidates and / or temporal candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N1 is a positive integer.
[0526] Item 67. The method according to any one of items 62 to 65, wherein N2 temporal candidates are selected from candidates that have been reordered with temporal candidate types, and the selected N2 temporal candidates are reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N2 is a positive integer.
[0527] Item 68. The method according to any one of items 62 to 65, wherein N3 block vector-guided candidates are selected from candidates that have been reordered with a block vector-guided candidate type, and the selected N3 block vector-guided candidates are reordered together with adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or temporal candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N3 is a positive integer.
[0528] Item 69. The method according to any one of items 62 to 68, wherein if a candidate is reordered more than once, the reordering criteria of the candidate are reused, the reordering criteria including template matching cost.
[0529] Item 70. The method according to any one of Items 1 to 69, wherein the block vector guided block vector prediction is used for at least one of: bidirectional prediction intra-block copy (IBC) prediction or multi-hypothesis IBC prediction.
[0530] Item 71. The method according to Item 70, wherein at least one prediction is derived from a common IBC prediction and at least one prediction is derived from a block vector prediction derived from a block vector.
[0531] Item 72. The method according to Item 71, wherein an index is indicated from the encoder to the decoder to indicate IBC Merge or AMVP BV information for the common IBC prediction.
[0532] Item 73. The method according to Item 71, wherein the block vector-guided block vector prediction is derived from the indicated IBC Merge or AMVP BV information.
[0533] Item 74. The method according to Item 71, wherein the common IBC prediction is derived from at least one of the following: regular IBC Merge prediction, regular IBC AMVP prediction, IBC-TM Merge prediction, IBC-TM AMVP prediction, RR-IBCMerge prediction, RR-IBC AMVP prediction, IBC-MBVD prediction, IBC-CIIP prediction, IBC-GPM prediction, encoder-only BV prediction, string copy vector prediction, intra-template matching prediction (IntraTMP) Merge prediction, or further BV prediction.
[0534] Item 75. The method according to any one of Items 1 to 74, wherein whether at least one of bidirectional prediction IBC prediction or multi-hypothesis IBC prediction is used is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
[0535] Item 76. The method according to any one of items 1 to 75, wherein the BV is replaced by the MV, the BVP list is replaced by the MVP list, the IBC prediction is replaced by the inter-frame prediction, the IBC Merge list is replaced by the Merge list, the IBC AMVP list is replaced by the AMVP list, and the vector guided by the block vector is replaced by the vector guided by the MV.
[0536] Item 77. The method according to any one of items 1 to 75, wherein the MV-guided MV prediction or candidate includes MV... 0,n+1 The cascaded MVP or candidate guided by the nth-order motion vector, and the MV 0,n+1 This is derived from the following: MV 0,n+1 =MV 0,1 +MV 1,2 +…+MV n-1,n +MV n,n+1 MV 0,1 MV represents the first guiding motion vector. 1,2 MV represents the first reference motion vector. n-1,n Let represent the (n-1)th reference motion vector, which points to the nth reference block in the first reference image, and MV n,n+1 This represents the nth reference motion vector, which points to the (n+1)th reference block in the second reference image.
[0537] Item 78. The method according to Item 77, wherein MV 0,n+1 and MV n,n+1 Referring to the same reference image, MV i,i+1 and MV j,j+1 Refers to different reference images.
[0538] Item 79. The method according to any one of items 1 to 78, wherein the MV-guided MV prediction or candidate includes MV... n,n+1 The direct MVP or candidate guided by the nth-order motion vector, where MVn,n+1 points from the nth reference block in the first reference image to the (n+1)th reference block in the second reference image.
[0539] Item 80. The method according to any one of items 1 to 79, wherein the syntax elements in the bit stream are binarized into one of the following: a flag, a fixed-length code, an exponential Golomb (x) (EG(x)) code, a unary code, a rounded unary code, or a rounded binary code, and the syntax elements are signed or unsigned.
[0540] Item 81. The method according to any one of items 1 to 80, wherein the syntax elements in the bitstream are encoded or decoded using at least one context model or are bypassed.
[0541] Item 82. The method according to any one of items 1 to 81, wherein a syntax element is included in the bitstream based on at least one condition, the at least one condition including a condition for a function associated with the syntax element to be applicable to the transformation.
[0542] Item 83. The method according to any one of items 1 to 82, wherein the syntax element is included at one of the following: block level, sequence level, picture group level, picture level, stripe level, slice group level, or codec structure, said codec structure including one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.
[0543] Item 84. The method according to any one of items 1 to 83, wherein the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0544] Item 85. The method according to any one of items 1 to 84, wherein whether and / or how the method is applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, stripe level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header or slice group header.
[0545] Item 86. The method according to any one of items 1 to 84, wherein whether and / or how the method is applied includes: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a strip, a slice, a sub-picture, or a region containing more than one sample point or pixel.
[0546] Item 87. The method according to any one of items 1 to 84, wherein whether and / or how the method is applied is based on encoded information, said encoded information including at least one of the following: block size, color format, single-tree or dual-tree segmentation, color components, stripe type or picture type.
[0547] Item 88. The method according to any one of items 1 to 87, wherein the conversion includes encoding the current video block into the bitstream.
[0548] Item 89. The method according to any one of items 1 to 87, wherein the conversion includes decoding the current video block from the bitstream.
[0549] Item 90. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 89.
[0550] Item 91. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 89.
[0551] Item 92. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method comprises: determining a first guide vector associated with a current video block of the video, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV); determining a first reference vector of a first reference block located based on the first guide vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction of the current video block based at least on the first guide vector and the first reference vector, the vector prediction including a block vector prediction or a motion vector prediction; and generating the bitstream based on the vector prediction.
[0552] Item 93. A method for storing a bitstream of video, comprising: determining a first guide vector associated with a current video block of the video, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV); determining a first reference vector for a first reference block located based on the first guide vector, the first reference vector including a first reference block vector or a first reference motion vector; determining a vector prediction of the current video block based at least on the first guide vector and the first reference vector, the vector prediction including a block vector prediction or a motion vector prediction; generating the bitstream based on the vector prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0553] Example device Figure 27 A block diagram of a computing device 2700 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2700 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0554] It should be understood that, Figure 27 The computing device 2700 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0555] like Figure 27 As shown, computing device 2700 includes general-purpose computing device 2700. Computing device 2700 may include at least one or more processors or processing units 2710, memory 2720, storage unit 2730, one or more communication units 2740, one or more input devices 2750, and one or more output devices 2760.
[0556] In some embodiments, the computing device 2700 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2700 can support any type of interface to the user (such as "wearable" circuit systems, etc.).
[0557] Processing unit 2710 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2700. Processing unit 2710 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0558] Computing device 2700 typically includes various computer storage media. Such media can be any media accessible by computing device 2700, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2730 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2700.
[0559] The computing device 2700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 27 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0560] Communication unit 2740 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 2700 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0561] Input device 2750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2760 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2740, computing device 2700 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2700 can also communicate with one or more devices that enable a user to interact with computing device 2700, or any device that enables computing device 2700 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0562] In some embodiments, some or all components of computing device 2700 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from regular servers or installed directly or otherwise on client devices.
[0563] In embodiments of this disclosure, computing device 2700 can be used to implement video encoding / decoding. Memory 2720 may include one or more video codec modules 2725 having one or more program instructions. These modules can be accessed and executed by processing unit 2710 to perform the functions of the various embodiments described herein.
[0564] In an example embodiment of performing video encoding, input device 2750 may receive video data as input 2770 to be encoded. The video data may be processed, for example, by video codec module 2725 to generate an encoded bitstream. The encoded bitstream may be provided as output 2780 via output device 2760.
[0565] In an example embodiment of performing video decoding, input device 2750 may receive an encoded bitstream as input 2770. The encoded bitstream may be processed, for example, by a video codec module 2725 to generate decoded video data. The decoded video data may be provided as output 2780 via output device 2760.
[0566] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, a first guiding vector associated with the current video block is determined, the first guiding vector including a first guiding block vector (BV) or a first guiding motion vector (MV). A first reference vector is determined based on the first guide vector to locate the first reference block, and the first reference vector includes a first reference block vector or a first reference motion vector; The vector prediction of the current video block is determined based at least on the first guiding vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction; and The transformation is performed based on the vector prediction.
2. The method of claim 1, wherein the block vector prediction comprises at least one of the following: Regular Intra-Block Copy (IBC) Merge Prediction, Regular IBC Advanced Motion Vector Prediction (AMVP) Prediction, IBC with Template Matching (IBC-TM) Merge Prediction, IBC-TM AMVP Prediction, Reconstruction Reordering IBC (RR-IBC) Merge Prediction, RR-IBC AMVP Prediction, IBC Merge Mode with Block Vector Difference (IBC-MBVD) Prediction, Combination of Intra-Block Copy and Intra-Prediction (IBC-CIIP), IBC with Geometric Segmentation (IBC-GPM) Prediction, Encoder-Only BV Prediction, String Copy Vector Prediction, Intra-Template Matching Prediction (IntraTMP) Merge Prediction or Further BV Prediction, and The motion vector prediction mentioned above includes at least one of the following: Regular inter-frame merge prediction, regular inter-frame advanced motion vector prediction (AMVP) prediction, inter-frame template matching (TM) merge prediction, or inter-frame-TM AMVP prediction, inter-frame merge mode with motion vector difference (MMVD), intra-frame inter-frame joint prediction (CIIP), geometric segmentation (GPM) prediction, encoder-only MV prediction, or further MV prediction.
3. The method according to claim 1 or 2, wherein determining the vector prediction of the current video block comprises: The sum of the first guiding block vector and the first reference block vector is determined as the block vector prediction.
4. The method according to any one of claims 1 to 3, wherein the vector prediction is a candidate in a vector candidate list, the vector candidate list including a block vector candidate list or a motion vector candidate list.
5. The method of claim 4, wherein the vector candidate list comprises at least one of the following: Regular Intra-Block Copy (IBC) Merge List Regular inter-frame merge list, List of standard IBC Advanced Motion Vector Prediction (AMVP) List of standard inter-frame AMVPs With template matching IBC (IBC-TM) Merge list, Inter-frame template matching (TM) Merge list, IBC-TM AMVP List Inter-Frame-TM AMVP List Reconstruct and reorder the IBC (RR-IBC) Merge list. RR-IBC AMVP List Basic candidate list of IBC Merge modes with block vector difference (IBC-MBVD) Basic candidate list of inter-frame merge mode (MMVD) with motion vector difference. Merge list for Intra-Block Copying and Intra-Prediction Combination (IBC-CIIP) Intra-frame and inter-frame joint prediction (CIIP) Merge list IBC-CIIP AMVP List CIIP AMVP List IBC (IBC-GPM) Merge List with Geometric Partitioning GPM Merge list IBC-GPM AMVP List GPM AMVP List Intra-Template Matching Prediction (IntraTMP) Merge List BV candidate list for encoder only Encoder-only MV candidate list Further BV candidate list, or Further list of music video candidates.
6. The method according to any one of claims 1 to 5, wherein the vector prediction comprises at least one vector-guided cascaded prediction or candidate, and the at least one vector-guided cascaded prediction or candidate comprises a first-order vector-guided cascaded prediction or candidate. The first-order vector-guided cascaded prediction or candidate is the sum of the first guiding vector and the first reference vector of the first reference block, wherein the first guiding vector points to the first reference block and the first reference vector points to the second reference block.
7. The method of claim 6, wherein the first reference block has been or partially reconstructed within the current image.
8. The method of claim 6, wherein the sum of the first guiding vector and the first reference vector is valid for the current video block.
9. The method of claim 6, wherein the at least one vector-guided cascaded prediction or candidate includes an n-order vector-guided cascaded prediction or candidate BV. 0,n+1 n is an integer greater than or equal to 1, where BV 0,n+1 It was determined through the following: BV 0,n+1 = BV 0,1 +BV 1,2 +…+BV n-1,n +BV n,n+1 , Among them BV 0,1 BV represents the first guiding vector. 1,2 Let BV represent the first reference vector, ..., BV. n-1,n Let BV represent the (n-1)th reference vector. n,n+1 This represents the nth reference vector. Among them BV 1,2 From the first reference block to the second reference block, BV n-1,n Pointing from the (n-1)th reference block to the nth reference block, and BV n,n+1 Point from the nth reference block to the (n+1)th reference block.
10. The method according to any one of claims 1 to 5, wherein the vector prediction comprises at least one vector-guided direct prediction or candidate, the at least one vector-guided direct prediction or candidate comprising a first-order vector-guided direct prediction or candidate. The first guiding vector points to the first reference block, and the direct prediction or candidate guided by the first-order vector is the first reference vector pointing from the first reference block to the second reference block.
11. The method of claim 10, wherein the first reference block has been or partially reconstructed within the current image.
12. The method of claim 10, wherein the first-order vector-guided direct prediction or candidate is valid for the current video block.
13. The method of claim 10, wherein the at least one vector-guided direct prediction or candidate includes an n-order vector-guided direct prediction or candidate BV. n,n+1 n is an integer greater than or equal to 1, where BV n,n+1 From the nth reference block to the (n+1)th reference block.
14. The method of claim 13, wherein BV is determined based on the nth reference block. n,n+1 include: Examine at least one position associated with the nth reference block to derive the vector.
15. The method of claim 14, wherein the at least one location includes a center location. Assuming the top-left position within a block of width W and height H is (0, 0), the center position of the block is represented by one of the following: (W>>1, H>>1) ((W>>1)-1, H>>1), ((W>>1), (H>>1)-1), or ((W>>1)-1, (H>>1)-1).
16. The method of claim 14, wherein the at least one position includes the upper left position of the nth reference block.
17. The method of claim 14, wherein the at least one position includes at least one of the following: upper left position, upper right position, center position, lower left position, lower right position, or some predefined positions.
18. The method of claim 17, wherein the inspection order is the center position, the upper left position, the upper right position, the lower left position, and the lower right position.
19. The method of claim 15, wherein all available vectors are used to derive vector-guided vector candidates, or wherein a first available vector is used to derive the vector-guided vector candidates.
20. The method according to any one of claims 14 to 19, wherein checking at least one location comprises: Inspect the motion mesh covering the position of the nth reference block; as well as If the motion mesh is determined to be available and has BV information, the BV information is used for block vector-guided BV candidate derivation.
21. The method according to any one of claims 1 to 20, wherein the first guiding vector associated with the current video block is determined according to at least one of the following: Existing vector candidates in the vector candidate list prior to determining the vector prediction; or At least one predefined vector candidate from the list of vector candidates before determining the vector prediction.
22. The method of claim 21, wherein the existing vector candidate includes vector-guided vector candidates, and the vector-guided vector candidate includes at least one of the following: vector-guided cascaded candidates or vector-guided direct candidates.
23. The method of claim 21, wherein the at least one predefined vector candidate includes the first N vector candidates or the selected N vector candidates, where N is a positive integer, or The at least one predefined vector candidate includes at least one candidate type of vector candidate.
24. The method of claim 23, wherein the at least one candidate type comprises at least one of the following: Adjacent airspace BV candidates, Non-adjacent airspace BV candidates Historical motion vector prediction (HMVP) BV candidates, Adjacent temporal BV candidates, Non-adjacent temporal BV candidates Paired average BV candidates, or Block vector-guided BV candidates.
25. The method according to any one of claims 6 to 24, wherein a first-order vector-guided cascaded vector candidate or a first-order vector-guided direct vector candidate is used for the transformation.
26. The method according to any one of claims 6 to 24, wherein a first-order vector-guided cascaded or direct vector candidate and a second-order vector-guided cascaded or direct vector candidate are used for the transformation.
27. The method according to any one of claims 6 to 24, wherein cascaded or direct vector candidates guided by the first N order vectors are used for the transformation, where N is a positive integer.
28. The method according to any one of claims 1 to 27, wherein the number of vector candidates guided by the vector is not limited.
29. The method according to any one of claims 1 to 27, wherein the number of vector-guided vector candidates is less than or equal to the maximum number.
30. The method of claim 29, wherein the maximum number is 5 or 4.
31. The method of claim 29, wherein the number of unique vector candidates guided by the vector after complete deduplication is less than or equal to the maximum number, which is 5, 4 or 3.
32. The method according to any one of claims 1 to 31, wherein when deriving vector candidates guided by vectors, redundancy checks or deduplication are performed, the vector candidates including block vector candidates or motion vector candidates.
33. The method of claim 32, wherein when deriving the vector-guided vector candidates, complete deduplication is performed to ensure that candidates with the same or similar motion information are excluded from the vector candidate list.
34. The method of claim 32, wherein partial deduplication is performed when deriving the vector-guided vector candidates.
35. The method according to any one of claims 1 to 34, wherein the position of the block vector-guided BV candidate in the BV candidate list is based on: inserting the block vector-guided BV candidate after the history-based motion vector prediction (HMVP) BV candidate.
36. The method according to any one of claims 1 to 34, wherein the position of the block vector-guided BV candidate in the BV candidate list is based on: inserting the block vector-guided BV candidate after a non-adjacent spatial BV candidate.
37. The method according to any one of claims 1 to 35, wherein the position of the block vector-guided BV candidate in the BV candidate list is based on one of the following: Insert block vector-guided BV candidates before the history-based motion vector prediction (HMVP) BV candidates. Insert a portion of the block vector-guided BV candidate before the HMVP BV candidate, and insert the remainder of the block vector-guided BV candidate after the HMVP BV candidate. Insert the block vector-guided BV candidate before the non-adjacent spatial BV candidate, or A portion of the block vector-guided BV candidate is inserted before the non-adjacent spatial BV candidate, and the remainder of the block vector-guided BV candidate is inserted after the non-adjacent spatial BV candidate.
38. The method according to any one of claims 1 to 37, wherein whether vector-guided vector prediction is used is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
39. The method according to any one of claims 1 to 38, further comprising: A reordering or refinement process is applied to derive a vector candidate list, which includes a BV candidate list or an MV candidate list.
40. The method of claim 39, wherein the reordering or the refinement process is based on template matching cost.
41. The method of claim 39 or 40, wherein the BV candidate list is constructed by: The derivation is at least partially derived by completely deduplicating at least one of the following candidates: N1 adjacent spatial candidates, N2 temporal candidates, N3 non-adjacent spatial candidates, N4 HMVP candidates, N5 block vector-guided BV candidates, N6 pairwise averaged candidates, or N7 predefined BV candidates, where N1, N2, N3, N4, N5, N6, and N7 are positive integers. Reordering the derived candidate data; as well as The top N candidates with the lowest cost are selected as the final candidates in the BV candidate list, where N is a positive integer.
42. The method according to claim 41, wherein N is 6, N1 is 5, N2 is 10, N3 is 35, N4 is 25, N5 is 10, N6 is 1, and N7 is 6.
43. The method of claim 41, wherein the number of unique BV candidates after complete deduplication is less than or equal to the maximum number.
44. The method of claim 43, wherein the maximum number is 20 or 28.
45. The method of claim 41, wherein the adjacent spatial candidates consist of at least one of the following: Candidates for the left airspace, the upper airspace, the upper right airspace, the lower left airspace, or the upper left airspace.
46. The method of claim 41, wherein the time-domain BV candidate is composed of adjacent time-domain BV candidates or non-adjacent time-domain BV candidates.
47. The method of claim 41, wherein the non-adjacent spatial domain candidates consist of non-adjacent spatial domain candidates between regular frames.
48. The method of claim 41, wherein the number of HMVP BV candidates and / or the HMVP table size is increased to 25.
49. The method of claim 41, wherein the block vector-guided BV candidate comprises at least one block vector-guided cascade or block vector-guided direct BV candidate.
50. The method of claim 41, wherein paired BV candidates are generated by averaging the candidates of predefined pairs in a motion candidate list, the predefined pairs being in a set denoted as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, wherein 0, 1, 2, and 3 represent motion candidate indices in the motion candidate list.
51. The method of claim 41, wherein the predefined BV candidate is located in the intra-block copy reference region.
52. The method according to any one of claims 41 to 51, wherein the derived candidate reordering comprises: Adaptive reordering of BV candidates based on BV candidate type (ARMC) is applied to reorder the BV candidates with at least one candidate type based on criteria.
53. The method of claim 52, wherein M candidates with the lowest cost having a candidate type are selected from N reordered candidates having the candidate type, where M and N are positive integers.
54. The method of claim 53, wherein M is based on the candidate type and / or encoding / decoding mode of the current video block.
55. The method of claim 53, wherein the candidate type includes adjacent spatial BV candidates, M is 4, and N is 5.
56. The method of claim 53, wherein the candidate type includes non-adjacent spatial BV candidates, M is 6, and N is 35.
57. The method of claim 53, wherein the candidate type includes a time-domain BV candidate, M is 4, and N is 10.
58. The method of claim 53, wherein the candidate type includes an HMVP BV candidate, M is 10, and N is 25.
59. The method of claim 53, wherein the candidate type includes block vector guided BV candidates, M is 4, and N is 10.
60. The method of claim 53, wherein the candidate type comprises pairwise averaged BV candidates, M is 1, and N is 6.
61. The method of claim 53, wherein the candidate type includes predefined BV candidates, M is 1, and N is 6.
62. The method according to any one of claims 41 to 61, wherein the derived candidate reordering comprises: Reorder the BV candidates of multiple BV candidate types together.
63. The method of claim 62, wherein when constructing the BV candidate list, M candidates with the lowest cost having at least one of the plurality of BV candidate types are selected from N reordered candidates of the plurality of BV candidate types, wherein M is a combination based on the plurality of BV candidate types and / or the encoding / decoding mode of the current video block.
64. The method of claim 63, wherein the plurality of BV candidate types include at least one of the following: adjacent spatial candidates, temporal candidates, non-adjacent spatial candidates, HMVP BV candidates, block vector guided candidates, pairwise averaged candidates, or predefined BV candidates. Where M is 6 and N is 20.
65. The method according to any one of claims 62 to 64, wherein at least one BV candidate type of BV candidate is first reordered using ARMC based on the BV candidate type.
66. The method according to any one of claims 62 to 65, wherein N1 HMVP candidates are selected from the reordered candidates having HMVP candidate types, and the selected N1 HMVP candidates are reordered together with the adjacent spatial candidates and / or temporal candidates and / or non-adjacent spatial candidates and / or block vector-guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N1 is a positive integer.
67. The method according to any one of claims 62 to 65, wherein N2 temporal candidates are selected from the reordered candidates having the temporal candidate type, and the selected N2 temporal candidates are reordered together with the adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or block vector guided candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N2 is a positive integer.
68. The method according to any one of claims 62 to 65, wherein N3 block vector-guided candidates are selected from the reordered candidates having the block vector-guided candidate type, and the selected N3 block vector-guided candidates are reordered together with the adjacent spatial candidates and / or HMVP candidates and / or non-adjacent spatial candidates and / or temporal candidates and / or pairwise averaged candidates and / or predefined BV candidates, where N3 is a positive integer.
69. The method according to any one of claims 62 to 68, wherein if a candidate is reordered more than once, the reordering criterion of the candidate is reused, the reordering criterion including template matching cost.
70. The method according to any one of claims 1 to 69, wherein the block vector guided block vector prediction is used for at least one of: bidirectional prediction intra-block copy (IBC) prediction or multi-hypothesis IBC prediction.
71. The method of claim 70, wherein at least one prediction is derived from a common IBC prediction, and at least one prediction is derived from a block vector prediction guided by the block vector.
72. The method of claim 71, wherein an index is indicated from the encoder to the decoder to indicate IBC Merge or AMVP BV information for the common IBC prediction.
73. The method of claim 71, wherein the block vector-guided block vector prediction is derived from the indicated IBCMeerge or AMVP BV information.
74. The method of claim 71, wherein the common IBC prediction is derived from at least one of the following: Conventional IBC Merge forecast, Standard IBC AMVP prediction IBC-TM Merge Prediction IBC-TM AMVP Prediction RR-IBC Merge Prediction RR-IBC AMVP Prediction IBC-MBVD prediction IBC-CIIP Prediction IBC-GPM forecast BV prediction of encoder only String copy vector prediction Intra-Template Matching Prediction (IntraTMP) Merge Prediction, or Further BV forecasts.
75. The method according to any one of claims 1 to 74, wherein whether at least one of bidirectional prediction IBC prediction or multi-hypothesis IBC prediction is used is indicated by one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice group header.
76. The method according to any one of claims 1 to 75, wherein the BV is replaced by the MV, the BVP list is replaced by the MVP list, the IBC prediction is replaced by the inter-frame prediction, the IBC Merge list is replaced by the Merge list, the IBC AMVP list is replaced by the AMVP list, and the vector guided by the block vector is replaced by the vector guided by the MV.
77. The method according to any one of claims 1 to 75, wherein the MV-guided MV prediction or candidate includes MV... 0,n+1 The cascaded MVP or candidate guided by the nth-order motion vector, and the MV 0,n+1 This is derived from the following: MV 0,n+1 = MV 0,1 +MV 1,2 +…+MV n-1,n +MV n,n+1 , MV 0,1 MV represents the first guiding motion vector. 1,2 MV represents the first reference motion vector. n-1,n Let represent the (n-1)th reference motion vector, which points to the nth reference block in the first reference image, and MV n,n+1 This represents the nth reference motion vector, which points to the (n+1)th reference block in the second reference image.
78. The method of claim 77, wherein MV 0,n+1 and MV n,n+1 Referring to the same reference image, MV i,i+1 and MV j,j+1 Refers to different reference images.
79. The method according to any one of claims 1 to 78, wherein the MV-guided MV prediction or candidate includes MV... n,n+1 The direct MVP or candidate guided by the nth-order motion vector, where MV n,n+1 Point from the nth reference block in the first reference image to the (n+1)th reference block in the second reference image.
80. The method according to any one of claims 1 to 79, wherein the syntax elements in the bitstream are binarized into one of the following: flags, fixed-length codes, exponential Golomb (x) (EG(x)) codes, unary codes, rounded unary codes, or rounded binary codes, and the syntax elements are signed or unsigned.
81. The method according to any one of claims 1 to 80, wherein the syntax elements in the bitstream are encoded or decoded using at least one context model or are bypassed.
82. The method according to any one of claims 1 to 81, wherein a syntax element is included in the bitstream based on at least one condition, the at least one condition including a condition for a function associated with the syntax element to be applicable to the transformation.
83. The method according to any one of claims 1 to 82, wherein the syntax element is included at one of the following: block level, sequence level, picture group level, picture level, stripe level, slice group level, or codec structure, said codec structure including one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.
84. The method according to any one of claims 1 to 83, wherein the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
85. The method according to any one of claims 1 to 84, wherein whether and / or how the method is applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, stripe level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header or slice group header.
86. The method according to any one of claims 1 to 84, wherein whether and / or how the method is applied is included in one of the following: Prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-image, or region containing more than one sample point or pixel.
87. The method according to any one of claims 1 to 84, wherein whether and / or how the method is applied is based on encoded information, said encoded information including at least one of the following: block size, color format, single-tree or dual-tree segmentation, color components, stripe type or picture type.
88. The method according to any one of claims 1 to 87, wherein the conversion comprises encoding the current video block into the bitstream.
89. The method according to any one of claims 1 to 87, wherein the conversion comprises decoding the current video block from the bitstream.
90. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 89.
91. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 89.
92. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: A first guide vector is determined to be associated with the current video block of the video, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV); Determine a first reference vector for a first reference block located based on the first guide vector, wherein the first reference vector includes a first reference block vector or a first reference motion vector; The vector prediction of the current video block is determined based at least on the first guiding vector and the first reference vector, the vector prediction including block vector prediction or motion vector prediction; and The bit stream is generated based on the vector prediction.
93. A method for storing a bitstream of video, comprising: A first guide vector is determined to be associated with the current video block of the video, the first guide vector including a first guide block vector (BV) or a first guide motion vector (MV); Determine a first reference vector for a first reference block located based on the first guide vector, wherein the first reference vector includes a first reference block vector or a first reference motion vector; The vector prediction of the current video block is determined based at least on the first guiding vector and the first reference vector, and the vector prediction includes block vector prediction or motion vector prediction; The bitstream is generated based on the vector prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.