Method and device for video processing and medium
By adjusting the variable values associated with the Merge candidate between video blocks and bitstreams, based on template size, sequence resolution, or block size, the problem of insufficient encoding and decoding efficiency in existing video encoding and decoding technologies is solved, achieving more efficient encoding and decoding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-09-20
- Publication Date
- 2026-04-17
AI Technical Summary
There is room for improvement in the efficiency of existing video encoding and decoding technologies, especially in video compression technology, where the effectiveness and efficiency of encoding and decoding need to be improved.
More efficient encoding and decoding can be achieved by determining the value of the first variable associated with the Merge candidate between the video block and the bitstream, and adjusting the value of the variable based on the template size, sequence resolution, or block size, or by adjusting the value of the variable by comparing the template matching cost.
It improves the effectiveness and efficiency of video encoding and decoding, and optimizes the video processing process.
Smart Images

Figure CN121890079A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to the adjustment of inherited variables. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: determining the value of a first variable associated with a merge candidate of the current video block for a conversion between a current video block and the video bitstream; determining adjustment information for the first variable based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and performing a conversion based on the adjustment information. In this way, whether and / or how to adjust the value of the first variable can be determined based on the template size, sequence resolution, and / or block size. Therefore, encoding / decoding efficiency and encoding / decoding effectiveness can be improved.
[0005] In a second aspect, another method for video processing is proposed. The method includes: determining a value of a first variable associated with a merge candidate of the current video block for a conversion between a current video block and a bitstream of the video; determining a first template matching cost when the first variable has a first value and a second template matching cost when the first variable has a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; and performing a conversion based on the adjusted value of the first variable. According to the method of the second aspect of this disclosure, the value of the first variable is adjusted based on a comparison of template matching costs, and the template matching cost is updated based on a scaling factor.
[0006] In a third aspect, another method for video processing is proposed. This method includes: for a conversion between a current video block and a video bitstream, determining whether a derived value of a template matching cost of a first variable based on the motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in a high-level motion vector prediction mode; indicating in the bitstream whether the derived value is the same as the selected value; and performing the conversion based on the indication. The method according to the third aspect of this disclosure includes an indication in the bitstream whether the derived value is the same as the selected value, rather than including the selected value in the bitstream.
[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.
[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first, second, or third aspect of this disclosure.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining the value of a first variable associated with a merge candidate of a current video block of the video; determining adjustment information of the first variable based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and generating a bitstream based on the adjustment information.
[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a value of a first variable associated with a merge candidate of a current video block of the video; determining a first template matching cost when the first variable has a first value and a second template matching cost when the first variable has a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; and generating a bitstream based on the adjusted value of the first variable.
[0011] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: for a conversion between a current video block and the bitstream of the video; determining whether a derived value of a template matching cost of a first variable based on the motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in a high-level motion vector prediction mode; indicating in the bitstream whether the derived value is the same as the selected value; and generating the bitstream based on the indication.
[0012] In a ninth aspect, a method for storing a bitstream of video is proposed. The method includes: determining the value of a first variable associated with a Merge candidate of a current video block of the video; determining adjustment information for the first variable based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; generating a bitstream based on the adjustment information; and storing the bitstream in a non-transitory computer-readable recording medium.
[0013] In a tenth aspect, a method for storing a bitstream of video is proposed. The method includes: determining the value of a first variable associated with a Merge candidate of a current video block of the video; determining a first template matching cost when the first variable has a first value and a second template matching cost when the first variable has a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; generating a bitstream based on the adjusted value of the first variable; and storing the bitstream in a non-transitory computer-readable recording medium.
[0014] In the eleventh aspect, a method for storing a bitstream of video is proposed. The method includes: for the conversion between a current video block and the bitstream of the video, determining whether a derived value of a template matching cost of a first variable based on the motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in a high-level motion vector prediction mode; determining an indication in the bitstream based on the determination; generating a bitstream based on the indication, the indication indicating whether the derived value is the same as the selected value; and storing the bitstream in a non-transitory computer-readable recording medium.
[0015] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0016] The above and other objects, features and advantages of the exemplary embodiments of this disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0017] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 A schematic diagram showing the locations of the spatial merge candidates is provided. Figure 5 A schematic diagram is shown of the candidate pairs considered in the redundancy check for spatial merge candidates. Figure 6 A schematic diagram of motion vector scaling for temporal Merge candidates is shown; Figure 7 A schematic diagram of candidate positions C0 and C1 for time-domain Merge candidates is shown; Figure 8 This diagram illustrates the VVC spatial neighboring blocks of the current block; Figure 9 A schematic diagram of the virtual block in the i-th round of search is shown; Figure 10 The spatial neighboring blocks used to derive spatial merge candidates are shown; Figure 11 This illustrates the non-adjacent temporal neighbor blocks used to derive non-adjacent temporal merge candidates; Figure 12A and Figure 12B The MMVD search points are shown respectively; Figure 13 The top and left neighbor blocks used in the CIIP weight derivation are shown; Figure 14 This demonstrates performing template matching on the search area surrounding the initial MV; Figure 15A and Figure 15B The methods for dividing angle patterns are shown respectively; Figure 16A and Figure 16B The affine motion models based on control points are shown respectively; Figure 17 The affine MVF for each sub-block is shown; Figure 18The location of the inherited affine motion prediction value is shown; Figure 19 This demonstrates the inheritance of control point motion vectors; Figure 20 The locations of candidate positions for the constructed affine Merge pattern are shown; Figure 21 The first HPT and the second HPT are shown; Figure 22 Spatial nearest neighbors for deriving affine Merge / AMVP candidates are shown: (a) candidates for deriving inheritance and (b) candidates for deriving the first type of construction. Figure 23 Affine Merge / AMVP candidates are shown, ranging from non-nearest neighbors to the first type of construction; Figure 24 A schematic diagram of the neighboring 4x4 sub-blocks used for RMVF parameter derivation is shown, where W and H are the width and height of the current CU; Figure 25A and Figure 25B The SbTMVP procedure in VVC is shown, where Figure 25A The spatial neighbor blocks used by ATMVP are shown, and Figure 25B The paper demonstrates how to derive the motion field of a sub-CU by applying motion displacements from spatial neighbors and scaling motion information from the corresponding co-located sub-CUs. Figure 26 This shows the spatial proximity locations used in the construction of the IBC Merge / AMVP list; Figure 27 The filling candidates for replacing the zero vector in the IBC list are shown; Figure 28 The IBC reference area is shown, depending on the current CU location; Figure 29 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 30 A flowchart of another method for video processing according to an embodiment of the present disclosure is shown; Figure 31 A flowchart of another method for video processing according to embodiments of the present disclosure is shown; and Figure 32 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0018] In all the accompanying drawings, the same or similar reference numerals usually indicate the same or similar elements. Detailed Implementation
[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art to understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0020] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0022] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0024] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 may also be referred to as a video encoding device, and the target device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the target device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded and decoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to target device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by target device 120.
[0027] Target device 120 may include I / O interface 126, video decoder 124, and display device 122. I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120, which is configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0029] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0035] The mode selection unit 203 can, for example, select one of several codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame / inter-frame joint prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0037] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images in lists 0 and 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0041] In one example, motion estimation unit 204 may indicate a value to video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0042] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information based on the entropy-decoded video data. This motion information includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0056] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0057] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially neighboring blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term "video processing" includes video encoding / decoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0062] 1. Brief Overview This disclosure relates to image / video codecs, and in particular to the adjustment of LIC directives. This disclosure can be applied to existing video codec standards such as HEVC or standard VVC (Video Codec for Multipurpose). This disclosure may also be applicable to future video codec standards or video codecs.
[0063] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding.
[0064] To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. JVET meetings are held quarterly, and the new video codec standard was officially named Multifunctional Video Codec (VVC) at the April 2018 JVET meeting, where the first version of the VVC Test Model (VTM) was also released. The VVC working draft and the VTM test model were subsequently updated after each meeting. The VVC project achieved Technical Completeness (FDIS) at the July 2020 meeting.
[0065] In January 2021, JVET established an exploratory experiment (EE) with the goal of achieving enhanced compression efficiency beyond VVC capabilities using novel conventional algorithms. Shortly thereafter, ECM was built as a general software foundation for long-term exploratory work on next-generation video codec standards.
[0066] 2.1. Extended Merge Forecast In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: 1) Airspace MVP from the airspace adjacent to the CU.
[0067] 2) Temporal MVP from the same CU.
[0068] 3) History-based MVP from FIFO table.
[0069] 4) Pair average MVP.
[0070] 5) Zero MV.
[0071] The size of the Merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU encoded in Merge mode, the index of the best Merge candidate is encoded using rounding univariate binarization (TU). The first bit of the Merge index is encoded using the context, and bypass encoding is used for the remaining bits.
[0072] This section provides the derivation process for various merge candidates. Similar to HEVC, VVC also supports parallel derivation of the merge candidate list for all CUs within a specific size region.
[0073] 2.1.1 Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Figure 4 From the candidates shown, a maximum of four Merge candidates can be selected. Figure 4 The locations of spatial merge candidates are shown. The derivation order is B1, A1, B0, A0, and B2. Location B2 is considered only when one or more CUs at locations B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or when intra-frame encoding / decoding is used. After adding the candidate at location A1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only... Figure 5 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information. Figure 5 The candidate pairs considered for redundancy checks of spatial Merge candidates are shown.
[0074] 2.1.2 Time-domain candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU is explicitly transmitted via signal transmission in the strip header. Figure 6 The scaling of motion vectors for temporal Merge candidates is shown. For example... Figure 6 As shown by the dashed lines, the scaled motion vectors for the temporal merge candidates are obtained by scaling the motion vectors from the co-located CUs using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index for the temporal merge candidates is set to 0.
[0075] like Figure 7 As shown, the position of the time-domain candidate is selected between candidate C0 and C1. Figure 7 A schematic diagram of candidate positions C0 and C1 for the temporal merge candidate is shown. Position C1 is used if the CU at position C0 is unavailable, intra-frame encoded / decoded, or outside the current line of the CTU. Otherwise, position C0 is used for the derivation of the temporal merge candidate.
[0076] 2.1.3 Historical Merge Candidate Derivation Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0077] HMVP Table Size S Setting it to 6 indicates that a maximum of 5 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, the first-in, first-out (FIFO) rule of the constraint is utilized, where a redundancy check is first applied to find if a duplicate HMVP already exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are shifted forward, with the duplicate HMVP inserted as the last entry in the table.
[0078] HMVP candidates can be used during the Merge candidate list construction process. The latest HMVP candidates in the table are checked sequentially and inserted into the candidate list, following the TMVP candidates. Redundancy checks are performed on HMVP candidates and spatial or temporal Merge candidates.
[0079] To reduce the number of redundant check operations, the following simplifications are introduced: 1. The last two entries in the table are checked for redundancy with the A1 and B1 airspace candidates, respectively.
[0080] 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the process of building the Merge candidate list from the HMVP is terminated.
[0081] 2.1.4 Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices of the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.
[0082] If the Merge list is not full after adding pairwise average Merge candidates, a zero MVP will be inserted at the end until the maximum number of Merge candidates is reached.
[0083] 2.2 New Merge Candidates 2.2.1 Derivation of Non-Adjacent Merge Candidates In VVC, Figure 8 The five spatial neighbor blocks and one temporal nearest neighbor shown were used to derive the Merge candidate. Figure 8 This diagram illustrates the VVC spatial neighboring blocks of the current block.
[0084] We propose using the same pattern as in VVC to derive additional Merge candidates from positions not adjacent to the current block. To achieve this, for each search round i, a virtual block is generated based on the current block, as follows: First, the relative position of the virtual block and the current block is calculated using the following formula: Offsetx =-i×gridX, Offsety = -i×gridY, Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.
[0085] Second, the width and height of the virtual block are calculated using the following formula: newWidth = i×2×gridX+ currWidthnewHeight = i×2×gridY +currHeight.
[0086] Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.
[0087] gridX and gridY are currently set to currWidth and currHeight, respectively.
[0088] Figure 9 This illustrates the relationship between the virtual block and the current block. Figure 9 A schematic diagram of the virtual block in the i-th round of search is shown.
[0089] After the virtual block is generated, block A i B i C i D i and E iThese can be considered VVC spatial neighboring blocks, and their positions are obtained using the same pattern as the pattern in the VVC. Clearly, if the search round i is 0, the virtual block is the current block. In this case, block A... i B i C i D i and E i It is a spatial neighbor block used in VVC Merge mode.
[0090] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are utilized.
[0091] Non-adjacent spatial domain merge candidates are inserted into the merge list after the temporal domain merge candidates in the order B1->A1->C1->D1->E1.
[0092] 2.2.2 Non-adjacent airspace candidates Non-adjacent airspace merge candidates are inserted after the TMVP in the regular merge candidate list. The style of airspace merge candidates is as follows: Figure 10 It is shown in the middle. Figure 10 The spatial neighboring blocks used to derive spatial merge candidates are shown. The distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. Line buffer limits are not applied.
[0093] 2.2.3 Non-adjacent time domain candidates Figure 11 This illustrates the non-adjacent temporal neighbor blocks used to derive non-adjacent temporal merge candidates. For example... Figure 11 As shown, non-adjacent temporal locations are introduced, where the non-adjacent temporal MVP location is situated in the same reference frame as the adjacent TMVP. The distance between the non-adjacent temporal candidate and the current codec block is based on the width and height of the current codec block.
[0094] 2.3. Merge Pattern with MVD (MMVD) In addition to the Merge mode (where implicitly derived motion information is directly used for generating prediction samples for the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is transmitted via signaling immediately after the regular Merge flag is sent to indicate whether the MMVD mode is used for the CU.
[0095] In MMVD, after selecting a Merge candidate, the Merge candidate is further refined using MVD information transmitted via signals. This further information includes a Merge candidate flag, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, one of the top two candidates in the Merge list is selected as the MV basis. The mmvd candidate flag is transmitted via signals to specify which one to use between the first and second Merge candidates.
[0096] Figure 12A and Figure 12B The MMVD search points are shown separately.
[0097] The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 12A and Figure 12B As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.
[0098] Table 1 Relationship between Distance Index and Predefined Offset
[0099] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 2. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is an unpredicted MV or a bidirectional predicted MV, and both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbols in Table 2 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional predicted MV, and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), and the POC difference in list 0 is greater than the POC difference in list 1, the symbols in Table 2 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and have the opposite value for the sign of the list 1 MV. Otherwise, if the POC difference in List 1 is greater than the POC difference in List 0, then the sign in Table 2 specifies the sign of the MV offset of the List 1 MV component added to the starting MV, and has the opposite value for the sign of the List 0 MV.
[0100] MVD is scaled based on the difference in POC in each direction. If the difference in POC is the same in both lists, no scaling is needed. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, then MVD for list 1 is scaled (by defining the difference in POC in L0 as td and the difference in POC in L1 as tb), as shown below. Figure 6 As shown. If the POC difference of L1 is greater than the POC difference of L0, then the MVD for list 0 is scaled in the same way. If the initial MV is unidirectionally predicted, then the MVD is added to the available MV.
[0101] Table 2. Signs of MV offsets defined by direction indices
[0102] 2.4. Inter-frame and Intra-frame Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., CU width multiplied by CU height equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether Inter-Frame Intra-Frame Joint Prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines inter-frame prediction signals with intra-frame prediction signals. The inter-frame prediction signal in CIIP mode... The inter-frame prediction process is derived using the same procedure as the regular Merge mode; and the intra-frame prediction signal... The conventional intra-frame prediction process utilizing a planar pattern is derived. Then, a weighted average is used to combine the intra-frame prediction signal and the inter-frame prediction signal, where the values are based on the top neighbor block and the left neighbor block (in...). Figure 13 The weight values are calculated based on the encoding / decoding mode (described in the image), as follows: – If the top nearest neighbor is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. – If the left nearest neighbor is available and is intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. – If (isIntraLeft + isIntraTop) equals 2, then wt is set to 3; Otherwise, if (isIntraLeft + isIntraTop) equals 1, then wt is set to 2; Otherwise, set wt to 1.
[0103] Figure 13 The top neighbor block and left neighbor block used in the CIIP weight derivation are shown.
[0104] The CIIP predictions are formed as follows: .
[0105] 2.5. Template Matching (TM) Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring block of the current CU) and a block in the reference image (i.e., the same size as the template). Figure 14 This illustrates performing template matching on the search area surrounding the initial MV. (Example) Figure 14 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. The template matching method is used with the following modifications: the search step size is determined based on the AMVR mode, and in Merge mode, the TM can be cascaded with the bilateral matching process.
[0106] In TM-AMVP mode, MVP candidates are determined based on template matching error, selecting the one that achieves the minimum difference between the current block template and the reference block template. TM is then performed only on that specific MVP candidate for MV refinement. TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. AMVP candidates can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially according to the AMVR mode specified in Table 3. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the AMVR mode after the TM process. During the search process, the search terminates if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold (equal to the area of the block).
[0107] Table 3. AMVR search patterns and search patterns using AMVR's Merge mode.
[0108] In TM-Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 3, TM can proceed up to 1 / 8 pixel MVD precision, or skip those precisions beyond half-pixel MVD precision, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check.
[0109] 2.6. Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter-prediction signal using CIIP-TM Merge candidate prediction and the intra-prediction signal using the intra-prediction mode derived using TIMD. This method is only applied to codec blocks with an area of 1024 or less.
[0110] The TIMD derivation method was used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list was selected and mapped to one of 67 regular intra-prediction modes.
[0111] Furthermore, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. Figure 15A and Figure 15B The partitioning methods for angle patterns are shown separately. For near-horizontal patterns (2 <= angle pattern index < 34), the current block is partitioned vertically, as follows: Figure 15A As shown; for near-vertical mode (34 <= angle mode index <= 66), the current block is divided horizontally, as... Figure 15B As shown.
[0112] Table 4 shows the (wIntra, wInter) values for different sub-blocks.
[0113] Table 4 Weights used for angle mode modification
[0114] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.
[0115] 2.7. Two-sided matching AMVP-Merge mode Bidirectional predictions consist of an AMVP prediction in one direction and a Merge prediction in the other. This mode can be enabled for codec blocks when the selected Merge and AMVP predictions satisfy the DMVR condition, where there is at least one reference image from the past and one reference image from the future relative to the current image, and both reference images are equidistant from the current image. Bidirectional matching MV refinement is applied to the Merge MV candidate and the AMVP MVP as starting points. Otherwise, if template matching is enabled, template matching MV refinement is applied to either the Merge or AMVP prediction, which has a higher template matching cost.
[0116] The AMVP portion of this mode is signaled as a regular one-way AMVP, i.e., the reference index and MVD are signaled, and it has a derived MVP index if template matching is used, or the MVP index is signaled when template matching is disabled.
[0117] For the AMVP direction LX, where X can be 0 or 1, the Merge portion in the other direction (1 – LX) is implicitly derived by minimizing the bilateral matching cost between the AMVP prediction and the Merge prediction, i.e., for a pair of AMVP and Merge motion vectors. For each Merge candidate in the list of Merge candidates with this other direction (1 – LX) motion vector, the bilateral matching cost is computed using the Merge candidate MV and the AMVP MV. The Merge candidate with the minimum cost is selected. Bilateral matching refinement is applied to the codec block with the selected Merge candidate MV and AMVP MV as starting points.
[0118] The third pass of the multi-pass DMVR (i.e., the 8x8 sub-PU BDOF refinement of the multi-pass DMVR) enables blocks encoded and decoded in AMVP-Merge mode.
[0119] This mode is indicated by a flag, and if this mode is enabled, the AMVP direction LX is further indicated by a flag.
[0120] When the bilateral matching (BM) AMVP-Merge pattern is used for the current block and template matching is enabled, MVD is not signaled. An additional pair of AMVP-Merge MVPs is introduced. The Merge candidate list is sorted in ascending order based on the BM cost. An index (0 or 1) is signaled to indicate which Merge candidate from the sorted Merge candidate list is used. When there is only one candidate in the Merge candidate list, a pair of AMVP MVPs and Merge MVPs without bilateral matching MV refinement is populated.
[0121] 2.8. Adaptive Decoder-Side Motion Vector Refinement The Adaptive Decoder-Side Motion Vector Refinement (ADMVR) method is an extension of multi-pass DMVR, consisting of two new Merge modes for refining the motion vectors (MVs) only in one direction (L0 or L1) of the bidirectional predictions of Merge candidates that satisfy the DMVR conditions. The multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vectors; however, in the first pass (i.e., PU level) of DMVR, either MVD0 or MVD1 is set to zero.
[0122] Merge candidates for the new Merge pattern are derived from spatially adjacent encoded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates, similar to the regular Merge pattern. The difference is that only those satisfying the DMVR conditions are added to the candidate list. Both new Merge patterns use the same Merge candidate list. If the BM candidate list contains inherited BCW weights, the DMVR process remains unchanged, except for the use of MRSAD or MRSATD for distortion calculation when the weights are unequal and bidirectional prediction is weighted with BCW weights. The encoding and decoding of the Merge index is the same as in the regular Merge pattern.
[0123] 2.9. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 16A and Figure 16B Affine motion models based on control points are shown respectively. Figure 16A A 4-parameter affine model is shown. Figure 16B A 6-parameter affine model is shown. (Example) Figure 16A and Figure 16B As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0124] For a 4-parameter affine motion model, the location of the sample point in the block ( x, y The motion vector at point () is derived as: (2-1).
[0125] For a 6-parameter affine motion model, the location of the sample points in the block ( x, y The motion vector at point () is derived as: (2-2).
[0126] in( mv 0x , mv 0y ) is the motion vector of the upper left control point, ( mv 1x , mv 1y ) is the motion vector of the upper right control point, and ( mv 2x , mv 2y ) is the motion vector of the lower left control point.
[0127] To simplify motion compensation prediction, block-based affine transformation prediction is applied. Figure 17 The affine MVF for each sub-block is shown. To derive the motion vector for each 4x4 luma sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 17 (As shown), and rounded to 1 / 16 fractional precision. Then a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4x4. The MV of the 4x4 chroma sub-block is calculated as the average of the MV of the upper-left luminance sub-block and the lower-right luminance sub-block in the corresponding 8x8 luminance region.
[0128] Similar to the inter-frame prediction for translational motion, there are two other inter-frame prediction modes for affine motion: affine Merge mode and affine AMVP mode.
[0129] 2.9.1 Affine Merge Prediction The AF_MERGE mode can be applied to CUs with a width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMV candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPMV candidates are used to form the affine merge candidate list: – Affine Merge candidates inferred from the CPMV of neighboring CUs.
[0130] – The constructed affine Merge candidate CPMVP is derived using the translational MV of the neighboring CU.
[0131] – Zero MV.
[0132] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 18 As shown. Figure 18 The positions of the inherited affine motion predictions are shown. For the predictions on the left, the scan order is A0->A1, and for the predictions above, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between candidates from two inherited sides. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list for the current CU. Figure 19 The inheritance of control point motion vectors is shown. For example... Figure 19 As shown, if the nearest lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on... Computed. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... Calculated.
[0133] The constructed affine candidates are built by combining the nearest-neighbor translational motion information of each control point. The motion information for the control points is derived from... Figure 20 The spatial and temporal nearest neighbors shown are derived. Figure 20 The positions of candidate locations for the constructed affine Merge pattern are shown. CPMV k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If the TMVP is available, it is used as CPMV4.
[0134] After obtaining the motion signatures (MVs) of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3}.
[0135] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0136] After the inherited affine Merge candidate and the constructed affine Merge candidate are checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0137] 2.9.2 Affine Model Inheritance and Non-Adjacent Affine Modes Based on Historical Parameters History-based Affine Model Inheritance (HAMI) allows affine models to be inherited from previously affine-encoded blocks that may not be adjacent to the current block. Similar to the enhanced regular Merge mode, Non-Adjacent Affine Mode (NA-AFF) is introduced.
[0138] The first history parameter table (HPT) is established. Each entry in the first HPT stores a set of affine parameters: a , b , c and d Each affine parameter is represented by a 16-bit signed integer. Entries in the HPT are categorized by reference list and reference index. Five reference indices are supported for each reference list in the HPT. The HPT category (denoted as HPTCat) is calculated in a formulaic manner. HPTCat (RefList, RefIdx) = 5×RefList + min (RefIdx, 4), RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively. A maximum of 7 entries can be stored for each category, resulting in a total of 70 entries in the HPT. At the beginning of each CTU line, the number of entries for each category is initialized to zero. Decoding involves using the reference list RefList... cur and RefIdx cur After the affine codec (CU), the affine parameters are used to update the category HPTCat(RefList) in a manner similar to HMVP table updates. cur ,RefIdx cur ) entries.
[0139] Figure 21 The first and second HPTs are shown. Candidate HPTs based on historical affine parameters (HAPCs) are derived from... Figure 21 The MV of one of the seven neighboring 4x4 blocks, represented as A0, A1, A2, B0, B1, B2, or B3, and a set of affine parameters stored in the corresponding entries in the first HPT are derived. The MV of the neighboring 4x4 blocks is used as the base MV. In a formulaic manner, the current block at position ( x , y The MV at point ) is calculated as: , in( mv h base , mv v base ) represents the MV of the adjacent 4x4 block, ( x base , y base () indicates the center position of the adjacent 4x4 block. x , yThe MV can be the top left, top right, or bottom left corner of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0140] A second historical parameter table (HPT) containing basic MV information is also added. The second HPT contains nine entries, each including the basic MV, reference index, four affine parameters for each reference list, and the base position. An additional Merge HAPC can be generated from the second HPT using the basic MV information stored in the entries and the corresponding affine model. The difference between the first and second HPT lies in... Figure 21 It is shown in the middle.
[0141] Furthermore, paired affine merge candidates are generated from two affine merge candidates, either historically derived or non-historically derived. Paired affine merge candidates are generated by averaging the CPMV of existing affine merge candidates in the list.
[0142] In response to the introduction of the new HAPC, the size of the sub-block-based Merge candidate list was increased from 5 to 15, all of which are involved in the ARMC process.
[0143] In NA-AFF, the pattern for obtaining the nearest neighbors in non-adjacent spatial domains is as follows: Figure 22 As shown. Figure 22 Spatial neighbors used to derive affine Merge / AMVP candidates are shown: (a) candidates for deriving inheritance, and (b) candidates for deriving the first type of construction. Similar to existing non-adjacent regular Merge candidates, the distance between non-adjacent spatial neighbors in NA-AFF and the current codec block is also defined based on the width and height of the current CU.
[0144] Figure 22 Motion information of non-adjacent spatial neighbors in the VVC is used to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for inherited candidates, the same derivation process for inherited affine Merge / AMVP candidates in the VVC remains unchanged, except that CPMV inherits from non-adjacent spatial neighbors. Non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from nearest to farthest). At a specific distance, for the derivation of inherited candidates, only the first available neighbors (encoded in affine mode) from each side (e.g., left and top) of the current block are included. Figure 22 As shown by the dashed arrows in (a), the order of checking the left and upper neighbors is from bottom to top and from right to left, respectively.
[0145] For candidates of the first type of construction, such as Figure 22As shown in (b), the positions of non-adjacent spatial neighbors on the left and top are first determined independently; then, the positions of the upper left neighbors can be determined accordingly, which can form a rectangular virtual block together with the non-adjacent neighbors on the left and top. Then, as... Figure 23 As shown, motion information from three non-adjacent neighbors is used to form a CPMV at the top left (A), top right (B), and bottom left (C) of the virtual block, which is then projected onto the current CU to generate corresponding candidates for construction. Figure 23 Affine Merge / AMVP candidates are shown, ranging from non-nearest neighbors to the first type of construction.
[0146] NA-AFF candidates are inserted into the existing affine Merge candidate list and affine AMVP candidate list in the following order: Affine Merge Mode: 1. SbTMVP candidate, if available.
[0147] 2. Inherit from the nearest neighbor.
[0148] 3. Inherit from non-nearest neighbors.
[0149] 4. Build from neighboring neighbors.
[0150] 5. Affine candidates constructed from the first type of non-near neighbors.
[0151] 6. Zero MV.
[0152] Affine AMVP mode: 1. Inherit from the nearest neighbor.
[0153] 2. Build from neighboring neighbors.
[0154] 3. Translation MV from neighboring units.
[0155] 4. Translation MV from temporal nearest neighbor.
[0156] 5. Inherit from non-adjacent neighbors.
[0157] 6. Affine candidates constructed from the first type of non-near neighbors.
[0158] 7. Zero MV.
[0159] The size of the affine Merge candidate list has increased from 5 to 15 due to the inclusion of additional candidates generated by NA-AFF. The size of the subgroups for ARMCs in the affine Merge pattern has increased from 3 to 15.
[0160] In NA-AFF: 1. The regions from which non-adjacent neighbors originate are restricted to the current CTU (i.e., there are no additional storage requirements for the row cache).
[0161] 2. The storage granularity for affine motion information (including CPMV and reference index) is reduced from 8x8 to 16x16 (i.e., only affine motion from the top-left 8x8 block is stored). Additionally, the stored CPMV is projected onto each 16x16 block before storage, eliminating the need for position and size information.
[0162] 3. Only the top left and top right CPMVs are stored (i.e., always using the 4-parameter affine model for NA-AFF).
[0163] 2.9.3 Derivation of Affine Candidates Based on Regression During the standardization process of VVC, a regression-based method for deriving the motion vector field (RMVF) was proposed, which provides novel sub-block-based merge candidates. For example... Figure 24 As shown, the motion vectors and center positions of the neighboring sub-blocks from the current CU are used as inputs to the linear regression process to derive a set of linear model parameters. Figure 24 A schematic diagram of the neighboring 4x4 sub-blocks used for RMVF parameter derivation is shown, where W and H are the width and height of the current CU.
[0164] A regression-based affine candidate derivation method is proposed. The sub-block motion field from the previously encoded / decoded affine CU and the motion vectors of neighboring sub-blocks from the current CU are used as inputs to the regression process. The difference compared to the regression process in the RMVF derivation method is that the predicted CPMV of the current block is derived instead of the sub-block motion field as the output. The proposed method will be tested in EE2.
[0165] This contribution reports the EE test results of the proposed method on ECM-5.0. A total of 3 tests were performed.
[0166] In test a, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from previously encoded and decoded affine CUs and the motion information of neighboring sub-blocks from the current CU are used as inputs to the regression process to derive the proposed affine candidates.
[0167] Previously encoded and decoded affine CUs can be identified by scanning non-adjacent locations and the affine HMVP table.
[0168] The information of adjacent sub-blocks of the current CU is as follows: Figure 24 The 4x4 sub-blocks, represented by the gray area, are captured. For each sub-block, given a list of references, the corresponding motion vector and center coordinates of the sub-block can be used.
[0169] For each affine CU, at most two affine candidates can be derived. One has neighboring subblock information, and the other does not. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, the ARMC process based on TM cost is applied. Subsequently, when N affine CUs are found, at most N candidates generated by linear regression are added to the affine merge list.
[0170] In test b, the number of affine candidates for ARMC is increased from 15 to 30, while the output list size remains at 15. Finally, in test c, the diversity criterion for ARMC ranking tested in EE2-2.5 is applied on top of test b.
[0171] 2.10. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-image to improve motion vector prediction and merge patterns for CUs in the current image. The same co-image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level; – TMVP obtains temporal motion vectors from co-op blocks in the co-op image (the co-op block is the lower right or center block relative to the current CU), and SbTMVP applies motion displacement before obtaining temporal motion information from the co-op image, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks from the current CU.
[0172] The SbTVMP process is as follows: Figure 25A and Figure 25B As shown. Figure 25A The spatial neighbor block used by ATMVP is shown. Figure 25B This demonstrates how to derive the sub-CU motion field by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, it examines... Figure 25A The spatial nearest neighbor A1 is selected. If A1 has a motion vector that uses a co-located image as its reference image, then that motion vector is chosen as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).
[0173] In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to the position of the current block. Figure 25BThe corresponding image shown obtains motion information (motion vectors and reference indices) at the sub-CU level. Figure 25B The example assumes the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-location image is used to derive the motion information for the sub-CU. After the motion information of the co-location sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.
[0174] In VVC, a sub-block-based combined Merge list containing both SbTVMP candidates and affine Merge candidates is used for signaling in sub-block-based Merge mode. SbTVMP mode is enabled / disabled via Sequence Parameter Set (SPS) flags. If SbTVMP mode is enabled, the SbTVMP prediction is added as the first entry to the sub-block-based Merge candidate list, followed by the affine Merge candidate. The size of the sub-block-based Merge list is transmitted via signaling in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5.
[0175] The sub-CU size used in SbTMVP is fixed at 8x8, and like the affine Merge pattern, the SbTMVP pattern only applies to CUs with a width and height greater than or equal to 8.
[0176] The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, that is, for each CU in the P-strip or B-strip, an additional RD check is performed to determine whether to use the SbTMVP candidate.
[0177] 2.11. Affine MMVD This contribution employs the previously proposed affine MMVD mode. The concept of applying range offsets to MMVD of the fundamental candidate motion is extended to the CPMV of the affine Merge mode. In the affine MMVD mode, there is only one fundamental vector candidate, which is the first affine candidate in the sub-block Merge candidate list, and therefore there is no fundamental vector candidate flag to be transmitted through the signal. There are three sets of range offsets to be selected based on the sequence resolution, as follows: • For sequences smaller than 1280x720, {1 / 16, 1 / 8, 1 / 4, 1 / 2, 1}; • For sequences smaller than 2560x1600, {1 / 4, 1 / 2, 1, 2, 4}; • For sequences greater than or equal to 2560x1600, {1 / 2, 1, 2, 4, 8}.
[0178] Similar to MMVD, affine MMVD also supports four directions for each of its distance offsets (i.e., (+, 0), (-, 0), (0, +), (0, -)). When the base candidate is a unidirectional prediction, the offset vector (i.e., the distance offset multiplied by the direction) is added equally to each CPMV. For bidirectional prediction candidates, the offset vector is added equally to the L0 CPMV, while the offset vector is first mirrored based on the POC distance and then added to the L1 CPMV.
[0179] 2.12. IBC Merge / AMVP List Construction The IBC Merge / AMVP list construction has been modified as follows: • It can be inserted into the IBC Merge / AMVP candidate list only if the IBC Merge / AMVP candidate is valid.
[0180] • Candidate airspace for the upper right, lower left, and upper left (e.g.) Figure 26 The candidates shown (B0, A0, and B2) and a pairwise average candidate can be added to the IBC Merge / AMVP candidate list. Figure 26 The spatial proximity locations used in the construction of the IBC Merge / AMVP list are shown.
[0181] • Template-based adaptive reordering (ARMC-TM) is applied to the IBC Merge list.
[0182] The HMVP table size for IBC is increased to 25. After deriving a maximum of 20 IBC Merge candidates using full deduplication, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list.
[0183] Zero vector candidates filling the IBC Merge / AMVP list are replaced by a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in the IBC Merge mode, and therefore are discarded as BVPs in the IBC candidate list.
[0184] Three candidates are located at the nearest corner of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), their coordinates determined by the width and height of the current block and the ΔX and ΔY parameters, as shown below. Figure 27 As shown. Figure 27 The filling candidates for replacing the zero vector in the IBC list are shown.
[0185] 2.13. IBC with Template Matching Template matching is used in both the IBC Merge mode and the IBC AMVP mode in IBC.
[0186] Compared to the IBC-TM Merge list used in the regular IBC Merge mode, the IBC-TM Merge list is modified so that candidates are selected based on a deduplication method that uses the motion distance between candidates as in the regular TM Merge mode. The zero-motion implementation at the end is replaced by motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0187] In IBC-TM Merge mode, the selected candidate is refined using a template matching method before the RDO or decoding process. IBC-TM Merge mode competes with the regular IBC Merge mode, and the TM-Merge flag is transmitted via signaling.
[0188] In the IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM Merge list. Each of these three selected candidates is refined using a template matching method and ranked according to its resulting template matching cost. Then, only the top two are considered in the motion estimation process as usual.
[0189] Since the IBC motion vector is constrained to (i) integers and (ii) in the form of Figure 28 Within the reference area shown, template matching refinement for both IBC-TM Merge and AMVP modes is quite straightforward. Figure 28 The IBC reference region, depending on the current CU position, is shown. Therefore, in IBC-TM Merge mode, all thinning is performed with integer precision, and in IBC-TMAMVP mode, they are performed with either integer precision or 4-pixel precision, depending on the AMVR value. Such thinning only accesses samples that are not interpolated. In both cases, the motion vectors for thinning and the template used in each thinning step must adhere to the constraints of the reference region.
[0190] 2.14. IBC Merge Mode with Block Vector Difference (IBC-MBVD) Affine-MMVD and GPM-MMVD have been adopted by ECM as extensions to the regular MMVD mode. Extending the MMVD mode to the IBC Merge mode is a natural progression.
[0191] In IBC-MBVD, the distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD direction is two horizontal directions and two vertical directions.
[0192] The base candidates are selected from the top five candidates in the reordered IBC Merge list. Then, all possible MBVD refinement positions (20×4) for each base candidate are reordered based on the SAD cost between the template (the row above and column to the left of the current block) and its reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are retained as available positions for MBVD index encoding and decoding. The MBVD index is binarized using rice codes with a parameter equal to 1.
[0193] Blocks encoded and decoded by IBC-MBVD do not inherit the flip type from neighboring blocks encoded and decoded by RR-IBC.
[0194] 2.15. Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique used to model the local illumination variation between the current block and its predicted block as a function of the local illumination variation between the current block template and the reference block template. The parameters of this function can be scaled... α and offset β This means that they form linear equations, i.e. α *p[x]+ β To compensate for lighting changes, p[x] is a reference sample point at position x on the reference image, pointed to by MV. When surround motion compensation is enabled, MV must account for surround offset for limiting. Because α and β It can be derived based on the current block template and the reference block template, so there is no signaling overhead for them except for signaling the LIC flag to indicate the use of LIC for AMVP mode.
[0195] Local illumination compensation was used for unidirectional prediction of inter-frame CU, and the following modifications were made.
[0196] • Intra-frame neighbor samples can be used for LIC parameter derivation; • For blocks with fewer than 32 luminance samples, LIC is disabled; • For both non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than based on the partial template block samples corresponding to the first top-left 16x16 cell; • Samples of the reference block template are generated using a MC with a block MV, without rounding them to integer pixel precision.
[0197] 2.16. Improvements to Local Illumination Compensation 2.16.1 Bidirectional Predictive LIC In this method, the LIC model is extended to bidirectional prediction of the CU. Specifically, two different linear models are applied to two prediction blocks, and then the two prediction blocks are combined to generate bidirectional prediction samples for the current CU, i.e., , and , , in ,as well as These indicate scaling and offset in L0 and L1, respectively; The weights (as indicated by the CU-level BCW index) are used to indicate the weighted combination used for L0 and L1 predictions. The same derivation scheme for the LIC model is reused and applied iteratively to derive the L0 and L1 LIC parameters. Specifically, the method first minimizes the L0 template prediction... With template The difference between them is used to derive the L0 parameter, and The sample points in the sample are subtracted The corresponding sample points in the dataset are used to update the data. Then, the minimum L1 template prediction is calculated. The L1 parameter represents the difference between the updated template and the current template. Finally, the L0 parameter is refined again in the same way.
[0198] Following the current LIC design, a flag indicating the LIC mode is transmitted via signal transmission for AMVP bidirectional prediction CUs, while this flag is inherited for Merge-related inter-frame CUs. Additionally, the LIC is disabled when decoder-side motion vector refinement (DMVR) (including multi-pass DMVR, adaptive DMVR, and affine DMVR) and bidirectional optical flow (BDOF) are applied.
[0199] 2.16.2 OBMC with LIC In this method, OBMC is enabled for inter-frame blocks encoded and decoded in LIC mode. Furthermore, to reduce complexity, OBMC is only applied to the top CU boundary and the left CU boundary, while the boundaries of internal sub-blocks within a LIC CU are always disabled. Additionally, when a neighboring block is encoded and decoded using LIC, its LIC parameters are applied to generate corresponding prediction samples for the OBMC of the current block.
[0200] 2.17. IBC with local illumination compensation (IBC-LIC) Intra-Block Copy with Local Illumination Compensation (IBC-LIC) is an encoding / decoding tool that uses linear equations to compensate for local illumination variations within an image between a CU encoded using IBC and its predicted blocks. Except for the use of block vectors to generate a reference template in IBC-LIC, the derivation of the parameters of the linear equations is the same as for LIC for inter-frame prediction. IBC-LIC can be applied to both IBC AMVP and IBC Merge modes. For IBC AMVP mode, an IBC-LIC flag is transmitted via signaling to indicate the use of IBC-LIC. For IBC Merge mode, the IBC-LIC flag is inferred from the Merge candidate.
[0201] 2.18. Improvements to IBC-LIC Three additional modes were proposed for IBC-LIC to further improve encoding and decoding performance.
[0202] The first two modes relate to using different template shapes in parameter derivation. In the current ECM, IBC-LIC uses both the top and left templates to derive parameters. A proposed method allows IBC-LIC to derive model parameters using only the top template, only the left template, or both the top and left templates.
[0203] In addition, an extension of MMLM to IBC-LIC is proposed, which allows IBC-LIC to have two linear models in a single CU.
[0204] Finally, a method to remove the large block size constraint for IBC-LIC is proposed. In this contribution, IBC-LIC and the proposed additional mode can be applied to CUs with block sizes greater than 32.
[0205] In the AMVP model, the signaling of the proposed method is summarized in Table 5.
[0206] Table 5 shows the IBC-LIC signaling in the proposed method.
[0207] 2.19. Adaptive OBMC Control OBMC may not be efficient for text and graphics with motion content (such as TGM content). This paper proposes some methods to improve OBMC patterns.
[0208] 1) Method #1: This paper proposes an adaptive determination of the OBMC (On-Board Mode Control) for CUs encoded and decoded via inter-merge mode based on block characteristics. Before OBMC, the number of principal gradients of the CUs encoded and decoded via inter-merge mode is calculated using a gradient histogram based on predicted samples. CUs with a finite number of principal gradients are considered typical SCC (Synchronous Compatibility Control) content, and OBMC is not applied to such CUs. Furthermore, if there are neighboring blocks encoded and decoded using SCC mode, they are also considered typical SCC content, and OBMC is not applied to them.
[0209] 2) Method #2: In this test, the OBMC tool was disabled at the sequence level.
[0210] If there are neighboring blocks encoded or decoded using IBC mode, Palette mode, or BDPCM mode, then the block is considered SCC content.
[0211] 3. Problem For inter-frame merge / affine merge / IBC merge modes, the inter-frame-LIC flag / affine-LIC flag / IBC-LIC flag is inherited from the merge candidate. This may not be entirely accurate.
[0212] 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0213] In the following text, inter-frame merge mode may include, but is not limited to, at least one of the following: regular merge mode, MMVD mode, CIIP mode, TM-Merge mode, CIIP-TM merge mode, ADMVR merge mode, and bilateral matching AMVP-Merge mode; affine merge mode may include, but is not limited to, at least one of the following: affine merge mode, affine-TM merge mode, and affine MMVD mode; IBC merge mode may include, but is not limited to, at least one of the following: IBC merge mode, IBC-TM merge mode, IBC-MBVD mode, IBC-CIIP merge mode, and IBC-GPM merge mode.
[0214] In the following text, a Merge candidate can be any of the Merge patterns listed above.
[0215] In the following text, the LIC flag can be an inter-frame LIC flag, an affine LIC flag, or an IBC-LIC flag.
[0216] In the following text, the LIC indication can be an inter-frame LIC indication, an affine LIC indication, or an IBC-LIC indication.
[0217] In the following text, the width and height of the current block are represented as W and H.
[0218] In the following text, if both the top and left templates of the current block are available, the template size is W+H; if only the top template of the current block is available, the template size is W; if only the left template of the current block is available, the template size is H.
[0219] In one example, a block can be considered SCC content under certain conditions (e.g., if there are neighboring blocks encoded and decoded using IBC mode, Palette mode, or BDPCM mode).
[0220] 1. In one example, the first variable inherited by the Merge candidate can be adjusted according to the criteria.
[0221] a. In one example, the first variable could be the LIC flag.
[0222] b. In one example, the first variable could be the LIC index (e.g., the IBC-LIC indication transmitted via signaling in Section 2.18).
[0223] c. In one example, the criteria may be based on the template of the current block (current template) and / or the template of the reference block (reference template).
[0224] (a) In one example, the criterion could be template matching cost.
[0225] (i). In one example, the cost function between the current template and the reference template can be SAD / MR-SAD, SATD / MR-SATD, SSD / MR-SSD, SSE / MR-SSE, weighted SAD / weighted MR-SAD, weighted SATD / weighted MR-SATD, weighted SSD / weighted MR-SSD, weighted SSE / weighted MR-SSE.
[0226] d. In one example, the first variable (such as the LIC flag for the Merge candidate) can be derived by comparing the template matching costs when the first variable (such as the LIC flag) is true (i.e., LIC on) or false (i.e., LIC off). The template cost corresponding to the LIC flag that is to be false is denoted as tmCost0, and the template cost corresponding to the LIC flag that is to be true is denoted as tmCost1.
[0227] (a) In one example, the reference template may not perform the LIC operation. For LIC on, MRSAD is used as the template matching cost; for LIC off, SAD is used as the template matching cost.
[0228] (b) In one example, SAD is used as the template matching cost for both LIC on and LIC off. And for LIC on, the reference template may need to perform a LIC operation.
[0229] 1) In one example, the top template and the left template can share a set of LIC parameters.
[0230] 2) In one example, the top template has one set of LIC parameters, and the left template has another set of LIC parameters.
[0231] (c) In one example, the template matching cost can be modified.
[0232] 1) In one example, the template matching cost (e.g., tmCost) can be multiplied by a factor.
[0233] i. In one example, the template matching cost of an inherited LIC flag (e.g., tmCostInherited) can be multiplied by a factor (i.e., tmCostInheritedFactor).
[0234] (i) In one example, the factor can be less than 1.
[0235] (ii) In one example, the factor can be 3 / 4, 1 / 2, 1 / 4 or 1 / 8.
[0236] (iii) In one example, at least multiplication and shifting can be used to apply factors.
[0237] a) In one example, tmCost' = (tmCost*F +Offset)>>S, where tmCost' is the template cost after factoring, and F, Offset, and S are integers. tmCost' is the modified template cost.
[0238] a. In one example, Offset = 1 << (S-1).
[0239] b. In one example, tmCostInheritedFactor can be defined as (F+Offset)>>S.
[0240] ii. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different image types.
[0241] (i) In one example, the image type can be a low-latency image and a non-low-latency image.
[0242] (ii) In one example, the image type may depend on whether backward inter-frame prediction is applied to the image.
[0243] a) In one example, backward inter-frame prediction can refer to a current image having a reference image whose POC is greater than that of the current image.
[0244] b) In one example, an image with backward inter-frame prediction can be defined as a non-low-latency image.
[0245] c) In one example, an image without backward inter-frame prediction can be defined as a low-latency image.
[0246] (iii) In one example, the template matching cost factor for the inherited LIC flag for non-low latency images may be less than the template matching cost factor for the inherited LIC flag for low latency images.
[0247] (iv) In one example, for unidirectional prediction of Merge candidates in inter-frame Merge mode, the template matching cost factor for inherited LIC flags for non-low-latency images can be smaller than the template matching cost factor for inherited LIC flags for low-latency images.
[0248] (v) In one example, for one-way prediction of merge candidates in inter-frame merge mode, the template matching cost of the inherited LIC flag can be multiplied by a factor of 1 / 2 for non-low-latency images and by a factor of 3 / 4 for low-latency images.
[0249] (vi) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different image types.
[0250] iii. In one example, the factor (e.g., tmCostInheritedFactor) may have different prediction directions for different Merge candidates.
[0251] (i) In one example, the factor (e.g., tmCostInheritedFactor) can be different for one-way prediction of Merge candidates and two-way prediction of Merge candidates.
[0252] (ii) In one example, the bidirectional factor predicting the Merge candidate (e.g., tmCostInheritedFactor) can be smaller than the unidirectional factor predicting the Merge candidate (e.g., tmCostInheritedFactor).
[0253] (iii) In one example, the factor for bidirectional prediction of Merge candidates (e.g., tmCostInheritedFactor) can be 1 / 4; the factor for unidirectional prediction of Merge candidates (e.g., tmCostInheritedFactor) can be 1 / 2.
[0254] (iv) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different prediction directions.
[0255] iv. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different Merge patterns.
[0256] (i) In one example, the factor (e.g., tmCostInheritedFactor) can be different for inter-frame merge / affine merge / IBC merge modes.
[0257] (ii) In one example, the factor (e.g., tmCostInheritedFactor) may be different for at least two of the following: regular Merge pattern, MMVD pattern, CIIP pattern, TM-Merge pattern, CIIP-TM Merge pattern, ADMVR Merge pattern, and bilateral matching AMVP-Merge pattern.
[0258] (iii) In one example, the factor (e.g., tmCostInheritedFactor) may be different for at least two of the affine Merge mode, affine-TM Merge mode, and affine MMVD mode.
[0259] (iv) In one example, the factor (e.g., tmCostInheritedFactor) may be different for at least two of the IBC Merge mode, IBC-TM Merge mode, and IBC-MBVD mode.
[0260] (v) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different Merge patterns.
[0261] v. In one example, the LIC flags for a factor (e.g., tmCostInheritedFactor) can be different for different inheritances.
[0262] (i) In one example, the tmCostInheritedFactor of a truly inherited LIC flag can be greater than the tmCostInheritedFactor of a falsely inherited LIC flag.
[0263] (ii) Alternatively, the LIC flags for factors (e.g., tmCostInheritedFactor) can be the same for different inheritances.
[0264] vi. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different color components.
[0265] (i) In one example, the factor (e.g., tmCostInheritedFactor) can be different for luminance and chrominance.
[0266] a) In one example, the luminance tmCostInheritedFactor can be greater than the chrominance tmCostInheritedFactor.
[0267] (ii) In one example, the factor (e.g., tmCostInheritedFactor) can be different for Y, Cb, and Cr.
[0268] (iii) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different color components.
[0269] vii. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different block contents.
[0270] (i) In one example, the factor (e.g., tmCostInheritedFactor) can be different for SCC content and non-SCC content.
[0271] a) In one example, the tmCostInheritedFactor of non-SCC content can be greater than the tmCostInheritedFactor of SCC content.
[0272] (ii) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different block contents.
[0273] viii. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different sequence resolutions.
[0274] (i) In one example, the factor for sequence resolution > S can be smaller than the factor for sequence resolution ≤ S.
[0275] a) In one example, S can be 1920x1080.
[0276] b) In one example, the factor for sequence resolution > S can be 1 / 2 or 1 / 4; the factor for sequence resolution ≤ S can be 23 / 32 or 18 / 32.
[0277] (ii) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different sequence resolutions.
[0278] ix. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different block sizes.
[0279] (i) In one example, the factor for block size > M can be smaller than the factor for block size ≤ M.
[0280] a) In one example, M can be 96.
[0281] b) In one example, the factor for block size > M can be 1 / 2 or 1 / 4; the factor for block size ≤ M can be 23 / 32 or 18 / 32.
[0282] (ii) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different block sizes.
[0283] x. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different template sizes.
[0284] (i) In one example, the factor for template size > N can be smaller than the factor for template size ≤ N.
[0285] a) In one example, N can be 96.
[0286] b) In one example, the factor for template size > N can be 1 / 2 or 1 / 4; the factor for template size ≤ N can be 23 / 32 or 18 / 32.
[0287] (ii) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different template sizes.
[0288] xi. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different strip types.
[0289] (iii) Alternatively, the factor (e.g., tmCostInheritedFactor) may be the same for different strip types.
[0290] xii. In one example, the factor (e.g., tmCostInheritedFactor) can be different for different codec configurations.
[0291] (iv) Alternatively, the factor (e.g., tmCostInheritedFactor) can be the same for different codec configurations.
[0292] 2) In one example, the modified tmCostInherited, denoted as tmCostInherited', can be derived as f(tmCostInherited), where f is a function.
[0293] i. tmCostInherited' = tmCostInherited - RightShift(tmCostInherited,4) - RightShift(tmCostInherited, 5).
[0294] ii. tmCostInherited' = tmCostInherited - RightShift(tmCostInherited,3) - RightShift(tmCostInherited, 4).
[0295] iii. tmCostInherited' = tmCostInherited - RightShift(tmCostInherited,2), in which case tmCostInheritedFactor can be considered as 3 / 4.
[0296] iv. tmCostInherited' = tmCostInherited - RightShift(tmCostInherited, 1). In this case, tmCostInheritedFactor can be regarded as 1 / 2.
[0297] v. tmCostInherited' = RightShift(tmCostInherited, 1).
[0298] vi. tmCostInherited' = RightShift(tmCostInherited, 2).
[0299] vii. tmCostInherited' = RightShift(tmCostInherited, 3).
[0300] viii. In one example, RightShift(X, S) = (X + offset) >> S, where the offset is an integer, for example, offset = 1 << (S - 1).
[0301] (d) Whether and / or how to modify the template matching cost can depend on the inherited LIC flag.
[0302] 1) For example, if the inherited LIC flag of the Merge candidate is false, then tmCost0 is modified to tmCost0' = tmCost0 * tmCostInheritedFactor, where tmCostInheritedFactor < M. M is an integer, for example, 1.
[0303] 2) For example, if the inherited LIC flag of the Merge candidate is false, then tmCost1 is modified to tmCost1' = tmCost1 * tmCostInheritedFactor, where tmCostInheritedFactor > M. M is an integer, for example, 1.
[0304] 3) For example, if the inherited LIC flag of the Merge candidate is true, then tmCost0 is modified to tmCost'0 = tmCost0 * tmCostInheritedFactor, where tmCostInheritedFactor > M. M is an integer, for example, 1.
[0305] 4) For example, if the inherited LIC flag of the Merge candidate is true, then tmCost1 is modified to tmCost1' = tmCost1 * tmCostInheritedFactor, where tmCostInheritedFactor < M. M is an integer, for example, 1.
[0306] (e) In one example, if the inherited LIC flag of the Merge candidate is false, the LIC flag of the Merge candidate can be derived by comparing the template matching costs with LIC on and LIC off; if the inherited LIC flag of the Merge candidate is true, the LIC flag of the Merge candidate can be set to its inherited LIC flag.
[0307] 1) Alternatively, if the inherited LIC flag of the Merge candidate is true, the LIC flag of the Merge candidate can be derived by comparing the template matching costs with LIC on and LIC off; if the inherited LIC flag of the Merge candidate is false, the LIC flag of the Merge candidate can be set to its inherited LIC flag.
[0308] 2. In one example, whether and / or how to adjust a first variable such as the LIC flag of the Merge candidate can depend on some factors.
[0309] a. In one example, whether and / or how to adjust a first variable such as the LIC flag of the Merge candidate can depend on color components.
[0310] (a) In one example, the inherited LIC flag of the chrominance component may not be adjusted.
[0311] 1) Alternatively, the inherited LIC flag of the chrominance component can be adjusted.
[0312] (b) In one example, the inherited LIC flag of the luminance component can be adjusted. <0oo00939>
[0313] b. In one example, whether and / or how to adjust a first variable such as the LIC flag of the Merge candidate can depend on picture type.
[0314] (a) In one example, the inherited LIC flag of a non-low-latency picture may not be adjusted.
[0315] 1) Alternatively, the inherited LIC flag of a non-low-latency picture can be adjusted.
[0316] (b) In one example, the inherited LIC flag of a low-latency picture can be adjusted.
[0317] c. In one example, whether and / or how to adjust the first variable, such as the LIC flag for the Merge candidate, can depend on the Merge pattern.
[0318] (a) In one example, the inherited LIC flags of the affine Merge pattern may not be adjusted.
[0319] 1) Alternatively, the inherited LIC flag of the affine Merge pattern can be adjusted.
[0320] (b) In one example, the inherited LIC flag for inter-frame Merge mode can be adjusted.
[0321] (c) In one example, the inherited LIC flag of the IBC Merge pattern can be adjusted.
[0322] d. In one example, whether and / or how to adjust the first variable, such as the LIC flag of the Merge candidate, can depend on the prediction direction of the Merge candidate.
[0323] (a) In one example, the LIC flag for bidirectional prediction of Merge candidates may not be adjusted.
[0324] 1) Alternatively, the LIC flag for the inheritance of the two-way prediction Merge candidate can be adjusted.
[0325] (b) In one example, the LIC flag for unidirectional prediction of Merge candidate inheritance can be adjusted.
[0326] e. In one example, whether and / or how to adjust the first variable of a LIC flag such as a Merge candidate can depend on the inherited LIC flag.
[0327] (a) In one example, the LIC flags that are truly inherited may not be adjusted.
[0328] 1) Alternatively, the true inherited LIC flag can be adjusted.
[0329] (b) In one example, the LIC flag of a fake inheritance can be adjusted.
[0330] f. In one example, whether and / or how to adjust the first variable, such as the LIC flag for the Merge candidate, can depend on the block content.
[0331] (a) In one example, the inherited LIC flags for SCC content may not be adjusted.
[0332] 1) Alternatively, the LIC flag for SCC content inheritance can be adjusted.
[0333] (b) In one example, the inherited LIC flag for non-SCC content can be adjusted.
[0334] g. In one example, the first variable of whether and / or how to adjust the LIC flag such as the Merge candidate can depend on the picture LIC flag.
[0335] (a) In one example, if the picture LIC flag is false, the inherited LIC flag may not be adjusted.
[0336] 1) Alternatively, if the picture LIC flag is false, the inherited LIC flag can be adjusted.
[0337] (b) In one example, if the picture LIC flag is true, the inherited LIC flag can be adjusted.
[0338] h. In one example, the first variable of whether and / or how to adjust the LIC flag such as the Merge candidate can depend on the sequence resolution.
[0339] (a) In one example, for sequence resolution > R1 the inherited LIC flag may not be adjusted.
[0340] 1) In one example, R1 can be 1920x1080.
[0341] 2) Alternatively, for sequence resolution > R1 the inherited LIC flag can be adjusted.
[0342] (b) In one example, for sequence resolution ≤ R2 the inherited LIC flag can be adjusted.
[0343] 1) Alternatively, for sequence resolution < R2, the inherited LIC flag may not be adjusted.
[0344] i. In one example, the first variable of whether and / or how to adjust the LIC flag such as the Merge candidate can depend on the block size.
[0345] (a) In one example, for block size > P1 the inherited LIC flag may not be adjusted.
[0346] 1) Alternatively, for block size > P1 the inherited LIC flag can be adjusted.
[0347] (b) In one example, for block size ≤ P2The inherited LIC flag can be adjusted.
[0348] 1) Alternatively, the inherited LIC flag for a block size < P2 may not be adjusted.
[0349] (a) In one example, for a block size > P1 or a block size < P2, the inherited LIC flag may not be adjusted.
[0350] 1) In one example, P1 is 96; P2 is 16.
[0351] j. In one example, whether and / or how to adjust a first variable of a LIC flag such as a Merge candidate may depend on the template size.
[0352] (a) In one example, the inherited LIC flag for a template size > Q1 may not be adjusted.
[0353] 1) Alternatively, the inherited LIC flag for a template size > Q1 may be adjusted.
[0354] (b) In one example, the inherited LIC flag for a template size ≤ Q2 may be adjusted.
[0355] 1) Alternatively, the inherited LIC flag for a template size < Q2 may not be adjusted. <--
[0356] (c) In one example, the inherited LIC flag for a template size > Q1 or a template size < Q2 may not be adjusted.
[0357] 1) In one example, Q1 is 96; Q2 is 16.
[0358] k. In one example, whether and / or how to adjust a first variable of a LIC flag such as a Merge candidate may depend on the template size and / or the block size.
[0359] (a) In one example, the inherited LIC flag for a block size > T1 or a template size < T2 may not be adjusted.
[0360] 1) In one example, T1 is 96; T2 is 16.
[0361] l. In one example, whether and / or how to adjust a first variable of a LIC flag such as a Merge candidate may depend on the stripe type. [[ID=:45]]
[0362] m. In one example, whether and / or how to adjust a first variable of a LIC flag such as a Merge candidate may depend on the coding configuration.
[0363] n. In one example, whether and / or how to adjust the first variable, such as the LIC flag for the Merge candidate, can depend on any combination of the above items.
[0364] 3. In one example, the LIC flag for the Merge candidate can be adjusted in a predefined step.
[0365] a. In one example, the LIC flag for a Merge candidate can be adjusted directly after the Merge candidate list is built.
[0366] b. In one example, if the Merge candidate list has a reordering operation, the LIC flag of the Merge candidate can be adjusted directly after the Merge candidate list is reordered.
[0367] c. In one example, if the Merge candidate is bidirectionally predicted, the LIC flag of the Merge candidate can be adjusted directly after the BCW index of the Merge candidate is adjusted.
[0368] d. In one example, if the Merge candidate has a DMVR / multiple-pass DMVR procedure, the LIC flag of the Merge candidate can be adjusted before the DMVR / multiple-pass DMVR procedure.
[0369] (a) Alternatively, if the Merge candidate has a DMVR / multiple-pass DMVR procedure, the LIC flag of the Merge candidate may be adjusted after the DMVR / multiple-pass DMVR procedure.
[0370] e. In one example, the LIC flag of a TM-Merge candidate can be adjusted before the template matching refinement process of the TM-Merge candidate.
[0371] (a) Alternatively, the LIC flag of the TM-Merge candidate can be adjusted after the template matching refinement process of the TM-Merge candidate.
[0372] f. In one example, the LIC flag of an IBC-TM Merge candidate can be adjusted before the template matching refinement process of the IBC-TM Merge candidate.
[0373] (a) Alternatively, the LIC flag of an IBC-TM Merge candidate can be adjusted after the template matching refinement process of the IBC-TM Merge candidate.
[0374] g. In one example, the LIC flags for all Merge candidates can be adjusted at both the encoder and decoder.
[0375] h. In one example, at the encoder, the LIC flags of all Merge candidates can be adjusted; at the decoder, the LIC flags of only the selected Merge candidate corresponding to the Merge candidate index transmitted via signaling can be adjusted.
[0376] 4. In one example, for the AMVP mode, if the LIC flag derived by comparing the template matching cost of the final MV is the same as the selected LIC flag of the final MV, then an indicator flag (e.g., true) can be signaled to the decoder instead of the selected LIC flag; otherwise, an opposite indicator flag (e.g., false) can be signaled to the decoder instead of the selected LIC flag.
[0377] a. In one example, the LIC flag of the final MV can be derived in the same way as the LIC flag of the Merge candidate.
[0378] b. In one example, the LIC flag of the final MV can be derived in a slightly different way than the LIC flag of the Merge candidate.
[0379] 5. In one example, whether to adjust the LIC flags of the Merge candidate can be predefined.
[0380] a. Alternatively, whether to adjust the LIC flag of the Merge candidate can be transmitted via signaling.
[0381] 6. In one example, the methods presented in the above projects can be applied to other LIC directives.
[0382] a. In one example, the LIC indication can be a LIC flag, a LIC index, or another form of LIC indication.
[0383] 7. In one example, the methods presented in the above projects can be combined in any way.
[0384] 8. In one example, the methods presented in the above projects can be applied to both the Merge pattern and the AMVP pattern.
[0385] General information 9. The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. These can be signed or unsigned.
[0386] 10. The grammatical elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded using a bypass method.
[0387] 11. The syntax elements (SE) disclosed above can be transmitted conditionally via signals.
[0388] a. SE is transmitted via signal only when the corresponding function is applicable.
[0389] 12. The syntax elements disclosed above can be transmitted via signals at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0390] 13. In the above examples, a block can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU row / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / sub-block of a block / sub-region within a block / any other region containing more than one sample point or pixel.
[0391] 14. Whether and / or how the methods disclosed above can be applied to transmit signals at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0392] 15. Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0393] 16. Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / image type.
[0394] Further embodiments will be described below. Figure 29 A flowchart of a method 2900 for video processing according to an embodiment of the present disclosure is shown. Method 2900 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.
[0395] At box 2910, for the conversion between the current video block and the video bitstream, determine the value of the first variable associated with the Merge candidate of the current video block.
[0396] At box 2920, adjustment information for the first variable is determined based on at least one of the following: template size, sequence resolution, or block size. The adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable.
[0397] At box 2930, a conversion is performed based on the adjustment information. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.
[0398] Method 2900 enables the determination of whether and / or how to adjust the value of a first variable, such as an inherited LIC flag, based on template size, sequence resolution, and / or block size. Encoding / decoding efficiency and / or encoding / decoding effectiveness can thus be improved.
[0399] In some embodiments, the first variable includes parameters inherited from the Local Illumination Compensation (LIC), and the parameters inherited from the LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0400] In some embodiments, the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and the first variable is not adjusted.
[0401] In some embodiments, the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and a first variable is adjusted.
[0402] In some embodiments, the adjustment information is determined based on the template size, wherein the template size is less than or equal to a second threshold, and the first variable is adjusted.
[0403] In some embodiments, the adjustment information is determined based on the template size, and wherein the template size is less than a second threshold, and the first variable is not adjusted.
[0404] In some embodiments, the adjustment information is determined based on the template size, wherein the template size is greater than a first threshold or less than a second threshold, and the first variable is not adjusted.
[0405] In some embodiments, the first threshold is 96 and the second threshold is 16.
[0406] In some embodiments, the sequence resolution is greater than the first threshold resolution, and the first variable is not adjusted.
[0407] In some embodiments, the sequence resolution is greater than a first threshold resolution, and the first variable is adjusted.
[0408] In some embodiments, the first threshold resolution is 1920x1080.
[0409] In some embodiments, the sequence resolution is less than or equal to the second threshold resolution, and the first variable is adjusted.
[0410] In some embodiments, the sequence resolution is less than the second threshold resolution, and the first variable is not adjusted.
[0411] In some embodiments, the block size of the current video block is greater than the first threshold size, and the first variable is not adjusted for the current video block.
[0412] In some embodiments, the block size of the current video block is greater than a first threshold size, and a first variable is adjusted for the current video block.
[0413] In some embodiments, the block size of the current video block is less than or equal to the second threshold size, and the first variable is adjusted for the current video block.
[0414] In some embodiments, the block size of the current video block is smaller than the second threshold size, and the first variable is not adjusted for the current video block.
[0415] In some embodiments, the block size of the current video block is greater than a first threshold size or less than a second threshold size, and the first variable is not adjusted for the current video block.
[0416] In some embodiments, the block size of the current video block is greater than the third threshold size or the template size of the current video block is less than the fourth threshold size, and the first variable is not adjusted for the current video block.
[0417] In some embodiments, the third threshold size is 96, and the fourth threshold size is 16.
[0418] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, the value of a first variable for a merge candidate of a current video block is determined. Adjustment information for the first variable is determined based on at least one of: template size, sequence resolution, or block size. The adjustment information includes at least one of: whether to adjust the value of the first variable, or how to adjust the value of the first variable. The bitstream is generated based on the adjustment information.
[0419] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, the value of a first variable for a merge candidate of the current video block is determined. Adjustment information for the first variable is determined based on at least one of the following: template size, sequence resolution, or block size. The adjustment information includes at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable. The bitstream is generated based on the adjustment information. The bitstream is stored in a non-transitory computer-readable recording medium.
[0420] Figure 30 A flowchart of a method 3000 for video processing according to an embodiment of the present disclosure is shown. Method 3000 is implemented during the conversion between video units or video blocks of a video and a bitstream of the video.
[0421] At box 3010, for the conversion between the current video block and the video bitstream, determine the value of the first variable associated with the Merge candidate of the current video block.
[0422] At box 3020, the first template matching cost when the first variable has a first value and the second template matching cost when the first variable has a second value are determined.
[0423] At box 3030, at least one of the first template matching cost or the second template matching cost is updated based on at least one scaling factor.
[0424] At box 3040, the value of the first variable is adjusted based on a comparison of the first template matching cost and the second template matching cost.
[0425] At box 3050, the conversion is performed based on the adjusted value of the first variable. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.
[0426] Method 3000 enables updating (multiple) template matching costs based on scaling factors and adjusting the value of a first variable based on a comparison of template matching costs.
[0427] In some embodiments, the first variable includes parameters inherited from the Local Illumination Compensation (LIC), and the parameters inherited from the LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0428] In some embodiments, the first scaling factor of the first variable for non-low latency images is smaller than the second scaling factor of the first variable for low latency images.
[0429] In some embodiments, for unidirectional prediction of Merge candidates in inter-frame Merge mode, the first scaling factor of the first variable for non-low-latency images is smaller than the second scaling factor of the first variable for low-latency images.
[0430] In some embodiments, a first scaling factor of the first sequence resolution is less than a second scaling factor of the second sequence resolution, the first sequence resolution is greater than a threshold resolution, and the second sequence resolution is less than or equal to the threshold resolution.
[0431] For example, the threshold resolution is 1920x1080. The first scaling factor can be 1 / 2 or 1 / 4, and the second scaling factor can be 23 / 32 or 18 / 32.
[0432] In some embodiments, the first scaling factor of the first block size is less than the second scaling factor of the second block size, the first block size is greater than a threshold size, and the second block size is less than or equal to the threshold size.
[0433] For example, the threshold size is 96. In some embodiments, the first scaling factor may be 1 / 2 or 1 / 4, and the second scaling factor may be 23 / 32 or 18 / 32.
[0434] In some embodiments, the first scaling factor for the first template size is different from the second scaling factor for the second template size.
[0435] In some embodiments, the first scaling factor of the first template size is less than the second scaling factor of the second template size, the first template size is greater than the threshold size, and the second template size is less than or equal to the threshold size.
[0436] In some embodiments, the threshold size is 96. In some embodiments, the first scaling factor may be 1 / 2 or 1 / 4, and the second scaling factor may be 23 / 32 or 18 / 32.
[0437] In some embodiments, the first scaling factor for the first template size is the same as the second scaling factor for the second template size.
[0438] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus. In this method, the value of a first variable associated with a merge candidate of a current video block of the video is determined. A first template matching cost when the first variable has a first value and a second template matching cost when the first variable has a second value are determined. At least one of the first template matching cost or the second template matching cost is updated based on at least one scaling factor. The value of the first variable is adjusted based on a comparison of the first template matching cost and the second template matching cost. The bitstream is generated based on the adjusted value of the first variable.
[0439] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, the value of a first variable associated with a merge candidate of a current video block of the video is determined. A first template matching cost for a first value and a second template matching cost for a second value of the first variable are determined. At least one of the first template matching cost or the second template matching cost is updated based on at least one scaling factor. The value of the first variable is adjusted based on a comparison of the first template matching cost and the second template matching cost. The bitstream is generated based on the adjusted value of the first variable. The bitstream is stored in a non-transitory computer-readable recording medium.
[0440] Figure 31 A flowchart of a method 3100 for video processing according to an embodiment of the present disclosure is shown. Method 3100 is implemented during the conversion between video units of a video and a bitstream of a video.
[0441] At box 3110, for the conversion between the current video block and the video bitstream, it is determined whether the derived value of the template matching cost of the first variable based on the motion vector of the current video block is the same as the selected value of the first variable of the motion vector, and the current video block is in advanced motion vector prediction mode.
[0442] At box 3120, the indication is based on the determination indicated in the bitstream. This indication indicates whether the derived value of the first variable is the same as the selected value of the first variable.
[0443] At box 3130, the conversion is performed based on an instruction. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.
[0444] Method 3100 enables the inclusion of an indication in the bitstream as to whether the derived value of the first variable is the same as the selected value of the first variable, rather than including the selected value of the first variable in the bitstream. In this way, signaling overhead can be reduced.
[0445] In some embodiments, the bitstream does not include the selected value of the first variable.
[0446] In some embodiments, the first variable includes parameters inherited from the Local Illumination Compensation (LIC), and the parameters inherited from the LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0447] In some embodiments, the derived value of the first variable of the motion vector is determined in a manner used to determine the value of the first variable for the Merge candidate.
[0448] In some embodiments, the derived value of the first variable of the motion vector is determined in a different manner than that used to determine the value of the first variable for the Merge candidate.
[0449] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus. In this method, for a conversion between a current video block and the video bitstream, it is determined whether a derived value of a template matching cost of a first variable based on the motion vector of the current video block is the same as a selected value of the first variable of the motion vector. The current video block is in an advanced motion vector prediction mode. An indication is determined and indicated in the bitstream. This indication indicates whether the derived value is the same as the selected value. The bitstream is generated based on the indication.
[0450] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, for the conversion between a current video block and the video bitstream, it is determined whether a derived value of a template matching cost of a first variable based on the motion vector of the current video block is the same as a selected value of the first variable of the motion vector. The current video block is in an advanced motion vector prediction mode. An indication is determined and indicated in the bitstream. This indication indicates whether the derived value is the same as the selected value. The bitstream is generated based on the indication. The bitstream is stored in a non-transitory computer-readable recording medium.
[0451] In some embodiments, method 2900 and / or method 3000 and / or method 3100 may be applied to at least one of the following: Merge mode, advanced motion vector prediction mode, or intra template matching prediction (intra-TMP) mode.
[0452] In some embodiments, syntax elements in the bitstream are binarized into one of the following: flags, fixed-length codes, exponential Golomb (x) (EG(x)) codes, unary codes, rounded unary codes, or rounded binary codes, and the syntax elements are signed or unsigned.
[0453] In some embodiments, syntax elements in the bitstream are either bypassed or encoded / decoded using at least one context model.
[0454] In some embodiments, a syntax element is included in the bitstream based on at least one condition, which includes the following condition: the function associated with the syntax element is suitable for conversion.
[0455] In some embodiments, syntax elements are included at one of the following: block level, sequence level, picture group level, picture level, strip level, slice group level, or codec structure, wherein the codec structure includes one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0456] In some embodiments, the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of a block, sub-region within a block, or region containing more than one sample point or pixel.
[0457] In some embodiments, whether and / or how method 2900 and / or method 3000 and / or method 3100 are applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, strip level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0458] In some embodiments, whether and / or how to apply method 2900 and / or method 3000 and / or method 3100 is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.
[0459] In some embodiments, whether and / or how to apply method 2900 and / or method 3000 and / or method 3100 is based on encoded information, which includes at least one of the following: block size, color format, single-tree segmentation or dual-tree segmentation, color components, stripe type or picture type.
[0460] Embodiments of this disclosure can be described according to the following entries, and its features can be combined in any reasonable manner.
[0461] Item 1. A method for video processing, comprising: for a conversion between a current video block and a bitstream of the video, determining a value of a first variable associated with a Merge candidate of the current video block; determining adjustment information of the first variable based on at least one of: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and performing the conversion based on the adjustment information.
[0462] Item 2. The method according to Item 1, wherein the first variable includes parameters inherited from Local Illumination Compensation (LIC), and the parameters inherited from LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0463] Item 3. The method according to Item 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and the first variable is not adjusted.
[0464] Item 4. The method according to Item 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and the first variable is adjusted.
[0465] Item 5. The method according to Item 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is less than or equal to a second threshold, and the first variable is adjusted.
[0466] Item 6. The method according to Item 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is less than a second threshold, and the first variable is not adjusted.
[0467] Item 7. The method according to Item 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold or less than a second threshold, and the first variable is not adjusted.
[0468] Item 8. The method according to Item 7, wherein the first threshold is 96 and the second threshold is 16.
[0469] Item 9. The method according to Item 1 or 2, wherein the sequence resolution is greater than the first threshold resolution, and the first variable is not adjusted.
[0470] Item 10. The method according to Item 1 or 2, wherein the sequence resolution is greater than a first threshold resolution, and the first variable is adjusted.
[0471] Item 11. The method according to Item 9 or 10, wherein the first threshold resolution is 1920x1080.
[0472] Item 12. The method according to Item 1 or 2, wherein the sequence resolution is less than or equal to the second threshold resolution, and the first variable is adjusted.
[0473] Item 13. The method according to Item 1 or 2, wherein the sequence resolution is less than the second threshold resolution and the first variable is not adjusted.
[0474] Item 14. The method according to Item 1 or 2, wherein the block size of the current video block is greater than a first threshold size, and the first variable is not adjusted for the current video block.
[0475] Item 15. The method according to Item 1 or 2, wherein the block size of the current video block is greater than a first threshold size, and the first variable is adjusted for the current video block.
[0476] Item 16. The method according to Item 1 or 2, wherein the block size of the current video block is less than or equal to the second threshold size, and the first variable is adjusted for the current video block.
[0477] Item 17. The method according to Item 1 or 2, wherein the block size of the current video block is less than the second threshold size, and the first variable is not adjusted for the current video block.
[0478] Item 18. The method according to Item 1 or 2, wherein the block size of the current video block is greater than a first threshold size or less than a second threshold size, and the first variable is not adjusted for the current video block.
[0479] Item 19. The method according to Item 1 or 2, wherein the block size of the current video block is greater than a third threshold size or the template size of the current video block is less than a fourth threshold size, and the first variable is not adjusted for the current video block.
[0480] Item 20. The method according to Item 18 or 19, wherein the third threshold size is 96 and the fourth threshold size is 16.
[0481] Item 21. A method for video processing, comprising: for a conversion between a current video block and a bitstream of the video, determining a value of a first variable associated with a merge candidate of the current video block; determining a first template matching cost when the first variable is a first value and a second template matching cost when the first variable is a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; and performing the conversion based on the adjusted value of the first variable.
[0482] Item 22. The method according to Item 21, wherein the first variable includes parameters inherited from Local Illumination Compensation (LIC), and the parameters inherited from LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0483] Item 23. The method according to Item 21 or 22, wherein the first scaling factor of the first variable for the non-low-latency image is less than the second scaling factor of the first variable for the low-latency image.
[0484] Item 24. The method according to Item 21 or 22, wherein for a one-way predicted merge candidate in an inter-frame merge mode, the first variable is scaled by a first scaling factor for a non-low-latency image less than the first variable is scaled by a second scaling factor for a low-latency image.
[0485] Item 25. The method according to Item 21 or 22, wherein a first scaling factor of a first sequence resolution is less than a second scaling factor of a second sequence resolution, the first sequence resolution is greater than a threshold resolution, and the second sequence resolution is less than or equal to the threshold resolution.
[0486] Item 26. The method according to Item 25, wherein the threshold resolution is 1920x1080.
[0487] Item 27. The method according to Item 25 or 26, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
[0488] Item 28. The method according to Item 21 or 22, wherein a first scaling factor of a first block size is less than a second scaling factor of a second block size, the first block size is greater than a threshold size, and the second block size is less than or equal to the threshold size.
[0489] Item 29. The method according to Item 28, wherein the threshold size is 96.
[0490] Item 30. The method according to Item 28 or 29, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
[0491] Item 31. The method according to Item 21 or 22, wherein the first scaling factor for the first template size is different from the second scaling factor for the second template size.
[0492] Item 32. The method according to Item 31, wherein the first scaling factor of the first template size is less than the second scaling factor of the second template size, the first template size is greater than a threshold size, and the second template size is less than or equal to the threshold size.
[0493] Item 33. The method according to Item 32, wherein the threshold size is 96.
[0494] Item 34. The method according to Item 32 or 33, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
[0495] Item 35. The method according to Item 21 or 22, wherein the first scaling factor for the first template size is the same as the second scaling factor for the second template size.
[0496] Item 36. A method for video processing, comprising: for a conversion between a current video block and a bitstream of the video, determining whether a derived value of a template matching cost of a first variable based on a motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in an advanced motion vector prediction mode; indicating an indication in the bitstream based on the determination, the indication indicating whether the derived value is the same as the selected value; and performing the conversion based on the indication.
[0497] Item 37. The method according to Item 36, wherein the selected value of the first variable is not included in the bitstream.
[0498] Item 38. The method according to Item 36 or 37, wherein the first variable includes parameters inherited from Local Illumination Compensation (LIC), and the parameters inherited from LIC include at least one of the following: LIC flag, LIC index, or LIC indication, wherein the LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
[0499] Item 39. The method according to any one of items 36 to 38, wherein the derived value of the first variable of the motion vector is determined in a manner used to determine the value of the first variable for the Merge candidate.
[0500] Item 40. The method according to any one of items 36 to 38, wherein the derived value of the first variable of the motion vector is determined in a manner different from that used to determine the value of the first variable for the Merge candidate.
[0501] Item 41. The method according to any one of Items 1 to 40, wherein the method is applied to at least one of: Merge mode, advanced motion vector prediction mode, or intra template matching prediction (intra-TMP) mode.
[0502] Item 42. The method according to any one of items 1 to 41, wherein the syntax elements in the bitstream are binarized into one of the following: a flag, a fixed-length code, an exponential Golomb (x) (EG(x)) code, a unary code, a rounded unary code, or a rounded binary code, and the syntax elements are signed or unsigned.
[0503] Item 43. The method according to any one of items 1 to 42, wherein the syntax elements in the bitstream are either bypassed or encoded / decoded using at least one context model.
[0504] Item 44. The method according to any one of items 1 to 43, wherein a syntax element is included in the bitstream based on at least one condition, the at least one condition including the condition that the function associated with the syntax element is applicable to the conversion.
[0505] Item 45. The method according to any one of items 1 to 44, wherein the syntax element is included at one of the following: block level, sequence level, picture group level, picture level, stripe level, slice group level, or codec structure, said codec structure including one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.
[0506] Item 46. The method according to any one of items 1 to 45, wherein the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
[0507] Item 47. The method according to any one of items 1 to 46, wherein whether and / or how the method is applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, stripe level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header or slice group header.
[0508] Item 48. The method according to any one of items 1 to 46, wherein whether and / or how the method is applied includes a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a strip, a slice, a sub-picture, or a region containing more than one sample point or pixel.
[0509] Item 49. The method according to any one of items 1 to 46, wherein whether and / or how the method is applied is based on encoded information, said encoded information including at least one of the following: block size, color format, single-tree segmentation or dual-tree segmentation, color components, stripe type or picture type.
[0510] Item 50. The method according to any one of items 1 to 49, wherein the conversion includes encoding the current video block into the bitstream.
[0511] Item 51. The method according to any one of items 1 to 49, wherein the conversion includes decoding the current video block from the bitstream.
[0512] Item 52. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 51.
[0513] Item 53. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 51.
[0514] Item 54. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining a value of a first variable associated with a merge candidate of a current video block of the video; determining adjustment information of the first variable based on at least one of: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and generating the bitstream based on the adjustment information.
[0515] Item 55. A method for storing a bitstream of video, comprising: determining a value of a first variable associated with a Merge candidate of a current video block of the video; determining adjustment information of the first variable based on at least one of: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of: whether to adjust the value of the first variable, or how to adjust the value of the first variable; generating the bitstream based on the adjustment information; and storing the bitstream in a non-transitory computer-readable recording medium.
[0516] Item 56. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of an apparatus for video processing, wherein the method includes: determining a value of a first variable associated with a merge candidate of a current video block of the video; determining a first template matching cost when the first variable is a first value and a second template matching cost when the first variable is a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; and generating the bitstream based on the adjusted value of the first variable.
[0517] Item 57. A method for storing a bitstream of video, comprising: determining a value of a first variable associated with a merge candidate of a current video block of the video; determining a first template matching cost when the first variable is a first value and a second template matching cost when the first variable is a second value; updating at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; adjusting the value of the first variable based on a comparison of the first template matching cost and the second template matching cost; generating the bitstream based on the adjusted value of the first variable; and storing the bitstream in a non-transitory computer-readable recording medium.
[0518] Item 58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: for a conversion between a current video block of the video and the bitstream of the video; determining whether a derived value of a first variable based on a template matching cost of a motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in an advanced motion vector prediction mode; indicating an indication in the bitstream based on the determination, the indication indicating whether the derived value is the same as the selected value; and generating the bitstream based on the indication.
[0519] Item 59. A method for storing a bitstream of video, comprising: for a conversion between a current video block and the bitstream of the video, determining whether a derived value of a template matching cost of a first variable based on a motion vector of the current video block is the same as a selected value of the first variable of the motion vector, the current video block being in an advanced motion vector prediction mode; indicating an indication in the bitstream based on the determination, the indication indicating whether the derived value is the same as the selected value; generating the bitstream based on the indication; and storing the bitstream in a non-transitory computer-readable recording medium.
[0520] Example device Figure 32 A block diagram of a computing device 3200 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3200 may be implemented as a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300).
[0521] It should be understood that, Figure 32 The computing device 3200 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0522] like Figure 32 As shown, computing device 3200 includes general-purpose computing device 3200. Computing device 3200 may include at least one or more processors or processing units 3210, memory 3220, storage unit 3230, one or more communication units 3240, one or more input devices 3250, and one or more output devices 3260.
[0523] In some embodiments, the computing device 3200 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3200 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0524] Processing unit 3210 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 3220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 3200. Processing unit 3210 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0525] Computing device 3200 typically includes various computer storage media. Such media can be any media accessible by computing device 3200, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 3220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 3230 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 3200.
[0526] The computing device 3200 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 32 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0527] Communication unit 3240 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 3200 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 3200 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0528] Input device 3250 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3260 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 3240, computing device 3200 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 3200 can also communicate with one or more devices that enable a user to interact with computing device 3200, or any device that enables computing device 3200 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0529] In some embodiments, some or all components of computing device 3200 may not be integrated into a single device, but may be deployed in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to perform the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to the user. Thus, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers, or directly installed or otherwise installed on client devices.
[0530] In embodiments of this disclosure, computing device 3200 can be used to implement video encoding / decoding. Memory 3220 may include one or more video codec modules 3225 having one or more program instructions. These modules can be accessed and executed by processing unit 3210 to perform the functions of the various embodiments described herein.
[0531] In an example embodiment of performing video encoding, input device 3250 may receive video data as input 3270 to be encoded. The video data may be processed, for example, by video codec module 3225 to generate an encoded bitstream. The encoded bitstream may be provided as output 3280 via output device 3260.
[0532] In an example embodiment of performing video decoding, input device 3250 may receive an encoded bitstream as input 3270. The encoded bitstream may be processed, for example, by a video codec module 3225 to generate decoded video data. The decoded video data may be provided as output 3280 via output device 3260.
[0533] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine the value of a first variable associated with the Merge candidate of the current video block; The adjustment information for the first variable is determined based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and The conversion is performed based on the adjustment information.
2. The method of claim 1, wherein the first variable includes inherited parameters of local illumination compensation (LIC), and the inherited parameters of the LIC include at least one of the following: LIC mark LIC index, or LIC instruction, The LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
3. The method according to claim 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and the first variable is not adjusted.
4. The method according to claim 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold, and the first variable is adjusted.
5. The method according to claim 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is less than or equal to a second threshold, and the first variable is adjusted.
6. The method according to claim 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is less than a second threshold, and the first variable is not adjusted.
7. The method according to claim 1 or 2, wherein the adjustment information is determined based on the template size, and wherein the template size is greater than a first threshold or less than a second threshold, and the first variable is not adjusted.
8. The method of claim 7, wherein the first threshold is 96 and the second threshold is 16.
9. The method according to claim 1 or 2, wherein the sequence resolution is greater than the first threshold resolution, and the first variable is not adjusted.
10. The method of claim 1 or 2, wherein the sequence resolution is greater than the first threshold resolution, and the first variable is adjusted.
11. The method of claim 9 or 10, wherein the first threshold resolution is 1920x1080.
12. The method of claim 1 or 2, wherein the sequence resolution is less than or equal to the second threshold resolution, and the first variable is adjusted.
13. The method according to claim 1 or 2, wherein the sequence resolution is less than the second threshold resolution, and the first variable is not adjusted.
14. The method according to claim 1 or 2, wherein the block size of the current video block is greater than the first threshold size, and the first variable is not adjusted for the current video block.
15. The method of claim 1 or 2, wherein the block size of the current video block is greater than a first threshold size, and the first variable is adjusted for the current video block.
16. The method of claim 1 or 2, wherein the block size of the current video block is less than or equal to the second threshold size, and the first variable is adjusted for the current video block.
17. The method according to claim 1 or 2, wherein the block size of the current video block is less than the second threshold size, and the first variable is not adjusted for the current video block.
18. The method according to claim 1 or 2, wherein the block size of the current video block is greater than a first threshold size or less than a second threshold size, and the first variable is not adjusted for the current video block.
19. The method according to claim 1 or 2, wherein the block size of the current video block is greater than a third threshold size or the template size of the current video block is less than a fourth threshold size, and the first variable is not adjusted for the current video block.
20. The method of claim 18 or 19, wherein the third threshold size is 96 and the fourth threshold size is 16.
21. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine the value of a first variable associated with the Merge candidate of the current video block; Determine the first template matching cost when the first variable has a first value and the second template matching cost when the first variable has a second value; Update at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; The value of the first variable is adjusted based on a comparison between the first template matching cost and the second template matching cost; as well as The transformation is performed based on the adjusted value of the first variable.
22. The method of claim 21, wherein the first variable includes inherited parameters of local illumination compensation (LIC), and the inherited parameters of the LIC include at least one of the following: LIC mark LIC index, or LIC instruction, The LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
23. The method of claim 21 or 22, wherein the first scaling factor of the first variable for the non-low-latency image is less than the second scaling factor of the first variable for the low-latency image.
24. The method of claim 21 or 22, wherein for the unidirectional prediction of merge candidates in the inter-frame merge mode, the first scaling factor of the first variable for non-low-latency images is smaller than the second scaling factor of the first variable for low-latency images.
25. The method of claim 21 or 22, wherein a first scaling factor of the first sequence resolution is less than a second scaling factor of the second sequence resolution, the first sequence resolution is greater than a threshold resolution, and the second sequence resolution is less than or equal to the threshold resolution.
26. The method of claim 25, wherein the threshold resolution is 1920x1080.
27. The method of claim 25 or 26, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
28. The method of claim 21 or 22, wherein a first scaling factor of the first block size is less than a second scaling factor of the second block size, the first block size is greater than a threshold size, and the second block size is less than or equal to the threshold size.
29. The method of claim 28, wherein the threshold size is 96.
30. The method of claim 28 or 29, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
31. The method of claim 21 or 22, wherein the first scaling factor for the first template size is different from the second scaling factor for the second template size.
32. The method of claim 31, wherein the first scaling factor of the first template size is less than the second scaling factor of the second template size, the first template size is greater than a threshold size, and the second template size is less than or equal to the threshold size.
33. The method of claim 32, wherein the threshold size is 96.
34. The method of claim 32 or 33, wherein the first scaling factor is 1 / 2 or 1 / 4, and the second scaling factor is 23 / 32 or 18 / 32.
35. The method of claim 21 or 22, wherein the first scaling factor for the first template size is the same as the second scaling factor for the second template size.
36. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine whether the derived value of the template matching cost of the first variable based on the motion vector of the current video block is the same as the selected value of the first variable of the motion vector, and the current video block is in advanced motion vector prediction mode; Based on the determination, an indication is given in the bitstream, the indication indicating whether the derived value is the same as the selected value; and The conversion is performed based on the instructions.
37. The method of claim 36, wherein the selected value of the first variable is not included in the bitstream.
38. The method of claim 36 or 37, wherein the first variable includes inherited parameters of local illumination compensation (LIC), and the inherited parameters of the LIC include at least one of the following: LIC mark LIC index, or LIC instruction, The LIC includes at least one of the following: inter-frame LIC, affine LIC, or intra-block copy (IBC) with LIC (IBC-LIC).
39. The method of any one of claims 36 to 38, wherein the derived value of the first variable of the motion vector is determined in a manner used to determine the value of the first variable for the Merge candidate.
40. The method of any one of claims 36 to 38, wherein the derived value of the first variable of the motion vector is determined in a manner different from that used to determine the value of the first variable for the Merge candidate.
41. The method according to any one of claims 1 to 40, wherein the method is applied to at least one of: Merge mode, advanced motion vector prediction mode, or intra-frame template matching prediction (intra-frame TMP) mode.
42. The method according to any one of claims 1 to 41, wherein the syntax elements in the bitstream are binarized into one of the following: flags, fixed-length codes, exponential Golomb (x) (EG(x)) codes, unary codes, rounded unary codes, or rounded binary codes, and the syntax elements are signed or unsigned.
43. The method according to any one of claims 1 to 42, wherein the syntax elements in the bitstream are either bypassed or encoded / decoded using at least one context model.
44. The method according to any one of claims 1 to 43, wherein a syntax element is included in the bitstream based on at least one condition, the at least one condition including the condition that the function associated with the syntax element is applicable to the conversion.
45. The method according to any one of claims 1 to 44, wherein the syntax element is included at one of the following: block level, sequence level, picture group level, picture level, stripe level, slice group level, or codec structure, said codec structure including one of the following: codec tree unit (CTU), codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.
46. The method according to any one of claims 1 to 45, wherein the current video block refers to one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), block, sub-block of block, sub-region within block, or region containing more than one sample point or pixel.
47. The method according to any one of claims 1 to 46, wherein whether and / or how the method is applied is included in the bitstream at one of the following: sequence level, picture group level, picture level, stripe level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header or slice group header.
48. The method according to any one of claims 1 to 46, wherein whether and / or how the method is applied is included in one of the following: Prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-image, or region containing more than one sample point or pixel.
49. The method according to any one of claims 1 to 46, wherein whether and / or how the method is applied is based on encoded information, the encoded information including at least one of the following: block size, color format, single-tree segmentation or dual-tree segmentation, color components, stripe type or picture type.
50. The method of any one of claims 1 to 49, wherein the conversion comprises encoding the current video block into the bitstream.
51. The method according to any one of claims 1 to 49, wherein the conversion comprises decoding the current video block from the bitstream.
52. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 51.
53. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 51.
54. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine the value of the first variable associated with the Merge candidate of the current video block of the video; The adjustment information for the first variable is determined based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; and The bit stream is generated based on the adjustment information.
55. A method for storing a bitstream of video, comprising: Determine the value of the first variable associated with the Merge candidate of the current video block of the video; The adjustment information for the first variable is determined based on at least one of the following: template size, sequence resolution, or block size, wherein the adjustment information indicates at least one of the following: whether to adjust the value of the first variable, or how to adjust the value of the first variable; The bitstream is generated based on the adjustment information; and The bitstream is stored in a non-transitory computer-readable recording medium.
56. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Determine the value of the first variable associated with the Merge candidate of the current video block of the video; Determine the first template matching cost when the first variable has a first value and the second template matching cost when the first variable has a second value; Update at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; The value of the first variable is adjusted based on a comparison between the first template matching cost and the second template matching cost; as well as The bitstream is generated based on the adjusted value of the first variable.
57. A method for storing a bitstream of video, comprising: Determine the value of the first variable associated with the Merge candidate of the current video block of the video; Determine the first template matching cost when the first variable has a first value and the second template matching cost when the first variable has a second value; Update at least one of the first template matching cost or the second template matching cost based on at least one scaling factor; The value of the first variable is adjusted based on a comparison between the first template matching cost and the second template matching cost; The bitstream is generated based on the adjusted value of the first variable; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
58. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: For the conversion between the current video block and the bitstream of the video, determine whether the derived value of the first variable of the template matching cost based on the motion vector of the current video block is the same as the selected value of the first variable of the motion vector, and the current video block is in advanced motion vector prediction mode; Based on the determination, an indication is given in the bitstream, the indication indicating whether the derived value is the same as the selected value; as well as The bit stream is generated based on the instruction.
59. A method for storing a bitstream of video, comprising: For the conversion between the current video block and the bitstream of the video, determine whether the derived value of the template matching cost of the first variable based on the motion vector of the current video block is the same as the selected value of the first variable of the motion vector, and the current video block is in advanced motion vector prediction mode; Based on the determination, an indication is given in the bitstream, the indication indicating whether the derived value is the same as the selected value; The bit stream is generated based on the instruction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.