Methods for constructing merge candidate list

By optimizing the insertion order and number of spatial merge candidates in video coding, the method enhances coding efficiency and compression performance in advanced standards like VVC.

JP2025178392APending Publication Date: 2025-12-05ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025162850
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-19
Filing Date
2025-09-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in optimizing the construction of merge candidate lists, particularly in advanced standards like VVC, which can impact coding efficiency and compression performance.

Method used

The method involves inserting spatial merge candidates into a merge candidate list in a specific order based on coding modes and picture types, such as upper, left, and upper-left neighboring blocks, and adjusting the number of candidates based on numeric limits, to enhance the construction of spatial merge candidates for improved coding efficiency.

Benefits of technology

This approach improves coding efficiency and compression performance by refining the order and number of spatial merge candidates, leading to better video coding outcomes in standards like VVC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025178392000001_ABST
    Figure 2025178392000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for constructing a merge candidate list used for video processing.SOLUTION: A method includes: inserting a set of spatial merge candidates to a merge candidate list of a coding block, where the set of spatial merge candidates are inserted according to the order of a top neighboring block, a left neighboring block, a top neighboring block, a left neighboring block and an above-left neighboring block. The method can further include adding to the merge candidate list at least one of: a temporal merge candidate from collocated coding units; a history-based motion vector predictor (HMVP) from a First-In, First-Out (FIFO) table; a pairwise average candidate; or a zero motion vector.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure is incorporated herein by reference in its entirety. This application claims priority to U.S. Provisional Patent Application No. 62 / 902,790, filed on the same date. [Background technology]

[0002] background

[0002] Video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is commonly referred to as encoding, and the decompression process is commonly referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify particular video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0003] Disclosure Overview

[0003] Embodiments of the present disclosure provide a method for constructing a merge candidate list. According to some embodiments, one exemplary method includes inserting a set of spatial merge candidates into a merge candidate list of a coded block, where the set of spatial merge candidates is inserted in the following order: upper neighboring block, left neighboring block, upper neighboring block, left neighboring block, and upper-left neighboring block.

[0004] According to some embodiments, one exemplary method is to calculate the number of times ... The method includes inserting a set of spatial merge candidates into a merge candidate list for the coded block based on the order of upper neighbor block, left neighbor block if the numeric limit is 2, and inserting a set of spatial merge candidates into a merge candidate list based on the order of upper neighbor block, left neighbor block, top neighbor block if the numeric limit is 3.

[0005] According to some embodiments, one exemplary method includes: The method includes inserting a set of spatial merge candidates into the complementary list, wherein when a first coding mode is applied to the coding block, the set of spatial merge candidates is inserted according to a first construction order, and when a second coding mode is applied to the coding block, the set of spatial merge candidates is inserted according to a second construction order, the first construction order being different from the second construction order.

[0006] According to some embodiments, one exemplary method includes: The method includes inserting a set of spatial merge candidates into the complementary list, where if the coded block is part of a low-delay picture, the set of spatial merge candidates is inserted according to a first construction order, and if the coded block is part of a non-low-delay picture, the set of spatial merge candidates is inserted according to a second construction order, and the first construction order is different from the second construction order.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and various aspects of the present disclosure are set forth in the following detailed description and accompanying drawings. The various features shown therein are not drawn to scale. [Brief explanation of the drawings]

[0008] [Figure 1]8 illustrates the structure of an exemplary video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0009] 1 shows a schematic diagram of an exemplary encoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B]

[0010] 1 shows a schematic diagram of another exemplary encoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3A]

[0011] 1 shows a schematic diagram of an exemplary decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3B]

[0012] 1 shows a schematic diagram of another exemplary decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 4A]

[0013] 1 shows a block diagram of an exemplary device for encoding or decoding video consistent with embodiments of the present disclosure. [Figure 4B]

[0014] 1 illustrates exemplary locations of spatial merge candidates consistent with embodiments of the present disclosure. [Figure 4C]

[0015] 10 illustrates exemplary locations of temporal merge candidates consistent with embodiments of the present disclosure. [Figure 5]

[0016] 10 illustrates an example scaling of temporal merge candidates consistent with embodiments of the present disclosure. [Figure 6]

[0017] 1 illustrates an example relationship between distance index and a predetermined offset in merge mode with motion vector difference (MMVD), consistent with embodiments of the present disclosure. [Figure 7]

[0018] 1 illustrates exemplary locations of spatial merge candidates consistent with embodiments of the present disclosure. [Figure 8]

[0019] 1 shows exemplary experimental results compared to VTM-6 under a random access (RA) configuration, consistent with embodiments of the present disclosure. [Figure 9]

[0020] 1 shows exemplary experimental results compared to VTM-6 under a low latency (LD) configuration, consistent with embodiments of the present disclosure. [Figure 10]

[0021] 10 shows exemplary experimental results compared to VTM-6 under RA configuration, consistent with embodiments of the present disclosure. [Figure 11]

[0022] 10 shows exemplary experimental results compared to VTM-6 under LD configuration, consistent with embodiments of the present disclosure. [Figure 12]

[0023] 1 illustrates an example syntax structure of a slice header, consistent with embodiments of the present disclosure. [Figure 13]

[0024] 1 illustrates an example syntax structure of a sequence parameter set (SPS), consistent with embodiments of the present disclosure. [Figure 14]

[0025] 1 illustrates an example syntax structure of a Picture Parameter Set (PPS), consistent with embodiments of this disclosure. [Figure 15]

[0026] 10 shows exemplary experimental results compared to VTM-6 under RA configuration, consistent with embodiments of the present disclosure. [Figure 16]

[0027] 10 shows exemplary experimental results compared to VTM-6 under LD configuration, consistent with embodiments of the present disclosure. [Figure 17]

[0028] 10 shows exemplary experimental results compared to VTM-6 under RA configuration, consistent with embodiments of the present disclosure. [Figure 18]

[0029] 10 shows exemplary experimental results compared to VTM-6 under LD configuration, consistent with embodiments of the present disclosure. [Figure 19]

[0030] 1 shows a flowchart of an exemplary video processing method consistent with embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description

[0031] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. This description refers to the accompanying drawings in which, unless otherwise indicated, like numerals in different figures represent the same or similar elements. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with aspects related to the present disclosure, as recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0010]

[0032] As mentioned above, images are frames arranged in chronological order to help us remember visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display these pictures in chronological order. Furthermore, in some applications, the video capture device can transmit the captured videos to a video playback device (e.g., a computer with a monitor) in real time, such as for surveillance, conferencing, or live broadcasting.

[0011]

[0033] To reduce the storage space and transmission bandwidth required for such applications, the video is stored in a The video may be compressed before transmission and decompressed before display. This compression and decompression may be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." The encoder and decoder may collectively be referred to as a "codec." The encoder and decoder may be implemented as various suitable hardware, software, or combinations thereof. For example, a hardware implementation of an encoder and decoder may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. A software implementation of an encoder and decoder may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be referred to as a "transcoder."

[0012]

[0034] The video encoding process produces useful information that can be used to reconstruct the picture. Information that is not important for reconstruction can be identified and preserved, and information that is not important for reconstruction can be ignored. If the ignored, unimportant information cannot be perfectly reconstructed, such an encoding process can be called "lossy." Otherwise, such an encoding process can be called "lossless." Most encoding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0013]

[0035] The useful information of the picture being coded (called the "current picture") is The changes include changes to a picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are the most important. Changes in the positions of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.

[0014]

[0036] A picture that is coded without reference to another picture (i.e., such a picture A picture that is coded using a past picture as a reference picture (i.e., the reference is "bidirectional") is called an "I-picture." A picture that is coded using a past picture as a reference picture is called a "P-picture." A picture that is coded using both past and future pictures as reference pictures (i.e., the reference is "bidirectional") is called a "B-picture."

[0015]

[0037] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth To achieve this, JVET is using the Joint Exploration Model (JEM) reference software to develop technology that surpasses HEVC. Because coding techniques are incorporated into JEM, JEM achieves significantly higher coding performance than HEVC.

[0016]

[0038] The VVC standard continues to include more coding techniques that give better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263. In VVC, a merge candidate list can be constructed, including new merge candidates. Different merge list sizes apply for different inter modes. Embodiments of the present disclosure take into account new merge candidates (e.g., HMVP, pairwise average) and new inter modes (e.g., MMVD, TPM) within VVC. For example, the order of spatial merge candidates can be refined, and the number of spatial merge candidates can be adjusted. Furthermore, the construction of spatial merge candidates is fixed for normal mode, MMVD mode, and TPM mode, and the construction of spatial merge candidates is fixed for normal mode, MMVD mode, and TPM mode.

[0017]

[0039] FIG. 1 illustrates an example structure of a video sequence 100 according to some embodiments of the present disclosure. 1 shows a video sequence 100. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0018]

[0040] As shown in FIG. 1, a video sequence 100 includes pictures 102, 104, 106, 108, 109, 110A, 110B, 110C, 110D, 110E, 110F, 110H, 110I, 110J, 110K, 110KN, 110KP, 110KR, 110 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the example shown in FIG.

[0019]

[0041] Typically, video codecs do not encode or decode the entire picture at once. This is because such a task is computationally complex. Instead, video codecs can divide pictures into elementary segments and encode or decode pictures segment by segment. In this disclosure, such elementary segments are referred to as Basic Processing Units ("BPUs"). For example, structure 110 in Figure 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units may have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail one wishes to preserve within the basic processing unit.

[0020]

[0042] The basic processing unit is in computer memory (e.g., in a video frame buffer) A basic processing unit may be a logical unit that may include various types of video data stored in a memory. For example, a basic processing unit for a color picture may include a luma component (Y) that represents achromatic luminance information, one or more chroma components (e.g., Cb and Cr) that represent color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.

[0021]

[0043] Video encoding involves several stages of operation, examples of which are shown in FIGS. 2A-2B and 3. 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size as the basic processing units. Like the basic processing units, the basic processing sub-units are also logical units that may contain various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division may be made to further levels depending on the processing needs. It should also be noted that different stages may use different schemes to divide the basic processing unit.

[0022]

[0044] For example, in the mode decision stage (one example of which is detailed in FIG. 2B), the basic process The encoder can decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a unit, and a basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.

[0023]

[0045] In another example, during the prediction stage (one example of which is detailed in FIGS. 2A-2B), The encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC). At this level, prediction operations can be performed.

[0024]

[0046] In another example, in the conversion step (one example of which is detailed in FIGS. 2A-2B), the code The encoder can perform a transform operation on the residual elementary processing sub-unit (e.g., CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder can further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks for the same CU may have different sizes and numbers.

[0025]

[0047] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.

[0026]

[0048] In some implementations, video encoding and decoding is provided with parallel processing and error resilience capabilities. To provide this functionality, a picture can be divided into regions for processing, thereby enabling the encoding or decoding process for a region of a picture to be independent of information about any other region of the picture. In other words, each region of a picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thus increasing coding efficiency. Furthermore, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that various pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0027]

[0049] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118. 1 are defined within structure 110, the boundaries of which are shown as solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are merely examples, and the present disclosure is not limited to such embodiments.

[0028]

[0050] FIG. 2A illustrates an example of an encoding process 200A consistent with embodiments of the present disclosure. 2A shows a simplified diagram. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode video sequence 202 into video bitstream 228 according to process 200A. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of the basic processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions (e.g., regions 114-118) of each original picture of video sequence 202.

[0029]

[0051] In FIG. 2A, the encoder processes the basic processing units of the original pictures of the video sequence 202. (referred to as the "original BPU") may be fed to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path." During process 200A, after quantization stage 214, the encoder may feed quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to predicted BPU 208 to generate a prediction reference 224 used in prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0030]

[0052] The encoder performs process 200A iteratively to extract the original peaks (in the forward path). The encoder may encode each original BPU of the original picture and generate a predicted reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0031]

[0053] Referring to process 200A, the encoder receives a video capture device (e.g., a camera) As used herein, the term "receive" may refer to receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any other action to input data.

[0032]

[0054] In the prediction step 204, in the current iteration, the encoder uses the original BPU and the prediction reference 224 and may perform a prediction operation to generate predicted data 206 and predicted BPU 208. Prediction reference 224 may be generated from a reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting predicted data 206, which may be used to reconstruct the original BPU as predicted BPU 208 from prediction data 206 and prediction reference 224.

[0033]

[0055] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract the values ​​(e.g., grayscale or RGB values) of pixels of predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant loss of quality.

[0034]

[0056] To further compress the residual BPU 210, in the transform stage 212, the encoder , the spatial redundancy of the residual BPU 210 can be reduced by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each of which is associated with a "transform coefficient" The basis patterns may have the same size (e.g., the size of the residual BPU 210). Each basis pattern may represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition may decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0035]

[0057] Different transformation algorithms can use different base patterns. Various transform algorithms can be used in transform stage 212, such as a discrete cosine transform or a discrete sine transform. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely operating the transform (referred to as an "inverse transform"). For example, to reconstruct pixels of residual BPU 210, the inverse transform can multiply the values ​​of corresponding pixels in a basis pattern by their associated coefficients and add the products to produce a weighted sum. In video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Thus, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they can be used to reconstruct residual BPU 210 without significant loss of quality. Thus, residual BPU 210 is further compressed.

[0036]

[0058] The encoder may further compress the transform coefficients in a quantization step 214 . In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally good at recognizing low-frequency fluctuations, an encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients in an inverse operation of quantization (referred to as "dequantization").

[0037]

[0059] The encoder ignores the remainder of such division in the rounding operation, so the quantization step 214 may be lossy. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0038]

[0060] In the binary encoding step 226, the encoder performs, for example, an entropy coder. The predicted data 206 and the quantized transform coefficients 216 may be encoded using a binary coding technique, such as coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the encoder may encode other information in a binary coding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. The output data of stage 226 may be used to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0039]

[0061] Referring to the reconstruction path of process 200A, in the inverse quantization step 218, the code In the inverse transform stage 220, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of the process 200A.

[0040]

[0062] Other variations of process 200A for encoding video sequence 202 include: It should be noted that various options may be used. In some embodiments, the encoder may perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be separated into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.

[0041]

[0063] FIG. 2B illustrates another example encoding process 200B consistent with embodiments of the present disclosure. 2 shows a schematic diagram. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0042]

[0064] Broadly speaking, forecasting techniques can be classified into two types: spatial and temporal forecasting. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels of one or more already coded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions of one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0043]

[0065] Referring to process 200B, in the forward path, the encoder performs spatial prediction. Prediction operations are performed in step 2042 and temporal prediction step 2044. For example, in spatial prediction step 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being coded, prediction reference 224 may include one or more neighboring BPUs that have been coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation may be adjacent to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below-left, below-right, above-left, or above-right of the original BPU), or any direction specified within the video coding standard used. In intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0044]

[0066] In another example, in the temporal prediction step 2044, the encoder may perform inter prediction. For an original BPU of a current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, a reference picture may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (called a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location with the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. If the encoder finds a region similar to the original BPU within the search window (e.g., using a pel-recursive algorithm, a block matching algorithm, etc.), the encoder may search for a matching region within the search window. Once a matching region is identified (e.g., by using a matching algorithm, etc.), the encoder can determine the region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., picture 106 in FIG. 1), the encoder can find the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region for each matching reference picture.

[0045]

[0067] Motion estimation distinguishes between different types of motion, e.g., translation, rotation, scaling, etc. In inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.

[0046]

[0068] To generate the predicted BPU 208, the encoder performs a "motion compensation" operation. Motion compensation can be performed. Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder can shift the matching regions of the reference picture according to the motion vectors, within which the encoder can predict the original BPU of the current picture. If multiple reference pictures are used (e.g., such as picture 106 in FIG. 1), the encoder can shift the matching regions of the reference pictures according to their respective motion vectors and average the pixel values ​​of the matching regions. In some embodiments, if the encoder assigns weights to the pixel values ​​of the matching regions of the respective matching reference pictures, the encoder can add a weighted sum of the pixel values ​​of the shifted matching regions.

[0047]

[0069] In some embodiments, inter prediction can be unidirectional or bidirectional. Inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (i.e., picture 102) precedes picture 104. The pictures 104 and 108 are bidirectional inter-predicted pictures in both temporal directions relative to the picture 104.

[0048]

[0070] Continuing with the forward pass of process 200B, spatial prediction step 204 After the temporal prediction step 2044, in a mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of the process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0049]

[0071] In the reconstruction path of process 200B, intra prediction mode is used in the forward path. If the inter prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been coded and reconstructed within the current picture), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the inter prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs have been coded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused by inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as inter-predicted reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in a binary coding stage 226.

[0050]

[0072] FIG. 3A illustrates an example of a decoding process 300A consistent with embodiments of the present disclosure. 2A-2B , the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder decodes a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region of each picture (eg, regions 114-118) encoded in video bitstream 228.

[0051]

[0073] In FIG. 3A, the decoder processes the basic processing unit of a coded picture (the “coded A portion of the video bitstream 228 associated with a particular bitstream (referred to as a "BPU") may be fed to a binary decoding stage 302. In the binary decoding stage 302, the decoder decodes the portion as The decoder may decode the predicted data 206 into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the predicted reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0052]

[0074] The decoder performs process 300A iteratively to generate a The coded BPU can be decoded to generate a predicted reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0053]

[0075] In the binary decoding step 302, the decoder decodes the binary coding used by the encoder. The decoder may perform an inverse operation of the technique (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, the prediction mode, parameters of the prediction operation, the type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.

[0054]

[0076] FIG. 3B is a schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure. 3 shows a schematic diagram. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.

[0055]

[0077] In process 300B, the coded picture being decoded ("current picture") is For a coded basic processing unit (referred to as a "current BPU") of a frame stream (referred to as a "frame stream"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, orientations of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.

[0056]

[0078] Based on the prediction mode indicator, the decoder performs spatial prediction in a spatial prediction step 2042. In a temporal prediction step 2044, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) or temporal prediction (e.g., inter prediction). Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.

[0057]

[0079] In process 300B, the decoder performs a prediction operation in the next iteration of process 300B. The predicted reference 224 for performing the above may be fed to the spatial prediction stage 2042 or the temporal prediction stage 2044. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as inter-prediction reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).

[0058]

[0080] FIG. 4A illustrates a mechanism for encoding or decoding video consistent with embodiments of the present disclosure. 4A is a block diagram of an example device 400. As shown in FIG. 4A, device 400 may include a processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for encoding or decoding video. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped as a single logical entity. For example, as shown in FIG. 4A, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0059]

[0081] The device 400 stores data (e.g., instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4A, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequences). The memory 404 may include a memory array (e.g., a video sequence 202, a video bitstream 228, or a video stream 304). The processor 402 can access program instructions and data for processing (e.g., via a bus 410) and execute the program instructions to operate on or process the data for processing. The memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a collection of memories (not shown in FIG. 4A ) grouped together as a single logical entity.

[0060]

[0082] Internal bus (e.g., CPU memory bus), external bus (e.g., universal serial bus) Bus 410 , such as a Real Bus Port, Peripheral Component Interconnect Express Port, etc., may be a communication device that transfers data between components within device 400 .

[0061]

[0083] For the sake of clarity and simplicity, this disclosure will refer to processor 40 2 and other data processing circuitry are collectively referred to as the "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module, or may be fully or partially combined within any other component of device 400.

[0062]

[0084] The device 400 is connected to a network (e.g., the Internet, an intranet, a local area network, etc.). The network interface 406 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may be any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, or the like. The chip may include any combination of chips, processors, cellular network chips, etc.

[0063]

[0085] In some embodiments, a peripheral device for providing a connection to one or more peripheral devices. Optionally, device 400 may further include a device interface 408. As shown in Figure 4A, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), and the like.

[0064]

[0086] Video codec (e.g., process 200A, 200B, 300A, or 300 It should be noted that the codecs (e.g., codecs executing processes 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0065]

[0087] For CUs coded using inter prediction, the previously decoded picture ( A reference block in a current picture (i.e., a reference picture) is identified as a predictor. The relative position between the reference block in the reference picture and the coded block in the current picture is called a motion vector (MV). The motion information of the current CU is specified by a predictor, a reference picture index, and the number of corresponding MVs. After obtaining a prediction by motion compensation based on the motion information, the residual between the predicted signal and the original signal can be further subjected to transformation, quantization, and entropy coding before being packed into the output bitstream.

[0066]

[0088] In some situations, the motion information of spatially and temporally neighboring CUs of the current CU is used. The motion information of the current CU can be predicted using the motion information of the current CU. To reduce the coding bits of the motion information, a merge mode can be adopted. In the merge mode, motion information is derived from spatially or temporally neighboring blocks, and a merge index can be signaled to indicate from which neighboring block the motion information is derived.

[0067]

[0089] In HEVC, a merge candidate list can be constructed based on the following candidates: do.

[0068]

[0090] (1) Up to four spatial merge candidates derived from five spatially adjacent blocks .

[0069]

[0091] (2) One temporal merge candidate derived from the temporally collocated block .

[0070]

[0092] (3) Additional merge candidates including combined bi-prediction candidates and zero motion vector candidates.

[0071]

[0093] The first candidate in the merge candidate list is the spatial neighbor. The position of each candidate is indicated by the order {A1, B1, B0, A0, B2}. The availability of each candidate position is checked according to the order {A1, B1, B0, A0, B2}. If a spatially neighboring block is intra-predicted or its position is outside the current slice or tile, the spatially neighboring block can be considered unavailable as a merge candidate. In addition, some redundancy checks can be performed to ensure that the motion data from neighboring blocks is as unique as possible. To reduce the complexity caused by the redundancy checks, only limited redundancy checks can be performed, and uniqueness is not necessarily guaranteed. For example, given the order {A1, B1, B0, A0, B2}, B0 checks only B1, A0 checks only A1, and B2 checks only A1 and B1.

[0072]

[0094] For temporal merge candidates, if available, the reference picture The bottom right position C0 directly outside the collocated block of the slice is used. Otherwise, the center position C1 can be used instead. Which reference picture list is used for the collocated reference picture can be indicated by an index signaled in the slice header. As shown in Figure 5, the MV of the collocated block can be scaled based on the Picture Order Count (POC) difference before being inserted into the merge list.

[0073]

[0095] The maximum number of merge candidates, C, can be specified in the slice header. If the number of possible merge candidates (including temporal candidates) exceeds C, only the first C-1 spatial and temporal candidates are retained. Otherwise, if the number of available merge candidates is less than C, additional candidates are generated until the number equals C. This configuration can simplify parsing and make the parsing more robust, as the ability to parse the coded data does not depend on the number of available merge candidates. In the common experimental conditions (CTC), the maximum number of merge candidates C is set to 5.

[0074]

[0096] For B slices, the default order for reference picture lists 0 and 1 is followed. Additional merge candidates are generated by combining two available candidates according to the list 0. For example, the first candidate generated uses the first merge candidate for list 0 and the second merge candidate for list 1. HEVC defines a total of 12 predefined pairs of two motion vectors in the already constructed merge candidate list as (0,1), (1,0), (0,2), (2,0), (1,2), (2,1), (0,3), (3,0), (1,3), (3,1), (2,3), and (3,2), in the above order, where (i,j) represents the index of the available merge candidate. Of them, a maximum of five candidates can be included after removing redundant entries.

[0075]

[0097] If the slice is a P slice or the number of merge candidates is still less than C, If so, the zero motion vector associated with the reference index from zero to the number of reference pictures minus one is used to fill any remaining entries in the merge candidate list.

[0076]

[0098] In VVC, the merge candidate list is created by including the following five candidates in order: is constructed: spatial merge candidates from spatially adjacent CUs, Temporal merge candidates from collocated CUs, History-based motion vector predictor (HMVP) from a FIFO table, Pairwise average candidates, and Zero MV.

[0077]

[0099] The definitions of spatial and temporal merge candidates are the same as in HEVC. After the spatial and temporal merge candidates, HMVP merge candidates are added to the merge list. In HMVP, motion information of previously coded blocks is stored in a table and used as a motion vector predictor for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). If there is a non-subblock inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0078]

[0100] In VVC, the size of the HMVP table can be set to 6, i.e., up to six HMVP candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule can be used, and a redundancy check is first applied to find whether there is an identical HMVP in the table. If found, the identical HMVP can be removed from the table and all subsequent HMVP candidates can be moved forward.

[0079]

[0101] During the merge candidate list construction process, the last few HMVP candidates in the table are examined in order and inserted into the merge candidate list after the temporal motion vector predictor (TMVP) candidates. Redundancy checks can be applied to examine HMVP candidates against spatial or temporal merge candidates.

[0080]

[0102] After inserting the HMVP candidate, if the merge candidate list is still not full, add a pairwise average candidate. The pairwise average candidate is generated by averaging a predefined pair of candidates in the existing merge candidate list. The predefined pair is defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge index in the merge candidate list. The averaged motion vector is calculated separately for each reference picture list. If both motion vectors are available in one list, the two motion vectors are averaged even if they point to different reference pictures. If only one motion vector is available, the available one is used directly. If no motion vectors are available, the list is considered invalid.

[0081]

[0103] If the merge list is still not full after adding the pairwise average merge candidates, zero motion vectors are inserted at the end until the maximum number of merge candidates is reached.

[0082]

[0104] In VVC, in addition to the normal merge mode, the construction of the merge candidate list can also be used for the merge mode with motion vector difference (MMVD) and the triangulation mode (TPM).

[0083]

[0105] In MMVD, a merge candidate is first selected from a merge candidate list and further refined by signaled motion vector difference (MVD) information. The size of the MMVD merge candidate list is set to 2. A merge candidate flag may be signaled to specify which of two MMVD candidates is used as the base motion vector (MV). MVD information may be signaled by a distance index and a direction index. The distance index specifies motion magnitude information and indicates a predefined offset from the base MV. The relationship between the distance index and the predefined offset is shown in the example of Figure 6. The direction index specifies the sign of the offset added to the base MV, for example, 0 indicates a positive sign and 1 indicates a negative sign.

[0084]

[0106] In TPM, a CU is divided equally into two triangular partitions using diagonal or anti-diagonal partitioning. Each triangular partition within a CU can be inter-predicted using its own motion. Only uni-prediction is allowed for each partition; that is, each partition has one motion vector and one reference index. Similar to bi-prediction, uni-prediction motion constraints are applied to ensure that only two motion-compensated predictions are required for each CU. When a triangulation mode is used for the current CU, a flag indicating the triangulation direction (diagonal or anti-diagonal) and two merge indices (one per partition) may also be signaled. After predicting each triangular partition, a blending process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal edges. This corresponds to a prediction signal for the entire CU, and the transform and quantization process can be applied to all CUs as in normal inter mode. A merge candidate list can be constructed. The maximum number of TPM merge candidates is explicitly signaled in the slice header and can be set to 5 in the CTC.

[0085]

[0107] In VVC, a merge candidate list is constructed, which includes spatial candidates, temporal candidates, HMVP, and pairwise average candidates. Different merge list sizes are applied to different inter modes. For example, spatial merge candidates can be inserted into the merge list according to the order {A1, B1, B0, A0, B2}. However, the construction process of spatial merge candidates has not changed from HEVC to VVC, and it does not take into account new merge candidates (e.g., HMVP, pairwise average) and new inter modes (e.g., MMVD, TPM) in VVC. This leads to various shortcomings of the current spatial merge candidates.

[0086]

[0108] For example, the order of spatial merge candidates can be improved. The number of spatial merge candidates can be adjusted. Furthermore, the construction of spatial merge candidates is fixed for normal mode, MMVD mode, and TPM mode, which limits the potential of the merging method. In addition, the construction of spatial merge candidates is fixed for low-latency pictures and non-low-latency pictures, which reduces flexibility. To address the above and other problems, this disclosure provides various solutions.

[0087]

[0109] For example, in some embodiments, the order of spatial merge candidates may be changed. A new order of spatial merge candidates {B1, A1, B0, A0, B2} may be applied. The positions of spatially neighboring blocks B1, A1, B0, A0, and B2 are shown in Figure 7. In some embodiments, this order can be changed to {B1, A1, B0, A0, B2}.

[0088]

[0110] The new order corresponds to the neighboring block above, the neighboring block to the left, the neighboring block above, the neighboring block to the left, and the neighboring block above and to the left in succession, and this order alternates between the neighbors above and the neighbors to the left. Furthermore, the new order of spatial merging candidates can achieve higher coding performance. As shown in Figures 8 and 9, according to some embodiments, the proposed method can obtain an average coding gain of 0.05% and 0.21% compared to VTM-6 under random access (RA) and low delay (LD) configurations, respectively.

[0089]

[0111] In some embodiments, the number of spatial merge candidates can be changed. To achieve a better trade-off between computational complexity and coding performance, various embodiments of the present disclosure propose and apply a reduction in the number of spatial merge candidates. When limiting the number of spatial merge candidates to two, a construction order {B1, A1} can be applied. For example, neighboring block B1 can be examined and inserted into the merge list if available. Then, neighboring block A1 can be examined and inserted into the merge list if available and not the same as B1. After inserting spatial merge candidate {B1, A1}, subsequent TMVP, HMVP, and pairwise average candidates can be added into the merge list.

[0090]

[0112] If we limit the number of spatial merge candidates to 3, we can apply the construction order {B1, A1, B0}. The inspection order of neighboring blocks is B1->A1->B0, and if available and not redundant, the corresponding MV can be inserted into the merge list.

[0091]

[0113] When using spatial merging candidates {B1, A1, B0}, experimental results compared with VTM-6 are shown in Figures 10 and 11. As shown in Figures 10 and 11, according to some embodiments, the proposed technique can achieve coding gains of 0.00% and 0.10% under RA and LD configurations.

[0092]

[0114] Furthermore, in some VVC techniques, the total number of merge candidates may be signaled. Some embodiments of the present disclosure propose further signaling the number of spatial merge candidates to gain more flexibility in constructing the merge candidate list. The number of spatial merge candidates may be set to various values ​​considering the prediction structure of the current picture. If the current picture is a non-low latency picture, the number of spatial merge candidates may be set to a first value. A non-low latency picture may refer to a picture that is coded using reference pictures from both the past and future according to display order. Otherwise, if the current picture is a low latency picture, the number of spatial merge candidates may be set to a second value. A low latency picture may refer to a picture that is coded using only reference pictures from the past according to display order. The first value may be greater than the second value. The first and second values ​​may be explicitly signaled in the bitstream, for example, in a slice header. An example is shown in Figure 12.

[0093]

[0115] Syntax element num_spatial_merge_cand_minus2 (e.g., element 1201 in Figure 12) may indicate the number of spatial merge candidates used for the current slice. The value of num_spatial_merge_cand_minus2 may be in the range of 0 to 3 (inclusive). If there are no elements of num_spatial_merge_cand_minus2, num_spatial_merge_cand_minus2 is assumed to be equal to 0. It can be argued.

[0094]

[0116] Depending on the reference picture used to code the current slice, the slice can be classified as low-delay or non-low-delay and can use a different number of merge candidates. The value of num_spatial_merge_cand_minus2 is set by the encoder accordingly. It can be transmitted in a bit stream.

[0095]

[0117] Alternatively, instead of signaling one syntax element in the slice header, two syntax elements num_spatial_merge_cand_minus2_non_lowdelay and num_spatial_merge_cand_minus2_lowdelay can be signaled in the Picture Parameter Set (PPS) or Sequence Parameter Set (SPS) as shown in Figure 13 (e.g., element 1301) and Figure 14 (e.g., element 1401). Furthermore, at the slice level, a corresponding number of spatial merge candidates can be used depending on the type of slice.

[0096]

[0118] The values ​​of num_spatial_merge_cand_minus2_non_lowdelay and num_spatial_merge_cand_minus2_lowdelay may indicate the number of spatial merge candidates used for non-low-delay and low-delay slices, respectively. The values ​​of num_spatial_merge_cand_minus2_non_lowdelay and num_spatial_merge_cand_minus2_lowdelay may be in the range of 0 to 3, inclusive. If num_spatial_merge_cand_minus2_non_lowdelay or num_spatial_merge_cand_minus2_lowdelay are missing, they can be inferred to be equal to 0.

[0097]

[0119] In some embodiments, different construction orders of spatial merge candidates may be applied for different inter modes. For example, two construction orders of spatial merge candidates may be considered, including {B1, A1, B0, A0, B2} and {A1, B1, B0, A0, B2}. Different construction orders may be adopted for normal merge mode, MMVD mode, and TPM mode. In some embodiments, we propose using {B1, A1, B0, A0, B2} for normal merge mode and TPM mode, and {A1, B1, B0, A0, B2} for MMVD mode. Experimental results of the exemplary embodiment are shown in Figures 15 and 16. It has been observed that the proposed method can achieve an average coding gain of 0.07% and 0.16% under RA and LD configurations, respectively.

[0098]

[0120] Based on this disclosure, one skilled in the art will understand that other combinations of spatial merge candidate orders and merge modes can be used, for example, {B1, A1, B0, A0, B2} can be used for normal merge mode only, and {A1, B1, B0, A0, B2} can be used for MMVD mode and TMP mode.

[0099]

[0121] In some embodiments, an adaptive construction order of spatial merge candidates may be applied based on the type of frame. For example, different spatial merge candidate construction methods may be applied to different types of inter-coded pictures, such as low-latency pictures and non-low-latency pictures. In some embodiments, for low-latency pictures, the spatial merge candidate construction order {B1, A1, B0, A0, B2} may be used for normal merge mode, TPM mode, and MMVD mode. For non-low-latency pictures, the spatial merge candidate construction order {B1, A1, B0, A0, B2} may be used for normal merge mode and TPM mode, and the spatial merge candidate construction order {A1, B1, B0, A0, B2} may be used for MMVD mode. Experimental results of the exemplary embodiment are shown in Figures 17 and 18. It has been observed that the proposed method can achieve an average coding gain of 0.08% and 0.21% under RA and LD configurations, respectively.

[0100]

[0122] FIG. 19 illustrates a flowchart of an exemplary video processing method 1900 consistent with embodiments of the present disclosure. In some embodiments, method 1900 may be performed by one or more software or hardware components of an encoder, decoder, device (e.g., device 400 of FIG. 4A). For example, a processor (e.g., processor 402 of FIG. 4A) may perform method 1900. In some embodiments, method 1900 may be performed by a computer. The present invention may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (eg, device 400 of FIG. 4A).

[0101]

[0123] In step 802, a set of spatial merge candidates may be inserted into a merge candidate list for a coded block, for example, by one or more software or hardware components of an encoder, decoder, or device (e.g., device 400 of FIG. 4A). In VVC, the order of spatial merge candidates may be improved. The set of spatial merge candidates may be inserted according to the order of the top neighboring block, the left neighboring block, the top neighboring block, the left neighboring block, and the top-left neighboring block. For example, a new order of spatial merge candidates {B1, A1, B0, A0, B2} is shown in FIG. 7.

[0102]

[0124] The number of spatial merge candidates can be adjusted. In step 804, a preset numerical limit for spatial merge candidates is determined.

[0103]

[0125] In step 806, if the numerical limit is 2, a set of spatial merge candidates is inserted into the merge candidate list based on the order of top neighbor block, left neighbor block. When limiting the number of spatial merge candidates to 2, a construction order {B1, A1} can be applied. For example, neighbor block B1 can be examined and inserted into the merge list if available. Then, neighbor block A1 can be examined and inserted into the merge list if available and not the same as B1. After inserting spatial merge candidate {B1, A1}, subsequent TMVP, HMVP, and pairwise average candidates can be added into the merge list.

[0104]

[0126] In step 808, if the numerical limit is 3, insert a set of spatial merge candidates into the merge candidate list based on the order of upper neighbor block, left neighbor block, and upper neighbor block. If the number of spatial merge candidates is limited to 3, a construction order {B1, A1, B0} can be applied. The inspection order of neighbor blocks is B1->A1->B0, and if available and not redundant, the corresponding MV can be inserted into the merge list.

[0105]

[0127] In some embodiments, at least one of a temporal merge candidate from a collocated coding unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector may be added to the merge candidate list.

[0106]

[0128] In HMVP, motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for the current coded unit. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). If there is a non-sub-block inter-coded coded unit, the motion information associated with the non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate.

[0107]

[0129] Pairwise average candidates are generated by averaging pairs of candidates in the merge candidate list, depending on whether the merge candidate list is full, and are added to the merge candidate list after one or more HMVPs have been added to the merge candidate list.

[0108]

[0130] If the merge list is still not full after adding the pairwise average merge candidates, insert zero motion vectors at the end of the merge candidate list until the maximum number of merge candidates is reached. Enter.

[0109]

[0131] In step 810, a determination is made, e.g., by an encoder or decoder, whether to apply a first coding mode or a second coding mode to the coding block. The first coding mode is different from the second coding mode. In some embodiments, each of the first coding mode and the second coding mode may be one of a normal merge mode, a merge mode with motion vector difference (MMVD), and a triangulation mode (TPM).

[0110]

[0132] In step 812, if a first coding mode is applied to the coded block, a set of spatial merge candidates is inserted according to a first construction order. For example, in MMVD, a merge candidate is first selected from a merge candidate list and refined by signaled motion vector difference (MVD) information, and a merge candidate flag is signaled to specify which of two MMVD candidates is used as the base motion vector. MVD information can be signaled by a distance index and a direction index. The distance index specifies motion magnitude information and indicates a predefined offset from the base MV. The relationship between the distance index and the predefined offset is shown in the example of Figure 6. The direction index specifies the sign of the offset added to the base MV, for example, 0 indicates a positive sign and 1 indicates a negative sign.

[0111]

[0133] In step 814, if a second coding mode is applied to the coding block, the set of spatial merge candidates is inserted according to a second construction order. For example, in TPM, the coding unit is evenly divided into two triangular partitions using at least one of a diagonal partition or an anti-diagonal partition. Each triangular partition within a CU can be inter-predicted using its own motion. Only uni-prediction is allowed for each partition. That is, each partition has one motion vector and one reference index.

[0112]

[0134] Step 816 determines whether the coded block is part of a low latency picture or a non-low latency picture.

[0113]

[0135] In step 818, if the coded block is part of a low-latency picture, the set of spatial merge candidates is inserted according to a third construction order. In some embodiments, for low-latency pictures, the construction order of spatial merge candidates {B1, A1, B0, A0, B2} can be used for normal merge mode, TPM mode, and MMVD mode.

[0114]

[0136] In step 820, if the coded block is part of a non-low delay picture, the set of spatial merge candidates is inserted according to a fourth construction order. The third construction order is different from the fourth construction order. The third construction order and the fourth construction order are used for MMVD. In some embodiments, for non-low delay pictures, the construction order of spatial merge candidates {B1, A1, B0, A0, B2} can be used for normal merge mode and TPM mode, and the construction order of spatial merge candidates {A1, B1, B0, A0, B2} can be used for MMVD mode.

[0115]

[0137] Those skilled in the art will appreciate that one or more of the above methods may be used in combination or separately, consistent with this disclosure. For example, techniques employing spatial merge candidate reduction may be used in combination with the proposed method that uses separate construction orders of spatial merge candidates for different inter modes.

[0116]

[0138] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, the instructions being transmitted by an apparatus (such as the disclosed encoders and decoders) for performing the above-described methods. Typical non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0117]

[0139] It should be noted that relational terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, terms such as "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive list of such items or to be limited only to the items they list.

[0118]

[0140] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if a database is stated to include A or B, the database can include A or B, or A and B, unless otherwise specified or impracticable. As a second example, if a database is stated to include A, B, or C, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise specified or impracticable.

[0119]

[0141] It will be understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0120]

[0142] Embodiments may be further described using the following clauses: 1. A video processing method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; A video processing method, wherein the set of spatial merge candidates is inserted according to the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block, and upper-left neighboring block. 2. The method of clause 1, further comprising adding at least one of a temporal merge candidate from a collocated coding unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 3. The method of clause 2, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for the current coding unit. 4. The motion information associated with the non-subblock inter-coded coded unit is added as a new HMVP candidate to the last entry of the FIFO table. 10. The method according to any one of the preceding claims. 5. The method of clause 2, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, depending on whether the merge candidate list is not full, one or more HMVPs are added to the merge candidate list after the HMVP is added to the merge candidate list. 6. The method of clause 2, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 7. A video processing method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block based on a preset numerical limit for the spatial merge candidates; If the numeric limit is 2, the spatial merge candidate pairs are inserted into the merge candidate list based on the order of top neighbor block, left neighbor block; and A method of processing video, wherein when the numerical limit is 3, a set of spatial merge candidates is inserted into a merge candidate list based on the order of upper neighbor block, left neighbor block, upper neighbor block. 8. In response to the current picture being coded using past reference pictures and future reference pictures according to display order, the number of spatial merge candidates is set to a first value; and The method of clause 7, wherein the number of spatial merging candidates is set to a second value less than the first value in response to the current picture being coded using past reference pictures according to display order. 9. The method of clause 7, further comprising signaling the number of spatial merge candidates inserted into the merge candidate list. 10. The method of any one of clauses 7 to 9, further comprising adding at least one of a temporal merge candidate from a collocated coding unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 11. The method of clause 10, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 12. The method of any one of clauses 10 and 11, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 13. The method of clause 10, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, depending on whether the merge candidate list is not full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 14. The method of clause 10, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 15. A video processing method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; When a first coding mode is applied to a coding block, the set of spatial merge candidates is inserted according to a first construction order; and When a second coding mode is applied to the coded block, the set of spatial merge candidates is inserted according to a second construction order; A video processing method, wherein the first construction order is different from the second construction order. 16. The method according to clause 15, wherein the first coding mode and the second coding mode are two different modes selected from a normal merge mode, a merge mode with motion vector difference (MMVD), and a triangulation mode (TPM). 17. In MMVD, a merge candidate is first selected from the merge candidate list and refined by the signaled motion vector difference (MVD) information, and the merge candidate flag is 17. The method of clause 16, wherein signaling is performed to specify which of two MMVD candidates is used as the base motion vector. 18. The method of clause 16, wherein in the TPM, the coding unit is divided equally into two triangular partitions using at least one of a diagonal division or an anti-diagonal division. 19. The method of any one of clauses 15 and 16, further comprising adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor from a FIFO table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 20. The method of clause 19, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 21. The method of any one of clauses 19 and 20, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 22. The method of clause 19, wherein pairwise average candidates are generated by averaging pairs of candidates in the merge candidate list, and, depending on whether the merge candidate list is not full, are added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 23. The method of clause 19, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 24. A video processing method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; If the coded block is part of a low-delay picture, the set of spatial merging candidates is inserted according to a first construction order; and If the coded block is part of a non-low delay picture, the set of spatial merging candidates is inserted according to a second construction order; A video processing method, wherein the first construction order is different from the second construction order. 25. The method of clause 24, wherein the first construction order and the second construction order are used for merge mode with motion vector difference (MMVD). 26. The method of clause 24, further comprising adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 27. The method of clause 26, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 28. The method of any one of clauses 26 and 27, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 29. The method of clause 26, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, depending on whether the merge candidate list is not full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 30. The method of clause 26, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 31. A video processing device, a memory for storing a set of instructions; and one or more processors, the one or more processors comprising: configured to execute a set of instructions to cause the device to insert the set of spatial merge candidates into a merge candidate list for the coded block; A video processing device, wherein the set of spatial merge candidates is inserted in the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block, and upper left neighboring block. 32. The device of clause 31, wherein the one or more processors are configured to execute a set of instructions to further cause the device to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 33. The apparatus of clause 32, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 34. The device of any one of clauses 32 and 33, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 35. The apparatus of clause 32, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 36. The apparatus of clause 32, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 37. A video processing device, a memory for storing a set of instructions; and one or more processors, the one or more processors comprising: configured to execute a set of instructions to cause the device to insert a set of spatial merge candidates into a merge candidate list for the coded block based on a preset numerical limit for the spatial merge candidates; If the numeric limit is 2, the spatial merge candidate pairs are inserted into the merge candidate list based on the order of top neighbor block, left neighbor block; and A video processing device, wherein when the numerical limit is 3, a set of spatial merge candidates is inserted into a merge candidate list based on the order of upper neighbor block, left neighbor block, upper neighbor block. 38. In response to the current picture being coded using past reference pictures and future reference pictures according to display order, the number of spatial merge candidates is set to a first value; and 38. The apparatus of clause 37, wherein the number of spatial merging candidates is set to a second value that is less than the first value in response to the current picture being coded using past reference pictures according to display order. 39. One or more processors may: 38. The apparatus of clause 37, configured to execute a set of instructions to further cause the apparatus to signal a number of spatial merge candidates inserted into the merge candidate list. 40. The device of any one of clauses 37 to 39, wherein the one or more processors are configured to execute a set of instructions to further cause the device to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 41. The apparatus of clause 40, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 42. The device of any one of clauses 40 and 41, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 43. The apparatus of clause 40, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 44. The apparatus of clause 40, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 45. A video processing device, a memory for storing a set of instructions; and one or more processors, the one or more processors comprising: configured to execute a set of instructions to cause the device to insert the set of spatial merge candidates into a merge candidate list for the coded block; When a first coding mode is applied to a coding block, the set of spatial merge candidates is inserted according to a first construction order; and When a second coding mode is applied to the coded block, the set of spatial merge candidates is inserted according to a second construction order; The first construction order is different from the second construction order of the video processing device. 46. ​​The device of clause 45, wherein the first coding mode and the second coding mode are two different modes selected from a normal merge mode, a merge mode with motion vector difference (MMVD), and a triangulation mode (TPM). 47. The device of clause 46, wherein in MMVD, a merge candidate is first selected from a merge candidate list and refined by signaled motion vector difference (MVD) information, and a merge candidate flag is signaled to specify which of the two MMVD candidates is used as the base motion vector. 48. A device according to any one of clauses 45 and 46, wherein in the TPM, the coding unit is divided equally into two triangular sections using at least one of a diagonal division or an anti-diagonal division. 49. The device of clause 46, wherein at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor from a FIFO table, a pairwise average candidate, or a zero motion vector is added to the merge candidate list. 50. The apparatus of clause 49, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 51. The device of any one of clauses 49 and 50, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 52. The apparatus of clause 49, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 53. The apparatus of clause 49, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 54. A video processing device, a memory for storing a set of instructions; and one or more processors, the one or more processors comprising: configured to execute a set of instructions to cause the device to insert the set of spatial merge candidates into a merge candidate list for the coded block; If the coded block is part of a low-delay picture, the set of spatial merging candidates is inserted according to a first construction order; and If the coded block is part of a non-low delay picture, the set of spatial merging candidates is inserted according to a second construction order; The first construction order is different from the second construction order of the video processing device. 55. The device according to clause 54, wherein the first construction order and the second construction order are used in merge mode with motion vector difference (MMVD). 56. The device described in clause 54, wherein the one or more processors are configured to execute a set of instructions to further cause the device to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, and a zero motion vector to a merge candidate list. 57. The apparatus of clause 56, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 58. The device of any one of clauses 56 and 57, wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate. 59. The apparatus of clause 56, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 60. The device of clause 56, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 61. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; 10. A non-transitory computer-readable medium, wherein the set of spatial merge candidates is inserted according to the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block, and upper-left neighboring block. 62. The non-transitory computer-readable medium of clause 61, wherein the set of instructions is executable by a computer to further cause the computer to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list. 63. The non-transitory computer-readable medium of clause 62, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 64. The non-transitory computer-readable medium of any one of clauses 62 and 63, wherein motion information associated with a non-subblock inter-coded coding unit is added to the last entry of the FIFO table as a new HMVP candidate. 65. The non-transitory computer-readable medium of clause 62, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 66. The non-transitory computer-readable medium of clause 62, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 67. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block based on a preset numerical limit for the spatial merge candidates; If the numeric limit is 2, the marking is done based on the order of the upper adjacent block, the left adjacent block. The set of spatial merge candidates is inserted into the list of merge candidates; and A non-transitory computer-readable medium, wherein if the numeric limit is 3, a set of spatial merge candidates is inserted into a merge candidate list based on the order of upper neighbor block, left neighbor block, and upper neighbor block. 68. In response to the current picture being coded using past reference pictures and future reference pictures according to display order, the number of spatial merge candidates is set to a first value; and 68. The non-transitory computer-readable medium of clause 67, wherein the number of spatial merging candidates is set to a second value less than the first value in response to the current picture being coded using past reference pictures according to display order. 69. The non-transitory computer-readable medium of clause 67, wherein the set of instructions is executable by a computer to further cause the computer to signal the number of spatial merge candidates inserted into the merge candidate list. 70. A non-transitory computer-readable medium described in any one of clauses 67 to 69, wherein at least one processor is configured to execute a set of instructions to further cause the computer to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, and a zero motion vector to a merge candidate list. 71. The non-transitory computer-readable medium of clause 70, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 72. The non-transitory computer-readable medium of any one of clauses 70 and 71, wherein motion information associated with a non-subblock inter-coded coding unit is added to the last entry of the FIFO table as a new HMVP candidate. 73. The non-transitory computer-readable medium of clause 70, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 74. The non-transitory computer-readable medium of clause 70, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 75. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; When a first coding mode is applied to a coding block, the set of spatial merge candidates is inserted according to a first construction order; and When a second coding mode is applied to the coded block, the set of spatial merge candidates is inserted according to a second construction order; A non-transitory computer-readable medium, wherein the first construction order is different from the second construction order. 76. The non-transitory computer-readable medium of clause 75, wherein the first coding mode and the second coding mode are two different modes selected from a normal merge mode, a merge mode with motion vector difference (MMVD), and a triangulation mode (TPM). 77. The non-transitory computer-readable medium of clause 76, wherein in MMVD, a merge candidate is first selected from a merge candidate list and refined by signaled motion vector difference (MVD) information, and a merge candidate flag is signaled to specify which of the two MMVD candidates is used as the base motion vector. 78. The non-transitory computer-readable medium of clause 76, wherein in the TPM, the coding unit is divided evenly into two triangular partitions using at least one of a diagonal division or an anti-diagonal division. 79. The non-transitory computer-readable medium of any one of clauses 75 and 76, further comprising adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor from a FIFO table, a pairwise average candidate, and a zero motion vector to a merge candidate list. 80. The non-transitory computer-readable medium of clause 79, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 81. The non-transitory computer-readable medium of any one of clauses 79 and 80, wherein motion information associated with a non-subblock inter-coded coding unit is added to the last entry of the FIFO table as a new HMVP candidate. 82. The non-transitory computer-readable medium of clause 79, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 83. The non-transitory computer-readable medium of clause 79, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached. 84. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; If the coded block is part of a low-delay picture, the set of spatial merging candidates is inserted according to a first construction order; and If the coded block is part of a non-low delay picture, the set of spatial merging candidates is inserted according to a second construction order; A non-transitory computer-readable medium, wherein the first construction order is different from the second construction order. 85. The non-transitory computer-readable medium of clause 84, wherein the first construction order and the second construction order are used for merge mode with motion vector difference (MMVD). 86. The non-transitory computer-readable medium of clause 84, wherein the set of instructions is executable by a computer to further cause the computer to add at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a FIFO table, a pairwise average candidate, and a zero motion vector to the merge candidate list. 87. The non-transitory computer-readable medium of clause 86, wherein motion information of previously coded blocks is stored in a FIFO table and used as a motion vector predictor for a current coded unit. 88. The non-transitory computer-readable medium of any one of clauses 86 and 87, wherein motion information associated with a non-subblock inter-coded coding unit is added to the last entry of the FIFO table as a new HMVP candidate. 89. The non-transitory computer-readable medium of clause 86, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and, in response to the merge candidate list not being full, is added to the merge candidate list after one or more HMVPs are added to the merge candidate list. 90. The non-transitory computer-readable medium of clause 86, wherein zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached.

[0121]

[0143] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications to the described embodiments may be made. Other embodiments may be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims. The order of the steps is for illustrative purposes only and is not intended to limit the steps to a particular order, as one skilled in the art will recognize that the steps may be performed in different orders while implementing the same method.

[0122]

[0144] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications to those embodiments may be made. Accordingly, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. A video processing method, comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; The image processing method, wherein the set of spatial merge candidates is inserted in the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block and upper left neighboring block.

2. 2. The method of claim 1, further comprising adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list.

3. The method of claim 2 , wherein motion information of previously coded blocks is stored in the FIFO table and used as the motion vector predictor for a current coded unit.

4. The method of claim 3 , wherein motion information associated with a non-sub-block inter-coded coded unit is added to the last entry of the FIFO table as a new HMVP candidate.

5. 3. The method of claim 2, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and is added to the merge candidate list after one or more HMVPs are added to the merge candidate list in response to the merge candidate list not being full.

6. The method of claim 2 , wherein the zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached.

7. A video processing device, a memory for storing a set of instructions; and one or more processors, wherein the one or more processors: configured to execute the set of instructions to cause the device to insert a set of spatial merge candidates into a merge candidate list for a coded block; The set of spatial merge candidates is inserted in the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block and upper left neighboring block.

8. the one or more processors adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list; 8. The device of claim 7, configured to execute the set of instructions to further cause the device to:

9. The apparatus of claim 8 , wherein motion information of previously coded blocks is stored in the FIFO table and used as the motion vector predictor for a current coded unit.

10. 10. The apparatus of claim 9, wherein motion information associated with a non-sub-block inter-coded coded unit is added as a new HMVP candidate to a last entry of the FIFO table.

11. 9. The apparatus of claim 8, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and is added to the merge candidate list after one or more HMVPs are added to the merge candidate list in response to the merge candidate list not being full.

12. The apparatus of claim 8 , wherein the zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached.

13. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a video processing method, the method comprising: inserting a set of spatial merge candidates into a merge candidate list for the coded block; The set of spatial merging candidates are inserted according to the order of upper neighboring block, left neighboring block, upper neighboring block, left neighboring block, and upper-left neighboring block.

14. The set of instructions adding at least one of a temporal merge candidate from a collocated coded unit, a history-based motion vector predictor (HMVP) from a first-in-first-out (FIFO) table, a pairwise average candidate, or a zero motion vector to the merge candidate list; 14. The non-transitory computer-readable medium of claim 13, executable by the computer to further cause the computer to:

15. 15. The non-transitory computer-readable medium of claim 14, wherein motion information of previously coded blocks is stored in the FIFO table and used as the motion vector predictor for a current coded unit.

16. 16. The non-transitory computer-readable medium of claim 15, wherein motion information associated with a non-sub-block inter-coded coded unit is added to a last entry of the FIFO table as a new HMVP candidate.

17. 15. The non-transitory computer-readable medium of claim 14, wherein the pairwise average candidate is generated by averaging pairs of candidates in the merge candidate list, and is added to the merge candidate list after one or more HMVPs are added to the merge candidate list in response to the merge candidate list not being full.

18. The non-transitory computer-readable medium of claim 14 , wherein the zero motion vectors are inserted at the end of the merge candidate list until a maximum number of merge candidates is reached.

Citation Information

Patent Citations

  • Image decoding apparatus, image decoding method and image decoding program

    JP2013016931A

  • Image encoding device, image encoding method, image encoding program, transmission device, transmission method, and transmission program

    JP2016054551A

  • Image encoding device, image encoding method, image encoding program, image decoding device, image decoding method, and image decoding program

    WO2020184459A1