Method and non-transitory computer-readable storage medium for performing sub-block-based inter prediction

Sub-block temporal motion prediction for normal inter mode addresses the challenge of high compression efficiency in video coding standards like VVC/H.266 by optimizing motion information derivation for sub-blocks, enhancing compression performance and reducing bandwidth.

JP2025540521APending Publication Date: 2025-12-15ALIBABA (CHINA) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025524548
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-16
Filing Date
2023-10-26
Publication Date
2025-12-15

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in achieving high compression efficiency, particularly in the context of the emerging VVC/H.266 standard, where improved techniques are needed to reduce bandwidth requirements while maintaining subjective quality.

Method used

Implementing sub-block temporal motion prediction (sbAmvp mode) for normal inter mode, which involves dividing coding units into sub-blocks and deriving motion information for each sub-block based on signaled CU-level motion information.

Benefits of technology

Enhances coding efficiency by allowing for improved compression performance, aligning with the goals of the VVC/H.266 standard by reducing bandwidth requirements without compromising visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025540521000001_ABST
    Figure 2025540521000001_ABST
Patent Text Reader

Abstract

A video decoding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter mode. The method includes receiving a bitstream including one or more syntax elements that signal CU-level motion information associated with a coding unit (CU), the CU being coded using normal inter mode, dividing the CU into a plurality of sub-blocks, and performing sbAmvp mode on the plurality of sub-blocks. Performing sbAmvp mode includes deriving motion information for a sub-block of the plurality of sub-blocks based on the signaled CU-level motion information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 381,146, filed October 27, 2022, U.S. Provisional Patent Application No. 63 / 477,569, filed December 29, 2022, U.S. Provisional Patent Application No. 63 / 493,301, filed March 30, 2023, and U.S. Provisional Patent Application No. 63 / 587,190, filed October 2, 2023, and claims the benefit of U.S. Patent Application Publication No. 18 / 487,387, filed October 16, 2023. All of the above applications are expressly incorporated herein by reference in their entirety.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] This disclosure relates generally to video processing, and more particularly to methods and systems for performing sub-block based inter prediction. [Background technology]

[0003] background

[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video may be compressed before storage or transmission and decompressed before display. The compression process is usually referred to as encoding, and decompression is usually referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Universal Video Coding (VVC / H.266) standard, and the AVS standard, are developed by standardization organizations to define specific video coding formats. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is becoming increasingly improved. Summary of the Invention [Means for solving the problem]

[0004] overview

[0004] This disclosure provides a video decoding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter mode. The method includes receiving a bitstream including one or more syntax elements signaling CU-level motion information associated with a coding unit (CU), the CU being coded using normal inter mode, dividing the CU into multiple sub-blocks, and performing the sbAmvp mode on the multiple sub-blocks. Performing the sbAmvp mode includes deriving motion information for a sub-block of the multiple sub-blocks based on the signaled CU-level motion information.

[0005]

[0005] In some embodiments, a video encoding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter mode is provided. The method includes receiving a video sequence, encoding one or more pictures of the video sequence, and generating a bitstream. The encoding includes encoding a coding unit (CU) using normal inter mode, signaling CU-level motion information associated with the CU, dividing the CU into multiple sub-blocks, and performing the sbAmvp mode on the multiple sub-blocks. The performing the sbAmvp mode includes deriving motion information for a sub-block of the multiple sub-blocks based on the signaled CU-level motion information.

[0006] In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated by operations including encoding a coding unit (CU) using a normal inter mode, signaling CU-level motion information associated with the CU, dividing the CU into a plurality of sub-blocks, and performing an sbAmvp mode on the plurality of sub-blocks. Performing the sbAmvp mode includes deriving motion information for a sub-block of the plurality of sub-blocks based on the signaled CU-level motion information.

[0007] In some embodiments, a video encoding apparatus is provided, the video encoding apparatus including a memory configured to store instructions and one or more processors configured to execute the instructions to cause the apparatus to perform a method according to the above-described embodiments.

[0008] In some embodiments, a video decoding apparatus is provided, the video decoding apparatus including a memory configured to store instructions and one or more processors configured to execute the instructions to cause the apparatus to perform a method according to the above-described embodiments.

[0009]

[0009] In some embodiments, a computer program product is provided, the computer program product including computer program instructions that enable a computer to perform a method according to the above-described embodiments.

[0010]

[0010] In some embodiments, a computer program is provided, the computer program including a computer program that enables a computer to perform a method according to the above-described embodiments.

[0011] BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the following detailed description and accompanying drawings, embodiments and various aspects of the present disclosure are set forth. The various features shown in the figures are not drawn to scale. [Brief explanation of the drawings]

[0012] [Figure 1]

[0012] FIG. 1 is a schematic diagram illustrating an exemplary system for pre-processing and coding image data according to some embodiments of the present disclosure. [Figure 2A]

[0013] 1 is a schematic diagram illustrating an example encoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 2B]

[0014] FIG. 2 is a schematic diagram illustrating another exemplary encoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 3A]

[0015] 1 is a schematic diagram illustrating an example decoding process for a hybrid video coding system consistent with an embodiment of the present disclosure. [Figure 3B]

[0016] FIG. 2 is a schematic diagram illustrating another exemplary decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 4]

[0017] 1 is a block diagram of an exemplary apparatus for pre-processing or coding image data, consistent with some embodiments of the present disclosure. [Figure 5]

[0018] FIG. 1 is a schematic diagram illustrating a bitstream structure according to some embodiments of the present disclosure. [Figure 6]

[0019] FIG. 2 is a schematic diagram illustrating an example process for deriving motion vectors for sub-blocks in sub-block-based temporal motion vector prediction (SbTMVP) mode, according to some embodiments of the present disclosure. [Figure 7]

[0020] 1 illustrates collocated blocks used for temporal motion, according to some embodiments of the present disclosure. [Figure 8]

[0021] FIG. 1 is a schematic diagram illustrating an example of normal inter-prediction according to some embodiments of the present disclosure. [Figure 9]

[0022] 1 is a schematic diagram illustrating an example of sub-blocks within a coding unit (CU), according to some embodiments of the present disclosure. [Figure 10]

[0023] FIG. 10 is a schematic diagram illustrating an example of motion derivation for sub-blocks according to some embodiments of the present disclosure. [Figure 11]

[0024] 1 shows a flowchart illustrating an example method of using sub-block temporal motion prediction (sbAmvp mode) for normal inter modes, according to some embodiments of the present disclosure. [Figure 12]

[0025] 1 shows a flowchart illustrating an example method for performing a collocated block based sbAmvp mode in accordance with some embodiments of the present disclosure. [Figure 13]

[0026] 1 illustrates an example of template matching according to some embodiments of the present disclosure. [Figure 14]

[0027] 10 illustrates an exemplary co-located picture coded using a non-inter mode, according to some embodiments. [Figure 15]

[0028] 10 illustrates an example averaged motion of the top and left sub-blocks according to some embodiments of the present disclosure. [Figure 16]

[0029] 10 illustrates an exemplary averaged motion of surrounding sub-blocks according to some embodiments of the present disclosure. [Figure 17]

[0030] FIG. 10 is a schematic diagram illustrating an example of sub-blocks within a CU with motion shift, according to some embodiments of the present disclosure. [Figure 18]

[0031] FIG. 10 is a schematic diagram illustrating an example of motion derivation for sub-blocks based on motion shifting, according to some embodiments of the present disclosure. [Figure 19]

[0032] 10 shows a flowchart illustrating an example method for performing a motion-shift based sbAmvp mode, according to some embodiments of the present disclosure. [Figure 20]

[0033] 1 shows a flowchart illustrating an example method for MVP derivation, according to some embodiments of the present disclosure. [Figure 21]

[0034] 1 illustrates an exemplary prediction process for a reference sample, according to some embodiments of the present disclosure. [Figure 22]

[0035] 1 illustrates an example of a direction of movement according to some embodiments of the present disclosure. [Figure 23]

[0036] 10 illustrates exemplary candidates with fixed motion direction, according to some embodiments of the present disclosure. [Figure 24]

[0037] 1 illustrates two sets of exemplary movement directions, according to some embodiments of the present disclosure. [Figure 25]

[0038] 25 illustrates exemplary candidates with two sets of motion directions shown in FIG. 24, according to some embodiments of the present disclosure. [Figure 26]

[0039] 1 illustrates an exemplary first group of motion magnitudes and motion directions according to some embodiments of the present disclosure. [Figure 27]

[0040] 10 illustrates a second group of exemplary motion magnitudes and motion directions, according to some embodiments of the present disclosure. [Figure 28]

[0041] 10 illustrates a third group of exemplary motion magnitudes and motion directions, according to some embodiments of the present disclosure. [Figure 29]

[0042] 1 shows a flowchart illustrating an example method for determining which groups of motion magnitudes and directions to use, according to some embodiments of the present disclosure. [Figure 30]

[0043] 10 shows another flowchart illustrating an example method for determining which groups of motion magnitudes and directions to use, according to some embodiments of the present disclosure. [Figure 31]

[0044] 1 illustrates an exemplary process for the derivation of α and β, according to some embodiments of the present disclosure. [Figure 32]

[0045] 1 illustrates LIC scale and offset according to one embodiment of the present disclosure. [Figure 33]

[0045] Figure 1 shows LIC scale and offset according to another embodiment of the present disclosure. [Figure 34]

[0045] Figure 1 shows LIC scale and offset according to another embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Description of the embodiment

[0046] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which like numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present invention as set forth in the appended claims. Certain aspects of the present disclosure are described in further detail below. In the event that terms and definitions provided herein conflict with terms or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0014]

[0047] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) are currently developing the General Purpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.

[0015]

[0048] To achieve the same subjective quality as HEVC / H.265 while using half the bandwidth, JVET is developing technology that goes beyond HEVC using the Joint Search Model (JEM) reference software. Because the coding technology is embedded within JEM, JEM achieves significantly higher coding performance than HEVC.

[0016]

[0049] The VVC standard has recently been developed and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in recent video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0017]

[0050] A video is a set of static pictures (or "frames") arranged in time sequence to preserve visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display such pictures in time sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as research, conferencing, or live broadcasting.

[0018]

[0051] To reduce the storage space and transmission bandwidth required by such applications, video may be compressed before storage and transmission and decompressed before display. Compression and decompression may be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or by dedicated hardware. A module for compression is commonly referred to as an "encoder," and a module for decompression is commonly referred to as a "decoder." Encoders and decoders may collectively be referred to as a "codec." Encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be referred to as a "transcoder."

[0019]

[0052] A video encoding process can identify and retain useful information that can be used to reconstruct a picture and discard information that is not important for reconstruction. If the discarded, unimportant information cannot be perfectly reconstructed, such an encoding process may be called "lossy." Otherwise, it may be called "lossless." Most encoding processes are lossy, and lossiness is a trade-off to reduce the required storage space and transmission bandwidth.

[0020]

[0053] Useful information of an encoded picture (called the "current picture") includes changes relative to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes may include changes in pixel position, brightness, or color, of which position changes are the most important. Changes in the position of a group of pixels representing an object may reflect the movement (motion) of the object between the reference picture and the current picture.

[0021]

[0054] A picture coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture is called a "P-picture" if some or all of the blocks in the picture (e.g., blocks that generally reference portions of a video picture) are predicted using intra- or inter-prediction with one reference picture (e.g., uni-predictive). A picture is called a "B-picture" if at least one block within it is predicted using two reference pictures (e.g., bi-predictive).

[0022]

[0055] 1 is a block diagram illustrating a system 100 for pre-processing and coding image data according to some disclosed embodiments. The image data may include one image (also referred to as a "picture" or "frame"), multiple images, or video. An image is a static picture. The multiple images may or may not be related spatially or temporally. A video is a set of images arranged in time sequence.

[0023]

[0056] 1 , system 100 includes a source device 120 that provides encoded video data to be later decoded by a destination device 140. Consistent with disclosed embodiments, each of source device 120 and destination device 140 may include any of a variety of devices, including a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smart watch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, or the like. Source device 120 and destination device 140 may be equipped for wireless or wired communication.

[0024]

[0057] Referring to FIG. 1 , source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. Destination device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. Image / video preprocessor 122 preprocesses image data, i.e., one or more images or videos, to generate an input bitstream for image / video encoder 124. Image / video encoder 124 encodes the input bitstream and outputs encoded bitstream 162 via output interface 126. Encoded bitstream 162 is transmitted over communication medium 160 and received by input interface 142. Image / video decoder 144 then decodes encoded bitstream 162 to generate decoded data, which can be utilized by machine vision application 146.

[0025]

[0058] More specifically, source device 120 may further include various devices (not shown) for providing source image data to be pre-processed by image / video preprocessor 122. The devices for providing source image data may include an image / video capture device such as a camera, an image / video archive or storage device containing pre-captured images / video, or an image / video feed interface for receiving images / video from an image / video content provider.

[0026]

[0059] The image / video encoder 124 and the image / video decoder 144 may each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the encoding or decoding is implemented partially in software, the image / video encoder 124 or the image / video decoder 144 may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform techniques consistent with this disclosure. The image / video encoder 124 or the image / video decoder 144 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in an individual device.

[0027]

[0060] Image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Generic Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to an existing standard. Although not shown in FIG. 1 , in some embodiments, image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and data, respectively, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams.

[0028]

[0061] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 directly from source device 120 to destination device 140 in real time. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.

[0029]

[0062] Communication medium 160 may include a transitory medium, such as a wireless broadcast or a wired network transmission. For example, communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, communication medium 160 may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded bitstream 162 from source device 120 and provide encoded bitstream 162 to destination device 140, for example, via a network transmission.

[0030]

[0063] Communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray® disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded image data from source device 120 and produce discs containing the encoded video data.

[0031]

[0064] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.

[0032]

[0065] The machine vision application 146 may include various hardware or software for utilizing the decoded image data generated by the image / video decoder 144. For example, the machine vision application 146 may include a display device that displays the decoded image data to a user, which may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. As another example, the machine vision application 146 may include one or more processors configured to use the decoded image data to perform various machine vision applications, such as object recognition and tracking, facial recognition, image matching, image / video retrieval, augmented reality, robotic vision and navigation, autonomous driving, three-dimensional structure construction, stereo correspondence, motion tracking, etc.

[0033]

[0066] Exemplary image data encoding and decoding techniques (such as those utilized by image / video encoder 124 and image / video decoder 144) will now be described with reference to Figures 2A-2B and 3A-3B.

[0034]

[0067] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A may be performed by an encoder, such as image / video encoder 124 of FIG. 1. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in time order. Each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, encoding one basic processing unit per iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions of each original picture of video sequence 202.

[0035]

[0068] 2A , an encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original image of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward pass.” In process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0036]

[0069] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward pass) and to generate prediction references 224 for encoding the next original BPU of the original picture (in the reconstruction pass). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0037]

[0070] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may mean receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or any act in any manner for inputting data.

[0038]

[0071] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a predicted reference 224 and perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The predicted reference 224 may be generated from a reconstruction pass of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting predicted data 206, which may be used to reconstruct the original BPU from the predicted data 206 and the predicted reference 224 as the predicted BPU 208.

[0039]

[0072] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 typically differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values ​​(e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values ​​of the original BPU. Each pixel of residual BPU 210 may have a residual value as a result of such subtraction between the corresponding pixel of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.

[0040]

[0073] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "base patterns," each associated with a "transform coefficient." The base patterns may have the same size (e.g., the size of the residual BPU 210). Each base pattern may represent a variation frequency (e.g., frequency of luminance variation) component of the residual BPU 210. None of the base patterns can be regenerated from any combination (e.g., a linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to a discrete Fourier transform of a function, the base patterns are analogous to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0041]

[0074] Different transform algorithms may use different base patterns. For example, various transform algorithms, such as a discrete cosine transform, a discrete sine transform, or the like, may be used in transform stage 212. The transform in transform stage 212 is invertible. That is, the encoder can recover residual BPU 210 by the inverse operation of the transform (referred to as an "inverse transform"). For example, to recover a pixel of residual BPU 210, the inverse transform may multiply the value of the corresponding pixel of the base pattern by each associated coefficient and add the products to generate a weighted sum. In a video coding standard, both the encoder and the decoder may use the same transform algorithm (and therefore the same base pattern). Thus, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the base pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have a smaller number of bits, which can be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0042]

[0075] The encoder may further compress the transform coefficients in the quantization stage 214. In the transform process, different base patterns may represent different fluctuation frequencies (e.g., frequencies of luminance fluctuations). Because the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without significant quality degradation during decoding. For example, in the quantization stage 214, the encoder may generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency base pattern may be converted to zero and some transform coefficients of the low-frequency base pattern may be converted to smaller integers. The encoder may ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also invertible; the quantized transform coefficients 216 can be reconstructed into transform coefficients through the inverse operation of quantization (referred to as "dequantization").

[0043]

[0076] Because the encoder ignores the remainder of such division in a rounding operation, the quantization stage 214 may be lossy. Typically, the quantization stage 214 may contribute the majority of the information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0044]

[0077] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a transform type in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may use the output data of the binary coding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0045]

[0078] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 used in the next iteration of process 200A.

[0046]

[0079] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by an encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.

[0047]

[0080] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.

[0048]

[0081] In general, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already-coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already-coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the coded picture. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0049]

[0082] Referring to process 200B, in the forward pass, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. When an original BPU of a picture is encoded, the prediction reference 224 may include one or more neighboring BPUs already encoded (in the forward pass) and reconstructed (in the reconstructed pass) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below the left, below the right, above the left, or above the right of the original BPU), or any direction defined by the video coding standard used. In the case of intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, or the like.

[0050]

[0083] For another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward pass) and reconstructed (in the reconstructed pass). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed image as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference picture (referred to as the "search window"). The location of the search window within the reference picture may be determined based on the location of the original BPU within the current picture. For example, the search window may be centered at a location in the current picture that has the same coordinates as the original BPU in the reference picture, and may extend outward a predetermined distance. If the encoder identifies a region within the search window that is similar to the original BPU (e.g., using a Pell regression algorithm, a block matching algorithm, or the like), the encoder may determine such a region as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller, equal, larger, or a different shape). Because the reference picture and the current picture are temporally separated on a timeline, the matching region may be expected to "move" to the location of the original BPU over time. The encoder may record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used, the encoder may search for the matching region and determine its associated motion vector for each reference picture. In some embodiments, the encoder may assign weights to pixel values ​​of the matching region in each matching reference picture.

[0051]

[0084] For example, motion estimation can be used to identify various types of motion, such as translation, rotation, zoom, or the like. In the case of inter prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, or the like.

[0052]

[0085] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, and the encoder may predict the original BPU of the current picture. If multiple reference pictures are used, the encoder may shift the matching region of the reference picture according to the individual motion vectors and the average pixel value of the matching region. In some embodiments, if the encoder assigns weights to the pixel values ​​of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values ​​of the shifted matching region.

[0053]

[0086] In some embodiments, inter prediction may be unidirectional or bidirectional. Unidirectional inter prediction may use one or more reference pictures in the same temporal direction relative to the current picture. Unidirectional inter prediction uses a reference picture that precedes the current picture. Bidirectional inter prediction may use one or more reference pictures in both temporal directions relative to the current picture.

[0054]

[0087] Further referring to the forward pass of process 200B, after spatial prediction 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, where the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0055]

[0088] In the reconstruction pass of process 200B, if an intra-prediction mode is selected in the forward pass, after generating the prediction reference 224 (e.g., the current BPU coded and reconstructed within the current picture), the encoder may directly supply the prediction reference 224 to spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current picture). If an inter-prediction mode is selected in the forward pass, after generating the prediction reference 224 (e.g., the current picture coded and reconstructed for all BPUs), the encoder may supply the prediction reference 224 to loop filter stage 232, at which point the encoder may apply a loop filter to the prediction reference 224 to reduce or remove distortions (e.g., blocking artifacts) introduced by inter-prediction. The encoder may apply various loop filter techniques in loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop-filtered reference pictures may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., used as inter-prediction reference pictures for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in the binary coding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.

[0056]

[0089] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A consistent with an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (e.g., image / video decoder 144 of FIG. 1) may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode one fundamental processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions of each picture encoded in video bitstream 228.

[0057]

[0090] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add a reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 to perform a prediction operation in a next iteration of the process 300A.

[0058]

[0091] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a predictive reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to video stream 304 for display and proceed to decode the next coded picture in video bitstream 228.

[0059]

[0092] In binary decoding stage 302, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if video bitstream 228 was packetized and transmitted over a network, the decoder may depacketize video bitstream 228 before providing it to binary decoding stage 302.

[0060]

[0093] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder conforming to a hybrid video coding standard (e.g., the H.26x series). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.

[0061]

[0094] In process 300B, for a coded elementary processing unit (referred to as the “current BPU”) of a coded picture (referred to as the “current picture”) being decoded, prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on the prediction mode used by the encoder to code the current BPU. For example, if intra prediction is used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. Parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. For another example, if inter prediction is used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, or the like. Parameters for the inter-prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights individually associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors individually associated with the matching regions, or the like.

[0062]

[0095] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may sum the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0063]

[0096] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may provide the prediction reference 224 to the loop filter stage 232 to reduce or remove distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., used as inter-prediction reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).

[0064]

[0097] Referring again to FIG. 1 , each of the image / video preprocessor 122, image / video encoder 124, and image / video decoder 144 may be implemented as any suitable hardware, software, or combination thereof. FIG. 4 is a block diagram of an exemplary device 400 for processing image data according to an embodiment of the present disclosure. For example, the device 400 may be a preprocessor, an encoder, or a decoder. As shown in FIG. 4 , the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may be a dedicated device for preprocessing, encoding, or decoding image data. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), application specific integrated circuits (ASICs), or any combination thereof. In some embodiments, processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0065]

[0098] The device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, the memory 404 may include any combination of any number of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. The memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical component.

[0066]

[0099] Bus 410 may be a communication device that transfers data between components internal to device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0067]

[0100] For ease of explanation and without ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module, or may be combined in whole or in part with any other component of the device 400.

[0068]

[0101] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, or a mobile communications network, or the like). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communications ("NFC") adapters, cellular network chips, or the like.

[0069]

[0102] In some embodiments, device 400 may further include a peripheral interface 408 to provide connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., an input interface coupled to a camera or video archive), or the like.

[0070]

[0103] It should be noted that a video codec (e.g., a codec that executes process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., an FPGA, an ASIC, an NPU, or the like).

[0071]

[0104] Figure 5 is a schematic diagram illustrating a bitstream structure according to some embodiments of the present disclosure. In some embodiments, the structure of bitstream 500 may be applied for video bitstream 162 shown in Figure 1. In Figure 5, bitstream 500 includes a video parameter set (VPS) 510, a sequence parameter set (SPS) 520, a picture parameter set (PPS) 530, a picture header 540, and slices 550-570 separated by synchronization markers M1-M7. Each slice 550-570 includes a corresponding header block (e.g., header 552) and data block (e.g., data 554), and each data block includes one or more CTUs (e.g., CTU1-CTUn in data 554). Furthermore, each CTU further includes multiple CUs (e.g., CU1-CUn in Figure 5).

[0072]

[0105] According to some embodiments, bitstream 500, which is a sequence of bits in the form of network abstraction layer (NAL) units or byte streams, forms one or more coded video sequences (CVSs). A CVS includes one or more coding layer video sequences (CLVSs). In some embodiments, a CLVS is a sequence of picture units (PUs), each of which includes one coded picture. Specifically, a PU includes one coded picture that includes zero or one picture header NAL unit (e.g., picture header 540) that includes a picture header syntax structure as payload, and one or more video coding layer (VCL) NAL units, optionally one or more other non-VCL NAL units. A VCL NAL unit is, in some embodiments, a collective term for coded slice NAL units (e.g., slices 550-570) and a subset of NAL units with a reserved value of NAL unit type that are classified as VCL NAL units. A coded slice NAL unit includes a slice header and slice data blocks (eg, header 552 and data 554).

[0073]

[0106] In some embodiments of the present disclosure, a layer may be a set of video coding layer (VCL) NAL units with a particular value of NAL layer ID and one or more associated non-VCL-NAL units. Among these layers, inter-layer prediction can be applied between different layers to achieve high compression performance.

[0074]

[0107] As mentioned above, in general video coding (e.g., VVC / H.266) standards, a picture may be partitioned into a set of CTUs, each of which is further partitioned into coding units (CUs) using a quad tree, a binary tree, or a ternary tree. Multiple CTUs may form a tile, a slice, or a subpicture. If a picture includes three sample arrays to store three color components (e.g., one luma component and two chroma components), a CTU may include N×N (N is an integer) blocks of luma samples, with each block of luma samples associated with two blocks of chroma samples. In some embodiments, an output layer set (OLS) may be defined to support decoding of some but not all layers. An OLS is a set of layers that includes a particular set of layers defined such that one or more layers in the layer set are output layers. Thus, an OLS may include one or more output layers and other layers required to decode one or more output layers for inter-layer prediction.

[0075]

[0108] An improved compression model (ECM) has been proposed and is being used as a new software basis for developing tools that go beyond the VVC standard.

[0076]

[0109] ECM supports a mode called sub-block-based temporal motion vector prediction (SbTMVP). In SbTMVP mode, a CU is divided into 4x4 sub-blocks, and each sub-block obtains its own motion vector from the motion field in a collocated picture. When obtaining the motion vector for each sub-block, a motion shift is applied, which is derived using the motion vector of the neighboring block at the bottom left of the CU. Figure 6 is a schematic diagram illustrating an example process for deriving a motion vector for a sub-block in sub-block-based temporal motion vector prediction (SbTMVP) mode according to some embodiments of the present disclosure. As shown in Figure 6, the motion shift MV A1is obtained from the neighboring block A1 of the CU in the current picture and used to derive the motion field in the co-located picture. For each sub-block, the motion information of its corresponding sub-block in the co-located picture is used to derive the motion information of the sub-block in the current picture. After the motion information of the co-located sub-block 602 is identified, it is converted into a motion vector and reference index of the current sub-block 602'. The reference index of the sub-block is selected from any one of the reference pictures in the reference picture list. The selected reference picture is the one whose scaling factor is closest to 1. After the reference index is identified, temporal motion scaling is applied. Note that if the corresponding sub-block in the co-located picture is non-inter-coded, such as intra-coded or intra-block copy (IBC) coded, the motion information of the central sub-block of the CU is used. As shown in Figure 6, the corresponding sub-block 601 is non-inter-coded, i.e., the sub-block 601 does not contain any motion information. Therefore, the motion information of the central sub-block 602 is used. As a result, the motion information of the sub-block 601' in the current picture corresponding to the sub-block 601 is set to be the same as the motion information of the sub-block 602' corresponding to the sub-block 602 in the co-located picture.

[0077]

[0110] In VVC, temporal motion is used as one of the candidates for merge mode and advanced motion vector prediction (AMVP) mode. Figure 7 shows a co-located block used for temporal motion according to some embodiments of the present disclosure. As shown in Figure 7, temporal motion is derived from the co-located picture and is obtained from the bottom right C0 or center C1 of the co-located block 701. Typically, only one co-located picture is allowed in VVC. This co-located picture is selected by the encoder and signaled in the picture and slice headers in the bitstream.

[0078]

[0111] In ECM, another co-located picture is introduced to further improve the accuracy of temporal motion prediction: two co-located pictures are used, which are two reference frames with different minimum picture order counts (POC) for the frame to be coded.

[0079]

[0112] In the ECM, the reference picture lists for each block are reordered according to template matching (TM) cost. For uni-predictive blocks, all reference pictures in reference picture list 0 and reference picture list 1 are combined to generate a joint list. Then, the TM cost is calculated for each reference picture. The joint list is reordered based on the ascending order of TM cost. For bi-predictive blocks, a list of pairs of reference pictures from list 0 and list 1 is generated and similarly reordered based on the TM cost. The index of the selected pair is signaled.

[0080]

[0113] In the current design, sbTMVP is only allowed in merge mode. However, it cannot be applied to normal inter mode, which is allowed to signal motion vector differentials. Therefore, the advantage of using temporal motion within sub-blocks cannot be fully utilized. Figure 8 is a schematic diagram illustrating an example of normal inter prediction according to some embodiments of the present disclosure. As shown in Figure 8, in normal inter mode, a motion vector 801 is used to predict the entire CU 802, and the CU 802 is not divided into sub-blocks.

[0081]

[0114] In this disclosure, sub-block temporal motion prediction (hereinafter referred to as sbAmvp mode) for regular inter modes is proposed.

[0082]

[0115] According to some embodiments, temporal motion information associated with a sub-block may be obtained from a co-located block. Specifically, a CU is predicted by sub-block temporal motion and signaled motion. FIG. 9 is a schematic diagram illustrating an example of sub-blocks within a CU according to some embodiments of the present disclosure. FIG. 10 is a schematic diagram illustrating an example of motion derivation for a sub-block according to some embodiments of the present disclosure. Referring to FIGS. 9 and 10, a CU 910 is divided into multiple sub-blocks. Each sub-block in the current CU 910 derives its motion information from a co-located block 920 and signaled CU-level motion information. For a sub-block 901, first, its corresponding co-located sub-block 901′ is identified. Then, the motion information of the corresponding co-located sub-block (hereinafter referred to as co-located motion) is scaled to a reference picture 930 (shown in FIG. 10). The scaled motion is added to the signaled CU-level motion and used as the motion for the sub-block 901. It should be noted that the reference pictures of each sub-block may be different or the same from each other, and the signaled CU level motion information may include any subset of the following: inter prediction direction (hereinafter referred to as interDir), reference picture index (hereinafter referred to as refIdx), motion vector predictor (hereinafter referred to as MVP), motion vector differential (hereinafter referred to as MVD).

[0083]

[0116] FIG. 11 shows a flowchart illustrating an example method 1100 of using sub-block temporal motion prediction (sbAmvp mode) for normal inter mode according to some embodiments of the present disclosure. Method 1100 may be performed by a decoder (e.g., process 300A of FIG. 3A or 300B of FIG. 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, one or more processors (e.g., processor 402 of FIG. 4) may perform method 1100. In some embodiments, method 1100 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 11, method 1100 may include the following steps 1102-1106.

[0084]

[0117] CU-level motion information associated with a coding unit (CU) is received in step 1102. The CU is coded using normal inter mode.

[0085]

[0118] In step 1104, the CU is divided into a plurality of sub-blocks. For example, the CU is divided into 4×4 sub-blocks.

[0086]

[0119] In step 1106, the sbAmvp mode is performed on the plurality of sub-blocks. Performing the sbAmvp includes deriving motion information of a sub-block among the plurality of sub-blocks based on signaled CU-level motion information. The signaled CU-level motion information may include any subset of: an inter-prediction direction (hereinafter referred to as interDir), a reference picture index (hereinafter referred to as refIdx), a motion vector predictor (hereinafter referred to as MVP), and a motion vector differential (hereinafter referred to as MVD).

[0087]

[0120] This disclosure provides some examples below for deriving temporal motion information associated with a sub-block from a co-located block: A co-located block is located within a co-located picture.

[0088]

[0121] FIG. 12 shows a flowchart illustrating an example method 1200 for performing sbAmvp mode based on collocated blocks according to some embodiments of the present disclosure. Method 1200 may be performed by a decoder (e.g., process 300A of FIG. 3A or 300B of FIG. 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform method 1200. In some embodiments, method 1200 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 12, method 1200 may include the following steps 1202-1206.

[0089]

[0122] In step 1202, a corresponding co-located sub-block in the co-located picture is identified for the sub-block.

[0090]

[0123] In step 1204, the motion information of the co-located sub-block is scaled to the reference picture, so that the scaled motion can be obtained.

[0091]

[0124] In step 1206, the scaled motion is added to the CU level motion to obtain the sub-block motion.

[0092]

[0125] In step 1208, motion information of the sub-block is obtained, including the motion of the sub-block. The motion information of the sub-block includes the obtained motion of the sub-block, the identified reference picture of the sub-block, and the inter-prediction direction inherited from the CU-level motion information.

[0093]

[0126] In some embodiments, the reference picture for each sub-block is the same as the other sub-blocks. That is, the reference pictures for all of the sub-blocks are the same. In some embodiments, the reference picture for each sub-block is the same as the other sub-blocks, and the signaled CU-level motion information includes interDir, refIdx, MVP, and MVD. Deriving the motion information for the sub-block may further include determining a reference picture for the sub-block using refIdx of the signaled CU-level motion information, determining CU-level motion by adding MVD to MVP, scaling the co-located motion to the reference picture to obtain scaled motion, and adding the scaled motion to the CU-level motion to obtain sub-block motion.

[0094]

[0127] In some embodiments, the reference picture of each sub-block is the same as that of other sub-blocks, and the signaled CU-level motion information includes interDir, MVP, and MVD, and the signaled CU-level motion information is directed to a predefined reference picture. The predefined reference picture may be a co-located reference picture, which means that the reference picture is a co-located picture. The pre-defined reference picture may also be signaled at the slice level, picture level, PPS level, or SPS level. Deriving the motion information of the sub-block may further include determining the CU-level motion by adding the MVD to the MVP, scaling the co-located motion to the pre-defined reference picture to obtain scaled motion, and adding the scaled motion to the CU-level motion.

[0095]

[0128] In some embodiments, if the signaled CU level motion information does not include MVD, the CU level motion vector is set to the motion vector predictor. For example, the reference picture of each sub-block is the same as that of other sub-blocks, and the signaled CU level motion information includes interDir, refIdx, and MVP. In this example, deriving the motion information of the sub-block may further include: determining a reference picture for the sub-block using refIdx of the signaled CU level motion information; setting the CU level motion to be MVP; obtaining scaled motion by scaling the co-located motion to the reference picture; and adding the scaled motion to the CU level motion.

[0096]

[0129] In some embodiments, the reference picture of each block is the same as that of other sub-blocks, and the signaled CU-level motion information includes interDir and MVD, and the reference picture and MVP may be determined using the TM cost. More specifically, for each reference picture, a motion vector predictor candidate is derived in a manner similar to the AMVP process of regular inter prediction. Figure 13 shows an example of template matching according to some embodiments of the present disclosure. As shown in Figure 13, a TM cost is calculated for each motion vector predictor candidate. A template 1310 includes neighboring reconstructed samples to the left 1311 and / or top 1312 of a CU 1320. Reference samples 1330 of the template are generated by each motion vector predictor candidate. Therefore, the reference picture and MVP of the signaled CU-level motion information are set to the picture with the smallest TM cost. In this example, deriving motion information for the sub-block may further include determining a reference picture and MVP for the sub-block using the TM cost, determining CU-level motion by adding the MVD to the MVP, obtaining scaled motion by scaling the co-located motion to the reference picture, and adding the scaled motion to the CU-level motion.

[0097]

[0130] In some embodiments, continuing from the above example, when calculating the TM cost, instead of using only the motion vector predictor candidate, the MVD also participates in the TM cost calculation. Specifically, the reference sample of the template is generated by the motion vector obtained by the motion vector predictor candidate + MVD.

[0098]

[0131] In some embodiments, the reference picture for each sub-block is different from the other sub-blocks, i.e., the reference pictures for multiple sub-blocks are different.

[0099]

[0132] In some embodiments, the reference picture of each sub-block is different from other sub-blocks, and the signaled CU-level motion information includes interDir, refIdx, MVP, and MVD. Derivation of the motion information of a sub-block may further include determining a co-located motion and a reference picture for each sub-block, where the reference picture is selected from any one of the reference pictures in the reference picture list, the selected reference picture having a scaling factor closest to 1, obtaining scaled motion by scaling the co-located motion to the determined reference picture, determining CU-level motion by adding MVD to the MVP, scaling the CU-level motion to the determined reference picture, and adding the scaled motion to the scaled CU-level motion.

[0100]

[0133] In some embodiments, the reference picture of each sub-block is different from other sub-blocks, and the signaled CU level motion information includes interDir, MVP, and MVD, and the signaled CU level motion information is directed to a predefined reference picture. The predefined reference picture may be a co-located reference picture. The pre-defined reference picture may be signaled at the slice level, the picture level, the PPS level, or the SPS level. Deriving the motion information of the sub-block may further include determining a co-located motion and reference picture for each sub-block, where the reference picture is selected from multiple reference pictures in a reference picture list, the selected reference picture having a scaling factor closest to 1; obtaining scaled motion by scaling the co-located motion to the determined reference picture; determining CU level motion by adding MVD to the MVP; scaling the CU level motion to the determined reference picture; and adding the scaled motion to the scaled CU level motion.

[0101]

[0134] In some embodiments, the co-located sub-blocks may be coded using a non-inter mode, where there is no co-located motion. Figure 14 shows an example co-located picture coded using a non-inter mode, according to some embodiments. Figure 15 shows an example averaged motion of the top and left sub-blocks, according to some embodiments of the present disclosure. Figure 16 shows an example averaged motion of the surrounding sub-blocks, according to some embodiments of the present disclosure. As shown in Figure 14, a co-located sub-block 1401 in a co-located picture is non-inter-coded. This co-located sub-block is coded using zero motion, the motion of the center sub-block of the CU (e.g., sub-block 1402 in Figure 14), the averaged motion of the top and left sub-blocks, i.e., (MV T +MV L) / 2, the averaged motion of the surrounding sub-blocks (Fig. 16), i.e. (MV T +MV L +MV R +MV B ) / 4, or one of the averaged motions of the CU. In some embodiments, when calculating the averaged motion of a sub-block, the motion of each sub-block is first scaled to the same reference picture, which may be indicated by the signaled CU-level motion information or may be a co-located reference picture.

[0102]

[0135] In some embodiments, the co-located motion is subtracted by the averaged motion before being added to the signaled CU-level motion information. Specifically, deriving the motion information of the sub-block may further include determining a reference picture for the sub-block using the refIdx of the signaled CU-level motion information, determining the CU-level motion by adding the MVD to the MVP, scaling the co-located motion to the reference picture to obtain scaled motion, calculating an averaged motion of all the scaled motions, subtracting the scaled co-located motion from the averaged motion to obtain differential motion, and adding the differential motion to the CU-level motion.

[0103]

[0136] This disclosure provides several examples below for deriving temporal motion information associated with a sub-block from a motion shift. Specifically, it is proposed to signal motion information at the CU level and use it as a motion shift to derive sub-block temporal motion information. Figure 17 is a schematic diagram illustrating an example of a sub-block in a CU with a motion shift according to some embodiments of the present disclosure. Figure 18 is a schematic diagram illustrating an example of motion derivation for a sub-block based on a motion shift according to some embodiments of the present disclosure. Referring to Figure 17, similar to the SbTMVP mode, a CU 1710 is divided into sub-blocks, such as sub-block 1711, and each sub-block obtains its own motion vector from a motion field 1721 in a co-located picture 1720. When obtaining the motion vector for each sub-block, a motion shift is applied. The motion shift is signaled at the CU level 1730. Then, referring to Figure 18, the motion in a co-located picture 1740 is converted into motion information for the sub-block 1711. The signaled CU-level motion information may include interDir, refIdx, MVP, and MVD.

[0104]

[0137] FIG. 19 shows a flowchart illustrating a method 1900 for performing a motion-shift-based sbAmvp mode according to some embodiments of the present disclosure. Method 1900 may be performed by a decoder (e.g., process 300A of FIG. 3A or 300B of FIG. 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, one or more processors (e.g., processor 402 of FIG. 4) may perform method 1900. In some embodiments, method 1900 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 19, method 1900 may include the following steps 1902-1906.

[0105]

[0138] In step 1902, a motion vector of a sub-block is obtained based on the motion shift.

[0106]

[0139] In step 1904, a corresponding co-located sub-block is identified within the co-located picture for the sub-block.

[0107]

[0140] In step 1906, the motion information of the corresponding co-located sub-block is transformed into motion information of the sub-block.

[0108]

[0141] In some embodiments, the signaled CU-level motion information includes an MVP and an MVD, where the MVP is directed to a co-located reference picture. Deriving the motion information of a sub-block may further include determining a motion shift by adding the MVD to the MVP, identifying a corresponding sub-block in the co-located picture with the motion shift, and converting the motion information of the corresponding sub-block into motion information of the sub-block.

[0109]

[0142] In some embodiments, the signaled CU-level motion information includes only the MVP, and the MVP is directed to the co-located reference picture. Deriving the motion information of the sub-block may further include setting a motion shift to the MVP, identifying a corresponding sub-block in the co-located picture with the motion shift, and converting the motion information of the corresponding sub-block into the motion information of the sub-block.

[0110]

[0143] In some embodiments, when converting the motion information of a corresponding sub-block into the motion information of the sub-block, the reference index of the sub-block is selected from one of the reference pictures in the reference picture list, and the selected reference picture is the one having a scaling factor closest to 1.

[0111]

[0144] In some embodiments, the signaled CU-level motion information includes interDir, refIdx, MVP, and MVD, where MVP is directed to a co-located reference picture. MVP and MVD are used to derive a motion shift, and interDir and refIdx are used to derive a reference picture for a sub-block. Deriving motion information for a sub-block may further include determining the motion shift by adding MVD to MVP, identifying a corresponding sub-block within the co-located picture with the motion shift, determining a reference picture for the sub-block using interDir and refIdx, and scaling the motion information of the corresponding sub-block to the reference picture.

[0112]

[0145] The sbAmvp mode proposed with the various examples above can be combined with local illumination compensation (LIC) mode, multiple hypothesis prediction (MHP) mode, overlapping block motion compensation (OBMC) mode, bi-prediction with CU level weights (BCW) mode, and / or adaptive motion vector resolution (AMVR) mode.

[0113]

[0146] The present disclosure also provides embodiments of MVP derivation and signaling.

[0114]

[0147] In some embodiments, the MVP of a CU can be derived using the same process as the AMVP-MVP derivation process in ECM. That is, first, an MVP candidate list is constructed using motion from spatially neighboring blocks, temporally co-located blocks, non-adjacent neighboring blocks, and a history-based motion vector prediction (HMVP) table. The HMVP table contains historical motion candidates that can be used to assist in predicting future motion and movement. By using these various motion sources and the HMVP table, the candidate list can capture a comprehensive range of potential MVPs. Then, a template matching (TM) cost is calculated for each MVP candidate, and the one with the smallest TM cost is selected and further refined by a template matching process. The TM cost is calculated using neighboring reconstructed samples and reference samples predicted using the MVP candidates.

[0115]

[0148] FIG. 20 shows a flowchart illustrating an example method 2000 of MVP derivation according to some embodiments of the present disclosure. Method 2000 may be performed by a decoder (e.g., process 300A of FIG. 3A or 300B of FIG. 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, one or more processors (e.g., processor 402 of FIG. 4) may perform method 2000. In some embodiments, method 2000 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 20, method 2000 may include the following steps 2002-2008.

[0116]

[0149] In step 2002, an MVP candidate list is constructed using motion from spatially neighboring blocks, temporally co-located blocks, non-neighboring neighboring blocks, and a history-based motion vector prediction (HMVP) table.

[0117]

[0150] In step 2004, the TM cost is calculated for each MVP candidate in the MVP candidate list.

[0118]

[0151] In step 2006, the MVP candidate with the smallest TM cost is selected.

[0119]

[0152] In step 2008, the MVP candidates are refined through a template matching process.

[0120]

[0153] In some embodiments, the reference sample is predicted using the corresponding co-located temporal motion instead of the MVP candidate. Figure 21 shows an example prediction process of the reference sample according to some embodiments of the present disclosure. As shown in Figure 21, the MVP 2110 is directed to the co-located picture 2120, and the corresponding neighboring sub-block 2121 is identified. The temporal motion is derived and used to predict the reference sample of the template 2131 for the current CU 2130 for each corresponding neighboring sub-block 2121.

[0121]

[0154] For MVD signaling, in one example, MVD is signaled in the same manner as AMVP-MVD in ECM, i.e., the absolute values ​​of MVD in the horizontal and vertical directions are binarized and signaled separately. The signs of MVD may be reordered using TM cost and indicated by the signaled index.

[0122]

[0155] In some embodiments, the MVD is constructed by an index indicating a motion magnitude and an index indicating a motion direction. The motion magnitude and motion direction are each selected from a predefined set. In one example, the predefined set for the motion magnitude is {¼ pel, ½ pel, 1 pel, 2 pels, 4 pels, 8 pels}, which is similar to that used in MMVD (Merge with Motion Vector Differential) mode. In another example, to obtain different TMVPs, TMVP (Temporal Motion Vector Prediction) considerations are stored in a 4x4 block grid, and the MVD is a multiple of 4 pels. Therefore, the predefined set for the motion magnitude may be {4 pels, 8 pels, 16 pels, 32 pels, 64 pels, 128 pels}. Figure 22 shows an example of a motion direction according to some embodiments of the present disclosure. The predefined set for the motion direction may be any subset of the directions shown in Figure 22.

[0123]

[0156] The predefined set for the motion magnitudes may be changed according to the sequence resolution, the motion direction, etc. In one example, for a sequence resolution greater than 4,096,000 luma samples, the predefined set for the motion magnitudes is set to {8 pels, 16 pels, 32 pels, 64 pels}. For a sequence resolution equal to or less than 4,096,000 luma samples, the predefined set for the motion magnitudes is set to {4 pels, 8 pels, 16 pels, 32 pels}. In another example, for a horizontal motion direction, the predefined set for the motion magnitudes is set to {4 pels, 8 pels, 16 pels, 32 pels}. For a vertical motion direction, the predefined set for the motion magnitudes is set to {4 pels, 8 pels, 12 pels, 16 pels}.

[0124]

[0157] The two indices of motion magnitude and motion direction can be signaled separately, or alternatively, the two indices are combined into one and signaled.

[0125]

[0158] In some embodiments, the predefined set for motion magnitude is set to {4 pels, 8 pels, 16 pels, 32 pels, 64 pels, 128 pels}. The predefined set for motion direction is set to {0, 4, 8, 12}, with reference to the motion directions shown in Figure 22. A parameter, such as an index or variable (whose value ranges from 0 to 23), is signaled to indicate the motion magnitude and motion direction, and the motion magnitude index and motion direction index are derived as follows:

[0126]

[0159] Movement magnitude index = signaled parameter / (number of movement directions = 4)

[0127]

[0160] Movement direction index = signaled parameter % (number of movement directions = 4)

[0128]

[0161] In another example, all possible combinations of motion magnitude and motion direction are reordered according to TM cost, and a parameter (eg, index) is signaled to indicate the combination to be used.

[0129]

[0162] In some embodiments, a high-level flag is signaled to indicate whether the proposed method is used, and the high-level flag may be an SPS flag, a PPS flag, a picture header level flag, or a slice level flag. In addition, the high-level flag may be signaled only if TMVP is enabled.

[0130]

[0163] In some embodiments, a CU-level flag is signaled to indicate whether the proposed sbAmvp mode is used. The CU-level flag is signaled if the CU is not coded using merge mode.

[0131]

[0164] Note that all proposed methods can be freely combined, for example, for CUs coded using non-merge mode. A CU-level flag is signaled to indicate whether the proposed method is used. If the proposed method is used, a first parameter is signaled to indicate whether the MVD is zero. If the MVD is non-zero, a second parameter indicating the motion magnitude and motion direction of the MVD is signaled. The motion magnitude index and motion direction index are derived as follows:

[0132]

[0165] Movement magnitude index = second parameter / number of movement directions

[0133]

[0166] Movement direction index = second parameter % number of movement directions

[0134] [Table 1]

[0135]

[0167] In some embodiments, if interDir in the CU-level motion information is not signaled, interDir in the motion information of the sub-block is inferred to be a reference direction containing a co-located picture. In some embodiments, if refIdx in the CU-level motion information is not signaled, refIdx in the motion information of the sub-block is set to be an index of a co-located picture. Furthermore, the AMVR, LIC, BCW, and MHP parameters are not signaled and are inferred to be disabled. An OBMC flag is signaled to indicate whether OBMC is applied to the proposed method.

[0136]

[0168] In some embodiments, when a CU is coded using the proposed sbAmvp mode, the CU is divided into sub-blocks, and each sub-block obtains its own motion vector from the motion field in the co-located picture. Deriving the motion information of a sub-block may further include determining a motion shift by adding the MVD to the MVP, identifying a corresponding sub-block in the co-located picture with the motion shift, and converting the motion information of the corresponding sub-block into the motion information of the sub-block.

[0137]

[0169] In some embodiments, if the motion information of the corresponding sub-block is not available, the motion information of the central sub-block is used. If the motion information of the central sub-block is not available, a motion shift is used.

[0138]

[0170] In some embodiments, motion direction candidates may be changed according to a motion magnitude index. FIG. 23 shows exemplary candidates with fixed motion directions according to some embodiments of the present disclosure. FIG. 24 shows exemplary two sets of motion directions according to some embodiments of the present disclosure. As shown in FIG. 24, instead of fixed motion directions (an example is shown in FIG. 23), two sets of motion directions are proposed. In the first set 2410, horizontal and vertical directions are supported. In the second set 2420, four diagonal directions are supported. If the motion magnitude index is an even number, the first set of motion directions 2410 is used. If the motion magnitude index is an odd number, the second set of motion directions 2420 is used. FIG. 25 shows exemplary candidates with two sets of motion directions shown in FIG. 24 according to some embodiments of the present disclosure. As shown in FIG. 25, MVD candidates in the two sets of motion directions are shown.

[0139]

[0171] Instead of using the same number of motion magnitudes and motion directions for all pictures in a sequence, they can be changed for each picture because the characteristics of each picture may differ from each other. Therefore, this disclosure proposes to determine the number of motion magnitudes and motion directions for each picture using an implicit or explicit method. The proposed method can be similarly applied to existing ECM coding modes such as MMVD and SbTMVP.

[0140]

[0172] Figure 26 shows a first group of exemplary motion magnitudes and motion directions according to some embodiments of the present disclosure. Figure 27 shows a second group of exemplary motion magnitudes and motion directions according to some embodiments of the present disclosure. Figure 28 shows a third group of exemplary motion magnitudes and motion directions according to some embodiments of the present disclosure. As shown in Figure 26, the first group 2600 of motion magnitudes and motion directions includes four motion magnitudes with two sets of motion directions. As shown in Figure 27, the second group 2700 of motion magnitudes and motion directions includes eight motion magnitudes with two sets of motion directions. As shown in Figure 28, the third group 2800 of motion magnitudes and motion directions includes four motion magnitudes with eight fixed motion directions.

[0141]

[0173] In some embodiments, the number of motion magnitudes and motion directions is determined according to whether the current picture or current slice is a low-latency picture. For example, if the current picture or slice is a low-latency picture, eight motion magnitudes with two sets of motion directions (second group 2700 in FIG. 27) are used. If the current picture / slice is a non-low-latency picture, eight motion magnitudes with eight fixed motion directions (third group 2800 in FIG. 28) are used.

[0142]

[0174] In some embodiments, the number of motion magnitudes and motion directions is determined according to the POC difference between the current picture and its co-located picture. For example, if the POC difference between the current picture / slice and its co-located picture is greater than a certain positive integer value (e.g., 2), four motion magnitudes with two sets of motion (first group 2600 in FIG. 26) are used. If the POC difference between the current picture / slice and its co-located picture is less than or equal to the positive integer value (e.g., 2), eight motion magnitudes with eight fixed motion directions (third group 2800 in FIG. 28) are used.

[0143]

[0175] In some embodiments, the number of motion magnitudes and motion directions is determined according to the temporal layer. For example, if the temporal layer of the current picture / slice is greater than a certain non-negative integer value (e.g., 3), four motion magnitudes with two sets of motion directions (first group 2600 in FIG. 26) are used. If the temporal layer of the current picture / slice is equal to or less than the certain non-negative integer value (e.g., 3), eight motion magnitudes with eight fixed motion directions (third group in FIG. 28) are used.

[0144]

[0176] In some embodiments, the number of motion magnitudes and motion directions is determined according to the parity of the POC number of the current picture / slice. For example, if the POC of the current picture / slice is even, eight motion magnitudes (third group in FIG. 28) with eight fixed motion directions are used. If the POC of the current picture / slice is odd, eight motion magnitudes (second group 2700 in FIG. 27) with two sets of motion directions are used.

[0145]

[0177] The magnitudes of the movements associated with Figures 26-28 are listed below in Tables 1-3, respectively.

[0146] [Table 2]

[0147] [Table 3]

[0148] [Table 4]

[0149]

[0178] In some embodiments, the motion magnitude and the number of motion directions are determined according to the quantization parameter (QP) of the current picture / slice.

[0150]

[0179] The disclosed methods for determining the magnitude of motion and the number of motion directions can be freely combined.

[0151]

[0180] In some embodiments, the motion magnitudes and the number of motion directions are determined according to the POC difference between the current picture and its co-located picture and whether the current picture is a low-latency picture. Figure 29 shows a flowchart 2900 for determining a group of motion magnitudes and motion directions to use according to some embodiments of the present disclosure. As shown in Figure 29, in step 2902, it is determined whether the POC difference between the current picture and its corresponding co-located picture is greater than a positive integer. If the POC difference between the current picture and its co-located picture is greater than the positive integer value (e.g., 2), in step 2904, four motion magnitudes (first group 2600 in Figure 26) together with two sets of motion directions are used. If the POC difference between the current picture and its co-located picture is less than or equal to the positive integer value (e.g., 2), the process proceeds to step 2906, where it is determined whether the current picture is a low-latency picture. If the current picture is a low latency picture, then eight motion magnitudes (second group 2700 in Figure 27) along with two sets of motion directions are used in step 2908. If the current picture is a non-low latency picture, then eight motion magnitudes (third group 2800 in Figure 28) along with eight fixed motion directions are used in step 2910.

[0152]

[0181] In some embodiments, the number of motion magnitudes and motion directions is determined according to the temporal layer and whether the current picture is a low-latency picture. Figure 30 shows another flowchart 3000 for determining the group of motion magnitudes and motion directions to use according to some embodiments of the present disclosure. As shown in Figure 30, in step 3002, it is determined whether the current picture is a low-latency picture. If the current picture is a low-latency picture in step 3004, eight motion magnitudes (second group 2700 in Figure 27) along with two sets of motion directions are used. If the current picture is a non-low-latency picture, proceed to step 3006 to determine whether the temporal layer of the current picture is greater than a certain non-negative integer (e.g., 4). If the temporal layer of the current picture is greater than the non-negative integer (e.g., 4), in step 3008, four motion magnitudes (first group 2600 in Figure 26) along with two sets of motion directions are used. If the temporal layer of the current picture is less than or equal to the non-negative integer, then in step 3010, eight motion magnitudes (third group 2800 in FIG. 23) are used along with eight fixed motion directions.

[0153]

[0182] In some embodiments, instead of implicitly determining the motion magnitude and number of motion directions per picture, the settings can be explicitly signaled in the picture or slice header.

[0154]

[0183] In some embodiments, two flags are signaled to indicate the number of motion magnitudes and the number of motion directions, respectively. For example, a first flag is signaled to indicate whether the number of motion magnitudes is equal to 4 or 8, and a second flag is signaled to indicate whether the number of motion directions is fixed at 8 or adaptively switches between two sets of motion directions.

[0155]

[0184] In some embodiments, only three types of settings for the number of motion magnitudes and motion directions are supported (FIGS. 26, 27, and 28). A first flag is signaled to indicate whether eight motion magnitudes with eight fixed motion directions (third group in FIG. 28) are used. If eight motion magnitudes with eight fixed motion directions are not used, a second flag is signaled to indicate whether four motion magnitudes with two sets of motion directions (first group 2600 in FIG. 26) or eight motion magnitudes with two sets of motion directions (second group 2700 in FIG. 27) are used.

[0156]

[0185] In some embodiments, the first and second flags are signaled only for non-low latency pictures. For low latency pictures, this is fixed to always be two sets of motion directions and eight motion magnitudes (second group 2700 in Figure 27).

[0157]

[0186] In some embodiments, at the encoder side, a fast algorithm may be applied: basically, the proposed sbAmvp mode may not be tested according to the best coding mode of the current block or the best coding mode of a co-located history block having the same size as the current block.

[0158]

[0187] For example, the number of MVDs tested for the proposed sbAmvp mode is reduced by half if the best coding mode of the current block is not SbTMVP mode, and the history block is also not coded by SbTMVP mode or the proposed method.

[0159]

[0188] To reduce the complexity of the encoder and decoder, in some embodiments, a slice header flag is signaled to indicate whether the proposed sbAmvp mode is enabled for the current slice. The slice header flag is determined according to the enabled areas of the proposed sbAmvp mode and SbTMVP mode in the previous decoded picture. If the enabled area is above a predefined threshold, the proposed mode is enabled for the current slice, and the slice header flag is set to true. The enabled area means a number of samples / pixels coded using sbAmvp / SbTMVP mode. Otherwise, if the enabled area is equal to or less than a predefined threshold, the proposed mode is disabled for the current slice, and the slice header flag is set to false. In some embodiments, the predefined threshold may be changed according to one or more of the QP of the current slice, the temporal layer, the low latency state, etc.

[0160]

[0189] In some embodiments, the proposed sbAmvp mode is disabled for blocks whose width or height is less than a predefined positive value N. For example, the value N may be set equal to 8.

[0161]

[0190] In some embodiments, the proposed sbAmvp mode is disabled for blocks whose width or height is greater than a predefined positive value M. For example, the value M may be set equal to 64.

[0162]

[0191] According to some embodiments, multiple co-located pictures can be used to derive temporal motion. As described above for the temporal motion derivation used in VVC, additional co-located pictures are introduced in the ECM. The additional co-located pictures are different from the original co-located pictures used in VVC. Several methods are proposed herein for selecting co-located pictures for the proposed sbAmvp.

[0163]

[0192] In some embodiments, the co-located picture for sbAmvp mode is always fixed to be the original co-located picture used in VVC.

[0164]

[0193] In some embodiments, when a block is coded using sbAmvp mode, a flag is signaled to indicate which of the two co-located pictures is used to derive the temporal motion. If the flag is equal to 0, the original co-located picture is used. Otherwise, if the flag is equal to 1, the further co-located picture is used.

[0165]

[0194] In some embodiments, if a block is placed in a low-delay picture, a flag is signaled to indicate which of the two co-located pictures is used to derive the temporal motion. If the block is placed in a non-low-delay picture, the co-located picture is always fixed to be the original co-located picture used in VVC.

[0166]

[0195] In some embodiments, the block-level reference picture list reordering technique described above can be applied to reorder the two co-located pictures (the original co-located picture and the further co-located picture). A flag is signaled to indicate which of the co-located pictures (the one with index 0 or the one with index 1) will be used to derive the temporal motion for the proposed sbAmvp mode.

[0167]

[0196] In some embodiments, first, two motion vector predictors of two co-located pictures are derived. Then, the TM costs of the two motion vector predictors are calculated. The one with the smallest TM cost is selected, and its corresponding co-located picture is used to derive the temporal motion for the proposed sbAmvp mode.

[0168]

[0197] In some embodiments, when a block is coded using the proposed sbAmvp mode, a flag is signaled to indicate the inter-prediction direction of the block. The co-located picture used to derive temporal motion is determined according to the inter-prediction direction of the block. The co-located picture is selected from the reference picture list to which the inter-prediction direction of the block is directed.

[0169]

[0198] In some embodiments, the number of co-located pictures is determined according to whether the current picture is a low-latency picture. If the current picture is a low-latency picture, two co-located pictures (the original co-located picture used in VVC and the further co-located picture mentioned above) are used, and a flag is signaled for the sbAmvp-coded block to indicate which of the co-located pictures is used. If the current picture is a non-low-latency picture, only the original co-located picture is used.

[0170]

[0199] According to some embodiments, various methods can be used to select a reference picture. In the sbTMVP mode design, the reference index of a subblock is selected from one of the reference pictures in the reference picture list. The selected reference picture is the one whose scaling factor is closest to 1. Temporal motion scaling is applied after the reference index is identified.

[0171]

[0200] In some embodiments, in the proposed sbAmvp mode, the reference indices of sub-blocks are derived using the same scheme as those in sbTMVP mode, i.e., the reference indices of sub-blocks can be different from each other.

[0172]

[0201] In some embodiments, it is proposed to fix the reference indices of sub-blocks within a block coded using sbAmvp mode to reduce discontinuities between sub-blocks. As an example, the reference indices of sub-blocks are fixed to 0. As another example, first, the reference indices for each sub-block are calculated. For each sub-block, the reference index whose scaling factor is closest to 1 is selected. A histogram is constructed using the selected reference indices of the sub-blocks. Then, the reference index with the largest magnitude is selected for all sub-blocks, and temporal motion scaling is applied to scale all temporal motions of the sub-blocks to the reference index with the largest magnitude.

[0173]

[0202] According to some embodiments, the proposed sbAmvp mode may be combined with adaptive motion vector resolution (AMVR) mode. In normal inter mode, motion vector differentials may be signaled at differential resolutions including {1 / 4 pel, 1 pel, 4 pel, 1 / 2 pel}. To fully utilize the benefits of AMVR mode, the present disclosure proposes combining the proposed sbAmvp mode with AMVR mode.

[0174]

[0203] In some embodiments, when a block is coded using the sbAmvp mode, a parameter (e.g., an index) is signaled to indicate the MVD magnitude resolution. The parameter may be signaled in the same manner as that of the regular inter mode. Considering that TMVP is stored in a 4x4 block grid, the MVD magnitude may be a multiple of 4 pels to obtain different TMVPs. Therefore, when combining the proposed sbAmvp mode with AMVR, the MVD magnitude resolution of sbAmvp is {4 pels, 16 pels, 64 pels, 8 pels}.

[0175]

[0204] In some embodiments, the resolution of the MVD magnitude of sbAmvp can be any subset of {4 pels, 16 pels, 64 pels, 8 pels}. For example, the resolution of the MVD magnitude of sbAmvp is {4 pels, 16 pels}. In another example, the resolution of the MVD magnitude of sbAmvp is {4 pels, 64 pels}.

[0176]

[0205] In some embodiments, the combination of sbAmvp and AMVR modes is enabled for larger sequence resolutions because AMVR mode achieves greater gains for these sequences. For example, if the sequence resolution is 720p or higher (i.e., 1280x720), the combination of sbAmvp and AMVR modes is enabled. If a block is coded using sbAmvp mode, a parameter is signaled to indicate the MVD magnitude resolution. If the sequence resolution is less than 720p, the combination of sbAmvp and AMVR modes is disabled. The MVD magnitude resolution of sbAmvp is always fixed at 4 pels.

[0177]

[0206] According to some embodiments, the proposed sbAmvp mode can be combined with the multiple hypothesis prediction (MHP) mode. In the ECM design, when a block is coded using a bi-prediction mode with unequal weights, the multiple hypothesis prediction (MHP) mode can be applied. When the MHP mode is applied, one or two additional motion information bits are signaled to improve the prediction quality.

[0178]

[0207] In some embodiments, MHP is applied to the sbAmvp mode. When a block is coded using the sbAmvp mode, a flag is signaled to indicate whether MHP is applied. When MHP is applied to an sbAmvp-coded block, a first set of predictive samples is generated using motion information of each sub-block, and a second set of predictive samples is generated using additional motion information. The first and second sets of predictive samples are then blended together.

[0179]

[0208] According to some embodiments, the proposed sbAmvp mode may be combined with local illumination compensation (LIC). LIC is employed in ECM to adjust for illumination changes between a reference picture and a current picture. The predicted samples of an LIC-coded block are modified according to a linear equation: α*p[x]+β, where α is the scale, β is the offset, and p[x] is the predicted sample. Figure 31 shows an example process for deriving α and β according to some embodiments of the present disclosure. Referring to Figure 31, α and β are derived using a reconstructed template 3120 adjacent to the current CU 3110 and its corresponding reference template 3130.

[0180]

[0209] To adjust for illumination changes in the sbAmvp mode, the present disclosure proposes applying the LIC mode to the sbAmvp. In some embodiments, a flag is signaled to indicate whether the LIC mode is applied to an sbAmvp-coded block. When LIC is applied, linear equations are derived at the subblock level. More specifically, in the sbAmvp mode, a block is divided into multiple subblocks. Only subblocks located on CU boundaries can be coded using the LIC mode. For each subblock located on a CU boundary, the LIC scale and subblock offset are derived using its own neighboring reconstructed template and the subblock's corresponding reference template. That is, the LIC scale and offset for each subblock can be different from each other. Figures 32 to 34 illustrate different LIC scales and offsets according to some embodiments of the present disclosure. As shown in Figure 32, for the top-left subblock 3201, the top and left reconstructed samples 3202 are used to derive the LIC parameters. As shown in Figure 33, for the upper right sub-block 3301, only the upper reconstructed sample 3302 is used because the adjacent sample on the left side of the upper right sub-block 3301 has not yet been reconstructed. Similarly, as shown in Figure 34, for the lower left sub-block 3401, only the left reconstructed sample 3402 is used because the upper adjacent sample on the left side of the lower left sub-block 3401 has not yet been reconstructed.

[0181]

[0210] According to some embodiments, merge mode with MVD (MMVD) can be applied to SbTMVP candidates. In ECM, merge mode with MVD (MMVD) is applied to tightly combine merge candidates, but not to SbTMVP candidates. In this disclosure, we propose extending MMVD to SbTMVP mode.

[0182]

[0211] In some embodiments, when a CU is coded using SbTMVP with MMVD mode, an index is signaled to indicate the MVD. The motion shift for the SbTMVP mode is then derived by adding the MVD to the motion derived from a neighboring block (such as the neighboring block on the left (e.g., A1 in FIG. 6)).

[0183]

[0212] It should be noted that all of the proposed methods can be freely combined.

[0184]

[0213] In some embodiments, a non-transitory computer-readable storage medium is provided that stores a video bitstream. The bitstream includes one or more syntax elements that signal the above-mentioned coding unit (CU)-level motion information. In some embodiments, the bitstream includes one or more syntax elements that signal parameters (e.g., flags or indices) included in the above-proposed SbAmvp mode. In some embodiments, the bitstream includes one or more syntax elements that signal parameters (e.g., flags or indices) included in the above-mentioned method.

[0185]

[0214] The embodiments may be further described using the following clauses. 1. A video decoding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter modes, comprising: receiving a bitstream including one or more syntax elements signaling CU-level motion information associated with a coding unit (CU), the CU being coded using a normal inter mode; Dividing a CU into a plurality of sub-blocks; Executing sbAmvp mode for multiple sub-blocks sbAmvp mode includes Deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information. 1. A video decoding method comprising:

[0186] 2. The method of clause 1, wherein the signaled CU-level motion information includes one or more of an inter-prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

[0187] 3. The method of clause 2, wherein deriving the motion information of the sub-blocks is further based on motion shifting.

[0188] 4. Running sbAmvp mode is Obtaining a motion vector of a sub-block based on the motion shift; identifying a corresponding co-located sub-block within a co-located picture of the sub-block; converting the motion information of the corresponding co-located sub-block into the motion information of the sub-block; 3. The method of clause 3, further comprising:

[0189] 5. The signaled CU-level motion information includes MVP and MVD, and the MVP is directed to the co-located reference picture. Implementing the sbAmvp mode is Determine the motion shift by adding the MVD to the MVP 4. The method of clause 4, further comprising:

[0190] 6. The signaled CU-level motion information includes only MVP, and the MVP is directed to a co-located reference picture. Implementing the sbAmvp mode is Setting the MVP as a movement shift 4. The method of clause 4, further comprising:

[0191] 7. The method of clause 4, wherein the reference index for the sub-block is selected from any one of the reference pictures in the reference picture list.

[0192] 8. The signaled CU-level motion information includes interDir, refIdx, MVP, and MVD, where MVP is directed to a co-located reference picture, and implementing sbAmvp mode is determining a motion shift by adding the MVD to the MVP; determining the reference picture of the sub-block using interDir and refIdx; 4. The method of clause 4, further comprising:

[0193] 9. The method of clause 4, wherein the sbAmvp mode is combined with one or more of a local lighting compensation (LIC) mode, a multiple hypothesis prediction (MHP) mode, an overlapping block motion compensation (OBMC) mode, a bi-prediction with CU level weights (BCW) mode, or an adaptive motion vector resolution (AMVR) mode.

[0194] 10. The method of clause 4, further comprising, in response to interDir not being signaled, determining interDir to be a reference direction that includes a co-located picture.

[0195] 11. The method of clause 4, further comprising determining, in response to refIdx not being signaled, that refIdx is an index of a co-located picture.

[0196] 12. Determining the motion shift by adding the MVD to the MVP; identifying corresponding sub-blocks in a co-located picture having a motion shift; converting the motion information of the corresponding sub-block into motion information of the sub-block; 3. The method of clause 3, further comprising:

[0197] 13. Performing the sbAmvp mode before obtaining the motion vector of the sub-block based on the motion shift identifying a corresponding co-located sub-block within a co-located picture of the sub-block; determining whether motion information of the corresponding co-located sub-block is available; determining whether motion information of the central sub-block is available in response to the motion information of the corresponding co-located sub-block being unavailable; responsive to the motion information of the central sub-block being available, converting the motion information of the central sub-block into motion information of the sub-block; 4. The method of clause 4, further comprising:

[0198] 14. The method of clause 2, wherein deriving the motion information of the sub-block is further based on a co-located block of the CU, the co-located block being located within the co-located picture.

[0199] 15. Running sbAmvp mode is identifying a corresponding co-located sub-block within a co-located picture of the sub-block; scaling the motion information of the co-located sub-block to a reference picture; adding the scaled motion to the CU level motion to obtain the sub-block motion; obtaining motion information of the sub-block including motion of the sub-block; The method of clause 14 further includes:

[0200] 16. The method of clause 15, wherein all reference pictures for the multiple sub-blocks are the same.

[0201] 17. The signaled CU-level motion information includes interDir, refIdx, MVP, and MVD, and performing the sbAmvp mode includes: determining a reference picture using the refIdx of the CU-level motion information; Adding MVD to MVP determines CU level movement. The method of clause 16 further includes:

[0202] 18. The signaled CU level motion information includes interDir, MVP, and MVD, and the signaled CU level motion information is directed to a predefined reference picture, and executing the sbAmvp mode is setting the reference picture to be a predefined reference picture; Adding MVD to MVP determines CU level movement. The method of clause 16 further includes:

[0203] 19. The signaled CU-level motion information includes interDir, refIdx, and MVP, and the reference picture is determined using refIdx. Implementing the sbAmvp mode involves: Setting CU level behavior to be MVP The method of clause 16 further includes:

[0204] 20. The signaled CU-level motion information includes interDir and MVD, and executing the sbAmvp mode determining a reference picture and an MVP using a template matching (TM) cost; Adding MVD to MVP determines CU level movement. The method of clause 16 further includes:

[0205] 21. The method of clause 20, wherein the TM cost is calculated using candidate motion vector predictors.

[0206] 22. The method of clause 20, where the TM cost is generated by the motion vector predictor candidate + MVD.

[0207] 23. The method of clause 20, wherein the MVP of the reference picture and the signaled CU-level motion information is set to the picture with the smallest TM cost.

[0208] 24. The method of clause 15, wherein the reference pictures for multiple sub-blocks are different.

[0209] 25. The signaled CU-level motion information includes interDir, refIdx, MVP, and MVD, and executing the sbAmvp mode determining motion information of the co-located sub-blocks and reference pictures for each sub-block, wherein the reference picture is selected from one of a plurality of reference pictures in a reference picture list, and the selected reference picture has a scaling factor that is closest to 1; Adding MVD to MVP determines CU level movement. The method of clause 24 further includes:

[0210] 26. The signaled CU level motion information includes interDir, MVP, and MVD, and the signaled CU level motion information is directed to a predefined reference picture, and executing the sbAmvp mode is setting a co-located reference picture to be a pre-defined reference picture; determining motion information of the co-located sub-blocks and reference pictures for each sub-block, where the reference picture is selected from any one of the reference pictures in the reference picture list, and the selected reference picture has a scaling factor closest to 1; Adding MVD to MVP determines CU level movement. The method of clause 24 further includes:

[0211] 27. The method of clause 26, wherein the predefined reference picture is signaled at the slice level, the picture level, the PPS level, or the SPS level.

[0212] 28. The method of clause 26, wherein the collocated sub-block is coded using a non-inter mode, and the motion information of the collocated sub-block is selected from one of zero motion, motion of a center sub-block of a CU, averaged motion of a top sub-block and a left sub-block, averaged motion of surrounding sub-blocks, or averaged motion of a CU.

[0213] 29. The method of clause 28, wherein the average motion is obtained by scaling the motion of each sub-block to the same reference picture, and the reference picture is indicated by signaled CU-level motion information or is a co-located reference picture.

[0214] 30. Running sbAmvp mode is determining a reference picture for the sub-block using the refIdx of the signaled CU-level motion information; determining CU level movement by adding the MVD to the MVP; obtaining scaled motion by scaling motion information of the co-located sub-block to a reference picture; calculating an average motion of all the scaled motions; subtracting the scaled motion by the averaged motion to obtain a differential motion; Adding the differential motion to the CU level motion The method of clause 28 further includes:

[0215] 31. The method of clause 1, wherein a slice header flag is signaled to indicate whether sbAmvp mode is enabled or disabled.

[0216] 32. The method of clause 1, wherein whether sbAmvp mode is enabled or disabled for the current slice is determined according to the enabled areas of sbAmvp mode and sub-block based temporal motion vector prediction (SbTMVP) mode in the previous decoded picture.

[0217] 33. Determining whether the activated area is above a predefined threshold; In response to the enabled area being above a predefined threshold, enabling sbAmvp mode for the current slice; Disabling the sbAmvp mode for the current slice in response to the enabled area being less than or equal to a predefined threshold. The method of clause 32 further comprises:

[0218] 34. The method of clause 33, wherein the predefined threshold is determined according to one or more of a quantization parameter (QP) of the current slice, a temporal tier of the current slice, and a low latency state.

[0219] 35. The method of clause 1, wherein sbAmvp mode is disabled for a block if the width or height of the block is less than a first predefined positive value.

[0220] 36. The method of clause 1, wherein sbAmvp mode is disabled for a block if the width or height of the block is greater than a second predefined positive value.

[0221] 37. The method of clause 1, wherein the co-located picture for deriving motion information of the sub-block is the original co-located picture used in VVC.

[0222] 38. Decoding a flag indicating a co-located picture for deriving motion information of the sub-block; If the flag is equal to a first value, determining the co-located picture to be a first co-located picture used in VVC; determining the co-located picture to be a second co-located picture different from the first co-located picture if the flag is equal to a second value; 2. The method of clause 1, further comprising:

[0223] 39. The method of clause 1, further comprising decoding a flag indicating which of a first co-located picture used for VVC and a second co-located picture different from the first co-located picture is used to derive motion information of a sub-block, wherein the first co-located picture and the second co-located picture are reordered based on a block-level reference picture recording method.

[0224] 40. Deriving a first motion vector predictor for a first co-located picture for use in generic video coding (VVC); deriving a second motion vector predictor for a second co-located picture different from the first co-located picture; calculating a template matching (TM) cost for each of the first motion vector predictor and the second motion vector predictor; determining the co-located picture used to derive the motion information of the sub-block to be the corresponding co-located picture with the smallest TM cost; 2. The method of clause 1, further comprising:

[0225] 41. Decoding a flag indicating the inter prediction direction of the block; determining a co-located picture to be used for deriving motion information of the sub-block according to an inter-prediction direction of the block; 2. The method of clause 1, further comprising:

[0226] 42. The method of clause 41, wherein the co-located picture is selected from the reference picture to which the inter prediction direction of the block is directed.

[0227] 43. Determining the number of co-located pictures according to whether the current picture is a low-latency picture; In response to the current picture being a low-latency picture, determining that the co-located pictures include a first co-located picture used in VVC and a second co-located picture different from the first co-located picture, and decoding a flag indicating which of the co-located pictures is used to derive motion information of the sub-block; In response to the current picture not being a low-latency picture, determining that the co-located picture used to derive motion information of the sub-block is a first co-located picture used in VVC; 2. The method of clause 1, further comprising:

[0228] 44. Identifying a reference index for the sub-block from one of the reference pictures in the reference picture list, the selected reference picture having a scaling factor closest to 1; After the reference index is identified, performing temporal motion scaling; 2. The method of clause 1, further comprising:

[0229] 45. The method of clause 44, wherein the reference indexes of the multiple subblocks are different.

[0230] 46. ​​The method of clause 1, wherein the reference index of a sub-block within a CU is fixed.

[0231] 47. The method of clause 46, wherein the reference index is fixed to 0.

[0232] 48. Determining a plurality of reference indexes for a plurality of sub-blocks, one reference index per sub-block, wherein a scaling factor of the reference index is closest to 1; constructing a histogram using the determined reference indices of the plurality of sub-blocks; selecting a reference index with the largest magnitude for all sub-blocks; performing temporal motion scaling to scale all temporal motions of the plurality of sub-blocks to the selected reference index; The method of clause 46 further includes:

[0233] 49. The method of clause 1, wherein the sbAmvp mode is combined with an adaptive motion vector resolution (AMVR) mode.

[0234] 50. The method of clause 49, wherein the resolution of the MVD magnitude is {4 pels, 16 pels, 64 pels, 8 pels}.

[0235] 51. The method of clause 49, wherein the sbAmvp mode is combined with an adaptive motion vector resolution (AMVR) mode, and the resolution of the MVD magnitude is selected from one subset of {4 pels, 16 pels, 64 pels, 8 pels}.

[0236] 52. The method of clause 49, wherein whether to enable a combination of sbAmvp mode and AMVR mode is determined according to sequence resolution.

[0237] 53. If the sequence resolution is 720p or higher, a combination of sbAmvp mode and AMVR mode is enabled, and a parameter indicating the resolution of the MVD magnitude is decoded; If the sequence resolution is 720p or less, the combination of sbAmvp mode and AMVR mode is disabled, and the MVD size resolution is fixed at 4 pels. The method of clause 52 further includes:

[0238] 54. The method of clause 1, wherein the sbAmvp mode is combined with a multiple hypothesis prediction (MHP) mode.

[0239] 55. Decoding a flag indicating whether MHP mode is applied; responsive to applying the MHP mode, generating a first set of prediction samples using motion information for each difference block and generating a second set of prediction samples using further motion information; blending the first set of prediction samples with the second set of prediction samples; The method of clause 54 further includes:

[0240] 56. The method of clause 1, wherein the sbAmvp mode is combined with a local illumination compensation (LIC) mode.

[0241] 57. Decoding a flag indicating whether LIC applies; deriving a linear equation at a sub-block level in response to applying LIC, the linear equation being α*p[x]+β, where α is a scale, β is an offset, and p[x] is a predicted sample; The method of clause 56 further includes:

[0242] 58. The method of clause 56, in which the LIC mode is applied to sub-blocks located on CU boundaries.

[0243] 59. The method of clause 58, wherein the LIC scale and offset of a sub-block are derived using the sub-block's corresponding adjacent reconstructed template and the corresponding reference template.

[0244] 60. A method of decoding a bitstream to output one or more pictures for a video stream, comprising: receiving a bitstream; Decoding one or more pictures using coded information of the bitstream; and the decoding includes: Determining the number of motion magnitudes and the number of motion directions for a picture A method comprising:

[0245] 61. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to whether the current picture is a low latency picture.

[0246] 62. The method of clause 61, wherein eight motion magnitudes together with two sets of motion directions are used if the current picture is a low latency picture, and eight motion magnitudes together with eight fixed motion directions are used if the current picture is a non-low latency picture.

[0247] 63. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to a picture order count (POC) difference between the current picture and the corresponding co-located picture.

[0248] 64. The method of clause 63, wherein four motion magnitudes along with two sets of motion directions are used if the POC difference between the current picture and the corresponding co-located picture is greater than a positive integer value, and eight motion magnitudes along with eight fixed motion directions are used if the POC difference between the current picture and the corresponding co-located picture is less than or equal to the positive integer value.

[0249] 65. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to the temporal layer of the current picture.

[0250] 66. The method of clause 65, wherein four motion magnitudes are used together with two pairs of motion directions when the temporal layer of the current picture is greater than a certain non-negative integer value, and eight motion magnitudes are used together with eight fixed motion directions when the temporal layer of the current picture is less than or equal to the certain non-negative integer value.

[0251] 67. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to the parity of the POC number of the current picture.

[0252] 68. The method of clause 67, wherein four motion magnitudes together with two pairs of motion directions are used if the POC of the current picture is odd, and eight motion magnitudes together with eight fixed motion directions are used if the POC of the current picture is even.

[0253] 69. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to the QP of the current picture.

[0254] 70. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to a POC difference between the current picture and the current corresponding co-located picture and whether the current picture is a low latency picture.

[0255] 71. Determining whether a POC difference between the current picture and the corresponding co-located picture of the current picture is greater than a positive integer; In response to a POC difference between the current picture and the corresponding co-located picture being greater than the positive integer, four motion magnitudes are used together with two sets of motion directions; determining whether the current picture is a low latency picture in response to a POC difference between the current picture and the corresponding co-located picture being less than or equal to the positive integer; in response to the current picture being a low latency picture, eight magnitudes are used together with two sets of motion directions; In response to the current picture being a non-low latency picture, eight motion magnitudes are used along with eight fixed motion directions; The method of clause 70 further includes:

[0256] 72. The method of clause 60, wherein the number of motion magnitudes and the number of motion directions are determined according to the temporal layer of the current picture and whether the current picture is a low latency picture.

[0257] 73. Determining whether the current picture is a low latency picture; responsive to the current picture being a low latency picture, eight motion magnitudes are used together with two sets of motion directions; In response to the current picture not being a low latency picture, determining whether a temporal tier of the current picture is greater than a non-negative integer; In response to the temporal layer of the current picture being greater than the non-negative integer, four motion magnitudes are used together with two sets of motion directions; In response to the temporal layer of the current picture being less than or equal to the non-negative integer, eight motion magnitudes are used together with eight fixed motion directions. The method of clause 72 further includes:

[0258] 74. Decoding a first flag indicating a number of motion magnitudes; decoding a second flag indicating the number of motion directions; The method of clause 60 further includes:

[0259] 75. The method of clause 74, wherein a first flag indicates whether the number of motion magnitudes is equal to 4 or 8, and a second flag indicates whether the number of motion directions is fixed at 8 or adaptively switches between two sets of motion directions.

[0260] 76. Decoding a first flag indicating whether eight motion magnitudes are used along with eight fixed motion directions; in response to eight motion magnitudes not being used with eight fixed motion directions, decoding a second flag indicating whether four motion magnitudes with two sets of motion directions or eight motion magnitudes with two sets of motion directions will be used; The method of clause 60 further includes:

[0261] 77. The method of clause 76, wherein the first flag and the second flag are signaled in the case of a non-low delay picture.

[0262] 78. The method of clause 60, wherein eight motion magnitudes along with two sets of motion directions are used for low-latency pictures.

[0263] 79. A video decoding method comprising: decoding an index indicating a motion vector differential (MVD); Deriving the motion shift for sub-block-based temporal motion vector prediction (SbTMVP) mode by adding the MVD to the motion derived from neighboring blocks; 1. A video decoding method comprising:

[0264] 80. A video coding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter modes, comprising: receiving a video sequence; encoding one or more pictures of a video sequence; Generating a bitstream and encoding includes: encoding a coding unit (CU) using a normal inter mode; signaling CU-level motion information associated with the CU; Dividing a CU into a plurality of sub-blocks; Executing sbAmvp mode for multiple sub-blocks sbAmvp mode includes Deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information. A video encoding method comprising:

[0265] 81. The method of clause 80, wherein the signaled CU-level motion information includes one or more of an inter prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

[0266] 82. The method of clause 81, wherein deriving the motion information of the sub-blocks is further based on motion shifting.

[0267] 83. Running sbAmvp mode is Obtaining a motion vector of a sub-block based on the motion shift; identifying a corresponding co-located sub-block within a co-located picture of the sub-block; converting the motion information of the corresponding co-located sub-block into the motion information of the sub-block; The method of clause 82 further includes:

[0268] 84. The signaled CU-level motion information includes MVP and MVD, and the MVP is directed to a co-located reference picture. Implementing the sbAmvp mode involves: Determine the motion shift by adding the MVD to the MVP The method of clause 83 further includes:

[0269] 85. The signaled CU-level motion information includes only MVP, and the MVP is directed to a co-located reference picture. Implementing the sbAmvp mode: Setting the MVP as a movement shift The method of clause 83 further includes:

[0270] 86. The method of clause 83, wherein the reference index for the subblock is selected from one of the reference pictures in the reference picture list.

[0271] 87. The signaled CU-level motion information includes interDir, refIdx, MVP, and MVD, where MVP is directed to a co-located reference picture, and performing the sbAmvp mode: determining a motion shift by adding the MVD to the MVP; determining the reference picture of the sub-block using interDir and refIdx; The method of clause 83 further includes:

[0272] 88. The method of clause 83, wherein the sbAmvp mode is combined with one or more of a local lighting compensation (LIC) mode, a multiple hypothesis prediction (MHP) mode, an overlapping block motion compensation (OBMC) mode, a bi-prediction with CU level weights (BCW) mode, or an adaptive motion vector resolution (AMVR) mode.

[0273] 89. The method of clause 83, further comprising, in response to interDir not being signaled, determining interDir to be a reference direction that includes a co-located picture.

[0274] 90. The method of clause 83, further comprising, in response to refIdx not being signaled, determining refIdx to be an index of a co-located picture.

[0275] 91. Determining the motion shift by adding the MVD to the MVP; identifying corresponding sub-blocks in a co-located picture having a motion shift; converting the motion information of the corresponding sub-block into motion information of the sub-block; The method of clause 82 further includes:

[0276] 92. Performing the sbAmvp mode before obtaining the motion vector of a sub-block based on the motion shift identifying a corresponding co-located sub-block within a co-located picture of the sub-block; determining whether motion information of the corresponding co-located sub-block is available; determining whether motion information of the central sub-block is available in response to the motion information of the corresponding co-located sub-block being unavailable; responsive to the motion information of the central sub-block being available, converting the motion information of the central sub-block into motion information of the sub-block; The method of clause 83 further includes:

[0277] 93. A non-transitory computer-readable storage medium that stores a bitstream generated by an operation, the operation comprising: encoding a coding unit (CU) using a normal inter mode; signaling CU-level motion information associated with the CU; Dividing a CU into a plurality of sub-blocks; Executing sbAmvp mode for multiple sub-blocks sbAmvp mode includes Deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information. 1. A non-transitory computer-readable storage medium comprising:

[0278] 94. A non-transitory computer-readable storage medium of 93, wherein the signaled CU-level motion information includes one or more of an inter-prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

[0279] 95. The non-transitory computer-readable storage medium of clause 94, wherein deriving motion information for the sub-blocks is further based on motion shifts.

[0280] To run 96.sbAmvp mode, Obtaining a motion vector of a sub-block based on the motion shift; identifying a corresponding co-located sub-block within a co-located picture of the sub-block; converting the motion information of the corresponding co-located sub-block into the motion information of the sub-block; 95. The non-transitory computer-readable storage medium of claim 95, further comprising:

[0281] 97. The signaled CU-level motion information includes MVP and MVD, and the MVP is directed to a co-located reference picture. Implementing the sbAmvp mode involves: Determine the motion shift by adding the MVD to the MVP 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0282] 98. The signaled CU-level motion information includes only MVP, and the MVP is directed to a co-located reference picture. Implementing the sbAmvp mode: Setting the MVP as a movement shift 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0283] 99. The non-transitory computer-readable storage medium of clause 96, wherein the reference index for the sub-block is selected from one of the reference pictures in the reference picture list.

[0284] 100. The signaled CU-level motion information includes interDir, redIdx, MVP, and MVD, where MVP is directed to a co-located reference picture, and performing the sbAmvp mode includes: determining a motion shift by adding the MVD to the MVP; determining the reference picture of the sub-block using interDir and refIdx; 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0285] 101. The non-transitory computer-readable storage medium of clause 83, wherein the sbAmvp mode is combined with one or more of a local lighting compensation (LIC) mode, a multiple hypothesis prediction (MHP) mode, an overlapping block motion compensation (OBMC) mode, a bi-prediction with CU level weights (BCW) mode, or an adaptive motion vector resolution (AMVR) mode.

[0286] 102. Actions are determining, in response to interDir not being signaled, that interDir is a reference direction that includes the co-located picture; 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0287] 103. Actions are determining, in response to refIdx not being signaled, that refIdx is an index of a co-located picture; 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0288] 104. Actions are determining a motion shift by adding the MVD to the MVP; identifying corresponding sub-blocks in a co-located picture having a motion shift; converting the motion information of the corresponding sub-block into motion information of the sub-block; 95. The non-transitory computer-readable storage medium of claim 95, further comprising:

[0289] 105. Performing the sbAmvp mode before obtaining the motion vector of a sub-block based on the motion shift includes: identifying a corresponding co-located sub-block within a co-located picture of the sub-block; determining whether motion information of the corresponding co-located sub-block is available; determining whether motion information of the central sub-block is available in response to the motion information of the corresponding co-located sub-block being unavailable; responsive to the motion information of the central sub-block being available, converting the motion information of the central sub-block into motion information of the sub-block; 96. The non-transitory computer-readable storage medium of claim 96, further comprising:

[0290] 106. A video encoding method comprising: receiving a video sequence; encoding one or more pictures of a video sequence; Generating a bitstream and encoding includes: Determining the number of motion magnitudes and motion directions of the picture A video encoding method comprising:

[0291] 107. A non-transitory computer-readable storage medium that stores a bitstream generated by an operation, the operation comprising: Determining the number of motion magnitudes and the number of motion directions for a picture 1. A non-transitory computer-readable storage medium comprising:

[0292] 108. A video encoding method comprising: receiving a video sequence; encoding one or more pictures of a video sequence; Generating a bitstream and encoding includes: deriving a motion shift for a sub-block-based temporal motion vector prediction (SbTMVP) mode by adding a motion vector differential (MVD) to a motion derived from a left neighboring block; Encoding an index indicating the MVD A video encoding method comprising:

[0293] 109. A non-transitory computer-readable storage medium that stores a bitstream generated by an operation, the operation comprising: deriving a motion shift for a sub-block-based temporal motion vector prediction (SbTMVP) mode by adding a motion vector differential (MVD) to a motion derived from a left neighboring block; Encoding an index indicating the MVD 1. A non-transitory computer-readable storage medium comprising:

[0294] 110. A video decoding device, comprising: a memory configured to store instructions; one or more processors configured to execute instructions to cause the apparatus to perform a video compression method according to any one of clauses 1 to 79; 1. A video decoding device comprising:

[0295] 111. A video encoding device, comprising: a memory configured to store instructions; one or more processors configured to execute instructions to cause an apparatus to perform a video decoding method according to any one of clauses 80-92, 106 and 108; 1. A video encoding device comprising:

[0296] 112. A computer program product comprising computer program instructions, the computer program instructions enabling a computer to perform a video decoding method according to any one of clauses 1 to 79.

[0297] 113. A computer program product comprising computer program instructions, the computer program instructions enabling a computer to perform a video encoding method according to any one of clauses 80 to 92, 106 and 108.

[0298] 114. A computer program enabling a computer to carry out a video decoding method according to any one of clauses 1 to 79.

[0299] 115. A computer program enabling a computer to carry out the video encoding method according to any one of clauses 80 to 92, 106 and 108.

[0300]

[0215] In some embodiments, a non-transitory computer-readable storage medium is also provided. In some embodiments, the medium can store all or a portion of a video bitstream with one or more flags indicating applied resampling, such as temporal resampling and spatial resampling. In some embodiments, the medium can store all or a portion of a video bitstream with indices indicating resampling coefficients. In some embodiments, the medium can store instructions that can be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, RPROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. The apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, or memory.

[0301]

[0216] It should be noted that relationship terms such as "first" and "second" are used only to distinguish one entity or operation from another and do not require or imply any actual relationship or ordering between those entities or operations. Furthermore, the terms "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and open-ended in that the item or items following any one of these terms are not intended to be an exhaustive list of such item or items or to be limited only to the listed item or items.

[0302]

[0217] As used herein, the term "or" encompasses all possible combinations except where infeasible, unless specifically stated otherwise. For example, if it is stated that a database may include A or B, then the database may include A, B, A and B, unless specifically stated otherwise or infeasible. As a second example, if it is stated that a database may include A, B or C, then the database may include A, B, C, A and B, A and C, B and C, A and B and C, unless specifically stated otherwise or infeasible.

[0303]

[0218] It should be understood that the above-described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The arithmetic units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0304]

[0219] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. The order of steps depicted in the figures is also intended for illustrative purposes only and is not intended to be limited to any particular order of steps. Thus, one skilled in the art will appreciate that these steps may be performed in different orders while implementing the same method.

[0305]

[0220] In the drawings and this specification, illustrative embodiments are disclosed. However, many variations and modifications to these embodiments may be made. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A video decoding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter modes, comprising: receiving a bitstream including one or more syntax elements signaling CU-level motion information associated with a coding unit (CU), the CU being coded using a normal inter mode; Dividing the CU into a plurality of sub-blocks; performing the sbAmvp mode on the plurality of sub-blocks; and executing the sbAmvp mode includes: deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information.

1. A video decoding method comprising:

2. 2. The method of claim 1, wherein the signaled CU-level motion information includes one or more of an inter-prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

3. The method of claim 2 , wherein deriving the motion information for the sub-blocks is further based on a motion shift.

4. Executing the sbAmvp mode includes: obtaining a motion vector of a sub-block based on the motion shift; identifying a corresponding co-located sub-block within a co-located picture of the sub-block; converting motion information of a corresponding co-located sub-block into motion information of said sub-block; The method of claim 3 further comprising:

5. The signaled CU level motion information includes an MVP and an MVD, and the MVP is directed to a co-located reference picture, and performing the sbAmvp mode includes: determining the motion shift by adding the MVD to the MVP; The method of claim 4 further comprising:

6. The signaled CU-level motion information includes only MVP, and the MVP is directed to a co-located reference picture, and performing the sbAmvp mode includes: Setting the MVP as the motion shift. The method of claim 4 further comprising:

7. The method of claim 4 , wherein the reference index for the sub-block is selected from one of the reference pictures in a reference picture list.

8. The signaled CU level motion information includes interDir, refIdx, MVP, and MVD, and the MVP is directed to a co-located reference picture, and performing the sbAmvp mode includes: determining the motion shift by adding the MVD to the MVP; determining a reference picture for the sub-block using the interDir and refIdx; The method of claim 4 further comprising:

9. 9. The method of claim 4, wherein the sbAmvp mode is combined with one or more of a local illumination compensation (LIC) mode, a multiple hypothesis prediction (MHP) mode, an overlapped block motion compensation (OBMC) mode, a bi-prediction with CU level weights (BCW) mode, or an adaptive motion vector resolution (AMVR) mode.

10. The method of claim 4 , further comprising: in response to interDir not being signaled, determining the interDir to be the reference direction that contains the co-located picture.

11. The method of claim 4 , further comprising: in response to the refIdx not being signaled, determining the refIdx to be the index of the co-located picture.

12. determining a motion shift by adding the MVD to the MVP; identifying corresponding sub-blocks within the co-located picture having the motion shift; converting the motion information of the corresponding sub-block into the motion information of the sub-block; The method of claim 3 further comprising:

13. performing the sbAmvp mode before obtaining the motion vector of the sub-block based on the motion shift, identifying a corresponding co-located sub-block within a co-located picture of the sub-block; determining whether motion information of the corresponding co-located sub-block is available; and determining whether motion information of a central sub-block is available in response to the motion information of the corresponding collocated sub-block being unavailable; responsive to the motion information of a central sub-block being available, converting the motion information of the central sub-block into motion information of the sub-block; The method of claim 4 further comprising:

14. 1. A video coding method using sub-block temporal motion prediction (sbAmvp mode) for normal inter modes, comprising: receiving a video sequence; encoding one or more pictures of the video sequence; Generating a bitstream wherein said encoding comprises: encoding a coding unit (CU) using a normal inter mode; signaling CU-level motion information associated with the CU; Dividing the CU into a plurality of sub-blocks; performing the sbAmvp mode on the plurality of sub-blocks; and executing the sbAmvp mode includes: deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information. A video encoding method comprising:

15. 15. The method of claim 14, wherein the signaled CU-level motion information includes one or more of an inter-prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

16. The method of claim 15 , wherein deriving the motion information for the sub-blocks is further based on a motion shift.

17. Executing the sbAmvp mode includes: obtaining a motion vector of a sub-block based on the motion shift; identifying a corresponding co-located sub-block within a co-located picture of the sub-block; converting motion information of a corresponding co-located sub-block into motion information of said sub-block; 17. The method of claim 16, further comprising:

18. 1. A non-transitory computer-readable storage medium that stores a bitstream generated by an operation, the operation comprising: encoding a coding unit (CU) using a normal inter mode; signaling CU-level motion information associated with the CU; Dividing the CU into a plurality of sub-blocks; performing an sbAmvp mode on the plurality of sub-blocks; and executing the sbAmvp mode includes: deriving motion information for a sub-block among the plurality of sub-blocks based on the signaled CU-level motion information.

1. A non-transitory computer-readable storage medium comprising:

19. 20. The non-transitory computer-readable storage medium of claim 18, wherein the signaled CU-level motion information includes one or more of an inter-prediction direction (interDir), a reference picture index (refIdx), a motion vector predictor (MVP), or a motion vector differential (MVD).

20. The non-transitory computer-readable storage medium of claim 19 , wherein deriving the motion information for the sub-blocks is further based on a motion shift.

21. 1. A video decoding device, comprising: a memory configured to store instructions; one or more processors configured to execute said instructions to cause said device to perform a video compression method according to any one of claims 1 to 13; 1. A video decoding device comprising:

22. 1. A video encoding device, comprising: a memory configured to store instructions; one or more processors configured to execute said instructions to cause said device to perform a video decoding method according to any one of claims 14 to 17; 1. A video encoding device comprising:

23. A computer program product comprising computer program instructions, the computer program instructions enabling a computer to carry out the video compression method according to any one of claims 1 to 13.

24. A computer program product comprising computer program instructions, the computer program instructions enabling a computer to carry out the video decoding method according to any one of claims 14 to 17.

25. A computer program enabling a computer to carry out the video compression method according to any one of claims 1 to 13.

26. A computer program enabling a computer to carry out the video decoding method according to any one of claims 14 to 17.

Citation Information

Cited By

  • Method, electronic device and computer program for sub-block motion vector prediction

    JP2026502489A