Bilinear model for video coding

CN122845812APending Publication Date: 2026-09-29ALIBABA INNOVATION PRIVATE LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610379145.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2026-03-06
Filing Date
2026-03-26
Publication Date
2026-09-29

Smart Images

  • Figure CN122845812A_ABST
    Figure CN122845812A_ABST
Patent Text Reader

Abstract

A method for encoding video data includes determining a motion vector predictor (MVP) for a target block based on a control point motion vector (CPMV), wherein a bilinear model is characterized by the CPMV; determining motion vector differences (MVDs), wherein each motion vector difference is based on a difference between a corresponding MVP for each control point and its CPMV; and setting the MVDs in a bitstream in response to using the bilinear model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Application No. 63 / 778,615, entitled “Bilinear Model for Video Coding,” filed March 27, 2025, and U.S. Patent Application No. 19 / 559,654, filed March 6, 2026, both of which are incorporated herein by reference in their entirety. Technical Field

[0002] This disclosure relates generally to video processing, and more specifically to bilinear models for video coding. Background Technology

[0003] Video consists of a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. This compression process is typically called encoding, while the decompression process is typically called decoding. There are many video coding formats that use standardized video coding techniques, the most common being based on prediction, transform, quantization, entropy coding, and loop filtering. Standardization organizations have developed video coding standards that specify particular video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard. As video standards adopt increasingly advanced video coding techniques, the coding efficiency of new video coding standards also increases. Summary of the Invention

[0004] Embodiments of this disclosure provide methods and apparatus for bilinear models used in video coding.

[0005] According to some embodiments, a method for encoding video data includes: determining a motion vector predictor (MVP) for a target block based on a control point motion vector (CPMV), wherein a bilinear model is characterized by the CPMV; determining a motion vector difference (MVD), wherein each motion vector difference is based on the difference between the MVP corresponding to each control point and its CPMV; and setting the MVD in a bitstream in response to using the bilinear model.

[0006] According to some embodiments, a method for decoding a video bitstream includes: receiving a bitstream comprising encoded video data of a target block; determining motion vector difference (MVD) based on the bitstream, wherein each motion vector difference is based on the difference between a predicted motion vector value (MVP) corresponding to each control point of the target block and its control point motion vector (CPMV); and reconstructing the target block based on the CPMV derived from the MVD in response to using a bilinear model, wherein the bilinear model is characterized by the CPMV.

[0007] According to some embodiments, a method for storing a bitstream includes: receiving a video sequence comprising one or more images; generating a bitstream by the following steps, the bitstream comprising encoded information associated with the video sequence: determining motion vector prediction values ​​(MVPs) for target blocks of the one or more images based on control point motion vectors (CPMVs), wherein a bilinear model is characterized by the CPMVs; determining motion vector differences (MVDs), wherein each motion vector difference is based on the difference between the MVP and its CPMV corresponding to each control point; and encoding the MVDs in the bitstream in response to using the bilinear model; and storing the bitstream in a non-transitory computer-readable medium. Attached Figure Description

[0008] Embodiments and aspects of this disclosure are illustrated in the following detailed description and accompanying drawings. Various features shown in the figures are not drawn to scale.

[0009] Figure 1 The structure of an exemplary video sequence according to some embodiments of this disclosure is shown.

[0010] Figure 2 A schematic diagram of an exemplary encoder of a video encoding system according to some embodiments of the present disclosure is shown.

[0011] Figure 3 A block diagram of an exemplary decoder for a video encoding system according to some embodiments of the present disclosure is shown.

[0012] Figure 4 This is a block diagram of an exemplary apparatus for encoding or decoding video according to some embodiments of the present disclosure.

[0013] Figure 5A and Figure 5B These are two schematic diagrams illustrating a control point-based affine model for a block according to some embodiments of the present disclosure.

[0014] Figure 6 This is a schematic diagram illustrating the motion vector of the center sample of each sub-block of a block according to some embodiments of the present disclosure.

[0015] Figure 7 The control point motion vector inheritance according to some embodiments of this disclosure is illustrated.

[0016] Figure 8A and Figure 8B The present disclosure illustrates spatial neighboring blocks for deriving affine merging according to some embodiments, and affine advanced motion vector prediction (AMVP) candidates for deriving inheritance candidates and first-type construction candidates.

[0017] Figure 9 The distribution of candidate positions for constructing an affine merging pattern is shown according to some embodiments of the present disclosure.

[0018] Figure 10 The first type of affine merge / AMVP candidate is shown according to some embodiments of this disclosure.

[0019] Figure 11A and Figure 11B This is a schematic diagram of a first history parameter table (HPT) and a second history parameter table (HPT) according to some embodiments of the present disclosure.

[0020] Figure 12 The illustration shows some embodiments of the present disclosure for affine merging candidate derivation of adjacent sub-blocks based on regression.

[0021] Figure 13 Sub-blocks MV and pixels are shown according to some embodiments of this disclosure.

[0022] Figure 14 An example 3×3 square search pattern is shown according to some embodiments of this disclosure.

[0023] Figure 15 An example 3×3 cross-search pattern is shown according to some embodiments of the present disclosure.

[0024] Figure 16 This is a schematic diagram illustrating the adaptive search steps in an example 3×3 cross-search pattern according to some embodiments of the present disclosure.

[0025] Figure 17 This is a schematic diagram illustrating example sub-block level pre-interpolation according to some embodiments of the present disclosure.

[0026] Figure 18 This is a schematic diagram illustrating example samples on integer and fractional search points according to some embodiments of the present disclosure.

[0027] Figure 19AThe illustration shows how affine parameters are corrected by fixing the motion vector (CPMV) of the upper left control point as the base MV according to some embodiments of the present disclosure.

[0028] Figure 19B The illustration shows how affine parameters are modified by fixing the upper right corner CPMV as the base MV according to some embodiments of the present disclosure.

[0029] Figure 19C The illustration shows how affine parameters are modified by fixing the lower left corner CPMV as the base MV according to some embodiments of the present disclosure.

[0030] Figure 20 The upper and left templates for affine motion compensation are shown according to some embodiments of the present disclosure.

[0031] Figure 21 The image shows the MV of a sub-template of an affine motion coding block according to some embodiments of the present disclosure.

[0032] Figure 22 The image shows the MV of a sub-template of an affine motion coding block according to some embodiments of the present disclosure.

[0033] Figure 23A An integer template matching (TM) search process is illustrated according to some embodiments of the present disclosure.

[0034] Figure 23B The half-pixel™ search process according to some embodiments of the present disclosure is illustrated.

[0035] Figure 24 This is a flowchart of an affine merging pattern process according to some embodiments of the present disclosure.

[0036] Figure 25 This is a schematic diagram illustrating the inheritance of control point motion vectors according to some embodiments of the present disclosure.

[0037] Figure 26 This is a schematic diagram illustrating the construction of a first type of affine merge / AMVP candidate according to some embodiments of the present disclosure.

[0038] Figures 27A-27D This is a schematic diagram illustrating an example of bilinear parameter correction according to some embodiments of the present disclosure.

[0039] Figure 28 The top and left templates for a current bilinear mode coded block are shown according to some embodiments of the present disclosure.

[0040] Figure 29 The MV of a sub-template of a bilinear mode coded block is shown according to some embodiments of the present disclosure.

[0041] Figure 30 The MV of a sub-template of a bilinear mode coded block is shown according to some embodiments of the present disclosure.

[0042] Figure 31 This is a flowchart of an example method for encoding a video bitstream according to some embodiments of the present disclosure.

[0043] Figure 32 This is a flowchart of an example method for decoding a video bitstream according to some embodiments of the present disclosure. Detailed Implementation

[0044] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise stated, the same numbers in different figures characterize the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with relevant aspects of this disclosure as set forth in the appended claims. Specific aspects of this disclosure are described below in more detail. In the event of any conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0045] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) is currently developing the Universal Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0046] To achieve this goal, since 2015, JVET has been continuously developing technologies that surpass HEVC using the Joint Exploration Model (JEM) reference software. With the incorporation of various coding techniques into JEM, it has achieved significantly higher coding performance than HEVC. In October 2017, VCEG and MPEG issued a joint call for proposals (CfP), officially launching the development of a next-generation video compression standard that surpasses HEVC. Responses to the CfP were evaluated at the JVET meeting in San Diego in April 2018, and formal development of the VVC standard began in April 2018.

[0047] Since April 2018, the VVC standard has progressed smoothly, continuously incorporating more coding technologies to provide better compression performance. VVC adopts the hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263. In July 2020, the first version of the VVC standard was finalized and officially released as an international standard. Subsequently, JVET began exploring new coding tools to further improve the coding performance of the VVC standard. In January 2021, the Enhanced Compression Model (ECM) was proposed and used as the foundation for developing new software that surpasses the VVC standard.

[0048] Figure 1 The structure of an exemplary video sequence according to some embodiments of this disclosure is shown. Video sequence 100 may be live video or video that has already been captured and archived. Video sequence 100 may be real-scene video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real-scene video with augmented reality effects). Video sequence 100 may originate from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider. Figure 1 As shown, video sequence 100 may include a series of images arranged chronologically along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108.

[0049] When video is being compressed or decompressed, useful information about the image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes can include variations in pixel position, brightness, or color. For example, changes in the position of a set of pixels can reflect the motion of a target represented by these pixels between two images (e.g., the reference image and the current image).

[0050] For example, such as Figure 1 As shown, image 102 is an I-image, using itself as a reference image. Image 104 is a P-image, using image 102 as its reference image, as indicated by the arrow. Image 106 is a B-image, using images 104 and 108 as its reference images, as indicated by the arrow. In some embodiments, the reference image of an image may or may not be immediately before or after the image. For example, the reference image of image 104 may be an image preceding image 102, that is, an image not immediately preceding image 104. Figure 1 The reference images 102-106 shown above are merely examples and are not intended to limit this disclosure.

[0051] Due to computational complexity, in some embodiments, a video codec may split an image into multiple basic segments and encode or decode the image segment by segment. That is, the video codec does not encode or decode the entire image at once. Such basic segments are referred to herein as basic processing units ("BPUs"). For example, Figure 1 An exemplary structure 110 for images (e.g., any one of images 102-108) of video sequence 100 is also shown. For example, structure 110 can be used to segment image 108. Figure 1 As shown, image 108 is divided into 4×4 basic processing units. In some embodiments, the basic processing unit may be called a "coding tree unit" (CTU) in some video coding standards (e.g., AVS3, H.265 / HEVC, or H.266 / VVC), or a "macroblock" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC). In AVS3 or VVC, the coding tree unit (CTU) can be the largest block unit and can be as large as 128×128 luminance samples (plus the corresponding chroma samples according to the chroma format).

[0052] Figure 1The basic processing units in the image are for illustrative purposes only. These basic processing units can have variable sizes in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any arbitrary shape and size of pixels. The size and shape of the basic processing units can be selected for the image based on a balance between coding efficiency and the level of detail to be maintained within the basic processing units.

[0053] The basic processing unit can be a logical unit that may include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luminance component (Y) representing non-color luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, wherein the size of the luminance and chrominance components may be the same as that of the basic processing unit. In some video coding standards, the luminance and chrominance components may be referred to as a "coding tree block" (CTB). Operations performed on a basic processing unit may be repeated on its luminance and chrominance components.

[0054] Video coding involves multiple operational stages, and the size of the basic processing unit may still be too large to process. Therefore, it can be further divided into multiple segments referred to herein as "basic processing subunits." For example, in the mode decision stage, the encoder can divide the basic processing unit into multiple basic processing subunits and determine the prediction type for each individual basic processing subunit. Figure 1 As shown, the basic processing unit 112 in structure 110 is further divided into 4×4 basic processing sub-units. For example, the coding tree unit CTU can be further divided into coding units (CUs) using a quadtree, binary tree, or extended binary tree. Figure 1The basic processing subunits described herein are for illustrative purposes only. Different basic processing units for the same image can be divided into basic processing subunits using different schemes. These basic processing subunits may be referred to as “coding units” (“CUs”) in some video coding standards (e.g., AVS3, H.265 / HEVC, or H.266 / VVC), or as “blocks” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC). The size of a basic processing subunit may be equal to or smaller than the size of a basic processing unit. Similar to the basic processing unit, a basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., stored in a video frame buffer). Operations performed on a basic processing subunit can be repeatedly performed on its luminance and chrominance components. Depending on processing needs, this division can be performed at deeper levels, and different schemes may be used to divide the basic processing units at different stages. At the leaf nodes of the partitioning structure, encoding information such as the encoding mode (e.g., intra-frame prediction mode or inter-frame prediction mode), motion information required for the corresponding encoding mode (reference index, motion vector (MV) difference, etc.) (if inter-frame coded), and quantization residual coefficients are transmitted.

[0055] In some cases, the basic processing subunit may still be too large to be processed in certain operational stages of video coding, such as the prediction or transform stage. Therefore, the encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs"), at which prediction operations can be performed. Similarly, the encoder can further divide the basic processing subunit into smaller segments (e.g., called "transform blocks" or "TBs"), at which transform operations can be performed. The partitioning scheme of the same basic processing subunit in the prediction and transform stages can differ. For example, the prediction blocks (PBs) and transform blocks (TBs) of the same CU can have different sizes and numbers. The operations in the mode decision stage, prediction stage, and transform stage will be discussed in later paragraphs. Figure 2 and Figure 3 The examples provided are described in detail.

[0056] Figure 2A schematic diagram of an exemplary encoder 200 for a video coding system (e.g., AVS3 or H.26x series) according to some embodiments of this disclosure is shown. The input video is processed block by block. As discussed above, in some coding standards (e.g., VVC), the Code Tree Unit (CTU) is the largest block unit and can be as large as 128 × 128 luma samples (plus corresponding chroma samples according to the chroma format). A CTU can be further divided into multiple CUs using quadtrees, binary trees, or ternary trees. Reference Figure 2 The encoder 200 can receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term "receive" can refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or otherwise inputting data. The encoder 200 can encode the video sequence 202 into a video bitstream 228. Similar to... Figure 1 Video sequence 100 and video sequence 202 may include a set of images (referred to as "original images") arranged in chronological order. Similar to... Figure 1 In structure 110, encoder 200 can divide any raw image of video sequence 202 into multiple basic processing units, multiple basic processing sub-units, or multiple regions for processing. In some embodiments, encoder 200 can perform the process at the level of basic processing units for the raw images of video sequence 202. For example, encoder 200 can perform the process iteratively. Figure 2 In the process, encoder 200 can encode a basic processing unit in one iteration of the process. In some embodiments, encoder 200 can target a region of the original image of video sequence 202 (e.g., Figure 1 The process (parts 114-118) is executed in parallel.

[0057] Components 202, 2042, 2044, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path." In Figure 2 In this process, encoder 200 can feed the basic processing unit (referred to as "raw BPU") of the original image of video sequence 202 to two prediction stages, namely intra-frame prediction (also known as "intra-image prediction" or "spatial prediction") stage 2042 and inter-frame prediction (also known as "inter-image prediction," "motion compensation," "motion compensation prediction," or "temporal prediction") stage 2044, to perform prediction operations and generate corresponding prediction data 206 and prediction BPU 208. In particular, encoder 200 can receive the raw BPU and prediction reference 224, which can be generated from the reconstruction path of the previous iteration of the process.

[0058] The purpose of intra-frame prediction stage 2042 and inter-frame prediction stage 2044 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224 that can be used to reconstruct the original BPU into the predicted BPU 208. In some embodiments, intra-frame prediction can use pixels from one or more encoded neighboring BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in the intra-frame prediction may include neighboring BPUs, such that spatially adjacent samples can be used to predict the current block. The intra-frame prediction can reduce the inherent spatial redundancy of the image.

[0059] In some embodiments, inter-frame prediction can use regions from one or more encoded images (“reference images”) to predict the current BPU. That is, prediction reference 224 in the inter-frame prediction can include the encoded images. The inter-frame prediction can reduce the inherent temporal redundancy of the images.

[0060] In the forward path, encoder 200 performs prediction operations in intra-frame prediction phase 2042 and inter-frame prediction phase 2044. For example, in intra-frame prediction phase 2042, encoder 200 may perform the intra-frame prediction. For a given original BPU of an image being encoded, prediction reference 224 may include one or more adjacent BPUs that have already been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. Encoder 200 may generate predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, encoder 200 may perform the extrapolation at the pixel level, for example, extrapolating the corresponding pixel value for each pixel of predicted BPU 208. The adjacent BPU used for extrapolation can be located from various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., at the lower left, lower right, upper left, or upper right corner of the original BPU), or any direction defined in the video coding standard used. For the intra-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPU used, the size of the adjacent BPU used, the extrapolation parameters, the orientation of the adjacent BPU used relative to the original BPU, etc.

[0061] In another example, during the inter-frame prediction phase 2044, encoder 200 may perform the inter-frame prediction. For a given original BPU of the current image, prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a BPU-by-BPU basis. For example, encoder 200 may add reconstructed residual BPU 222 to prediction BPU 208 to generate a reconstructed BPU. After all reconstructed BPUs of the same image have been generated, encoder 200 may generate a reconstructed image as a reference image. Encoder 200 may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference image. The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When encoder 200 identifies a region similar to the original BPU in the search window (e.g., using a pixel recursive algorithm, block matching algorithm, etc.), encoder 200 can determine such a region as a matching region. The matching region may have different specifications than the original BPU (e.g., less than, equal to, or greater than the original BPU, or have a different shape). Because the reference image and the current image are temporally separated in the timeline (e.g., as...), Figure 1 As shown), the matching region can be considered to "move" to the position of the original BPU over time. The encoder 200 can record the direction and distance of this movement as a "motion vector (MV)". In other words, MV is the positional difference between a reference block in the reference image and the current block in the current image. In inter-frame prediction, the reference block is used as a predictor for the current block; therefore, the reference block is also referred to as the prediction block. When using multiple reference images (e.g., ...), ... Figure 1 When using image 106 in the reference image, encoder 200 can search for matching regions for each reference image and determine the associated MV of the matching regions. In some embodiments, encoder 200 can assign weights to the pixel values ​​of the matching regions of each matching reference image.

[0062] The motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, a reference index, the position (e.g., coordinates) of the matching region, the MV associated with the matching region, the number of reference images, the weights associated with the reference images, or other motion information.

[0063] To generate the predicted BPU 208, encoder 200 can perform a "motion compensation" operation. This motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., MV) and prediction reference 224. For example, encoder 200 can move the matching region of the reference image according to the MV, where encoder 200 can predict the original BPU of the current image. When using multiple reference images (e.g., ... Figure 1 When the encoder 200 moves the matching region of the reference image (image 106) according to the respective MV and average pixel value of the matching region, the encoder 200 can move the matching region of the reference image. In some embodiments, if the encoder 200 has already assigned weights to the pixel values ​​of the matching regions of each matching reference image, the encoder 200 can perform a weighted summation of the pixel values ​​of the moved matching region.

[0064] In some embodiments, the inter-frame prediction may utilize unidirectional or bidirectional prediction, and may be either unidirectional or bidirectional. Unidirectional inter-frame prediction may use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in the image is a one-way inter-frame predicted image, wherein the reference image (i.e., image 102) precedes image 104. In one-way prediction, only one MV pointing to a reference image is used to generate the prediction signal for the current block.

[0065] On the other hand, bidirectional inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 Image 106 in the image is a bidirectional inter-frame predicted image, wherein the reference images (e.g., images 104 and 108) are relative to image 104 in two temporal directions. In bidirectional prediction, two MVs (each pointing to its respective reference image) are used to generate the predicted signal for the current block. After generating video bitstream 228, the MVs and reference indices can be sent to the decoder in video bitstream 228 to identify where one or more predicted signals for the current block originate.

[0066] For inter-frame prediction CU, motion parameters may include MV, reference image index, and reference image list usage index, or other additional information required for the coding features to be used. Motion parameters can be signaled explicitly or implicitly. In some embodiments, under certain inter-frame coding modes (e.g., skip mode or direct mode), motion parameters (e.g., MV difference and reference image index) are not encoded or signaled in the video bitstream 228. Instead, the motion parameters can be derived at the decoding end using the same rules defined in encoder 200. Details of the skip mode and the direct mode will be discussed in the following paragraphs.

[0067] Following the intra-frame prediction phase 2042 and the inter-frame prediction phase 2044, in the mode decision phase 230, the encoder 200 can select a prediction mode (e.g., one of the intra-frame prediction or the inter-frame prediction) for the current iteration of the process. For example, the encoder 200 can execute a rate-distortion optimization method, wherein the encoder 200 selects a prediction mode based on the bit rate of a candidate prediction mode and the distortion of a reference image reconstructed under the candidate prediction mode to minimize the value of the cost function. Based on the selected prediction mode, the encoder 200 can generate a corresponding prediction BPU 208 (e.g., a prediction block) and prediction data 206.

[0068] In some embodiments, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder 200 can subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210, also referred to as the prediction residual.

[0069] For example, encoder 200 can subtract the pixel value (e.g., grayscale or RGB value) of predicted BPU 208 from the value of the corresponding pixel in the original BPU. Each pixel in residual BPU 210 can have a residual value generated by this subtraction between the corresponding pixel value of the original BPU and the predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.

[0070] After generating the residual BPU 210, the encoder 200 can feed the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized residual coefficients 216. To further compress the residual BPU 210, in the transform stage 212, the encoder 200 can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a two-dimensional set of "fundamental patterns," each fundamental pattern associated with a "transform coefficient." The fundamental patterns can have the same size (e.g., the size of the residual BPU 210). Each fundamental pattern can characterize a frequency-varying component of the residual BPU 210 (e.g., the frequency of brightness variation). No fundamental pattern can be reproduced by any combination of any other fundamental patterns (e.g., a linear combination). In other words, the decomposition decomposes the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the fundamental patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0071] Different transformation algorithms can use different base modes. Various transformation algorithms can be used in transformation stage 212, for example, discrete cosine transform, discrete sine transform, etc. The transformation in transformation stage 212 is reversible. That is, encoder 200 can recover the residual BPU 210 through the inverse operation of the transformation (called the "inverse transform"). For example, to recover a pixel of the residual BPU 210, the inverse transform can be to multiply the value of the corresponding pixel in the base mode by the corresponding correlation coefficient and sum the products to produce a weighted sum. For video coding standards, encoder 200 and the corresponding decoder (e.g., Figure 3 The decoder 300 can use the same transform algorithm (and thus the same base pattern). Therefore, the encoder 200 can record only the transform coefficients, and the decoder 300 can reconstruct the residual BPU 210 based on these transform coefficients without receiving the base pattern from the encoder 200. Compared to the residual BPU 210, the transform coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0072] Encoder 200 can further compress the transform coefficients in quantization stage 214. During the transform process, different fundamental modes can represent different frequencies of change (e.g., brightness change frequency). Because the human eye is generally better at identifying low-frequency changes, encoder 200 can ignore information about high-frequency changes without causing a significant degradation in decoding quality. For example, in quantization stage 214, encoder 200 can generate quantization residual coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and performing rounding on the quotient. After this operation, some transform coefficients of the high-frequency fundamental modes can be converted to zero, while the transform coefficients of the low-frequency fundamental modes can be converted to smaller integers. Encoder 200 can ignore the zero-valued quantization residual coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, wherein the quantization residual coefficients 216 can be reconstructed into the transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0073] Because encoder 200 ignores the remainder of such division in the rounding operation, quantization stage 214 may be lossy. Typically, quantization stage 214 may constitute the primary source of information loss in the encoding process. The greater the information loss, the fewer bits the quantization residual coefficients 216 may require. To obtain different levels of information loss, encoder 200 can use different values ​​of the quantization parameters, or any other parameters of the quantization process.

[0074] Encoder 200 can feed the prediction data 206 and quantization residual coefficients 216 to the binary encoding stage 226 to generate a video bitstream 228 to complete the forward path. In the binary encoding stage 226, encoder 200 can use binary encoding techniques to encode the prediction data 206 and quantization residual coefficients 216, such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding (CABAC), or any other lossless or lossy compression algorithm.

[0075] For example, the CABAC encoding process in binary encoding stage 226 may include a binarization step, a context modeling step, and a binary arithmetic encoding step. If the syntax element is not binary, encoder 200 first maps the syntax element to a binary sequence. Encoder 200 may choose a context encoding mode or a bypass encoding mode for encoding. In some embodiments, for the context encoding mode, a probability model for the binary symbol (bin) to be encoded is selected via the “context,” which refers to previously encoded syntax elements. The binary symbol and the selected context model are then passed to an arithmetic encoding engine, which encodes the binary symbol and updates the corresponding probability distribution of the context model. In some embodiments, for the bypass encoding mode, the probability model is not selected via the “context,” but the binary symbol is encoded with a fixed probability (e.g., a probability equal to 0.5). In some embodiments, the bypass encoding mode is selected for a specific binary symbol to accelerate the entropy encoding process, and the encoding efficiency loss is negligible.

[0076] In some embodiments, in addition to the prediction data 206 and the quantization residual coefficients 216, the encoder 200 may also encode other information in the binary encoding stage 226, such as the prediction mode selected in the prediction stage (e.g., intra-frame prediction stage 2042 or inter-frame prediction stage 2044), the parameters of the prediction operation (e.g., intra-frame prediction mode, motion parameters, etc.), the transform type of the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. That is, the encoded information can be sent to the binary encoding stage 226 to further reduce the bitrate before being packaged into the video bitstream 228. The encoder 200 can use the output data of the binary encoding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 can be further packaged for network transmission.

[0077] Components 218, 220, 222, 224, 232, and 234 may be referred to as "reconstruction paths." These reconstruction paths can be used to ensure the encoder 200 and its corresponding decoder (e.g., Figure 3 The decoder 300 in the middle uses the same reference data for prediction.

[0078] During the process, after quantization phase 214, encoder 200 can feed quantized residual coefficients 216 to inverse quantization phase 218 and inverse transform phase 220 to generate reconstructed residual BPU 222. In inverse quantization phase 218, encoder 200 can perform inverse quantization on quantized residual coefficients 216 to generate reconstructed transform coefficients. In inverse transform phase 220, encoder 200 can generate reconstructed residual BPU 222 based on the reconstructed transform coefficients. Encoder 200 can add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which will be used in prediction phases 2042, 2044 for the next iteration of the process.

[0079] In the reconstruction path, if an intra-prediction mode has been selected in the forward path, after generating prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), encoder 200 can directly feed prediction reference 224 to intra-prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU in the current image). If an inter-prediction mode has been selected in the forward path, after generating prediction reference 224 (e.g., the current image where all BPUs have been encoded and reconstructed), encoder 200 can feed prediction reference 224 to loop filtering stage 232, where encoder 200 can apply loop filtering to prediction reference 224 to reduce or eliminate distortion (e.g., blockiness) introduced by inter-prediction. Encoder 200 can apply various loop filtering techniques in loop filtering stage 232, such as deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. In SAO, after the deblocking filter, a nonlinear amplitude mapping is introduced within the inter-frame prediction loop to reconstruct the original signal amplitude using a lookup table, which is described by a small number of additional parameters determined at the encoding end through histogram analysis.

[0080] The loop-filtered reference image can be stored in buffer 234 (or “decoded image buffer”) for subsequent use (e.g., as an inter-frame prediction reference image for future images of video sequence 202). Encoder 200 can store one or more reference images in buffer 234 for use in inter-frame prediction stage 2044. In some embodiments, encoder 200 can encode the loop-filtered parameters (e.g., loop-filter strength), as well as the quantization residual coefficients 216, prediction data 206, and other information in binary encoding stage 226.

[0081] Encoder 200 can iteratively execute the above process to encode each raw BPU of the original image (in the forward path) and generate prediction reference 224 for encoding the next raw BPU of the original image (in the reconstruction path). After encoding all raw BPUs of the original image, encoder 200 can continue to encode the next image in video sequence 202.

[0082] It should be noted that other variations of the encoding process can be used to encode the video sequence 202. In some embodiments, the stages of the process can be performed by the encoder 200 in different orders. In some embodiments, one or more stages of the encoding process can be combined into a single stage. In some embodiments, a single stage of the encoding process can be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the encoding process may include additional stages ( Figure 2 (Not shown in the image). In some embodiments, the encoding process may be omitted. Figure 2 One or more stages in the process.

[0083] For example, in some embodiments, encoder 200 can operate in a transform skip mode. In transform skip mode, transform stage 212 is bypassed, and a transform skip flag is sent for TB signaling. This can improve compression performance for certain types of video content, such as computer-generated images or graphics mixed with camera-captured content (e.g., scrolling text). Furthermore, encoder 200 can also operate in a lossless mode. In lossless mode, transform stage 212, quantization stage 214, and other processes affecting the decoded image (e.g., SAO and deblocking filtering) are bypassed. The residual signal from the intra-frame prediction stage 2042 or inter-frame prediction stage 2044 is fed into binary encoding stage 226, using the same neighborhood context as the quantized transform coefficients. This allows for mathematically lossless reconstruction. Therefore, both transform and transform skip residual coefficients are encoded within non-overlapping CGs. That is, each CG may include one or more transform residual coefficients or one or more transform skip residual coefficients.

[0084] Figure 3 A block diagram of an exemplary decoder 300 for a video encoding system (e.g., AVS3 or H.26x series) according to some embodiments of the present disclosure is shown. Decoder 300 can perform operations related to... Figure 2 The compression process described above corresponds to the decompression process. The corresponding stages of the compression and decompression processes are described in... Figure 2 and Figure 3 The same numbers are used to mark them.

[0085] In some embodiments, the decompression process can be similar to Figure 2 The reconstruction path described in [the original text]. Decoder 300 can correspondingly decode video bitstream 228 into video stream 304. Video stream 304 can be very similar to [the original text]. Figure 2 The video sequence 202 in the video. However, due to the compression and decompression process (e.g., Figure 2 The information loss in the quantization stage 214) means that video stream 304 may differ from video sequence 202. Similar to... Figure 2 The encoder 200 and decoder 300 in the video bitstream 228 can perform the decoding process at the level of a basic processing unit (BPU) for each image encoded in the video bitstream 228. For example, the decoder 300 can perform the process iteratively, wherein the decoder 300 can decode the basic processing unit in one iteration. In some embodiments, the decoder 300 can perform the decoding process in parallel for a region (e.g., slices 114-118) of each image encoded in the video bitstream 228.

[0086] exist Figure 3 In the binary decoding stage 302, decoder 300 can feed a portion of the video bitstream 228 associated with the basic processing unit (referred to as the "encoded BPU") of the encoded image to binary decoding stage 302. In binary decoding stage 302, decoder 300 can decode the video bitstream into prediction data 206 and quantization residual coefficients 216. Decoder 300 can use the prediction data 206 and quantization residual coefficients to reconstruct the video stream 304 corresponding to the video bitstream 228.

[0087] In the binary decoding stage 302, the decoder 300 may perform the inverse operation of the binary encoding technique used by the encoder 200 (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantization residual coefficients 216, the decoder 300 may also decode other information in the binary decoding stage 302, such as prediction mode, parameters of the prediction operation, transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over the network in the form of data packets, the decoder 300 may unpack the video bitstream 228 before feeding it to the binary decoding stage 302.

[0088] Decoder 300 can feed quantized residual coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. Decoder 300 can feed prediction data 206 to intra-frame prediction stage 2042 and inter-frame prediction stage 2044 to generate prediction BPU 208. Specifically, for an encoded basic processing unit (referred to as the "current BPU") of an encoded image being decoded (referred to as the "current image"), the prediction data 206 decoded by decoder 300 from binary decoding stage 302 can include various types of data depending on the prediction mode used by encoder 200 to encode the current BPU. For example, if encoder 200 uses intra-frame prediction to encode the current BPU, the prediction data 206 can include encoded information, such as a prediction mode indicator (e.g., a flag value) indicating the intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, the sizes of the neighboring BPUs, extrapolation parameters, and the orientation of the neighboring BPUs relative to the original BPU. In another example, if the encoder 200 encodes the current BPU using inter-frame prediction, the prediction data 206 may include encoding information, such as prediction mode indicators (e.g., flag values) indicating the inter-frame prediction, and the parameters of the inter-frame prediction operation. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with each reference image, the positions (e.g., coordinates) of one or more matching regions in each reference image, and one or more MVs associated with each matching region.

[0089] Therefore, the prediction mode indicator can be used to select whether to invoke the inter-frame prediction module or the intra-frame prediction module. Then, parameters for the corresponding prediction operation can be sent to the corresponding prediction module to generate one or more prediction signals. Specifically, based on the prediction mode indicator, the decoder 300 can determine whether to perform intra-frame prediction in the intra-frame prediction phase 2042 or inter-frame prediction in the inter-frame prediction phase 2044. Figure 2 The details of performing such intra-frame or inter-frame prediction are described in the previous section and will not be repeated below. After performing such intra-frame or inter-frame prediction, the decoder 300 can generate a prediction BPU 208.

[0090] After generating the prediction BPU 208, the decoder 300 can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder 300 can feed the prediction reference 224 to the intra-frame prediction stage 2042 and the inter-frame prediction stage 2044 for performing prediction operations in the next iteration.

[0091] For example, if the current BPU is decoded using the intra-frame prediction in the intra-frame prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder 300 can directly feed the prediction reference 224 to the intra-frame prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU of the current image). If the current BPU is decoded using the inter-frame prediction in the inter-frame prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference image where all BPUs have been decoded), the decoder 300 can feed the prediction reference 224 to the loop filtering stage 232 to reduce or eliminate distortion (e.g., blockiness). Additionally, the prediction data 206 may further include loop filtering parameters (e.g., loop filtering strength). Therefore, the decoder 300 can, as... Figure 2 The method described herein applies loop filtering to prediction reference 224. For example, loop filtering such as deblocking, SAO, or ALF can be applied to form a loop-filtered reference image, which is stored in buffer 234 (e.g., a decoded picture buffer (DPB) in computer memory) for subsequent use (e.g., in the inter-frame prediction stage 2044, for prediction of a future encoded image of the video bitstream 228). In some embodiments, the reconstructed image from buffer 234 can also be sent to a display device, such as a TV, PC, smartphone, or tablet, for viewing by an end user.

[0092] Decoder 300 can iteratively execute the decoding process to decode each encoded BPU of the encoded image and generate a prediction reference 224 for encoding the next encoded BPU of the encoded image. After decoding all encoded BPUs of the encoded image, decoder 300 can output the image to video stream 304 for display and continue decoding the next encoded image in video bitstream 228.

[0093] Figure 4 This is a block diagram of an exemplary apparatus 400 for encoding or decoding video according to some embodiments of this disclosure. Figure 4As shown, device 400 may include processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit system capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), and any combination thereof. In some embodiments, processor 402 may also be a set of processors grouped into individual logic components. For example, such as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b and processor 402n.

[0094] The device 400 may also include a memory 404 configured to store data (e.g., an instruction set, computer code, intermediate data, etc.). For example, such as Figure 4 As shown, the stored data may include program instructions (e.g., instructions for implementing...). Figure 2 and Figure 3The processor 402 can access the program instructions (e.g., via bus 410) and the data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and the data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 may include any combination of any number of random-access memory (RAM), read-only memory (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped into a single logical component. Figure 4 (Not shown in the image).

[0095] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), etc.

[0096] For ease of explanation and to avoid ambiguity, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be wholly or partially integrated into any other component of device 400.

[0097] Device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transceivers, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near-field communication (NFC) adapters, cellular network chips, etc.

[0098] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4As shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), etc.

[0099] It should be noted that the video codec (e.g., the codec that performs the process of encoder 200 or decoder 300) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of the process of encoder 200 or decoder 300 can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of the process of encoder 200 or decoder 300 can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0100] exist Figure 2 and Figure 3 In the inter-frame prediction stage 2044, a reference index is used to indicate which previously encoded image the reference block comes from. The motion vector (MV), i.e., the positional difference between the reference block in the reference image and the current block in the current image, is used to indicate the position of the reference block in the reference image. For bidirectional prediction (e.g., Figure 1 Image 106 in the image uses two reference blocks to generate a combined prediction block. One reference block comes from a reference image in reference image list 0 (e.g., image 104), and the other reference block comes from a reference image in reference image list 1 (e.g., image 108). Therefore, bidirectional prediction requires two reference indices (e.g., list 0 reference index and list 1 reference index) and two motion vectors (e.g., list 0 motion vector and list 1 motion vector). The motion vectors are determined by the encoder and signaled to the decoder. In some embodiments, to save signaling overhead, motion vector difference (MVD) is instead signaled in the bitstream. For the decoder, a motion vector predictor (MVP) can be derived based on spatial and temporal neighboring block motion information, and the MV can be obtained by adding the MVD parsed from the bitstream to the MVP.

[0101] As discussed above, different modes can be used to implement the video encoding or decoding process. In some conventional inter-frame coding modes, encoder 200 can explicitly signal one or more MVs, a reference image index corresponding to each reference image list, a reference image list usage flag, or other information for each CU. On the other hand, when the CU is encoded in skip mode or direct mode, motion information including reference indices and motion vectors is not signaled to decoder 300 in video bitstream 228. Instead, decoder 300 can derive the motion information using the same rules as encoder 200. The skip mode and the direct mode share the same motion information derivation rules and therefore have the same motion information. The difference between the two modes is that in the skip mode, the signaling of the prediction residual is skipped by setting the residual to zero. In the direct mode, the prediction residual is still signaled in the bitstream.

[0102] For example, when a CU is encoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded MV differences, or a reference image index. In skip mode, signaling transmission of the residual data can be skipped by setting the residual to zero. In direct mode, the residual data is transmitted while deriving motion information and partitioning.

[0103] On the other hand, in inter-frame mode, when the motion vector difference and reference index are signaled to decoder 300, encoder 200 can select any allowed motion vector and reference index values. Compared to inter-frame mode where the motion information is signaled, bits dedicated to the motion information can be saved in skip mode or direct mode. However, encoder 200 and decoder 300 need to follow the same rules to derive the motion vector and reference index to perform inter-frame prediction 2044. In some embodiments, the derivation of the motion information may be based on the spatially or temporally adjacent blocks. Therefore, skip mode and direct mode are applicable when the motion information of the current block is close to the motion information of the spatially or temporally adjacent blocks of the current block.

[0104] For example, the skip mode or the direct mode allows the motion information to be inherited from spatial or temporal (co-located) neighboring blocks (e.g., reference index, MV, etc.). A candidate list of motion candidates can be generated from these neighboring blocks. In some embodiments, to derive the motion information for inter-frame prediction 2044 in skip mode or direct mode, encoder 200 may first derive the candidate list of motion candidates and select one of the motion candidates to perform inter-frame prediction 2044. When signaling is sent to video bitstream 228, encoder 200 can send the index of the selected candidate via signaling. At the decoding end, decoder 300 can obtain the index parsed from video bitstream 228, derive the same candidate list, and use the same motion candidate (including motion vectors and reference image indexes) to perform inter-frame prediction 2044.

[0105] In some embodiments, different skip and direct modes exist, including a regular skip and direct mode, a final motion vector representation mode, an angle-weighted prediction mode, an enhanced temporal motion vector prediction mode, and an affine motion compensation skip / direct mode. The motion candidate list may include multiple candidates obtained based on different methods. For example, for the regular skip and direct mode, the motion candidate list may have 12 candidates, including temporal motion vector predictor (TMVP) candidates (i.e., temporal candidates), one or more spatial motion vector predictor (SMVP) candidates (i.e., spatial candidates), one or more motion vector angular predictor (MVAP) candidates (i.e., sub-block-based spatial candidates), and one or more history-based motion vector predictor (HMVP) candidates (i.e., history-based candidates). In some embodiments, the encoder or decoder may first derive the TMVP and SMVP candidates and add them to the candidate list. After adding the TMVP and SMVP candidates, the encoder or decoder derives and adds the MVAP and HMVP candidates. In some embodiments, the number of MVAP candidates added to the candidate list can vary depending on the number of one or more directions available during the MVAP process. For example, the number of the one or more MVAP candidates can be between 0 and a maximum number (e.g., 5). After adding one or more MVAP candidates, one or more HMVP candidates can be added to the candidate list until the total number of candidates reaches a target number, and the maximum number can also be signaled in the bitstream.

[0106] In some embodiments, the first candidate is the TMVP derived from the MV of a co-located block in a predefined reference frame. The predefined reference frame is defined as the reference frame with reference index 0 in list 1 for B-frames, or as the reference frame with reference index 0 in list 0 for P-frames. When the MV of the co-located block is unavailable, the MV prediction value (MVP) derived from the MV of spatially adjacent blocks is used as the block-level TMVP.

[0107] Next, we describe affine motion compensation. In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied.

[0108] Figure 5A and Figure 5B These are two schematic diagrams illustrating control point-based affine models for blocks 500a and 500b according to some embodiments of the present disclosure. In some embodiments, Figure 5A Control points 510a and 520a in the middle and Figure 5B Control points 510b, 520b, and 530b are respectively located at the corners of blocks 500a and 500b. For example... Figure 5A and Figure 5B As shown, the affine motion field of a block is described using the motion information of two control points in a 4-parameter affine motion model or the motion vectors of three control points in a 6-parameter affine motion model. Figure 5A As shown, for a 4-parameter affine motion model, two control points, 510a and 520a, are required. Figure 5B As shown, a 6-parameter affine motion model requires three control points: 510b, 520b, and 530b. To reduce the computational complexity of the model and the bandwidth requirements of motion compensation, the affine motion compensation is performed at the sub-block level rather than the sample level. In some coding standards, 4×4 or 8×8 luma sub-block affine motion compensation is used, where each 4×4 or 8×8 sub-block has a motion vector to perform motion compensation.

[0109] To derive the motion vector for each 8×8 or 4×4 luminance sub-block, the motion vector at the center position of each sub-block is calculated based on two or three control points (CPs) and rounded to 1 / 16 pixel precision.

[0110] For a 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived using the following formula: (1) For a 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived using the following formula: (2) Among them, (mv 0x ,mv 0y (mv) is the motion vector of the top-left control point. 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point.

[0111] Figure 6 This is a schematic diagram illustrating the motion vectors of the center samples of each sub-block of block 600 according to some embodiments of the present disclosure. Specifically, Figure 6 An example of a six-parameter affine model is given, where the motion vector of each sub-block can be derived from the motion vectors CPMV0, CPMV1, and CPMV2 of three control points 610-630. After deriving the sub-block motion vectors, the motion compensation is performed to generate a predicted block of the sub-block using the derived motion vectors.

[0112] To simplify motion compensation prediction, block-based affine transformation prediction can be applied. To derive the motion vector for each 4×4 brightness sub-block, the motion vector of the center sample of each sub-block is calculated according to the formula described above, such as... Figure 6 As shown, it is rounded to 1 / 16 fractional precision.

[0113] Similar to translational motion inter-frame prediction, affine motion inter-frame prediction also has two modes: affine merging mode and affine AMVP mode. In some embodiments, affine merging prediction can be used.

[0114] Affine merging mode (AF_MERGE) can be applied to code blocks with a width and height greater than or equal to 8. In this mode, a control point motion vector (CPMV) for the current code block is generated based on the motion information of spatially adjacent code blocks. Up to 15 affine candidates can exist, and a signaling transmission index indicates which affine candidate will be used for the current code block.

[0115] In some embodiments, the following eight types of candidates are used to form the affine merge candidate list: (a) candidates inherited from neighboring blocks; (b) candidates inherited from non-neighboring blocks; (c) candidates constructed from neighboring blocks; (d) a second type of affine candidate constructed from non-neighboring blocks; (e) a first type of affine candidate constructed from non-neighboring blocks; (f) regression-based affine merge candidates; (g) pairwise affine; and (h) zero MV.

[0116] Figure 7 The inheritance of control point motion vectors according to some embodiments of this disclosure is illustrated. The inherited affine candidates are derived from the affine motion models of adjacent or non-adjacent blocks. When an adjacent or non-adjacent affine coding block is identified, its control point motion vectors are used to derive the CPMVP candidates in the current CU 710's affine merging list. Figure 7 As shown, if the adjacent lower left corner block A is encoded in an affine pattern, then the motion vectors of the upper left, upper right, and lower left corners of the CU 720 containing block A are derived. , and When encoding block A using a 4-parameter affine model, according to and Calculate the two CPMVs of the current CU 710. When encoding block A using a 6-parameter affine model, according to... , and Calculate the three CPMVs of the current CU 710.

[0117] Figure 8A and Figure 8B The illustration shows spatial neighbor blocks for deriving affine merging and affine advanced motion vector prediction (AMVP) candidates according to some embodiments of the present disclosure. Specifically, Figure 8A The spatial neighbor blocks used to derive inheritance candidates are shown. Figure 8B The spatial neighbor blocks used to derive the first type of build candidate are shown.

[0118] for Figure 8A Candidates for inheritance from non-adjacent blocks are determined based on the distance between the non-adjacent spatial neighbor blocks and the current block 810, i.e., the non-adjacent spatial neighbor blocks are checked in order from nearest neighbor to farthest neighbor. At a specific distance, only the first available neighbor block from each side (e.g., the left and top) of the current block 810 (encoded using an affine pattern) is selected for the inheritance candidate derivation. Figure 8A As indicated by the dashed arrows, the inspection order for the left and top adjacent blocks is from bottom to top and from right to left, respectively.

[0119] Affine candidates constructed from neighboring blocks are candidates built by fusing the translational motion information of neighboring blocks at each control point. Figure 9 The distribution of candidate positions for the constructed affine merging pattern for the current block 910 is shown according to some embodiments of this disclosure.

[0120] The motion information of the control point can be obtained from Figure 9 The specified spatial and temporal neighbor blocks T shown are derived in CPMV. k(k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check blocks B2, B3, and A2, and use the MV of the first available block. For CPMV2, check blocks B1 and B0, and for CPMV3, check blocks A1 and A0. If TMVP is available, it can be used as CPMV4.

[0121] After deriving the motion values ​​(MVs) of the four control points, affine merging candidates are constructed based on that motion information. These are constructed sequentially using combinations of the following control point MVs: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.

[0122] Combining three CPMVs constructs a 6-parameter affine merge candidate, while combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, if the reference indices of the control points are different, the associated control point MV combinations are discarded.

[0123] For the first type of construction candidate, such as Figure 8B As shown, firstly, the positions of a non-adjacent spatial block on the left and above are independently determined; then, the position of the adjacent block in the upper left corner can be determined accordingly, which can surround a rectangular virtual block together with the non-adjacent blocks on the left and above.

[0124] Figure 10 This illustrates a first type of affine merge / AMVP candidate based on some embodiments of this disclosure. (e.g.) Figure 10 As shown, motion information from three non-adjacent blocks can be used to construct the CPMV at the top left (A), top right (B), and bottom left (C) corners of the virtual block 1020. The virtual block is then projected onto the current coding block 1010 to generate a corresponding construction candidate.

[0125] For the second type of construction candidate, non-translation affine parameters are inherited from non-adjacent spatial neighboring blocks. Specifically, the second type of affine construction candidate is generated by a combination of the following two items: (1) translation affine parameters of adjacent 4×4 blocks; and (2) ... Figure 8A The non-translational affine parameters defined and inherited from the non-adjacent spatial neighbor blocks.

[0126] Figure 11A and Figure 11B This is a schematic diagram illustrating a first history parameter table (HPT) 1110 and a second history parameter table (HPT) 1120 according to some embodiments of the present disclosure. Figure 11AAs shown, in some embodiments, a first HPT 1110 can be established. The entries of the first HPT 1110 store a set of affine parameters (a0, b0, c0, and d0), and each affine parameter can be represented by a 16-bit signed integer. The entries in the first HPT 1110 can be categorized according to reference lists and reference indices. In some embodiments, five reference indices are supported for each reference list L0 or L1 in the first HPT 1110. The category (denoted as HPTCat) of the first HPT 1110 can be calculated based on the following formula: (3) RefList and RefIdx represent the list of reference images (e.g., L0 or L1) and the reference index, respectively.

[0127] For each category, a maximum of seven entries can be stored, thus the first HPT 1110 has a total of 70 entries. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding the affine-coded blocks using reference lists RefListcur and RefIdxcur, the affine parameters are used to update the category HPTCat(RefList) in a manner similar to HMVP table updates. cur RefIdx cur The entries in ) . The history-affine-parameter-based candidate (HAPC) is selected from one of the seven adjacent 4×4 blocks (e.g. Figure 11A The values ​​shown are represented as A0, A1, A2, B0, B1, B2, or B3), and derived from an affine parameter set stored in the corresponding entry in the first HPT 1110. The MV of adjacent 4×4 blocks is used as the base MV. Formulatically, the MV of the current block at position (x, y) can be calculated based on the following formula: (4) in,( , ) represents the MV of the adjacent 4×4 blocks, ( , The position (x, y) represents the center position of the adjacent 4×4 blocks. The position (x, y) can be the top left corner, top right corner, and bottom left corner of the current block to obtain the corner-position MV (CPMV) of the current block, or it can be the center of the current block to obtain the regular MV of the current block.

[0128] like Figure 11BAs shown, in some embodiments, a second HPT 1120 with base MV information may also be attached. The second HPT 1120 contains nine entries, and each entry includes a base MV ( ), a reference index and four affine parameters for each reference list, and a base position ( An additional merged HAPC can be generated based on the second HPT 1120, using the basic MV information and corresponding affine model stored in a certain entry of the second HPT. The difference between the first HPT 1110 and the second HPT 1120 is... Figure 11A and 11B As shown. Furthermore, paired affine merge candidates can be generated from two affine merge candidates, either historically derived or non-historically derived. A paired affine merge candidate can be generated by averaging the CPMV of existing affine merge candidates in the list.

[0129] Figure 12 The illustration shows some embodiments according to this disclosure, in which adjacent 4×4 sub-blocks 1220 are used to derive regression-based affine merging candidates. For example... Figure 12 As shown, the 4×4 sub-block 1210 in the current CU is surrounded by adjacent 4×4 sub-blocks 1220, and the adjacent 4×4 sub-blocks 1220 are formed as follows: Figure 12 The gray area shown represents... (e.g.) Figure 12 As shown, the sub-block motion fields from previously encoded affine blocks and the motion vectors of neighboring sub-blocks from the current block are used as inputs to the regression process. The predicted CPMV of the current block is derived as output. Regression-based affine merging candidates are derived and added to the affine merging list. The proposed affine candidates are derived by using the sub-block motion fields from previously encoded affine blocks and the motion information of neighboring sub-blocks from the current block as inputs to the regression process. The previously encoded affine blocks can be identified by scanning non-adjacent locations and the affine HMVP table.

[0130] The neighboring sub-block information of the current coding block is obtained from 4×4 sub-blocks 1220. For each sub-block, given a reference list, the motion vector and center coordinates corresponding to that sub-block can be used. For each affine coding block, a maximum of two affine candidates can be derived, one with neighboring sub-block information and the other without. All candidates generated by linear regression are pruned and merged into a candidate subgroup. When ARMC is enabled, the ARMC process based on TM cost is applied. Then, when N affine coding blocks are found, up to N linear regression-generated candidates are added to the affine merge list. For example, the number of affine candidates for ARMC is 30, and the output list size is 15.

[0131] In some embodiments, affine candidates derived from the temporally co-located image are added to the affine merge candidate list. The same sampling pattern from the regular inter-frame merge mode is reused to scan predefined locations in the co-located image, thereby deriving temporally affine candidates. If the scanned location belongs to an affine-coded block, its CPMV is scaled to the current coded block based on its position and block size in the co-located image to derive a new affine candidate. This new affine candidate is then inserted into the existing affine merge list and reordered along with other affine merge candidates via the ARMC process.

[0132] After inserting all the above candidates into the candidate list, if the list is still not full, zero MV will be inserted at the end of the list.

[0133] To reduce the implementation cost of non-adjacent affine candidates, the following constraints can be imposed. First, the non-adjacent blocks are restricted to the current CTU (i.e., the row buffer has no additional storage requirements). Second, the storage granularity for affine motion information (including CPMV and reference index) is reduced from 8×8 to 16×16 (i.e., only the affine motion from the top-left 8×8 block is stored). Furthermore, the stored CPMV is projected onto each 16×16 block before storage, eliminating the need to store position and size information. Third, only the CPMVs from the top-left and top-right corners are stored (i.e., the 4-parameter affine model used for NA-AFF is always used).

[0134] Next, the Affine Advanced Motion Vector Prediction (AMVP) mode is described. In some embodiments, the Affine Advanced Motion Vector Prediction (AMVP) mode can be used. As in the regular AMVP mode, in the Affine AMVP mode, an AMVP candidate list is constructed using various types of candidates. The encoder selects one of the candidates and signals the index of the selected candidate in the AMVP candidate list in the bitstream. An AMVP candidate contains two CPMV predictions for a 4-parameter affine model or three CPMV predictions for a 6-parameter affine model. In the AMVP mode, the CPMV of the current coded block is not directly inherited from the AMVP candidate, but is determined by motion estimation in the encoder. Furthermore, the difference between the CPMV determined by the motion estimation and the CPMV predictions of the selected AMVP candidate is signaled in the bitstream. Here, the difference signaled in the bitstream is called the Motion Vector Difference (MVD). For a 6-parameter affine model defined by three CPMVs, three MVDs are signaled in the bitstream. For a 4-parameter affine model defined by two CPMVs, two MVDs are signaled in the bitstream. At the decoding end, after decoding the index and MVD of the AMVP candidate, the AMVP candidate is determined according to the index. Then, the decoded MVD is added to the CPMV prediction value of the AMVP candidate indicated by the index to obtain the CPMV used in the motion compensation of the current block.

[0135] A maximum of two affine AMVP candidates may exist, and a signaling sending index is used to identify the candidate to be used by the current CU. Similar to the affine merging mode, the following types of candidates may also be used to construct the affine AMVP candidate list: (1) candidates inherited from neighboring blocks, (2) candidates constructed from neighboring blocks, (3) translation MV from neighboring blocks, (4) translation MV from temporally neighboring blocks, (5) candidates inherited from non-neighboring blocks, (6) second type of affine candidates constructed from non-neighboring blocks, (7) first type of affine candidates constructed from non-neighboring blocks, (8) regression-based affine merging candidates, (9) pairwise affine, and (10) zero MV.

[0136] In some embodiments, adaptive affine sub-block size and pixel-based affine motion compensation are considered. In the ECM, the sub-block size is adaptively determined. If the motion vector difference between two adjacent luma sub-blocks is less than a threshold, the luma sub-blocks are merged into a larger sub-block. If the motion vector difference of the larger sub-block is still less than the threshold, the merging of the larger sub-blocks continues until the motion vector difference between two adjacent sub-blocks is greater than the threshold, or until the sub-block is equal to a coding unit. For both luma and chroma components, the minimum affine sub-block size can be 1x1, and a 1x1 sub-block size allows for pixel-based affine MC. When the affine sub-block width or height is less than 4, prediction refinement with optical flow (PROF) is disabled.

[0137] After determining the sub-block size for affine motion compensation, a motion compensation interpolation filter is applied to generate a prediction for each sub-block using the derived motion vectors. The sub-block size for the chroma component depends on the size of the luma sub-block. The MV of the chroma sub-block is calculated by averaging the MVs of the top-left and bottom-right luma sub-blocks in the co-located luma region.

[0138] Next, optical flow-based prediction correction for affine modes is described. In some embodiments, optical flow-based prediction correction for affine modes can be used. Compared to pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but it leads to a decrease in prediction accuracy. To achieve finer motion compensation granularity, optical flow-based prediction correction (PROF) is used to correct sub-block-based affine motion compensation predictions without increasing the memory access bandwidth used for motion compensation. In VVC, after performing sub-block-based affine motion compensation, the brightness prediction samples are corrected by adding the difference derived from the optical flow formula. The PROF is described as the following four steps.

[0139] In step 1, sub-block-based affine motion compensation is performed to generate sub-block predictions. .

[0140] In step 2, at each sample location, a 3-tap filter [−1, 0, 1] is used to calculate the spatial gradient of the sub-block prediction. and The gradient calculation is exactly the same as the gradient calculation in BDOF, and is based on the following formula: (5) (6) in, This is used to control the precision of the gradient. For gradient calculation, the sub-block (i.e., 4×4) prediction is expanded by one sample on each side. To avoid additional storage bandwidth overhead and additional interpolation calculations, those expanded samples located on the expansion boundaries are copied from the nearest integer pixel position in the reference image.

[0141] In step 3, the brightness prediction correction is calculated using the following optical flow formula: (7) like Figure 13 As shown, where, At the sample location The sample MV calculated at (represented as) ) and the sample The difference between the MV values ​​of the sub-blocks belonging to the same sub-block. Figure 13 Subblock MV V is shown according to some embodiments of this disclosure. SB And pixel Δv(i,j). The... Quantization is performed using a 1 / 32 luminance sample precision unit.

[0142] Since the affine model parameters and the sample position relative to the center of the sub-block remain unchanged across the sub-blocks, calculations can be performed for the first sub-block. Then it is reused in other sub-blocks within the same CU. Assume... and Sample locations To the center of the sub-block The horizontal and vertical offsets, then It can be exported using the following formula: (8) (9) To maintain accuracy, the sub-block center Through The calculated values ​​are given, where WSB and HSB represent the width and height of the sub-block, respectively. For the 4-parameter affine model, the coefficients are calculated based on the following formula. C , D , E and F : (10) For a 6-parameter affine model, the coefficients are calculated based on the following formula. C , D , E and F : (11) in, , , These are the motion vectors of the control points at the top left, top right, and bottom left corners, respectively. and These are the width and height of the CU, respectively.

[0143] Finally, in step 4, the brightness prediction is corrected. Add to sub-block prediction The final prediction I' is generated using the following formula: (12) For affine-coded CUs, PROF is not applicable in two cases. First, PROF is not applicable when all control points MV are the same, indicating that the CU only has translational motion. Second, PROF is not applicable when the affine motion parameters are greater than a specified limit, because the sub-block-based affine MC is downgraded to a CU-based MC to avoid large storage access bandwidth requirements.

[0144] Next, a correction based on bi-directional optical flow (BDOF) is described. In some embodiments, a BDOF-based correction for affine subblocks can be used. The BDOF-based correction can also be applied to affine coded blocks with subblocks MC when the BDOF condition is met.

[0145] Affine-coded blocks (e.g., affine regular merge mode, affine BM merge mode, affine AMVP mode) derive the MV of each 4×4 sub-block from the affine model. The BDOF process begins by grouping 4×4 sub-blocks with the same MV together. The first iteration of BDOF MV correction is performed in an 8×8 sub-block grid, as specified in ECM-10.0. When the size of the grouped sub-blocks is less than 256, the second iteration of BDOF MV correction is performed in a 4×4 sub-block grid; otherwise, it is performed in an 8×8 sub-block grid. When the size of the grouped sub-blocks is 4×N or N×4, the first iteration of BDOF MV correction is bypassed.

[0146] Next, decoder-side motion vector refinement (DMVR) is described. In some embodiments, decoder-side motion vector refinement (DMVR) for affine merging of coded blocks can be used. In some embodiments, basic MV refinement can be used.

[0147] In the affine model, the motion vector at the sample position (x, y) can be formulated as: (13) in, It is the motion vector derived at the sample position (x, y). The fundamental MV in the model is called the motion vector at the sample position (0, 0), and a, b, c, d are the parameters of the affine model, which can be derived based on the motion vectors at two other sample positions in the plane. Typically, the fundamental MV in the model can be the motion vector at any sample position, not necessarily the one at (0, 0). If the motion vector at sample position (w, h) is chosen as the fundamental MV (denoted as...), then... Then the motion vector at the sample position (x, y) can be formulated as: (14) For a 4-parameter affine model, b equals -c, and d equals a. Therefore, the 4-parameter affine model can be formulated as: (15) In theory, all parameters of the affine model, including a, b, c, d, and... , All of these can be corrected within the DMVR. However, to limit complexity, some embodiments of this disclosure propose fixing the affine parameters a, b, c, and d, and correcting only the underlying MV. In other words, the template only undergoes translational movement during the search process. At each search position, all sub-templates have the same MV offset compared to the initial MV. Therefore, the three CPMVs and the sub-block MV also have the same MV offset after correction.

[0148] Similar to conventional DMVR, all current DMVR search methods can be applied. The difference lies in that motion compensation is performed at the sub-block level when calculating the SAD or SATD cost between the two predictions of L0 and L1. Therefore, an affine model can be used to derive the MV for each sub-block, and motion compensation can be performed at the sub-block level to obtain the predictions for the entire current affine-coded block. Furthermore, the SAD and SATD of the two predictions for the current affine-coded block (e.g., one from the L0 reference image and the other from the L1 reference image) are calculated, from which the cost of the current search point is derived.

[0149] In some embodiments, the control point motion vector (CPMV) is corrected. That is, an initial set of CPMVs corresponds to an initial position, and then MV offsets are added to all CPMVs to obtain the surrounding search points according to the following formula: (16) (17) (18) (19) (20) (twenty one) Wherein, CPMVx_l0 is the x-th l0 CPMV, and CPMVx_l1 is the x-th l1 CPMV. MV_offset is the motion vector offset of the search point, which is the difference between the initial CPMV and the corrected CPMV. After the CPMV is corrected, the affine model can be applied to calculate the MV of each sub-block, and then sub-block-level motion compensation can be applied to obtain the predicted value of the current block.

[0150] Figure 14 An example 3×3 square search pattern 1400 is shown according to some embodiments of the present disclosure. Figure 15 An example 3×3 cross-search pattern 1500 according to some embodiments of the present disclosure is shown. In some embodiments, conventional search schemes may be applied. For example, such as Figure 14 As shown, a 3×3 square search scheme can be applied to obtain the optimal integer MV offset. Then, a fractional search and fractional error surface estimation method can be applied to derive the optimal MV offset. Figure 14 As shown, point P0 is the location indicated by the initial MV. Therefore, the search first searches for points P1-P8 around the initial location, calculating the cost of each location. If point P7 has the minimum cost, it is set as the search center, and points P9-P11 are searched. If the cost of point P10 is less than the cost of point P7, the search center shifts to point P10, and points P12-P14 are searched. If point P12 has the minimum cost among points P6-P14, it is set as the new search center. If the costs of points P10, P11, P13P, and P15-P19 around point P12 are all greater than the cost of point P12, then point P12 is the optimal location, and the search process stops.

[0151] For another example, such as Figure 15 As shown, a cross-search scheme is used to reduce the number of searches. For example... Figure 15 As shown, point P0 is the location pointed to by the initial MV and is set as the first search center. Therefore, points P1 to P4 around the initial location are searched first, and the cost of each location is calculated.

[0152] If point P4 has the minimum cost, then point P4 is set as the second search center, and points P5-P7 are searched. Next, if the cost of point P7 is less than the costs of points P4, P5, and P6, then point P7 is set as the third search center, and points P8 to P10 are searched. Next, when point P9 has the minimum cost among points P7 to P10, point P9 is set as the fourth search center. Finally, if it is found that the costs of points P6, P16, and P18 surrounding point P9 are all greater than the cost of point P9, then point P9 is the optimal position, and the search process stops.

[0153] In some other examples, 3×3 square search and 3×3 cross search can be combined. A square search scheme can be used in the first k rounds, followed by a cross search to determine the optimal point; alternatively, a cross search scheme can be used first, followed by a square search to determine the optimal point.

[0154] In some other examples, an adaptive search step can be used to speed up the search process. That is, in the first k rounds of the search, the search step size is set to a large number (e.g., 2), and later, the search step size can be changed to a smaller number (e.g., 1).

[0155] Figure 16 This is a schematic diagram illustrating the adaptive search steps in an example 3×3 cross-search pattern 1600 according to some embodiments of the present disclosure. Figure 16 As shown, point P0 is the location indicated by the initial MV and is set as the first search center. First, points P1 to P8 surrounding the initial location are searched with a search step size of 2, so the distance between points P1-P8 and point P0 is 2 pixels, using Manhattan distance instead of Euclidean distance. Then, in the second round of search, since point P7 has the lowest cost, the search center is point P7, and the search step size remains 2. Therefore, search points P9, P10, and P11 are all 2 pixels away from the search center point P7. When point P10 has the lowest cost, in the third round of search, the search center is point P10, and the search step size is set to 1. Therefore, search points P12 to P14 are 1 pixel away from point P10. Assuming point P13 has the minimum cost, then in the fourth round of search, the search center is P13, the search step size is 1, and points P11, P15, P16, P17, and P18 are 1 pixel away from the search center point P13. In various embodiments, the adaptive search step can also be applied to cross-search mode or other search modes.

[0156] For each search point, the sub-block MV is recalculated using the current search point's CPMV, and sub-block-level motion compensation is performed using the recalculated sub-block MV to derive the predicted value of the block. Furthermore, the current method can also be applied to calculate the cost for each search point. For example, the cost can be calculated based on the following formula: (twenty two) Wherein, sadCost is the SAD between the l0 prediction value and the l1 prediction value of the current block, and mvDistanceCost is based on the distance between the search point and the initial point (i.e., the difference between the corrected CPMV and the initial CPMV).

[0157] To control the complexity of the correction, PROF may or may not be applied before calculating SAD during the search process. If PROF is applied, then for each search point, PROF is applied after obtaining the predicted value to correct the predicted value, and SAD between the two PROF-corrected predicted values ​​is calculated. If PROF is not applied, then for each search point, SAD can be calculated directly after obtaining the l0 and l1 predicted values.

[0158] Furthermore, to further reduce the complexity of the search process, a 2-tap bilinear interpolation filter can be used instead of an 8-tap or 12-tap interpolation filter to generate the predicted value for each search point.

[0159] In some other embodiments, the correction is applied to the sub-block MV. That is, each sub-block MV is derived using the initial CPMV, and then the MV offset is added to the sub-block MV to correct the sub-block MV according to the following formula: (twenty three) (twenty four) in, and Let l0 and l1 be the MVs of sub-block X, respectively. The MV_offset of the current search point is added to all sub-block MVs to obtain the corrected sub-block MVs. Then, for each sub-block MV, motion compensation is performed at the sub-block level to obtain two predictions for the entire CU. The SAD or SATD between the two predictions is calculated, and the cost of the current search is obtained. The search point with the minimum cost is considered the optimal point, and the corresponding MV_offset is obtained as the optimal MV_offset. The search method and cost calculation method used in the CPMV correction method can also be applied to the sub-block MV correction method.

[0160] Figure 17This is a schematic diagram illustrating an example sub-block-level pre-interpolation 1700 according to some embodiments of the present disclosure. In some embodiments, to reduce search complexity, samples of predicted values ​​can first be interpolated at each search point within the search window and stored in a buffer. Then, for each search point, the predicted value for each sub-block can be directly retrieved from the buffer without interpolation. Figure 17 As shown in the embodiment, a coding block 1710 is divided into 16 4×4 sub-blocks SB0-SB15 for affine motion compensation. The initial CPMV is used to calculate the initial sub-block MV for each sub-block. Then, for each sub-block, a reference sub-block (i.e., the predictor of the sub-block) can be located in the reference image 1720 using the initial sub-block MV. Since the correction is applied to the sub-block MV, all reference sub-blocks are offset by the same MV offset for each search point. Therefore, for each sub-block, pre-interpolation can be performed on the samples within each search window, such as... Figure 17 The gray area is indicated by [the map]. Then, for a search point, the sample of the reference block at that search point can be directly obtained from the search window without interpolation. Therefore, a significant amount of interpolation can be saved.

[0161] Pre-interpolation can be applied to both integer and fractional search processes. For the integer search process, pre-interpolated samples can be spaced one pixel apart. In the fractional search process, depending on the phase, the pre-interpolated samples can be stored individually. Figure 18 This is a schematic diagram illustrating example samples on integer and fractional search points according to some embodiments of the present disclosure. Figure 18 As shown, square marker 1810 represents samples of integer search points that can be pre-interpolated; cross marker 1820 represents samples at a 1 / 2 pixel position in the horizontal direction but an integer position in the vertical direction; triangle marker 1830 represents samples at a 1 / 2 pixel position in the vertical direction but an integer position in the horizontal direction; and circle markers represent samples at a 1 / 2 pixel position in both the horizontal and vertical directions. Therefore, if only the 1 / 2 pixel search process is considered, there may be three different phases of samples, and the distance between fractional samples of the same phase is one pixel. Therefore, fractional samples of the same type can be pre-interpolated together, while samples of different phases can be stored separately.

[0162] Following the CU-level base MV correction, a sub-CU-level base MV correction can also be applied. For example, in affine motion compensation, a CU can be divided into multiple sub-CUs, where each sub-CU is larger than a sub-block (e.g., a 16×16 sub-CU). Then, for each sub-CU, the base MV is further corrected. That is, sub-blocks within a sub-CU share the same base MV offset. Therefore, the final MV for each sub-block can be formulated as: (25) (26) in, and These are the l0 and l1 MVs of sub-block X, respectively. MV_offset is the offset obtained during the CU-level basic MV correction process, while MV_offset(sbCUx) is the MV offset obtained for the sub-CU containing sub-block X during the sub-CU-level basic MV correction. and This is the final corrected MV for sub-block X. Then, motion compensation is performed at the sub-block level using the corrected sub-block MV.

[0163] To achieve a good balance between computational complexity and performance, the search range can be set based on the PU size or the quantization parameter (QP). For example, a larger CU may require larger corrections and therefore a larger search range, while a smaller CU may have a smaller search range. In another example, a larger quantization parameter may lead to more distortion and thus may require a larger search range. Therefore, a larger search range can be used for a larger quantization parameter, and a smaller search range for a smaller quantization parameter. However, to reduce encoding time, setting a small search range for a larger quantization parameter can significantly reduce the number of search points. Therefore, in some cases, a smaller search range can be chosen for a larger quantization parameter, and a larger search range for a smaller quantization parameter.

[0164] In some embodiments, restrictions can be imposed on the size of the PU to further reduce complexity. That is, for certain CU sizes, the affine DMVR process can be skipped. For example, affine DMVR is not applied to CUs smaller than 8×8 or 16×16, or to CUs larger than 64×64 or 128×128.

[0165] To reduce complexity, the DMVR process for affine blocks can enable early termination. For example, in the affine merging candidate list, if a search point has already been checked during the search process of the initial motion vectors of previous candidates, that search point can be skipped. In another example, if the SAD at a search point is less than a predefined threshold, the search process can be terminated, and that search point can be used as the optimal position after correction.

[0166] In some embodiments, to reduce the complexity of the encoder and decoder, a fast algorithm can be used. For example, if the difference between the current affine merge candidate (e.g., the current initial MV) and a previously checked affine merge candidate (e.g., the previous initial MV) is less than a threshold, the DMVR process can be skipped for the current affine merge candidate. That is, the current affine merge candidate can be directly used for motion compensation of the current block without correction. In some embodiments, the threshold can depend on the current block size, and a first threshold for larger blocks can be greater than a second threshold for smaller blocks. For another example, DMVR is disabled for affine coded blocks with a size smaller than (or greater than) a certain threshold. That is, the small block (or the large block) does not use DMVR to correct the motion. In various applications, the threshold can be fixed or signaled in the bitstream.

[0167] In some embodiments, when the DMVR procedure is applied to an affine block, a bilateral matching cost for each sub-block is calculated. Then, using the sub-block bilateral matching costs and the corrected sub-block MV, the optimal corrected CPMV for the entire affine block is determined. More specifically, the CPMV is corrected according to the following steps.

[0168] In the first step, integer-pixel bilateral matching is performed on the sub-blocks. The bilateral matching costs of the sub-blocks are accumulated to determine the optimal integer-pixel MV offset. Then, in the second step, a half-pixel bilateral matching search is performed, using the optimal integer MV offset as the initial offset to output the optimal MV offset that minimizes the bilateral matching cost for the same sub-block set in the first step. Then, in the third step, linear regression is performed, using the corrected sub-block MV from the first step as input to output a set of control point motion vectors. Then, in the fourth step, the bilateral matching costs output by the second and third steps are compared to select the one with the minimum cost.

[0169] Furthermore, a CPMV search process can be added to the affine DMVR. For each control point, bilateral matching can be performed independently on the block centered on the control point to derive a revised CPMV. The revised CPMV is then used to derive an optimized set of CPMVs that minimizes the bilateral matching cost of the current block.

[0170] In some embodiments, affine non-translation parameter corrections can be used. In addition to the base MV, the parameters of the affine model can also be corrected. In some embodiments, as in the affine model represented in the above formula, offset values ​​offset_a, offset_b, offset_c, and offset_d are added to parameters a, b, c, and d to correct these parameters of the affine model. Both the encoder and the decoder can search for offset values ​​offset_a, offset_b, offset_c, and offset_d to reduce the SAD or SATD between the L0 and L1 prediction values ​​of the affine coded block. For example, the parameters can be corrected according to the following formula: (27) (28) (29) (30) (31) (32) (33) (34) in, a_l0, b_l0, c_l0 and d_l0 These are the parameters of the affine model referenced in list 0, and a_l1、b_l1、c_l1 and d_l1 These are the parameters of the affine model referenced in Listing 1. All four non-translation parameters have been corrected. In embodiments of this disclosure, these corrections may be referred to as 4-parameter corrections.

[0171] After the parameter is corrected, based on having the corrected parameter and The affine model formula is used to derive the sub-block MVs of reference lists 0 and 1, and then the sub-block MVs are derived, with motion compensation performed at the sub-block level. Bilateral matching costs (e.g., SAD or SATD) can be calculated at the sub-block level or the CU level. If the bilateral matching cost is calculated at the sub-block level, the cost of each sub-block is calculated after motion compensation, and then the CU-level cost is calculated by summing the costs of all sub-blocks after obtaining the costs of the sub-blocks. If the bilateral matching cost is calculated at the CU level, motion compensation of all sub-blocks is first performed to obtain the L1 and L0 predictions for the entire CU, and then the cost of the entire CU can be calculated. To reduce computational complexity, in some embodiments, only a subset of sub-blocks or a subset of samples in the CU is considered in the calculation of the bilateral matching cost. That is, only the differences between a subset of sub-blocks or a subset of samples are calculated, thus allowing motion compensation for sub-blocks or samples not considered in the cost calculation to be skipped.

[0172] For a 4-parameter affine model, since parameter b equals -c and parameter d equals parameter a, the correction can also follow this constraint. That is, the offset value offset_b equals -offset_c, and the offset value offset_d equals the offset value offset_a. Therefore, the encoder and the decoder only need to search for offset values ​​offset_a and offset_b, and then derive offset values ​​offset_b and offset_d based on offset values ​​offset_a and offset_b. In embodiments of this disclosure, the above correction can be referred to as a 2-parameter correction.

[0173] Since searching for two parameters is simpler than searching for four parameters, the following constraints can also be applied in the DMVR process of a six-parameter affine model: offset_b equals -offset_c and offset_d equals offset_a. That is, the correction follows the formula: (35) (36) (37) (38) (39) (40) (41) (42) On the other hand, the 4-parameter correction may be superior to the 2-parameter correction and can also be applied to the 4-parameter model. That is, the same correction method described in Equations 27 to 34 can be applied to both the 4-parameter affine model and the 6-parameter affine model.

[0174] The MV search method described above can be applied to parameter searches. For example, for 2-parameter correction, such as... Figure 14 or Figure 15 As shown, a 3×3 square search scheme or a 3×3 cross search scheme can be applied to obtain the desired parameter offset. For 4-parameter correction, the search is performed in 4-dimensional space. A 3×3×3×3 square search scheme or a 3×3×3×3 cross search scheme can be applied to obtain the desired parameter offset. For the 3×3×3×3 square search scheme, each center position has 80 neighboring positions to be searched, while for the 3×3×3×3 cross search scheme, each center position has 8 neighboring positions to be searched. The number of positions to be searched in the cross search is much smaller than the number of positions to be searched in the 3×3×3×3 square search. Assume the parameter offset of the current center position is ( offset_a, offset_b, offset_c, offset_d ), then in the 3×3×3×3 cross-search scheme, the eight adjacent positions to be searched are ( offset_a + step_a, offset_b, offset_ c, offset_d) 、( offset_a – step_a, offset_b, offset_c, offset_d ), ( offset_a, offset_b + step_b, offset_c, offset_d ), ( offset_a, offset_b – step_b offset_c, offset_d ), ( offset_a, offset_b, offset_c + step_c, offset_d ), ( offset_a, offset_b, offset_c – step_c, offset_d ), ( offset_a, offset_b, offset_c, offset_ d + step_d ), ( offset_a, offset_b, offset_c, offset_d – step_d ),in, step_a , step_b , step_c and step_d Parameters a , b , c and d The search step size.

[0175] After obtaining the desired parameter offset, an offset estimation based on the error surface can be applied to further refine the parameters with higher accuracy. The search step size can be set to a fixed value. For example, since the MV accuracy in the ECM is 1 / 16, and the basic sub-block used for affine motion compensation is 4×4, the search step size can be 1 / 64, making the MV difference between two adjacent sub-blocks 1 / 16 (i.e., 1 / 64 × 4), which is the minimum MV difference. The search step size can be greater than 1 / 64. A larger search step size reduces the number of search rounds, thus saving search time, but at the cost of correction accuracy. In some other examples, the search step size depends on the CU size. For a CU with width w and height h, the parameters... a and c The search step size (denoted as step_ac) and parameters d and b The search step size (denoted as step_db) can satisfy the following formula: (43) Here, T1 and T2 are two thresholds, which can be 1 / 16, 1 / 8, 1 / 4, or other values. The thresholds define the MV difference of the samples in the current encoding, the samples that are furthest from the sample with the base MV, which can be generated in each search step. In some embodiments, different parameters can have different search step sizes.

[0176] The cost for each search point can also be considered the difference in parameter offsets. That is, the cost, expressed as bilCost, can be a weighted sum of the SAD or SATD between two predictions of the coded block and the parameter offset, as shown in the following formula: (44) Where w is the weight, sadCost is the SAD / SATD or mean-reduced SAD / SATD cost of the predicted value, and ParameterOffsetCost is the cost depending on the parameter offset of the corrected parameter. When the weight w equals 0, only sadCost is considered.

[0177] When the affine parameters are determined, the base MV can be fixed. Theoretically, the MV of any point in the plane can be fixed as the base MV. In some embodiments, the CPMV is fixed as the base MV. Figure 19A The illustration shows how affine parameters are corrected by fixing the motion vector (CPMV) of the upper left control point as the base MV according to some embodiments of the present disclosure. Figure 19B The illustration shows how affine parameters are modified by fixing the upper right corner CPMV as the base MV according to some embodiments of the present disclosure. Figure 19C The illustration shows how affine parameters are modified by fixing the lower left corner CPMV as the base MV according to some embodiments of the present disclosure.

[0178] For example, such as Figure 19A As shown, the top-left CPMV 1910 is fixed, and the affine parameters are corrected. As the parameters change, the coded block 1900 rotates and scales up / down, thus the top-right CPMV 1920 and the bottom-left CPMV 1930 also change. Then, the sub-block MV is derived using the corrected parameters and the new CPMV, and motion compensation is performed. Figure 19B and Figure 19C Examples are provided, respectively, of fixing the top-right CPMV 1920 and the bottom-left CPMV 1930 to the base MV and correcting the four affine parameters. Similar to... Figure 19A As the affine parameters are corrected, the encoding block 1900 rotates and scales up / down, thus changing the non-fixed CPMV. In some embodiments, different CPMVs are sequentially fixed to the base MV. That is, the search process is divided into several steps. In the first step, as... Figure 19A As shown, the top-left corner CPMV 1910 is fixed as the base MV, and the parameters are searched. Using the obtained optimal parameters, the top-right corner CPMV 1920 can be calculated. Then, in the second step, as... Figure 19B As shown, the corrected upper right corner CPMV 1920 is fixed, and the parameters are corrected again. Using the optimal parameters obtained in the second step, the lower left corner CPMV 1930 can be calculated. Then in the third step, as... Figure 19C As shown, the corrected lower left corner CPMV 1930 is fixed, and the parameters are corrected again. These steps can be repeated several times. That is, the third step can be followed by the first step, where the new upper left corner CPMV 1910 is fixed as the base MV. And the process can continue until specific conditions are met. For example, the conditions can be, but are not limited to: (1) a preset number of iterations; (2) the SAD or SATD between the l0 predicted value and the l1 predicted value is less than a threshold; (3) the current fixed CPMV is the same as or similar to the fixed CPMV in the previous iteration; (4) the parameter offset in this round of search is less than a threshold.

[0179] In some embodiments, the CPMV is not used as the underlying MV, but a zero MV can be found in the plane and used as the underlying MV. First, the following formula can be solved to find the point at (x,y) where the MV is zero.

[0180] (45) Assuming the solution is (x1, y1), the affine model can be represented by zero-foundation MV as follows: (46) Then, search the affine parameters. a , b , c and d To find the corrected value. All of the above correction methods can be applied to the embodiments of this disclosure.

[0181] As described above, the affine parameter correction process can be similar to the basic MV correction process. The search process can be performed in multiple rounds, and for each round, if the bilateral matching cost of the center position is less than that of all adjacent positions, the current center position is taken as the optimal position, and the search process terminates. Otherwise, the adjacent position with the minimum bilateral matching cost is set as the new center position, and the search proceeds to the next round. To control search complexity, a maximum number of search rounds is set at both the encoding and decoding ends. Therefore, the search process terminates in two cases: the center position has the minimum cost, or the number of search rounds reaches a preset maximum. A larger maximum number of search rounds results in higher encoding performance, but also longer encoding and decoding times. To achieve the desired trade-off between complexity and performance, the maximum number of search rounds may depend on factors such as a fixed basic MV, QP, time layer, and CU size. For example, as... Figures 19A-19C The search process shown in the diagram involves the following steps: In the first step, the top-left CPMV 1910 is fixed as the base MV. Since this is the first time the parameters are adjusted, the maximum number of search rounds is set to a larger value, which better utilizes coding performance. Then, in the second step, the top-right CPMV 1920 is fixed as the base MV, and the maximum number of search rounds can be set to a smaller value because the parameters were adjusted in the first step, and a smaller number of search rounds reduces coding time. Then, in the third step, the bottom-left CPMV 1930 is fixed as the base MV, and the maximum number of search rounds can be set to a smaller value to further reduce coding time. Therefore, in these embodiments, the maximum number of search rounds is initially set to a larger value and then changed to a smaller value in subsequent steps.

[0182] In some embodiments, the maximum number of search rounds in subsequent steps depends on the actual number of search rounds in the previous step. For example, in the first step, when the top-left corner CPMV is set to the base MV, the maximum number of search rounds is set to N. In the first step, the search process terminates in the kth search round because the center position has the minimum bilateral matching cost before the search process reaches the Nth round. Then, in the second step, the maximum number of search rounds can be set to k / 2 (or another value that depends on k and is less than the original maximum number of search rounds P in the second round). If the search process reaches the maximum number of search rounds in the search rounds of the first step, then in the second step, the maximum number of search rounds is set to P, which is a value less than N. A similar method can be applied to the third step. If the actual number of search rounds in the second step reaches the maximum, then the maximum number of search rounds in the third step is set to L, where L is less than P. If the actual number of search rounds is t before the search process reaches the Pth round, then the maximum number of search rounds in the third step is set to t / 2. Therefore, the maximum number of search rounds can be adaptively determined based on the previous search process.

[0183] In some embodiments, to reduce complexity, the number of neighboring positions searched in a particular search round can be adaptively reduced based on the previous search round. For example, in a 3×3×3×3 cross-search scheme, each search round checks eight neighboring positions. Assume the current center is ( a, b, c, d ), then the eight adjacent positions to be checked are respectively and The bilateral matching costs for the eight adjacent positions are denoted as cost_pa0, cost_pa1, cost_pb0, cost_pb1, cost_pc0, cost_pc1, cost_pd0, and cost_pd1, respectively. In some embodiments, comparisons can be made with parameters. a The two associated costs are cost_pa0 and cost_pa1, and if cost_pa0 is less than cost_pa1, then in the next round, the parameters are considered. a Only positive offsets are considered. If cost_pa0 is greater than cost_pa1, then in the next round, the parameters are considered. a Only negative offsets are considered. Similarly, comparisons can be made with parameters. b The two associated costs are cost_pb0 and cost_pb1, and if cost_pb0 is less than cost_pb1, then in the next round, the parameters are considered. b Only positive offsets are considered. If cost_pb0 is greater than cost_pb1, then in the next round, the parameters will be considered. b Only negative offsets are considered. Similarly, comparisons can be made with parameters.c The two associated costs are cost_pc0 and cost_pc1, and if cost_pc0 is less than cost_pc1, then in the next round, the parameters are considered. c Only positive offsets are considered. If cost_pc0 is greater than cost_pc1, then in the next round, the parameters will be considered. c Only negative offsets are considered. Similarly, comparisons can be made with parameters. d The two associated costs are cost_pd0 and cost_pd1, and if cost_pd0 is less than cost_pd1, then in the next round, the parameters are considered. d Only positive offsets are considered. If cost_pd0 is greater than cost_pd1, then in the next round, the parameters will be considered. d Considering only negative offsets. Assuming that for the current search round, cost_pa0 is less than cost_pa1, cost_pb0 is greater than cost_pb1, cost_pc0 is less than cost_pc1, and cost_pd0 is greater than cost_pd1, then the four adjacent positions to be checked in the next round are... and ,in, It is the central location for the next round of searching.

[0184] In some embodiments, the minimum bilateral matching cost of the current search round is compared with the minimum bilateral matching cost of the previous search round. If the reduction in minimum cost is small, the search process terminates. For example, if the cost of the previous search round is A, which means the cost of the current search center is also A, then the minimum cost of the adjacent position is B at position posb, where B is less than A. According to the search rules, the search proceeds to the next round, with its search center at posb. However, in some embodiments, if AB is less than K or B is greater than A×f, the search process terminates, and posb is selected as the optimal position for this search step, where K and f are preset thresholds. For example, f can be a factor less than 1, such as 0.95, 0.9, or 0.8.

[0185] The quantization parameter (QP) controls quantization in video coding. A higher QP requires a larger quantization step size, thus introducing more distortion. Higher QPs require more rounds of search during correction, increasing coding time. To reduce overall coding time, in these embodiments, it is recommended to apply a smaller maximum number of search rounds at higher QPs than at lower QPs. Other methods for reducing complexity can also be used at high QPs. For example, the number of adjacent positions to be searched can be reduced based on previous search processes, the number of search rounds can be adaptively reduced, or the search process can be terminated early. Therefore, in these embodiments, different search strategies can be employed at different QPs.

[0186] In some embodiments, inter-frame coded frames, such as B-frames and P-frames, may have one or more reference frames. The temporal distance between the current frame and the reference frames affects the accuracy of inter-frame prediction. In video coding, this temporal distance between two frames is typically represented by the picture order count (POC) distance. Generally, the longer the POC distance, the lower the accuracy of inter-frame prediction and the lower the accuracy of motion information, thus requiring more correction. In some embodiments, the search process depends on the POC distance between the current frame and the reference frame. For hierarchical B-frames, frames with higher temporal layers have shorter POC distances to the reference frames, while frames with lower temporal layers have longer POC distances. Therefore, the search process can also depend on the temporal layer of the current frame. For example, affine parameter correction for higher temporal layers can be disabled because the POC distance to the reference frames is shorter and may not require correction. In another example, for higher temporal layer frames, a smaller number of search rounds or fewer adjacent search positions can be set. Furthermore, other methods for reducing the complexity of parameter correction can also be used for higher temporal layer frames. In these embodiments, the parameter correction process depends on the time distance or the POC distance between the current frame and the reference frame.

[0187] In some embodiments, template matching (TM)-based corrections can be used for affine coding blocks. In some embodiments, basic MV corrections can be used. In some embodiments, the template matching-based corrections are applied to affine merging patterns to improve the accuracy of affine motion inherited from previous coding blocks.

[0188] To apply the template-matching-based correction, a TM affine merge list is first derived. In one example, the TM affine merge list is the same as the regular affine merge list. That is, the same candidates are used for both the regular and TM affine merge modes. For the regular affine merge mode, one of the candidates is selected and indicated in the bitstream. The motion of the selected candidate is used for motion compensation. For the TM affine merge mode, the motion of the candidate can be corrected by TM, and the corrected motion of the selected candidate is used for motion compensation.

[0189] As a TM merging mode, the motion is corrected by TM. Therefore, in another example, a different affine merging candidate list is constructed by considering the TM effect. In this TM merging candidate list, the similarity of the candidates is checked. If a candidate to be inserted in the list is similar to an existing candidate in the list, the candidate is not inserted because similar candidates may produce the same motion after TM correction. To check the similarity of two affine merging candidates, the difference in CPMV of the affine candidates is calculated and compared to a threshold. Assume the first affine merging candidate has three CPMVs, respectively... Furthermore, the second merge candidate has three CPMVs, namely... The first affine candidate and the second affine candidate are similar to each other if the following conditions are met: (47) in, TH0_x, TH0_y, TH1_x, TH1_y, TH2_x, TH2_y This can depend on a threshold of the coded block size. All types of affine candidates include candidates inherited from neighboring and non-neighboring blocks, candidates constructed from neighboring blocks, first-type candidates constructed from non-neighboring blocks, second-type candidates constructed from non-neighboring blocks, regression-based candidates, and pairwise affine candidates.

[0190] Since affine motion compensation is performed at the sub-block level, the template for affine merging candidates also includes several sub-templates. Figure 20 The diagram illustrates an upper template 2010 and a left template 2020 for affine motion compensation according to some embodiments of the present disclosure. (See diagram for details.) Figure 20 As shown, if the template size used in TM is equal to Ts, and the size of the affine merging candidate sub-blocks is equal to Ws × Hs, then the upper template 2010 includes several sub-templates 2012, 2014, 2016, and 2018 with a size of Ws × Ts, and the left template 2020 includes several sub-templates 2022, 2024, 2026, and 2028 with a size of Ts × Hs. Figure 20 As shown, the white area is the coding block of the current affine motion coding, which includes 16 sub-blocks 2030. The dotted area is the template, which includes four upper sub-templates 2012, 2014, 2016, and 2018 with a size of Ws × Ts, and four left sub-templates 2022, 2024, 2026, and 2028 with a size of Ts × Hs. Ts is the size of the template. For example, Ts can be 1, 2, 3, or 4.

[0191] To obtain the template of the reference block, the MV of each sub-block template needs to be derived. In some embodiments, the MV of each sub-template can be borrowed from the boundary sub-block. That is, the MV of the sub-template is the same as the MV of the adjacent sub-block within the current coding block. Figure 21The image shows the MV of a sub-template of an affine motion-coded block according to some embodiments of this disclosure. For example... Figure 21 As shown, the white sub-block is sub-block 2110 of the current encoded block, and the dotted sub-block is the sub-template 2120. The MV values ​​(e.g., MV0, MV1, MV2, MV3, MV4, MV8, and MV12) of the boundary sub-block 2110 are the same as the MV of the corresponding sub-template 2120. Therefore, the sub-template 2130 of the reference block is adjacent to the reference sub-block 2140 marked "ref". In some other embodiments, the MV of the sub-template can be derived from an affine model based on the coordinates of each sub-template. Therefore, each sub-template has its own MV, which may be different from the MV of the boundary sub-blocks. Figure 22 The image shows the MV of a sub-template of an affine motion-coded block according to some embodiments of this disclosure. For example... Figure 22 As shown, the boundary sub-block 2210 has MV values ​​labeled MV0, MV1, MV2, MV3, MV4, MV8, and MV12, and the sub-template 2220 has MV values ​​labeled MV16 to MV23. Since the sub-template MV can be different from the boundary sub-block MV, the sub-template 2230 of the reference block can be separated from the reference sub-block 2240.

[0192] For an affine model, the motion vector at the sample position (x, y) can be formulated as follows: (48) in, It is the motion vector derived at the sample position (x, y). This is referred to as the fundamental MV in the model, which is the motion vector at the sample location (0, 0). a , b , c , d These are the parameters of the affine model, which can be derived based on the motion vectors of two other sample positions in the plane. Typically, the underlying MV in the model can be the motion vector at any sample position, not necessarily the motion vector at position (0, 0). If the motion vector at sample position (w, h) is chosen as the underlying MV (denoted as...), then... Then the motion vector at the sample position (x, y) can be formulated based on the following formula.

[0193] (49) For a 4-parameter affine model, the parameters b equal c And parameters d equal to parameter aTherefore, the 4-parameter affine model can be formulated based on the following formula.

[0194] (50) In theory, all parameters of the affine model, including the parameters a , b , c , d and basic MV ( , All of these can be corrected in the DMVR. However, to limit complexity, some embodiments of this disclosure propose fixing the affine parameters. a , b , c , d And only the base MV is modified. In other words, the template only undergoes translational movement during the search process. At each search position, all sub-templates have the same MV offset compared to the initial MV. Therefore, the three CPMVs and the sub-block MV also have the same MV offset after correction. If the three initial CPMVs are denoted as CPMV0, CPMV1, and CPMV2, the sub-block MV before correction is denoted as sbMV, and the three corrected CPMVs are denoted as CPMV0', CPMV1', and CPMV2', then their relationship can be expressed by the following formula: (51) (52) Among them, MV offset It is the MV correction amount in the TM correction process (i.e., the MV offset that produces the optimal TM cost).

[0195] All search modes are available, including cross search, 8-point diamond search, and 16-point diamond search. Figure 23A An integer template matching (TM) search process 2300A according to some embodiments of the present disclosure is illustrated. For example, to reduce search complexity, during the integer TM search process, in Figure 23A The search only considers 20 positions 2320 surrounding the initial position 2310. In the integer search, the position with the minimum TM cost is selected as the optimal position and set as the initial position in the subsequent fractional search. Figure 23BA half-pixel TM search process 2300B according to some embodiments of the present disclosure is illustrated. In some embodiments, to reduce complexity, eight half-pixel positions 2340 surrounding an optimal integer position 2330 are searched during the fractional search process, wherein the optimal integer position 2330 is obtained during the integer search process. In the fractional search, the position with the minimum TM cost is taken as the optimal position and output as the best position, and its corresponding MV is called the corrected MV.

[0196] Affine TM corrections can also be applied to affine coding blocks in conjunction with affine DMVRs. In this case, the TM correction process can be performed before the DMVR, or after the basic MV correction of the affine DMVR but before the affine model parameter correction of the affine DMVR, or after the affine DMVR.

[0197] Template-based reordering of merging candidates can also be performed on TM merging candidates. For example, after constructing the TM affine merging candidate list, the candidates are reordered based on the template. The TM correction is then applied to the candidates in the list, and after the TM correction, another template-based reordering and candidate similarity check can be performed to remove redundant candidates. A second TM correction can then be applied. Figure 24 This is a flowchart of a process 2400 for an affine merging pattern according to some embodiments of the present disclosure. Figure 24 Provide an example of the processing order for the affine merging pattern. For example... Figure 24 As shown, process 2400 includes steps 2410-2450. In step 2410, a TM affine merging candidate list is constructed. In step 2420, the TM affine merging candidates are reordered based on a template. In step 2430, preliminary affine TM-based correction is performed. In step 2440, the TM affine merging candidates are reordered based on the template, and a similarity check is performed to remove redundant candidates. In step 2450, final affine TM-based motion correction is performed.

[0198] In some embodiments, affine non-translation parameters can be used for correction. In some embodiments, to further improve the accuracy of the affine model, the non-translation parameters are also corrected in the TM correction. One way to correct the non-translation parameters is to add an offset to the initial parameters to obtain corrected non-translation parameters, and then derive the CPMV, sub-block MV, or sub-template MV from the corrected non-translation parameters. The template matching cost is obtained by calculating the difference between the template of the current block and the template of the reference block, the template of which is obtained based on the sub-template MV.

[0199] In some embodiments, an affine non-translational parameter search is performed. For affine models: (53) Searching for the non-translation parameter in the parameter space a, b, c and d For the search positions where the parameter values ​​are a', b', c', and d', it can be represented as: (54) Here, offset_a, offset_b, offset_c, and offset_d are the parameter offsets searched during the TM correction process. After obtaining... After the value, it can be used according to The affine model represents the sub-block MV and sub-template MV. The TM cost can be calculated using the difference between the template of the reference block and the template of the current block. This is achieved through comparison. offset_a, offset_b, offset_c and offset_d The different values ​​of the TM cost correspond to the optimal non-translation parameter, which can be used to obtain the optimal non-translation parameter. As a corrected non-translation parameter, the corresponding CPMV can be calculated as the corrected CPMV.

[0200] To reduce search complexity, a two-parameter search can be applied. That is, offset_b is The constraint is equal to - offset_c ,and offset_d Constrained to be equal to offset_a Therefore, the encoder and the decoder only need to search. offset_a and offset_b and according to offset_a and offset_b Export offset_b and offset_d In this disclosure, it is referred to as 2-parameter correction.

[0201] The MV search method described above can be applied to parameter searches. For example, for 2-parameter correction, such as... Figure 14 and Figure 15 As shown, a 3×3 square search or a 3×3 cross search scheme can be applied to obtain the optimal parameter offset. For 4-parameter correction, the search is performed in 4-dimensional space. A 3×3×3×3 square search or a 3×3×3×3 cross search scheme can be used to obtain the optimal parameter offset. For the 3×3×3×3 square search scheme, each center position has 80 neighboring positions to be searched, while for the 3×3×3×3 cross search scheme, each center position has 8 neighboring positions to be searched. The number of positions to be searched in the cross search is much smaller than that in the 3×3×3×3 square search. Assume the parameter offset of the current center position is ( offset_a, offset_b, offset_c, offset_d ), then in the 3×3×3×3 cross-search scheme, the eight adjacent positions to be searched are (offset_a + step_a, offset_b, offset_c, offset_d ), ( offset_a – step_a, offset_b, offset_c, offset_d ), ( offset_a, offset_b + step_ b, offset_c, offset_d ), ( offset_a, offset_b, - step_b offset_c, offset_d ), ( offset_a, offset_b, offset_c + step_c, offset_d ), ( offset_a, offset_b, offset_c – step_c, offset_d ), ( offset_a, offset_b, offset_c, offset_d + step_ d ), ( offset_a, offset_b, offset_c, offset_d – step_d ),in, step_a , step_b , step_ c and step_d Parameters a b c and d The search step size. After obtaining the optimal parameter offset, an offset estimation based on the error surface can be applied to further refine the parameters with higher accuracy. The search step size can be set to a fixed value. For example, since the MV accuracy in the ECM is 1 / 16 and the basic sub-block used for affine motion compensation is 4×4, the search step size can be 1 / 64, such that the MV difference between two adjacent sub-blocks is 1 / 64 × 4 = 1 / 16, which is the minimum difference in MV. The search step size can be greater than 1 / 64. A larger search step size reduces the number of search rounds, thus saving search time, but at the cost of correction accuracy. In another example, the search step size depends on the CU size. For a CU with a width of w and a height of h, the parameters... a and c The search step size (denoted as step_ac), and parameters d and b The search step size (denoted as step_db) can satisfy the following formula: (55) Here, T1 and T2 are two thresholds, which can be 1 / 16, 1 / 8, 1 / 4, or other values. These thresholds define the MV difference between samples in the current coding block and the samples furthest from the base MV, which can be generated in each search step. Note that in this example, different parameters have different search step sizes.

[0202] For the cost at each search point, the difference in parameter offsets can also be considered. That is, the cost, expressed as TMCost, can be a weighted sum of SAD or SATD between the template of the reference block and the template of the current block, as shown in the following formula: (56) Where w is the weight, sadCost is the SAD / SATD cost or mean-reduced SAD / SATD cost of the template, and ParameterOffsetCost is the cost depending on the parameter offset of the corrected parameter. When the weight w equals 0, only sadCost is considered.

[0203] When the affine parameters are determined, the underlying MV can be fixed, as described above. Figures 19A-19C The details discussed in the embodiments are as described above, and therefore, for the sake of brevity, they will not be repeated here.

[0204] In some embodiments, similar to minimum bilateral matching cost, the minimum template matching cost of the current search round can be compared with the minimum template matching cost of the previous search round. If the reduction in minimum cost is small, the search process terminates. Detailed explanations have been discussed in the above embodiments, and therefore will not be repeated here for the sake of brevity.

[0205] As discussed in the embodiments above, the quantization parameter (QP) controls quantization in video coding. Therefore, in these embodiments, different search strategies can be employed for different QPs. Detailed explanations have already been discussed in the embodiments above, and therefore will not be repeated here for the sake of brevity.

[0206] In some embodiments, since higher QP introduces more distortion, which requires more correction, a smaller maximum search round number can be set for lower QP and a larger maximum search round number for higher QP to maintain coding efficiency while reducing complexity. Other methods for reducing complexity can also be used at lower QP, as fewer corrections may be needed.

[0207] The number of search rounds may also depend on the sequence resolution. For example, for video sequences with high resolution, the maximum number of search rounds or the number of adjacent positions to be searched in each round is set to a larger value, while for video sequences with low resolution, the maximum number of search rounds or the number of adjacent positions to be searched in each round is set to a smaller value.

[0208] In some embodiments, inter-frame coded frames, such as B-frames and P-frames, have one or more reference frames. The temporal distance between the current frame and the reference frames affects the accuracy of inter-frame prediction. Therefore, in these embodiments, the parameter correction process may depend on the temporal distance or the POC distance between the current frame and the reference frames. Detailed explanations have been discussed in the above embodiments, and therefore will not be repeated here for the sake of brevity.

[0209] In some embodiments, a CPMV search can be performed. The correction is not performed directly on the non-translation parameters, but rather on the CPMVs themselves. Because the non-translation parameters are corrected, each CPMV may have a different offset in the correction, unlike the base MV correction—in which all CPMVs have the same offset. If the three initial CPMVs are denoted as CPMV0, CPMV1, and CPMV2, and the three corrected CPMVs are denoted as CPMV0', CPMV1', and CPMV2', their relationship can be expressed by the following formula: (57) in, MV offset0 , MV offset1 and MV offset2 These are the three MV offsets searched for for the three CPMVs during the TM correction process. For each search position, a sub-template MV is derived based on the CPMV corresponding to the search position, and the TM cost is calculated accordingly. The CPMV that produces the minimum TM cost is considered the corrected CPMV output by the TM correction process.

[0210] All search methods and complexity reduction methods used in non-translation parameter search can be used in the CPMV search.

[0211] In existing affine motion compensation techniques, the affine model is a linear model acquired by two or three CPMVs. Typically, the top-left, top-right, and bottom-left corners of the current coding block are used as control points, and all sub-block MVs are derived from these three CPMVs based on the linear affine model corresponding to a plane in a three-dimensional coordinate system. However, coding blocks have four corners, and the affine model does not use the fourth corner. For example, for a 6-parameter affine model represented by the top-left, top-right, and bottom-left CPMVs, such as... Figure 6 As shown, motion information at the bottom right corner is lost during the export of the sub-block MV, which leads to inaccurate exported sub-block MVs, especially for those sub-blocks near the bottom right corner of the encoded block. Furthermore, while affine motion models can capture motions more complex than translation, such as rotation, scaling, and shear mapping, they still cannot handle nonlinear motion.

[0212] This disclosure provides a solution to one or more of the above problems. In some embodiments, a bilinear model can be used. To solve the above problems, it is proposed to extend the affine model to a bilinear model, which is represented by four CPMVs and is a nonlinear model, which can be expressed by the following formula: (58) in, It is the basic MV of the bilinear model, with parameters , , , , and These are the non-translation parameters of the bilinear model, where, , , and It is a linear parameter. and These are the nonlinear parameters that define the nonlinear motion field of the bilinear model. and These are the width and height of the encoded block, respectively. The sample is in coordinates ( , The MV at point () is used. In the above formula, the MV at the top left corner is set as the base MV, therefore the origin of the coordinate system is the top left corner of the coded block. The bilinear model is a nonlinear model, and the nonlinear parameters change when the origin of the coordinate system changes. and It will also change accordingly.

[0213] Similar to the affine pattern described above, the bilinear model can be represented by four CPMVs, for example, the top-left, top-right, bottom-left, and bottom-right corners of the current coding block. For each of these... , , , The four CPMVs, the non-translation parameters can be derived based on the following formula: (59) In some embodiments, a bilinear AMVP pattern can be used. Since the bilinear model is represented by four CPMVs, signaling sends four motion vector differences (MVDs) to support the bilinear model in the AMVP pattern. First, signaling sends a bilinear model flag to indicate whether the bilinear model is used. If the flag is on, signaling sends four MVDs. In another example, signaling sends a model index to indicate which of the 4-parameter affine pattern, 6-parameter affine pattern, and 8-parameter bilinear pattern to use. When the bilinear pattern is used, signaling sends four MVDs.

[0214] Before the signaling is sent for the MVD, the motion vector prediction values ​​(MVPs) of the four CPMVs are first derived. The MVPs of the first three CPMVs can be derived using the current method. For the fourth CPMV, since it is located in the lower right corner of the coding block, there is almost no usable spatial adjacency coding information. Therefore, the MVP of the lower right corner CPMV can be derived based on the temporal motion vector prediction values ​​(TMVPs) of the motion information from the coding block in the reference image. In another example, the MVP of the lower right corner CPMV can be derived based on the MVPs of the other three CPMVs. The MVPs of the upper left, upper right, and lower left corner CPMVs are respectively... , , The MVP of the CPMV in the lower right corner It can be derived based on the following formula: (60) After deriving the MVP, for each control point, the MVD can be calculated using the difference between the MVP and the MV determined by motion estimation at the encoder. When signaling transmits the MVD, the MVD can be predicted again, and the residual of the MVD is transmitted via signaling. In some embodiments, the first, second, and third MVDs can be predicted and signaled via the current method used in the 6-parameter affine mode, and the MVD of the fourth CPMV can be predicted based on the first, second, and third MVDs. For example, the decoded MVDs of the top-left, top-right, and bottom-left CPMVs are respectively... , , The predicted value of MVD of the CPMV in the lower right corner It can be derived based on the following formula: (61) Then, at the encoding end, the difference between the predicted value of the MVD of the lower right corner CPMV and the predicted value of the MVD of the lower right corner CPMV is transmitted in the bitstream via signaling, and is represented as... ,in, and At the decoding end, after decoding the difference between the predicted value of the MVD of the lower right corner CPMV and the predicted value of the MVD of the lower right corner CPMV, the MVD of the lower right corner CPMV can be derived based on the following formula: (62) Furthermore, the CPMV in the lower right corner can be derived based on the following formula: (63) in, It is the CPMV in the lower right corner. It is the predicted value of CPMV in the bottom right corner.

[0215] In some embodiments, a bilinear merging mode can be used. For example, the bilinear mode can also be used in a merging mode. When the bilinear mode is used in a merging mode, the model information of the current coding block is inherited from the previous coding block, and after decoding, the model information is stored and used by other coding blocks in the future.

[0216] In some embodiments, a mode flag is introduced for bilinear mode. If the current block is encoded using bilinear mode, the flag is set to true and stored in memory. The bilinear mode flag for the current merge candidate can also be inherited from the previously encoded block during merge candidate derivation.

[0217] In some other embodiments, there is no explicit flag for the bilinear pattern. In this case, the bilinear pattern is considered as an 8-parameter affine model containing four CPMVs. Therefore, if the current block is encoded using the bilinear pattern, all four CPMVs are stored and used for future encoded blocks. During the merge candidate derivation, if the previous encoded block contains four CPMVs, the bilinear model information is extracted and inherited.

[0218] To support bilinear mode, the construction of the affine merging candidate is extended. For inherited candidates, the inherited bilinear candidates are derived from the bilinear models of adjacent or non-adjacent blocks. When an adjacent or non-adjacent bilinear coding block is identified, its control point motion vector is used to derive the CPMVP candidate in the affine merging list for the current coding block. Figure 25 This is a schematic diagram illustrating the inheritance of control point motion vectors according to some embodiments of the present disclosure. For example... Figure 25 As shown, if the adjacent lower left corner block 2510 is encoded in bilinear mode, then the motion vectors of the upper left, upper right, lower left, and lower right corners of the encoded block containing block 2510 can be derived. , , and When encoding block 2510 using a bilinear model, according to , , and Calculate the four CPMVs of the current coded block 2520 with angular positions. , , and .

[0219] For candidates inherited from non-adjacent blocks, the non-adjacent spatial neighbors are examined from nearest to farthest based on their distance from the current block. At a specific distance, only the first available neighbor from each side (e.g., left and top) of the current block (encoded using an affine pattern) is selected for the inheritance candidate derivation. Figure 8A As indicated by the dashed arrows, the checking order of the adjacent blocks on the left and top is from bottom to top and from right to left, respectively. When encoding a check block in bilinear mode, the four CPMVs of the current block are derived using the four CPMVs of the current block, based on the bilinear model equations.

[0220] The bilinear candidate constructed from neighboring blocks is a candidate built by fusing the adjacent translational motion information of each control point. The motion information of the control points is derived from... Figure 9 The specified spatial and temporal neighbor blocks shown are derived from CPMV. k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks B2, B3, and A2 are checked in sequence, and the MV of the first available block is used. For CPMV2, blocks B1 and B0 are checked in sequence, and for CPMV3, blocks A1 and A0 are checked in sequence. If the TMVP is available, it is used as CPMV4. After deriving the MVs of the four control points, affine merge candidates and bilinear merge candidates are constructed based on that motion information. The following combinations of control point MVs are used in sequence for construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}, and {CPMV1, CPMV2, CPMV3, CPMV4}.

[0221] Combinations of four CPMVs construct a bilinear merge candidate. Combinations of three CPMVs construct a six-parameter affine merge candidate, while combinations of two CPMVs construct a four-parameter affine merge candidate. To avoid motion scaling, related control point MV combinations are discarded if the reference indices of the control points are different.

[0222] For the first type of construction candidate, such as Figure 8B As shown, first, the positions of a non-adjacent spatial block to the left and above are determined independently. Then, the position of the upper-left neighboring block can be determined accordingly, which can surround a rectangular virtual block together with the upper-left non-adjacent block.

[0223] Figure 26This is a schematic diagram illustrating the construction of a first type of affine merge / AMVP candidate according to some embodiments of this disclosure. Figure 26 As shown, the CPMV is constructed at the top left (A), top right (B), bottom left (C), and bottom right (D) corners of virtual block 2620 using motion information from four non-adjacent blocks. The virtual block is then projected onto the current coding block 2610 to generate a corresponding construction candidate.

[0224] For the second type of construction candidate, non-translation parameters are inherited from non-adjacent spatial neighboring blocks. Specifically, the second type of bilinear construction candidate is generated by a combination of the following two items: (1) the base MV of adjacent 4×4 blocks, and (2) ... Figure 8A The defined non-translation parameters are inherited from non-adjacent spatial neighbor blocks. The first History Parameter Table (HPT) is expanded to store more parameters of the bilinear model. Each entry in the first HPT stores a set of non-translation parameters: a , b , c , d , e and f Each parameter is represented by a 16-bit signed integer. Entries in the HPT can be categorized based on reference lists and reference indices. Each reference list in the HPT supports five reference indices. The HPT category (denoted as HPTCat) can be calculated formulaically based on the following formula: (64) Here, RefList and RefIdx represent the list of reference images (0 or 1) and the reference index, respectively.

[0225] For each category, a maximum of seven entries can be stored, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. After decoding affine-coded blocks or bilinear-mode-coded blocks using reference lists RefListcur and RefIdxcur, the category HPTCat(Ref) is updated using non-translation parameters in a manner similar to updating a history-based motion vector prediction (HMVP) table. Listcur RefIdx cur Entries from seven adjacent 4×4 blocks (represented as A0, A1, A2, B0, B1, B2, B3 or T, such as...). Figure 9One of the following (as shown), and a set of non-translation parameters stored in the corresponding entry in the first HPT, are used to derive a history-non-translation-parameter-based candidate (HNTPC). The MV of adjacent 4×4 blocks is used as the base MV. The position of the current block is calculated in a formulaic manner based on the following formula (as shown). x, y ) at the MV: (65) Among them, (mv h base , mv v base ) represents the MV of adjacent 4×4 blocks, (x base , y base ) indicates the center position of the adjacent 4×4 blocks. x, y The MV can be the top left, top right, bottom left, and bottom right corners of the current block to obtain the corner position MV (CPMV) of the current block, or it can be the center of the current block to obtain the regular MV of the current block.

[0226] A second historical parameter table (HPT) containing basic MV information is also included. This second HPT contains nine entries, each including a basic MV, a reference index, up to six non-translation parameters for each reference list, and a base position. A merged HAPC can be generated based on this second HPT, using the basic MV information stored in one of the entries of the second HPT, along with the corresponding affine or bilinear model. Furthermore, paired bilinear merge candidates can be generated from two bilinear merge candidates, either historically derived or non-historically derived. A paired bilinear merge candidate is generated by averaging the CPMV of existing bilinear merge candidates in the list.

[0227] For regression-based bilinear merge candidates, such as Figure 12 As shown, the motion vectors and center positions of neighboring sub-blocks from the current coded block are used as inputs to a linear regression process to derive a model parameter set. The sub-block motion fields from previously coded affine or bilinear pattern coded blocks and the motion vectors from neighboring sub-blocks of the current coded block are used as inputs to the regression process. The predicted CPMV of the current block is derived as output. The regression-based bilinear merge candidates are derived and added to the affine merge list. The previously coded affine or bilinear pattern coded blocks can be identified by scanning non-adjacent locations and the affine HMVP table. The neighboring sub-block information of the current coded block is obtained from... Figure 12The data is obtained from 4×4 sub-blocks represented by the dotted regions depicted. For each sub-block, given a reference list, the motion vector and center coordinates corresponding to that sub-block can be used. For each bilinear coding block, a maximum of two bilinear candidates can be derived: one with neighboring sub-block information and the other without. All candidates generated by linear regression are pruned and merged into a candidate subgroup, and a ranking process based on TM cost can be applied. Then, a maximum of N candidates generated by linear regression are added to the affine merge list.

[0228] In some embodiments, bilinear candidates derived from the temporally co-located image are added to the affine merge candidate list. The same sampling pattern from the regular inter-frame merge mode is reused to scan predefined locations in the co-located image, thereby deriving temporally bilinear candidates. If the scanned location belongs to a bilinearly coded block, its CPMV is scaled to the current coded block based on its location and block size in the co-located image to derive a new bilinear candidate. This new bilinear candidate is then inserted into the existing affine merge list and reordered along with other affine merge or bilinear merge candidates through a TM cost-based reordering process.

[0229] After inserting all the above candidates into the candidate list, if the list is still not full, zero MV can be inserted at the end of the list.

[0230] In some embodiments, a DMVR for a bilinear model can be used. The decoder-side motion vector correction process can also be applied to bilinear mode coded blocks. Since the bilinear model contains fundamental MV and non-translation parameters similar to those of an affine model, the DMVR for bilinear mode coded blocks can also be divided into fundamental MV correction and non-translation parameter correction.

[0231] In some embodiments, a base MV correction can be used. In a base MV correction, the base MV... Corrected, and non-translation parameter a, b, c, d, e and f It is fixed. In this stage, the optimal MV offset for the base MV is searched by minimizing the BM cost (SAD or SATD between L0 and L1 predictions). After correction, the optimal MV offset is added to the base MV, and the sub-block MV is derived using the corrected base MV based on the bilinear model formula. The search also follows the following symmetry rules: (66) in, and These are the initial base MVs of the bilinear models for reference image list 0 (RPL0) and reference image list 1 (RPL1), respectively. and These are the corrected base MVs of the bilinear models RPL0 and RPL1, respectively. It is the base MV offset obtained during the search process.

[0232] Adding the MV offset to the base MV is equivalent to adding the same MV offset to all CPMVs of the bilinear model. Therefore, the base MV search process can also be implemented using a CPMV search process, which can be described based on the following formula: (67) in, , It refers to RPL0 MV and RPL1 MV at control point 0. , These are RPL0 MV and RPL1 MV at control point 1. , It refers to RPL0 MV and RPL1 MV at control point 2. , It refers to RPL0 MV and RPL1 MV at control point 3. It is the MV offset obtained during the search process.

[0233] The basic MV correction method for the affine model can also be used for the basic MV correction of the bilinear model for calculating the BM cost and the detailed search process. The difference is that when correcting the affine model, the MV offset is applied to three CPMVs, while when correcting the bilinear model, the MV offset is applied to four CPMVs. Similarly, the basic MV correction can also be applied at the sub-block level, where the sub-block MVs are derived based on the bilinear model and the same MV offset is added during the search process.

[0234] To reduce complexity, the basic MV correction can also be implemented at the sub-block level. Furthermore, all early termination methods and other complexity reduction methods used for basic MV correction in affine models can be applied to basic MV correction in bilinear models. These methods include, but are not limited to, setting an early termination threshold and applying an adaptive search process based on video resolution, quantization parameters, etc.

[0235] In some embodiments, non-translation parameter MV correction can be used. In addition to the basic MV, the parameters of the bilinear model can also be corrected. Similar to the non-translation correction of the affine model, the correction of the non-translation parameters of the bilinear model can be described based on the following formula: (68) in a_l0 , b_l0 , c_l0 , d_l0, e_l0 and f_l0 These are the parameters of the bilinear model of RPL0. a_l1 , b_l1 , c_l1 , d_l1 , e_l1 and f_l1 These are the parameters of the bilinear model of RPL1. offset_a, offset_b, offset_c, offset_d, offset_e and offset_f Add it to the parameters in a symmetrical manner to reduce the cost of BM. , and These are the corrected non-translation parameters for RPL0 and RPL1, respectively.

[0236] For each search point, a sub-block MV is derived, and motion compensation is performed at the sub-block level. The bilateral matching cost (SAD or SATD) can be calculated at either the sub-block or the coding block level. If the bilateral matching cost is calculated at the sub-block level, the cost of each sub-block is calculated after motion compensation, and then the coding block level cost is calculated by summing the costs of all sub-blocks after obtaining the costs of the sub-blocks. If the bilateral matching cost is calculated at the coding block level, motion compensation of all sub-blocks can be performed first to obtain the RPL0 and RPL1 predictions for the entire coding block, and then the cost of the entire coding block can be calculated. To reduce computational complexity, in some embodiments, only a subset of sub-blocks or samples of the coding block are considered in the calculation of the bilateral matching cost. That is, only the differences between a subset of sub-blocks or samples are calculated, thus allowing for the skipping of motion compensation for sub-blocks or samples not considered in the cost calculation.

[0237] For the search process described above, all search methods for non-translation parameter correction of affine models can be used for non-translation parameter correction of bilinear models.

[0238] Because there are 6 non-translation parameters that need to be corrected. In some embodiments, the search is performed in six-dimensional space. For example, a 3×3×3×3×3×3 square search or a 3×3×3×3×3×3 cross search scheme can be applied to obtain the optimal parameter offset. For the 3×3×3×3×3×3 square search scheme, each center position has 728 neighboring positions to be searched, while for the 3×3×3×3×3×3 cross search scheme, each center position has 12 neighboring positions to be searched. The number of positions to be searched in the cross search is much smaller than that in the 3×3×3×3×3×3 square search. The parameter offset for the current center position is ( offset_a, offset_b, offset_c, offset_d, offset_ e, offset_f In a 3×3×3×3×3×3 cross-search scheme, the adjacent positions to be searched are ( offset_a + step_a, offset_b, offset_c, offset_d, offset_e, offset_f ), ( offset_a – step_a, offset_b, offset_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b + step_ b, offset_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b – step_b, offset_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c + step_ c, offset_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c – step_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d + step_ d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d – step_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e + step_ e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e – step_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e, offset_f + step_ f ), ( offset_a, offset_b, offset_c, offset_d, offset_e, offset_f – step_f ),in step_a, step_b, step_c, step_d , step_e and step_f These are parameters a, b, c, d, e and f The search step size can be determined by the minimum sub-block size, the maximum coding block size, and the MV precision. After obtaining the optimal parameter offset, an offset estimation based on the error surface can be applied to further refine the parameters with higher precision.

[0239] Similar to the non-translation parameter correction of affine models, an iterative process can be applied. Since the bilinear model has four CPMVs, the iterative process is as follows: Figures 27A-27D As shown. Figures 27A-27D This is a schematic diagram illustrating example bilinear parameter correction according to some embodiments of the present disclosure. In the first step, as... Figure 27AAs shown, the top-left corner CPMV 2710 is fixed as the base MV, and the parameters are searched. Using the obtained optimal parameters, the top-right corner CPMV, bottom-left corner CPMV, and bottom-right corner CPMV can be calculated. Then, in the second step, as... Figure 27B As shown, the corrected upper right corner CPMV 2720 is fixed, and the parameters are corrected again. Using the optimal parameters obtained in the second step, the upper left corner CPMV, lower left corner CPMV, and lower right corner CPMV can be calculated. Then in the third step, as... Figure 27C As shown, the corrected lower left corner CPMV 2730 is fixed, and the parameters are corrected again. Using the optimal parameters obtained in the third step, the upper left corner CPMV, upper right corner CPMV, and lower right corner CPMV can be calculated. In the fourth step, as... Figure 27D As shown, the bottom right CPMV 2740 is fixed as the base MV, and the parameters are searched. Using the obtained optimal parameters, the top left CPMV, top right CPMV, and bottom left CPMV can be calculated. The above steps can be repeated multiple times. That is, the fourth step can be followed by the first step, where the new top left CPMV 2710 is fixed as the base MV. And the process can be repeated until specific conditions are met. For example, the conditions can be, but are not limited to: (1) a preset number of iterations; (2) the SAD or SATD between the RPL0 prediction value and the RPL1 prediction value is less than a threshold; (3) the current fixed CPMV is the same as or similar to the fixed CPMV in the previous iteration; (4) the offset of the parameters searched in this round is less than a threshold.

[0240] In some embodiments, to reduce complexity, non-translation parameters a, b, c, d, e and f The correction is divided into two steps. In the first step, the parameters are corrected. a, b, c and d In the second step, the parameters are corrected. e and f Therefore, the first step of the correction can use a non-translation parameter correction method for a 6-parameter affine model. And in the second step, only the parameters will be corrected. e and f The search process is performed in two-dimensional space. For example, it can be done using... Figure 14 and Figure 15 The 3×3 square search or 3×3 cross search shown is an example. It can also be applied to... Figures 27A-27D The iterative process is shown. The top-left, top-right, bottom-left, and bottom-right CPMV of RPL0 are represented as follows: , , and The top-left, top-right, bottom-left, and bottom-right CPMV of RPL1 are respectively represented as... , , and In the first iteration, based on the following formula, according to the parameters e and f Update the CPMV in the bottom right corner with the new value: (69) In the second iteration, based on the following formula, according to the parameters e and f Update the CPMV in the lower left corner with the new value: (70) In the third iteration, based on the following formula, according to the parameters e and f Update the CPMV in the upper right corner with the new value: (71) In the fourth iteration, based on the following formula, according to the parameters... e and f Update the top-left CPMV with the new value: (72) in, , , and These are the corrected CPMVs located at the top left, top right, bottom left, and bottom right corners of RPL0. , , and These are the corrected CPMVs in the top left, top right, bottom left, and bottom right corners of RPL1.

[0241] In some embodiments, a non-translation parameter MV can be used for correction. The difference in parameter offsets can also be considered for the cost at each search point. That is, the cost can be a weighted sum, as shown in the following formula: (73) Where w is the weight, sadCost is the SAD / SATD or mean-reduced SAD / SATD between the RPL0 and RPL1 predicted values, and ParameterOffsetCost is the cost of the parameter offset depending on the corrected parameters. When the weight w equals 0, only sadCost is considered.

[0242] In some embodiments, the affine model can be extended to a bilinear model. The modification of the bilinear model can also be applied to affine coding blocks. When applied to affine coding blocks, the modification of the affine model, including the basic MV modification and the non-translation modification, can be applied first. Then, the non-translation parameters are modified. e and f Similar to affine models, the parameters... e and f All are equal to 0. Therefore, the parameter e and f The initial value is set to 0. After correction, the parameter... e and f It can have non-zero values. If the parameter... e and f If one of the values ​​is not equal to 0, then the affine model is extended to the bilinear model. The coded block is a bilinear mode coded block.

[0243] In some embodiments, when applying bilinear model correction to an affine coded block, the lower right corner CPMV can be initialized based on the following formula: (74) in, It's the exported CPMV file in the top left corner. It's the exported CPMV file in the upper right corner. It is the exported CPMV in the lower left corner and It's the exported CPMV file in the bottom right corner.

[0244] After deriving the initial value of the lower right corner CPMV, all bilinear model correction methods can be applied. After the correction, the affine model is expanded to a bilinear model. Therefore, the current block is expanded into a block encoded in bilinear mode.

[0245] In some embodiments, template matching (TM)-based corrections to the bilinear model can be used. For example, TM-based corrections can also be applied to the bilinear model coding blocks. Figure 28 The diagram illustrates an upper and left template for a current bilinear mode coded block according to some embodiments of this disclosure. Since bilinear motion compensation is performed at the sub-block level, the template for bilinear merging candidates also includes several sub-templates, which are related to... Figure 20 The templates for the affine merging candidates shown are identical. If the template size used in TM is equal to Ts, and the sub-block size of the bilinear merging candidate is equal to Ws × Hs, then the upper template 2810 includes several sub-templates 2812, 2814, 2816, and 2818 with a size of Ws × Ts, and the left template 2820 includes several sub-templates 2822, 2824, 2826, and 2828 with a size of Ts × Hs, as shown in the figure. Figure 28As shown, the white area represents the encoded block of the current bilinear mode, comprising 16 sub-blocks 2830. The dotted area represents a template comprising four upper sub-templates 2812, 2814, 2816, and 2818 with dimensions Ws × Ts, and four left sub-templates 2822, 2824, 2826, and 2828 with dimensions Ts × Hs. Ts is the size of the template. For example, Ts can be 1, 2, 3, or 4.

[0246] Similar to the MV of a sub-template in an affine motion coding block, the MV of each sub-block template needs to be derived to obtain the template of a reference block. In some embodiments, the MV of each sub-template can be borrowed from a boundary sub-block. That is, the MV of the sub-template is the same as the MV of the adjacent sub-blocks within the current coding block. Figure 29 The image shows the MV of a sub-template of a bilinear mode coded block according to some embodiments of this disclosure. For example... Figure 29 As shown, the white sub-block is sub-block 2910 of the current encoded block, and the dotted sub-block is the sub-template 2920. The MV values ​​(e.g., MV0, MV1, MV2, MV3, MV4, MV8, and MV12) of the boundary sub-block 2910 are the same as the MV of the corresponding sub-template 2920. Therefore, the sub-template 2930 of the reference block is adjacent to the reference sub-block 2940 marked "ref". In some other embodiments, the MV of the sub-template can be derived based on the coordinates of each sub-template using a bilinear model. Therefore, each sub-template has its own MV, which may differ from the MV of the boundary sub-blocks. Figure 30 The image shows the MV of a sub-template of a bilinear mode coded block according to some embodiments of this disclosure. For example... Figure 30 As shown, the boundary sub-block 3010 has MV values ​​labeled MV0, MV1, MV2, MV3, MV4, MV8, and MV12, and the sub-template 3020 has MV values ​​labeled MV16 to MV23. Since the sub-template MV can be different from the boundary sub-block MV, the sub-template 3030 of the reference block can be separated from the reference block 3040.

[0247] In some embodiments, a basic MV correction can be used. Similar to the TM-based correction for affine coded blocks, the TM correction for bilinear mode coded blocks can also be divided into basic MV correction and non-translation parameter correction. In the basic MV correction, the basic MV... Corrected, and non-translation parameter a , b , c , d , e and fIt is fixed. In this stage, the optimal MV offset of the base MV is searched by minimizing the TM cost (the difference between the template of the current coded block and the template of the reference block). After correction, the optimal MV offset is added to the base MV, and the sub-block MV is derived using the corrected base MV based on the bilinear model formula. The base MV correction can be described by the following formula: (75) in, It is the initial basic MV of the bilinear model, and It is the modified base MV of the bilinear model. It is the base MV offset obtained during the search process.

[0248] Adding the MV offset to the base MV is equivalent to adding the same MV offset to all CPMVs of the bilinear model. Therefore, the base MV search process can also be implemented using a CPMV search process, which can be described based on the following formula: (76) in, It is the MV of control point 0. It is the MV of the first control point. It is the MV of the second control point. It is the MV of the third control point. It is the MV offset obtained during the search process.

[0249] The base MV correction is also equivalent to adding the same MV offset to all sub-block MVs, which can be described based on the following formula: (77) in, It is the initial MV of the sub-blocks within the coded block. It is the corrected MV of the sub-block, and It is the MV offset obtained during the search process.

[0250] All search patterns and methods used in TM correction and DMVR can be used for the base MV correction of the bilinear model. All early termination methods and fast algorithms used in TM correction and DMVR can be used here to reduce complexity. For other detailed designs, the current TM used for affine or translational motion-coded blocks can also be used.

[0251] In some embodiments, non-translation parameter correction can be used. Similar to the non-translation correction of DMVR in the bilinear model, in TM-based correction, the non-translation parameter can also be corrected based on the following formula: (78) in, a , b , c , d , e and f These are the initial non-translation parameters of the bilinear model. , , , , and These are the corrected non-translation parameters of the bilinear model. offset_a, offset_b, offset_c , offset_d, offset_e and offset_f It is the parameter offset found during the TM correction process.

[0252] In obtaining After the value, it can be based on the adoption The bilinear model of the parameters derives the sub-block MV and sub-template MV. The TM cost can be calculated by the difference between the template of the reference block and the template of the current block. This is achieved through comparison. offset_a , offset_b , offset_c , offset_ d , offset_e and offset_f The different values ​​of the TM cost correspond to the optimal non-translation parameter, which can be used to obtain the optimal non-translation parameter. As a corrected non-translation parameter, the corresponding CPMV can be calculated as the corrected CPMV.

[0253] For the search process, all search methods for non-translation parameter correction of affine models can be used for non-translation parameter correction of bilinear models. This is because there are 6 non-translation parameters that need correction in the bilinear model. In some embodiments, the search is performed in six-dimensional space. For example, a 3×3×3×3×3×3 square search or a 3×3×3×3×3×3 cross search scheme can be applied to obtain the optimal parameter offset. For the 3×3×3×3×3×3 square search scheme, each center position has 728 neighboring positions to be searched, while for the 3×3×3×3×3×3 cross search scheme, each center position has 12 neighboring positions to be searched. The number of positions to be searched in the cross search is much smaller than that in the 3×3×3×3×3×3 square search. The parameter offset for the current center position is ( offset_a, offset_b, offset_c, offset_d, offset_e, offset_fIn a 3×3×3×3×3×3 cross-search scheme, the adjacent positions to be searched are ( offset_a + step_a, offset_b, offset_c, offset_d, offset_e, offset_f ), ( offset_a – step_a, offset_b, offset_c, offset_d, offset_ e, offset_f ), ( offset_a, offset_b + step_b, offset_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b – step_b, offset_c, offset_d, offset_e, offset_ f ), ( offset_a, offset_b, offset_c + step_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c – step_c, offset_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d + step_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d – step_d, offset_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e + step_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e – step_e, offset_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e, offset_f + step_f ), ( offset_a, offset_b, offset_c, offset_d, offset_e, offset_f – step_f ),in step_a, step_b, step_c, step_d , step_e and step_f These are parameters a, b, c, d, e and f The search step size can be determined by the minimum sub-block size, the maximum coding block size, and the MV precision. After obtaining the optimal parameter offset, an offset estimation based on the error surface can be applied to further refine the parameters with higher precision.

[0254] Similar to the non-translation parameter correction in affine models, an iterative process can be applied. Since the bilinear model has four CPMVs, the iterative process is as follows: Figures 27A-27D As shown above, this has already been discussed, so for the sake of brevity, it will not be repeated here.

[0255] In some embodiments, to reduce complexity, non-translation parameters a, b, c, d, e and f The correction is divided into two steps. In the first step, the parameters are corrected. a, b, c and d In the second step, the parameters are corrected. e and f Therefore, the first step of the correction can use a non-translation parameter correction method for a 6-parameter affine model. And in the second step, only the parameters will be corrected. e and f The detailed operations have been discussed in the above embodiments, so they will not be repeated here for the sake of brevity. The top-left corner CPMV, top-right corner CPMV, bottom-left corner CPMV, and bottom-right corner CPMV are respectively represented as... , , and In the first iteration, based on the following formula, according to the parameters e and f The new value updates the CPMV in the bottom right corner: (79) In the second iteration, based on the following formula, according to the parameters e and f Update the CPMV in the lower left corner with the new value: (80) In the third iteration, based on the following formula, according to the parameters e and f Update the CPMV in the upper right corner with the new value: (81) In the fourth iteration, based on the following formula, according to the parameters... e and f Update the top-left CPMV with the new value: (82) in, , , and It refers to the corrected CPMV of the top left, top right, bottom left, and bottom right corners.

[0256] For the cost at each search point, the difference in parameter offsets can also be considered. That is, the cost can be a weighted sum of the parameter offset and the SAD or SATD between the template of the reference block and the template of the current block, as shown in the following formula: (83) Where w is the weight, sadCost is the SAD / SATD or mean-reduced SAD / SATD cost of the template, and ParameterOffsetCost is the cost depending on the parameter offset of the corrected parameter. When the weight w equals 0, only sadCost is considered.

[0257] To reduce search complexity, the early termination method or other adaptive search methods used in affine TM correction can also be used in bilinear TM correction.

[0258] The correction to the bilinear model can also be applied to affine coded blocks. When applied to affine coded blocks, the correction to the affine model can be applied first, including corrections to the basic MV and non-translation parameters. Then, the non-translation parameters are corrected. e and f Similar to affine models, the parameters... e and f All are equal to 0. Therefore, the parameter e and f The initial value is set to 0. After correction, the parameter... e andf It can have non-zero values. If the parameter... e and f If one of the values ​​is not equal to 0, then the affine model can be extended to the bilinear model. The coded block is a bilinear mode coded block.

[0259] In some embodiments, when applying bilinear model correction to an affine coding block, the lower right corner CPMV is initialized based on the following formula: (84) in, It's the exported CPMV file in the top left corner. It's the exported CPMV file in the upper right corner. It is the exported CPMV in the lower left corner and It's the exported CPMV file in the bottom right corner.

[0260] After deriving the initial value of the lower right corner CPMV, all bilinear model correction methods can be applied. After the correction, the affine model is expanded to a bilinear model. Therefore, the current block can be expanded into a block encoded in bilinear mode.

[0261] Figure 31 This is a flowchart of an example method 3100 for encoding a video bitstream according to some embodiments of the present disclosure. The method 3100 may be generated by an encoder (e.g., Figure 2 The encoder 200 performs the operation to encode the video bitstream. For example, the encoder may be implemented to encode the bitstream (e.g., ...). Figure 2 A device for encoding the video bitstream (228) in order to reconstruct video frames or video sequences (e.g., Figure 4 One or more software or hardware components of the device 400 in the middle. For example, a processor (e.g., Figure 4 The processor 402 in the processor can execute the method 3100. Figure 31 As shown, the method 3100 includes the following steps 3110-3150.

[0262] In step 3110, the encoder determines multiple motion vector predictions (MVPs) for the target block based on multiple control point motion vectors (CPMVs). The bilinear model is characterized by multiple CPMVs. The bilinear model includes multiple linear parameters, multiple nonlinear parameters, and a basic motion vector (MV).

[0263] In step 3120, the encoder determines multiple motion vector differences (MVDs). Each motion vector difference is based on the difference between the MVP and its CPMV for each control point. For example, four motion vector differences (MVDs) can be associated with the top-left CPMV, the top-right CPMV, the bottom-left CPMV, and the bottom-right CPMV, respectively.

[0264] In step 3130, the encoder sets a bilinear model flag in the bitstream to indicate whether a bilinear model is used, or a model index to indicate which model is used. For example, if the bilinear model flag is on, four MVDs are set and signaled. In another example, a model index may be set and signaled to indicate which of the following is used: a 4-parameter affine mode, a 6-parameter affine mode, or an 8-parameter bilinear mode.

[0265] In step 3140, in response to using the bilinear model, the encoder sets multiple motion vector differences (MVDs) in the bitstream. For example, since the bilinear model is represented by four CPMVs, in order to support the bilinear model in AMVP mode, four motion vector differences (MVDs) associated with the top-left CPMV, top-right CPMV, bottom-left CPMV, and bottom-right CPMV can be set.

[0266] In step 3150, the encoder sets a difference between a corresponding MVD and a corresponding MVP associated with a corresponding CPMV in the bitstream. In some embodiments, the encoder may determine the corresponding MVP based on a first MVD, a second MVD, and a third MVD among the plurality of MVDs. For example, the difference may be set and signaled in the bitstream, and the difference is the difference between the MVD of the lower right CPMV and the predicted value of the MVD of the lower right CPMV. For example, the predicted value of the MVD of the lower right CPMV... It can be based on the decoded MVD of the top left corner CPMV, top right corner CPMV, and bottom left corner CPMV, i.e. , , To export. In some embodiments, the encoder transmits the bitstream via a public or private network.

[0267] In some embodiments, in method 3100, the encoder may further construct bilinear merging candidates by using a combination of multiple CPMVs, thereby using the bilinear model in the merging mode. The operation of the bilinear merging mode has been discussed in detail above in the embodiments of this disclosure, and therefore will not be repeated here for the sake of brevity.

[0268] In some embodiments, in method 3100, the encoder may further perform a basic MV correction on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed, or perform a parameter correction on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters. The operations of the basic MV correction and the parameter correction for the bilinear model have been discussed in detail in the embodiments disclosed above, and therefore will not be repeated here for the sake of brevity.

[0269] Figure 32 This is a flowchart of an example method 3200 for decoding a video bitstream according to some embodiments of the present disclosure. In some embodiments, the method 3200 may be performed by a decoder (e.g., Figure 3 The decoder (300) is used to perform this operation. For example, the decoder may be implemented for processing the bitstream (e.g., ...). Figure 3 The video bitstream 228 in the video stream is decoded to reconstruct the video frames or video sequence of the bitstream (e.g., Figure 3 The device (e.g., the video stream 304 in the video stream) Figure 4 One or more software or hardware components of the device 400 in the middle. For example, a processor (e.g., Figure 4 The processor 402 in the processor can execute the method 3200. Figure 32 As shown, the method 3200 may include steps 3610-3640.

[0270] In step 3210, the decoder receives a bitstream comprising encoded video data of the target block. In step 3220, the decoder determines multiple motion vector differences (MVDs) based on the bitstream. Each MVD is based on the difference between the predicted motion vector (MVP) corresponding to each control point of the target block and its control point motion vector (CPMV).

[0271] In step 3230, the decoder decodes from the bitstream a bilinear model flag indicating whether a bilinear model is used, or a model index indicating which model is used. The bilinear model includes multiple linear parameters and multiple nonlinear parameters, as well as a fundamental motion vector (MV). Details of the bilinear model flag or the model index are... Figure 31 The details of step 3130 are the same or similar, so for the sake of brevity, they will not be repeated here.

[0272] In step 3240, in response to using the bilinear model, the decoder reconstructs the target block based on multiple CPMVs derived from the multiple MVDs. The bilinear model is characterized by the multiple CPMVs. In some embodiments, the decoder can decode the difference between a corresponding MVD and a corresponding MVP associated with a corresponding CPMV from the bitstream. The decoder can then derive the corresponding MVD based on the decoded difference and the corresponding MVP, and derive the corresponding CPMV based on the corresponding MVD and the corresponding MVP. For example, after decoding the difference between the MVD of the lower right CPMV and the predicted value of the MVD of the lower right CPMV, the MVD of the lower right CPMV can be derived. The lower right CPMV can then be derived accordingly.

[0273] In some embodiments, in method 3200, the decoder may also construct bilinear merge candidates by using a combination of the plurality of CPMVs, thereby using the bilinear model in the merge mode. The operation of the bilinear merge mode has been discussed in detail above in the embodiments of this disclosure, and therefore will not be repeated here for the sake of brevity.

[0274] In some embodiments, in method 3200, the decoder may further perform a basic MV correction on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed, or perform a parameter correction on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters. The operations of the basic MV correction and the parameter correction for the bilinear model have been discussed in detail in the embodiments disclosed above, and therefore will not be repeated here for the sake of brevity.

[0275] The embodiments described in this disclosure can be freely combined.

[0276] In some embodiments, a non-transitory computer-readable storage medium is also provided on which a bitstream is stored. The bitstream can be encoded and decoded based on the disclosed bilinear model for video coding.

[0277] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (e.g., the disclosed encoder and decoder) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAMs, caches, registers, any other memory chips or cassette tapes, and their networking versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0278] In some embodiments, a method for storing a bitstream is provided. The method includes receiving a video sequence comprising one or more images, generating a bitstream comprising encoded information associated with the video sequence, and storing the bitstream in a non-transitory computer-readable medium. The operations for generating the bitstream include... Figure 31 The steps in method 3100 are the same or similar, therefore, for the sake of brevity, they will not be repeated here.

[0279] It should be noted that the relational terms such as "first," "second," etc., used in this document are only used to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be identical in meaning and open-ended, because one or more items following any of these words do not imply an exhaustive list of such one or more items, nor do they imply limitation to only the listed one or more items.

[0280] As used herein, unless otherwise specified, the term "or" covers all possible combinations unless impractical. For example, if it is specified that a database may include A or B, then unless otherwise specified or impractical, the database may include A, or B, or A and B. As a second example, if it is specified that a database may include A, B, or C, then unless otherwise specified or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0281] The embodiments may be further described using the following terms: 1. A method for encoding video data, the method comprising: Based on multiple control point motion vectors (CPMV), multiple motion vector prediction values ​​(MVP) of the target block are determined, wherein the bilinear model is characterized by the multiple CPMVs; Multiple motion vector differences (MVDs) are determined, where each MVD is based on the difference between the MVP and the CPMV corresponding to each control point; and In response to using the bilinear model, the plurality of MVDs are set in the bitstream.

[0282] 2. The method according to Clause 1 further includes: In the bitstream, a bilinear model flag is set to indicate whether the bilinear model is used, or a model index is set to indicate which model is used.

[0283] 3. The method according to clause 1 or 2 further includes: Based on the first MVD, second MVD, and third MVD among the multiple MVDs, determine the corresponding MVP; and In the bitstream, the difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV is set.

[0284] 4. The method according to any one of clauses 1 to 3, further comprising: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

[0285] 5. The method according to any one of clauses 1 to 4, wherein the bilinear model comprises: a plurality of linear parameters and a plurality of nonlinear parameters, and a fundamental motion vector (MV).

[0286] 6. The method according to Clause 5 further includes: A basic MV correction is performed on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed.

[0287] 7. The method described according to Clause 5 or 6 further includes: Parameter correction is performed on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters.

[0288] 8. The method according to any one of clauses 1 to 7, further comprising: The bit stream is transmitted via a public or private network.

[0289] 9. A method for decoding a video bitstream, the method comprising: Receive a bitstream, the bitstream comprising encoded video data of the target block; Based on the bitstream, multiple motion vector differences (MVDs) are determined, wherein each motion vector difference is based on the difference between the predicted motion vector (MVP) and the control point motion vector (CPMV) corresponding to each control point of the target block; and In response to the use of a bilinear model, the target block is reconstructed based on a plurality of CPMVs derived from the plurality of MVDs, wherein the bilinear model is characterized by the plurality of CPMVs.

[0290] 10. The method according to Clause 9 further includes: From the bitstream, decode the bilinear model flag used to indicate whether the bilinear model is used, or the model index used to indicate which model is used.

[0291] 11. The method according to clause 9 or 10, further comprising: Decode from the bitstream the difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV; and Based on the decoded difference and the corresponding MVP, derive the corresponding MVD; and Based on the corresponding MVD and the corresponding MVP, the corresponding CPMV is derived.

[0292] 12. The method according to any one of clauses 9 to 11, further comprising: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

[0293] 13. The method according to any one of clauses 9 to 12, wherein the bilinear model comprises: a plurality of linear parameters and a plurality of nonlinear parameters, and a fundamental motion vector (MV).

[0294] 14. The method according to Clause 13 further includes: A basic MV correction is performed on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed.

[0295] 15. The method described pursuant to Clause 13 or 14, further comprising: Parameter correction is performed on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters.

[0296] 16. A method for storing a bit stream, the method comprising: Receive a video sequence, the video sequence comprising one or more images; A bitstream is generated through the following steps, the bitstream including encoded information associated with the video sequence: Based on multiple control point motion vectors (CPMV), multiple motion vector prediction values ​​(MVP) of target blocks in one or more images are determined, wherein a bilinear model is characterized by the multiple CPMVs; Multiple motion vector differences (MVDs) are determined, where each MVD is based on the difference between the MVP and the CPMV corresponding to each control point; and In response to using the bilinear model, the plurality of MVDs are encoded in the bitstream; and The bitstream is stored in a non-transitory computer-readable medium.

[0297] 17. The method according to Clause 16, wherein the encoded information includes: A bilinear model flag used to indicate whether the bilinear model is used, or a model index used to indicate which model is used.

[0298] 18. The method according to clause 16 or 17, wherein the encoded information includes: The difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV.

[0299] 19. The method according to any one of clauses 16 to 18, wherein generating the bitstream further comprises: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

[0300] 20. The method according to any one of clauses 16 to 19, wherein the bilinear model comprises: a plurality of linear parameters and a plurality of nonlinear parameters, and a fundamental motion vector (MV).

[0301] It should be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. When executed by the processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art should also understand that multiple modules / units described above can be combined into one module / unit, and each module / unit described above can be further divided into multiple sub-modules / sub-units.

[0302] In the foregoing specification, numerous specific details have been described with reference to embodiments, which may vary depending on the implementation. Certain adjustments and modifications may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the appended claims. The sequence of steps shown in the figures is also to be considered for illustrative purposes only and is not intended to limit one to any particular sequence of steps. Therefore, those skilled in the art will understand that these steps may be performed in a different order while implementing the same method.

[0303] Exemplary embodiments have been disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terminology has been used, it is used in a general and descriptive sense only and is not intended to be limiting.

Claims

1. A method for encoding video data, comprising: Based on the multiple control point motion vectors (CPMV), multiple motion vector prediction values ​​(MVP) of the target block are determined, wherein the bilinear model is characterized by the multiple CPMVs; Multiple motion vector differences (MVDs) are determined, where each MVD is based on the difference between the MVP and the CPMV corresponding to each control point; and In response to using the bilinear model, the plurality of MVDs are set in the bitstream.

2. The method according to claim 1, further comprising: In the bitstream, a bilinear model flag is set to indicate whether the bilinear model is used, or a model index is set to indicate which model is used.

3. The method according to claim 1, further comprising: Based on the first MVD, second MVD, and third MVD among the multiple MVDs, determine the corresponding MVP; as well as In the bitstream, the difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV is set.

4. The method according to claim 1, further comprising: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

5. The method according to claim 1, wherein, The bilinear model includes multiple linear parameters and multiple nonlinear parameters, as well as a basic motion vector MV.

6. The method of claim 5, further comprising: A basic MV correction is performed on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed.

7. The method according to claim 5, further comprising: Parameter correction is performed on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters.

8. The method according to claim 1, further comprising: The bit stream is transmitted via a public or private network.

9. A method for decoding a video bitstream, comprising: Receive a bitstream, the bitstream comprising encoded video data of the target block; Based on the bitstream, multiple motion vector differences (MVDs) are determined, wherein each motion vector difference is based on the difference between the motion vector prediction value (MVP) corresponding to each control point of the target block and its control point motion vector (CPMV); and In response to the use of a bilinear model, the target block is reconstructed based on a plurality of CPMVs derived from the plurality of MVDs, wherein the bilinear model is characterized by the plurality of CPMVs.

10. The method of claim 9, further comprising: From the bitstream, decode the bilinear model flag used to indicate whether the bilinear model is used, or the model index used to indicate which model is used.

11. The method of claim 9, further comprising: From the bitstream, decode the difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV; as well as Based on the decoded difference and the corresponding MVP, the corresponding MVD is derived; as well as Based on the corresponding MVD and the corresponding MVP, the corresponding CPMV is derived.

12. The method of claim 9, further comprising: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

13. The method according to claim 9, wherein, The bilinear model includes multiple linear parameters and multiple nonlinear parameters, as well as a basic motion vector MV.

14. The method of claim 13, further comprising: A basic MV correction is performed on the bilinear model to correct the basic motion vector of the bilinear model while keeping the plurality of linear parameters and the plurality of nonlinear parameters fixed.

15. The method of claim 13, further comprising: Parameter correction is performed on the bilinear model to correct the plurality of linear parameters and the plurality of nonlinear parameters.

16. A method for storing a bit stream, comprising: Receive a video sequence, the video sequence comprising one or more images; A bitstream is generated through the following steps, the bitstream including encoded information associated with the video sequence: Based on multiple control point motion vectors (CPMV), multiple motion vector prediction values ​​(MVPs) of target blocks in one or more images are determined, wherein the bilinear model is characterized by the multiple CPMVs; Multiple motion vector differences (MVDs) are determined, where each MVD is based on the difference between the MVP and the CPMV corresponding to each control point; and In response to using the bilinear model, the plurality of MVDs are encoded in the bitstream; and The bitstream is stored in a non-transitory computer-readable medium.

17. The method according to claim 16, wherein, The encoded information includes: A bilinear model flag used to indicate whether the bilinear model is used, or a model index used to indicate which model is used.

18. The method according to claim 16, wherein, The encoded information includes: The difference between the corresponding MVD and the corresponding MVP associated with a corresponding CPMV.

19. The method according to claim 16, wherein generating the bitstream further comprises: Bilinear merge candidates are constructed by combining the multiple CPMVs, and the bilinear model is used in the merge mode.

20. The method of claim 16, wherein, The bilinear model includes multiple linear parameters and multiple nonlinear parameters, as well as a basic motion vector MV.