Motion compensation using sparse optical flow representation

CN116134817BActive Publication Date: 2026-09-25HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180044096.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-22
Filing Date
2021-04-21
Publication Date
2026-09-25
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

但是,压缩率相当低

Benefits of technology

[0198]本发明的一些优点如下所述:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116134817B_ABST
    Figure CN116134817B_ABST
Patent Text Reader

Abstract

The invention provides a method and apparatus for estimating motion vectors of a dense motion field from a sub-sampled sparse motion field. The sparse motion field comprises two or more motion vectors and their respective starting positions. For each of the motion vectors, a transformation is derived that transforms the motion vector from its starting point to a target point. The transformed motion vector then contributes to the motion vector estimation at the target position. The contribution of each motion vector is weighted. Such motion estimation can be easily used for video encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention generally relate to the field of video processing, and more particularly to motion compensation and methods and apparatus for video processing. Background Technology

[0002] Video decoding (video encoding and decoding) is widely used in digital video applications, such as broadcast digital television, video transmission based on the Internet and mobile networks, real-time conversational applications such as video chat and video conferencing, DVD and Blu-ray discs, video content capture and editing systems, and portable cameras for security applications.

[0003] Even relatively short videos require a significant amount of video data to describe, which can be challenging when streaming or otherwise transmitting data over communication networks with limited bandwidth. Therefore, video data is typically compressed before transmission over modern telecommunications networks. The size of the video can also be an issue when storing it on storage devices due to potentially limited memory resources. Video compression devices typically encode video data using software and / or hardware at the source side before transmitting or storing it, reducing the amount of data required to represent a digital video image. Video decompression devices then decode the video data and receive the compressed data at the destination side. Given limited network resources and the growing demand for higher video quality, there is a need to improve compression and decompression techniques to increase compression ratios with minimal impact on image quality.

[0004] Generally, image compression can be lossless or lossy. In lossless image compression, the original image can be perfectly reconstructed from the compressed image. However, the compression ratio is quite low. In contrast, lossy image compression can achieve a high compression ratio, but the drawback is that it cannot perfectly reconstruct the original image. Especially when used at low bitrates, lossy image compression introduces visible spatial compression artifacts. Summary of the Invention

[0005] The present invention relates to a method and apparatus for estimating motion vectors at a given target location.

[0006] This invention is defined by the scope of the independent claims. Some advantageous embodiments are provided in the dependent claims.

[0007] Specifically, embodiments of the present invention provide an efficient method for estimating a motion vector at a given target location from a sparse motion field representation. This is performed by weighting contributing motion vectors, wherein the contributing motion vectors are obtained by performing at least two different transformations on motion vectors belonging to the sparse motion field representation.

[0008] According to one aspect, a method for estimating a motion vector at a target location is provided. The method includes: acquiring two or more starting positions and two or more motion vectors respectively starting from the two or more starting positions; for each of the two or more starting positions, acquiring a corresponding transformation for transforming the motion vectors starting from the starting position to another position; determining two or more contributing motion vectors by transforming each of the two or more motion vectors from the starting position to the target position of the corresponding transformation using the corresponding transformation; and estimating the motion vector at the target location, wherein the motion vector includes a weighted average of the two or more contributing motion vectors.

[0009] The method described above can use a more complex motion model to represent a larger region while using fewer parameters to describe it. These parameters can be predicted from the optical flow available at the encoder, rather than using most well-known, complex rate-distortion optimization (RDO) methods. With a simpler motion model, the encoder can perform more subsampling. A more complex motion model can be used to describe motion within a larger region of the predicted frame, which reduces signaling overhead.

[0010] In some implementations, the nonlinear function is a Gaussian distribution function. This nonlinear motion model can densify any sparse representation of the motion field. For example, distance corresponds to the square norm. The square norm is easy to compute, especially when used in conjunction with a Gaussian distribution function. Since the Gaussian distribution has a quadratic term for distance, the square root calculation required to compute the norm is unnecessary.

[0011] According to one embodiment, the acquisition of the transformation includes: acquiring a motion vector starting from the other position; and estimating parameters of the affine transformation based on the affine transformation from the motion vector starting from the initial position to the motion vector starting from the other position. Affine transformations can cover a wide range of motion types commonly found in natural videos, such as scaling, rotation, or translation.

[0012] For example, the two or more starting positions belong to a set of Ns starting positions, where Ns > 2, and the starting positions are arranged in a predefined order; for starting position j, 0 ≤ j ≤ Ns, the other position is position j+1 in the predefined order. The ordering of positions and possibly associated motion vectors can efficiently store and / or transmit these edge information parameters.

[0013] For example, the weight of a contribution vector depends on the starting position of the corresponding transformed motion vector within the predefined order. In this way, the association between weights and position / motion vectors can be stored or transmitted without explicit indication.

[0014] According to one embodiment, the two or more starting positions are sample positions within image slices, wherein the image comprises multiple slices, each slice being a group of image samples smaller than the image itself. Motion vector estimation and transformation based on these slices can better adapt to the image content and enable some form of parallel processing.

[0015] In some exemplary implementations, the method includes the step of reconstructing the motion vector field of the image segments, including estimating a motion vector starting from each (e.g., each integer) sample target position P(x,y) of the segment, wherein the motion vector does not belong to two or more starting positions available for the corresponding motion vector. This enables the reconstruction of a dense motion field from a sparse (subsample) motion field, thus constituting an approximate optical flow. Optical flow can be used for prediction in video codecs, etc. It should be noted that, in addition to estimating the motion vector starting from each position that does not belong to two or more starting positions, in some embodiments, two or more starting positions may also be estimated.

[0016] According to one embodiment, the two or more starting positions and the two or more motion vectors starting from the two or more starting positions are obtained by parsing from a bitstream associated with the slice of the image; the weights used in the weighted average are determined based on one or more parameters parsed from the bitstream. Providing these parameters in the bitstream enables the encoder and decoder to transmit these parameters (of the motion vector field and / or video image).

[0017] Furthermore, in some implementations, the two or more starting positions within the slice of the image are determined based on features of the slice decoded from the bitstream; the two or more motion vectors starting from the two or more starting positions are obtained by parsing from the bitstream associated with the slice; and the weights used in the weighted average are determined based on one or more parameters parsed from the bitstream. Indicating side information to specify the weighting function in the bitstream can adapt the weights to the image content, thus achieving more accurate reconstruction.

[0018] For example, the two or more starting positions and the two or more motion vectors starting from the two or more starting positions are obtained by determining a motion vector field and by subsampling the obtained motion vector field, wherein the motion vector field includes the motion vector at each (e.g., each integer) sample position of the slice of the image; and / or the weights of the corresponding contributing motion vectors are determined by rate-distortion optimization or machine learning.

[0019] According to one aspect, a method for decoding an image is provided. The method includes: estimating a motion vector at a target location of a sample as described in the above embodiments and examples; predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image; and reconstructing the sample at the target location based on the prediction. The decoder performing motion vector estimation can reconstruct the motion vector of any part of the image using very few parameters. Therefore, speed can be used efficiently when parameters are passed to the decoder.

[0020] According to one aspect, a method for encoding an image is provided. The method includes: estimating a motion vector at a target location according to the above embodiments and examples; predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image; and encoding the sample at the target location based on the prediction. The encoder performing motion vector estimation can efficiently encode a motion vector field.

[0021] According to one aspect, an apparatus for estimating a motion vector at a target location is provided. The apparatus includes processing circuitry comprising: circuitry for acquiring two or more starting positions and two or more motion vectors respectively starting from the two or more starting positions; circuitry for: for each of the two or more starting positions, acquiring a corresponding transformation for transforming the motion vectors starting from the starting positions to another position; circuitry for using the corresponding transformation to transform each of the two or more motion vectors from the starting positions to the target position of the corresponding transformation, determining two or more contributing motion vectors; and circuitry for estimating a motion vector at the target location, wherein the motion vector includes a weighted average of the two or more contributing motion vectors.

[0022] According to one aspect, an encoding apparatus for encoding an image is provided. The apparatus includes: means for estimating a motion vector at a target location according to any of the above embodiments and examples; a sample predictor for predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image; and a bitstream generator for encoding the sample at the target location based on the prediction.

[0023] According to one aspect, a decoding apparatus for decoding an image is provided. The apparatus includes: means for estimating a motion vector at a target location according to any of the above embodiments and examples; a sample predictor for predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image; and a sample reconstructor for reconstructing the sample at the target location based on the prediction.

[0024] Furthermore, the present invention also provides a method corresponding to the steps performed by the above-described processing circuit.

[0025] According to one aspect, a computer product is provided, comprising program code for performing the methods described above. The computer product may be provided in a non-transitory medium and includes instructions that, when executed on one or more processors, perform the steps of the methods described above.

[0026] The above-mentioned device can be implemented on an integrated chip.

[0027] Any of the above embodiments and exemplary implementations can be combined together. Attached Figure Description

[0028] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0029] Figure 1 This is a block diagram of an exemplary encoding device for encoding video.

[0030] Figure 2 This is a block diagram of an exemplary decoding device for decoding video.

[0031] Figure 3 It is a schematic diagram of various sports field transformations.

[0032] Figure 4 This is a diagram that roughly compares multiple parameters based on the order of the underlying motion model.

[0033] Figure 5 This is a block diagram of a codec (encoding end) that can implement some embodiments.

[0034] Figure 6 This is a flowchart of an exemplary method for estimating motion vectors.

[0035] Figure 7 This is a schematic diagram of two transformations of the motion vector.

[0036] Figure 8 This is a schematic diagram of two transformations of the motion vector.

[0037] Figure 9 This is a schematic diagram illustrating two transformations of motion vectors and obtaining the final motion vector.

[0038] Figure 10 This is a schematic diagram of the weight distribution.

[0039] Figure 11 This is a schematic diagram illustrating the process of obtaining estimated motion vectors through weighted averaging.

[0040] Figure 12 This is a flowchart of a method for estimating the motion vector at a given target location.

[0041] Figure 13 It is a block diagram of a device for estimating the motion vector at a given target position.

[0042] Figure 14 This is a block diagram of an example of a video decoding system that can be used to implement some embodiments.

[0043] Figure 15 This is a block diagram of another example of a video decoding system that can be used to implement some of the embodiments.

[0044] Figure 16 It is a block diagram of an example of an encoding or decoding device;

[0045] Figure 17 This is a block diagram of another example of an encoding or decoding device. Detailed Implementation

[0046] In the following description, reference is made to the accompanying drawings, which form part of this invention, which illustrate by way of description specific aspects of embodiments of the invention or aspects in which embodiments of the invention may be used. It should be understood that embodiments of the invention may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.

[0047] For example, it should be understood that the disclosure relating to the described method also applies to the corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., functional units) to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units perform one or more of a plurality of steps respectively), even if such one or more units are not explicitly described or shown in the figures. On the other hand, for example, if a specific apparatus is described according to one or more units (e.g., functional units), the corresponding method may include a step to perform the function of one or more units (e.g., one step performs the function of one or more units, or multiple steps perform the function of one or more of a plurality of units respectively), even if such one or more steps are not explicitly described or shown in the figures. Furthermore, it should be understood that, unless otherwise expressly stated, features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0048] Video decoding generally refers to the processing of image sequences that form a video or video sequence. In the field of video decoding, the terms "frame" and "picture / image" can be used synonymously. Video decoding (or generally referred to as decoding) consists of two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing (e.g., by compression) the raw video image to reduce the amount of data required to represent the video image (thus storing and / or transmitting it more efficiently). Video decoding is performed on the destination side and typically involves inverse processing relative to the encoder to reconstruct the video image. The "decoding" of video images (or generally referred to as images) in the embodiments should be understood as involving the "encoding" or "decoding" of video images or respective video sequences. The encoding and decoding parts are also collectively referred to as encoding and decoding (encoding and decoding).

[0049] In lossless video decoding, the original video image can be reconstructed, meaning the reconstructed video image has the same quality as the original (assuming no transmission loss or other data loss during storage or transmission). In lossy video decoding, further compression is performed through quantization to reduce the amount of data representing the video image, and the decoder cannot completely reconstruct the video image, meaning the quality of the reconstructed video image is lower or worse than the quality of the original video image.

[0050] Several video decoding standards belong to the "lossy hybrid video codec" group (i.e., combining spatial and temporal predictions in the sample domain with 2D transform coding for quantization in the transform domain). Each image in a video sequence is typically segmented into a set of non-overlapping blocks, and decoding is usually performed at the block level. In other words, the encoder side typically processes the video at the block (video block) level, i.e., encoding, by generating prediction blocks, for example, using spatial (intra-frame) predictions and / or temporal (inter-frame) predictions, subtracting the prediction blocks from the current block (the block currently being processed / to be processed) to obtain residual blocks, transforming and quantizing the residual blocks in the transform domain to reduce the amount of data to be transmitted (compressed); while the decoder side applies the inverse processing portion relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. Furthermore, the encoder replicates the decoder processing loop, such that the encoder and decoder generate the same predictions (e.g., intra-frame and inter-frame predictions) and / or reconstructions for processing subsequent blocks, i.e., decoding.

[0051] The following explains some technical terms used in at least some of the embodiments described herein.

[0052] A reference frame is a frame used as a reference (sometimes also called a reference image) for purposes such as prediction. Prediction here can be inter-frame prediction, meaning a temporal prediction made based on samples in another frame for some samples in the current frame.

[0053] A motion vector is a vector that specifies the spatial distance between two corresponding points in two different frames, typically represented as v = [v_x, v_y]. Such a motion vector can be a 2D motion vector. However, for 2D motion vectors, it is usually assumed that a reference image (frame) is known. Generally, the reference image (frame) can also be an indicator of the motion vector's coordinates, and these coordinates can be 3D. There can be many more coordinates.

[0054] The coordinates here sometimes refer to the position of a pixel (sample) or the position of the origin of the motion vector in the image, denoted as p = [px, py].

[0055] A motion field is a set of {p,v} pairs, abbreviated as MF, and sometimes called a motion vector field (MVF). In other words, a motion field is a set of motion vectors with different origins within an image.

[0056] Optical flow indicates the distribution of the apparent velocity of luminance pattern motion in an image (frame). Specifically, this optical flow can be represented / indicated by a motion field.

[0057] A dense motion field is a motion field that covers every sample in an image (e.g., an integer sample or any sample including subsamples in a desired sample grid). Here, when storing or transmitting a dense motion field, p is redundant if the image size is known because the motion vectors can be arranged in the row scan order for each sample (for each p), such as scanning from left to right and from top to bottom, or any other way.

[0058] A sparse motion field is a motion field that does not cover all pixels (e.g., any sample in an integer sample or a desired sample grid). Here, knowing p is necessary when storing or transmitting the sparse motion field, and when p can vary. It is also possible to organize the locations p into predetermined patterns (e.g., every four samples or a certain grid, etc.), thus eliminating the need to store or transmit these locations.

[0059] Sub-sampling is the process of transforming dense data (such as dense motion fields) into sparse representations (such as sparse motion fields).

[0060] The current frame here refers to the frame that is currently being processed (e.g., to determine the motion vectors or motion field of the frame, to encode or decode, to predict, to filter, or to otherwise process the frame). The term "currently" usually refers to a part currently under consideration when describing a process or apparatus, such as a frame.

[0061] A predicted frame is a frame that includes samples estimated using processed information (e.g., encoded, decoded, transmitted, received, filtered, etc.). Such information can be, for example, a reference frame and / or other transmitted side information.

[0062] A residual frame is the difference between a predicted frame and the current frame. Residual frames can be used, for example, to compensate for prediction errors. Specifically, residual frames can be encoded and transmitted to compensate for prediction errors, for example, by being added to the predicted frame at the decoder / receiver.

[0063] Motion compensation is a term that refers to the generation of a predicted image using a reference image and motion information.

[0064] Inter-prediction is a prediction technique in video coding where motion information is indicated to the decoder (or derived rather than indicated on the decoder side) so that the decoder can use previously decoded frames to generate a predicted image.

[0065] A bell-shaped function is a Gaussian distribution (or any similar normal distribution) function. In general, a bell-shaped function is a function with a quadratic exponent. The term "bell" refers to the graphical representation of a bell-shaped curve.

[0066] Known video decoders use motion estimation and motion compensation for inter-frame prediction to take advantage of temporal redundancy. Motion vectors represent the way pixels in a reference frame must be shifted to obtain a better prediction for a pixel in the current frame. This is typically done block-by-block, assigning the same motion vector to each pixel within a block. This process is often inaccurate and produces block artifacts. On the other hand, there are usually very few motion vectors to be transmitted.

[0067] Figure 1 This is a schematic diagram of an exemplary video encoder 20. The video encoder 20 can also be modified or configured to implement the techniques of this invention, as described below with reference to further figures and embodiments. Figure 1 In the example, the video encoder 20 includes an input terminal 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal 272 (or output interface 272). The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown can also be called a hybrid video encoder or a video encoder based on a hybrid video codec.

[0068] The residual calculation unit 204, transform processing unit 206, quantization unit 208, and mode selection unit 260 can form the forward signal path of the encoder 20, while the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, buffer 216, loop filter 220, decoded picture buffer (DPB) 230, inter-frame prediction unit 244, and intra-frame prediction unit 254 can form the backward signal path of the video encoder 20. The backward signal path of the video encoder 20 corresponds to the decoder (see [link to decoder]). Figure 3The signal path of the video decoder 30 in the video encoder 20. The inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter-frame prediction unit 244 and intra-frame prediction unit 254 also constitute the "built-in decoder" of the video encoder 20.

[0069] Encoder 20 can be used to receive image 17 (or image data 17) via input terminal 201, etc. Image 17 can be an image in a sequence of images that make up a video or video sequence. The received image or image data can also be a preprocessed image 19 (or preprocessed image data 19). For simplicity, image 17 will be used in the following description. Image 17 can also be referred to as the current image or the image to be decoded (especially in video decoding to distinguish the current image from other images (e.g., previously encoded and / or decoded images) in the same video sequence (i.e., a video sequence that also includes the current image).

[0070] A (digital) image is, or can be viewed as, a two-dimensional array or matrix composed of samples with intensity values. Samples in the array are also called pixels (short for image elements). The number of samples in the array or image along the horizontal and vertical directions (or axes) defines the image size and / or resolution. To represent color, three color components are typically used; that is, an image can be represented as, or may include, an array of three samples. In RBG format or color space, an image includes corresponding red, green, and blue sample arrays. However, in video decoding, each pixel is typically represented in a luminance and chrominance format or color space, such as YCbCr, including the luminance component represented by Y (sometimes also L) and two chrominance components represented by Cb and Cr. The luminance component Y represents the brightness or grayscale intensity (e.g., both are the same in grayscale images), while the two chrominance components Cb and Cr represent the chrominance or color information components. Therefore, a YCbCr format image consists of a luminance sample array composed of luminance sample values ​​(Y) and two chrominance sample arrays composed of chrominance values ​​(Cb and Cr). An RGB format image can be converted or transformed to YCbCr format, and vice versa. This process is also called color transformation or conversion. If the image is black and white, it may only include the luminance sample array. Accordingly, the image can be, for example, a black and white format luminance sample array or a 4:2:0, 4:2:2, and 4:4:4 color format luminance sample array and two corresponding chrominance sample arrays.

[0071] Embodiments of the video encoder 20 may include an image segmentation unit ( Figure 1(Not shown in the image) is used to segment image 17 into multiple (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (in H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (in H.265 / HEVC and VVC). The image segmentation unit can be used to apply the same block size and a corresponding grid with a defined block size to all images in a video sequence, or to vary the block size between images, subsets of images, or groups of images, and to segment each image into multiple corresponding blocks.

[0072] In other embodiments, the video encoder may be used to directly receive blocks 203 in image 17, such as one, several, or all of the blocks that make up image 17. Image block 203 may also be referred to as the current image block or the image block to be decoded.

[0073] Similar to image 17, image block 203 is also, or can be considered as, a two-dimensional array or matrix composed of samples with intensity values ​​(sample values), but the size of image block 203 is smaller than that of image 17. In other words, block 203 may include, for example, a single sample array (e.g., a luminance array in the case of black and white image 17, or a luminance or chrominance array in the case of a color image) or three sample arrays (e.g., one luminance array and two chrominance arrays in the case of color image 17) or any other number and / or type of array, depending on the color format used. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, a block may be, for example, an M×N (M columns × N rows) sample array, or an M×N transform coefficient array, etc. Figure 1 The embodiment of the video encoder 20 shown can be used to encode the image 17 block by block, for example, by performing encoding and prediction in blocks 203.

[0074] Figure 1 The embodiment of the video encoder 20 shown can also be used to segment and / or encode an image using slices (also referred to as video slices). An image can be segmented into one or more (typically non-overlapping) slices or encoded using one or more (typically non-overlapping) slices, each slice may include one or more blocks (e.g., CTUs). Figure 1The embodiment of the video encoder 20 shown can also be used to segment and / or encode an image using tile groups (also referred to as video tile groups) and / or blocks (also referred to as video blocks). An image can be segmented into one or more (typically non-overlapping) tile groups or encoded using one or more (typically non-overlapping) tile groups; each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more blocks; each block may, for example, be rectangular and may include one or more complete or partial blocks (e.g., CTUs), etc.

[0075] The residual calculation unit 204 can be used to calculate the residual block 205 (also referred to as residual 205) based on the image block 203 and the prediction block 265 (more detailed description of the prediction block 265 is provided later): for example, subtracting the sample value of the prediction block 265 from the sample value of the image block 203 on a sample-by-sample (pixel-by-pixel) basis to obtain the residual block 205 in the sample domain.

[0076] Transform processing unit 206 can be used to apply a discrete cosine transform (DCT) or a discrete sine transform (DST) or their integer approximations to the sample values ​​of residual block 205 to obtain transform coefficients 207 in the transform domain. Transform coefficients 207 can also be called transform residual coefficients and represent residual block 205 in the transform domain. Embodiments of video encoder 20 (correspondingly, transform processing unit 206) can be used to directly output or encode or compress transform parameters (e.g., one or more transform types) via entropy coding unit 270, so that video decoder 30 can receive and use the transform parameters for decoding, etc.

[0077] Quantization unit 208 can be used to quantize transform coefficients 207 by applying scalar quantization or vector quantization to obtain quantized coefficients 209. Quantized coefficients 209 can also be called quantized transform coefficients 209 or quantized residual coefficients 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, n-bit transform coefficients can be rounded down to m-bit transform coefficients during quantization, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different degrees of scaling can be used to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. A suitable quantization step size can be represented by the quantization parameter (QP). For example, the quantization parameter can be a set of predefined indices for suitable quantization step sizes. For example, smaller quantization parameters can correspond to fine quantization (smaller quantization step size), larger quantization parameters can correspond to coarse quantization (larger quantization step size), and vice versa. Quantization can include dividing by the quantization step size, while corresponding and / or dequantization performed by the dequantization unit 210, etc., can include multiplying by the quantization step size. Quantization is a lossy operation, where the larger the quantization step size, the greater the loss. Embodiments of the video encoder 20 (correspondingly, the quantization unit 208) can be used to directly output or encode quantization parameters (QPs) through the entropy coding unit 270, so that the video decoder 30 can receive and use the quantization parameters for decoding, etc.

[0078] The dequantization unit 210 is used to apply the dequantization of the quantized coefficients to the quantized coefficients by the quantization unit 208 to obtain the dequantized coefficients 211, for example, by applying a dequantization scheme opposite to the quantization scheme applied by the quantization unit 208, using or according to the same quantization step size as the quantization unit 208. The dequantized coefficients 211 can also be called dequantized residual coefficients 211, corresponding to the transform coefficients 207, but due to the loss caused by quantization, they are usually different from the transform coefficients.

[0079] The inverse transform processing unit 212 applies an inverse transform opposite to the transform applied by the transform processing unit 206, such as the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST) or other inverse transforms, to obtain the reconstructed residual block 213 (or the corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.

[0080] Reconstruction unit 214 (e.g., adder or summer 214) is used to add transform block 213 (i.e. reconstructed residual block 213) to prediction block 265 in such a way as to obtain reconstructed block 215 in the sample domain: for example, adding the sample values ​​of reconstructed residual block 213 to the sample values ​​of prediction block 265 one sample at a time.

[0081] Loop filter unit 220 (or simply "loop filter 220") is used to filter reconstructed block 215 to obtain filtered block 221, or generally to filter reconstructed samples to obtain filtered sample values. For example, the loop filter unit is used to smoothly perform pixel transitions or otherwise improve video quality. Although the loop filter unit 220 is used in… Figure 1 The loop filter unit 220 is shown as an in-loop filter, but in other configurations, it can be implemented as a post-loop filter. The filtered block 221 can also be referred to as the filtered reconstructed block 221.

[0082] The decoded picture buffer (DPB) 230 can be a memory that stores reference images or reference image data, which are used to encode video data by the video encoder 20. The DPB 230 can be formed from any of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer (DPB) 230 can be used to store one or more filtered blocks 221. The decoded picture buffer 230 can also be used to store other previously filtered blocks (e.g., previously filtered reconstructed blocks 221) in the same current image or different images (e.g., previous reconstructed images), and can provide previously fully reconstructed (i.e., decoded) images (and corresponding reference blocks and samples) and / or partially reconstructed current images (and corresponding reference blocks and samples) for inter-frame prediction, etc. The decoded picture buffer (DPB) 230 can also be used to: store one or more unfiltered reconstructed blocks 215 or normally store unfiltered reconstructed samples if the reconstructed block 215 is not filtered by the loop filter unit 220, or to store any other blocks or samples obtained after further processing of the reconstructed blocks or samples.

[0083] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254, and is used to receive or acquire raw image data such as the raw block 203 (current block 203 in the current image 17) and reconstructed image data (e.g., filtered and / or unfiltered reconstructed samples or blocks from the same (current) image and / or one or more previously decoded images) from the decoded image buffer 230 or other buffers (e.g., a line buffer, not shown in the figure), etc. The reconstructed image data is used as reference image data for prediction such as inter-frame prediction or intra-frame prediction to obtain prediction block 265 or prediction value 265.

[0084] The mode selection unit 260 can be used to determine or select a segmentation method for the current block prediction mode (including no segmentation) and to determine or select a prediction mode (e.g., intra-frame or inter-frame prediction mode), and generate the corresponding prediction block 265. The prediction block 265 is used to calculate the residual block 205 and reconstruct the reconstructed block 215.

[0085] An embodiment of the mode selection unit 260 can be used to select a segmentation method and a prediction mode (e.g., from those modes supported or available by the mode selection unit 260). The segmentation method and prediction mode provide an optimal match or minimum residual (minimum residual implies better compression in transmission or storage), or provide minimum signaling overhead (minimum signaling overhead implies better compression in transmission or storage), or consider or balance both. The mode selection unit 260 can be used to determine the segmentation method and prediction mode based on rate distortion optimization (RDO), i.e., selecting the prediction mode that provides minimum rate distortion. The terms "optimal," "minimum," and "best" as used herein do not necessarily refer to "best," "minimum," or "best" overall, but can also refer to situations that meet termination or selection criteria. For example, values ​​exceeding or falling below a threshold or other constraints may lead to a "suboptimal choice," but reduce complexity and processing time. In other words, the segmentation unit 262 can be used to segment block 203 into smaller block partitions or sub-blocks (reforming blocks) in the following ways: for example, iteratively using quad-tree (QT) segmentation, binary-tree (BT) segmentation, or triple-tree (TT) segmentation or any combination thereof; and to perform predictions on each block partition or sub-block, etc., wherein mode selection includes selecting the tree structure of the segmented block 203, and the prediction mode is applied to each block partition or sub-block.

[0086] As described above, the term "block" as used herein can be a portion of an image, particularly a square or rectangular portion. Referring to HEVC and VVC, a block can be or may correspond to a coding tree unit (CTU), coding unit (CU), prediction unit (PU), and transform unit (TU), and / or correspond to multiple corresponding blocks, such as a coding tree block (CTB), coding block (CB), transform block (TB), or prediction block (PB). For example, a coding tree unit (CTU) can be or may include one CTB consisting of luminance samples from an image with three sample arrays and two corresponding CTBs consisting of chrominance samples from that image, or it can be or may include one CTB consisting of samples from a black and white image or an image decoded using three separate color planes and syntax structures. These syntax structures are used to decode the aforementioned samples. Correspondingly, a coding tree block (CTB) can be an N×N sample block, where N can be set to a value such that a component is divided into multiple CTBs; this is one segmentation method. A coding unit (CU) can be or can include a coding block consisting of luminance samples from an image with three sample arrays and two corresponding coding blocks consisting of chrominance samples from the same image; or it can be or can include a coding block consisting of samples from a black and white image or an image decoded using three separate color planes and syntax structures. These syntax structures are used to decode the aforementioned samples. Correspondingly, a coding block (CB) can be an M×N sample block, where M and N can be set to a value such that a CTB is divided into multiple coding blocks; this is another segmentation method.

[0087] In an embodiment, for example according to HEVC, a coding tree unit (CTU) can be divided into multiple CUs by a quadtree structure represented as a coding tree. Whether to use inter-frame (temporal) prediction or intra-frame (spatial) prediction to decode the image region is determined at the CU level. Each CU can be further divided into one, two, or four PUs according to the PU partitioning type. The same prediction process is performed within a PU, and relevant information is transmitted to the decoder on a PU-by-PU basis. After obtaining residual blocks through the prediction process according to the PU partitioning type, the CU can be partitioned into transform units (TUs) according to other quadtree structures similar to the coding tree of that CU.

[0088] As described above, the video encoder 20 is used to determine or select the best or optimal prediction mode from (e.g., a predetermined) set of prediction modes. For example, the set of prediction modes may include intra-frame prediction modes and / or inter-frame prediction modes, etc. Specifically, mode selection may also include selecting a prediction mode according to the present invention, as detailed below with reference to specific embodiments of motion information derivation, representation, and indication.

[0089] The intra-prediction mode set may include, for example, 35 different intra-prediction modes, such as non-directional or directional modes like DC (or mean) mode and planar mode defined in HEVC, or it may include 67 different intra-prediction modes, such as non-directional or directional modes like DC (or mean) mode and planar mode defined in VVC. Intra-prediction unit 254 is used to generate intra-prediction block 265 using reconstructed samples of neighboring blocks of the same current image, based on the intra-prediction modes within the intra-prediction mode set. Intra-prediction unit 254 (or commonly referred to as mode selection unit 260) is also used to output intra-prediction parameters (or commonly referred to as information representing the selected intra-prediction mode of the block) to entropy coding unit 270 in the form of syntax element 266 to include the intra-prediction parameters in the encoded image data 21, so that video decoder 30 can receive and use the prediction parameters for decoding, etc.

[0090] The set of possible inter-frame prediction modes depends on the available reference image (i.e., at least a portion of the decoded image stored in the DBP 230, etc.) and other inter-frame prediction parameters, such as whether the entire reference image or only a portion of the reference image (e.g., a search window region around the current block's region) is used to search for the best-matching reference block, and / or whether pixel interpolation is performed, such as half-pixel interpolation and / or quarter-pixel interpolation. Inter-frame prediction modes may include modes that operate in conjunction with motion field determination and representation, as described in the embodiments below. Such a mode may be one of several inter-frame modes.

[0091] In addition to the prediction modes mentioned above, skip mode and / or direct mode can also be used.

[0092] Inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in... Figure 2(Not shown in the image). The motion estimation unit can be used to receive or acquire image block 203 (current image block 203 in current image 17) and decoded image 231, or at least one or more previous reconstructed blocks (e.g., reconstructed blocks in one or more other / different previous decoded images 231) for motion estimation. For example, the video sequence may include the current image and the previous decoded image 231, or in other words, the current image and the previous decoded image 231 may be part of or constitute an image sequence that forms a video sequence.

[0093] For example, encoder 20 can be used to select a reference block from multiple reference blocks of the same or different images in multiple other images, and provide the motion estimation unit with the offset (spatial offset) between the position (x-coordinate, y-coordinate) of the reference image (or reference image index) and / or the position of the reference block and the position of the current block as inter-frame prediction parameters. This offset is also called a motion vector (MV). As shown in some detailed embodiments, motion information in some inter-frame prediction modes does not necessarily need to be provided block by block. Motion information may include motion vectors and possible reference images, which are usually indicated separately from the motion vectors. However, in general, the reference image may be a part of the motion vector; for example, in addition to the two spatial components of the motion vector, the reference image may be a third (temporal) component.

[0094] The motion compensation unit is used to acquire (e.g., receive) inter-frame prediction parameters and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain inter-frame prediction blocks 265, or, in general, to perform prediction on some samples in the current image. Motion compensation performed by the motion compensation unit may include extracting or generating prediction blocks (prediction samples) based on motion / block vectors determined by motion estimation, and may also include performing interpolation to achieve sub-pixel accuracy. Interpolation filtering can generate other pixel samples based on known pixel samples, potentially increasing the number of candidate prediction blocks / samples that can be used to decode image blocks. Upon receiving the motion vector corresponding to the PU of the current image block, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists. The motion compensation unit can also generate syntax elements associated with blocks, sample regions, and video strips for use by the video decoder 30 when decoding image blocks in the video strips. In addition to strips and corresponding syntax elements, or as an alternative to strips and corresponding syntax elements, block groups and / or blocks and their corresponding syntax elements can be generated or used.

[0095] Entropy coding unit 270 is used to apply or not apply entropy coding algorithms or schemes (e.g., variable length coding (VLC), context adaptive VLC (CAVLC), arithmetic coding schemes, binary coding, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to (uncompressed) quantized coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements to obtain encoded image data 21 that can be output through output terminal 272 in the form of encoded bitstream 21, etc., so that video decoder 30 can receive and use these parameters for decoding, etc. The encoded bitstream 21 can be transmitted to video decoder 30 or stored in memory for later transmission or retrieval by video decoder 30.

[0096] Other structural variations of the video encoder 20 can be used to encode video streams. For example, a non-transform-based encoder 20 can directly quantize residual signals for certain blocks or frames without a transform processing unit 206. In another implementation, the encoder 20 may include a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.

[0097] Figure 2 An example of a video decoder 30 that can be modified or configured to implement the technology of the present invention is shown. The video decoder 30 is used to receive, for example, encoded image data 21 (e.g., encoded bitstream 21) encoded by encoder 20 to obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, such as data and associated syntax elements representing image blocks of encoded video stripes (and / or chunks or blocks).

[0098] exist Figure 2In the example, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter-frame prediction unit 344, and an intra-frame prediction unit 354. The inter-frame prediction unit 344 may be or may include a motion compensation unit. In some examples, video decoder 30 may perform substantially the same functions as the reference unit. Figure 1 The video encoder 100 described in the text is the inverse of the encoding process and the decoding process.

[0099] As described with reference to encoder 20, the inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter-frame prediction unit 344, and intra-frame prediction unit 354 also constitute the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 can be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 can be functionally identical to inverse transform processing unit 212, reconstruction unit 314 can be functionally identical to reconstruction unit 214, loop filter 320 can be functionally identical to loop filter 220, and decoded picture buffer 330 can be functionally identical to decoded picture buffer 230. Therefore, the explanation of the corresponding units and functions of video encoder 20 is correspondingly applicable to the corresponding units and functions of video decoder 30.

[0100] Entropy decoding unit 304 is used to parse bitstream 21 (or commonly referred to as encoded image data 21) and perform entropy decoding on encoded image data 21 to obtain quantization coefficients 309 and / or decoded encoding parameters. Figure 3 (Not shown in the image) For example, inter-frame prediction parameters (e.g., reference image index and motion vector), intra-frame prediction parameters (e.g., intra-frame prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements, etc. The entropy decoding unit 304 can be used to apply a decoding algorithm or scheme corresponding to the encoding scheme described by the entropy coding unit 270 in the reference encoder 20. The entropy decoding unit 304 can also be used to provide inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360, and to provide other parameters to other units in the decoder 30. The video decoder 30 can receive video strip-level and / or video block-level syntax elements. In addition to stripes and corresponding syntax elements, or as a substitute for stripes and corresponding syntax elements, it can receive and / or use chunk groups and / or chunks and corresponding syntax elements.

[0101] The dequantization unit 310 can be used to receive quantization parameters (QP) (or commonly referred to as dequantization-related information) and quantization coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.), and dequantize the decoded quantization coefficients 309 according to these quantization parameters to obtain dequantized coefficients 311. The dequantized coefficients 311 can also be referred to as transform coefficients 311. The dequantization process may include determining the degree of quantization using the quantization parameters determined by the video encoder 20 for each video block in a video strip (or chunk or group of chunks), and also determining the degree of dequantization to be applied.

[0102] The inverse transform processing unit 312 can be used to receive the dequantized coefficients 311 (also called transform coefficients 311) and transform the dequantized coefficients 311 to obtain the reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 can also be called transform block 313. The transform can be an inverse transform, such as inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be used (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.) to receive transform parameters or corresponding information from the encoded image data 21 to determine the transform to be applied to the dequantized coefficients 311.

[0103] Reconstruction unit 314 (e.g., adder or summer 314) can be used to add reconstructed residual block 313 to prediction block 365 to obtain reconstructed block 315 in the sample domain by, for example, adding the sample value of reconstructed residual block 313 to the sample value of prediction block 365.

[0104] Loop filter unit 320 (in or after the decoding loop) is used to filter the reconstructed block 315 to obtain the filtered block 321, thereby facilitating pixel transformation or otherwise improving video quality, etc. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, a co-filter, or any combination thereof. In some configurations, loop filter unit 320 may be implemented as a post-loop filter.

[0105] The decoded video block 321 from one image is then stored in the decoded image buffer 330. The decoded image buffer 330 stores the decoded image 331 as a reference image for subsequent motion compensation and / or output or display of other images. The decoder 30 is used to output the decoded image 311 through the output terminal 312, etc., to present to the user or for the user to view.

[0106] Inter-frame prediction unit 344 may be functionally identical to inter-frame prediction unit 244 (particularly to motion compensation unit), and intra-frame prediction unit 354 may be functionally identical to inter-frame prediction unit 254. Both units perform segmentation or partitioning decisions and execute predictions based on (e.g., parsed and / or decoded by entropy decoding unit 304, etc.) the segmentation method and / or prediction parameters or corresponding information received from the encoded image data 21. Pattern application unit 360 may be used to perform predictions (intra-frame or inter-frame predictions) on a block or sample basis based on the reconstructed image, blocks, or corresponding samples (filtered or unfiltered) to obtain prediction blocks 365.

[0107] When a video strip is decoded into an intra-decoded (I) strip, the intra-prediction unit 354 in the mode application unit 360 generates prediction blocks 365 for image blocks in the current video strip based on the indicated intra-prediction mode and data from previously decoded blocks in the current image. When a video image is decoded into an inter-decoded (e.g., B or P) strip, the inter-prediction unit 344 (e.g., motion compensation unit) in the mode application unit 360 generates prediction blocks 365 for video blocks in the current video strip based on motion vectors and other syntax elements received from the entropy decoding unit 304. For inter-prediction, these prediction blocks can be generated based on one of the reference images in one of the reference image lists. The video decoder 30 can construct reference frame lists 0 and 1 using the default construction technique based on the reference images stored in the DPB 330. The same or similar processes can be applied to or applied by embodiments using chunk groups (e.g., video chunk groups) and / or chunks (e.g., video chunks) in addition to or as an alternative to stripes (e.g., video stripes). For example, video can be decoded using I, P, or B blocks and / or chunks.

[0108] The pattern application unit 360 is used to determine prediction information for video blocks in the current video stripe by parsing motion vectors or related information and other syntax elements, and to generate prediction blocks for the current video block being decoded using the prediction information. For example, the pattern application unit 360 uses some received syntax elements to determine the prediction mode (e.g., intra-frame or inter-frame prediction) for decoding video blocks in the video stripe, the inter-frame prediction stripe type (e.g., B-strip, P-strip, or GPB-strip), construction information for one or more reference image lists for the stripe, motion vectors for each inter-frame coded video block of the stripe, inter-frame prediction state for each inter-frame decoded video block of the stripe, and other information to decode video blocks in the current video stripe. In addition to or as an alternative to stripes (e.g., video stripes), the same or similar processes can be applied to or applied by embodiments using chunk groups (e.g., video chunk groups) and / or chunks (e.g., video chunks). For example, video can be decoded using I, P, or B chunk groups and / or chunks.

[0109] Figure 2 The embodiment of the video decoder 30 shown can be used to segment and / or decode images by stripes (also referred to as video stripes). An image can be segmented into one or more (typically non-overlapping) stripes or decoded using one or more (typically non-overlapping) stripes, each stripe may include one or more blocks (e.g., CTUs).

[0110] Figure 2 The embodiment of the video decoder 30 shown can be used to segment and / or decode an image using chunk groups (also called video chunk groups) and / or chunks (also called video chunks). An image can be segmented into one or more (typically non-overlapping) chunk groups or decoded using one or more (typically non-overlapping) chunk groups; each chunk group may include one or more blocks (e.g., CTUs) or one or more chunks, etc.; each chunk may be rectangular, etc., and may include one or more complete or partial blocks (e.g., CTUs, etc.).

[0111] Other variations of the video decoder 30 can be used to decode the encoded image data 21. For example, the decoder 30 can generate an output video stream without the loop filter unit 320. For example, the non-transform-based decoder 30 can directly dequantize the residual signal for certain blocks or frames without the inverse transform processing unit 312. In another implementation, the video decoder 30 may include a dequantization unit 310 and an inverse transform processing unit 312 combined into a single unit.

[0112] Some optical flow algorithms generate dense motion fields. These fields consist of numerous motion vectors, with each pixel in the image corresponding to one motion vector. Predictions using these motion fields typically yield good prediction quality. However, since the dense motion field contains as many motion vectors as the image contains pixels, the total representation to be transmitted or stored is considerable. Therefore, the dense motion field must be subsampled and quantized to reduce the amount of data to be transmitted / stored. The decoder then interpolates the missing motion vectors and uses the reconstructed dense motion field for motion compensation.

[0113] In most cases, a sparse representation of a motion field can maintain a low bitrate, but this sparse representation can be obtained using some interpolation technique. The motion field can be subsampled using spatial cells / blocks with the same region in a regular pattern (e.g., a regular grid with uniformly distributed nodes, as used for JPEG image compression), regardless of the content. Motion vectors do not need to be interpolated within these regions; that is, all pixels have the same motion vector and are shifted together. This results in many sampled points being located in suboptimal positions. Regions with uniform motion that require few motion vectors include the same number of motion vectors as regions with different motion that require many support points. This increases the bitrate of the residual data (greater than the desired bitrate), leading to poor prediction quality due to the need for more motion vectors.

[0114] Another approach is to transmit the parameters of the higher-order motion model, and only at locations where a good reconstruction of the flow field is necessary. This way, regions with uniform motion do not require a high rate, while regions with complex motion are sampled densely enough. However, since only the encoder knows the entire motion field, the locations must be indicated in some way. When we refer to transmission in this paper, it generally means transferring information from the encoder to one or more possible decoders. This is typically done by carrying side information in the bitstream provided along with the encoded data (image / video data).

[0115] Some video codecs perform implicit subsampling using block-based motion estimation and compensation. Modern codecs (such as HEVC or VVC) to some extent use different block sizes for content-adaptive sampling. These codecs explicitly indicate block partitioning as quadtrees and ternary trees, as shown above. Figure 1 and Figure 2 As shown, increasingly adaptive partitioning has been established, leading to significant improvements in the coding efficiency of corresponding codecs. AVC and HEVC codecs use simple translation models, while recently developed codecs (such as VVC, EVC, and AV1) employ higher-order affine transformations with up to six parameters per prediction unit (block).

[0116] Figure 3An exemplary non-translational motion model is shown. (From left to right in the figures) Specifically, rotation, scaling, and combinations of rotation, scaling, and translation are illustrated. In particular, these three images show the dense motion fields of the corresponding transformations.

[0117] It should be noted that increasing the order of the motion model may disrupt the balance between rate (indicating parameters of the motion model) and distortion (which is reduced due to better predictions resulting from motion compensation). Figure 4 The diagram illustrates the trend of the number of parameters to be indicated (y-axis) increasing with the order of the motion model (x-axis). The maximum representation of intra-frame motion corresponds to optical flow, causing the number of parameters to increase to twice the number of pixels in a high-precision frame (the x and y components of the motion vector in the 2D motion vector space) (the data range required to represent the MV components). It should be noted that... Figure 4 This is merely an illustrative illustration, intended to roughly compare the impact of motion precision representation on the number of parameters that must be provided along with the encoded video.

[0118] Non-block-based motion compensation is rarely used in video decoding. The main reason is that the entire motion field must be transmitted. Natural video does not represent a linear motion model. Only certain regions of a frame can usually be described using a simple model, for which the aforementioned codec (see above) provides an example. Figure 1 and Figure 2 This could be efficient. On the other hand, more complex nonlinear motion models could increase the amount of indication, which is currently indicated by segmentation, motion model parameters, and the residuals of the corresponding predictions. Existing solutions that support nonlinear motion features of natural content may require high bitrate overhead for indication and may also introduce some block artifacts at block boundaries.

[0119] Some embodiments of the present invention can use more complex motion models to represent larger regions, but with fewer parameters to describe them. The parameters can be easily predicted from the optical flow available at the encoder, without using most well-known, complex rate-distortion optimization (RDO) methods. When the motion model is simple, more subsampling can be performed. More complex motion models can be used to describe motion within larger regions of the predicted frame, which reduces signaling overhead.

[0120] Specifically, optical flow is reconstructed using a sparse representation based on a set of affine transformations. Specifically, the affine transformations can be constant shape transformations.

[0121] Figure 5 This is a block diagram of an encoder provided in one embodiment of the present invention. In this embodiment, the method for increasing the density of the sports field is part of motion compensation. Figure 5The encoder in the above Figure 1 The encoder 10 shown has some similarities. The motion estimation and motion compensation described herein can be similarly applied. Figure 1 encoder and Figure 2 The decoder in the image, for example, as a specific inter-frame prediction mode.

[0122] exist Figure 5 In this process, either the intra-frame prediction module 520 or the inter-frame prediction module 530 predicts the input video frame 510 (also referred to as the current frame). The intra-frame prediction module 520 generates a predicted frame, which is then subtracted from the input video frame 510 to obtain a residual signal 550. Furthermore, the predicted frame is also used for reconstruction, i.e., added to the reconstructed residual. Additionally, the intra-frame prediction module 520 generates side information that is sent to the entropy encoder 570, which generates an output (encoded) bitstream 575. The residual frame 550, corresponding to the difference between the predicted frame and the current frame, passes through the quality control module 560. The quality control module 560 may include transformation (e.g., transformation to the spectral domain) and / or quantization. Quantization may perform lossy compression, resulting in a certain degree of quality degradation. Simplifying the residual signal in other ways to obtain a compact representation in the bitstream 575 may incur further losses.

[0123] The inter-frame prediction module 530 comprises two sub-modules: a motion estimation (ME) sub-module 532 and motion compensation (MC) sub-modules 534 to 540. The purpose of the ME sub-module 532 is to find the most suitable parameters for the motion model (defined and used in the video codec) and provide these parameters to the entropy encoder 570 so that they can be included in the bitstream as side information. The side information inserted into the bitstream can typically include various other parameters. For example, the side information may carry segmentation / segmentation information, motion vector fields 534, and other side information such as weighting or other additional control parameters. The purpose of the side information is to transmit parameters to the decoder to assist the decoder in performing reconstruction. Using the parameters of the motion model and reference frames in the decoded pictures buffer (DPB) 595, the decoder can reconstruct the encoded image in exactly the same manner as the encoded image.

[0124] Codec 500 uses a sparse representation of the motion vector field 534 to reduce signaling overhead. This can be achieved through segmentation, subsampling, segmentation, and corresponding motion model parameters. If the motion model has more parameters than the translation model (e.g., six parameters are used in some known codecs), interpolation is required to obtain a dense motion field (538, 540) associated with a specific reference frame and corresponding to the prediction frame. In some embodiments of the invention, densification 538 is part of the motion compensation process. Densification 538, using a sparse representation of the dense motion vector field 534 as input 536 and providing a dense motion vector field as output 540, can provide suitable motion estimation using a finite number of parameters.

[0125] The following text refers to Figures 6 to 11 Describe a specific, detailed example of how this motion estimation is implemented.

[0126] The parameters known on the encoder and decoder sides in this exemplary implementation are listed below. These parameters can be fixed in advance (i.e., predefined by a standard) or transmitted (transmitted in the bitstream).

[0127] S s A segment with index 's'. For example, an image to be encoded can be divided into different segments, each of which can be assigned an index. Generally, a segment is any defined set of samples. A segment can be a prediction unit of any shape, such as a rectangle, triangle, ellipse, hexagon, etc. It should be noted that, for the purpose of this invention, segments do not necessarily have to be continuous. Spatially disconnected segments can be used. These segments can represent objects. However, it should be noted that this invention is also directly applicable to unsegmented images.

[0128] N s The decoding end anticipates the number of motion vectors for the slice with index S. It should be noted that the number N... S It doesn't necessarily depend on S, meaning it's not necessarily specific to a piece. For example, the number of motion vectors can be defined as N, meaning all pieces have the same number of motion vectors.

[0129] P s,i Segmentation S s The position of the i-th motion vector (Px) s,i ,Py s,i ), where i = 0…N s .

[0130] MV s,i With S-shards s The i-th motion vector (Vx) in the corresponding list of motion vectors s,i Vy s,i ), where i = 0…N sHere, "corresponding to" means "belonging to" or "associated with".

[0131] w s,j (x,y) is the weighting function for the j-th pair of motion vectors, where j = 0…N s -1. The weighting function can depend on the position (x, y) in the dense motion field. Alternatively, the weighting function can depend on the position j in a list (or pair) of motion vectors, which can be used to derive the weights. Different slices can have different weighting functions. For example, the weighting function can depend on the size of the slice, such as the horizontal and / or vertical dimensions of the slice, or on the number of image samples in each slice.

[0132] a s,j σ s,j c s,j The parameters of the weighting function. In some embodiments, the weighting function is a nonlinear function. For example, the weighting function could be a Gaussian distribution function. It should be noted that these parameters are merely examples. Weighted functions can have more or fewer such parameters.

[0133] Figure 6 This is a flowchart of an example of dense motion field estimation based on the above parameters.

[0134] In step 610, the initialization method is performed. For each of the M fragments S... S The method is executed starting from S = 0. In step 615, index j is initialized to j = 0. In this particular example, index j passes through N of the current shard (the shard being processed with index S). S -1 motion vector.

[0135] Steps 620 to 640 construct a loop with index j on the motion vectors in a set of sparse motion vectors (for the current slice S with index s). In each iteration of the loop, in step 620, the j-th motion vector is obtained, which is determined by its starting position (Px). s,i ,Py s,i ) and the motion vector (Vx) associated with that location s,i Vy s,i The parameter 's' is specified. Additionally, weighting parameters are retrieved, which specify the weighting function for the current partition S at index 's'. Specifically, parameter 'a' is retrieved. s,j σ s,j c s,j It should be noted that weighting parameters can be specified for the current slice S and the current motion vector j. In step 630, it is checked whether the method has been applied to all N (generally N). SThe method iterates through the j-th motion vector. If no (yes in step 630), the method proceeds to step 640. In step 640, the affine parameter affX of the j-th motion vector of the current slice is derived. s,j and affY s,j Now, let's explain step 640 in detail.

[0136] Figure 7 An example of a restricted affine transformation 700 representing a combination of scaling, rotation, and some translational motion is shown. The term "restricted" here means that this exemplary transformation can only represent the three motions mentioned above (scaling, rotation, and translation). However, the invention is not limited thereto; in general, different numbers and types of transformations can be performed.

[0137] In sub-image (a), the starting position P is shown. s,0 (701) and the first motion vector MV of the motion displacement toward position 705 s,0 (703), this vector represents the origin of the transformation (i.e., the motion vector to be transformed). Second motion vector MV s,1 (713) has a starting position P s,1 (710) and the movement displacement to position 715.

[0138] These two motion vectors can completely define the affine transformation of space while preserving the shape of the object transformed by this transformation. In this example, from Figure 7 As can be seen in sub-image (b), rectangle 720 is transformed into rectangle 730 while maintaining the same aspect ratio of the rectangle sides. In other words, for the original triangle 720 and the transformed triangle 730, the aspect ratio of the rectangle sides is the same. From sub-image (b), it can be seen that the two initial positions P... s,0 (701) and P s,1 (710) Connecting line 725 is linearly transformed to connecting line 735, which connects target positions 705 and 715 after the transformation. In this specific case, motion vector 703 is the j-th motion vector, and motion vector 713 is the (j+1)-th motion vector of the s-th segment; these two vectors together form the j-th pair of motion vectors. The distance between position 701 of motion vector 703 and position 710 of motion vector 713 is determined by its x-component px. s,j (x) and y components py s,j (y) is given, as shown in the following expression:

[0139]

Expression 1

[0140]

Expression 2

[0141] The j-th pair of motion vectors MV s,j and MV s,j+1 The parameters of the affine transformation affX s,j and affY s,j It can then be written as:

[0142]

Expression 3

[0143] affX s,j =((Vx) s,j+1 ·px s,j +Vy s,j+1 ·py s,j )-(Vx s,j ·px s,j +Vy s,j ·py s,j )) / (px s,j 2 +py s,j 2 )

[0144]

Expression 4

[0145] affY s,j =((Vx) s,j+1 ·py s,j -Vy s,j+1 ·px s,j )-(Vx s,j ·py s,j -Vy s,j ·px s,j )) / (px s,j 2 +py s,j 2 )

[0146] Affine transformations such as Figure 7 The sub-image (c) is shown in the diagram. Specifically, dashed line 740 illustrates how some points in rectangle 720 are transformed into corresponding points in rectangle 730, so line 740 corresponds to motion vectors that can now be derived at any point via an affine transformation. Here, it is assumed that a predefined order exists, and a pair of motion vectors is a pair consisting of the j-th motion vector and the (j+1)-th motion vector. However, in general, j can be replaced by i, and j+1 can be replaced by j, representing a pair of motion vectors consisting of the i-th motion vector and the j-th motion vector. The spatial transformation is then performed by an affine transformation (affX) calculated for the first pair of motion vectors (j=0). s,0 ,affY s,0The motion vector (mvx) is completely determined (both inside and outside the rectangle). In other words, the interpolated motion vector is... s,0 (x,y),mvy s,0 (x,y))740 can be derived at any location in 2D space.

[0147] Generally speaking, (mvx) s,j (x,y),mvy s,j (x,y) is obtained by partitioning S s The components of the 2D motion vector obtained by performing an affine transformation on a point at position (x, y) represented by the j-th pair of motion vectors are shown below:

[0148]

Expression 5

[0149] mvx s,j (x,y)=Vx s,j +dx s,j (x,y)·affX s,j +dy s,j (x,y)·affY s,j

[0150]

Expression 6

[0151] mvy s,j (x,y)=Vy s,j +dx s,j (x,y)·(-affY s,j )+dy s,j (x,y)·affX s,j

[0152] In addition, dx s,j (x,y) and dy s,j (x, y) is the position P of point (x, y) and the first motion vector in the j-th pair of motion vectors. s,j Components of the distance between:

[0153]

Expression 7

[0154]

Expression 8

[0155] In this example, the restricted affine transformation supports combinations of scaling, rotation, and translation, and is represented by expressions 5 and 6. It should be noted that, in general, this invention is not limited to any particular transformation. The restricted affine transformation described herein is merely an example. In general, simpler or more complex transformations can be performed, such as transformations that support some form of nonlinear motion, and so on.

[0156] It should be noted that, Figure 7 The special case of j=0 is shown. Figure 8 Another example is shown where j=1 (second iteration). Specifically, Figure 8 An exemplary combination of two restricted affine transformations defined by three motion vectors is shown, namely the first transformation (affX) determined when j=0. s,j ,affY s,j The second transformation is determined when j=1.

[0157] exist Figure 8 In the middle, the additional motion vector MV s,2 (803) and motion vector MV s,1 Together they form a second pair of motion vectors (MV) s,1 ,MV s,2 Motion Vector MV s,2 (803) has a starting point P s,2 (801), and defines the movement displacement towards position 805. Corresponding to Figure 7 The motion vector MV of motion vector 713 in s,1 (813) Starting from position 810, and defining the displacement towards position 815. These two motion vectors (the second pair of motion vectors) can completely define another (second type) restricted affine transformation of space. Another rectangle 820 is transformed into a rectangle 830 with the same aspect ratio.

[0158] In this specific case, the first affine transformation (affX) s,0 ,affY s,0 The target rectangle 730 (in) Figure 8 The affine transformation (represented as 820) is the second type of affine transformation (affX). s,1 ,affY s,1 The original rectangle 820 is 840. (e.g.) Figure 8 As shown, the target rectangle 830 for the second transformation is obtained. Two starting positions P s,1 (810) and P s,2 The connecting line 825 of (801) is linearly transformed into line 835 connecting the new positions 815 and 805 after the second affine transformation 840. The interpolated motion vector (mvx) s,1 (x,y),mvy s,1 (x,y))(840) can be derived at any location in 2D space.

[0159] In the loops from steps 620 to 640, the affine transformation parameters (affX) are determined for all pairs of motion vectors. s,j ,affY s,jThen, if there are no more vector pairs in the slice ("No" in step 630), the method proceeds to step 650.

[0160] Figure 9 Two juxtaposed affine transformations in a 2D space are illustrated. Three motion vectors 910, 911, and 912 (corresponding to two pairs of motion vectors) specify the two restricted affine transformations. Here, the term "juxtaposed" means that the transformation is restricted to motion vectors of the same piece. Therefore, the motion vector at each point of the piece can be obtained through any juxtaposed transformation. The motion vector at a point obtained through different transformations may be different. Figure 9 As shown, two motion vectors (e.g., 940 and 945) are associated with each sample location in 2D space (e.g., exemplary location 950). The first motion vector 940 at point 950 is interpolated using a first affine transformation, while the second motion vector 945 is interpolated using a second affine transformation. To obtain a predicted value for a sample at location 950, a sample from a reference frame should be used, thus requiring a motion vector to be provided to capture that sample. To obtain each motion vector, and given two or more motion vectors, a normalized weighted sum is calculated for each interpolated MV location, as follows: Figure 10 As shown.

[0161] Figure 10 Two juxtaposed affine transformations are illustrated. Isolines are used to describe the values ​​of the weighting factors used to weight the two contributing motion vectors obtained through the corresponding two affine transformations. The term "isoline" refers to a line where the weighting factors have the same value. The term "contribution" here refers to the contribution (weighted using the corresponding weights) made by each of these motion vectors to the final motion vector at position 950. Specifically, the two attraction points P... s,0 (1001) and P s,1(1025) is located at the starting point of the first motion vector in each pair of motion vectors. The two attraction points 1001 and 1025 correspond to the center of the corresponding 2D "bell"-shaped function depicted by contour lines. Contour lines 1071, 1072, 1073, and 1074 correspond to the attraction point 1001 of the first pair of motion vectors, while contour lines 1061, 1062, 1063, and 1064 correspond to the attraction point 1025 of the second pair of motion vectors. It can be seen that the weighting factor used to weight motion vector 1045 is smaller than the weighting factor of motion vector 1040 because, according to the distance from the center of the corresponding attraction position (also called the attraction location), the contour line 1072 corresponding to the weighting factor is farther than the contour line of the other interpolated motion vector 1040. In other words, in this example, the motion vector at the target point 950 is a weighted sum of the contributing motion vectors. The contributing motion vector at the target point 950 is obtained through a predetermined transformation of the motion vectors with specific control. Different contributing motion vectors are obtained through (different) transformations of different corresponding control motion vectors. Here, the term "different transformations" refers to the same transformation rule with different parameters. However, the invention is not limited to this, and different transformations with corresponding different transformation rules can be performed. The weighting function shown in some embodiments is nonlinear. If the first motion vector is obtained by transforming the third motion vector, and the third motion vector is closer to the target position than the fourth motion vector that produces the second motion vector, then the weight of the first contributing motion vector is greater than the weight of the second motion vector. This nonlinearity makes the weight of the first contributing motion vector greater than the linear distribution of the weights between the first and second motion vectors. One example of such a nonlinear function is the Gaussian distribution function. However, other functions, such as cosine, Laplace distribution, etc., can be used. The sum of the weights of all contributing motion vectors (two in this example, but more than two in some embodiments) reaches 1 (e.g., the same as the probability distribution function).

[0162] exist Figure 6 In the loops of steps 650 and 660, a dense motion field is interpolated for each point in the juxtaposed 2D space. Specifically, in step 650, it is checked whether the current segment S has already been interpolated. S Interpolation is performed at each position (x, y) in the current segment. If no ("No" in step 650), the loop continues by interpolating the current position (x, y) for step 660. By performing step 660 on all positions (x, y) in the current segment, a dense motion field is obtained.

[0163] Expressions 9 and 10 below show an example of deriving the dense motion vector field (MVF) corresponding to step 660 performed on position (x,y):

[0164]

Expression 9

[0165]

[0166]

Expression 10

[0167]

[0168] In this specific example, the weight w s,j (x, y) is determined by a bell-shaped function, such as the Gaussian distribution function shown in expression 11:

[0169]

Expression 11

[0170]

[0171] Where, d s,j (x,y) represents the position P of point (x,y) and the first (origin) motion vector in the j-th pair of motion vectors. s,j Distance between:

[0172]

Expression 11

[0173]

[0174] Here, for example, k = 2. However, the invention is not limited to Euclidean distance (using the square norm). Instead, k can be 1 or 3, or any other metric. It should also be noted that the bell curve does not necessarily originate from circular contour lines; contour lines may form ellipses. Furthermore, the invention is not limited to using bell curves, but other functions can be applied to the weight distribution determined in the piecewise space. These other functions can be nonlinear to highlight closer areas or origins around the first motion vector. However, the invention can use other functions.

[0175] After performing a loop on all positions (x, y) in step 660, the dense motion vectors are obtained in step 670. It should be noted that... Figure 6Step 670 in the loop shows the loop result and may include storing the dense motion field in memory or other storage. However, step 670 does not need to appear explicitly in the method because step 660 in the loop already provides the dense motion field. In step 680, it is checked whether all slices have been processed. If there are still some slices to be processed (the "yes" in step 680 indicates that not all segments of the M slices have been processed), the loop on the slice continues to call 615, and the following steps are performed again as described above. If all assignments in the image have been processed, the method ends. It should be noted that, in principle, not all slices need to be processed when processing slices in an image (frame). For example, in some methods and applications, it may be desirable to provide a dense motion vector representation only for a portion of the image. Figure 5 As shown and referenced Figure 5 The encoder described can select to use the dense motion field representation / reconstruction only for some parts of the image (e.g., as a prediction mode). Other parts of the image can be processed (encoded) using different prediction modes, or not predicted at all.

[0176] An example of image segmentation is multi-reference inter-frame prediction. Therefore, one or more parts of an image are predicted from one reference image, while other parts are predicted from another reference image. A special case of multi-reference prediction is bidirectional image prediction. For example, after different transformations based on different reference frames, the same image segment can be predicted as a weighted sum of predicted samples (in the sample value domain). In other words, the present invention also applies to multi-reference prediction cases.

[0177] Figure 11 An example of weighting two motion vectors 1140 and 1145 is shown. Motion vector 1140 is weighted to motion vector 1141, with a weight greater than that of motion vector 1145. Motion vector 1145 is then weighted to motion vector 1146. Finally, the sum of the two weighted MVs 1141 and 1146 is represented by the resulting motion vector 1148.

[0178] In other words, in some embodiments of the invention, motion vectors are interpolated between available control points having control motion vectors. Here, the terms "control" point and "control" motion vector refer to positions and motion vectors originating from those positions, which are the inputs to the aforementioned dense motion field reconstruction. For example, control positions and control motion vectors can be identified on the encoder side and interpolated into the bitstream provided to the decoder to reconstruct the motion field. In other words, control positions and control motion vectors control the construction (reconstruction) of the dense motion field, that is, the construction (reconstruction) of any other motion vectors in the slice (motion vectors at any other position within the slice).

[0179] The embodiments of the present invention described above illustrate the construction of a motion field (approximate optical flow) based on control position and control motion vectors. This construction can be readily used... Figure 1 or Figure 5 The encoder shown or Figure 2 The decoder shown is an example. However, the application of this invention is not limited to video encoding and decoding. Any application requiring efficient storage or transmission of optical streams (sports fields) can utilize embodiments of this invention.

[0180] The following processing components can be derived using motion vectors:

[0181] – The motion model is based on the optical flow of slices (which can be called prediction units when used as prediction modes in video encoding / decoding). This is achieved by deriving a transformation (the affine transformation in the previous example) from the control position and the control motion vector starting from the control position.

[0182] – A small number of control points are assigned to each slice. This can be viewed as subsampling of the optical flow. How subsampling is performed is not important for the implementation of this invention. One advantage of the embodiments and examples of this invention is that any regular or irregular subsampling can be used. Specifically, subsampling does not have to follow any regular pattern, and different numbers and / or locations of control points can be provided within each slice. This allows subsampling to adapt to the content, thereby further improving compression efficiency by enhancing the reconstruction quality (rate-distortion relationship) provided for a given bitstream size.

[0183] – The nonlinear motion vector (i.e. the motion vector used for nonlinear motion) can be derived from the optical flow and is associated with each point (sample location).

[0184] – In some embodiments, each pair of consecutive motion vectors (MV0, MV1), (MV1, MV2), etc., can determine the affine optical flow. Here, MV i It has a starting position P s,i and its incremental MV s,I The motion vector, where i = 0…N s (N may depend on the size of the slices). This is a concrete example where the motion vectors are ordered, and the target motion vector of one transformation is the origin motion vector of another transformation. However, the invention is not limited to this; affine transformations can be derived using any (and unrelated) pairs of motion vectors. Furthermore, the transformation does not necessarily need to be determined from a pair of motion vectors. Instead, the transformation can be given in another way and / or associated with only a single motion vector. This may be advantageous for simple types of motion (e.g., translational motion, etc.).

[0185] – The weighted sum of several affine motion fields provides the final representation of the dense motion vector field. In other words, each affine transformation specifies the motion vector at any location within a piece. Several affine transformations result in several corresponding motion fields. To obtain a final motion field, these motion fields are weighted. There can be two or more such motion fields corresponding to the respective affine transformations.

[0186] – A bell-shaped (Gaussian) distribution can be used to determine the weights of the corresponding motion field. The parameters of the Gaussian distribution function can be indicated or determined by an encoder and decoder that follow the same rules (e.g., unsupervised learning, such as Gaussian Mixture Models, GMMs).

[0187] One embodiment takes a weighted sum of two or more motion vectors at each sampling location. The parameters can be adjusted based on the image content. The "bell"-shaped function a s,j The intensity can increase the nonlinear effect on a specific pair of motion vectors. Parameter σ s,j It has an impact on distribution and diffusion and c s,j It is a linear parameter, so diffusion can still be controlled, but in a linear manner.

[0188] When storing or indicating parameters for constructing (reconstructing) a motion field, it is best to ensure that the encoder and decoder use the same method to parse (syntactic) and interpret (semantic) the parameters. This can be achieved through some favorable syntactic and / or semantic rules. For example, based on predetermined interpretation rules, a list of motion vectors can be provided, and the reconstruction method can then proceed as follows.

[0189] – Use pairs of motion vectors consecutively, such as (MV0, MV1), then (MV1, MV2), and so on. This has been illustrated above, where the target MV of the first transformation is the origin MV of the second transformation.

[0190] – To use paired motion vectors (MV0, MV1), (MV2, MV3), etc. independently (this is equivalent to setting the weight of every second vector to 0 in the previous options), where MV i Yes (P) s,i ,MV s,i ).

[0191] It should be noted that in some exemplary implementations, there may be one (or more) pairs of motions, each with one motion vector in the first slice and another in the second slice, where the first and second slices are different. This approach creates dependencies between slices. Such dependencies do not exist in other implementations.

[0192] – In order to support translational motion models by using only one motion vector in a specific slice / cell, N s It can be equal to 1. In this case, only one pair of MVs can be generated by copying the first MV, and their positions have a small shift of (1,0)+P. s,i Then, the first method described above can be used. Generally, a position and a motion vector are indicated. The second position is implicitly derived by adding a small offset vector `mvOff` (here, a horizontal shift of 1 sample (1,0) is used as an example, but any horizontal and / or vertical shift can typically be performed). The direction / magnitude of the motion vector can remain the same.

[0193] The order of the motion vectors in the list is important for the encoder and decoder to interpret them in the same way. An example order can be used as follows.

[0194] • The first few items in the list have higher weights than the last few items. In other words, the index j in the list (or pair) of motion vectors can be used to deduce the weights of the motion vectors.

[0195] • Weight can depend on the distance from the top items in the list (the greater the distance, the greater the weight, and vice versa).

[0196] The weights can depend on the distribution of control points and can be trained using artificial intelligence (such as neural networks).

[0197] – The distance in Expression 11 can be estimated based on the position of the second motion vector in a pair of motion vectors, or estimated as the minimum distance to the point where the line segment connecting the first and second motion vectors in a pair of motion vectors connects. The example above using Expression 11 estimates the distance based on the position of the first motion vector.

[0198] Some advantages of the present invention are as follows:

[0199] • It can support high-order motion models for video compression while maintaining the lowest possible signaling overhead.

[0200] • It can be predicted through reverse motion and avoid segment overlap and discontinuity based on optical flow, etc.

[0201] • It can combine connected segments that may have different movements (such as the body and hands).

[0202] • Applicable to any dimension, such as 2D (two MVs), 3D (two MVs + rotation angle) or higher.

[0203] Reference Figure 6 The described methods include detailed information, which can be omitted or replaced with alternatives. Figure 12This is a flowchart of another method provided in one embodiment.

[0204] According to this embodiment, a method 1200 is provided for estimating the motion vector at a target location (e.g., 950). Advantageously, this method is provided to estimate the motion vector at each target location out of all locations in an image. These locations can be some or all integer samples (pixels) in the image. However, it should be noted that, through the motion field derivation described above, motion vectors at non-integer locations can be derived in the same way. It should be noted that the term "dense motion field" refers to all samples in the desired sample grid (desired resolution).

[0205] The method described above includes step 1210: acquiring two or more starting positions (e.g., 901, 925) and two or more motion vectors (e.g., 910, 911) starting from each of the two or more starting positions (e.g., 901, 925). Acquiring 1210 may include: reading the position and corresponding motion vector from memory, acquiring the position and corresponding motion vector from a result received by an application or from a previously determined result (e.g., subsampling of a dense motion field), or parsing the position and corresponding motion vector from memory or a bitstream received through a channel, etc. Similarly, step 1220 includes: for each of the two or more starting positions (e.g., 901, 925), acquiring a corresponding transformation (e.g., 740, 840) for transforming the motion vector (e.g., 910, 911) starting from the starting position (e.g., 901, 925) to another position (e.g., 925, 927). The transformation acquired in step 1220 may be performed in any manner. For example, the position and motion vectors obtained in step 1210 can be used to calculate the transformation. However, the invention is not limited to this example, and the transformation can be defined (obtained) in another way. Here, transformation can refer to the parameters of the transformation. For example, as mentioned above, the transformation can be an affine transformation, where scaling, rotation, and translation can be modeled. Then, the parameters of this transformation are determined. Optionally or additionally, the type of transformation can also be obtained.

[0206] Method 1200 further includes: transforming each of the two or more motion vectors (e.g., 910, 911) from the starting position (e.g., 901, 925) to the target position 950 of the corresponding transformation (740, 840) using the corresponding transformation (e.g., 740, 840), to determine 1230 two or more contributing motion vectors (e.g., 1140, 1145). Furthermore, after determining the contributing motion vectors, the method further includes the step of: estimating 1240 the motion vector (e.g., 1148) at the target position 950, wherein the motion vector includes a weighted average of the two or more contributing motion vectors (e.g., 1140, 1145).

[0207] In method 1200, the weighted average can be calculated by weighting two or more contributing motion vectors using weights, where the weights are a nonlinear function of the distance between the starting position and the target position 950 (from which the contributing motion vectors are obtained through transformation). As mentioned above, it may be advantageous when the sum of all weights equals 1 to maintain appropriate size. In some embodiments, the nonlinear function is a function that has a maximum value of 0 (corresponding to a distance of 0 from the starting position of the MV, where the MV is transformed to the contributing MV). For example, the nonlinear function is a Gaussian distribution function.

[0208] In some embodiments, the distance corresponds to the square norm, that is, to the Euclidean distance in the space given by one or more (current) pieces. However, this does not limit the invention. The distance can be defined by other norms, such as absolute difference or higher norms or some other distance metric.

[0209] Obtaining the transformation corresponding to 1220 may include: obtaining motion vectors (e.g., 911, 912) starting from the other position (e.g., 925, 927); estimating the parameters of the affine transformation based on the affine transformation from the motion vectors (e.g., 901, 911) starting from the starting position (e.g., 901, 925) to the motion vectors (e.g., 911, 912) starting from the other position (e.g., 925, 927). As mentioned above, the affine transformation is just a specific example. The invention is not limited thereto; in general, any one or more transformations can be used. This transformation can model nonlinear motion, etc. For example, two or more starting positions belong to a set of Ns starting positions, where Ns > 2, and the starting positions are arranged in a predefined order; for starting position j, 0 ≤ j ≤ Ns, the other position is position j+1 in the predefined order. Specifically, the weight of the contribution vector may depend on the position of the starting position of the corresponding transformed motion vector within the predefined order.

[0210] In some embodiments, the two or more starting positions are sample positions within a slice of an image, wherein the image comprises multiple slices, each slice being a set of image samples smaller than the image itself. For example, a sample can be a pixel in the image (e.g., an integer pixel, or pixels within a desired pixel grid in a general sense). However, it should be noted that the invention is not limited to these examples. One advantage of these embodiments is that motion vectors can be predicted at any location, including subsample locations. In fact, using sub-pel precision can be beneficial for some applications such as video encoding and decoding. These slices can be predefined slices, such as units (blocks) of a predetermined size known to the encoder and decoder. Slices do not need to be rectangular or square. If the invention is used in a codec, the slices can correspond to other meaningful image regions supported by the codec. For example, these slices can correspond to the references above. Figure 1 and Figure 2 The invention mentions CTU, stripes, or blocks. However, different types of slicing can also be used, such as slices corresponding to objects. In this case, the image can be segmented into objects, and one of the objects can also correspond to the background. Furthermore, it should be noted that slicing is not necessarily continuous; it can also be distributed. In other words, portions of the same slice do not necessarily have boundaries with any other portions of the same slice. Specifically, in some embodiments, it may be advantageous if slicing is based on motion characteristics, so that image portions with similar motion characteristics are grouped into the same slice.

[0211] The above method may further include the following steps: reconstructing the motion vector field of the image segments, including estimating a motion vector starting from each sample target position P(x,y) of the segment, wherein the motion vector does not belong to two or more available starting positions for the corresponding motion vector. In other words, as shown in some of the embodiments and examples above, motion vectors can be derived to obtain a dense motion field of optical flow in an approximate image.

[0212] Specifically, the two or more starting positions and the two or more motion vectors starting from the two or more starting positions are obtained by parsing from the bitstream associated with the slice of the image. The weights used in the weighted average are determined based on one or more parameters parsed from the bitstream. This method is particularly suitable for deploying the invention as part of an encoder or encoding method.

[0213] Optionally or additionally, the two or more starting positions within the slice of the image are determined based on features of the slice decoded from the bitstream. The two or more motion vectors starting from the two or more starting positions are obtained by parsing from the bitstream associated with the slice. The weights used in the weighted average are determined based on one or more parameters parsed from the bitstream. This approach is particularly suitable for deploying the invention as part of a decoder or decoding method.

[0214] In one embodiment, the two or more starting positions and the two or more motion vectors starting from the two or more starting positions are obtained by determining a motion vector field and by subsampling the obtained motion vector field, wherein the motion vector field includes the motion vector at each sample position (as the target position) of the segment of the image. Optionally or additionally, the weights of the corresponding contributing motion vectors are determined by rate-distortion optimization or machine learning.

[0215] As described above, the motion vector estimation (derivation) of the present invention can be used in encoding and decoding methods. Therefore, such a decoding method can be provided to decode an image. The above method may include estimating the motion vector at the target location of a sample according to any of the above embodiments or examples. Furthermore, the above decoding method may also include the step of predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image. Additionally, the above decoding method may include reconstructing the sample at the target location based on the prediction.

[0216] Accordingly, the present invention provides a method for encoding an image. The method may include the following steps: estimating a motion vector at a target location based on any of the above embodiments or examples. The encoding method includes: predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image. The method further includes: encoding the sample at the target location based on the prediction.

[0217] The encoding method may include other steps, such as determining a dense motion field before the estimation; and subsampling the determined motion field to obtain at least two positions and corresponding motion vectors, i.e., motion vectors starting from the positions generated by the subsampling.

[0218] The present invention also provides an apparatus corresponding to the above-described method. That is, the apparatus is capable of and used to perform the steps of the above-described method.

[0219] Figure 13An apparatus 1300 for estimating motion vectors at a target location is shown. The apparatus includes processing circuitry 1301. This processing circuitry can be implemented by one or more processors. As described above, these processors can be general-purpose processors or special-purpose processors, such as digital signal processors, programmable hardware, and / or special-purpose circuitry, such as ASICs. It should be noted that the processing circuitry 1301 can also implement other functions different from those related to motion vector estimation. For example, the processing circuitry can also implement an encoder and / or decoder to encode / decode images (video frames) using motion estimation.

[0220] The processing circuit 1301 may further include a circuit 1320 for acquiring two or more starting positions and two or more motion vectors respectively starting from said two or more starting positions. Circuit 1320 can be considered a functional module or unit for acquiring two or more starting positions and two or more motion vectors. Such a module / unit / circuit 1320 may be a spatially separate part of the processing circuit 1301, or it may share the processing circuit with other functional modules / units. Module 1320 may be a memory read module. The memory may be part of the device 1300, but it may not be. The memory may be part of the processing circuit 1301, but it may not be. It should be noted that the processing circuit 1301 may be an integrated circuit on a single chip. However, other configurations are also possible. For example, the entire device 1300 may be integrated on a chip.

[0221] Furthermore, the processing unit may include circuitry 1330, configured to: for each of the two or more starting positions, obtain a corresponding transformation for transforming a motion vector starting from one starting position to another. Similar to module 1320, circuitry 1330 may correspond to a functional module / unit, such as a transformation determination unit. The transformation determination unit / module 1330 can calculate, based on the position and motion vector obtained from module 1320, a transformation according to the type of transformation to be determined (e.g., affine transformation). Determination may include determining parameters for such a transformation.

[0222] The processing circuit 1301 may further include a circuit 1340 for determining two or more contributing motion vectors by transforming each of the two or more motion vectors from the starting position to the target position of the corresponding transformation using the corresponding transformation. This circuit corresponds to a functional unit, such as a transformation module 1340. This module may be controlled by a transformation determination module 1330, which determines the transformation and provides it to the transformation module 1340 that uses the transformation.

[0223] The processing circuit 1301 may further include a circuit 1350 for estimating the motion vector at the target position, wherein the motion vector includes a weighted average of the two or more contributing motion vectors. This circuit may correspond to a functional unit such as a motion vector estimation module 1350. This module is used to estimate the motion vector at the target position based on the contributing motion vectors obtained from the transformation module 1340. Furthermore, the memory read circuit 1320 may also obtain the weighting function and / or weights of the corresponding contributing motion vectors and provide them to the MV estimation module 1350.

[0224] As described above, the present invention also provides an encoding apparatus (device) 20. The encoding apparatus may be an encoder (or a combined encoder and decoder) for encoding one image from a plurality of images in a video sequence. The encoder 20 may include means 1300 for estimating a motion vector at a target location. The encoder 20 may also include a sample predictor 244 for predicting a sample at the target location in the image based on the estimated motion vector and a corresponding reference image. Means 1300 may be part of the predictor 244. The encoder 20 may also include a bitstream generator 270 for encoding the sample at the target location based on the prediction.

[0225] The present invention may also provide a decoding device 30 for decoding an image. Such a decoder 30 may include means 1300 for estimating motion vectors at a target location, a sample predictor 344 for predicting samples at the target location in the image based on the estimated motion vectors and the corresponding reference image, and a sample reconstructor 314 for reconstructing samples at the target location based on the predictions.

[0226] Figure 14 This is a schematic block diagram of an exemplary decoding system 10, such as a video decoding system 10 (or simply decoding system 10) that can utilize the technology of the present invention. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) in the video decoding system 10 are two examples, i.e., devices that can be used to perform various technologies according to the various examples described in this application. Specifically, encoder 20 may correspond to... Figure 1 or Figure 5 The encoder shown. Decoder 30 can correspond to... Figure 2 The aforementioned decoder, etc.

[0227] like Figure 14As shown, the decoding system 10 includes a source device 12, which provides encoded image data 21 to a destination device 14, etc., for decoding the encoded image data 13. The source device 12 includes an encoder 20 and may additionally (optionally) include an image source 16, a preprocessor (or preprocessing unit) 18 (e.g., an image preprocessor 18), and a communication interface or communication unit 22. The image source 16 may include, or may be, any type of image capture device such as a camera for capturing real-world images, and / or any type of image generation device such as a computer graphics processor for generating computer-generated animated images, or any other device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory / storage for storing any of the aforementioned images.

[0228] To distinguish between the processing performed by the preprocessor 18 and the preprocessing unit 18, the image or image data 17 may also be referred to as the raw image or raw image data 17. The preprocessor 18 receives the (raw) image data 17 and performs preprocessing on it to obtain a preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It is understood that the preprocessing unit 18 may be an optional component.

[0229] The video encoder 20 is used to receive preprocessed image data 19 and provide encoded image data 21. The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and transmit the encoded image data 21 (or data obtained after further processing of the encoded image data 21) to another device, such as the destination device 14 or any other device, via the communication channel 13 for storage or direct reconstruction. The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally (optionally) include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0230] The communication interface 28 in the destination device 14 is used to receive encoded image data 21 (or data obtained after further processing of encoded image data 21) directly from the source device 12 or from any other source such as a storage device (e.g., an encoded image data storage device), and to provide the encoded image data 21 to the decoder 30. Communication interfaces 22 and 28 can be used to transmit or receive encoded image data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14 (e.g., a direct wired or wireless connection) or via any type of network (e.g., a wired network or a wireless network or any combination thereof, or any type of private and public network, or any combination thereof).

[0231] For example, communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format (e.g., data packets) and / or process the encoded image data using any type of transmission encoding or processing method for transmission over a communication link or network. For example, communication interface 28, corresponding to communication interface 22, can be used to receive transmitted data and process the transmitted data using any corresponding transmission decoding or processing method and / or decapsulation method to obtain the encoded image data 21. Both communication interface 22 and communication interface 28 can be configured as follows: Figure 14 The communication channel 13, from source device 12 to destination device 14, is a one-way communication interface indicated by the arrow, or configured as a two-way communication interface, and can be used to send and receive messages, etc., to establish connections, acknowledge and exchange any other information related to the communication link and / or data transmission (e.g., encoded image data transmission), etc.

[0232] Decoder 30 is used to receive encoded image data 21 and provide decoded image data 31 or decoded image 31. Postprocessor 32 in destination device 14 is used to postprocess the decoded image data 31 (also called reconstructed image data) (e.g., decoded image 31) to obtain post-processed image data 33 (e.g., post-processed image 33). Postprocessing performed by postprocessing unit 32 may include color format conversion (e.g., from YCbCr to RGB), color adjustment, trimming or resampling, or any other processing to provide the decoded image data 31 for display by display device 34, etc., etc.

[0233] The display device 34 in the destination device 14 is used to receive post-processed image data 33 in order to display the image to a user or viewer. The display device 34 can be or can include any type of display for representing the reconstructed image, such as an integrated or external display or screen. For example, the display can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display.

[0234] Although Figure 14 Source device 12 and destination device 14 are shown as separate devices, but device embodiments may also include both devices or the functions of both devices simultaneously, i.e., source device 12 or its corresponding functions and destination device 14 or its corresponding functions. In these embodiments, the functions of source device 12 or its corresponding functions and destination device 14 or its corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. It will be apparent to those skilled in the art based on the description that… Figure 14 The presence and (precise) functional division of different units or functions within the source device 12 and / or destination device 14 shown may vary depending on the actual device and application.

[0235] Encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30), or encoder 20 and decoder 30, can be transmitted via... Figure 15 The processing circuit shown is used to implement this. This processing circuit includes one or more microprocessors, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more discrete logic devices, one or more hardware devices, one or more dedicated video decoding processors, or any combination thereof. Encoder 20 can be implemented through processing circuitry 46 to include reference... Figure 1 The encoder 20 describes various modules and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented via processing circuitry 46 to include references... Figure 2The decoder 30 describes various modules and / or any other decoder system or subsystem described herein. Processing circuitry can be used to perform various operations discussed below. Figure 17 As shown, if the aforementioned technical components are implemented in software, a device can store the instructions of that software in a suitable non-transitory computer-readable storage medium, and these instructions can be executed in hardware by one or more processors to perform the techniques of the present invention. The video encoder 20 or video decoder 30 can be integrated into a single device as part of a combined codec (codec), for example, as... Figure 15 As shown.

[0236] Source device 12 and destination device 14 can include any of a variety of devices, including any type of handheld or fixed device, such as a laptop or notebook computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (e.g., content service server or content distribution server), broadcast receiver device, broadcast transmitter device, etc., and may or may not use any type of operating system. In some cases, source device 12 and destination device 14 can be used for wireless communication. Therefore, source device 12 and destination device 14 can be wireless communication devices.

[0237] In some cases, Figure 14The video decoding system 10 shown is merely an example, and the techniques in this application can be applied to video decoding setups (e.g., video encoding or video decoding) that do not necessarily involve any data communication between encoding and decoding devices. In other examples, data is retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data into memory and / or retrieve data from memory and decode it. For ease of description, embodiments of the invention are described herein with reference, for example, to reference software developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.

[0238] Figure 16 This is a schematic diagram of a video decoding device 400 provided according to one embodiment of the present invention. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the decoding device 400 may be a decoder (e.g., Figure 14 The video decoder 30 or encoder (e.g.) in the video decoder 30) or encoder Figure 14 The video encoder 20 is included in the video decoding device 400. The video decoding device 400 includes an input port 410 (or input port 410) and a receiving unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing the data, a transmission unit (Tx) 440 and an output port 450 (or output port 450) for transmitting the data, and a memory 460 for storing the data. The video decoding device 400 may also include optical-to-electrical (OE) components and electro-optical (EO) components coupled to the input port 410, receiving unit 420, transmission unit 440, and output port 450, serving as the output or input of optical or electrical signals.

[0239] Processor 430 is implemented through hardware and software. Processor 430 can be implemented as one or more CPU chips, one or more cores (e.g., a multi-core processor), one or more FPGAs, one or more ASICs, and one or more DSPs. Processor 430 communicates with ingress port 410, receiver 420, transmitter 440, egress port 450, and memory 460. Processor 430 includes decoding module 470. Decoding module 470 implements the embodiments disclosed above. For example, decoding module 470 performs, processes, prepares, or provides various decoding operations. Therefore, including decoding module 470 provides a substantial improvement to the functionality of video decoding device 400 and affects the transitions of video decoding device 400 to different states. Optionally, decoding module 470 is implemented with instructions stored in memory 460 and executed by processor 430.

[0240] Memory 460 may include one or more disks, one or more tape drives, and one or more solid-state drives, and may be used as an overflow data storage device to store programs as selected for execution, as well as instructions and data read during program execution. For example, memory 460 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0241] Figure 17 A simplified block diagram of a device 1700 provided for an exemplary embodiment. Device 1700 can be used as... Figure 14 The source device 12 and / or destination device 14 are included. The processor 1702 in the apparatus 1700 may be a central processing unit. Alternatively, the processor 1702 may be any other type of device or multiple devices, existing or to be developed in the future, capable of operating or processing information. While the disclosed implementation may be implemented using a single processor such as the processor 502 shown, using multiple processors can improve speed and efficiency.

[0242] In one implementation, the memory 1704 in device 1700 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as memory 1704. Memory 1704 may include code and data 1706 accessed by processor 1702 via bus 1712. Memory 1704 may also include an operating system 1708 and an application program 1710, which includes at least one program that causes processor 1702 to perform the methods described herein. For example, application program 1710 may include applications 1 through N, and also includes a video decoding application that performs the methods described herein.

[0243] Device 1700 may also include one or more output devices, such as display 1718. In one example, display 1718 may be a touch-sensitive display combining a display with a touch-sensitive element capable of sensing touch input. Display 1718 may be coupled to processor 1702 via bus 1712. Although bus 1712 in device 1700 is described herein as a single bus, bus 1712 may include multiple buses. Furthermore, auxiliary memory 1714 may be directly coupled to other components in device 1700 or may be accessed via a network, and may include a single integrated unit (e.g., a memory card) or multiple units (e.g., multiple memory cards). Therefore, device 1700 can be implemented in a variety of configurations.

[0244] In summary, this invention provides a method and apparatus for estimating motion vectors of a dense motion field based on a subsampled sparse motion field. The sparse motion field includes two or more motion vectors and their respective starting positions. For each motion vector, a transformation is derived to transform the motion vector from its starting point to a target point. The transformed motion vector is then contributed to the motion vector estimate at the target position. The contribution of each motion vector is weighted. This motion estimation can be readily used for video encoding and decoding.

Claims

1. A method for estimating the motion vector at a target location, characterized in that, The method comprises: acquiring two or more starting positions and two or more motion vectors respectively starting from the two or more starting positions; for each starting position in the two or more starting positions, acquiring a corresponding transformation for transforming the motion vector starting from said starting position to another position; wherein, said acquiring the corresponding transformation comprises: acquiring a motion vector starting from said another position; estimating parameters of an affine transformation according to the affine transformation from the motion vector starting from said starting position to the motion vector starting from said another position; transforming each motion vector of said two or more motion vectors from said starting position to a target position of said corresponding transformation by using said corresponding transformation, to determine two or more contribution motion vectors; estimating a motion vector at said target position, wherein said motion vector comprises a weighted average of said two or more contribution motion vectors.

2. The method according to claim 1, characterized in that, Said weighted average is calculated by weighting each contribution motion vector of said two or more contribution motion vectors by using a weight, wherein said weight is a non-linear function of a distance between said starting position and said target position.

3. The method according to claim 2, characterized in that, Said non-linear function is a Gaussian distribution function.

4. The method according to claim 2 or 3, characterized in that, Said distance is obtained by calculating a squared norm.

5. The method according to claim 1, characterized in that, said two or more starting positions belong to a set of Ns starting positions, wherein Ns>2, and said starting positions are arranged in a predefined order; for a starting position j, 0≤j<Ns, said another position is a position j+1 in said predefined order.

6. The method according to claim 5, characterized in that, the weight of a contribution motion vector depends on a position of a starting position of a corresponding transformed motion vector within said predefined order.

7. The method according to any one of claims 1 to 3, characterized in that, said two or more starting positions are sample positions in a tile of an image, wherein said image comprises a plurality of tiles, said tile is a set of image samples, and said set of image samples is smaller than said image.

8. The method according to claim 7, characterized in that, said method comprises the following step: reconstructing a motion vector field of said tile of said image, comprising estimating a motion vector starting from each sample target position P(x, y) of said tile, wherein each said sample target position does not belong to two or more starting positions for which corresponding motion vectors are available.

9. The method according to claim 7, characterized in that, said two or more starting positions and said two or more motion vectors respectively starting from said two or more starting positions are obtained by parsing from a code stream related to said tile (S) of said image; the weight used in said weighted average is determined according to one or more parameters parsed from said code stream.

10. The method according to claim 7, characterized in that, said two or more starting positions within said tile (S) of said image are determined according to characteristics of said tile (S) decoded from a code stream; The two or more motion vectors starting from the two or more starting positions are obtained by parsing from the bitstream associated with the segment (S); The weights used in the weighted average are determined based on one or more parameters parsed from the bitstream.

11. The method according to claim 7, characterized in that, The two or more starting positions and the two or more motion vectors starting from the two or more starting positions are obtained by determining a motion vector field and by subsampling the acquired motion vector field, wherein the motion vector field includes the motion vector of each sample target position of the segment of the image; and / or The weights of the corresponding contribution motion vectors are determined through rate-distortion optimization or machine learning.

12. A method for decoding an image, characterized in that, The method includes: The method according to any one of claims 1 to 10 is used to estimate the motion vector at the target location; Based on the estimated motion vector and the corresponding reference image, predict the sample at the target location in the image; Reconstruct the sample at the target location based on the prediction.

13. A method for encoding an image, characterized in that, The method includes: The method according to any one of claims 1 to 7 or 11 is used to estimate the motion vector at the target location; Based on the estimated motion vector and the corresponding reference image, predict the sample at the target location in the image; The sample at the target location is encoded based on the prediction.

14. An apparatus for estimating the motion vector at a target position, characterized in that, The device includes a processing circuit, the processing circuit comprising: A circuit for acquiring two or more starting positions and two or more motion vectors starting from the two or more starting positions respectively; A circuit is configured to: for each of the two or more starting positions, obtain a corresponding transformation for transforming a motion vector starting from the starting position to another position; wherein obtaining the corresponding transformation includes: obtaining a motion vector starting from the other position; and estimating the parameters of the affine transformation based on the affine transformation from the motion vector starting from the starting position to the motion vector starting from the other position. A circuit for: determining two or more contributing motion vectors by transforming each of the two or more motion vectors from the starting position to the target position of the corresponding transformation using the corresponding transformation; A circuit for estimating a motion vector at the target location, wherein the motion vector includes a weighted average of two or more contributing motion vectors.

15. An encoding device for encoding images, characterized in that, The device includes: The apparatus for estimating the motion vector at a target location according to claim 14; A sample predictor is used to predict a sample at the target location in the image based on the estimated motion vector and the corresponding reference image; A stream generator is used to encode samples at the target location based on the prediction.

16. A decoding device for decoding images, characterized in that, The device includes: The apparatus for estimating the motion vector at a target location according to claim 14; A sample predictor is used to predict a sample at the target location in the image based on the estimated motion vector and the corresponding reference image; A sample reconstructor is used to reconstruct samples at the target location based on the prediction.

17. A computer program product, characterized in that, The computer program product includes instruction code that, when executed on one or more processors, performs the steps of the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Devices and Methods for Sparse Representation of Dense Motion Vector Fields for Compression of Visual Pixel Data

    US20140049607A1