An affine motion estimation method, device, storage medium and terminal

By constructing an affine motion estimation method based on a three-layer image downsampling pyramid, the problem of large data access volume in existing technologies is solved, the encoder design is simplified, and the encoding complexity and time are reduced.

CN114666606BActive Publication Date: 2026-04-24HANGZHOU WEIMING XINKE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU WEIMING XINKE TECH CO LTD
Filing Date
2022-02-07
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing motion estimation methods require large amounts of data to access memory when encoding 3D images and ultra-high-definition videos, resulting in high encoding complexity, complex hardware design, and long encoding time.

Method used

An affine motion estimation method based on a three-layer image downsampling pyramid is adopted. By constructing a three-layer image downsampling pyramid, the motion vectors of control points are obtained, reduced and updated, and the final motion vector is calculated layer by layer, thereby reducing the amount of data access.

Benefits of technology

It reduces the amount of data access during motion compensation, simplifies the encoding complexity of the encoder, and reduces the design difficulty and encoding time of the hardware encoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114666606B_ABST
    Figure CN114666606B_ABST
Patent Text Reader

Abstract

The application relates to an affine motion estimation method, device, storage medium and terminal. The method comprises the following steps: constructing a three-layer image down-sampling pyramid; acquiring a control point motion vector of an original block on a zero-layer image; reducing the control point motion vector to obtain an initial motion vector of the original block on the zero-layer image; updating according to the initial motion vector to acquire a final motion vector of a prediction block on the zero-layer image; updating according to the final motion vector on the zero-layer image to acquire a final motion vector of a prediction block on a first-layer image; and updating according to the final motion vector on the first-layer image to output a final motion vector of a prediction block on a second-layer image. The application is based on the three-layer image down-sampling pyramid affine motion estimation; the data access amount in the motion compensation process can be reduced, and the access cost is further reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video coding technology, and more specifically, to an affine motion estimation method, apparatus, storage medium, and terminal. Background Technology

[0002] In today's information age, the demand for video services such as 3D imaging, ultra-high-definition video, and virtual reality is growing, and the encoding and transmission of high-definition video is a hot research topic.

[0003] Motion estimation is a widely used technique in video coding and video processing; excellent motion estimation algorithms can improve video coding efficiency.

[0004] The present invention proposes an affine motion estimation method, device, storage medium and terminal, which is based on a three-layer image downsampling pyramid. It can reduce the amount of data access during motion compensation and thus reduce the cost of accessing memory. Summary of the Invention

[0005] This application provides an affine motion estimation method, apparatus, storage medium, and terminal. To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general description, nor is it intended to identify key / important components or describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0006] In a first aspect, embodiments of this application provide an affine motion estimation method, the method comprising:

[0007] Construct a three-layer image downsampling pyramid; the three layers include the zeroth layer image, the first layer image, and the second layer image.

[0008] Obtain the motion vector of the control point of an original block on the zero-level image;

[0009] The motion vectors of the control points are reduced to obtain the initial motion vectors of the original blocks on the zero-layer image;

[0010] Update based on the initial motion vector to obtain the final motion vector of the prediction block on the zero-level image;

[0011] Update based on the final motion vector on the zero-level image, and obtain the final motion vector of a prediction block on the first-level image;

[0012] The final motion vector on the first layer image is updated, and the final motion vector of a predicted block on the second layer image is output.

[0013] Optionally, updates are performed based on the initial motion vector to obtain the final motion vector of a predicted block on the zero-level image, including:

[0014] Calculate the update amount of the initial motion vector on the zero-layer image based on the initial motion vector;

[0015] Based on the initial motion vector and the update amount of the initial motion vector, obtain the final motion vector of a prediction block on the zero-level image.

[0016] Optionally, the update amount of the initial motion vector on the zero-layer image is calculated based on the initial motion vector, including:

[0017] Based on the initial motion vector, a prediction block is obtained on the zero-th layer image;

[0018] An error block is formed on the zero-level image based on the predicted block and the original block on the zero-level image;

[0019] Obtain the error at any position on the error block of the zero-level image;

[0020] The pixel gradient at any location on the prediction block in the zero-level image is calculated using the Sobel operator;

[0021] The update amount of the initial motion vector is calculated based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0022] Optionally, updates are performed based on the final motion vector on the zero-level image, obtaining the final motion vector of a predicted block on the first-level image, including:

[0023] The update amount of the first layer motion vector on the first layer image is calculated based on the final motion vector on the zero layer image;

[0024] Based on the final motion vector on the zero-level image and the update amount of the motion vector on the first level, obtain the final motion vector of a prediction block on the first level image.

[0025] Optionally, the update amount of the first-layer motion vector on the first-layer image is calculated based on the final motion vector on the zero-layer image, including:

[0026] Based on the final motion vector on the zero-level image, a prediction block is obtained on the first-level image;

[0027] An error block is formed on the first layer image based on the prediction block on the first layer image and the prediction block on the zeroth layer image;

[0028] Obtain the error at any position on the error block in the first layer image;

[0029] The pixel gradient at any location on the prediction block in the first-layer image is calculated using the Sobel operator;

[0030] The update amount of the motion vector of the first layer is calculated based on the error at any position on the error block in the first layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0031] Optionally, based on the final motion vector on the first-layer image, the final motion vector of a predicted block on the second-layer image is updated, including:

[0032] The update amount of the second-layer motion vector in the second-layer image is calculated based on the final motion vector in the first-layer image;

[0033] Based on the final motion vector on the first layer image and the update amount of the motion vector on the second layer image, output the final motion vector of a prediction block on the second layer image.

[0034] Optionally, the update amount of the second-layer motion vector in the second-layer image is calculated based on the final motion vector in the first-layer image, including:

[0035] Based on the final motion vector on the first layer image, a predicted block on the second layer image is obtained;

[0036] Error blocks in the second layer image are formed based on the prediction blocks in the second layer image and the prediction blocks in the first layer image;

[0037] Obtain the error at any position on the error block in the second layer image;

[0038] The pixel gradient at any location on the prediction block in the second-layer image is calculated using the Sobel operator;

[0039] The update amount of the motion vector of the second layer is calculated based on the error at any position on the error block in the second layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0040] Secondly, embodiments of this application provide an affine motion estimation device, the device comprising:

[0041] The layer construction module is used to construct a three-layer image downsampling pyramid; the three layers include the zeroth layer image, the first layer image, and the second layer image.

[0042] The control point data acquisition module is used to acquire the motion vector of the control points of an original block on the zero-layer image;

[0043] The initial motion vector acquisition module is used to reduce the motion vector of the control points to obtain the initial motion vector of the original block on the zero-layer image;

[0044] The zero-layer motion vector acquisition module is used to update based on the initial motion vector and obtain the final motion vector of a prediction block on the zero-layer image;

[0045] The first-layer motion vector acquisition module is used to update based on the final motion vector on the zero-layer image and obtain the final motion vector of a prediction block on the first-layer image;

[0046] The second-layer motion vector output module is used to update the final motion vector of a prediction block on the second-layer image based on the final motion vector on the first-layer image.

[0047] Thirdly, embodiments of this application provide a computer storage medium storing multiple instructions adapted for loading and execution of the above-described method steps by a processor.

[0048] Fourthly, embodiments of this application provide a terminal that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed by the above-described method steps.

[0049] The technical solutions provided in this application embodiment may have the following beneficial effects:

[0050] In this embodiment, the affine motion estimation method first constructs a three-layer image downsampling pyramid to obtain the control point motion vector of an original block on the zero-layer image. Then, the control point motion vector is scaled down to obtain the initial motion vector of the original block on the zero-layer image. Next, the initial motion vector is updated to obtain the final motion vector of a predicted block on the zero-layer image. Then, the final motion vector on the zero-layer image is updated to obtain the final motion vector of a predicted block on the first-layer image. Finally, the final motion vector on the first-layer image is updated to output the final motion vector of a predicted block on the second-layer image. This application is based on affine motion estimation using a three-layer image downsampling pyramid; it can reduce the amount of data access during motion compensation, thereby reducing memory access costs.

[0051] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0052] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0053] Figure 1 This is a flowchart illustrating an affine motion estimation method provided in an embodiment of this application;

[0054] Figure 2This is a schematic diagram of an affine motion field for an affine motion estimation method provided in an embodiment of this application;

[0055] Figure 3 This is a schematic diagram of the predicted control point motion vector of an affine motion estimation method provided in an embodiment of this application;

[0056] Figure 4 This is a schematic diagram of a three-layer affine motion estimation method provided in the embodiments of this application;

[0057] Figure 5 This is a flowchart illustrating another affine motion estimation method provided in an embodiment of this application;

[0058] Figure 6 This is a schematic diagram of an affine motion estimation device provided in an embodiment of this application;

[0059] Figure 7 This is a schematic diagram of a terminal provided in an embodiment of this application. Detailed Implementation

[0060] The following description and accompanying drawings fully illustrate specific embodiments of the invention to enable those skilled in the art to practice them.

[0061] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0062] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of systems and methods consistent with some aspects of the invention as detailed in the appended claims.

[0063] In the description of this invention, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances. Furthermore, in the description of this invention, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0064] In inter-frame predictive coding, since there is a certain correlation between objects in neighboring frames of an image, the image can be divided into several original blocks, and the position of each original block in the neighboring frame image can be searched to obtain the relative offset between the two spatial positions. The obtained relative offset is the motion vector, and the process of obtaining the motion vector is called motion estimation.

[0065] Motion estimation can remove inter-frame redundancy, greatly reducing the number of bits transmitted in video. Motion estimation is an important component of video compression processing systems.

[0066] This invention proposes an affine motion estimation method, apparatus, storage medium, and terminal based on a three-layer image downsampling pyramid. This method reduces data access during motion compensation, thereby lowering memory access costs. It also reduces encoder complexity, simplifies hardware encoder design, and reduces encoding time for both hardware and software encoders.

[0067] The following will be combined with the appendix Figure 1 -Appendix Figure 5 This paper provides a detailed description of an affine motion estimation method provided in the embodiments of this application.

[0068] Please see Figure 1-4 This document provides a flowchart illustrating an affine motion estimation method as described in an embodiment of this application. Figure 1-4 As shown, the method in this application embodiment includes: S110, S120, S130, S140, S150 and S160.

[0069] In many inter-frame prediction tools, the motion vector is the difference between the current coding unit (CU) and the optimal matching block position on the reference image. The motion model can only represent translational motion. For more complex motions, such as scaling and rotation, it is necessary to divide the coding unit into multiple smaller units and represent them using translational motion vectors. In this application, the motion of a coding unit is represented as affine motion using the affine motion model introduced in the coding standards AVS3 and H.266 / VVC. In the affine motion model, the pixel motion vector at any position is determined by its position and the parameters of the affine motion model, as shown below:

[0070] mv x =n1x + n2y + n3

[0071] mv y = n4x + n5y + n6

[0072] Among them, mv x and MVy This represents the x and y components of the current pixel's motion vector, where x and y represent the current pixel's position, and n1 to n6 represent the parameters of the affine motion. Since having a different motion vector for each pixel would increase encoding complexity, the encoder ultimately uses sub-block-level affine motion. This means that a coding unit is divided into equally sized 4x4 original sub-blocks, and all pixels within each original sub-block share a single motion vector, which is the affine motion vector of the pixel at the center of the current original sub-block. The mapping relationship between the positions of these original sub-blocks with motion vectors and the motion vectors is called the Affine Motion Field (AMF). Based on the characteristics of affine motion vector fields, three non-linear motion vectors mv0, mv1, and mv2 are given in the affine motion vector field. The motion vector at any position can be linearly expressed by these three known motion vectors mv0, mv1, and mv2. In the encoder, the parameters of the affine motion field are represented by the three motion vectors mv0, mv1, and mv2 at the lower left, upper left, and upper right corners of a coding unit. The affine motion field of the entire coding unit is expressed through these parameters, as shown in the following representation: Figure 2 As shown.

[0073] Under the affine motion model, affine motion estimation (AME) can be categorized as a parameter estimation problem that minimizes the distortion between the current original block and the predicted block obtained according to the affine motion model. The distortion metric is a simplified rate distortion cost:

[0074] J = SATD + λR(MVD)

[0075] Where SATD represents the sum of absolute transformed differences (STD), R(MVD) is the representation cost of the differences between the three motion vectors, and λ is the Lagrange coefficient in rate-distortion optimization. The three motion vectors can be the final motion vectors on the second-layer image output in this embodiment of the application. and

[0076] The following section details the specific process of affine motion estimation, mainly involving image downsampling methods and affine motion estimation based on image downsampling pyramids; for example... Figure 4 As shown.

[0077] S110, construct a three-layer image downsampling pyramid; the three layers include a zero-layer image, a first-layer image, and a second-layer image; during the construction of the reference frame list, the image undergoes downsampling twice to form a three-layer image. The second-layer image is the original image, the first-layer image is a 1:4 downsampled image, and the zero-layer image is a 1:16 downsampled image.

[0078] In this embodiment, the image is downsampled twice using a Gaussian filter. The same downsampling operation can also be performed on the original frame of the current original block.

[0079] S120, obtain the motion vector of the control point of an original block on the zero-layer image; the original block can be a coding unit.

[0080] In this embodiment of the application, since motion has spatial correlation, it is necessary to predict the motion vector of the control points on the zero-layer image by the following method to generate a set of predicted control point motion vectors, namely pmv0, pmv1 and pmv2.

[0081] First, define the neighboring blocks A, B, C, D, F, and G of an original block E. The positions of the neighboring blocks A, B, C, D, F, and G relative to the original block E are as follows: Figure 3 As shown. For control point 0, its predicted motion vector pmv0 on the zero-th layer image is the encoded motion vector mv of the adjacent blocks A, B, and D at the top left corner of the original block E. A ,mv B and MV D The values ​​are obtained by summing and averaging; for control point 1, its predicted control point motion vector pmv1 on the zero-level image is obtained by summing the motion vectors mv of the adjacent blocks C and G at the upper right corner of the original block E. C and MV G The weighted average is used to obtain the motion vector pmv2 of the control point on the predicted zero-level image, which is obtained by the motion vector mv2 of the adjacent block F at the lower left corner of the original block E. F In the process of predicting the motion vector of the control point of an original block on the zeroth layer image, if the motion vector of an adjacent block is unavailable, the motion vector of that adjacent block can be represented by a 0 vector.

[0082] S130, the control point motion vectors are reduced to obtain the initial motion vectors of the original blocks on the zero-layer image. In this embodiment, the control point motion vectors on the zero-layer image are pmv0, pmv1, and pmv2. Since the zero-layer image is downsampled at a ratio of 1:16, the area of ​​the zero-layer image is reduced by 1 / 16, and the width and height of the zero-layer image are reduced by 1 / 4. The control point motion vectors on the zero-layer image can be reduced by 1 / 4 to obtain the initial motion vectors on the zero-layer image. and but:

[0083]

[0084]

[0085]

[0086] In this embodiment, the update amount of the initial motion vector on the zero-layer image can be calculated based on the initial motion vector of the original block on the zero-layer image. The final motion vector of a predicted block on the zero-layer image is obtained based on the initial motion vector and the update amount of the initial motion vector on the zero-layer image. For other layers, the update amount of the motion vector on that layer can be calculated based on the final motion vector of a predicted block on the previous layer image. The final motion vector of a predicted block on that layer is then calculated using the final motion vector on the previous layer image and the update amount of the motion vector on that layer. The previous layer refers to the layer before this layer. When the previous layer image is the zero-layer image, this layer image is the first layer image; when the previous layer image is the first layer image, this layer image is the second layer image.

[0087] S140, update based on the initial motion vector to obtain the final motion vector of a prediction block on the zero-level image. S140 includes:

[0088] S141, calculate the update amount of the initial motion vector on the zero-level image based on the initial motion vector. The specific process is as follows:

[0089] First, a prediction block on the zeroth layer image is obtained based on the initial motion vector. In this embodiment, the prediction block can be obtained based on the initial motion vector. and Calculate the motion vectors of each original sub-block within the original block E. Within an original block E of length h and width W, where length h is... Sub Width is w Sub The original sub-block, with position coordinates p(x,y) and motion vector mv Sub for:

[0090]

[0091] Based on the motion vector mv of the original sub-block Sub On the reference image of the corresponding zero-level image, find the motion vector mv of the original sub-block. Sub The pointed-to location is used to determine the size of the sub-block corresponding to that location, which is then used as a predicted sub-block of the original sub-block. The predicted sub-block on the zeroth layer image is located on the reference image of the zeroth layer image.

[0092] After performing the motion compensation described above on all the original sub-blocks, the resulting predicted sub-blocks are combined to form a predicted block of the original block.

[0093] Then, an error block is formed on the zero-layer image based on the predicted block and the original block; the error at any position on the error block on the zero-layer image is obtained. In this embodiment, the error block on the zero-layer image, also called the error plane on the zero-layer image, is formed based on the pixel range value at each corresponding position of the predicted block and the original block (i.e., the difference between the position coordinates on the predicted block and the corresponding position coordinates on the original block). The pixel range value is the error, and the error at any position point P on the error block is represented by u(p).

[0094] Next, the pixel gradient at any location of the prediction block on the zeroth layer image is calculated using the Sobel operator. In this embodiment, for any pixel at location point p(x,y)... Its horizontal pixel gradient is:

[0095]

[0096] The vertical pixel gradient is:

[0097]

[0098] Then the pixel gradient at the prediction block location p(x,y) on the zero-th layer image is:

[0099]

[0100] Finally, the update amount of the initial motion vector is calculated based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient. In this embodiment, the update amount of the three initial motion vectors is calculated using the following formula. and

[0101]

[0102]

[0103]

[0104] In the formula, m0(p) represents the update amount of the initial motion vector when acquiring control point 0. At that time, the position coefficient of position point P; m1(p) represents the update amount of the initial motion vector of control point 1. At that time, the position coefficient of position point P; m2(p) represents the update amount of the initial motion vector of control point 2. When, the position coefficient of position point P; m0(p), m1(p), and m2(p) are known functions;

[0105] m I (p) also represents the position coefficient of position point p, which is a known function, where I represents control points 0, 1, and 2;

[0106] R is the set of positions of each point in the original block, and T represents the transpose of the matrix;

[0107] This represents the pixel position q of the predicted block on the zero-level image. It is represented as the pixel gradient of the original sub-block at position p on the zero-level image, pointing to the reference frame at position q on the reference image.

[0108] The update values ​​of the three initial motion vectors can be obtained by solving the above system of three linear equations. and

[0109] S142, based on the initial motion vector and its update amount, obtain the final motion vector of a prediction block on the zero-layer image. In this embodiment, the final motion vector of a prediction block on the zero-layer image is calculated using the following formula. and

[0110]

[0111]

[0112]

[0113] `scale` represents the scaling factor, which is a fixed value. When calculating the final motion vector on the zeroth layer image, `scale = 2`.

[0114] S150, update based on the final motion vector on the zero-level image, and obtain the final motion vector of a predicted block on the first-level image. The final motion vector of a predicted block on the zero-level image can be used as the initial motion vector of an original block on the first-level image. S150 includes:

[0115] S151, calculate the update amount of the first-layer motion vector on the first-layer image based on the final motion vector on the zero-layer image. The specific process is as follows:

[0116] First, a prediction block on the first layer image is obtained based on the final motion vector on the zero-layer image. In this embodiment, the motion vectors of each sub-block within the prediction block on the zero-layer image can be calculated based on the final motion vector on the zero-layer image.

[0117] The calculation method for the motion vectors of each sub-block within the prediction block on the zero-layer image is similar to the calculation method for the motion vectors of the original sub-blocks on the zero-layer image, and will not be repeated here.

[0118] In this embodiment, based on the motion vectors of each sub-block on the reference image of the corresponding first layer image, the position pointed to by the motion vectors of each sub-block is found. The sub-block of the size corresponding to the pointed position is taken as the predicted sub-block of each sub-block divided on the zeroth layer image, that is, the predicted sub-block on the first layer image. The predicted sub-block on the first layer image is located on the reference image of the first layer image.

[0119] After performing motion compensation on all sub-blocks divided on the zero-layer image, the obtained prediction sub-blocks are combined into a prediction block on the zero-layer image, that is, the obtained prediction sub-blocks on the first layer image are combined into a prediction block on the first layer image.

[0120] Then, based on the prediction blocks on the first layer image and the prediction blocks on the zero layer image, an error block is formed on the first layer image; the error at any position on the error block on the first layer image is obtained.

[0121] In the embodiments of this application, an error block on the first layer image, also called an error plane on the first layer image, is formed based on the error at each corresponding position of the prediction block on the zero-layer image and the prediction block on the first layer image (i.e., the difference between the position coordinates on the prediction block of the zero-layer image and the corresponding position coordinates on the prediction block of the first layer image).

[0122] Next, the pixel gradient at any location on the prediction block of the first-layer image is calculated using the Sobel operator.

[0123] In this embodiment of the application, the method for calculating the pixel gradient at any position of the prediction block on the first layer image is similar to the method for calculating the pixel gradient at any position of the prediction block on the zero layer image. Both methods calculate the pixel gradient at any position by calculating the pixel gradient in the horizontal direction and the pixel gradient in the vertical direction.

[0124] Finally, the update amount of the motion vector of the first layer is calculated based on the error at any position on the error block in the first layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0125] In this embodiment of the application, the update amount of the first layer motion vector and The calculation method and the update amount of the above three initial motion vectors and The calculation method is similar; it obtains the update amount of the first-layer motion vector by constructing and solving a system of three linear equations. and

[0126] S152, based on the final motion vector on the zero-layer image and the update amount of the motion vector on the first layer, obtain the final motion vector of a prediction block on the first layer image.

[0127] In this embodiment, the final motion vector of a prediction block on the first layer image can be calculated using the following formula. and

[0128]

[0129]

[0130]

[0131] When obtaining the final motion vector on the first layer image, scale = 2 represents the scaling factor.

[0132] S160 updates the image based on the final motion vector in the first layer image, outputting the final motion vector of a predicted block in the second layer image. The final motion vector of a predicted block in the first layer image can be used as the initial motion vector of an original block in the second layer image. S160 includes:

[0133] S161, calculate the update amount of the second-layer motion vector in the second-layer image based on the final motion vector in the first-layer image. The specific process is as follows:

[0134] First, a prediction block on the second layer image is obtained based on the final motion vector on the first layer image. In this embodiment, the motion vectors of each sub-block within the prediction block on the first layer image can be calculated based on the final motion vector on the first layer image.

[0135] The calculation method for the motion vectors of each sub-block within the prediction block on the first layer image is similar to the calculation method for the motion vectors of the original sub-blocks on the zero layer image, and will not be repeated here.

[0136] In this embodiment, based on the motion vectors of each sub-block on the reference image of the corresponding second layer image, the position pointed to by the motion vectors of each sub-block is found. The sub-block of the size corresponding to the pointed position is taken as the predicted sub-block of each sub-block divided on the first layer image, that is, the predicted sub-block on the second layer image. The predicted sub-block on the second layer image is located on the reference image of the second layer image.

[0137] After performing motion compensation on all sub-blocks divided on the first layer image, the obtained predicted sub-blocks are combined into a prediction block on the first layer image, that is, the obtained predicted sub-blocks on the second layer image are combined into a prediction block on the second layer image.

[0138] Then, based on the prediction blocks on the second-layer image and the prediction blocks on the first-layer image, an error block is formed on the second-layer image; the error at any position on the error block on the second-layer image is obtained.

[0139] In the embodiments of this application, an error block on the second layer image, also called an error plane on the second layer image, is formed based on the error at each corresponding position of the prediction block on the first layer image and the prediction block on the second layer image (i.e., the difference between the position coordinates on the prediction block of the first layer image and the corresponding position coordinates on the prediction block of the second layer image).

[0140] Next, the pixel gradient at any location on the prediction block of the second-layer image is calculated using the Sobel operator.

[0141] In this embodiment of the application, the method for calculating the pixel gradient at any position of the prediction block on the second layer image is similar to the method for calculating the pixel gradient at any position of the prediction block on the zero layer image described above. Both methods calculate the pixel gradient at a certain position by calculating the pixel gradient in the horizontal direction and the pixel gradient in the vertical direction.

[0142] Finally, the update amount of the motion vector of the second layer is calculated based on the error at any position on the error block in the second layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0143] In this embodiment of the application, the update amount of the second layer motion vector and The calculation method and the update amount of the above three initial motion vectors and The calculation method is similar; it obtains the update amount of the second-layer motion vector by constructing and solving a system of three linear equations. and

[0144] S162, based on the final motion vector on the first layer image and the update amount of the second layer motion vector, output the final motion vector of a prediction block on the second layer image.

[0145] In this embodiment, the final motion vector of a predicted block on the second-layer image is calculated using the following formula. and

[0146]

[0147]

[0148]

[0149] When obtaining the final motion vector on the second layer image, scale = 1 represents the scaling factor.

[0150] Output the final motion vector on the second layer image. and

[0151] This application constructs a three-layer image downsampling pyramid to perform affine motion estimation from coarse to fine. The Taylor expansion in the coarse process is performed on the downsampled image. One coded unit of displacement (i.e., the update amount of the motion vector in this application) represents the distance of multiple pixels on the fine image. The convergence process is relatively fast, requiring only three iterations, and the update amount of the motion vector obtained in each iteration is relatively small. The data access required for affine motion estimation through three-layer image downsampling is also reduced according to the downsampling ratio, with the data access during motion compensation being only [amount missing]. Each pixel represents 18.75% of the original data access time, thus reducing the cost of accessing memory.

[0152] In this embodiment, the affine motion estimation method first constructs a three-layer image downsampling pyramid to obtain the control point motion vector of an original block on the zero-layer image. Then, the control point motion vector is reduced to obtain the initial motion vector of the original block on the zero-layer image. Next, the initial motion vector is updated to obtain the final motion vector of a predicted block on the zero-layer image. Then, the final motion vector on the zero-layer image is updated to obtain the final motion vector of a predicted block on the first-layer image. Finally, the final motion vector on the first-layer image is updated to output the final motion vector of a predicted block on the second-layer image. This application is based on affine motion estimation using a three-layer image downsampling pyramid. The affine motion estimation process requires only three iterations, which reduces the amount of data access during motion compensation, thereby reducing memory access costs.

[0153] Please see Figure 5This document provides a flowchart illustrating an affine motion estimation method as described in an embodiment of this application. Figure 5 As shown, the method in this application embodiment may include the following steps:

[0154] S201, construct a three-layer image downsampling pyramid; the three layers of images include the zeroth layer image, the first layer image, and the second layer image;

[0155] S202, Obtain the motion vector of the control point of an original block on the zero-layer image;

[0156] S203, reduce the motion vector of the control point to obtain the initial motion vector of the original block on the zero-layer image;

[0157] S204, Based on the initial motion vector, obtain a prediction block on the zero-th layer image;

[0158] S205, form an error block on the zero-level image based on the predicted block and the original block on the zero-level image;

[0159] S206, Obtain the error at any position on the error block of the zero-layer image;

[0160] S207, calculate the pixel gradient at any location of the prediction block on the zero-level image using the Sobel operator;

[0161] S208, calculate the update amount of the initial motion vector based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient;

[0162] S209, Based on the initial motion vector and the update amount of the initial motion vector, obtain the final motion vector of a prediction block on the zero-level image;

[0163] S210, Based on the final motion vector on the zero-layer image, obtain a prediction block on the first-layer image;

[0164] S211, an error block is formed on the first layer image based on the prediction block on the first layer image and the prediction block on the zeroth layer image;

[0165] S212, Obtain the error at any position on the error block in the first layer image;

[0166] S213, calculate the pixel gradient at any position of the prediction block on the first layer image using the Sobel operator;

[0167] S214, calculate the update amount of the motion vector of the first layer based on the error at any position on the error block in the first layer image, the pixel gradient at any position of the prediction block, and the position coefficient;

[0168] S215, based on the final motion vector on the zero-layer image and the update amount of the motion vector on the first layer, obtain the final motion vector of a prediction block on the first layer image;

[0169] S216, Based on the final motion vector on the first layer image, obtain a prediction block on the second layer image;

[0170] S217, form an error block on the second layer image based on the prediction block on the second layer image and the prediction block on the first layer image;

[0171] S218, Obtain the error at any position on the error block in the second layer image;

[0172] S219, calculate the pixel gradient at any location of the prediction block on the second-layer image using the Sobel operator;

[0173] S220, calculate the update amount of the motion vector of the second layer based on the error at any position on the error block on the second layer image, the pixel gradient at any position on the prediction block, and the position coefficient;

[0174] S221, based on the final motion vector on the first layer image and the update amount of the second layer motion vector, output the final motion vector of a prediction block on the second layer image.

[0175] In this embodiment, the affine motion estimation method first constructs a three-layer image downsampling pyramid to obtain the control point motion vector of an original block on the zero-layer image. Then, the control point motion vector is reduced to obtain the initial motion vector of the original block on the zero-layer image. Next, the initial motion vector is updated to obtain the final motion vector of a predicted block on the zero-layer image. Then, the final motion vector on the zero-layer image is updated to obtain the final motion vector of a predicted block on the first-layer image. Finally, the final motion vector on the first-layer image is updated to output the final motion vector of a predicted block on the second-layer image. This application is based on affine motion estimation using a three-layer image downsampling pyramid. The affine motion estimation process requires only three iterations, which reduces the amount of data access during motion compensation, thereby reducing memory access costs.

[0176] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the embodiments of the apparatus of the present invention, please refer to the embodiments of the method of the present invention.

[0177] Please see Figure 6The diagram illustrates a schematic representation of an affine motion estimation device according to an exemplary embodiment of the present invention. The device 1 includes: a layer construction module 10, a control point data acquisition module 20, an initial motion vector acquisition module 30, a zero-layer motion vector acquisition module 40, a first-layer motion vector acquisition module 50, and a second-layer motion vector output module 60.

[0178] Layer construction module 10 is used to construct a three-layer image downsampling pyramid; the three layers include a zero-layer image, a first-layer image, and a second-layer image.

[0179] Control point data acquisition module 20 is used to acquire the motion vector of the control point of an original block on the zero-layer image;

[0180] The initial motion vector acquisition module 30 is used to reduce the motion vector of the control point to obtain the initial motion vector of the original block on the zero-layer image;

[0181] The zero-layer motion vector acquisition module 40 is used to update the initial motion vector and obtain the final motion vector of a prediction block on the zero-layer image;

[0182] The first-layer motion vector acquisition module 50 is used to update based on the final motion vector on the zero-layer image and acquire the final motion vector of a prediction block on the first-layer image;

[0183] The second-layer motion vector output module 60 is used to update the final motion vector of a prediction block on the second-layer image based on the final motion vector on the first-layer image.

[0184] It should be noted that the affine motion estimation device provided in the above embodiments is only illustrated by the division of the functional modules described above when executing an affine motion estimation method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the affine motion estimation device and the affine motion estimation method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0185] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0186] In this embodiment, the affine motion estimation device first constructs a three-layer image downsampling pyramid to obtain the control point motion vector of an original block on the zero-layer image. Then, the control point motion vector is reduced to obtain the initial motion vector of the original block on the zero-layer image. Next, the initial motion vector is updated to obtain the final motion vector of a predicted block on the zero-layer image. Then, the final motion vector on the zero-layer image is updated to obtain the final motion vector of a predicted block on the first-layer image. Finally, the final motion vector on the first-layer image is updated to output the final motion vector of a predicted block on the second-layer image. This application is based on affine motion estimation using a three-layer image downsampling pyramid. The affine motion estimation process requires only three iterations, which reduces the amount of data access during motion compensation, thereby reducing memory access costs.

[0187] The present invention also provides a computer-readable medium having program instructions stored thereon, which, when executed by a processor, implement the affine motion estimation methods provided in the above-described method embodiments.

[0188] The present invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform an affine motion estimation method according to the various method embodiments described above.

[0189] Please see Figure 7 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Figure 7 As shown, terminal 1000 may include: at least one processor 1001, at least one network interface 1004, user interface 1003, memory 1005, and at least one communication bus 1002.

[0190] The communication bus 1002 is used to realize the connection and communication between these components.

[0191] The user interface 1003 may include a display screen and a camera. Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.

[0192] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0193] The processor 1001 may include one or more processing cores. The processor 1001 connects to various parts within the electronic device 1000 using various interfaces and lines. It executes various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. Optionally, the processor 1001 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1001 may integrate one or more of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 1001.

[0194] The memory 1005 may include random access memory (RAM) or read-only memory. Optionally, the memory 1005 may include a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 7 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for analyzing the availability of vehicle trajectory data.

[0195] exist Figure 7In the terminal 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and to obtain the user's input data; while the processor 1001 can be used to call the affine motion estimation application stored in the memory 1005 and specifically perform the following operations:

[0196] Construct a three-layer image downsampling pyramid; the three layers include the zeroth layer image, the first layer image, and the second layer image.

[0197] Obtain the motion vector of the control point of an original block on the zero-level image;

[0198] The motion vectors of the control points are reduced to obtain the initial motion vectors of the original blocks on the zero-layer image;

[0199] Update based on the initial motion vector to obtain the final motion vector of the prediction block on the zero-level image;

[0200] Update based on the final motion vector on the zero-level image, and obtain the final motion vector of a prediction block on the first-level image;

[0201] The final motion vector on the first layer image is updated, and the final motion vector of a predicted block on the second layer image is output.

[0202] In one embodiment, when the processor 1001 performs the update based on the initial motion vector to obtain the final motion vector of a prediction block on the zero-level image, it specifically performs the following operations:

[0203] Calculate the update amount of the initial motion vector on the zero-layer image based on the initial motion vector;

[0204] Based on the initial motion vector and the update amount of the initial motion vector, obtain the final motion vector of a prediction block on the zero-level image.

[0205] In one embodiment, when the processor 1001 performs the update of the initial motion vector on the zero-layer image based on the initial motion vector, it specifically performs the following operations:

[0206] Based on the initial motion vector, a prediction block is obtained on the zero-th layer image;

[0207] An error block is formed on the zero-level image based on the predicted block and the original block on the zero-level image;

[0208] Obtain the error at any position on the error block of the zero-level image;

[0209] The pixel gradient at any location on the prediction block in the zero-level image is calculated using the Sobel operator;

[0210] The update amount of the initial motion vector is calculated based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0211] In one embodiment, when the processor 1001 performs the update based on the final motion vector on the zero-layer image to obtain the final motion vector of a prediction block on the first-layer image, it specifically performs the following operations:

[0212] The update amount of the first layer motion vector on the first layer image is calculated based on the final motion vector on the zero layer image;

[0213] Based on the final motion vector on the zero-level image and the update amount of the motion vector on the first level, obtain the final motion vector of a prediction block on the first level image.

[0214] In one embodiment, when the processor 1001 performs the operation of calculating the update amount of the first layer motion vector on the first layer image based on the final motion vector on the zero layer image, it specifically performs the following operations:

[0215] Based on the final motion vector on the zero-level image, a prediction block is obtained on the first-level image;

[0216] An error block is formed on the first layer image based on the prediction block on the first layer image and the prediction block on the zeroth layer image;

[0217] Obtain the error at any position on the error block in the first layer image;

[0218] The pixel gradient at any location on the prediction block in the first-layer image is calculated using the Sobel operator;

[0219] The update amount of the motion vector of the first layer is calculated based on the error at any position on the error block in the first layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0220] In one embodiment, when the processor 1001 performs the operation of updating based on the final motion vector on the first layer image and outputting the final motion vector of a predicted block on the second layer image, it specifically performs the following operations:

[0221] The update amount of the second-layer motion vector in the second-layer image is calculated based on the final motion vector in the first-layer image;

[0222] Based on the final motion vector on the first layer image and the update amount of the motion vector on the second layer image, output the final motion vector of a prediction block on the second layer image.

[0223] In one embodiment, when the processor 1001 calculates the update amount of the second-layer motion vector on the second-layer image based on the final motion vector on the first-layer image, it specifically performs the following operations:

[0224] Based on the final motion vector on the first layer image, a predicted block on the second layer image is obtained;

[0225] Error blocks in the second layer image are formed based on the prediction blocks in the second layer image and the prediction blocks in the first layer image;

[0226] Obtain the error at any position on the error block in the second layer image;

[0227] The pixel gradient at any location on the prediction block in the second-layer image is calculated using the Sobel operator;

[0228] The update amount of the motion vector of the second layer is calculated based on the error at any position on the error block in the second layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

[0229] In this embodiment, the affine motion estimation method first constructs a three-layer image downsampling pyramid to obtain the control point motion vector of an original block on the zero-layer image. Then, the control point motion vector is reduced to obtain the initial motion vector of the original block on the zero-layer image. Next, the initial motion vector is updated to obtain the final motion vector of a predicted block on the zero-layer image. Then, the final motion vector on the zero-layer image is updated to obtain the final motion vector of a predicted block on the first-layer image. Finally, the final motion vector on the first-layer image is updated to output the final motion vector of a predicted block on the second-layer image. This application is based on affine motion estimation using a three-layer image downsampling pyramid. The affine motion estimation process requires only three iterations, which reduces the amount of data access during motion compensation, thereby reducing memory access costs.

[0230] Those skilled in the art will understand that implementing all or part of the processes in the above embodiments can be accomplished by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0231] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An affine motion estimation method, characterized in that, Includes the following steps: Construct a three-layer image downsampling pyramid; the three layers of images include a zero-layer image, a first-layer image, and a second-layer image. Obtain the motion vector of the control point of an original block on the zeroth layer image; The motion vector of the control point is reduced to obtain the initial motion vector of the original block on the zeroth layer image; Update the initial motion vector to obtain the final motion vector of a prediction block on the zeroth layer image; Update the image based on the final motion vector on the zeroth layer image to obtain the final motion vector of a prediction block on the first layer image; The step of updating based on the final motion vector on the zeroth layer image to obtain the final motion vector of a prediction block on the first layer image includes: The final motion vector on the zeroth layer image is determined as the initial motion vector of an original block in the first layer image. The final motion vector of a prediction block on the first layer image is determined based on the initial motion vector of the original block; Update the image based on the final motion vector on the first layer image, and output the final motion vector of a prediction block on the second layer image; The step of updating based on the initial motion vector to obtain the final motion vector of a prediction block on the zeroth layer image includes: Calculate the update amount of the initial motion vector on the zeroth layer image based on the initial motion vector; Based on the initial motion vector and the update amount of the initial motion vector, obtain the final motion vector of a prediction block on the zeroth layer image; The step of calculating the update amount of the initial motion vector on the zeroth layer image based on the initial motion vector includes: Based on the initial motion vector, a prediction block is obtained on the zeroth layer image; An error block is formed on the zeroth layer image based on the predicted block and the original block on the zeroth layer image; Obtain the error at any position on the error block of the zeroth layer image; The pixel gradient at any location of the prediction block on the zeroth layer image is calculated using the Sobel operator; The update amount of the initial motion vector is calculated based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

2. The affine motion estimation method according to claim 1, characterized in that, The step of updating based on the final motion vector on the zeroth layer image to obtain the final motion vector of a prediction block on the first layer image includes: The update amount of the first layer motion vector on the first layer image is calculated based on the final motion vector on the zero layer image; Based on the final motion vector on the zero-layer image and the update amount of the motion vector on the first layer, the final motion vector of a prediction block on the first layer image is obtained.

3. The affine motion estimation method according to claim 2, characterized in that, The step of calculating the update amount of the first layer motion vector on the first layer image based on the final motion vector on the zeroth layer image includes: Based on the final motion vector on the zeroth layer image, a prediction block on the first layer image is obtained; An error block is formed on the first layer image based on the prediction block on the first layer image and the prediction block on the zeroth layer image; Obtain the error at any position on the error block of the first layer image; The pixel gradient at any location of the prediction block on the first layer image is calculated using the Sobel operator; The update amount of the motion vector of the first layer is calculated based on the error at any position on the error block in the first layer image, the pixel gradient at any position of the prediction block, and the position coefficient.

4. The affine motion estimation method according to claim 1, characterized in that, The step of updating based on the final motion vector on the first layer image and outputting the final motion vector of a predicted block on the second layer image includes: The update amount of the second layer motion vector in the second layer image is calculated based on the final motion vector in the first layer image; Based on the final motion vector on the first layer image and the update amount of the second layer motion vector, output the final motion vector of a prediction block on the second layer image.

5. The affine motion estimation method according to claim 4, characterized in that, The step of calculating the update amount of the second layer motion vector in the second layer image based on the final motion vector in the first layer image includes: Based on the final motion vector on the first layer image, a prediction block on the second layer image is obtained; An error block is formed on the second layer image based on the prediction block on the second layer image and the prediction block on the first layer image; Obtain the error at any position on the error block in the second layer image; The pixel gradient at any location of the prediction block on the second-layer image is calculated using the Sobel operator; The update amount of the motion vector of the second layer is calculated based on the error at any position on the error block in the second layer image, the pixel gradient at any position of the prediction block, and the position coefficient.

6. An affine motion estimation device, characterized in that, include: A layer construction module is used to construct a three-layer image downsampling pyramid; the three layers include a zero-layer image, a first-layer image, and a second-layer image. The control point data acquisition module is used to acquire the motion vector of the control point of an original block on the zeroth layer image; An initial motion vector acquisition module is used to reduce the motion vector of the control point to obtain the initial motion vector of the original block on the zeroth layer image; The zero-layer motion vector acquisition module is used to update the initial motion vector and obtain the final motion vector of a prediction block on the zero-layer image; The first-layer motion vector acquisition module is used to update based on the final motion vector on the zero-layer image to obtain the final motion vector of a prediction block on the first-layer image; The first layer motion vector acquisition module is specifically used for: The final motion vector on the zeroth layer image is determined as the initial motion vector of an original block in the first layer image. The final motion vector of a prediction block on the first layer image is determined based on the initial motion vector of the original block; The second-layer motion vector output module is used to update the final motion vector of a prediction block on the first-layer image based on the final motion vector on the first-layer image, and output the final motion vector of a prediction block on the second-layer image. The zeroth layer motion vector acquisition module is specifically used for: Calculate the update amount of the initial motion vector on the zeroth layer image based on the initial motion vector; Based on the initial motion vector and the update amount of the initial motion vector, obtain the final motion vector of a prediction block on the zeroth layer image; The zeroth layer motion vector acquisition module is further specifically used for: Based on the initial motion vector, a prediction block is obtained on the zeroth layer image; An error block is formed on the zeroth layer image based on the predicted block and the original block on the zeroth layer image; Obtain the error at any position on the error block of the zeroth layer image; The pixel gradient at any location of the prediction block on the zeroth layer image is calculated using the Sobel operator; The update amount of the initial motion vector is calculated based on the error at any position on the error block in the zero-layer image, the pixel gradient at any position in the prediction block, and the position coefficient.

7. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method steps as claimed in any one of claims 1-5.

8. A terminal, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the method steps as claimed in any one of claims 1-5.

Citation Information

Patent Citations

  • Self-adaptively selecting global motion estimation method for panoramic video coding

    CN101771878A

  • Affine motion estimation method, device and equipment and storage medium

    CN113630601A