Picture encoding and decoding method and apparatus for video sequence
Patent Information
- Application Number
- ZA202103941
- Authority / Receiving Office
- ZA · ZA
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-30
- Filing Date
- 2021-06-08
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2039-07-29
AI Technical Summary
The computational complexity of DMVR technology in video encoding and decoding is too high, resulting in low efficiency.
By reducing the number of decoded or encoded prediction blocks obtained by motion search of the prediction reference image block, and sampling the pixel points in the image block, the calculation amount of the difference comparison is reduced, and different methods are used to calculate the difference value, including absolute value and sum of squares. Improve image encoding and decoding efficiency.
The computational complexity is reduced, the efficiency of image encoding and decoding is improved, and the amount of calculation and implementation complexity are reduced.
Abstract
Description
Image encoding and decoding method and apparatus for video sequences
[0001] This application claims priority to Chinese Patent Application No. 201811458677.4, filed on November 30, 2018, entitled "Image Encoding and Decoding Method and Apparatus for Video Sequences", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to video image technology, and more particularly to an image encoding and decoding method and apparatus for video sequences. Background Technology
[0003] In video coding and decoding frameworks, hybrid coding structures are commonly used for encoding and decoding video sequences. The encoding end of a hybrid coding structure typically includes a prediction module, a transform module, a quantization module, and an entropy coding module; the decoding end typically includes an entropy decoding module, an inverse quantization module, an inverse transform module, and a prediction compensation module. The combination of these encoding and decoding modules effectively removes redundant information from the video sequence and ensures that the encoded image of the video sequence is obtained at the decoding end. In video coding and decoding frameworks, the images in a video sequence are usually divided into image blocks for encoding; that is, a frame is divided into several image blocks, and the aforementioned modules are used for encoding and decoding based on these image blocks.
[0004] In the above modules, the prediction module and the prediction compensation module can adopt two techniques: intra-frame prediction and inter-frame prediction. In the inter-frame prediction technique, in order to effectively remove redundant information of the current image block in the image to be predicted, decoder-side motion vector refinement (DMVR) is used.
[0005] However, the computational complexity of DMVR technology is too high.
[0006] Summary of the Invention
[0007] This application provides an image encoding and decoding method and apparatus for video sequences to reduce the computational load of difference comparison and improve image encoding efficiency.
[0008] In a first aspect, this application provides an image decoding method for a video sequence, comprising:
[0009] The motion information of the block to be decoded is determined, including predicted motion vectors and reference image information, which is used to identify the predicted reference image block. A first decoding prediction block is obtained based on the motion information. A motion search with a first precision is performed within the predicted reference image block to obtain at least two second decoding prediction blocks, the search position of which is determined by the predicted motion vectors and the first precision. The first decoding prediction block is downsampled to obtain a first sampled pixel matrix. At least two second decoding prediction blocks are downsampled to obtain at least two second sampled pixel matrices. The difference between the first sampled pixel matrix and each second sampled pixel matrix is calculated, and the motion vector between the second decoding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be decoded is taken as the target predicted motion vector of the block to be decoded. The target decoding prediction block of the block to be decoded is obtained based on the target predicted motion vector, and the block to be decoded is then decoded based on the target decoding prediction block.
[0010] This application reduces the computational load of difference comparison and improves image decoding efficiency by reducing the number of second decoded prediction blocks obtained from motion search of the predicted reference image blocks and sampling the pixels in the image blocks.
[0011] In one possible implementation, the difference between the first sampled pixel matrix and each second sampled pixel matrix is calculated separately, including:
[0012] Calculate the sum of the absolute values of the pixel differences at corresponding positions in the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the absolute values of these pixel differences as the variance; or,
[0013] Calculate the sum of squares of the pixel differences between corresponding pixels in the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of squares of the pixel differences as the variance.
[0014] The advantage of this approach is that it uses different methods to determine the difference value, providing a variety of solutions with varying degrees of complexity.
[0015] In one possible implementation, at least two second decoded prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. Motion search with a first precision is performed within the prediction reference image block to obtain at least two second decoded prediction blocks, including:
[0016] When the difference corresponding to the upper block is less than the difference corresponding to the base block, the lower block is excluded from at least two second decoded prediction blocks. The base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on the first prediction vector, the position pointed to by the first prediction vector being obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset; the lower block is obtained within the prediction reference image block based on the second prediction vector, the position pointed to by the second prediction vector being obtained by offsetting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by a first precision; or...
[0017] If the difference corresponding to the current block is less than the difference corresponding to the base block, exclude the previous block from at least two second decoded prediction blocks; or...
[0018] When the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from at least two second decoded prediction blocks. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the right by a target offset. Alternatively...
[0019] When the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from at least two second decoded prediction blocks.
[0020] In one possible implementation, performing a motion search with first precision within a predicted reference image block to obtain at least two second decoded prediction blocks further includes:
[0021] When the next block and the right block are excluded, the upper left block is obtained within the prediction reference image block based on the fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and shifting it upward by the target offset. The upper left block is one of at least two second decoding prediction blocks.
[0022] When the next block and the left block are excluded, the upper right block is obtained within the prediction reference image block based on the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset. The upper right block is one of at least two second decoding prediction blocks.
[0023] When the upper block and the left block are excluded, the lower right block is obtained within the prediction reference image block according to the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset. The lower right block is one of at least two second decoding prediction blocks.
[0024] When the upper and right blocks are excluded, the lower left block is obtained within the prediction reference image block based on the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset. The lower left block is one of at least two second decoding prediction blocks.
[0025] The advantage of this approach is that it reduces the number of search points and lowers the complexity of implementation.
[0026] In one possible implementation, at least two second decoded prediction blocks include an upper block, a lower block, a left block, and a right block. Motion search with a first precision is performed within the predicted reference image block to obtain at least two second decoded prediction blocks, including:
[0027] When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the upper left block is obtained within the prediction reference image block based on the fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset and then offsetting it upwards by a target offset. The upper left block is one of at least two second decoding prediction blocks. The upper block is obtained within the prediction reference image block based on the first prediction vector, and the position pointed to by the first prediction vector is obtained by offsetting the position pointed to by the prediction motion vector upwards by a target offset. The lower block is obtained within the prediction reference image block based on the second prediction vector, and the position pointed to by the second prediction vector is obtained by offsetting the position pointed to by the prediction motion vector downwards by a target offset. The left block is obtained within the prediction reference image block based on the third prediction vector, and the position pointed to by the third prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on the fourth prediction vector, and the position pointed to by the fourth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset. The target offset is determined by the first precision. Alternatively...
[0028] When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the predicted reference image block based on the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset and upward by the target offset. The upper right block is one of at least two second decoded prediction blocks; or...
[0029] When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the lower right block is obtained within the predicted reference image block based on the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset and then offsetting it downwards by the target offset. The lower right block is one of at least two second decoded prediction blocks; or...
[0030] When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained within the prediction reference image block according to the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and downward by the target offset. The lower left block is one of at least two second decoding prediction blocks.
[0031] The advantage of this approach is that it reduces the number of search points and lowers the complexity of implementation.
[0032] In one possible implementation, after determining the motion information of the block to be decoded, the following steps are also included:
[0033] Round the motion vector down or up.
[0034] The advantage of this approach is that it eliminates the need for reference image interpolation when the motion vector is not an integer, thus reducing the complexity of implementation.
[0035] In one possible implementation, after performing a motion search with first precision within the predicted reference image block to obtain at least two second decoded prediction blocks, the method further includes:
[0036] The predicted motion vectors corresponding to at least two second decoding prediction blocks are rounded down or up.
[0037] The advantage of this approach is that it eliminates the need for reference image interpolation when the motion vector is not an integer, thus reducing the complexity of implementation.
[0038] Secondly, this application provides an image encoding method for a video sequence, including:
[0039] The motion information of the block to be encoded is determined, including predicted motion vectors and reference image information, which is used to identify the predicted reference image block. A first coded prediction block is obtained based on the motion information. A motion search with a first precision is performed within the predicted reference image block to obtain at least two second coded prediction blocks, the search position of which is determined by the predicted motion vectors and the first precision. The first coded prediction block is downsampled to obtain a first sampled pixel matrix. At least two second coded prediction blocks are downsampled to obtain at least two second sampled pixel matrices. The difference between the first sampled pixel matrix and each second sampled pixel matrix is calculated, and the motion vector between the second coded prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be encoded is taken as the target predicted motion vector of the block to be encoded. The target coded prediction block is obtained based on the target predicted motion vector, and the block to be encoded is encoded based on the target coded prediction block.
[0040] This application reduces the computational load of difference comparison and improves image coding efficiency by reducing the number of second coding prediction blocks obtained by motion search of the prediction reference image blocks and sampling the pixels in the image blocks.
[0041] In one possible implementation, the difference between the first sampled pixel matrix and each second sampled pixel matrix is calculated separately, including:
[0042] Calculate the sum of the absolute values of the pixel differences at corresponding positions in the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the absolute values of these pixel differences as the variance; or,
[0043] Calculate the sum of squares of the pixel differences between corresponding pixels in the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of squares of the pixel differences as the variance.
[0044] In one possible implementation, at least two second coded prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. A first-precision motion search is performed within the prediction reference image block to obtain at least two second coded prediction blocks, including:
[0045] When the difference corresponding to the upper block is less than the difference corresponding to the base block, the lower block is excluded from at least two second-coded prediction blocks. The base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on the first prediction vector, the position pointed to by the first prediction vector being obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset; the lower block is obtained within the prediction reference image block based on the second prediction vector, the position pointed to by the second prediction vector being obtained by offsetting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by a first precision; or...
[0046] If the difference corresponding to the current block is less than the difference corresponding to the base block, exclude the previous block from at least two second-coded prediction blocks; or...
[0047] When the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from at least two second-coded prediction blocks. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the right by a target offset. Alternatively...
[0048] When the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from at least two second-coded prediction blocks.
[0049] In one possible implementation, performing a motion search with first precision within a predicted reference image block to obtain at least two second coded prediction blocks further includes:
[0050] When the next block and the right block are excluded, the upper left block is obtained within the prediction reference image block based on the fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and shifting it upward by the target offset. The upper left block is one of at least two second coding prediction blocks.
[0051] When the next block and the left block are excluded, the upper right block is obtained within the prediction reference image block based on the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset. The upper right block is one of at least two second coding prediction blocks.
[0052] When the upper block and the left block are excluded, the lower right block is obtained within the prediction reference image block according to the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset. The lower right block is one of at least two second coding prediction blocks.
[0053] When the upper and right blocks are excluded, the lower left block is obtained within the prediction reference image block based on the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset. The lower left block is one of at least two second-coded prediction blocks.
[0054] In one possible implementation, at least two second coded prediction blocks include an upper block, a lower block, a left block, and a right block. A first-precision motion search is performed within the prediction reference image block to obtain at least two second coded prediction blocks, including:
[0055] When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the upper left block is obtained within the prediction reference image block based on the fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset and upwards by a target offset. The upper left block is one of at least two second-coded prediction blocks. The upper block is obtained within the prediction reference image block based on the first prediction vector, and the position pointed to by the first prediction vector is obtained by shifting the position pointed to by the prediction motion vector upwards by a target offset. The lower block is obtained within the prediction reference image block based on the second prediction vector, and the position pointed to by the second prediction vector is obtained by shifting the position pointed to by the prediction motion vector downwards by a target offset. The left block is obtained within the prediction reference image block based on the third prediction vector, and the position pointed to by the third prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on the fourth prediction vector, and the position pointed to by the fourth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the right by a target offset. The target offset is determined by the first precision. Alternatively...
[0056] When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the predicted reference image block based on the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset and upward by the target offset. The upper right block is one of at least two second-coded prediction blocks; or...
[0057] When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the lower right block is obtained within the predicted reference image block based on the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset and downwards by the target offset. The lower right block is one of at least two second-coded prediction blocks; or...
[0058] When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained within the prediction reference image block based on the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and downward by the target offset. The lower left block is one of at least two second coding prediction blocks.
[0059] In one possible implementation, after determining the motion information of the block to be encoded, the following steps are also included:
[0060] Round the motion vector down or up.
[0061] In one possible implementation, after performing a motion search with first precision within the predicted reference image block to obtain at least two second coded prediction blocks, the method further includes:
[0062] The predicted motion vectors corresponding to at least two second-coded prediction blocks are rounded down or up.
[0063] Thirdly, this application provides an image decoding apparatus for a video sequence, comprising:
[0064] The determination module is used to determine the motion information of the block to be decoded. The motion information includes predicted motion vectors and reference image information, which is used to identify the predicted reference image block.
[0065] The processing module is configured to: obtain a first decoding prediction block of the block to be decoded based on motion information; perform a motion search with a first precision within the prediction reference image block to obtain at least two second decoding prediction blocks, wherein the search position of the motion search is determined by the predicted motion vector and the first precision; downsample the first decoding prediction block to obtain a first sampled pixel matrix; downsample at least two second decoding prediction blocks to obtain at least two second sampled pixel matrices; calculate the difference between the first sampled pixel matrix and each second sampled pixel matrix, and take the motion vector between the second decoding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be decoded as the target predicted motion vector of the block to be decoded;
[0066] The decoding module is used to obtain the target decoding prediction block of the block to be decoded based on the target predicted motion vector, and to decode the block to be decoded based on the target decoding prediction block.
[0067] This application reduces the computational load of difference comparison and improves image decoding efficiency by reducing the number of second decoded prediction blocks obtained from motion search of the predicted reference image blocks and sampling the pixels in the image blocks.
[0068] In one possible implementation, the processing module is specifically used to calculate the sum of the absolute values of the pixel differences between corresponding positions of the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of the absolute values of the pixel differences as the difference; or, to calculate the sum of the squares of the pixel differences between corresponding positions of the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of the squares of the pixel differences as the difference.
[0069] In one possible implementation, at least two second decoding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to exclude the lower block from the at least two second decoding prediction blocks when the difference corresponding to the upper block is less than the difference corresponding to the base block. Specifically, the base block is obtained within the prediction reference image block based on a predicted motion vector; the upper block is obtained within the prediction reference image block based on a first predicted vector, the position pointed to by the first predicted vector being obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset; and the lower block is obtained within the prediction reference image block based on a second predicted vector, the position pointed to by the second predicted vector being obtained by offsetting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by a first precision. Alternatively, if the difference corresponding to the next block is less than the difference corresponding to the base block, the previous block is excluded from at least two second decoded prediction blocks; or, if the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from at least two second decoded prediction blocks, wherein the left block is obtained within the prediction reference image block based on the third prediction vector, and the position pointed to by the third prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset; the right block is obtained within the prediction reference image block based on the fourth prediction vector, and the position pointed to by the fourth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the right by a target offset; or, if the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from at least two second decoded prediction blocks.
[0070] In one possible implementation, the processing module is further configured to, when the next block and the right block are excluded, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector, wherein the position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset and then offsetting it upwards by a target offset, and the upper-left block is one of at least two second decoded prediction blocks; and when the next block and the left block are excluded, obtain an upper-right block within the prediction reference image block based on a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset and then offsetting it upwards by a target offset, and the upper-right block is one of at least two second decoded prediction blocks. One of the decoding prediction blocks; when the upper block and the left block are excluded, the lower right block is obtained within the prediction reference image block according to the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset. The lower right block is one of at least two second decoding prediction blocks; when the upper block and the right block are excluded, the lower left block is obtained within the prediction reference image block according to the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset. The lower left block is one of at least two second decoding prediction blocks.
[0071] In one possible implementation, at least two second decoding prediction blocks include an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset and then offsetting it upwards by a target offset. The upper-left block is one of at least two second decoding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by offsetting the position pointed to by the prediction motion vector upwards by a target offset. The lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by offsetting the position pointed to by the prediction motion vector downwards by a target offset. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset. The displacement is determined by a first precision; or, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the prediction reference image block according to the sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset, and the upper right block is one of at least two second decoding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the prediction reference image block according to the seventh prediction vector. The lower right block is obtained by offsetting the position of the seventh prediction vector to the right and downward by the target offset from the position of the predicted motion vector, and is one of at least two second decoding prediction blocks; or, if the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained in the prediction reference image block according to the eighth prediction vector, wherein the position of the eighth prediction vector is obtained by offsetting the position of the predicted motion vector to the left and downward by the target offset, and the lower left block is one of at least two second decoding prediction blocks.
[0072] In one possible implementation, the processing module is also used to round the motion vector down or up.
[0073] In one possible implementation, the processing module is further configured to round down or up the predicted motion vectors corresponding to at least two second decoding prediction blocks respectively.
[0074] Fourthly, this application provides an image encoding apparatus for a video sequence, comprising:
[0075] The determination module is used to determine the motion information of the block to be encoded. The motion information includes the predicted motion vector and reference image information. The reference image information is used to identify the predicted reference image block.
[0076] The processing module is configured to: obtain a first coding prediction block for the block to be encoded based on motion information; perform a motion search with a first precision within the prediction reference image block to obtain at least two second coding prediction blocks, wherein the search position for the motion search is determined by the predicted motion vector and the first precision; downsample the first coding prediction block to obtain a first sampled pixel matrix; downsample at least two second coding prediction blocks to obtain at least two second sampled pixel matrices; calculate the difference between the first sampled pixel matrix and each second sampled pixel matrix, and take the motion vector between the second coding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be encoded as the target predicted motion vector of the block to be encoded;
[0077] The encoding module is used to obtain the target encoding prediction block of the block to be encoded based on the target predicted motion vector, and to encode the block to be encoded based on the target encoding prediction block.
[0078] This application reduces the computational load of difference comparison and improves image coding efficiency by reducing the number of second coding prediction blocks obtained by motion search of the prediction reference image blocks and sampling the pixels in the image blocks.
[0079] In one possible implementation, the processing module is specifically used to calculate the sum of the absolute values of the pixel differences between corresponding positions of the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of the absolute values of the pixel differences as the difference; or, to calculate the sum of the squares of the pixel differences between corresponding positions of the first sampled pixel matrix and each second sampled pixel matrix, and use the sum of the squares of the pixel differences as the difference.
[0080] In one possible implementation, at least two second coding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to exclude the lower block from the at least two second coding prediction blocks when the difference corresponding to the upper block is less than the difference corresponding to the base block. Specifically, the base block is obtained within the prediction reference image block based on a predicted motion vector; the upper block is obtained within the prediction reference image block based on a first predicted vector, the position pointed to by the first predicted vector being obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset; and the lower block is obtained within the prediction reference image block based on a second predicted vector, the position pointed to by the second predicted vector being obtained by offsetting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by a first precision. Alternatively, if the difference corresponding to the next block is less than the difference corresponding to the base block, the previous block is excluded from at least two second-coded prediction blocks; or, if the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from at least two second-coded prediction blocks, wherein the left block is obtained within the prediction reference image block based on the third prediction vector, and the position pointed to by the third prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset; the right block is obtained within the prediction reference image block based on the fourth prediction vector, and the position pointed to by the fourth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset; or, if the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from at least two second-coded prediction blocks.
[0081] In one possible implementation, the processing module is further configured to, when the next block and the right block are excluded, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector, wherein the position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset and then offsetting it upwards by a target offset, and the upper-left block is one of at least two second coded prediction blocks; and when the next block and the left block are excluded, obtain an upper-right block within the prediction reference image block based on a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset and then offsetting it upwards by a target offset, and the upper-right block is one of at least two second coded prediction blocks. One of the coded prediction blocks; when the upper block and the left block are excluded, the lower right block is obtained within the prediction reference image block according to the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset. The lower right block is one of at least two second coded prediction blocks; when the upper block and the right block are excluded, the lower left block is obtained within the prediction reference image block according to the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset. The lower left block is one of at least two second coded prediction blocks.
[0082] In one possible implementation, at least two second coding prediction blocks include an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset and then offsetting it upwards by a target offset. The upper-left block is one of at least two second coding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by offsetting the position pointed to by the prediction motion vector upwards by a target offset. The lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by offsetting the position pointed to by the prediction motion vector downwards by a target offset. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by offsetting the position pointed to by the prediction motion vector to the right by a target offset. The displacement is determined by a first precision; or, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the prediction reference image block according to the sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset, and the upper right block is one of at least two second coding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, the upper right block is obtained within the prediction reference image block according to the seventh prediction vector. The lower right block is obtained by offsetting the position of the seventh prediction vector to the right and downward by the target offset from the position of the predicted motion vector, and is one of at least two second-coded prediction blocks; or, if the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained in the prediction reference image block according to the eighth prediction vector, wherein the position of the eighth prediction vector is obtained by offsetting the position of the predicted motion vector to the left and downward by the target offset, and the lower left block is one of at least two second-coded prediction blocks.
[0083] In one possible implementation, the processing module is also used to round the motion vector down or up.
[0084] In one possible implementation, the processing module is further configured to round down or up the predicted motion vectors corresponding to at least two second coded prediction blocks respectively.
[0085] Fifthly, this application provides a codec device, comprising:
[0086] One or more controllers;
[0087] Memory, used to store one or more programs;
[0088] When one or more programs are executed by one or more controllers, causing one or more controllers to implement the method as described in either the first or second aspect above.
[0089] Sixthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, are used to perform the method of any one of the first or second aspects described above.
[0090] In a seventh aspect, this application provides a computer program that, when executed by a computer, performs the method of any one of the first or second aspects described above. Attached Figure Description
[0091] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0092] Figure 1 is a flowchart of an embodiment of the image decoding method for video sequences in this application;
[0093] Figure 2 is a flowchart of a second embodiment of the image decoding method for video sequences in this application;
[0094] Figure 3 is a flowchart of an embodiment of the image encoding method for video sequences in this application;
[0095] Figure 4 is a flowchart of Embodiment 2 of the image encoding method for video sequences in this application;
[0096] Figure 5 is a schematic diagram of the structure of an embodiment of the image decoding device for video sequences in this application;
[0097] Figure 6 is a structural schematic diagram of an embodiment of the image encoding device for video sequences in this application;
[0098] Figure 7 is a structural schematic diagram of an embodiment of the encoding / decoding device of this application;
[0099] Figure 8 is a structural schematic diagram of an embodiment of the encoding / decoding system of this application;
[0100] Figure 9 is an exemplary schematic diagram of the determination of the search point location in an embodiment of this application. Detailed Implementation
[0101] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0102] The process of image coding method for inter-frame prediction in DMVR technology is as follows:
[0103] 1. Determine the motion information of the block to be encoded. This motion information includes the predicted motion vector and reference image information. The reference image information is used to identify the predicted reference image block.
[0104] Here, the block to be encoded is the current image block to be encoded. The predicted motion vectors in the motion information include forward predicted motion vectors and backward predicted motion vectors. The reference image information includes the reference frame index information of the forward and backward predicted reference image blocks (i.e., the image order count (POC) of the forward and backward predicted reference image blocks). The predicted reference image block of the block to be encoded can be obtained based on the predicted motion vectors or reference image information of the block to be encoded.
[0105] 2. Based on the motion information, perform bidirectional prediction to obtain the initial coding prediction block of the block to be coded, and obtain the first coding prediction block of the block to be coded based on the initial coding prediction block.
[0106] When the Proof of Concept (POC) of the predicted reference image block is not equal to the POC of the block to be encoded, bidirectional prediction is performed on the block to be encoded based on motion information, including forward prediction and backward prediction. Forward prediction based on motion information involves predicting the block to be encoded using forward motion information to obtain a forward initial coding prediction block for the block to be encoded. This includes predicting the block to be encoded using the forward-predicted reference image block from the motion information, or predicting the block to be encoded using the forward-predicted motion vector from the motion information. Backward prediction based on motion information involves predicting the block to be encoded using backward motion information to obtain a backward initial coding prediction block for the block to be encoded. This includes predicting the block to be encoded using the backward-predicted reference image block from the motion information, or predicting the block to be encoded using the backward-predicted motion vector from the motion information.
[0107] When obtaining the first coding prediction block of the block to be encoded based on the initial coding prediction block, there are three implementation methods: One is to obtain the first coding prediction block of the block to be encoded by weighted summation of the forward and backward initial coding prediction blocks. Another is to use the forward initial coding prediction block as the first coding prediction block of the block to be encoded. A third is to use the backward initial coding prediction block as the first coding prediction block of the block to be encoded.
[0108] Third, perform a motion search with first precision within the predicted reference image block to obtain at least one second coded prediction block.
[0109] The search position for motion search is determined by the predicted motion vector and a first precision. In one feasible implementation, the position for motion search is the position around the location identified by the predicted motion vector within the coverage area of the first precision. For example, if the position identified by the predicted motion vector is (1, 1) and the first precision is 1 / 2 pixel precision, then the position for motion search includes (1, 1) and nine positions centered on (1, 1): above (1, 1.5), below (1, 0.5), to the left (0.5, 1), to the right (1.5, 1), to the upper right (1.5, 1.5), to the lower right (1.5, 0.5), to the upper left (0.5, 1.5), to the lower left (0.5, 0.5), and (1, 1).
[0110] A first-precision motion search can be performed on the forward-predicted reference image block based on the forward-predicted motion vector. The forward-coded prediction block obtained in each search is used as the forward second-coded prediction block, resulting in at least one second-coded prediction block. Then, a first-precision motion search can be performed on the backward-predicted reference image block based on the backward-predicted motion vector. The backward-coded prediction block obtained in each search is used as the backward second-coded prediction block, resulting in at least one second-coded prediction block. The second-coded prediction block includes both the forward second-coded prediction block and the backward second-coded prediction block. The first precision can be an integer pixel precision, a 1 / 2 pixel precision, a 1 / 4 pixel precision, or a 1 / 8 pixel precision, and is not limited to any particular precision.
[0111] Fourth, calculate the difference between the first coding prediction block and each second coding prediction block, and take the motion vector between the second coding prediction block with the smallest difference and the block to be coded as the target prediction motion vector of the block to be coded.
[0112] Each forward second-order coding prediction block is compared with the first-order coding prediction block. The forward motion vector between the forward second-order coding prediction block with the smallest difference and the block to be encoded is taken as the target forward-order prediction motion vector. Similarly, each backward second-order coding prediction block is compared with the first-order coding prediction block. The backward motion vector between the backward second-order coding prediction block with the smallest difference and the block to be encoded is taken as the target backward-order prediction motion vector. The target prediction motion vector includes both the target forward-order prediction motion vector and the target backward-order prediction motion vector. When performing the difference comparison, the sum of the absolute values of all pixel differences in the two image blocks can be used as the difference value between the second-order coding prediction block and the first-order coding prediction block, or the sum of the squares of all pixel differences in the two image blocks can be used as the difference value between the second-order coding prediction block and the first-order coding prediction block.
[0113] 5. Based on the target predicted motion vector, perform bidirectional prediction on the block to be encoded to obtain the third coding prediction block of the block to be encoded.
[0114] The forward prediction of the target motion vector is used to perform forward prediction on the block to be encoded to obtain the forward third coding prediction block of the block to be encoded. Then, the backward prediction of the target motion vector is used to perform backward prediction on the block to be encoded to obtain the backward third coding prediction block of the block to be encoded.
[0115] 6. Obtain the target coding prediction block of the block to be coded based on the third coding prediction block, and encode the block to be coded based on the target coding prediction block.
[0116] The target coding prediction block of the block to be coded can be obtained by weighted summation of the forward third coding prediction block and the backward third coding prediction block, or the forward third coding prediction block can be used as the target coding prediction block of the block to be coded, or the backward third coding prediction block can be used as the target coding prediction block of the block to be coded.
[0117] The process of image decoding using inter-frame prediction in DMVR technology is as follows:
[0118] 1. Determine the motion information of the block to be decoded. This motion information includes the predicted motion vector and reference image information, which is used to identify the predicted reference image block.
[0119] Here, the block to be decoded is the current image block to be decoded. The predicted motion vectors in the motion information include forward predicted motion vectors and backward predicted motion vectors. The reference image information includes the reference frame index information of the forward and backward predicted reference image blocks (i.e., the image order count (POC) of the forward and backward predicted reference image blocks). The predicted reference image block of the block to be decoded can be obtained based on the predicted motion vectors or reference image information of the block to be decoded.
[0120] 2. Based on the motion information, perform bidirectional prediction to obtain the initial decoding prediction block of the block to be decoded, and obtain the first decoding prediction block of the block to be decoded based on the initial decoding prediction block.
[0121] When the Proof of Concept (POC) corresponding to the predicted reference image block is not equal to the POC of the block to be decoded, bidirectional prediction is performed on the block to be decoded based on motion information, including forward prediction and backward prediction. Forward prediction based on motion information involves predicting the block to be decoded using forward motion information to obtain a forward initial decoding prediction block for the block to be decoded. This includes predicting the block to be decoded using the forward-predicted reference image block in the motion information, or predicting the block to be decoded using the forward-predicted motion vector in the motion information. Backward prediction based on motion information involves predicting the block to be decoded using backward motion information to obtain a backward initial decoding prediction block for the block to be decoded. This includes predicting the block to be decoded using the backward-predicted reference image block in the motion information, or predicting the block to be decoded using the backward-predicted motion vector in the motion information.
[0122] When obtaining the first decoded prediction block of the block to be decoded based on the initial decoded prediction block, there are three implementation methods: One is to obtain the first decoded prediction block of the block to be decoded by weighted summation of the forward and backward initial decoded prediction blocks. Another is to use the forward initial decoded prediction block as the first decoded prediction block of the block to be decoded. A third is to use the backward initial decoded prediction block as the first decoded prediction block of the block to be decoded.
[0123] Third, perform a motion search with first precision within the predicted reference image block to obtain at least one second decoding prediction block.
[0124] The search position for motion search is determined by the predicted motion vector and a first precision. In one feasible implementation, the position for motion search is the position around the location identified by the predicted motion vector within the coverage area of the first precision. For example, if the position identified by the predicted motion vector is (1, 1) and the first precision is 1 / 2 pixel precision, then the position for motion search includes (1, 1) and nine positions centered on (1, 1): above (1, 1.5), below (1, 0.5), to the left (0.5, 1), to the right (1.5, 1), to the upper right (1.5, 1.5), to the lower right (1.5, 0.5), to the upper left (0.5, 1.5), to the lower left (0.5, 0.5), and (1, 1).
[0125] A first-precision motion search can be performed on the forward-predicted reference image block based on the forward-predicted motion vector. The forward-decoded prediction block obtained in each search is used as the forward second-decoded prediction block, resulting in at least one second-decoded prediction block. Then, a first-precision motion search can be performed on the backward-predicted reference image block based on the backward-predicted motion vector. The backward-decoded prediction block obtained in each search is used as the backward second-decoded prediction block, resulting in at least one second-decoded prediction block. The second-decoded prediction block includes both the forward second-decoded prediction block and the backward second-decoded prediction block. The first precision can be an integer pixel precision, a 1 / 2 pixel precision, a 1 / 4 pixel precision, or a 1 / 8 pixel precision, and is not limited to any particular precision.
[0126] Fourth, calculate the difference between the first decoding prediction block and each second decoding prediction block, and take the motion vector between the second decoding prediction block with the smallest difference and the block to be decoded as the target prediction motion vector of the block to be decoded.
[0127] Each forward second-decoding prediction block is compared with the first-decoding prediction block for difference. The forward motion vector between the forward second-decoding prediction block with the smallest difference and the block to be decoded is taken as the target forward-predicted motion vector. Similarly, each backward second-decoding prediction block is compared with the first-decoding prediction block for difference. The backward motion vector between the backward second-decoding prediction block with the smallest difference and the block to be decoded is taken as the target backward-predicted motion vector. The target predicted motion vector includes both the target forward-predicted motion vector and the target backward-predicted motion vector. When performing the difference comparison, the sum of the absolute values of all pixel differences in the two image blocks can be used as the difference value between the second-decoding prediction block and the first-decoding prediction block, or the sum of the squares of all pixel differences in the two image blocks can be used as the difference value between the second-decoding prediction block and the first-decoding prediction block.
[0128] 5. Based on the target predicted motion vector, perform bidirectional prediction on the block to be decoded to obtain the third decoding prediction block of the block to be decoded.
[0129] The forward prediction of the target motion vector is used to perform forward prediction on the block to be decoded to obtain the forward third decoding prediction block of the block to be decoded. Then, the backward prediction of the target motion vector is used to perform backward prediction on the block to be decoded to obtain the backward third decoding prediction block of the block to be decoded.
[0130] 6. Obtain the target decoding prediction block of the block to be decoded based on the third decoding prediction block, and decode the block to be decoded based on the target decoding prediction block.
[0131] The target decoding prediction block of the block to be decoded can be obtained by weighted summation of the forward third decoding prediction block and the backward third decoding prediction block, or the forward third decoding prediction block can be used as the target decoding prediction block of the block to be decoded, or the backward third decoding prediction block can be used as the target decoding prediction block of the block to be decoded.
[0132] However, the computational complexity of the aforementioned DMVR technology is too high. This application provides an image encoding and decoding method for video sequences, which will be described in detail below with specific embodiments.
[0133] Figure 1 is a flowchart of an embodiment of the image decoding method for video sequences of this application. As shown in Figure 1, the method of this embodiment may include:
[0134] Step 101: Determine the motion information of the block to be decoded. The motion information includes the predicted motion vector and the reference image information. The reference image information is used to identify the predicted reference image block.
[0135] As mentioned above, the predicted motion vector obtained in step one of the image decoding method for inter-frame prediction in existing DMVR technology is a coordinate information. This coordinate information may represent a non-pixel point (located between two pixels), which requires interpolation between two adjacent pixels to obtain the non-pixel point.
[0136] This application directly rounds the predicted motion vector down or up and maps it to the position of the pixel, thus eliminating the need for interpolation calculations and reducing computational load and complexity.
[0137] Step 102: Obtain the first decoding prediction block of the block to be decoded based on the motion information.
[0138] Step 102 of this application is similar in principle to step two of the image decoding method for inter-frame prediction in the above-mentioned DMVR technology, and will not be described again here.
[0139] Step 103: Perform a motion search with first precision within the predicted reference image block to obtain at least two second decoded prediction blocks. The search position for the motion search is determined by the predicted motion vector and the first precision.
[0140] As described above, step three of the existing DMVR technology's inter-frame prediction image decoding method involves performing motion search within a first precision coverage area around a base block in the prediction reference image block to obtain at least nine second decoding prediction blocks. These nine second decoding prediction blocks include a base block, an upper block, a lower block, a left block, a right block, an upper-left block, an upper-right block, a lower-right block, and a lower-left block. The base block is obtained within the prediction reference image block based on predicted motion vectors; the upper block is obtained within the prediction reference image block based on a first prediction vector; the lower block is obtained within the prediction reference image block based on a second prediction vector; the left block is obtained within the prediction reference image block based on a third prediction vector; the right block is obtained within the prediction reference image block based on a fourth prediction vector; the upper-left block is obtained within the prediction reference image block based on a fifth prediction vector; the upper-right block is obtained within the prediction reference image block based on a sixth prediction vector; the lower-right block is obtained within the prediction reference image block based on a seventh prediction vector; and the lower-left block is obtained within the prediction reference image block based on an eighth prediction vector. The position pointed to by the first predicted vector is obtained by offsetting the position pointed to by the predicted motion vector upwards by the target offset. The position pointed to by the second predicted vector is obtained by offsetting the position pointed to by the predicted motion vector downwards by the target offset. The position pointed to by the third predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left by the target offset. The position pointed to by the fourth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset. The position pointed to by the fifth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left and upwards by the target offset. The position pointed to by the sixth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right and upwards by the target offset. The position pointed to by the seventh predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right and downwards by the target offset. The target offset is determined by the first precision.
[0141] The implementation of step 103 in this application may include the following implementation methods:
[0142] The first method is based on at least two second decode prediction blocks, including a base block, an upper block, a lower block, a left block, and a right block. If the difference corresponding to the upper block is less than the difference corresponding to the base block, the lower block is excluded from the at least two second decode prediction blocks. Alternatively, if the difference corresponding to the lower block is less than the difference corresponding to the base block, the upper block is excluded from the at least two second decode prediction blocks.
[0143] If the difference corresponding to the left block is less than the difference corresponding to the base block, exclude the right block from at least two second decoded prediction blocks. Alternatively, if the difference corresponding to the right block is less than the difference corresponding to the base block, exclude the left block from at least two second decoded prediction blocks.
[0144] This involves selecting one of the top and bottom blocks, and one of the left and right blocks. If the difference between these blocks is less than the difference between the base block and the top block, then there's no need to acquire blocks in the opposite direction. In this case, only three positions of the second decoded prediction block need to be acquired. The difference in this step refers to the difference between the second and first decoded prediction blocks. Once it's found that the difference between one of the top and bottom blocks, or one of the left and right blocks, is less than the difference between the base block, it means that these two blocks are closer to the block to be decoded than the base block, and therefore, it's unnecessary to acquire blocks in the opposite offset direction.
[0145] Furthermore, when the down block and the right block are excluded, a top-left block is obtained within the prediction reference image block, and the top-left block is one of at least two second-decoding prediction blocks. When the down block and the left block are excluded, a top-right block is obtained within the prediction reference image block, and the top-right block is one of at least two second-decoding prediction blocks. When the up block and the left block are excluded, a bottom-right block is obtained within the prediction reference image block, and the bottom-right block is one of at least two second-decoding prediction blocks. When the up block and the right block are excluded, a bottom-left block is obtained within the prediction reference image block, and the bottom-left block is one of at least two second-decoding prediction blocks.
[0146] That is, based on the second decoding prediction blocks at the two positions mentioned above, one of the top left block, top right block, bottom right block, and bottom left block can be obtained, so a total of only four positions of the second decoding prediction blocks need to be obtained.
[0147] The second method is based on at least two second decoding prediction blocks, including an upper block, a lower block, a left block, and a right block. When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a top-left block is obtained within the prediction reference image block, and the top-left block is one of at least two second decoding prediction blocks. Alternatively, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a top-right block is obtained within the prediction reference image block, and the top-right block is one of at least two second decoding prediction blocks. Alternatively, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a bottom-right block is obtained within the prediction reference image block, and the bottom-right block is one of at least two second decoding prediction blocks. Alternatively, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a bottom-left block is obtained within the prediction reference image block, and the bottom-left block is one of at least two second decoding prediction blocks.
[0148] First, obtain the second decoding prediction blocks at four positions. Then, based on these four second decoding prediction blocks, obtain one of the top-left block, top-right block, bottom-right block, and bottom-left block. In this way, only five positions of the second decoding prediction blocks need to be obtained.
[0149] Both methods in this application are based on the difference between the base block and its surrounding blocks. Once it is found that the difference in one of the two opposite directions is smaller than the difference in the other, or that the difference in one of the two directions is smaller than the difference in the base block, it is not necessary to obtain a second decoding prediction block in the other direction, or even the angular direction. In this way, the number of second decoding prediction blocks obtained is definitely less than the nine second decoding prediction blocks in the prior art. The reduction in the number of second decoding prediction blocks means a reduction in the computational amount and complexity of subsequent pixel sampling and difference calculation.
[0150] In one feasible implementation, based on the second method described above, the difference values of five second decoding prediction blocks (up, down, left, right, and one diagonal direction) are compared with the difference value of the base block. The second decoding prediction block with the smallest difference value is selected as a candidate prediction block. Eight new second decoding prediction blocks are generated from the candidate prediction blocks (up, down, left, right, and four diagonal directions). Since some of these eight new second decoding prediction blocks have already had their difference values calculated (as shown in Figure 9, the new second decoding prediction block to the right is the original second decoding prediction block, the new second decoding prediction block to the bottom is the original second decoding prediction block, and the new second decoding prediction block to the lower right diagonal direction is the original base block), only the new second decoding prediction blocks for which difference value calculation has not yet been performed are used for difference value calculation and compared with the difference values of the candidate prediction blocks.
[0151] In one feasible implementation, when the difference values of the four second decoding prediction blocks (up, down, left, and right) are compared with the difference value of the base block and all are greater than or equal to the difference value of the base block, the base block is selected as the prediction block, and the motion vector corresponding to the base block is taken as the target motion vector.
[0152] Furthermore, as mentioned above, in the existing DMVR technology, after obtaining at least one second decoding prediction block in step three of the inter-frame prediction image decoding method, even if the predicted motion vector corresponding to the base block represents a pixel, if the precision of the motion search is not an integer pixel, for example, 1 / 2 pixel precision, 1 / 4 pixel precision, or 1 / 8 pixel precision, then the predicted motion vector corresponding to the offset second decoding prediction block is very likely to represent a non-pixel point.
[0153] This application rounds down or up the predicted motion vectors corresponding to at least two second decoding prediction blocks and maps them to the positions of pixels, thus eliminating the need for interpolation calculations and reducing computational load and complexity.
[0154] Step 104: Downsample the first decoding prediction block to obtain a first sampled pixel matrix, and downsample at least two second decoding prediction blocks to obtain at least two second sampled pixel matrices.
[0155] As mentioned above, in the existing DMVR technology, the fourth step of the inter-frame prediction image decoding method is to use the sum of the absolute values of the pixel differences between all pixels in the first decoding prediction block and all pixels in each second decoding prediction block as the difference between the two image blocks. Alternatively, the sum of the squares of the pixel differences between all pixels in the first decoding prediction block and all pixels in each second decoding prediction block can be used as the difference between the two image blocks. However, this requires a large amount of computation.
[0156] Because of the interconnectedness of the images, the pixel values of adjacent pixels in an image are not significantly different, and abrupt changes in pixel values are rare. Therefore, this application samples all pixels in an image block, that is, selects only a subset of pixels from all pixels. For example, pixels are sampled alternately by row, column, or both, without specific limitations. As long as the first decoding prediction block and each second decoding prediction block use the same sampling rule, the number of pixels to be calculated is reduced, and the corresponding computational load is greatly reduced.
[0157] Step 105: Calculate the difference between the first sampled pixel matrix and each second sampled pixel matrix, and take the motion vector between the second decoding prediction block and the block to be decoded corresponding to the second sampled pixel matrix with the smallest difference as the target prediction motion vector of the block to be decoded.
[0158] Since pixels in the first decoding prediction block and each of the second decoding prediction blocks are sampled, when calculating the difference between the two image blocks, the sum of the absolute values of the pixel differences at corresponding positions in the first sampled pixel matrix and each of the second sampled pixel matrices can be used as the difference between the first decoding prediction block and each of the second decoding prediction blocks. Alternatively, the sum of the squares of the pixel differences at corresponding positions in the first sampled pixel matrix and each of the second sampled pixel matrices can be used as the difference between the first decoding prediction block and each of the second decoding prediction blocks. In other words, the calculation for all pixels in the two image blocks is transformed into a calculation for a subset of pixels in the two image blocks, significantly reducing the computational load.
[0159] Step 106: Obtain the target decoding prediction block of the block to be decoded based on the target predicted motion vector, and decode the block to be decoded based on the target decoding prediction block.
[0160] As mentioned above, steps five and six of the existing DMVR technology's inter-frame prediction image decoding method perform motion search in two major directions: forward and backward. Moreover, two target prediction motion vectors are obtained in the forward and backward directions respectively. The same calculation process needs to be performed twice, which inevitably results in a large amount of computation.
[0161] This application can obtain a symmetrical backward target prediction vector based on the forward target prediction motion vector obtained from the forward target prediction motion vector after completing the forward motion search. That is, the backward target decoding prediction block is obtained by using the same value as the forward target prediction motion vector in the backward direction. In this way, the existing target prediction motion vector is directly used in the backward direction, so there is no need to perform a lot of repeated calculations, thus reducing the amount of computation.
[0162] This application reduces the computational load of difference comparison and improves image decoding efficiency by reducing the number of second decoded prediction blocks obtained by motion search of the predicted reference image blocks and sampling pixels in the image blocks.
[0163] Figure 2 is a flowchart of a second embodiment of the image decoding method for video sequences in this application. As shown in Figure 2, the method of this embodiment may include:
[0164] Step 201: Determine the motion information of the block to be decoded. The motion information includes the predicted motion vector and the reference image information. The reference image information is used to identify the predicted reference image block.
[0165] Step 201 of this application is similar in principle to step 101 above, and will not be repeated here.
[0166] Step 202: Obtain the base block, upper block, lower block, left block, and right block within the predicted reference image block.
[0167] Step 203: When the differences corresponding to the top block, bottom block, left block, and right block are all greater than the differences of the base block, only the base block is retained.
[0168] Step 204: Use the motion vector between the base block and the block to be decoded as the target prediction motion vector of the block to be decoded.
[0169] This application is based on at least two second decoding prediction blocks, including a base block, an upper block, a lower block, a left block, and a right block. First, the base block, upper block, lower block, left block, and right block are obtained. Then, the differences between the other blocks and the base block are compared. When the differences corresponding to the upper block, lower block, left block, and right block are all greater than the difference of the base block, only the base block is retained. That is, only five positions of the second decoding prediction blocks are needed to directly determine the target prediction motion vector used for the block to be decoded. Once it is found that the differences corresponding to the upper block, lower block, left block, and right block are all greater than the differences corresponding to the base block, it means that the second decoding prediction blocks at the offset positions are not as close to the block to be decoded as the base block. Therefore, the motion vector between the base block and the block to be decoded is directly used as the target prediction motion vector for the block to be decoded.
[0170] Step 205: Obtain the target decoding prediction block of the block to be decoded based on the target predicted motion vector, and decode the block to be decoded based on the target decoding prediction block.
[0171] Step 205 of this application is similar in principle to step 106 above, and will not be repeated here.
[0172] This application reduces the computational load of difference comparison and improves image decoding efficiency by reducing the number of second decoded prediction blocks obtained from motion search of the predicted reference image blocks.
[0173] Figure 3 is a flowchart of an embodiment of the image encoding method for video sequences of this application. As shown in Figure 3, the method of this embodiment may include:
[0174] Step 301: Determine the motion information of the block to be encoded. The motion information includes the predicted motion vector and the reference image information. The reference image information is used to identify the predicted reference image block.
[0175] As mentioned above, the predicted motion vector obtained in step one of the image coding method for inter-frame prediction in existing DMVR technology is a coordinate information. This coordinate information may represent a non-pixel point (located between two pixels), which requires interpolation between two adjacent pixels to obtain the non-pixel point.
[0176] This application directly rounds the predicted motion vector down or up and maps it to the position of the pixel, thus eliminating the need for interpolation calculations and reducing computational load and complexity.
[0177] Step 302: Obtain the first coding prediction block of the block to be coded based on the motion information.
[0178] Step 302 of this application is similar in principle to step two of the image coding method for inter-frame prediction in the above-mentioned DMVR technology, and will not be described again here.
[0179] Step 303: Perform a motion search with first precision within the predicted reference image block to obtain at least two second coded prediction blocks. The search position for the motion search is determined by the predicted motion vector and the first precision.
[0180] As described above, step three of the existing DMVR technology's inter-frame prediction image coding method involves performing motion search within a first precision coverage area around a base block in the prediction reference image block to obtain at least nine second coding prediction blocks. These nine second coding prediction blocks include a base block, an upper block, a lower block, a left block, a right block, an upper-left block, an upper-right block, a lower-right block, and a lower-left block. The base block is obtained within the prediction reference image block based on predicted motion vectors; the upper block is obtained within the prediction reference image block based on a first prediction vector; the lower block is obtained within the prediction reference image block based on a second prediction vector; the left block is obtained within the prediction reference image block based on a third prediction vector; the right block is obtained within the prediction reference image block based on a fourth prediction vector; the upper-left block is obtained within the prediction reference image block based on a fifth prediction vector; the upper-right block is obtained within the prediction reference image block based on a sixth prediction vector; the lower-right block is obtained within the prediction reference image block based on a seventh prediction vector; and the lower-left block is obtained within the prediction reference image block based on an eighth prediction vector. The position pointed to by the first predicted vector is obtained by offsetting the position pointed to by the predicted motion vector upwards by the target offset. The position pointed to by the second predicted vector is obtained by offsetting the position pointed to by the predicted motion vector downwards by the target offset. The position pointed to by the third predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left by the target offset. The position pointed to by the fourth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset. The position pointed to by the fifth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left and upwards by the target offset. The position pointed to by the sixth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right and upwards by the target offset. The position pointed to by the seventh predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right and downwards by the target offset. The target offset is determined by the first precision.
[0181] The implementation of step 303 in this application may include the following implementation methods:
[0182] The first method is based on at least two second-order coding prediction blocks, including a base block, an upper block, a lower block, a left block, and a right block. If the difference corresponding to the upper block is less than the difference corresponding to the base block, the lower block is excluded from the at least two second-order coding prediction blocks. Alternatively, if the difference corresponding to the lower block is less than the difference corresponding to the base block, the upper block is excluded from the at least two second-order coding prediction blocks.
[0183] If the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from at least two second-coded prediction blocks. Alternatively, if the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from at least two second-coded prediction blocks.
[0184] This involves selecting one of the upper and lower blocks, and one of the left and right blocks. If the difference is less than the difference of the base block, then there's no need to acquire blocks in the opposite direction; only three positions of the second coding prediction block need to be acquired. The difference in this step refers to the difference between the second coding prediction block and the first coding prediction block. Once it's found that the difference between one of the upper and lower blocks, or one of the left and right blocks, is less than the difference of the base block, it means that these two blocks are closer to the block to be encoded than the base block, and therefore, it's unnecessary to acquire blocks in the opposite offset direction.
[0185] Furthermore, when the down block and the right block are excluded, a top-left block is obtained within the prediction reference image block, and the top-left block is one of at least two second-coded prediction blocks. When the down block and the left block are excluded, a top-right block is obtained within the prediction reference image block, and the top-right block is one of at least two second-coded prediction blocks. When the up block and the left block are excluded, a bottom-right block is obtained within the prediction reference image block, and the bottom-right block is one of at least two second-coded prediction blocks. When the up block and the right block are excluded, a bottom-left block is obtained within the prediction reference image block, and the bottom-left block is one of at least two second-coded prediction blocks.
[0186] That is, based on the second coding prediction blocks at the above two positions, one of the top left block, top right block, bottom right block and bottom left block can be obtained, so only four positions of the second coding prediction blocks need to be obtained in total.
[0187] The second method is based on at least two second coding prediction blocks, including an upper block, a lower block, a left block, and a right block. When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a top-left block is obtained within the prediction reference image block, and the top-left block is one of at least two second coding prediction blocks. Alternatively, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a top-right block is obtained within the prediction reference image block, and the top-right block is one of at least two second coding prediction blocks. Alternatively, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a bottom-right block is obtained within the prediction reference image block, and the bottom-right block is one of at least two second coding prediction blocks. Alternatively, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a bottom-left block is obtained within the prediction reference image block, and the bottom-left block is one of at least two second coding prediction blocks.
[0188] First, obtain the second coding prediction blocks at four positions. Then, based on these four second coding prediction blocks, obtain one of the top-left block, top-right block, bottom-right block, and bottom-left block. In this way, only five positions of the second coding prediction blocks need to be obtained.
[0189] Both methods in this application are based on the difference between the base block and its surrounding blocks. Once it is found that the difference in one of the two opposite directions is smaller than the difference in the other, or that the difference in one of the two directions is smaller than the difference in the base block, it is not necessary to obtain a second coding prediction block in the other direction, or even the angular direction. In this way, the number of second coding prediction blocks obtained is definitely less than the nine second coding prediction blocks in the prior art. The reduction in the number of second coding prediction blocks means a reduction in the computational amount and complexity of subsequent pixel sampling and difference calculation.
[0190] In one feasible implementation, based on the second method described above, the difference values of five second coding prediction blocks (up, down, left, right, and one diagonal direction) are compared with the difference values of the base block. The second coding prediction block with the smallest difference value is selected as a candidate prediction block. Eight new second coding prediction blocks (up, down, left, right, and four diagonal directions) are then generated from the candidate prediction blocks. Since some of these eight new second coding prediction blocks have already had their difference values calculated (as shown in Figure 9, the new second coding prediction block to the right is the original second coding prediction block, the new second coding prediction block to the bottom is the original second coding prediction block, and the new second coding prediction block to the lower right diagonal direction is the original base block), only the new second coding prediction blocks for which difference values have not yet been calculated are used for difference value calculation and compared with the difference values of the candidate prediction blocks.
[0191] In one feasible implementation, when the difference values of the four second coding prediction blocks (up, down, left, and right) are compared with the difference value of the base block and all are greater than or equal to the difference value of the base block, the base block is selected as the prediction block, and the motion vector corresponding to the base block is taken as the target motion vector.
[0192] Furthermore, as mentioned above, in the existing DMVR technology, after obtaining at least one second coding prediction block in step three of the inter-frame prediction image coding method, even if the predicted motion vector corresponding to the base block represents a pixel, if the precision of the motion search is not an integer pixel, for example, 1 / 2 pixel precision, 1 / 4 pixel precision, or 1 / 8 pixel precision, then the predicted motion vector corresponding to the offset second coding prediction block is very likely to represent a non-pixel point.
[0193] This application rounds down or up the predicted motion vectors corresponding to at least two second coding prediction blocks and maps them to the positions of pixels, thus eliminating the need for interpolation calculations and reducing computational load and complexity.
[0194] Step 304: Downsample the first coding prediction block to obtain a first sampled pixel matrix, and downsample at least two second coding prediction blocks to obtain at least two second sampled pixel matrices.
[0195] As mentioned above, in the existing DMVR technology, the fourth step of the inter-frame prediction image coding method is to use the sum of the absolute values of the pixel differences between all pixels in the first coding prediction block and all pixels in each second coding prediction block as the difference between the two image blocks. Alternatively, the sum of the squares of the pixel differences between all pixels in the first coding prediction block and all pixels in each second coding prediction block can be used as the difference between the two image blocks. However, this requires a large amount of computation.
[0196] Because of the interconnectedness of the images, the pixel values of adjacent pixels in an image are not significantly different, and abrupt changes in pixel values are rare. Therefore, this application samples all pixels in an image block, that is, selects only a subset of pixels from all pixels. For example, pixels are sampled alternately by row, column, or both, without specific limitations. As long as the first coding prediction block and each second coding prediction block use the same sampling rule, the number of pixels to be calculated is reduced, and the corresponding computational load is greatly reduced.
[0197] Step 305: Calculate the difference between the first sampled pixel matrix and each second sampled pixel matrix, and take the motion vector between the second coding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be coded as the target prediction motion vector of the block to be coded.
[0198] Since pixels in the first and each second coding prediction block are sampled, when calculating the difference between the two image blocks, the sum of the absolute values of the pixel differences at corresponding positions in the first and each second sampled pixel matrix can be used as the difference between the first and each second coding prediction block. Alternatively, the sum of the squares of the pixel differences at corresponding positions in the first and each second sampled pixel matrix can be used as the difference between the first and each second coding prediction block. This transforms the calculation from dealing with all pixels in the two image blocks to dealing with a subset of pixels in the two image blocks, significantly reducing the computational load.
[0199] Step 306: Obtain the target coding prediction block of the block to be encoded based on the target predicted motion vector, and encode the block to be encoded based on the target coding prediction block.
[0200] As mentioned above, in the existing DMVR technology, the steps five and six of the inter-frame prediction image coding method perform motion search in two major directions: forward and backward. Moreover, two target prediction motion vectors are obtained in the forward and backward directions respectively. The same calculation process needs to be performed twice, which inevitably results in a large amount of computation.
[0201] This application can obtain a symmetrical backward target prediction vector based on the forward target prediction motion vector obtained from the forward target prediction motion vector after completing the forward motion search. That is, the backward target coding prediction block is obtained by using the same value as the forward target prediction motion vector in the backward direction. In this way, the existing target prediction motion vector is directly used in the backward direction, so there is no need to perform a lot of repeated calculations, thus reducing the amount of computation.
[0202] This application reduces the computational load of difference comparison and improves image coding efficiency by reducing the number of second coded prediction blocks obtained from motion search of the prediction reference image blocks and sampling pixels in the image blocks.
[0203] Figure 4 is a flowchart of a second embodiment of the image encoding method for video sequences according to this application. As shown in Figure 4, the method of this embodiment may include:
[0204] Step 401: Determine the motion information of the block to be encoded. The motion information includes the predicted motion vector and the reference image information. The reference image information is used to identify the predicted reference image block.
[0205] Step 401 of this application is similar in principle to step 301 above, and will not be repeated here.
[0206] Step 402: Obtain the base block, upper block, lower block, left block, and right block within the predicted reference image block.
[0207] Step 403: When the differences corresponding to the top block, bottom block, left block, and right block are all greater than the differences of the base block, only the base block is retained.
[0208] Step 404: Use the motion vector between the base block and the block to be encoded as the target prediction motion vector of the block to be encoded.
[0209] This application is based on at least two second coding prediction blocks, including a base block, an upper block, a lower block, a left block, and a right block. First, the base block, upper block, lower block, left block, and right block are obtained. Then, the differences between the other blocks and the base block are compared. When the differences corresponding to the upper block, lower block, left block, and right block are all greater than the difference of the base block, only the base block is retained. That is, only five positions of the second coding prediction blocks are needed to directly determine the target prediction motion vector used for the block to be encoded. Once it is found that the differences corresponding to the upper block, lower block, left block, and right block are all greater than the differences corresponding to the base block, it means that the second coding prediction blocks at the offset positions are not as close to the block to be encoded as the base block. Therefore, the motion vector between the base block and the block to be encoded is directly used as the target prediction motion vector of the block to be encoded.
[0210] Step 405: Obtain the target coding prediction block of the block to be encoded based on the target predicted motion vector, and encode the block to be encoded based on the target coding prediction block.
[0211] Step 405 of this application is similar in principle to step 306 above, and will not be repeated here.
[0212] This application reduces the computational cost of difference comparison and improves image coding efficiency by reducing the number of second coded prediction blocks obtained from motion search of the predicted reference image blocks.
[0213] Figure 5 is a schematic diagram of the structure of an embodiment of the image decoding device for video sequences according to this application. As shown in Figure 5, the device of this embodiment may include: a determining module 11, a processing module 12, and a decoding module 13. The determining module 11 is used to determine the motion information of the block to be decoded, the motion information including a predicted motion vector and reference image information, the reference image information being used to identify a predicted reference image block. The processing module 12 is used to obtain a first decoding prediction block of the block to be decoded based on the motion information; and to perform a motion search with a first precision within the predicted reference image block to obtain at least two second decoding prediction blocks, the search position of which is determined by the predicted motion vector and the... First precision is determined; the first decoding prediction block is downsampled to obtain a first sampled pixel matrix; the at least two second decoding prediction blocks are downsampled to obtain at least two second sampled pixel matrices; the difference between the first sampled pixel matrix and each of the second sampled pixel matrices is calculated respectively, and the motion vector between the second decoding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be decoded is taken as the target predicted motion vector of the block to be decoded; the decoding module 13 is used to obtain the target decoding prediction block of the block to be decoded according to the target predicted motion vector, and decode the block to be decoded according to the target decoding prediction block.
[0214] Based on the above technical solution, the processing module 12 is specifically used to calculate the sum of the absolute values of the pixel differences between the corresponding positions of the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the absolute values of the pixel differences as the difference; or, to calculate the sum of the squares of the pixel differences between the corresponding positions of the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the squares of the pixel differences as the difference.
[0215] Based on the above technical solution, the at least two second decoding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The processing module 12 is specifically used to exclude the lower block from the at least two second decoding prediction blocks when the difference corresponding to the upper block is less than the difference corresponding to the base block. The base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by shifting the position pointed to by the predicted motion vector upwards by a target offset; the lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by shifting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by the first precision; or... When the difference corresponding to the lower block is less than the difference corresponding to the base block, the upper block is excluded from the at least two second decoding prediction blocks; or, when the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from the at least two second decoding prediction blocks, wherein the left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset, and the right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the right by the target offset; or, when the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from the at least two second decoding prediction blocks.
[0216] Based on the above technical solution, the processing module 12 is further configured to, when the lower block and the right block are excluded, obtain an upper-left block within the prediction reference image block according to a fifth prediction vector, wherein the position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset, and the upper-left block is one of the at least two second decoding prediction blocks; when the lower block and the left block are excluded, obtain an upper-right block within the prediction reference image block according to a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and upward by the target offset, and the upper-right block is one of the at least two second decoding prediction blocks. One of the second decoding prediction blocks; when the upper block and the left block are excluded, a lower right block is obtained within the prediction reference image block according to the seventh prediction vector, the position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset, and the lower right block is one of the at least two second decoding prediction blocks; when the upper block and the right block are excluded, a lower left block is obtained within the prediction reference image block according to the eighth prediction vector, the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset, and the lower left block is one of the at least two second decoding prediction blocks.
[0217] Based on the above technical solution, the at least two second decoding prediction blocks include an upper block, a lower block, a left block, and a right block. The processing module 12 is specifically used to obtain an upper-left block within the prediction reference image block based on a fifth prediction vector when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset. The upper-left block is one of the at least two second decoding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector. The position pointed to by the predicted motion vector is obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset. The lower block is obtained within the predicted reference image block based on the second predicted vector, and the position pointed to by the second predicted vector is obtained by offsetting the position pointed to by the predicted motion vector downwards by the target offset. The left block is obtained within the predicted reference image block based on the third predicted vector, and the position pointed to by the third predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left by a target offset. The right block is obtained within the predicted reference image block based on the fourth predicted vector, and the position pointed to by the fourth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset. The target offset is determined by the first precision. Determine; or, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, obtain the upper right block within the prediction reference image block according to the sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset, and the upper right block is one of the at least two second decoding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, obtain the upper right block within the prediction reference image block according to the seventh prediction vector. The lower right block is obtained by offsetting the position pointed to by the seventh prediction vector to the right and downward by the target offset from the position pointed to by the prediction motion vector, and the lower right block is one of the at least two second decoding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained in the prediction reference image block according to the eighth prediction vector, wherein the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left and downward by the target offset from the position pointed to by the prediction motion vector, and the lower left block is one of the at least two second decoding prediction blocks.
[0218] Based on the above technical solution, the processing module 12 is also used to round down or up the motion vector.
[0219] Based on the above technical solution, the processing module 12 is further configured to round down or up the predicted motion vectors corresponding to the at least two second decoding prediction blocks respectively.
[0220] The apparatus of this embodiment can be used to execute the technical solution of the method embodiment shown in FIG1 or FIG2. Its implementation principle and technical effect are similar, and will not be described again here.
[0221] Figure 6 is a structural schematic diagram of an embodiment of the image encoding device for video sequences according to this application. As shown in Figure 6, the device of this embodiment may include: a determining module 21, a processing module 22, and an encoding module 23. The determining module 21 is used to determine the motion information of the block to be encoded, the motion information including a predicted motion vector and reference image information, the reference image information being used to identify a predicted reference image block. The processing module 22 is used to obtain a first encoding prediction block of the block to be encoded based on the motion information; and to perform a motion search with a first precision within the predicted reference image block to obtain at least two second encoding prediction blocks, the search position of which is determined by the predicted motion vector and the... First precision determination; downsampling the first coding prediction block to obtain a first sampled pixel matrix; downsampling the at least two second coding prediction blocks to obtain at least two second sampled pixel matrices; calculating the difference between the first sampled pixel matrix and each of the second sampled pixel matrices, and taking the motion vector between the second coding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be encoded as the target predicted motion vector of the block to be encoded; encoding module 23, used to obtain the target coding prediction block of the block to be encoded according to the target predicted motion vector, and to encode the block to be encoded according to the target coding prediction block.
[0222] Based on the above technical solution, the processing module 22 is specifically used to calculate the sum of the absolute values of the pixel differences between the corresponding positions of the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the absolute values of the pixel differences as the difference; or, to calculate the sum of the squares of the pixel differences between the corresponding positions of the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the squares of the pixel differences as the difference.
[0223] Based on the above technical solution, the at least two second coding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The processing module 22 is specifically used to exclude the lower block from the at least two second coding prediction blocks when the difference corresponding to the upper block is less than the difference corresponding to the base block. The base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by shifting the position pointed to by the predicted motion vector upwards by a target offset; the lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by shifting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by the first precision; or... When the difference corresponding to the lower block is less than the difference corresponding to the base block, the upper block is excluded from the at least two second coding prediction blocks; or, when the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from the at least two second coding prediction blocks, wherein the left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset, and the right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset; or, when the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from the at least two second coding prediction blocks.
[0224] Based on the above technical solution, the processing module 22 is further configured to, when the lower block and the right block are excluded, obtain an upper-left block within the prediction reference image block according to a fifth prediction vector, wherein the position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset, and the upper-left block is one of the at least two second coded prediction blocks; when the lower block and the left block are excluded, obtain an upper-right block within the prediction reference image block according to a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and upward by the target offset, and the upper-right block is one of the at least two second coded prediction blocks. One of the second coding prediction blocks; when the upper block and the left block are excluded, a lower right block is obtained within the prediction reference image block according to the seventh prediction vector, the position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset, and the lower right block is one of the at least two second coding prediction blocks; when the upper block and the right block are excluded, a lower left block is obtained within the prediction reference image block according to the eighth prediction vector, the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset, and the lower left block is one of the at least two second coding prediction blocks.
[0225] Based on the above technical solution, the at least two second coding prediction blocks include an upper block, a lower block, a left block, and a right block. The processing module 22 is specifically used to obtain an upper-left block within the prediction reference image block based on a fifth prediction vector when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset. The upper-left block is one of the at least two second coding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector. The position pointed to by the predicted motion vector is obtained by offsetting the position pointed to by the predicted motion vector upwards by a target offset. The lower block is obtained within the predicted reference image block based on the second predicted vector, and the position pointed to by the second predicted vector is obtained by offsetting the position pointed to by the predicted motion vector downwards by the target offset. The left block is obtained within the predicted reference image block based on the third predicted vector, and the position pointed to by the third predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the left by a target offset. The right block is obtained within the predicted reference image block based on the fourth predicted vector, and the position pointed to by the fourth predicted vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset. The target offset is determined by the first precision. Determine; or, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, obtain an upper right block within the prediction reference image block based on a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it upward by the target offset, and the upper right block is one of the at least two second coding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, obtain a right block within the prediction reference image block based on a seventh prediction vector. The lower right block is obtained by offsetting the position pointed to by the seventh prediction vector to the right and downward by the target offset from the position pointed to by the prediction motion vector, and the lower right block is one of the at least two second coding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained in the prediction reference image block according to the eighth prediction vector, wherein the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left and downward by the target offset from the position pointed to by the prediction motion vector, and the lower left block is one of the at least two second coding prediction blocks.
[0226] Based on the above technical solution, the processing module 22 is also used to round down or up the motion vector.
[0227] Based on the above technical solution, the processing module 22 is further configured to round down or up the predicted motion vectors corresponding to the at least two second coding prediction blocks respectively.
[0228] The apparatus of this embodiment can be used to execute the technical solution of the method embodiment shown in FIG3 or FIG4. Its implementation principle and technical effect are similar, and will not be described again here.
[0229] Figure 7 is a structural schematic diagram of an embodiment of the encoding / decoding device of this application, and Figure 8 is a structural schematic diagram of an embodiment of the encoding / decoding system of this application. The units in Figures 7 and 8 will be described below.
[0230] The encoding / decoding device 30 may be, for example, a mobile terminal or user equipment of a wireless communication system. It should be understood that embodiments of this application may be implemented in any electronic device or apparatus that may require encoding and decoding of video images, or encoding or decoding.
[0231] The codec device 30 may include a housing for incorporating and protecting the device. The codec device 30 may also include a display 31 in the form of a liquid crystal display. In other embodiments of this application, the display 31 may be any suitable display technology suitable for displaying images or videos. The codec device 30 may also include a keypad 32. In other embodiments of this application, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or a data entry system as part of a touch-sensitive display. The codec device 30 may also include a microphone 33 or any suitable audio input, which may be a digital or analog signal input. The codec device 30 may also include an audio output device, which in embodiments of this application may be any of the following: headphones 34, a speaker, or an analog or digital audio output connection. The codec device 30 may also include a battery; in other embodiments of this application, the codec device 30 may be powered by any suitable mobile energy device, such as a solar cell, fuel cell, or clock generator. The codec device 30 may also include an infrared port 35 for near-line-of-sight communication with other devices. In other embodiments, the codec device 30 may also include any suitable short-range communication solution, such as Bluetooth wireless connectivity or USB / FireWire wired connectivity.
[0232] The encoding / decoding device 30 may include a controller 36 or a controller for controlling the encoding / decoding device 30. The controller 36 may be connected to a memory 37, which in this embodiment may store data in the form of images and audio data, and / or may also store instructions for implementation on the controller 36. The controller 36 may also be connected to a codec 38 suitable for encoding and decoding audio and / or video data, or for auxiliary encoding and decoding implemented by the controller 36.
[0233] The encoding / decoding device 30 may also include a card reader 39 and a smart card 40 for providing user information and suitable for providing authentication information for network authentication and authorization of users.
[0234] The codec device 30 may further include a radio interface 41, which is circuitically connected to the controller and adapted to generate wireless communication signals, for example, for communication with a cellular communication network, a wireless communication system, or a wireless local area network. The codec device 30 may also include an antenna 42, which is connected to the radio interface 41 for transmitting radio frequency signals generated at the radio interface 41 to other devices(s) and for receiving radio frequency signals from other devices(s).
[0235] In some embodiments of this application, the encoding / decoding device 30 includes a camera capable of recording or detecting single frames, and the encoder / decoder 38 or controller receives and processes these single frames. In some embodiments of this application, the device can receive video image data to be processed from another device before transmission and / or storage. In some embodiments of this application, the encoding / decoding device 30 can receive images for encoding / decoding via a wireless or wired connection.
[0236] The encoding / decoding system 50 includes an image encoding device 51 and an image decoding device 52. The image encoding device 51 generates encoded video data. Therefore, the image encoding device 51 can be referred to as an encoding device. The image decoding device 52 decodes the encoded video data generated by the image encoding device 51. Therefore, the image decoding device 52 can be referred to as a decoding device. The image encoding device 51 and the image decoding device 52 can be examples of encoding / decoding devices. The image encoding device 51 and the image decoding device 52 can include a wide range of devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, smartphones and other handheld devices, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.
[0237] Image decoding device 52 can receive encoded video data from image encoding device 51 via channel 53. Channel 53 may include one or more media and / or devices capable of moving encoded video data from image encoding device 51 to image decoding device 52. In one example, channel 53 may include one or more communication media enabling image encoding device 51 to transmit encoded video data directly to image decoding device 52 in real time. In this example, image encoding device 51 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to image decoding device 52. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network (e.g., the Internet)). The one or more communication media may include routers, switches, base stations, or other devices facilitating communication from image encoding device 51 to image decoding device 52.
[0238] In another example, channel 53 may include a storage medium for storing encoded video data generated by image encoding device 51. In this example, image decoding device 52 may access the storage medium via disk access or card access. The storage medium may include various local access data storage media, such as Blu-ray discs, DVDs, CD-ROMs, flash memory, or other suitable digital storage media for storing encoded video data.
[0239] In another example, channel 53 may include a file server or another intermediate storage device storing the encoded video data generated by image encoding device 51. In this example, image decoding device 52 may access the encoded video data stored on the file server or other intermediate storage device via streaming or downloading. The file server may be a server type capable of storing and transmitting the encoded video data to image decoding device 52. Example file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, and local disk drives.
[0240] The image decoding device 52 can access the encoded video data via a standard data connection (e.g., an Internet connection). Examples of data connection types include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on a file server. The transmission of the encoded video data from the file server can be streaming, downloading, or a combination of both.
[0241] The technology described in this application is not limited to wireless application scenarios. For example, it can be applied to video encoding and decoding to support various multimedia applications, including: over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding video data stored on data storage media, decoding video data stored on data storage media, or other applications. In some instances, the encoding / decoding system 50 can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0242] In the example of Figure 8, the image encoding device 51 includes a video source 54, a video encoder 55, and an output interface 56. In some examples, the output interface 56 may include a modulator / demodulator (modem) and / or a transmitter. The video source 54 may include a video capture device (e.g., a video camera), a video archive containing previously captured video data, a video input interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of the above video data sources.
[0243] The video encoder 55 can encode video data from the video source 54. In some instances, the image encoding device 51 transmits the encoded video data directly to the image decoding device 52 via the output interface 56. The encoded video data can also be stored on storage media or a file server for later access by the image decoding device 52 for decoding and / or playback.
[0244] In the example of Figure 8, the image decoding device 52 includes an input interface 57, a video decoder 58, and a display device 59. In some examples, the input interface 57 includes a receiver and / or a modem. The input interface 57 can receive encoded video data via channel 53. The display device 59 can be integrated with the image decoding device 52 or can be external to the image decoding device 52. Generally, the display device 59 displays the decoded video data. The display device 59 can include various display devices, such as liquid crystal displays (LCDs), plasma displays, organic light-emitting diode (OLED) displays, or other types of display devices.
[0245] The video encoder 55 and video decoder 58 operate according to video compression standards (e.g., the High Efficiency Video Codec H.265 standard) and conform to the HEVC test model (HM). The textual description of the H.265 standard, ITU-TH.265(V3)(04 / 2015), was published on April 29, 2015, and can be downloaded from http: / / handle.itu.int / 11.1002 / 1000 / 12455. The entire contents of that document are incorporated herein by reference.
[0246] Alternatively, the video encoder 55 and video decoder 59 may operate according to other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), which include Scalable Video Coding (SVC) and Multi-View Video Coding (MVC) extensions. It should be understood that the technology in this application is not limited to any particular codec standard or technology.
[0247] Furthermore, Figure 8 is merely an example, and the technology of this application can be applied to video encoding and decoding applications that may not necessarily involve any data communication between the encoding and decoding devices (e.g., one-sided video encoding or video decoding). In other examples, data is retrieved from local memory, streamed over a network, or manipulated in a similar manner. The encoding device may encode data and store the data in memory, and / or the decoding device may retrieve data from memory and decode the data. In many examples, encoding and decoding are performed by multiple devices that do not communicate with each other but only encode data to memory and / or retrieve and decode data from memory.
[0248] The video encoder 55 and video decoder 59 can each be implemented as any of a variety of suitable circuits, such as one or more microcontrollers, digital signal controllers (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented partly or entirely in software, the device can store software instructions in a suitable non-transitory computer-readable storage medium, and one or more controllers can be used to execute instructions in hardware to perform the technology of this application. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more controllers. Each of the video encoder 55 and video decoder 59 can be included in one or more encoders or decoders, and any of them can be integrated as part of a combined encoder / decoder (codec (CODEC)) in other devices.
[0249] This application generally refers to the video encoder 55 “signaling” information to another device (e.g., video decoder 59). The term “signaling” generally refers to the transmission of syntax elements and / or encoded video data. This transmission can occur in real-time or near real-time. Alternatively, this communication can occur over a time span, for example, when syntax elements are stored as encoded binary data on a computer-readable storage medium during encoding, and the syntax elements can then be retrieved by the decoding device at any time after being stored on this medium.
[0250] In one possible implementation, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the methods in any of the embodiments shown in Figures 1-4 above.
[0251] In one possible implementation, this application provides a computer program that, when executed by a computer, performs the methods in any of the embodiments shown in Figures 1-4 above.
[0252] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for image decoding of a video sequence, characterized in that, include: Determine the motion information of the block to be decoded, the motion information including predicted motion vectors and reference image information, the reference image information being used to identify the predicted reference image block; The first decoding prediction block of the block to be decoded is obtained based on the motion information; A motion search with a first precision is performed within the predicted reference image block to obtain at least two second decoded prediction blocks, wherein the search position of the motion search is determined by the predicted motion vector and the first precision. The first decoding prediction block is downsampled to obtain the first sampled pixel matrix; The at least two second decoding prediction blocks are downsampled to obtain at least two second sampled pixel matrices; Calculate the difference between the first sampled pixel matrix and each of the second sampled pixel matrices, and take the motion vector between the second decoding prediction block corresponding to the second sampled pixel matrix with the smallest difference and the block to be decoded as the target predicted motion vector of the block to be decoded. The target decoding prediction block of the block to be decoded is obtained based on the target predicted motion vector, and the block to be decoded is decoded based on the target decoding prediction block.
2. The method according to claim 1, characterized in that, The step of calculating the difference between the first sampled pixel matrix and each of the second sampled pixel matrices includes: Calculate the sum of the absolute values of the pixel differences at corresponding positions in the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of the absolute values of the pixel differences as the difference; or, Calculate the sum of squares of the pixel differences between corresponding pixels in the first sampled pixel matrix and each of the second sampled pixel matrices, and use the sum of squares of the pixel differences as the difference.
3. The method according to claim 1 or 2, characterized in that, The at least two second decoding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The step of performing a motion search with a first precision within the prediction reference image block to obtain at least two second decoding prediction blocks includes: When the difference corresponding to the upper block is less than the difference corresponding to the base block, the lower block is excluded from the at least two second decoding prediction blocks. The base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by shifting the position pointed to by the predicted motion vector upwards by a target offset; the lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by shifting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by the first precision; or... When the difference corresponding to the lower block is less than the difference corresponding to the base block, the upper block is excluded from the at least two second decoded prediction blocks; or... When the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from the at least two second decoded prediction blocks. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the right by the target offset. Alternatively... When the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from the at least two second decoding prediction blocks.
4. The method according to claim 3, characterized in that, The step of performing a motion search with a first precision within the predicted reference image block to obtain at least two second decoded prediction blocks further includes: When the lower block and the right block are excluded, the upper left block is obtained within the prediction reference image block according to the fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and shifting it upward by the target offset. The upper left block is one of the at least two second decoding prediction blocks. When the lower block and the left block are excluded, the upper right block is obtained within the prediction reference image block according to the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offset upward by the target offset. The upper right block is one of the at least two second decoding prediction blocks. When the upper block and the left block are excluded, the lower right block is obtained within the prediction reference image block according to the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset. The lower right block is one of the at least two second decoding prediction blocks. When the upper block and the right block are excluded, the lower left block is obtained within the prediction reference image block according to the eighth prediction vector. The position pointed to by the eighth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and shifting it downward by the target offset. The lower left block is one of the at least two second decoding prediction blocks.
5. The method according to claim 1 or 2, characterized in that, The at least two second decoding prediction blocks include an upper block, a lower block, a left block, and a right block. The step of performing a motion search with first precision within the prediction reference image block to obtain at least two second decoding prediction blocks includes: When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a top-left block is obtained within the prediction reference image block based on a fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset. The top-left block is one of the at least two second decoding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by shifting the position pointed to by the prediction motion vector upward by the target offset. The lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by shifting the position pointed to by the prediction motion vector downward by the target offset. The left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector being obtained by shifting the position pointed to by the prediction motion vector to the right by the target offset. The target offset is determined by the first precision. Alternatively... When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a right upper block is obtained within the prediction reference image block based on the sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and upward by the target offset. The right upper block is one of the at least two second decoding prediction blocks; or... When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a lower right block is obtained within the prediction reference image block based on the seventh prediction vector. The position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and downward by the target offset. The lower right block is one of the at least two second decoding prediction blocks; or... When the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, a lower left block is obtained within the prediction reference image block according to the eighth prediction vector, wherein the position pointed to by the eighth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and shifting it downward by the target offset, and the lower left block is one of the at least two second decoding prediction blocks.
6. The method according to any one of claims 1-5, characterized in that, After determining the motion information of the block to be decoded, the method further includes: The motion vector is rounded down or up.
7. The method according to any one of claims 1-6, characterized in that, After performing a motion search with first precision within the predicted reference image block to obtain at least two second decoded prediction blocks, the method further includes: The predicted motion vectors corresponding to the at least two second decoding prediction blocks are rounded down or up.
8. An image decoding apparatus for a video sequence, characterized in that, include: A determination module is used to determine the motion information of the block to be decoded, the motion information including a predicted motion vector and reference image information, the reference image information being used to identify the predicted reference image block; The processing module is configured to obtain a first decoding prediction block of the block to be decoded based on the motion information; perform a motion search with a first precision within the prediction reference image block to obtain at least two second decoding prediction blocks, wherein the search position of the motion search is determined by the predicted motion vector and the first precision; and downsample the first decoding prediction block to obtain a first sampled pixel matrix. The at least two second decoding prediction blocks are downsampled to obtain at least two second sampling pixel matrices; the difference between the first sampling pixel matrix and each of the second sampling pixel matrices is calculated respectively, and the motion vector between the second decoding prediction block corresponding to the second sampling pixel matrix with the smallest difference and the block to be decoded is taken as the target predicted motion vector of the block to be decoded; The decoding module is used to obtain the target decoding prediction block of the block to be decoded based on the target predicted motion vector, and to decode the block to be decoded based on the target decoding prediction block.
9. The apparatus according to claim 8, characterized in that, The processing module is specifically used to calculate the sum of the absolute values of the pixel differences between the corresponding positions of the first sampled pixel matrix and each of the second sampled pixel matrices, and to use the sum of the absolute values of the pixel differences as the difference. Alternatively, the sum of squares of the pixel differences between corresponding pixels in the first sampled pixel matrix and each of the second sampled pixel matrices can be calculated, and the sum of squares of the pixel differences can be used as the difference.
10. The apparatus according to claim 8 or 9, characterized in that, The at least two second decoding prediction blocks include a base block, an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to exclude the lower block from the at least two second decoding prediction blocks when the difference corresponding to the upper block is less than the difference corresponding to the base block. Specifically, the base block is obtained within the prediction reference image block based on the predicted motion vector; the upper block is obtained within the prediction reference image block based on a first prediction vector, the position pointed to by the first prediction vector being obtained by shifting the position pointed to by the predicted motion vector upwards by a target offset; and the lower block is obtained within the prediction reference image block based on a second prediction vector, the position pointed to by the second prediction vector being obtained by shifting the position pointed to by the predicted motion vector downwards by a target offset, the target offset being determined by the first precision. Alternatively, when the lower block corresponds to… When the difference is less than the difference corresponding to the base block, the upper block is excluded from the at least two second decoding prediction blocks; or, when the difference corresponding to the left block is less than the difference corresponding to the base block, the right block is excluded from the at least two second decoding prediction blocks, wherein the left block is obtained within the prediction reference image block based on a third prediction vector, the position pointed to by the third prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by a target offset, and the right block is obtained within the prediction reference image block based on a fourth prediction vector, the position pointed to by the fourth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset; or, when the difference corresponding to the right block is less than the difference corresponding to the base block, the left block is excluded from the at least two second decoding prediction blocks.
11. The apparatus according to claim 10, characterized in that, The processing module is further configured to, when the lower block and the right block are excluded, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector, wherein the position pointed to by the fifth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset, and the upper-left block is one of the at least two second decoding prediction blocks; and when the lower block and the left block are excluded, obtain an upper-right block within the prediction reference image block based on a sixth prediction vector, wherein the position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and upward by the target offset, and the upper-right block is one of the at least two second decoding prediction blocks. One of the at least two second decoding prediction blocks; when the upper block and the left block are excluded, a lower right block is obtained within the prediction reference image block according to the seventh prediction vector, the position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and offsetting it downward by the target offset, and the lower right block is one of the at least two second decoding prediction blocks; when the upper block and the right block are excluded, a lower left block is obtained within the prediction reference image block according to the eighth prediction vector, the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the left by the target offset and offsetting it downward by the target offset, and the lower left block is one of the at least two second decoding prediction blocks.
12. The apparatus according to claim 8 or 9, characterized in that, The at least two second decoding prediction blocks include an upper block, a lower block, a left block, and a right block. The processing module is specifically configured to, when the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the left block is less than the difference corresponding to the right block, obtain an upper-left block within the prediction reference image block based on a fifth prediction vector. The position pointed to by the fifth prediction vector is obtained by shifting the position pointed to by the prediction motion vector to the left by the target offset and upward by the target offset. The upper-left block is one of the at least two second decoding prediction blocks. The upper block is obtained within the prediction reference image block based on a first prediction vector, and the position pointed to by the first prediction vector is determined by the prediction... The lower block is obtained by offsetting the position pointed to by the motion vector upwards by a target offset. The lower block is obtained within the prediction reference image block based on a second prediction vector, the position of which is obtained by offsetting the position pointed to by the predicted motion vector downwards by the target offset. The left block is obtained within the prediction reference image block based on a third prediction vector, the position of which is obtained by offsetting the position pointed to by the predicted motion vector to the left by a target offset. The right block is obtained within the prediction reference image block based on a fourth prediction vector, the position of which is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset. The target offset is determined by the first precision; or, when... When the difference corresponding to the upper block is less than the difference corresponding to the lower block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a right upper block is obtained within the prediction reference image block based on a sixth prediction vector. The position pointed to by the sixth prediction vector is obtained by offsetting the position pointed to by the prediction motion vector to the right by the target offset and then offsetting it upwards by the target offset. The right upper block is one of the at least two second decoding prediction blocks. Alternatively, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the right block is less than the difference corresponding to the left block, a right lower block is obtained within the prediction reference image block based on a seventh prediction vector. In the above, the position pointed to by the seventh prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the right by the target offset and downward by the target offset, and the lower right block is one of the at least two second decoding prediction blocks; or, when the difference corresponding to the lower block is less than the difference corresponding to the upper block, and the difference corresponding to the left block is less than the difference corresponding to the right block, the lower left block is obtained in the prediction reference image block according to the eighth prediction vector, wherein the position pointed to by the eighth prediction vector is obtained by offsetting the position pointed to by the predicted motion vector to the left by the target offset and downward by the target offset, and the lower left block is one of the at least two second decoding prediction blocks.
13. The apparatus according to any one of claims 8-12, characterized in that, The processing module is also used to round the motion vector down or up.
14. The apparatus according to any one of claims 8-13, characterized in that, The processing module is also used to round down or up the predicted motion vectors corresponding to the at least two second decoding prediction blocks respectively.