Image prediction method and apparatus, and codec
The image prediction method addresses high complexity in video compression by determining optimal reference blocks with mirror or proportional offset relationships, improving accuracy and reducing complexity in inter-frame prediction.
Patent Information
- Application Number
- JP2023070921
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-31
- Filing Date
- 2023-04-24
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2038-12-27
AI Technical Summary
Existing video compression technologies face challenges in reducing complexity while improving image prediction accuracy, particularly in inter-frame prediction processes where high search complexity is associated with template matching for motion vector refinement.
An image prediction method that determines optimal reference blocks in forward and backward reference images based on a mirror or proportional relationship of position offsets, avoiding template matching and simplifying the prediction process.
Improves image prediction accuracy and reduces complexity by determining optimal reference blocks without template matching, enhancing encoding performance.
Smart Images

Figure 0007712977000001 
Figure 0007712977000002 
Figure 0007712977000003
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding technology, and specifically, to a method and apparatus for image prediction, and a codec.
Background Art
[0002] By using video compression technologies such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC), and video compression technologies described in extensions of these standards, digital video information can be efficiently transmitted and received between devices. Generally, images in a video sequence are divided into image blocks for encoding or decoding.
[0003] In video compression technology, spatial prediction (intra prediction) and / or temporal prediction (inter prediction) based on image blocks are introduced to reduce or remove redundant information in a video sequence. The inter prediction mode may include, but is not limited to, the Merge Mode, the non-Merge Mode (for example, the advanced motion vector prediction mode (AMVP mode)), and the like. All inter predictions are performed by using a multi-motion information contention method.
[0004] In the inter-frame prediction process, a candidate motion information list (abbreviated as candidate list), which includes a plurality of groups of motion information (also referred to as a plurality of candidate motion information), is introduced. For example, the encoder may use the motion information (e.g., motion vector) of the current image block to be encoded or a group of motion information selected from the candidate list to predict it, to obtain a reference image block (i.e., reference sample) of the current image block to be encoded. Correspondingly, the decoder may decode the bitstream to obtain the instruction information and may obtain a group of motion information. Since the overhead of encoding motion information (i.e., the overhead of bits in the occupied bitstream) is limited in the inter-frame prediction process, this affects the accuracy of motion information to a certain extent and further affects the image prediction accuracy.
[0005] To improve the image prediction accuracy, existing decoder-side motion vector refinement (DMVR) techniques can be used to refine the motion information. However, when a DMVR solution is used for image prediction, a template matching block needs to be calculated, and the template matching block needs to be used to separately perform a search matching process in the forward reference image and the backward reference image. As a result, the search complexity becomes relatively high. Therefore, how to reduce the complexity during image prediction while improving the image prediction accuracy is a problem that needs to be solved. Summary of the Invention
[0006] Embodiments of the present application provide a method and apparatus for image prediction, and corresponding encoders and decoders, to improve the image prediction accuracy, reduce the image prediction complexity to a certain extent, and further improve the encoding performance. Means for Solving the Problems
[0007] According to a first aspect, an embodiment of the present application provides an image prediction method. The method includes obtaining initial motion information of a current image block, and determining positions of N forward reference blocks and positions of N backward reference blocks based on the initial motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1, determining, from positions of M pairs of reference blocks based on a matching cost criterion, that positions of a pair of reference blocks are positions of a target forward reference block of the current image block and a target backward reference block of the current image block, where positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block, and for positions of each pair of reference blocks, a first position offset and a second position offset are in a mirror image relationship, the first position offset represents an offset of the position of the forward reference block with respect to the position of an initial forward reference block, the second position offset represents an offset of the position of the backward reference block with respect to the position of an initial backward reference block, M is an integer greater than or equal to 1, and M is less than or equal to N, determining, and obtaining a predicted value of a pixel value of the current image block based on pixel values (samples) of the target forward reference block and pixel values (samples) of the target backward reference block.
[0008] In this embodiment of the present application, it should be particularly noted that positions of the N forward reference blocks include the position of one initial forward reference block and positions of (N - 1) candidate forward reference blocks, and positions of the N backward reference blocks include the position of one initial backward reference block and positions of (N - 1) candidate backward reference blocks. Therefore, the offset of the position of the initial forward reference block with respect to the position of the initial forward reference block is 0, and the offset of the position of the initial backward reference block with respect to the position of the initial backward reference block is 0. The offset 0 and the offset 0 also satisfy the condition of the mirror image relationship.
[0009] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. For the positions of each pair of reference blocks among the positions of the N pairs of reference blocks, there is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, the pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0010] Furthermore, it should be understood that the current image block (referred to as the current block) in this specification can be understood as the image block currently being processed. For example, in the encoding process, the current image block is an encoding block. In the decoding process, the current image block is Decryption a decoding block.
[0011] Furthermore, it should be understood that the reference block in this specification is a block that provides a reference signal for the current block. In the search process, it is necessary to traverse multiple reference blocks to find the optimal reference block. The reference block arranged in the forward reference image is referred to as the forward reference block. The reference block arranged in the backward reference image is referred to as the backward reference block.
[0012] Furthermore, it should be understood that the block that provides a prediction for the current block is referred to as the prediction block. For example, after traversing multiple reference blocks, the optimal reference block can be found. The optimal reference block provides a prediction for the current block and is referred to as the prediction block. The pixel value, sampling value, or sampling signal in the prediction block is referred to as the prediction signal.
[0013] Furthermore, it should be understood that the matching cost criterion in this specification is understood as a criterion for considering the matching cost between the paired forward reference block and the backward reference block. The matching cost may be understood as the difference between two blocks and may be considered as the cumulative difference of samples at corresponding positions in the two blocks. The difference is usually calculated based on the SAD (sum of absolute difference) criterion or another criterion, such as SATD (Sum of Absolute Transform Difference), MR-SAD (mean-removed sum of absolute difference), or SSD (sum of squared differences).
[0014] Furthermore, it should be noted that the initial motion information of the current image block in this embodiment of the present application may include a motion vector MV and reference picture indication information. Certainly, the initial motion information may alternatively include either one of the motion vector or the reference picture indication information, or both the motion vector and the reference picture indication information. For example, when both the encoder side and the decoder side are in agreement regarding the reference picture, the initial motion information may include only the motion vector MV. The reference picture indication information is used to indicate which one or more reconstructed images are used as the reference picture for the current block. The motion vector indicates the offset of the position of the reference block in the used reference picture with respect to the position of the current block, and generally includes a horizontal component offset and a vertical component offset. For example, (x, y) is used to represent the MV, where x represents the position offset in the horizontal direction and y represents the position offset in the vertical direction. The position of the reference block of the current block in the reference picture can be obtained by adding the MV to the position of the current block. The reference picture indication information may include a reference picture list and / or a reference picture index corresponding to the reference picture list. The reference picture index is used to identify the reference picture corresponding to the used motion vector in the specified reference picture list (RefPicList0 or RefPicList1). An image may sometimes be called a frame, and a reference image may sometimes be called a reference frame.
[0015] In this embodiment of the present application, the initial motion information of the current image block is initial bidirectional prediction motion information, that is, it includes the motion information used in the forward prediction direction and the motion information used in the backward prediction direction. In this specification, the forward and backward prediction directions are the two prediction directions of the bidirectional prediction mode. "Forward" and "backward" may be understood to correspond to reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1) of the current image, respectively.
[0016] Furthermore, it should be noted that the position of the initial forward reference block in this embodiment of the present application is the position of the reference block in the forward reference image, and is also the position obtained by adding the position of the current block to the offset represented by the initial MV. The position of the initial backward reference block in this embodiment of the present application is the position of the reference block in the backward reference image, and is also the position obtained by adding the position of the current block to the offset represented by the initial MV.
[0017] It should be understood that the method in this embodiment of the present application can be executed by an image prediction device. For example, the method can be executed by a video encoder, a video decoder, or an electronic device having a video encoding function. For example, the method can be particularly executed by an inter-frame prediction unit in a video encoder or a motion compensation unit in a video decoder.
[0018] Regarding the first aspect, in some implementations of the first aspect, the fact that the first position offset and the second position offset are in a mirror image relationship can be understood as the first position offset value being the same as the second position offset value. For example, the direction of the first position offset (also referred to as the vector direction) is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset.
[0019] In one example, the first position offset includes a first horizontal component offset and a first vertical component offset, and the second position offset includes a second horizontal component offset and a second vertical component offset. The direction of the first horizontal component offset is opposite to the direction of the second horizontal component offset, and the amplitude value of the first horizontal component offset is the same as the amplitude value of the second horizontal component offset. The direction of the first vertical component offset is opposite to the direction of the second vertical component offset, and the amplitude value of the first vertical component offset is the same as the amplitude value of the second vertical component offset.
[0020] In another example, both the first position offset and the second position offset are 0.
[0021] Regarding the first aspect, in some implementations of the first aspect, the method is to obtain updated motion information of the current image block, where the updated motion information further includes an updated forward motion vector and an updated backward motion vector, the updated forward motion vector points to the position of a target forward reference block, and the updated backward motion vector points to the position of a target backward reference block.
[0022] In a different example, the updated motion information of the current image block is obtained based on the position of the target forward reference block, the position of the target backward reference block, and the position of the current image block, or is obtained based on a first position offset and a second position offset corresponding to the determined positions of a pair of reference blocks.
[0023] It can be seen that the refined motion information of the current image block can be obtained in this embodiment of the present application. This improves the accuracy of the motion information of the current image block and also smooths the prediction of another image block. For example, it improves the prediction accuracy of the motion information of another image block.
[0024] Regarding the first aspect, in some implementations of the first aspect, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, or the positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0025] Note that the positions of the N pairs of reference blocks include the positions of the paired initial forward reference block and initial backward reference block, and the positions of the paired candidate forward reference block and candidate backward reference block. The offset of the position of the candidate forward reference block with respect to the position of the initial forward reference block in the forward reference image is in a mirror image relationship with the offset of the position of the candidate backward reference block with respect to the position of the initial backward reference block in the backward reference image.
[0026] Regarding the first aspect, in some implementations of the first aspect, the initial motion information includes forward prediction motion information and backward prediction motion information. Determining the positions of the N forward reference blocks and the N backward reference blocks based on the initial motion information and the position of the current image block is to determine the positions of the N forward reference blocks in the forward reference image based on the forward prediction motion information and the position of the current image block, where the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, and is to determine the positions of the N backward reference blocks in the backward reference image based on the backward prediction motion information and the position of the current image block, where the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0027] Regarding the first aspect, in some implementations of the first aspect, the initial motion information includes a first motion vector and a first reference image index in the forward prediction direction, and a second motion vector and a second reference image index in the backward prediction direction. Determining the positions of the N forward reference blocks and the N backward reference blocks based on the initial motion information and the position of the current image block Based on the first motion vector and the position of the current image block, using the position of the initial forward reference block as the first search starting point, determine the position of the initial forward reference block of the current image block in the forward reference image corresponding to the first reference image index, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, where the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of the (N - 1) candidate forward reference blocks, and determine. Based on the second motion vector and the position of the current image block, using the position of the initial backward reference block as the second search starting point, determine the position of the initial backward reference block of the current image block in the backward reference image corresponding to the second reference image index, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, where the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of the (N - 1) candidate backward reference blocks, and determine.
[0028] Regarding the first aspect, in some implementations of the first aspect, determining from the positions of M pairs of reference blocks based on the matching cost criterion that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block is determining from the positions of M pairs of reference blocks that the positions of the pair of reference blocks with the minimum matching error are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, or determining from the positions of M pairs of reference blocks that the positions of the pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, where M is less than or equal to N, and determining.
[0029] In one example, the matching cost criterion is a matching cost minimization criterion. For example, for the positions of M pairs of reference blocks, the difference between the pixel values of the forward reference block and the backward reference block is calculated for each pair of reference blocks, and from the positions of the M pairs of reference blocks, the positions of the pair of reference blocks with the pixel values having the minimum difference are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block.
[0030] In another example, the matching cost criterion is a matching cost minimization and early termination criterion. For example, for the position of the nth pair of reference blocks (one forward reference block and one backward reference block), the difference between the pixel values of the forward reference block and the backward reference block is calculated, where n is an integer greater than or equal to 1 and less than or equal to N, and when the pixel value difference is less than or equal to the matching error threshold, the positions of the nth pair of reference blocks (one forward reference block and one backward reference block) are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block.
[0031] Regarding the first aspect, in some implementations of the first aspect, the method is used to encode a current image block, and obtaining initial motion information of the current image block includes obtaining the initial motion information from a candidate motion information list of the current image block, or the method is used to decode a current image block, and before obtaining the initial motion information of the current image block, the method further includes obtaining indication information from the bitstream of the current image block, where the indication information is used to indicate the initial motion information of the current image block.
[0032] The image prediction method in this embodiment of the present application is applicable not only to the Merge prediction mode and / or the advanced motion vector prediction (AMVP) mode, but also to another mode in which a spatial reference block, a temporal reference block, and / or an inter-view reference block are used to predict the motion information of the current image block. This improves the coding information.
[0033] The second aspect of the present application provides an image prediction method, the method comprising obtaining initial motion information of a current image block, determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the initial motion information and the position of the current image block, wherein the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1; determining, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, wherein the position of each pair of reference blocks includes the position of the forward reference block and the position of the backward reference block, and for the position of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on a temporal region distance, the first position offset represents the offset of the position of the forward reference block with respect to the position of the initial forward reference block, the second position offset represents the offset of the position of the backward reference block with respect to the position of the initial backward reference block, M is an integer greater than or equal to 1, and M is less than or equal to N; and obtaining a predicted value of the pixel value of the current image block based on the pixel value of the target forward reference block and the pixel value of the target backward reference block.
[0034] In this embodiment of the present application, it should be particularly noted that the offset of the position of the initial forward reference block with respect to the position of the initial forward reference block is 0, and the offset of the position of the initial backward reference block with respect to the position of the initial backward reference block is 0. The offset 0 and the offset 0 also satisfy the condition of a mirror image relationship or the condition of a proportional relationship based on the time-domain distance. In other words, at the positions of the (N - 1) pairs of reference blocks, for the positions of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time-domain distance or a mirror image relationship. In this specification, the positions of the (N - 1) pairs of reference blocks do not include the position of the initial forward reference block or the position of the initial backward reference block.
[0035] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form N pairs of reference block positions. For the positions of each pair of reference blocks among the positions of the N pairs of reference blocks, there is a proportional relationship based on the time domain (also referred to as a mirror image relationship based on the time domain distance) between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (that is, the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (that is, the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0036] Regarding the second aspect, in some implementation forms of the second aspect, for each pair of reference blocks, the fact that the first position offset and the second position offset have a proportional relationship based on the time domain distance means that For each pair of reference blocks, the proportional relationship between the first position offset and the second position offset is determined based on the proportional relationship between the first time domain distance and the second time domain distance. The first time domain distance represents the time domain distance between the current image to which the current image block belongs and the forward reference image, and the second time domain distance represents the time domain distance between the current image and the backward reference image.
[0037] Regarding the second aspect, in some implementations of the second aspect, the fact that the first position offset and the second position offset are in a proportional relationship based on the time-domain distance means that if the first time-domain distance is the same as the second time-domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset, or if the first time-domain distance is different from the second time-domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the proportional relationship between the amplitude value of the first position offset and the amplitude value of the second position offset is based on the proportional relationship between the first time-domain distance and the second time-domain distance, the first time-domain distance represents the time-domain distance between the current image to which the current image block belongs and the forward reference image, and the second time-domain distance represents the time-domain distance between the current image and the backward reference image.
[0038] Regarding the second aspect, in some implementations of the second aspect, the method further includes obtaining updated motion information of the current image block, where the updated motion information includes an updated forward motion vector and an updated backward motion vector, the updated forward motion vector points to the position of the target forward reference block, and the updated backward motion vector points to the position of the target backward reference block.
[0039] It can be seen that the refined motion information of the current image block can be obtained in this embodiment of the present application. This improves the accuracy of the motion information of the current image block and also smooths the prediction of another image block. For example, it improves the prediction accuracy of the motion information of another image block.
[0040] Regarding the second aspect, in some implementations of the second aspect, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, or The positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0041] Regarding the second aspect, in some implementations of the second aspect, the positions of the N pairs of reference blocks include the positions of the paired initial forward reference block and the initial backward reference block, and the positions of the paired candidate forward reference block and the candidate backward reference block. There is a proportional relationship based on the temporal domain distance between the offset of the position of the candidate forward reference block with respect to the position of the initial forward reference block in the forward reference image and the offset of the position of the candidate backward reference block with respect to the position of the initial backward reference block in the backward reference image.
[0042] Regarding the second aspect, in some implementations of the second aspect, the initial motion information includes forward prediction motion information and backward prediction motion information. Determining the positions of the N forward reference blocks and the N backward reference blocks based on the initial motion information and the position of the current image block is determining the positions of the N forward reference blocks in the forward reference image based on the forward prediction motion information and the position of the current image block, where the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, and determining the positions of the N backward reference blocks in the backward reference image based on the backward prediction motion information and the position of the current image block, where the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0043] Regarding the second aspect, in some implementations of the second aspect, the initial motion information includes a first motion vector and a first reference image index in the forward prediction direction, and a second motion vector and a second reference image index in the backward prediction direction. Determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the initial motion information and the position of the current image block is Based on the first motion vector and the position of the current image block, using the position of the initial forward reference block as the first search starting point to determine the position of the initial forward reference block of the current image block in the forward reference image corresponding to the first reference image index, and determining the positions of (N - 1) candidate forward reference blocks in the forward reference image, where the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of the (N - 1) candidate forward reference blocks, and determining Based on the second motion vector and the position of the current image block, using the position of the initial backward reference block as the second search starting point to determine the position of the initial backward reference block of the current image block in the backward reference image corresponding to the second reference image index, and determining the positions of (N - 1) candidate backward reference blocks in the backward reference image, where the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of the (N - 1) candidate backward reference blocks, and including determining.
[0044] Regarding the second aspect, in some implementations of the second aspect, determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block based on the matching cost criterion is Determining from the positions of M pairs of reference blocks that the positions of the pair of reference blocks with the minimum matching error are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, or Determining that, from the positions of M pairs of reference blocks, the positions of a pair of reference blocks with a matching error less than or equal to a matching error threshold are the positions of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, where M is less than or equal to N, including determining.
[0045] In one example, the matching cost criterion is the matching cost minimization criterion. For example, for the positions of M pairs of reference blocks, the difference between the pixel values of the forward reference block and the pixel values of the backward reference block is calculated for each pair of reference blocks, and from the positions of the M pairs of reference blocks, the positions of the pair of reference blocks with the pixel values having the minimum difference are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block.
[0046] In another example, the matching cost criterion is the matching cost minimization and early termination criterion. For example, for the position of the nth pair of reference blocks (one forward reference block and one backward reference block), the difference between the pixel values of the forward reference block and the pixel values of the backward reference block is calculated, where n is an integer greater than or equal to 1 and less than or equal to N, and when the pixel value difference is less than or equal to the matching error threshold, the positions of the nth pair of reference blocks (one forward reference block and one backward reference block) are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block.
[0047] Regarding the second aspect, in some implementations of the second aspect, the method is used to encode the current image block, and obtaining the initial motion information of the current image block includes obtaining the initial motion information from a candidate motion information list of the current image block, or The method is used to decode the current image block. Before obtaining the initial motion information of the current image block, the method is to obtain indication information from the bitstream of the current image block, and the indication information is used to indicate the initial motion information of the current image block, and further includes obtaining.
[0048] The third aspect of the present application provides an image prediction method. The method is to obtain the i-th motion information of the current image block, and determine the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in the forward reference image, the N backward reference blocks are arranged in the backward reference image, and N is an integer greater than 1, determining; from the positions of M pairs of reference blocks based on the matching cost criterion, determining that the positions of a pair of reference blocks are the positions of the i-th target forward reference block of the current image block and the positions of the i-th target backward reference block of the current image block, where the positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block, and for the positions of each pair of reference blocks, the first position offset and the second position offset are in a mirror image relationship, the first position offset represents the offset of the position of the forward reference block relative to the position of the (i - 1)-th target forward reference block, the second position offset represents the offset of the position of the backward reference block relative to the position of the (i - 1)-th target backward reference block, M is an integer greater than or equal to 1, and M is less than or equal to N, determining; obtaining a predicted value of the pixel value of the current image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1, obtaining.
[0049] In this embodiment of the present application, it should be particularly noted that the offset of the position of the initial forward reference block with respect to the position of the initial forward reference block is 0, and the offset of the position of the initial backward reference block with respect to the position of the initial backward reference block is 0. The offset 0 and the offset 0 also satisfy the condition of the mirror image relationship.
[0050] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form N pairs of reference block positions. For each pair of reference block positions among the N pairs of reference block positions, there is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (that is, the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (that is, the optimal backward reference block / backward prediction block) of the current image block, thereby obtaining a predicted value of the pixel value of the current image block based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, in this embodiment of the present application, the accuracy of refining the motion vector MV can be further improved by using the iteration method, thereby further improving the coding performance.
[0051] Regarding the third aspect, in some implementation forms of the third aspect, if i = 1, the i-th motion information is the initial motion information of the current image block. Correspondingly, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance. Or the positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0052] If i > 1, the i-th motion information includes a forward motion vector pointing to the position of the (i - 1)-th target forward reference block and a backward motion vector pointing to the position of the (i - 1)-th target backward reference block. Correspondingly, the positions of the N forward reference blocks include the position of one (i - 1)-th target forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the (i - 1)-th target forward reference block is an integer pixel distance or a fractional pixel distance. Or the positions of the N backward reference blocks include the position of one (i - 1)-th target backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the (i - 1)-th target backward reference block is an integer pixel distance or a fractional pixel distance.
[0053] Note that when the method is used to encode the current image block, the initial motion information of the current image block is obtained by using a method of determining the initial motion information from the candidate motion information list of the current image block, or when the method is used to decode the current image block, the initial motion information of the current image block is obtained by using a method of obtaining the indication information from the bitstream of the current image block, where the indication information is used to indicate the initial motion information of the current image block.
[0054] Regarding the third aspect, in some implementations of the third aspect, obtaining a predicted value of the pixel value of the image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1, the obtaining is When the repetition end condition is satisfied, obtaining a predicted value of the pixel value of the image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1, includes the obtaining.
[0055] Regarding the third aspect, in some implementations of the third aspect, the first position offset and the second position offset being in a mirror image relationship includes that the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset.
[0056] Regarding the third aspect, in some implementations of the third aspect, the i-th motion information includes a forward motion vector, a forward reference image index, a backward motion vector, and a backward reference image index. Determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block is Based on the forward motion vector and the position of the current image block, the position of the (i - 1)-th target forward reference block is used as the starting point for the i-th f search to determine the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index, and to determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, where the positions of the N forward reference blocks include the position of the (i - 1)-th target forward reference block and the positions of the (N - 1) candidate forward reference blocks, and to determine; Based on the backward motion vector and the position of the current image block, the position of the (i - 1)-th target backward reference block is used as the starting point for the i-th b search to determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index, and to determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, where the positions of the N backward reference blocks include the position of the (i - 1)-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks, and to determine, which includes.
[0057] Regarding the third aspect, in some implementations of the third aspect, determining from the positions of M pairs of reference blocks based on the matching cost criterion that the positions of a pair of reference blocks are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block is determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks with the minimum matching error are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, or determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, where M is less than or equal to N, and to determine, which includes.
[0058] A fourth aspect of the present application provides an image prediction method, the method including obtaining the i-th motion information of the current image block, determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1; determining, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the positions of the i-th target forward reference block of the current image block and the positions of the i-th target backward reference block of the current image block, where the positions of each pair of reference blocks include the position of a forward reference block and the position of a backward reference block, and for the positions of each pair of reference blocks, a first position offset and a second position offset are in a proportional relationship based on a time-domain distance, the first position offset represents the offset of the position of the forward reference block with respect to the position of the (i - 1)-th target forward reference block in the forward reference image, the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i - 1)-th target backward reference block in the backward reference image, M is an integer greater than or equal to 1, and M is less than or equal to N; obtaining a predicted value of the pixel value of the current image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1.
[0059] In this embodiment of the present application, it should be particularly noted that the offset of the position of the initial forward reference block with respect to the position of the initial forward reference block is 0, and the offset of the position of the initial backward reference block with respect to the position of the initial backward reference block is 0. The offset 0 and the offset 0 also satisfy the conditions of a mirror image relationship or a proportional relationship based on the time domain distance. In other words, at the positions of (N - 1) pairs of reference blocks, for the positions of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time domain distance or a mirror image relationship. In this specification, the positions of (N - 1) pairs of reference blocks do not include the position of the initial forward reference block or the position of the initial backward reference block.
[0060] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. For the positions of each pair of reference blocks among the positions of the N pairs of reference blocks, there is a proportional relationship based on the time-domain distance between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, in this embodiment of the present application, the accuracy of refining the motion vector MV can be further improved by using an iterative method, thereby further improving the coding performance.
[0061] Regarding the fourth aspect, in some implementation forms of the fourth aspect, if i = 1, the i-th motion information is the initial motion information of the current image block, or if i > 1, the i-th motion information includes a forward motion vector pointing to the position of the (i - 1)-th target forward reference block and a backward motion vector pointing to the position of the (i - 1)-th target backward reference block.
[0062] Regarding the fourth aspect, in some implementation forms of the fourth aspect, obtaining a predicted value of the pixel value of an image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1. The obtaining includes When the repetition end condition is satisfied, obtaining a predicted value of the pixel value of an image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1. The obtaining includes
[0063] Regarding the fourth aspect, in some implementation forms of the fourth aspect, the fact that the first position offset and the second position offset have a proportional relationship based on the time domain distance means If the first time domain distance is the same as the second time domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset, or If the first time domain distance is different from the second time domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the proportional relationship between the amplitude value of the first position offset and the amplitude value of the second position offset is based on the proportional relationship between the first time domain distance and the second time domain distance, including The first time domain distance represents the time domain distance between the current image to which the current image block belongs and the forward reference image, and the second time domain distance represents the time domain distance between the current image and the backward reference image.
[0064] Regarding the fourth aspect, in some implementation forms of the fourth aspect, the i-th motion information includes a forward motion vector, a forward reference image index, a backward motion vector, and a backward reference image index. Determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block is Based on the forward motion vector and the position of the current image block, use the position of the (i - 1)-th target forward reference block as the search starting point for the i-th one to determine the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, where the positions of the N forward reference blocks include the position of the (i - 1)-th target forward reference block and the positions of the (N - 1) candidate forward reference blocks, and perform the determination. f Based on the backward motion vector and the position of the current image block, use the position of the (i - 1)-th target backward reference block as the search starting point for the i-th one to determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, where the positions of the N backward reference blocks include the position of the (i - 1)-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks, and perform the determination. Based on the forward motion vector and the position of the current image block, use the position of the (i - 1)-th target forward reference block as the search starting point for the i-th one to determine the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, where the positions of the N forward reference blocks include the position of the (i - 1)-th target forward reference block and the positions of the (N - 1) candidate forward reference blocks, and perform the determination. b Based on the backward motion vector and the position of the current image block, use the position of the (i - 1)-th target backward reference block as the search starting point for the i-th one to determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, where the positions of the N backward reference blocks include the position of the (i - 1)-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks, and perform the determination.
[0065] Regarding the fourth aspect, in some implementations of the fourth aspect, based on the matching cost criterion, determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block is determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks with the minimum matching error are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, or determining from the positions of M pairs of reference blocks that the positions of a pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, where M is less than or equal to N, and performing the determination.
[0066] A fifth aspect of the present application provides an image prediction apparatus including several functional units configured to implement any method in the first aspect. For example, the image prediction apparatus includes a first acquisition unit configured to acquire initial motion information of a current image block, and a first search unit that determines positions of N forward reference blocks and positions of N backward reference blocks based on the initial motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1. The first search unit determines, from positions of M pairs of reference blocks based on a matching cost criterion, that positions of a pair of reference blocks are positions of a target forward reference block of the current image block and positions of a target backward reference block of the current image block. Positions of each pair of reference blocks include positions of the forward reference block and positions of the backward reference block. For positions of each pair of reference blocks, a first position offset and a second position offset are in a mirror image relationship. The first position offset represents an offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents an offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer greater than or equal to 1 and M is less than or equal to N. The image prediction apparatus may further include a first prediction unit configured to obtain a predicted value of pixel values of the current image block based on pixel values of the target forward reference block and pixel values of the target backward reference block.
[0067] In different application scenarios, the image prediction apparatus is applied to, for example, a video encoding apparatus (video encoder) or a video decoding apparatus (video decoder).
[0068] A sixth aspect of the present application provides an image prediction apparatus including several functional units configured to implement any method in the second aspect. For example, the image prediction apparatus includes a second acquisition unit configured to acquire initial motion information of a current image block, and a second search unit that determines positions of N forward reference blocks and positions of N backward reference blocks based on the initial motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1. The second search unit determines, from positions of M pairs of reference blocks based on a matching cost criterion, that positions of a pair of reference blocks are positions of a target forward reference block and a target backward reference block of the current image block. Positions of each pair of reference blocks include positions of a forward reference block and a backward reference block, and for positions of each pair of reference blocks, a first position offset and a second position offset are in a proportional relationship based on a time-domain distance. The first position offset represents an offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents an offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer greater than or equal to 1 and M is less than or equal to N. The second search unit is further configured to obtain a predicted value of pixel values of the current image block based on pixel values of the target forward reference block and pixel values of the target backward reference block. The image prediction apparatus may further include a second prediction unit configured to obtain a predicted value of pixel values of the current image block based on pixel values of the target forward reference block and pixel values of the target backward reference block.
[0069] In different application scenarios, the image prediction apparatus is applied to, for example, a video encoding apparatus (video encoder) or a video decoding apparatus (video decoder).
[0070] A seventh aspect of the present application provides an image prediction apparatus including several functional units configured to implement any method in the third aspect. For example, the image prediction apparatus includes a third acquisition unit configured to acquire the i-th motion information of a current image block, and a third search unit that determines the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1; and determines, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the positions of the i-th target forward reference block of the current image block and the positions of the i-th target backward reference block of the current image block, where the positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block, and for the positions of each pair of reference blocks, a first position offset and a second position offset are in a mirror image relationship, the first position offset represents the offset of the position of the forward reference block with respect to the position of the (i - 1)-th target forward reference block, the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i - 1)-th target backward reference block, and M is an integer greater than or equal to 1 and M is less than or equal to N; and a third prediction unit configured to obtain a predicted value of the pixel value of the current image block based on the pixel values of the j-th target forward reference block and the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1.
[0071] In different application scenarios, the image prediction apparatus is applied to, for example, a video encoding apparatus (video encoder) or a video decoding apparatus (video decoder).
[0072] The eighth aspect of the present application provides an image prediction apparatus including several functional units configured to implement any method in the fourth aspect. For example, the image prediction apparatus includes a fourth acquisition unit configured to acquire the i-th motion information of the current image block, and a fourth search unit configured to determine the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in the forward reference image, the N backward reference blocks are arranged in the backward reference image, and N is an integer greater than 1; and determine, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the positions of the i-th target forward reference block of the current image block and the positions of the i-th target backward reference block of the current image block, where the positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block, and for the positions of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time-domain distance, the first position offset represents the offset of the position of the forward reference block with respect to the position of the (i-1)-th target forward reference block in the forward reference image, the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i-1)-th target backward reference block in the backward reference image, M is an integer greater than or equal to 1, and M is less than or equal to N; and a fourth prediction unit configured to obtain a predicted value of the pixel value of the current image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1.
[0073] In different application scenarios, the image prediction apparatus is applied to, for example, a video encoding device (video encoder) or a video decoding device (video decoder).
[0074] A ninth aspect of the present application provides an image prediction apparatus, the apparatus comprising a processor and a memory coupled to the processor. The processor is configured to execute the method according to the first aspect, the second aspect, the third aspect, the fourth aspect, or an implementation form of the foregoing aspects.
[0075] A tenth aspect of the present application provides a video encoder. The video encoder is configured to encode an image block, and includes an inter-frame prediction module that includes an image prediction apparatus according to the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, and is configured to obtain a predicted value of a pixel value of the image block through prediction, an entropy encoding module configured to encode instruction information into a bitstream, where the instruction information is used to indicate initial motion information of the image block, and a reconstruction module configured to reconstruct the image block based on the predicted value of the pixel value of the image block.
[0076] An eleventh aspect of the present application provides a video decoder. The video decoder is configured to decode a bitstream to obtain an image block, and includes an entropy decoding module configured to decode the bitstream to obtain instruction information, where the instruction information is used to indicate initial motion information of the currently obtained image block through decoding, an inter-frame prediction module that includes an image prediction apparatus according to the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, and is configured to obtain a predicted value of a pixel value of the image block through prediction, and a reconstruction module configured to reconstruct the image block based on the predicted value of the pixel value of the image block.
[0077] A twelfth aspect of the present application provides a video encoding device comprising a non-volatile memory medium and a processor. The non-volatile memory medium stores an executable program. The processor and the non-volatile memory medium are coupled to each other, and the processor executes an executable program implementing the first, second, third, or fourth aspect, or an implementation form of the first, second, third, or fourth aspect.
[0078] A thirteenth aspect of the present application provides a video decoding device comprising a non-volatile memory medium and a processor. The non-volatile memory medium stores an executable program. The processor and the non-volatile memory medium are coupled to each other, and the processor executes an executable program implementing the first, second, third, or fourth aspect, or an implementation form of the first, second, third, or fourth aspect.
[0079] A fourteenth aspect of the present application provides a computer-readable storage medium storing instructions. When the instructions are executed on a computer, the computer can execute the method in the first, second, third, or fourth aspect, or an implementation form of the first, second, third, or fourth aspect.
[0080] A fifteenth aspect of the present application provides a computer program product comprising instructions. When the instructions are executed on a computer, the computer can execute the method in the first, second, third, or fourth aspect, or an implementation form of the first, second, third, or fourth aspect.
[0081] A sixteenth aspect of the present application provides an electronic device comprising the video encoder in the tenth aspect, the video decoder in the eleventh aspect, or the image prediction device in the fifth, sixth, seventh, or eighth aspect.
[0082] It should be understood that the beneficial effects brought about by these aspects and the corresponding implementable design methods are similar, and thus will not be repeated here.
Brief Description of the Drawings
[0083]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0084] Next, while referring to the accompanying drawings in the embodiments of the present application, the technical solution in the embodiments of the present application will be clarified to description and explained.
[0085] FIG. 1 is a schematic block diagram of a video encoding system according to an embodiment of the present application. In the system, the video encoder 20 and the video decoder 30 predict predicted values of pixel values of image blocks based on examples of various image prediction methods provided in the present application, and refine motion information such as motion vectors of the current encoded or decoded image blocks, so as to be configured to further improve the encoding performance. As shown in FIG. 1, the system includes a source device 12 and a destination device 14. The source device 12 、Generate the encoded video data to be decoded by the destination device 14. The source device 12 and the destination device 14 can include any one of a wide range of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, telephone handsets such as "smart" phones, "smart" touch pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming transmission devices, or the like.
[0086] The destination device 14 can receive the encoded video data to be decoded by using the link 16. The link 16 can include any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one possible implementation, the link 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. It may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency spectrum or one or more physical transmission lines. The communication medium can be part of a packet-based network (e.g., a local area network, a wide area network, or the global network of the Internet). The communication medium can include routers, switches, base stations, or any other device that can be configured to facilitate communication from the source device 12 to the destination device 14.
[0087] Alternatively, the encoded data may be output from the output interface 22 to the storage device 24. Similarly, the encoded data may be accessed from the storage device 24 through the input interface. The storage device 24 may include any one of a plurality of distributed or local data storage media, such as a hard disk drive, a Blu-ray Disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media used to store encoded video data. In another possible implementation, the storage device 24 may correspond to a file server or another intermediate storage device that can store the encoded video data generated by the source device 12. The destination device 14 may access the video data stored in the storage device 24 through a streaming transmission function or a download function. The file server may be any type of server that can store the encoded video data and transmit the encoded video data to the destination device 14. In one possible implementation, the file server may include a web server, a file transfer protocol server, a network-attached storage device, or a local hard disk drive. The destination device 14 may access the encoded video data through any standard data connection, including an Internet connection. The data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a cable modem), or a combination of these applicable for accessing the encoded video data stored in the file server. The transmission of the encoded video data from the storage device 24 may be a streaming transmission, a download transmission, or a combination of these.
[0088] The technology in this application is not necessarily limited to wireless applications or settings. The technology can be applied to video decoding to support any one of a plurality of multimedia applications, such as television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), digital video encoding stored on a data storage medium, decoding of digital video stored on a data storage medium, or another application. In some possible implementations, the system can be configured to support one-way or two-way video transmission for applications such as streaming video transmission, video playback, video broadcasting, and / or video telephony.
[0089] In one possible implementation of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. In some applications, the output interface 22 may include a modulator / demodulator (modem) and / or a transmitter. In the source device 12, the video source 18 may include, for example, as a source, a video capture device (e.g., a video camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination thereof. In one possible implementation, when the video source 18 is a video camera, the source device 12 and the destination device 14 can form a camera phone or a video phone. For example, the technology described in this application may be applied to video decoding and can be applied to wireless and / or wired applications.
[0090] The video encoder 20 may be configured to encode video that is captured, pre-captured, or generated by a computer. The encoded video data may be transmitted directly to the destination device 14 through the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on the storage device 24 for subsequent access by the destination device 14 or another device for decoding and / or playback.
[0091] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. In some applications, the input interface 28 may include a receiver and / or a modem. The input interface 28 of the destination device 14 receives the encoded video data by using the link 16. The encoded video data transmitted to or provided to the storage device 24 by using the link 16 may be generated by the video encoder 20 and may include a plurality of syntax elements used by the video decoder to decode the video data. to 30 These syntax elements are included in the encoded video data transmitted on the communication medium and may be stored on a storage medium or stored in a file server.
[0092] The display device 32 may be integrated with the destination device 14 or disposed outside the destination device 14. In some possible implementations, the destination device 14 may include an integrated display device and may also be configured to connect to an interface of an external display device. In other possible implementations, the destination device 14 may be a display device. Generally, the display device 32 displays the decoded video data to the user and may include any one of a plurality of display devices, such as a liquid crystal display, a plasma display, an organic light emitting diode display, or another type of display device.
[0093] Video encoder 20 and video decoder 30 may operate, for example, according to the next-generation video coding compression standard (H.266) currently under development, and may conform to the H.266 test model (JEM). Alternatively, video encoder 20 and video decoder 30 may operate, for example, according to other proprietary or industry standards or extensions of the ITU-T H.265 standard or the ITU-T H.264 standard. The ITU-T H.265 standard is also referred to as the high efficiency video decoding standard, and the ITU-T H.264 standard is alternatively referred to as MPEG-4 Part 10, or advanced video coding (AVC). However, the technology of this application is not limited to any specific decoding standard. Other possible implementations of video compression standards include MPEG-2 and ITU-T H.263.
[0094] Although not shown in FIG. 1, in some embodiments, video encoder 20 and video decoder 30 may be integrated with an audio encoder and an audio decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software for encoding both audio and video in a common data stream or in separate data streams. If applicable, in some possible implementations, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).
[0095] Each of the video encoder 20 and the video decoder 30 can be implemented as any of a plurality of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these techniques are implemented partially as software, the apparatus can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in the form of hardware using one or more processors to implement the techniques of the present application. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of the one or more encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) within the corresponding apparatus.
[0096] This application may relate to another device in which, for example, a video encoder 20 transmits specific information as a signal to, for example, a video decoder 30. However, it should be understood that the video encoder 20 can transmit information as a signal by associating specific syntax elements with the encoded portion of the video data. That is, the video encoder 20 may store specific syntax elements in the header information of the encoded portion of the video data and transmit the data as a signal. In some applications, these syntax elements can be received by the video decoder 30 and stored encoded (e.g., stored in a storage system 34 or a file server 36) before being decoded. Therefore, the term "signal" may mean the transmission of syntax or other data used to decode compressed video data, regardless of whether the transmission is performed in real time, near real time, or within a certain period. For example, the transmission may be performed when the syntax elements are stored in the medium during encoding, and then the syntax elements can be retrieved by the decoding device at any time after being stored in the medium.
[0097] JCT-VC developed the H.265 (HEVC) standard. HEVC standardization is based on a development model of video decoding devices, and the model is referred to as the HEVC Test Model (HM). The latest H.265 standard document is available at http: / / www.itu.int / rec / T-REC-H.265. The latest version of the standard document is H.265(12 / 16), and the entire standard document is incorporated herein by reference. In HM, it is assumed that the video decoding device has some additional functions compared to the existing algorithms of ITU-T H.264 / AVC. For example, H.264 provides nine intra prediction and coding modes, while HM can provide up to 35 intra prediction and coding modes.
[0098] JVET is working on the development of the H.266 standard. The H.266 standardization process is based on the development model of video decoding devices, and the model is referred to as the H.266 test model. The description of the H.266 algorithm is available at http: / / phenix.int-evry.fr / jvet, and the description of the latest algorithm is included in JVET-F1001-v2. The entire description of this algorithm is incorporated herein by reference. Further, the reference software for the JEM test model is available from https: / / jvet.hhi.fraunhofer.de / svn / svn_HMJEMSoftware / , and this is also incorporated herein by reference in its entirety.
[0099] Generally, as described in the HM working model, a video frame or image can be divided into a series of tree blocks or largest coding units (LCUs) that include both luminance samples and chrominance samples. An LCU is also referred to as a CTU. A tree block has a function similar to that of a macroblock in the H.264 standard. A slice includes several consecutive tree blocks in decoding order. A video frame or image can be partitioned into one or more slices. Each tree block can be divided into coding units based on a quadtree. For example, a tree block acting as the root node of a quadtree can be divided into four child nodes, and each child node can act as a parent node and be divided into four other child nodes. The last non-divisible child node acting as the leaf node of the quadtree contains a decoded node, for example, a decoded image block. In the syntax data associated with the decoded bitstream, the maximum number of possible divisions of a tree block and the minimum size of a decoded node can be defined.
[0100] The symbolization unit includes a decoding node, a prediction unit (PU), and a transform unit (TU) associated with the decoding node. The CU should have a size corresponding to that of the decoding node and be square in shape. The size of the CU can range from 8×8 pixels up to a maximum of 64×64 pixels, or a larger tree block size. Each CU may include one or more PUs and one or more TUs. For example, the syntax data associated with a CU may describe partitioning one CU into one or more PUs. The partitioning pattern may vary when the CU is encoded in skip or direct mode, in intra-prediction mode, or in inter-prediction mode. The PUs obtained through partitioning may be non-square in shape. For example, the syntax data associated with a CU may also describe partitioning one CU into one or more TUs based on a quadtree. The TU may take on a square or non-square shape.
[0101] The HEVC standard allows TU-based transformation, and the TUs can be different for different CUs. The TU size is typically set based on the size of the PU within a given CU defined for the partitioned LCU. However, this is not always the case. The TU size is generally the same as or smaller than the PU size. In some possible implementations, a quadtree structure called "residual quadtree" (RQT) can be used to divide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can be referred to as TUs. The pixel differences associated with the TUs are transformed to generate transform coefficients, and the transform coefficients can be quantized. quadtree The quadtree structure called "residual quadtree" (RQT) can be used to divide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can be referred to as TUs. The pixel differences associated with the TUs are transformed to generate transform coefficients, and the transform coefficients can be quantized.
[0102] Generally, the transformation and quantization processes are used for TUs. A given CU having one or more PUs may also include one or more TUs. After prediction, video encoder 20 may calculate residual values corresponding to the PUs. The residual values include pixel differences, which may be transformed into transform coefficients, and the transform coefficients are quantized, subjected to TU scanning, and serialized transform coefficients for entropy decoding are generated. In this application, the term "image block" is generally used to represent the decoding node of a CU. In some specific applications, in this application, the term "image block" may also be used to represent a decoding node, a PU, and a TU, for example, a tree block including an LCU or a CU. In this embodiment of this application, to improve the coding performance by performing an inverse quantization process of the transform coefficients corresponding to the current image block (i.e., the current transform block), various method examples described by the adaptive inverse quantization method in video coding or decoding are described in detail below.
[0103] A video sequence generally includes a series of video frames or images. For example, a group of pictures (GOP) includes a series of video images, one video image, or multiple video images. The GOP may include syntax data in the GOP header information, in the header information of one or more of the images, or elsewhere, and the syntax data describes the number of images included in the GOP. Each slice of an image may include slice syntax data that describes the coding mode of the corresponding image. Video encoder 20 typically performs operations on image blocks in several video slices to encode video data. The image block may correspond to a decoding node within a CU. The size of the image block may be fixed or changeable and may vary according to the specified decoding standard.
[0104] In one possible implementation, HM supports predictions for various PU sizes. Assuming that the size of a given CU is 2N×2N, HM supports intra-frame prediction for PU sizes of 2N×2N or N×N, and inter-frame prediction for symmetric PU sizes of 2N×2N, 2N×N, N×2N, or N×N. HM also supports asymmetric partitioning of inter-frame prediction for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N. In asymmetric partitioning, one direction of the CU is not partitioned, while the other direction is partitioned into two parts, one part occupying 25% of the CU and the other part occupying 75% of the CU. The part that occupies 25% of the CU is indicated by an indicator that includes "n" followed by "Up", "Down", "Left", or "Right". Thus, for example, "2N×nU" refers to a 2N×2N CU partitioned horizontally, with the 2N×0.5N PU on top and the 2N×1.5N PU on the bottom.
[0105] In this application, "N×M" and "N multiplied by M" can be used interchangeably to indicate the pixel size of an image block in the vertical and horizontal dimensions, for example, 16×8 pixels or a number of pixels equal to 16 multiplied by 8. Generally, a 16×8 block has 16 pixels in the horizontal direction and 8 pixels in the vertical direction. In other words, the width of the image block is 16 pixels and the height of the image block is 8 pixels.
[0106] After in - frame or inter - frame prediction decoding of the PUs within the CU, the video encoder 20 may calculate the residual data of the TUs within the CU. A PU may include pixel data within a spatial region (also referred to as a pixel region). A TU may include coefficients within a transform region after a transform (e.g., discrete cosine transform (DCT), integer transform, wavelet transform, or other conceptually similar transform) has been performed on the residual video data. The residual data may correspond to the difference between the pixel values of the non - encoded image and the predicted pixel values corresponding to the PU. The video encoder 20 may generate a TU that includes the residual data of the CU and then transform the TU to generate CU transform coefficients.
[0107] In this embodiment of the present application, various method examples of the inter - frame prediction process in video encoding or decoding for obtaining the sampling values of the sampling points of the optimal forward - reference block and the optimal backward - reference block of the current image block and further predicting the sampling value of the sampling point of the current image block are described in detail below. An image block is a two - dimensional sampling - point array, which may be a square array or a rectangular array. For example, an image block of size 4×4 can be considered as a square sampling - point array formed by a total of 4×4 = 16 sampling points. The signal within the image block is the sampling value of the sampling points within the image block. Further, the sampling point may also be referred to as a sample or a pixel and should be used interchangeably in the present specification of the present invention. Correspondingly, the value of the sampling point may also be referred to as a pixel value and should be used interchangeably in the present application. An image may be represented as a two - dimensional sampling - point array and is represented by using a method similar to the method used for image blocks.
[0108] After performing the transformation to generate transformation coefficients, the video encoder 20 may quantize the transformation coefficients. The quantization means, for example, the process of quantizing the coefficients, reduces the amount of data used to represent the coefficients and implements further compression. The quantization process may reduce the bit depth associated with some or all of the coefficients. For example, during quantization, if n is greater than m, then the n-bit value may be reduced to an m-bit value.
[0109] The JEM model further improves the video image coding structure. In particular, a block coding structure called "quadtree plus binary tree" (QTBT) is introduced. Without using concepts such as CUs, PUs, and TUs in HEVC, the QTBT structure supports a more flexible CU partitioning shape. One CU can take on a square or rectangular shape. Quadtree partitioning is first performed on the CTU, and binary tree partitioning is further performed on the leaf nodes of the quadtree. Additionally, there are two binary tree partitioning modes, namely, symmetric horizontal partitioning and symmetric vertical partitioning. The leaf nodes of the binary tree are referred to as CUs. The CUs within JEM cannot be further partitioned during prediction and transformation. In other words, the CUs, PUs, and TUs in JEM have the same block size. In the existing JEM, the maximum CTU size is 256×256 luminance pixels.
[0110] In some possible implementations, the video encoder 20 may scan the quantized transformation coefficients in a predefined scan order to generate a serialized vector that can be entropy encoded. In other possible implementations, the video encoder 20 may perform an adaptive scan. After scanning the quantized transformation coefficients to form a one-dimensional vector, the video encoder 20 performs entropy encoding on the one-dimensional vector using context-based adaptive variable length coding (CAVLC), context-based adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval segmentation entropy (PIPE) encoding or another entropyencoding It can be executed by using the method. The video encoder 20 may further perform entropy encoding on the syntax elements associated with the encoded video data for decoding the video data by the video decoder 30.
[0111] FIG. 2A is a schematic block diagram of a video encoder 20 according to an embodiment of the present application. Referring also to FIG. 3, the video encoder 20 may be configured to perform an image prediction process. In particular, the motion compensation unit 44 within the video encoder 20 may perform the image prediction process.
[0112] As shown in FIG. 2A, the video encoder 20 may include a prediction module 41, an adder 50, a conversion module 52, a quantization module 54, and an entropy encoding module 56. In one example, the prediction module 41 may include a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction unit 46. The internal structure of the prediction module 41 is not limited to this embodiment of the present application. Optionally, for a video encoder with a hybrid architecture, the video encoder 20 may further include an inverse quantization module 58, an inverse conversion module 60, and an adder 62.
[0113] In one possible implementation of FIG. 2A, the video encoder 20 may further include a segmentation unit (not shown) and a reference image memory 64. It should be understood that the segmentation unit and the reference image memory 64 may alternatively be disposed outside the video encoder 20.
[0114] In another possible implementation, the video encoder 20 may further include a filter (not shown) that filters the block boundaries to remove block effect artifacts from the reconstructed video. When necessary, the filter typically performs filtering on the output of the adder 62.
[0115] As shown in FIG. 2A, video encoder 20 receives video data, and a segmentation unit segments the data into image blocks. Such segmentation may further include segmentation into slices, image blocks, or other larger units, for example, image block segmentation based on the quadtree structure of LCU and CU. Generally, a slice may be divided into a plurality of image blocks.
[0116] Prediction module 41 is configured to generate a prediction block of a current encoded image block. Prediction module 41 may select one of a plurality of possible decoding modes of the current image block, for example, one of a plurality of intra-frame decoding modes or one of a plurality of inter-frame decoding modes, based on the encoding quality and cost calculation results (e.g., rate-distortion cost, RDcost). Prediction module 41 provides the intra-frame decoded or inter-frame decoded block to adder 50 to generate residual block data, provides the intra-frame decoded or inter-frame decoded block to adder 62 to reconstruct the encoded block, and may use the reconstructed block as a reference image.
[0117] Motion estimation unit 42 and motion compensation unit 44 in prediction module 41 perform inter-frame prediction decoding on the current image block with respect to one or more prediction blocks in one or more reference images to provide temporal compression. Motion estimation unit 42 is configured to determine an inter-frame prediction mode for a video slice based on a preset mode of the video sequence. In the preset mode, a video slice in the sequence may be designated as a P slice, a B slice, or a GPB slice. Motion estimation unit 42 and motion compensation unit 44 may be tightly integrated but are described separately for illustrative purposes. Motion estimation performed by motion estimation unit 42 is a process of generating a motion vector to estimate an image block. For example, the motion vector may indicate the displacement of a PU of an image block in a current video frame or image with respect to a prediction block in a reference image.
[0118] The prediction block is a block within the PU that is known to exactly match the image block to be decoded based on pixel differences, and the pixel differences can be determined based on the sum of absolute differences (SAD), the sum of squared differences (SSD), or another difference metric. In some possible implementations, the video encoder 20 may calculate the values of the sub-integer pixel positions of the reference image stored in the reference image memory 64.
[0119] By comparing the position of the PU with the position of the prediction block of the reference image, the motion estimation unit 42 calculates the motion vector of the PU of the image block within the inter-frame decoded slice. The reference image can be selected from the first reference image list (list 0) or the second reference image list (list 1). Each item in the list identifies one or more reference images stored in the reference image memory 64. The motion estimation unit 42 entropy-codes the calculated motion vector module and sends it to the entropy coder 56 and the motion compensation unit 44.
[0120] The motion compensation performed by the motion compensation unit 44 may well include removing or generating a prediction block based on the motion vector determined through motion estimation, and sub-pixel level interpolation may be performed. After receiving the motion vector of the PU of the current image block, the motion compensation unit 44 may identify the position of the prediction block pointed to by the motion vector in one of the reference image lists. The video encoder 20 subtracts the pixel value of the prediction block from the pixel value of the current image block being decoded, obtains a residual image block, and obtains a pixel difference. The pixel difference forms the residual data of the block and may include a luminance difference component and a chrominance difference component. The adder 50 is one or more components that perform a subtraction operation. The motion compensation unit 44 may further generate syntax elements associated with the image block and the video slice, whereby the video decoder 30 may decode the image block of the video slice. Next, the image prediction process in the embodiments of the present application will be described in detail with reference to FIGS. 3, 10 to 12, and 14 to 17. Details are not described here.
[0121] The intra-frame prediction unit 46 within the prediction module 41 may perform intra-frame prediction decoding on the current image block with respect to one or more adjacent blocks within the same image or slice as the current block to be decoded in order to provide spatial compression. Thus, instead of performing inter-frame prediction (as described above) by the motion estimation unit 42 and the motion compensation unit 44, the intra-frame prediction unit 46 may perform intra-frame prediction on the current block. Specifically, the intra-frame prediction unit 46 may determine an intra-frame prediction mode for encoding the current block. In some possible implementations, the intra-frame prediction unit 46 may use various intra-frame prediction modes (for example) to encode the current block during a separate encoding traversal, and the intra-frame prediction unit 46 (or in some possible implementations, the mode selection unit 40) may select an appropriate intra-frame prediction mode from the test modes.
[0122] After the prediction module 41 generates a predicted block of the current image block by performing inter-frame prediction or intra-frame prediction, the video encoder 20 generates a residual image block by subtracting the predicted block from the current image block. The residual video data within the residual block is included in one or more TUs and can be applied to the transformation module 52. The transformation module 52 is configured to transform the residual between the original block of the current encoded image block and the predicted block of the current image block. The transformation module 52 transforms the residual data into residual transform coefficients, for example, by performing a discrete cosine transform (DCT) or a conceptually similar transform (e.g., discrete sine transform DST). The transformation module 52 can transform the residual video data from pixel domain data to transform the region (e.g., frequency domain) data.
[0123] The transformation module 52 may send the obtained transform coefficients to the quantization module 54. The quantization module 54 quantizes the transform coefficients to further reduce the bit rate. In some possible implementations, the quantization module 54 may continue to scan the matrix containing the quantized transform coefficients. Alternatively, the entropy encoding module 56 may perform the scan.
[0124] After quantization, the entropy encoding module 56 may perform entropy encoding on the quantized transform coefficients. For example, the entropy encoding module 56 may perform context-based adaptive variable length encoding (CAVLC), context-based adaptive binary arithmetic encoding (CABAC), syntax-based context adaptive binary arithmetic encoding (SBAC), probability interval partitioning entropy (PIPE) encodingOr another entropy encoding method or technique may be executed. The entropy encoding module 56 may also perform entropy encoding on the motion vectors and other syntax elements that are those of the currently encoded video slice. After the entropy encoding module 56 performs entropy encoding, the encoded bitstream may be transmitted to the video decoder 30 or stored by the video decoder 30 for subsequent transmission or search.
[0125] The inverse quantization module 58 and the inverse transform module 60 perform inverse quantization and inverse transform respectively to reconstruct the residual block in the pixel region as a reference block of the reference image. The adder 62 adds the reconstructed residual block to the prediction block generated by the prediction module 41 to generate a reconstructed block, uses the reconstructed block as a reference block, and stores it in the reference image memory 64. The reference block can be used as a reference block for performing inter-frame prediction on blocks in subsequent video frames or images by the motion estimation unit 42 and the motion compensation unit 44.
[0126] It should be understood that another structural modification of the video encoder 20 can also be used to encode the video stream. For example, for some image blocks or image frames, the residual signal may well be directly quantized by the video encoder 20 without being processed by the transform module 52, and correspondingly, the residual signal is the inverse transform module 60It does not need to be processed by. Alternatively, for some image blocks or image frames, the video encoder 20 does not generate residual data, and correspondingly, no processing needs to be performed by the conversion module 52, quantization module 54, inverse quantization module 58, and inverse conversion module 60. Alternatively, the reconstructed image block can be directly stored as a reference block by the video encoder 20 without being processed by the filter unit. Alternatively, the quantization module 54 and inverse quantization module 58 in the video encoder 20 can be integrated. Alternatively, the conversion module 52 and inverse conversion module 60 in the video encoder 20 can be integrated. Alternatively, the adder 50 and adder 62 may be integrated.
[0127] FIG. 2B is a schematic block diagram of a video decoder 30 according to an embodiment of the present application. Also referring to FIGS. 3, 10 to 12, and 14 to 17, the video decoder 30 may execute an image prediction process. In particular, the motion compensation unit 82 in the video decoder 30 may execute the image prediction process.
[0128] As shown in FIG. 2B, the video decoder 30 may include an entropy decoding module 80, a prediction processing module 81, an inverse quantization module 86, an inverse conversion module 88, and a reconstruction module 90. In one example, the prediction module 81 may include a motion compensation unit 82 and an intra-frame prediction unit 84. This is not limited to this embodiment of the present application.
[0129] In one possible implementation, the video decoder 30 may further include a reference image memory 92. It should be understood that the reference image memory 92 may alternatively be disposed outside the video decoder 30. In some possible implementations, the video decoder 30 may execute an exemplary decoding process opposite to the encoding process described in the video encoder 20 in FIG. 2A.
[0130] During decoding, video decoder 30 receives from video encoder 20 an encoded video bitstream representing an encoded video slice and associated syntax elements of an image block. Video decoder 30 may receive syntax elements at the video slice level and / or at the image block level. Entropy decoding module 80 of video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients and some syntax elements. Entropy decoding module 80 transfers the syntax elements to prediction module 81. In this application, in one example, the syntax elements of this specification may include inter-frame prediction data related to the current image block, and the inter-frame prediction data includes an index identifier block_based_index, whereby it may indicate which motion information (also referred to as the initial motion information of the current image block) is used by the current image block. Optionally, the inter-frame prediction data may further include a switch flag block_based_enable_flag, whereby it indicates whether to perform image prediction on the current image block by using FIG. 3 or FIG. 14 (in other words, it indicates whether to perform inter-frame prediction on the current image block by using the MVD mirroring constraint condition proposed in this application), or whether to perform image prediction on the current image block by using FIG. 12 or FIG. 16 (in other words, it indicates whether to perform inter-frame prediction on the current image block by using the proportional relationship proposed in this application based on the temporal distance).
[0131] When a video slice is decoded into an intra-frame (I) slice, the intra-frame prediction unit 84 of the prediction module 81 may generate a prediction block of an image block of the current video slice based on the intra-frame prediction mode notified by transmitting signals and data of previously decoded blocks from the current frame or image. When a video slice is decoded into an inter-frame (i.e., B or P) slice, the motion compensation unit 82 of the prediction module 81 determines an inter-frame prediction mode used to decode the current image block of the current video slice based on the syntax elements received from the entropy decoding module 80 and may decode the current image block based on the determined inter-frame prediction mode (e.g., perform inter-frame prediction on the current image block). In particular, the motion compensation unit 82 may determine which prediction method is used to predict the current image block of the current video slice. For example, the syntax elements may indicate that an image prediction method based on the MVD mirror constraint condition should be used to predict the current image block. The motion information of the current image block of the current video slice is predicted or refined, and thereby, a predicted block of the current image block is generated by using the predicted motion information of the current image block by using the motion compensation process. The motion information in this specification may include reference image information and a motion vector. The reference image information may include, but is not limited to, one-way / bidirectional prediction information, a reference image list number, and a reference image index corresponding to the reference image list. For inter-frame prediction, a prediction block may be generated from one of the reference images in one of the reference image lists. The video decoder 30 may construct reference image lists, i.e., list 0 and list 1, based on the reference images stored in the reference image memory 92. The reference frame index of the current image may be included in one or both of the reference frame lists 0 and reference frame list 1. In some cases, the video encoder 20 may transmit a signal for indicating which new image prediction method is used.
[0132] In this embodiment, the prediction module 81 is configured to generate a prediction block of the current encoded image block. In particular, when the video slice is decoded into an intra-frame decoding (I) slice, the intra-frame prediction unit 84 of the prediction module 81 may generate a prediction block of the image block of the current video slice based on the intra-frame prediction mode of the transmitted signaling frame and the data of the previously decoded image blocks from the current frame or image. When the video image is decoded into an inter-frame decoding (e.g., B, P, or GPB) slice, the motion compensation unit 82 of the prediction module 81 generates a prediction block of the image block of the current video image based on the motion vector and other syntax elements received from the entropy decryption module 80.
[0133] The inverse quantization module 86 performs inverse quantization on the quantized transform coefficients provided in the bitstream obtained by the entropy decoding module 80 through decoding, i.e., inverse quantizes them. The inverse quantization process may include determining the degree of quantization to be applied by using the quantization parameter calculated by the video encoder 20 for each image block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform module 88 performs an inverse transform, e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process, on the transform coefficients to generate a pixel domain residual block.
[0134] After the motion compensation unit 82 generates a prediction block for the current image block, the video decoder 30 adds the residual block from the inverse transform module 88 and the corresponding prediction block generated by the motion compensation unit 82 to obtain a reconstructed block, i.e., a decoded image block. The adder 90 represents a component that performs the addition operation. When necessary, a loop filter (within or after the decoding loop) can be further used to smooth the pixel transformation, or the video quality can be improved in another way. The filter unit (not shown) can represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Further, the decoded image blocks within a given frame or image may be further stored in the decoded image buffer 92, and the decoded image buffer 92 stores reference images used for subsequent motion compensation. The decoded image buffer 92 may be part of the memory and may further store the decoded video for subsequent display on a display device (e.g., the display device 32 of FIG. 1). Alternatively, the decoded image buffer 92 may be separated from such memory.
[0135] Another structural variation of the video decoder 30 should be understood to be usable for decoding an encoded video bitstream. For example, the video decoder 30 can generate an output video stream without the processing by the filter unit. Alternatively, for some image blocks or image frames, the entropy decoding module 80 of the video decoder 30 does not obtain the quantized coefficients through decoding, and correspondingly, the inverse quantization module 86 and the inverse transform module 88 are not required to process. For example, the inverse quantization module 86 and the inverse transform module 88 within the video decoder 30 can be integrated.
[0136] Figure 3 is a schematic flowchart of an image prediction method according to an embodiment of the present application. The method shown in Figure 3 can be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in Figure 3 can be used in an encoding process or a decoding process. More specifically, the method shown in Figure 3 can be used in an inter-frame prediction process during encoding or decoding. Process 300 may be executed by the video encoder 20 or the video decoder 30, and in particular, may be executed by the motion compensation unit of the video encoder 20 or the video decoder 30. For a video data stream having a plurality of video frames, the video encoder or the video decoder executes Process 300 including the following steps to predict a predicted value of the pixel values of the current image block of the current video frame for use.
[0137] The method shown in Figure 3 includes steps 301 to 304, which will be described in detail below with respect to steps 301 to 304.
[0138] 301: Obtain the initial motion information of the current image block.
[0139] The image block in this specification may be an image block in an image to be processed or a sub-image within the image to be processed. Further, the image block in this specification may be an image block to be encoded in an encoding process or an image block to be decoded in a decoding process.
[0140] Furthermore, the initial motion information may include indication information of a prediction direction (usually bi-directional prediction), a motion vector pointing to a reference image block (usually a motion vector of an adjacent block), and information of the image in which the reference image block is located (usually understood as reference image information). The motion vector includes a forward motion vector and a backward motion vector, and the reference image information includes reference frame index information of a forward prediction reference image block and a backward prediction reference image block.
[0141] The initial motion information of the image block can be obtained in multiple ways. For example, the initial motion information of the image block can be obtained by the following Method 1 and Method 2.
[0142] Method 1: Referring to FIGS. 4 and 5, in the merge mode of inter-frame prediction, the candidate motion information list is constructed based on the motion information of adjacent blocks of the current image block, and one candidate motion information is selected from the candidate motion information list as the initial motion information of the current image block. The candidate motion information list includes motion vectors, reference frame index information, and the like. For example, the motion information of adjacent block A0 (referring to the candidate motion information with index 0 in FIG. 5) is selected as the initial motion information of the current image block. In particular, the forward motion vector of A0 is used as the forward prediction motion vector of the current block, and the backward motion vector of A0 is used as the backward prediction motion vector of the current block.
[0143] Method 2: In the non-merge mode of inter-frame prediction, the motion vector prediction value list is constructed based on the motion information of adjacent blocks of the current image block, and the motion vector is selected from the motion vector prediction value list as the motion vector prediction value of the current image block. In this case, the motion vector of the current image block may be the motion vector value of the adjacent block or the sum of the differences between the motion vector of the selected adjacent block and the motion vector of the current image block. The motion vector difference is the difference between the motion vector obtained by performing motion estimation on the current image block and the motion vector of the selected adjacent block. For example, the motion vectors corresponding to indices 1 and 2 in the motion vector prediction value list are selected as the forward motion vector and the backward motion vector of the current image block.
[0144] It should be understood that the foregoing Method 1 and Method 2 are merely two specific methods for obtaining the initial motion information of an image block. In this application, image the method for obtaining initial the motion information of a block is not limited, and any method that can obtain the initial motion information of an image block shall fall within the protection scope of this application.
[0145] 302: Determine the positions of N forward reference blocks and the positions of N backward reference blocks based on the initial motion information and position of the current image block. The N forward reference blocks are arranged in the forward reference image, and the N backward reference blocks are arranged in the backward reference image. N is an integer greater than 1.
[0146] Referring to FIG. 6, the current image to which the current image block in this embodiment of this application belongs has two reference images, namely, a forward reference image and a backward reference image.
[0147] In one example, the initial motion information includes a first motion vector and a first reference image index in the forward prediction direction, and a second motion vector and a second reference image index in the backward prediction direction.
[0148] Correspondingly, step 302 includes Using the position of the first motion vector of the current image block and the position as the first search starting point (shown as (0, 0) in FIG. 8), determine the position of the initial forward reference block of the current image block in the forward reference image corresponding to the first reference image index, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image; Using the position of the second motion vector of the current image block and the position as the second search starting point, determine the position of the initial backward reference block of the current image block in the backward reference image corresponding to the second reference image index, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image.
[0149] In one example, referring to FIG. 7, the positions of the N forward reference blocks include the position of the initial forward reference block (indicated by (0, 0)) and the positions of (N - 1) candidate forward reference blocks (indicated by (0, -1), (-1, -1), (-1, 1), (1, -1), (1, 1), and the like), and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance (as shown in FIG. 8) or a fractional pixel distance, provided that N = 9. Or the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance, provided that N = 9.
[0150] Referring to FIG. 8, in the motion estimation or motion compensation process, the MV accuracy may be fractional pixel accuracy (e.g., 1 / 2 pixel accuracy or 1 / 4 pixel accuracy). If the image has only pixel values of integer pixels and the current MV accuracy is fractional pixel accuracy, in order to obtain the pixel value at the fractional pixel position, interpolation needs to be performed by using an interpolation filter and by using the pixel values at the integer pixel positions of the reference image, thereby obtaining the pixel value at the fractional pixel position, and the obtained pixel value is used as the value of the predicted block of the current block. A specific interpolation process is related to the interpolation filter used. Generally, the pixel values of the integer samples around the reference sample can be linearly weighted to obtain the value of the reference sample. Common interpolation filters include 4-tap, 6-tap, and 8-tap interpolation filters, and the like.
[0151] As shown in FIG. 7, Ai,j is a sample at an integer pixel position, and its bit width is bitDepth. a0,0, b0,0, c0,0, d0,0, h0,0, n0,0, e0,0, i0,0, p0,0, f0,0, j0,0, q0,0, g0,0, k0,0, and r0,0 are samples at fractional pixel positions. When an 8-tap interpolation filter is used, a0,0 can be obtained by calculating using the following formula. a0,0=(C0*A -3,0 +C1*A -2,0 +C2*A -1,0 +C3*A 0,0 +C4*A 1,0 +C5*A 2,0 +C6*A 3,0 +C7*A 4,0 )>> shift1
[0152] In the above formula, C k is the coefficient of the interpolation filter, where k = 0, 1,..., 7. If the sum of the coefficients of the interpolation filter is a power of 2, the gain of the interpolation filter is N. For example, N being 6 indicates that the gain of the interpolation filter is 6 bits. shift1 is the number of bits of right shift, and shift1 may be set to bitDepth - 8, where bitDepth is the target bit width. In this way, based on the above formula, the finally obtained bit width of the pixel value of the prediction block is bitDepth + 6 - shift1 = 14 bits.
[0153] 303: Based on the matching cost criterion, determine that the positions of a pair of reference blocks among the M pairs of reference block positions are the positions of the target forward reference block and the target backward reference block of the current image block. The positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block. For the positions of each pair of reference blocks, the first position offset and the second position offset are in a mirror image relationship. The first position offset represents the offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer greater than or equal to 1, and M is less than or equal to N.
[0154] Referring to FIG. 9, the offset of the position of the candidate forward reference block 904 in the forward reference image Ref0 with respect to the position of the initial forward reference block 902 (i.e., the forward search base point) is MVD0 (delta0x, delta0y). The offset of the position of the candidate backward reference block 905 in the backward reference image Ref1 with respect to the position of the initial backward reference block 903 (i.e., the backward search base point) is MVD1 (delta1x, delta1y).
[0155] MVD0 = -MVD1. Specifically, delta0x = -delta1x, and delta0y = -delta1y.
[0156] In different examples, step 303 is From the positions of M pairs of reference blocks (one forward reference block and one backward reference block), it may be determined that the positions of the pair of reference blocks with the minimum matching error are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, or from the positions of the M pairs of reference blocks, it may be determined that the positions of the pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, provided that M is less than or equal to N. Further, the difference between the pixel value of the forward reference block and the pixel value of the backward reference block may be measured by using the sum of absolute differences (SAD), the sum of absolute transformation differences (SATD), the sum of squared absolute differences, or the like.
[0157] 304: Obtain a predicted value of the pixel value of the current image block based on the pixel value of the target forward reference block and the pixel value of the target backward reference block.
[0158] In one example, in step 304, weighting processing is performed on the pixel value of the target forward reference block and the pixel value of the target backward reference block, thereby obtaining a predicted value of the pixel value of the current image block.
[0159] Optionally, in one embodiment, the method shown in FIG. 3 is to obtain updated motion information of the current image block, where the updated motion information includes an updated forward motion vector and an updated backward motion vector, the updated forward motion vector points to the position of the target forward reference block, and the updated backward motion vector points to the position of the target backward reference block, and further includes obtaining. The updated motion information of the current image block can be obtained based on the position of the target forward reference block, the position of the target backward reference block, and the position of the current image block, or is obtained based on a first position offset and a second position offset corresponding to the determined pair of reference block positions.
[0160] The motion vector of the image block is updated. In this way, another image block can be effectively predicted based on the image block during the next image prediction.
[0161] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. There is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (that is, the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (that is, the optimal backward reference block / backward prediction block) of the current image block. Thus, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0162] Next, the image prediction method in the embodiment of the present application will be described in detail with reference to FIG. 10.
[0163] FIG. 10 is a schematic flowchart of an image prediction method according to an embodiment of the present application. The method shown in FIG. 10 can be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 10 can be used in an encoding process or a decoding process. More specifically, the method shown in FIG. 10 can be used in an inter-frame prediction process during encoding or decoding.
[0164] The method shown in FIG. 10 includes steps 1001 to 1007, which will be described in detail below with respect to steps 1001 to 1007.
[0165] 1001: Obtain the initial motion information of the current block.
[0166] For example, for an image block whose inter-frame prediction / coding mode is merge, the group of motion information is obtained from the merge candidate list based on the merge index, and the motion information is the initial motion information of the current block. For example, for an image block whose inter-frame prediction / coding mode is AMVP, the MVP is obtained from the MVP candidate list based on the index of the AMVP mode, and the MV of the current block is obtained by obtaining the sum of the MVP and the MVD included in the bitstream. The initial motion information includes reference image indication information and a motion vector. The forward reference image and the backward reference image are determined by using the reference image indication information. The positions of the forward reference block and the backward reference block are determined by using the motion vector.
[0167] 1002: Determine the position of the start forward reference block of the current image block in the forward reference image, and the position of the start forward reference block is the search start point (also referred to as the search base point) in the forward reference image.
[0168] In particular, the search base point (hereinafter referred to as the first search base point) in the forward reference image is obtained based on the forward MV and position information of the current block. For example, the forward MV information is (MV0x, MV0y). The position information of the current block is (B0x, B0y). The first search base point in the forward reference image is (MV0x + B0x, MV0y + B0y).
[0169] 1003: Determine the position of the start backward reference block of the current image block in the backward reference image, and the position of the start backward reference block is the search start point in the backward reference image.
[0170] In particular, the search reference point in the backward reference image (hereinafter referred to as the second search reference point) is obtained based on the backward MV and position information of the current block. For example, the backward MV is (MV1x, MV1y), and the position information of the current block is (B0x, B0y). The second search reference point in the backward reference image is (MV1x + B0x, MV1y + B0y).
[0171] 1004: Based on the MVD mirroring constraint condition, determine the positions of a pair of most matching reference blocks (i.e., one forward reference block and one backward reference block), and obtain the optimal forward motion vector and the optimal backward motion vector.
[0172] The MVD mirroring constraint condition in this specification can be explained as follows. The offset of the block position in the forward reference image with respect to the forward search reference point is MVD0 (delta0x, delta0y), and the offset of the block position in the backward reference image with respect to the backward search reference point is MVD1 (delta1x, delta1y). The following relationship is satisfied. MVD0 = -MVD1, specifically, delta0x = -delta1x, and delta0y = -delta1y.
[0173] Referring to FIG. 7, in the forward reference image, motion search with integer pixel steps is performed by using the search base point (indicated by (0, 0)) as the starting point. The integer pixel step means that the offset of the position of the candidate reference block relative to the search base point is an integer pixel distance. It should be understood that regardless of whether the search base point is an integer sample (the starting point is an integer pixel or a sub-pixel, such as 1 / 2, 1 / 4, 1 / 8, or 1 / 16), the motion search with integer pixel steps may be first performed to obtain the position of the forward reference block of the current image block. When the search is performed by using integer pixel steps, the search starting point may be an integer pixel or a fractional pixel, for example, an integer pixel, 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, or 1 / 16 pixel.
[0174] As shown in FIG. 7, the point (0, 0) is used as the search base point, and eight search points with integer pixel steps around the search base point are searched to obtain the positions of the corresponding candidate reference blocks. FIG. 7 shows eight candidate reference blocks. If the offset of the position of the forward candidate reference block in the forward reference image relative to the position of the forward search base point is (-1, -1), then the offset of the position of the corresponding backward candidate reference block in the backward reference image relative to the position of the backward search base point is (1, 1). Thus, the positions of the paired forward candidate reference block and backward candidate reference block are obtained. For the obtained positions of a pair of reference blocks, the matching cost between the two corresponding candidate reference blocks is calculated. The forward reference block and backward reference block having the minimum matching cost are selected as the optimal forward reference block and the optimal backward reference block, and the optimal forward motion vector and the optimal backward motion vector are obtained.
[0175] 1005 and 1006: Execute the motion compensation process by using the optimal forward motion vectors obtained in step 1004 to obtain the pixel values of the optimal forward reference block, and execute the motion compensation process by using the optimal backward motion vectors obtained in step 1004 to obtain the pixel values of the optimal backward reference block.
[0176] 1007: Perform a weighting process on the obtained pixel values of the optimal forward reference block and the obtained pixel values of the optimal backward reference block to obtain a predicted value of the pixel values of the current image block.
[0177] In particular, the predicted value of the pixel values of the current image block can be obtained based on the following formula (2). predSamples'[x][y]=(predSamplesL0'[x][y]+predSamplesL1'[x][y]+1)>>1 (2)
[0178] In the above formula, predSamplesL0' is the optimal forward reference block, predSamplesL1' is the optimal backward reference block, predSamples' is the predicted block of the current image block, predSamplesL0'[x][y] is the pixel value of the optimal forward reference block at sample (x, y), predSamplesL1'[x][y] is the pixel value of the optimal backward reference block at sample (x, y), and predSamples'[x][y] is the final pixel value of the predicted block at sample (x, y).
[0179] Note that in this embodiment of the present application, the search method to be used is not limited, and any search method may be used. For each forward candidate block obtained through the search , previousThe difference between the forward candidate block and the corresponding backward candidate block is calculated, and the forward candidate block and backward candidate block having the minimum SAD, the forward motion vector corresponding to the forward candidate block, and the backward motion vector corresponding to the backward candidate block are each selected as the optimal forward reference block, the optimal backward reference block, the optimal forward reference block corresponding to the optimal forward motion vector, and the optimal backward reference block corresponding to the optimal backward motion vector. Alternatively, for each backward candidate block obtained through the search, the difference between the backward candidate block and the corresponding forward candidate block in step 4 is calculated, and the backward candidate block and forward candidate block having the minimum SAD, the backward motion vector corresponding to the backward candidate block, and the forward motion vector corresponding to the forward candidate block are each selected as the optimal backward reference block, the optimal forward reference block, the optimal backward motion vector corresponding to the optimal backward reference block, and the optimal forward motion vector corresponding to the optimal forward reference block.
[0180] Note that only an example of the search method based on integer pixel steps is presented in step 1004. In fact, in addition to the search using integer pixel steps, a search using fractional pixel steps can also be used. For example, in step 1004, after the search using integer pixel steps, a search using fractional pixel steps is executed. Alternatively, the search using fractional pixel steps is directly executed. The specific search method is not limited in this specification.
[0181] Note that in this embodiment of the present application, the method for calculating the matching cost is not limited. For example, the SAD criterion, the MR-SAD criterion, or another criterion may also be used. Further, the matching cost can be calculated by using only the luminance component or by using both the luminance component and the chrominance component.
[0182] Note that in the exploration process, when the matching cost is 0 or reaches a preset threshold value, the traversal operation or exploration operation can be terminated in advance. The early termination condition of the exploration method is not limited in this specification.
[0183] The sequences of step 1005 and step 1006 are not limited, and it should be understood that these can be executed simultaneously or sequentially.
[0184] In the existing method, it can be seen that the template matching block needs to be calculated first, and the forward search and backward search are executed separately by using the template matching block. However, in this embodiment of the present application, in the process of searching for the matching block, the matching cost is directly calculated by using the candidate block in the forward reference image and the candidate block in the backward reference image, thereby determining two blocks with the minimum matching cost. This simplifies the image prediction process, improves the image prediction accuracy, and reduces the complexity.
[0185] FIG. 11 is a schematic flowchart of an image prediction method according to an embodiment of the present application. The method shown in FIG. 11 can be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 11 includes steps 1101 to 1105. For steps 1101 to 1103 and step 1105, refer to the descriptions of steps 1001 to 1003 and step 1007 in FIG. 10. Details will not be described again here.
[0186] The difference between this embodiment of the present application and the embodiment shown in FIG. 10 is that the pixel values of the current optimal forward and backward reference blocks are retained and updated in the search process. After the search is completed, the predicted value of the pixel value of the current image block can be calculated by using the pixel values of the current optimal forward and backward reference blocks.
[0187] For example, the positions of N pairs of reference blocks need to be traversed. Costi is the i-th matching cost, and MinCost indicates the current minimum matching cost. Bfi is the pixel value of the forward reference block, Bbi is the pixel value of the backward reference block, and the pixel values are obtained at the i-th time. BestBf is the value of the current optimal forward reference block, and BestBb is the value of the current optimal backward reference block. CalCost(M, N) represents the matching cost between block M and block N.
[0188] When the search starts (i = 1), MinCost = Cost0 = CalCost(Bf0, Bb0), BestBf = Bf0, and BestBb = Bb0.
[0189] When the other pairs of reference blocks are then traversed, BestBf and BestBb is in real time update is performed. For example, when the i-th (i > 1) search is executed, if Costi < MinCost, then BestBf = Bfi and BestBb = Bbi; otherwise, no update is performed.
[0190] When the search is finished, BestBf and BestBb are used to obtain the predicted value of the pixel value of the current block.
[0191] FIG. 12 is a schematic flowchart of an image prediction method according to an embodiment of the present application. The method shown in FIG. 12 can be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 12 can be used in an encoding process or a decoding process. More specifically, the method shown in FIG. 12 can be used in an inter-frame prediction process during encoding or decoding. Process 1200 may be executed by video encoder 20 or video decoder 30, and in particular, may be executed by the motion compensation unit of video encoder 20 or video decoder 30. For a video data stream having a plurality of video frames, the video encoder or video decoder executes Process 1200, including the following steps, to obtain a predicted value of the pixel values of the current image block of the current video frame.
[0192] The method shown in FIG. 12 includes steps 1201 to 1204. For steps 1201, 1202, and 1204, refer to the descriptions of steps 301, 302, and 304 in FIG. 3. Details will not be described again here.
[0193] The differences between this embodiment of the present application and the embodiment shown in FIG. 3 are as follows. In step 1203, from the positions of M pairs of reference blocks based on the matching cost criterion, the positions of a pair of reference blocks are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block. The positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block. For the positions of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time-domain distance. The first position offset represents the offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer greater than or equal to 1, and M is less than or equal to N.
[0194] Referring to FIG. 13, the offset of the position of the candidate forward reference block 1304 in the forward reference image Ref0 with respect to the position of the initial forward reference block 1302 (i.e., the forward search base point) is MVD0 (delta0x, delta0y). The offset of the position of the candidate backward reference block 1305 in the backward reference image Ref1 with respect to the position of the initial backward reference block 1303 (i.e., the backward search base point) is MVD1 (delta1x, delta1y).
[0195] In the search process, the position offsets of the two matching blocks placement satisfy the condition of a mirror image relationship, and the time-domain interval needs to be considered in the mirror image relationship. In this specification, TC, T0, and T1 represent the time points of the current frame, the time point of the forward reference image, and the time point of the backward reference image, respectively. TD0 and TD1 indicate the time intervals between the two time points.
[0196] TD0 = TC - T0, and TD1 = TC - T1.
[0197] In a specific encoding process, TD0 and TD1 can be calculated by using the picture order count (POC). For example, it is as follows. TD0 = POCc - POC0 and TD1 = POCc - POC1.
[0198] In this specification, POCc, POC0, and POC1 represent the POC of the current picture, the POC of the forward reference picture, and the POC of the backward reference picture, respectively. TD0 represents the picture order count (POC) distance between the current picture and the forward reference picture, and TD1 represents the POC distance between the current picture and the backward reference picture.
[0199] delta0 = (delta0x, delta0y), and delta1 = (delta1x, delta1y).
[0200] The mirror image relationship considering the time domain interval is described as follows. delta0x = (TD0 / TD1) * delta1x, and delta0y = (TD0 / TD1) * delta1y, or delta0x / delta1x = (TD0 / TD1), and delta0y / delta1y = (TD0 / TD1).
[0201] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. There is a proportional relationship based on the time-domain distance between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0202] In the foregoing embodiment, the search process is executed once. Furthermore, the search can be executed multiple times by using an iterative method. In particular, after the forward reference block and the backward reference block are obtained in each round of the search, the search can be executed once or multiple times based on the current refined MV.
[0203] One process of the image prediction method in one embodiment of the present application will be described in detail below with reference to FIG. 14. Similar to the method shown in FIG. 3, the method shown in FIG. 14 can also be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 14 can be used in an encoding process or a decoding process. In particular, the method shown in FIG. 14 can be used in an inter-frame prediction process during encoding or decoding.
[0204] The method shown in FIG. 14 particularly includes the following steps 1401 to 1404.
[0205] 1401: Obtain the i-th motion information of the current image block.
[0206] The image block in this specification may be an image block in the image to be processed or a sub-image within the image to be processed. Further, the image block in this specification may be an image block to be encoded in the encoding process or an image block to be decoded in the decoding process.
[0207] If i = 1, the i-th motion information is the initial motion information of the current image block.
[0208] If i>1, the i-th motion information includes a forward motion vector indicating the position of the (i - 1)-th target forward reference block and a backward motion vector indicating the position of the (i - 1)-th target backward reference block.
[0209] Furthermore, the initial motion information may include indication information of the prediction direction (usually bi-directional prediction), a motion vector pointing to a reference image block (usually a motion vector of an adjacent block), and information of the image in which the reference image block is located (usually understood as reference image information). The motion vector includes a forward motion vector and a backward motion vector, and the reference image information includes reference frame index information of the forward prediction reference image block and the backward prediction reference image block.
[0210] The initial motion information of the image block can be obtained in a plurality of ways. For example, the initial motion information of the image block can be obtained by the following Method 1 and Method 2.
[0211] Method 1: Referring to FIGS. 4 and 5, in the merge mode of inter-frame prediction, the candidate motion information list is constructed based on the motion information of the adjacent blocks of the current image block, and one candidate motion information is selected from the candidate motion information list as the initial motion information of the current image block. The candidate motion information list includes motion vectors, reference frame index information, and the like. For example, the motion information of adjacent block A0 (referring to the candidate motion information with index 0 in FIG. 5) is selected as the initial motion information of the current image block. In particular, the forward motion vector of A0 is used as the forward prediction motion vector of the current block, and the backward motion vector of A0 is used as the backward prediction motion vector of the current block.
[0212] Method 2: In the non-merge mode of inter-frame prediction, the motion vector prediction value list is constructed based on the motion information of the adjacent blocks of the current image block, and the motion vector is selected from the motion vector prediction value list as the motion vector prediction value of the current image block. In this case, the motion vector of the current image block may be the motion vector value of the adjacent block or the sum of the differences between the motion vector of the selected adjacent block and the motion vector of the current image block. The motion vector difference is the difference between the motion vector obtained by performing motion estimation on the current image block and the motion vector of the selected adjacent block. For example, the motion vectors corresponding to indices 1 and 2 in the motion vector prediction value list are selected as the forward motion vector and the backward motion vector of the current image block.
[0213] It should be understood that the foregoing Method 1 and Method 2 are merely two specific methods for obtaining the initial motion information of an image block. In this application, the method for obtaining the motion information of a prediction block is not limited, and any method capable of obtaining the initial motion information of an image block shall fall within the protection scope of this application.
[0214] 1402: Based on the motion information of the i-th time and the position of the current image block, determine the positions of N forward reference blocks and the positions of N backward reference blocks. The N forward reference blocks are arranged in the forward reference image, and the N backward reference blocks are arranged in the backward reference image. N is an integer greater than 1.
[0215] In one example, the motion information of the i-th time includes a forward motion vector, a forward reference image index, a backward motion vector, and a backward reference image index.
[0216] Correspondingly, step 1402 is Based on the forward motion vector and the position of the current image block, use the position of the (i - 1)-th target forward reference block as the search start point of the i-th f to determine the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, and Based on the backward motion vector and the position of the current image block, use the position of the (i - 1)-th target backward reference block as the search start point of the i-th b to determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image.
[0217] In one example, referring to FIG. 7, the positions of the N forward reference blocks include the position of the i-th target forward reference block (indicated by (0, 0)) and the positions of the (N - 1) candidate forward reference blocks (indicated by (0, -1), (-1, -1), (-1, 1), (1, -1), (1, 1), and the like), and the offset of the position of each candidate forward reference block with respect to the position of the i-th target forward reference block is an integer pixel distance (as shown in FIG. 8) or a fractional pixel distance, provided that N = 9, or the positions of the N backward reference blocks include the position of the i-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the i-th target backward reference block is an integer pixel distance or a fractional pixel distance, provided that N = 9.
[0218] 1403: Based on the matching cost criterion, from the positions of the M pairs of reference blocks, it is determined that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block. The positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block. For the positions of each pair of reference blocks, the first position offset and the second position offset are in a mirror image relationship. The first position offset represents the offset of the position of the forward reference block with respect to the position of the (i - 1)-th target forward reference block, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i - 1)-th target backward reference block. M is an integer greater than or equal to 1, and M is less than or equal to N.
[0219] That the first position offset and the second position offset are in a mirror image relationship can be understood as follows. The direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset.
[0220] 9, the offset of the position of the candidate forward reference block 904 in the forward reference image Ref0 relative to the position of the (i-1)th target forward reference block 902 (i.e., the forward search base point) is MVD0(delta0x, delta0y). The offset of the position of the candidate backward reference block 905 in the backward reference image Ref1 relative to the position of the (i-1)th target backward reference block 903 (i.e., the backward search base point) is MVD1(delta1x, delta1y).
[0221] MVD0=-MVD1, specifically, delta0x=-delta1x, and delta0y=-delta1y.
[0222] In a different example, step 1403 includes: It may include determining that from the positions of M pairs of reference blocks (one forward reference block and one backward reference block), the positions of a pair of reference blocks with the smallest matching error are the positions of the i-th target forward reference block of the current image block and the i-th target backward reference block of the current image block, or determining that from the positions of M pairs of reference blocks, the positions of a pair of reference blocks with a matching error equal to or less than a matching error threshold are the positions of the i-th target forward reference block of the current image block and the i-th target backward reference block of the current image block, where M is equal to or less than N. Furthermore, the difference between the pixel values of the forward reference block and the pixel values of the backward reference block may be measured by using a sum of absolute differences (SAD), a sum of absolute transformation differences (SATD), a sum of absolute squared differences, or the like.
[0223] 1404: Obtain a predicted value of a pixel value of the current image block based on a pixel value of the target forward reference block and a pixel value of the target backward reference block.
[0224] In one example, in step 1404 , weighted processing is performed on the pixel values of the target forward reference block and the pixel values of the target backward reference block, thereby obtaining a predicted value of the pixel values of the current image block. Further, in the present application, the predicted value of the pixel values of the current image block can alternatively be obtained by using another method. This is not limited in the present application.
[0225] The motion vector of the image block is updated. For example, the initial motion information is updated to the second motion information, and the second motion information includes a forward motion vector pointing to the position of the first target forward reference block and a backward motion vector pointing to the first target backward reference block. In this way, another image block can be effectively predicted based on the image block during the next image prediction.
[0226] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. There is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, the accuracy of refining the MV can be further improved by increasing the number of repetitions, whereby the coding performance can be further improved.
[0227] A process of an image prediction method in an embodiment of the present application is described in detail below with reference to FIG. 15. The method shown in FIG. 15 can also be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 15 can be used in an encoding process or a decoding process. In particular, the method shown in FIG. 15 can be used in an inter-frame prediction process during encoding or decoding.
[0228] The method shown in FIG. 15 includes, in particular, steps 1501 to 1508, which will be described in detail below for steps 1501 to 1508.
[0229] 1501: Obtain the initial motion information of the current image block.
[0230] For example, in the first search, the initial motion information of the current block is used. For example, for an image block whose coding mode is merge, the motion information is obtained from the merge candidate list based on the index of the merge mode, and the motion information is the initial motion information of the current block. For example, for an image block whose coding mode is AMVP, the MVP is obtained from the MVP candidate list based on the index of the AMVP mode, and the MV of the current block is obtained by obtaining the sum of the MVP and MVD included in the bitstream. In a search that is not the first time, the MV information updated in the previous search is used. The motion information includes reference image indication information and motion vector information. The forward reference image and the backward reference image are determined by using the reference image indication information. The position of the forward reference block and the position of the backward reference block are determined by using the motion vector information.
[0231] 1502: Determine the search base point in the forward reference image.
[0232] The search base point in the forward reference image is determined based on the forward MV information and position information of the current block. The specific process is similar to the process in the embodiment of FIG. 10 or FIG. 11. For example, if the forward MV information is (MV0x, MV0y) and the position information of the current block is (B0x, B0y), the search base point in the forward reference image is (MV0x + B0x, MV0y + B0y).
[0233] 1503: Determine the search base point in the backward reference image.
[0234] The search base point in the backward reference image is determined based on the backward MV information and position information of the current block. The specific process is similar to the process in the embodiment of FIG. 10 or FIG. 11. For example, if the backward MV information is (MV1x, MV1y) and the position information of the current block is (B0x, B0y), the search base point in the backward reference image is (MV1x + B0x, MV1y + B0y).
[0235] 1504: In the forward reference image and the backward reference image, based on the MVD mirror constraint condition, determine the positions of a pair of most matching reference blocks (i.e., one forward reference block and one backward reference block), and obtain the refined forward motion vector and the refined backward motion vector of the current image block.
[0236] The specific search process is similar to the process in the embodiment of FIG. 10 or FIG. 11. Details are not described again here.
[0237] 1505: Determine whether the iteration end condition is satisfied. If the iteration end condition is not satisfied, execute steps 1502 and 1503. If the iteration end condition is satisfied, execute steps 1506 and 1507.
[0238] The design of the end condition of the iterative search is not limited in this specification. For example, the traversal can be executed based on a specified number of L iterations, or another iteration end condition is satisfied. For example, after the result of the current iterative operation is obtained, if MVD0 is close to or equal to 0, and MVD1 is close to or equal to 0, for example, if MVD0 = (0, 0) and MVD1 = (0, 0), the iterative operation can be terminated.
[0239] L is a preset value and is an integer greater than 1. The value of L can be a numerical value preset before the image is predicted, or the value of L can be set based on the accuracy of image prediction and the complexity in the search of prediction blocks, or L can be set based on past experience values, or L can be determined based on the verification of the results of the intermediate search process.
[0240] For example, in this embodiment, by using integer pixel steps, the search is performed a total of two times. During the first search, the position of the initial forward reference block may be used as the search base point, and the positions of (N - 1) candidate forward reference blocks are determined within the forward reference image (also referred to as the forward reference region). The position of the initial backward reference block is used as the search base point, and the positions of (N - 1) candidate backward reference blocks are determined within the backward reference image ( rear also referred to as the reference region). For one or more pairs of reference block positions among the positions of N pairs of reference blocks, the matching costs of two corresponding reference blocks are calculated. For example, the matching costs of the initial forward reference block and the initial backward reference block are calculated, and the matching costs of candidate forward reference blocks and candidate backward reference blocks that satisfy the MVD mirror constraint condition are calculated. In this way, the position of the first target forward reference block and the position of the first target backward reference block in the first search are obtained, and updated motion information is further obtained. The updated motion information includes a forward motion vector indicating that the position of the current block points to the position of the first target forward reference block, and a backward motion vector indicating that the position of the current image block points to the position of the first target backward reference block. It should be understood that the updated motion information and the initial motion information include the same reference frame index and the like. Next, the second search is performed. The position of the first target forward reference block is used as the search base point, and the positions of (N - 1) candidate forward reference blocks are determined within the forward reference image (also referred to as the forward reference region). The position of the first target backward reference block is used as the search base point, and the positions of the backward reference image ( rearWithin the (also referred to as the reference region), the positions of (N - 1) candidate backward reference blocks are determined. For the positions of one or more pairs of reference blocks among the positions of N pairs of reference blocks, the matching costs of two corresponding reference blocks are calculated. For example, the matching costs of the first target forward reference block and the first target backward reference block are calculated, and the matching costs of candidate forward reference blocks and candidate backward reference blocks that satisfy the MVD mirroring constraint condition are calculated. In this way, the positions of the second target forward reference block and the second target backward reference block in the second search are obtained, and updated motion information is further obtained. The updated motion information includes a forward motion vector indicating that the position of the current image block points to the position of the second target forward reference block, and a backward motion vector indicating that the position of the current image block points to the position of the second target backward reference block. It should be understood that the updated motion information and the initial motion information include other same information such as the reference frame index. When the preset number L of repetitions is 2, in the second search process in this specification, the second target forward reference block and the second target backward reference block are the finally obtained target forward reference block and target backward reference block (also referred to as the optimal forward reference block and the optimal backward reference block).
[0241] 1506 and 1507: Execute the motion compensation process by using the optimal forward motion vector obtained in step 1504 to obtain the pixel values of the optimal forward reference block, and execute the motion compensation process by using the optimal backward motion vector obtained in step 1504 to obtain the pixel values of the optimal backward reference block.
[0242] 1508: Obtain a predicted value of the pixel values of the current image block based on the pixel values of the optimal forward reference block and the pixel values of the optimal backward reference block obtained in steps 1506 and 1507.
[0243] In step 1504, a search (also referred to as motion search) within the forward reference image or the backward reference image is performed by using integer pixel steps, whereby the positions of at least one forward reference block and at least one backward reference block can be obtained. When the search is performed by using integer pixel steps, the search start point may be an integer pixel or a fractional pixel, for example, an integer pixel, 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, or 1 / 16 pixel.
[0244] Furthermore, in step 1504, the fractional pixel steps can also be directly used to search for the positions of at least one forward reference block and at least one backward reference block, or both the search by using integer pixel steps and the search by using fractional pixel steps are performed. The search method is not limited in this application.
[0245] In step 1504, for each pair of reference blocks position when the difference between the pixel value of the forward reference block and the pixel value of the corresponding backward reference block is calculated, the difference between the pixel value of each forward reference block and the pixel value of the corresponding backward reference block can be measured by using SAD, SATD, sum of absolute squared differences, or the like. However, this application is not limited thereto.
[0246] When the predicted value of the pixel value of the current image block is determined based on the optimal forward prediction block and the optimal backward prediction block, the weighting process may be performed on the pixel values of the optimal forward reference block and the optimal backward reference block obtained in step 1506 and step 1507, and the pixel value obtained after the weighting process is used as the predicted value of the pixel value of the current image block.
[0247] In particular, the predicted value of the pixel value of the current image block can be obtained based on the following formula (8). predSamples'[x][y]=(predSamplesL0'[x][y]+predSamplesL1'[x][y]+1)>>1 (8)
[0248] In the above formula, predSamplesL0'[x][y] is the pixel value of the optimal forward reference block at sample (x, y), predSamplesL1'[x][y] is the pixel value of the optimal backward reference block at sample (x, y), and predSamples'[x][y] is the pixel prediction value of the current image block at sample (x, y).
[0249] Referring to FIG. 11, the pixel values of the current optimal forward reference block and the current optimal backward reference block can be further retained and updated in the iterative search process in this embodiment of the present application. After the search is completed, the predicted value of the pixel value of the current image block is directly calculated by using the pixel values of the current optimal forward and backward reference blocks. In this implementation, steps 1506 and 1507 are optional steps.
[0250] For example, the positions of N pairs of reference blocks need to be traversed. Costi is the i-th matching cost, and MinCost represents the current minimum matching cost. Bfi is the pixel value of the forward reference block, Bbi is the pixel value of the backward reference block, and the pixel values are obtained at the i-th time. BestBf is the pixel value of the current optimal forward reference block, and BestBb is the pixel value of the current optimal backward reference block. CalCost(M, N) represents the matching cost between block M and block N.
[0251] When the search starts (i = 1), MinCost = Cost0 = CalCost(Bf0, Bb0), BestBf = Bf0, and BestBb = Bb0.
[0252] When other pairs of reference blocks are traversed thereafter, the update is performed in real time. For example, when the i-th (i > 1) search is performed, if Costi < MinCost, then BestBf = Bfi and BestBb = Bbi; otherwise, the update is not performed.
[0253] When the search is finished, BestBf and BestBb are used to obtain the predicted values of the pixel values of the current block.
[0254] In the foregoing embodiment shown in FIG. 12, the search process is performed once. Further, the search can be performed multiple times by using an iterative method. In particular, after the forward reference block and the backward reference block are obtained in each round of the search, the search can be performed once or multiple times based on the current refined MV.
[0255] A process of an image prediction method 1600 in an embodiment of the present application is described in detail below with reference to FIG. 16. The method shown in FIG. 16 can also be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 16 can be used in an encoding process or a decoding process. In particular, the method shown in FIG. 16 can be used in an inter-frame prediction process during encoding or decoding.
[0256] The method 1600 shown in FIG. 16 includes steps 1601 to 1604. For steps 1601, 1602, and 1604, refer to the descriptions of steps 1401, 1402, and 1404 in FIG. 14. Details will not be described again here.
[0257] The differences between this embodiment of the present application and the embodiment shown in FIG. 14 are as follows. In step 1603, from the positions of M pairs of reference blocks based on the matching cost criterion, the positions of a pair of reference blocks are determined as the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block. The positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block. For the positions of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time-domain distance. The first position offset represents the offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer greater than or equal to 1, and M is less than or equal to N.
[0258] Referring to FIG. 13, the offset of the position of the candidate forward reference block 1304 in the forward reference image Ref0 with respect to the position of the initial forward reference block 1302 (i.e., the forward search base point) is MVD0 (delta0x, delta0y). The offset of the position of the candidate backward reference block 1305 in the backward reference image Ref1 with respect to the position of the initial backward reference block 1303 (i.e., the backward search base point) is MVD1 (delta1x, delta1y).
[0259] In the search process, the position offsets of the two matching blocks satisfy the condition of mirror image relationship, and the time-domain interval needs to be considered in the mirror image relationship. In this specification, TC, T0, and T1 represent the time points of the current frame, the forward reference image, and the backward reference image, respectively. TD0 and TD1 indicate the time intervals between the two time points. placement TD0 = TC - T0, and
[0260] TD1 = TC - T1. TD1 = TC - T1.
[0261] In a specific encoding process, TD0 and TD1 can be calculated by using the picture order count (POC). For example, it is as follows. TD0 = POCc - POC0 and TD1 = POCc - POC1.
[0262] In this specification, POCc, POC0, and POC1 represent the POC of the current picture, the POC of the forward reference picture, and the POC of the backward reference picture, respectively. TD0 represents the picture order count (POC) distance between the current picture and the forward reference picture, and TD1 represents the POC distance between the current picture and the backward reference picture.
[0263] delta0 = (delta0x, delta0y), and delta1 = (delta1x, delta1y).
[0264] The mirror image relationship considering the time domain interval is described as follows. delta0x = (TD0 / TD1) * delta1x, and delta0y = (TD0 / TD1) * delta1y, or delta0x / delta1x = (TD0 / TD1), and delta0y / delta1y = (TD0 / TD1).
[0265] In different examples, step 1603 is From the positions of M pairs of reference blocks (one forward reference block and one backward reference block), it may be determined that the positions of the pair of reference blocks with the minimum matching error are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, or from the positions of the M pairs of reference blocks, it may be determined that the positions of the pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, provided that M is less than or equal to N. Further, the difference between the pixel values of the forward reference block and the pixel values of the backward reference block may be measured by using the sum of absolute differences (SAD), the sum of absolute transformation differences (SATD), the sum of absolute squared differences, or the like.
[0266] In this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form N pairs of reference block positions. For each pair of reference block positions among the N pairs of reference block positions, there is a proportional relationship based on the time-domain distance between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, the pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, the accuracy of refining the MV can be further improved by increasing the number of repetitions, whereby the coding performance can be further improved.
[0267] A process of an image prediction method in an embodiment of the present application is described in detail below with reference to FIG. 17. The method shown in FIG. 17 can also be executed by a video encoding device, a video codec, a video encoding system, or another device having a video encoding function. The method shown in FIG. 17 can be used in an encoding process or a decoding process. In particular, the method shown in FIG. 17 can be used in an inter-frame prediction process during encoding or decoding.
[0268] The method shown in FIG. 17 includes steps 1701 to 1708. For steps 1701 to 1703 and steps 1705 to 1708, refer to the descriptions of steps 1501 to 1503 and steps 1505 to 1508 in FIG. 15. For details, they will not be described again here.
[0269] The difference between this embodiment of the present application and the embodiment shown in FIG. 15 is as follows.
[0270] 1704: Based on the MVD mirroring constraint condition considering the time-domain distance, determine the positions of a pair of most-matching reference blocks (i.e., one forward reference block and one backward reference block), and obtain the refined forward motion vector and refined backward motion vector of the current image block.
[0271] The MVD mirroring constraint condition based on the time-domain distance can be described as follows in this specification. The position offsets MVD0 (delta0x, delta0y) of the block position in the forward reference image with respect to the forward search base point and the position offsets MVD1 (delta1x, delta1y) of the block position in the backward reference image with respect to the backward search base point satisfy the following relationship.
[0272] The position offsets of the two matching blocks placement satisfy the conditions of the mirroring relationship based on the time-domain distance. In this specification, TC, T0, and T1 represent the time points of the current frame, the forward reference image, and the backward reference image, respectively. TD0 and TD1 indicate the time intervals between the two time points.
[0273] TD0 = TC - T0, and TD1 = TC - T1.
[0274] In a specific encoding process, TD0 and TD1 can be calculated by using the picture order count (POC). For example, it is as follows. TD0 = POCc - POC0 and TD1 = POCc - POC1.
[0275] In this specification, POCc, POC0, and POC1 represent the POC of the current image, the POC of the forward reference image, and the POC of the backward reference image, respectively. TD0 represents the picture order count (POC) distance between the current image and the forward reference image, and TD1 represents the POC distance between the current image and the backward reference image.
[0276] delta0 = (delta0x, delta0y), and delta1 = (delta1x, delta1y).
[0277] The mirror relationship considering the temporal domain distance (also referred to as the temporal domain interval) is described as follows. delta0x = (TD0 / TD1) * delta1x, and delta0y = (TD0 / TD1) * delta1y, or delta0x / delta1x = (TD0 / TD1), and delta0y / delta1y = (TD0 / TD1).
[0278] A specific search process is similar to the process in the embodiment of FIG. 10 or FIG. 11. Details are not described here again.
[0279] In this embodiment of the present application, it should be understood that the temporal domain interval may or may not be considered in the mirror relationship. In actual use, whether the temporal domain interval is considered in the mirror relationship when motion vector refinement is performed on the current frame or the current block can be adaptively selected.
[0280] For example, the indication information may be added to sequence level header information (SPS), picture level header information (PPS), slice header, or block bitstream information indicating whether a time interval is considered in the mirroring relationship used for the current sequence, current picture, current slice, or current block.
[0281] Alternatively, based on the POC of the forward reference picture and the POC of the backward reference picture, it is adaptively determined whether a time interval is considered in the mirroring relationship used for the current block in the current block.
[0282] For example, if |POCc - POC0| - |POCc - POC1| > T, then the interval needs to be considered for the mirroring relationship used, otherwise the time interval is not considered for the mirroring relationship used. T is a preset threshold in this specification. For example, T = 2 or T = 3. The specific value of T is not limited in this specification.
[0283] In another example, it is assumed that the ratio of the larger value of |POCc - POC0| and |POCc - POC1| to the smaller value of |POCc - POC0| and |POCc - POC1| is greater than a threshold R, that is, (Max(|POCc - POC0|, |POCc - POC1|) / Min(|POCc - POC0|, |POCc - POC1|)) > R is true.
[0284] Max(A, B) indicates the larger value of A and B, and Min(A, B) indicates the smaller value of A and B.
[0285] In this case, the interval needs to be considered for the mirroring relationship being used. If the ratio of the larger value of |POCc - POC0| and |POCc - POC1| to the smaller value of |POCc - POC0| and |POCc - POC1| is less than or equal to the threshold R, the time interval is not considered for the mirroring relationship being used. R is a preset threshold in this specification. For example, R = 2 or R = 3. The specific value of R is not limited in this specification.
[0286] It should be understood that the image prediction method in this embodiment of the present application can be particularly executed by a motion compensation module in an encoder (for example, encoder 20) or a decoder (for example, decoder 30). Furthermore, the image prediction method in this embodiment of the present application can be executed by any electronic device or apparatus that needs to encode and / or decode video images.
[0287] Next, the image prediction apparatus in the embodiment of the present application will be described in detail with reference to FIGS. 18 to 21.
[0288] FIG. 18 is a schematic block diagram of an image prediction apparatus according to an embodiment of the present application 1800 It should be noted that the prediction apparatus 1800 is applicable to both inter-frame prediction for decoding video images and inter-frame prediction for encoding video images. It should be understood that the prediction apparatus 1800 in this specification may correspond to the motion compensation unit 44 in FIG. 2A or the motion compensation unit 82 in FIG. 2B. The prediction apparatus 1800 is a first acquisition unit 1801 configured to acquire initial motion information of the current image block, and A first search unit 1802 that determines the positions of N forward reference blocks and the positions of N backward reference blocks based on initial motion information and the position of a current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, and N is an integer greater than 1; and determines, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, where the position of each pair of reference blocks includes the position of the forward reference block and the position of the backward reference block, and for the position of each pair of reference blocks, a first position offset and a second position offset are in a mirror image relationship, the first position offset represents the offset of the position of the forward reference block with respect to the position of an initial forward reference block, the second position offset represents the offset of the position of the backward reference block with respect to the position of an initial backward reference block, M is an integer greater than or equal to 1, and M is less than or equal to N; and a first prediction unit 1803 configured to obtain a predicted value of the pixel value of the current image block based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. It may include a first prediction unit 1803 configured to obtain a predicted value of the pixel value of the current image block based on the pixel value of the target forward reference block and the pixel value of the target backward reference block.
[0289] The fact that the first position offset and the second position offset are in a mirror image relationship can be understood as the first position offset value being the same as the second position cf. For example, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset.
[0290] Preferably, within the apparatus 1800 in this embodiment of the present application, the first prediction unit 1803 is further configured to obtain updated motion information of the current image block, where the updated motion information includes an updated forward motion vector and an updated backward motion vector, the updated forward motion vector points to the position of the target forward reference block, and the updated backward motion vector points to the position of the target backward reference block.
[0291] It can be seen that the motion vector of the image block has been updated. In this way, another image block can be effectively predicted based on the image block during the next image prediction.
[0292] In the apparatus 1800 in this embodiment of the present application, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, or the positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0293] In the apparatus 1800 in this embodiment of the present application, the initial motion information includes a first motion vector and a first reference image index in the forward prediction direction, and a second motion vector and a second reference image index in the backward prediction direction.
[0294] In a manner of determining the positions of the N forward reference blocks and the positions of the N backward reference blocks based on the initial motion information and the position of the current image block, the first search unit Based on the first motion vector and the position of the current image block, determine the position of the initial forward reference block of the current image block in the forward reference image corresponding to the first reference image index, use the position of the initial forward reference block as the first search starting point, determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, and the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of the (N - 1) candidate forward reference blocks. Based on the second motion vector and the position of the current image block, determine the position of the initial backward reference block of the current image block in the backward reference image corresponding to the second reference image index, use the position of the initial backward reference block as the second search starting point, determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, and the positions of the N backward reference blocks are specifically configured to include the position of the initial backward reference block and the positions of the (N - 1) candidate backward reference blocks.
[0295] In the apparatus 1800 in this embodiment of the present application, in a manner of determining from the positions of M pairs of reference blocks based on the matching cost criterion that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, the first search unit 1802 determines from the positions of M pairs of reference blocks whether the positions of the pair of reference blocks with the minimum matching error are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, or is specifically configured to determine from the positions of M pairs of reference blocks that the positions of the pair of reference blocks with a matching error less than or equal to the matching error threshold are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, provided that M is less than or equal to N.
[0296] Device 1800 may be configured to execute the methods shown in FIGS. 3, 10, and 11, and it should be understood that device 1800 may in particular be a video encoding device, a video decoding device, a video encoding system, or another device having a video encoding function. Device 1800 may be configured not only to perform image prediction in the encoding process, but also to perform image prediction in the decoding process.
[0297] For details, refer to the description of the image prediction method in this specification. For the sake of brevity, the details are not described again here.
[0298] It can be seen that according to the prediction device in this embodiment of the present application, the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form N pairs of reference block positions. For each pair of reference block positions among the N pairs of reference block positions, there is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (that is, the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (that is, the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0299] FIG. 19 is a schematic block diagram of another image prediction apparatus according to an embodiment of the present application. It should be noted that the prediction apparatus 1900 is applicable to both inter-frame prediction for decoding a video image and inter-frame prediction for encoding a video image. It should be understood that the prediction apparatus 1900 in this specification may correspond to the motion compensation unit 44 in FIG. 2A or may correspond to the motion compensation unit 82 in FIG. 2B. The prediction apparatus 1900 a second acquisition unit 1901 configured to acquire initial motion information of a current image block, a second search unit 1902 that determines the positions of N forward reference blocks and the positions of N backward reference blocks based on the initial motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image and the N backward reference blocks are arranged in a backward reference image, N is an integer greater than 1, and based on a matching cost criterion, from the positions of M pairs of reference blocks, it is determined that the positions of a pair of reference blocks are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block. The position of each pair of reference blocks includes the position of the forward reference block and the position of the backward reference block. For the position of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on the time-domain distance. The first position offset represents the offset of the position of the forward reference block with respect to the position of the initial forward reference block, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the initial backward reference block. M is an integer of 1 or more and M is less than or equal to N, and a second search unit 1902 configured to perform the above operations, and a second prediction unit 1903 configured to obtain a predicted value of the pixel value of the current image block based on the pixel values of the target forward reference block and the target backward reference block.
[0300] For each pair of reference blocks, the fact that the first position offset and the second position offset are in a proportional relationship based on the time-domain distance means For each pair of reference blocks, the proportional relationship between the first position offset and the second position offset is determined based on the proportional relationship between the first time-domain distance and the second time-domain distance. The first time-domain distance represents the time-domain distance between the current image to which the current image block belongs and the forward reference image, and the second time-domain distance represents the time-domain distance between the current image and the backward reference image. It can be understood as such.
[0301] In one implementation, the fact that the first position offset and the second position offset are in a proportional relationship based on the time-domain distance means that if the first time-domain distance is the same as the second time-domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset, or if the first time-domain distance is different from the second time-domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the proportional relationship between the amplitude value of the first position offset and the amplitude value of the second position offset is based on the proportional relationship between the first time-domain distance and the second time-domain distance. This may be included.
[0302] The first time-domain distance represents the time-domain distance between the current image to which the current image block belongs and the forward reference image, and the second time-domain distance represents the time-domain distance between the current image and the backward reference image.
[0303] Optimally, within the device in this embodiment, the second prediction unit 1903 is further configured to obtain updated motion information of the current image block. The updated motion information includes an updated forward motion vector and an updated backward motion vector. The updated forward motion vector points to the position of the target forward reference block, and the updated backward motion vector points to the position of the target backward reference block.
[0304] It can be seen that the refined motion information of the current image block can be obtained in this embodiment of the present application. This improves the accuracy of the motion information of the current image block, smooths the prediction of another image block, and for example, improves the prediction accuracy of the motion information of another image block.
[0305] In one implementation, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, or the positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0306] In one implementation, the initial motion information includes forward prediction motion information and backward prediction motion information. In a manner of determining the positions of the N forward reference blocks and the N backward reference blocks based on the initial motion information and the position of the current image block, the second search unit 1902 determines the positions of the N forward reference blocks in the forward reference image based on the forward prediction motion information and the position of the current image block, and the positions of the N forward reference blocks include the position of the initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, and is specifically configured to determine the positions of the N backward reference blocks in the backward reference image based on the backward prediction motion information and the position of the current image block, and the positions of the N backward reference blocks include the position of the initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance.
[0307] In another implementation form, the initial motion information includes a first motion vector and a first reference image index in the forward prediction direction, and a second motion vector and a second reference image index in the backward prediction direction, In a mode of determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the initial motion information and the position of the current image block, the second search unit Based on the first motion vector and the position of the current image block, determine the position of the initial forward reference block of the current image block in the forward reference image corresponding to the first reference image index, use the position of the initial forward reference block as the first search start point, and determine the positions of (N - 1) candidate forward reference blocks in the forward reference image. The positions of the N forward reference blocks include the position of the initial forward reference block and the positions of the (N - 1) candidate forward reference blocks, Based on the second motion vector and the position of the current image block, determine the position of the initial backward reference block of the current image block in the backward reference image corresponding to the second reference image index, use the position of the initial backward reference block as the second search start point, and determine the positions of (N - 1) candidate backward reference blocks in the backward reference image. The positions of the N backward reference blocks are particularly configured to include the position of the initial backward reference block and the positions of the (N - 1) candidate backward reference blocks.
[0308] In one implementation form, in a mode of determining that the positions of a pair of reference blocks from the positions of M pairs of reference blocks based on the matching cost criterion are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, the second search unit 1902 Determine from the positions of the M pairs of reference blocks whether the positions of the pair of reference blocks with the minimum matching error are the position of the target forward reference block of the current image block and the position of the target backward reference block of the current image block, or Particularly configured to determine that, from the positions of M pairs of reference blocks, the positions of a pair of reference blocks with a matching error less than or equal to a matching error threshold are the positions of the target forward reference block and the target backward reference block of the current image block, provided that M is less than or equal to N.
[0309] In one example, the matching cost criterion is a matching cost minimization criterion. For example, for the positions of M pairs of reference blocks, the difference between the pixel values of the forward reference block and the backward reference block is calculated for each pair of reference blocks, and from the positions of the M pairs of reference blocks, the positions of the pair of reference blocks with the pixel values having the minimum difference are determined as the positions of the target forward reference block and the target backward reference block of the current image block.
[0310] In another example, the matching cost criterion is a matching cost minimization and early termination criterion. For example, for the position of the nth pair of reference blocks (one forward reference block and one backward reference block), the difference between the pixel values of the forward reference block and the backward reference block is calculated, where n is an integer greater than or equal to 1 and less than or equal to N, and when the pixel value difference is less than or equal to the matching error threshold, the positions of the nth pair of reference blocks (one forward reference block and one backward reference block) are determined as the positions of the target forward reference block and the target backward reference block of the current image block.
[0311] In one implementation, the second acquisition unit 1901 is configured to acquire initial motion information from the candidate motion information list of the current image block or acquire initial motion information based on instruction information, where the instruction information is used to indicate the initial motion information of the current image block. It should be understood that the initial motion information is related to the refined motion information.
[0312] Device 1900 may be configured to execute the method shown in FIG. 12, and it should be understood that device 1900 may be a video encoding device, a video decoding device, a video encoding system, or another device having a video encoding function. Device 1900 may be configured not only to perform image prediction in the encoding process, but also to perform image prediction in the decoding process.
[0313] For details, please refer to the description of the image prediction method in this specification. For the sake of brevity, the details will not be described again here.
[0314] It can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form N pairs of reference block positions by the prediction device in this embodiment of the present application. For the positions of each pair of reference blocks among the positions of the N pairs of reference blocks, there is a proportional relationship based on the time-domain distance between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, the pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel values of the target forward reference block and the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity.
[0315] FIG. 20 is a schematic block diagram of another image prediction apparatus according to an embodiment of the present application. It should be noted that the prediction apparatus 2000 is applicable to both inter-frame prediction for decoding a video image and inter-frame prediction for encoding a video image. It should be understood that the prediction apparatus 2000 of the present specification may correspond to the motion compensation unit 44 in FIG. 2A or the motion compensation unit 82 in FIG. 2B. The prediction apparatus 2000 a third acquisition unit 2001 configured to acquire the i-th motion information of the current image block, a third search unit 2002 that determines the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in the forward reference image and the N backward reference blocks are arranged in the backward reference image, N is an integer greater than 1, and based on a matching cost criterion, from the positions of M pairs of reference blocks, it is determined that the positions of a pair of reference blocks are the positions of the i-th target forward reference block of the current image block and the positions of the i-th target backward reference block of the current image block, and the positions of each pair of reference blocks include the position of the forward reference block and the position of the backward reference block, and for the positions of each pair of reference blocks, the first position offset and the second position offset are in a mirror image relationship, the first position offset represents the offset of the position of the forward reference block with respect to the position of the (i - 1)-th target forward reference block, the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i - 1)-th target backward reference block, M is an integer greater than or equal to 1, and M is less than or equal to N, and a third search unit 2002 configured to perform the above operations, a third prediction unit 2003 configured to obtain a predicted value of the pixel value of the current image block based on the pixel values of the j-th target forward reference block and the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1, and may include.
[0316] If i = 1, the i-th motion information is the initial motion information of the current image block. Correspondingly, the positions of the N forward reference blocks include the position of one initial forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the initial forward reference block is an integer pixel distance or a fractional pixel distance, or the positions of the N backward reference blocks include the position of one initial backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the initial backward reference block is an integer pixel distance or a fractional pixel distance. It should be noted.
[0317] If i > 1, the i-th motion information includes a forward motion vector pointing to the position of the (i - 1)-th target forward reference block and a backward motion vector pointing to the position of the (i - 1)-th target backward reference block. Correspondingly, the positions of the N forward reference blocks include the position of one (i - 1)-th target forward reference block and the positions of (N - 1) candidate forward reference blocks, and the offset of the position of each candidate forward reference block with respect to the position of the (i - 1)-th target forward reference block is an integer pixel distance or a fractional pixel distance, or the positions of the N backward reference blocks include the position of one (i - 1)-th target backward reference block and the positions of (N - 1) candidate backward reference blocks, and the offset of the position of each candidate backward reference block with respect to the position of the (i - 1)-th target backward reference block is an integer pixel distance or a fractional pixel distance.
[0318] In this embodiment of the present application, the third prediction unit 2003 is specifically configured to obtain a predicted value of the pixel value of the image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block when the repetition end condition is satisfied, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1. For the description of the repetition end condition, refer to other embodiments. Details will not be described again here.
[0319] In the apparatus according to this embodiment of the present application, the fact that the first position offset and the second position offset are in a mirror image relationship can be understood as the first position offset value being the same as the second position offset value. For example, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset.
[0320] In one implementation form, the i-th motion information includes a forward motion vector, a forward reference image index, a backward motion vector, and a backward reference image index. In a manner of determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, the third search unit 2002 Based on the forward motion vector and the position of the current image block, determine the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index, and use the position of the (i - 1)-th target forward reference block as the starting point of the i-th f search, determine the positions of (N - 1) candidate forward reference blocks in the forward reference image, and the positions of the N forward reference blocks include the position of the (i - 1)-th target forward reference block and the positions of the (N - 1) candidate forward reference blocks. Based on the backward motion vector and the position of the current image block, determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index, and use the position of the (i - 1)-th target backward reference block as the starting point of the i-th b search, determine the positions of (N - 1) candidate backward reference blocks in the backward reference image, and the positions of the N backward reference blocks are specifically configured to include the position of the (i - 1)-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks.
[0321] In one implementation form, in a manner of determining, from the positions of M pairs of reference blocks based on a matching cost criterion, that the positions of a pair of reference blocks are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, the third search unit 2002 determines, from the positions of M pairs of reference blocks, whether the positions of a pair of reference blocks with the minimum matching error are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, or determines, from the positions of M pairs of reference blocks, that the positions of a pair of reference blocks with a matching error less than or equal to a matching error threshold are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, provided that M is less than or equal to N, and is particularly configured to do so.
[0322] The apparatus 2000 may be configured to execute the methods shown in FIGS. 14 and 15. It should be understood that the apparatus 2000 may particularly be a video encoding device, a video decoding device, a video encoding system, or another device having a video encoding function. The apparatus 2000 may be configured not only to perform image prediction in the encoding process but also to perform image prediction in the decoding process.
[0323] For details, please refer to the description of the image prediction method in this specification. For the sake of brevity, the details will not be described again here.
[0324] According to the prediction device in this embodiment of the present application, it can be seen that the positions of N forward reference blocks in the forward reference image and the positions of N backward reference blocks in the backward reference image form N pairs of reference block positions. For each pair of reference block positions among the N pairs of reference block positions, there is a mirror image relationship between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the position of the backward reference block with respect to the position of the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (i.e., the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (i.e., the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, the accuracy of refining the MV can be further improved by increasing the number of repetitions, whereby the coding performance can be further improved.
[0325] FIG. 21 is a schematic block diagram of another image prediction device according to an embodiment of the present application. It should be noted that the prediction device 2100 is applicable to both inter-frame prediction for decoding a video image and inter-frame prediction for encoding a video image. It should be understood that the prediction device 2100 in this specification may correspond to the motion compensation unit 44 in FIG. 2A or may correspond to the motion compensation unit 82 in FIG. 2B. The prediction device 2100 is a fourth acquisition unit 2101 configured to acquire the i-th motion information of the current image block, and A fourth search unit 2102 that determines the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, where the N forward reference blocks are arranged in a forward reference image, the N backward reference blocks are arranged in a backward reference image, N is an integer greater than 1, and based on a matching cost criterion, from the positions of M pairs of reference blocks, it is determined that the positions of a pair of reference blocks are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block. The position of each pair of reference blocks includes the position of the forward reference block and the position of the backward reference block. For the position of each pair of reference blocks, the first position offset and the second position offset are in a proportional relationship based on a time-domain distance. The first position offset represents the offset of the position of the forward reference block with respect to the position of the (i - 1)-th target forward reference block in the forward reference image, and the second position offset represents the offset of the position of the backward reference block with respect to the position of the (i - 1)-th target backward reference block in the backward reference image. M is an integer greater than or equal to 1, and M is less than or equal to N. The fourth search unit 2102 is configured to perform the above operations. A fourth prediction unit 2103 that obtains a predicted value of the pixel value of the current image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1.
[0326] In the iterative search process, if i = 1, the i-th motion information is the initial motion information of the current image block.
[0327] If i > 1, the i-th motion information includes a forward motion vector pointing to the position of the (i - 1)-th target forward reference block and a backward motion vector pointing to the position of the (i - 1)-th target backward reference block.
[0328] In one implementation form, when the repetition end condition is satisfied, the fourth prediction unit 2103 obtains a predicted value of the pixel value of the image block based on the pixel value of the j-th target forward reference block and the pixel value of the j-th target backward reference block, where j is greater than or equal to i, and both i and j are integers greater than or equal to 1, and is particularly configured as such.
[0329] In the apparatus of this embodiment, the fact that the first position offset and the second position offset have a proportional relationship based on the time domain distance means that if the first time domain distance is the same as the second time domain distance, the direction of the first position offset is opposite to the direction of the second position offset, and the amplitude value of the first position offset is the same as the amplitude value of the second position offset, or if the first time domain distance is different from the second time domain distance, it can be understood that the direction of the first position offset is opposite to the direction of the second position offset, and the proportional relationship between the amplitude value of the first position offset and the amplitude value of the second position offset is based on the proportional relationship between the first time domain distance and the second time domain distance.
[0330] The first time domain distance represents the time domain distance between the current image to which the current image block belongs and the forward reference image, and the second time domain distance represents the time domain distance between the current image and the backward reference image.
[0331] In one implementation form, the i-th motion information includes a forward motion vector, a forward reference image index, a backward motion vector, and a backward reference image index. Correspondingly, in a manner of determining the positions of N forward reference blocks and the positions of N backward reference blocks based on the i-th motion information and the position of the current image block, the fourth search unit 2102 determines the position of the (i - 1)-th target forward reference block of the current image block in the forward reference image corresponding to the forward reference image index based on the forward motion vector and the position of the current image block, and sets the position of the (i - 1)-th target forward reference block as the i fIt is used as a search start point to determine the positions of (N - 1) candidate forward reference blocks in the forward reference image. The positions of the N forward reference blocks include the position of the (i - 1)-th target forward reference block and the positions of the (N - 1) candidate forward reference blocks. Based on the backward motion vector and the position of the current image block, determine the position of the (i - 1)-th target backward reference block of the current image block in the backward reference image corresponding to the backward reference image index. Use the position of the (i - 1)-th target backward reference block as the b search start point to determine the positions of (N - 1) candidate backward reference blocks in the backward reference image. The positions of the N backward reference blocks are particularly configured to include the position of the (i - 1)-th target backward reference block and the positions of the (N - 1) candidate backward reference blocks.
[0332] In one implementation, in the aspect of determining from the positions of M pairs of reference blocks based on the matching cost criterion that the positions of a pair of reference blocks are the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, the fourth search unit 2102 determines from the positions of the M pairs of reference blocks whether the position of the pair of reference blocks with the minimum matching error is the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, or determines from the positions of the M pairs of reference blocks that the position of the pair of reference blocks with a matching error less than or equal to the matching error threshold is the position of the i-th target forward reference block of the current image block and the position of the i-th target backward reference block of the current image block, provided that M is less than or equal to N. It is particularly configured in this way.
[0333] The device 2100 may be configured to execute the method shown in FIG. 16 or FIG. 17, and it should be understood that the device 2100 may be a video encoding device, a video decoding device, a video encoding system, or another device having a video encoding function. The device 2100 may be configured not only to perform image prediction in the encoding process, but also to perform image prediction in the decoding process.
[0334] For details, please refer to the description of the image prediction method in this specification. For the sake of brevity, the details are not described again here.
[0335] According to the prediction device in this embodiment of the present application, it can be seen that the positions of the N forward reference blocks in the forward reference image and the positions of the N backward reference blocks in the backward reference image form the positions of N pairs of reference blocks. For the positions of each pair of reference blocks among the positions of the N pairs of reference blocks, there is a proportional relationship based on the time-domain distance between the first position offset of the forward reference block with respect to the initial forward reference block and the second position offset of the backward reference block with respect to the initial backward reference block. Based on such a situation, the positions of a pair of reference blocks (for example, a pair of reference blocks with the minimum matching cost) are determined from the positions of the N pairs of reference blocks as the position of the target forward reference block (that is, the optimal forward reference block / forward prediction block) of the current image block and the position of the target backward reference block (that is, the optimal backward reference block / backward prediction block) of the current image block. Thereby, a predicted value of the pixel value of the current image block is obtained based on the pixel value of the target forward reference block and the pixel value of the target backward reference block. Compared with the prior art, the method in this embodiment of the present application avoids the process of pre-calculating the template matching block and the process of performing forward search matching and backward search matching by using the template matching block, and simplifies the image prediction process. This improves the image prediction accuracy and reduces the image prediction complexity. Furthermore, the accuracy of refining the MV can be further improved by increasing the number of repetitions, whereby the coding performance can be further improved.
[0336] FIG. 22 is a schematic block diagram of an implementation of a video encoding device or a video decoding device (abbreviated as decoding device 2200) according to an embodiment of the present application. The decoding device 2200 may include a processor 2210, a memory 2230, and a bus system 2250. The processor and the memory are connected by using the bus system. The memory is configured to store instructions. The processor is configured to execute the instructions stored in the memory. The memory of the encoding device stores program codes. The processor calls the program codes stored in the memory to execute the video encoding or decoding method described in the present application, particularly, the video encoding or decoding method of various inter-frame prediction modes or intra-frame prediction modes, and the motion information prediction method of various inter-frame prediction modes or intra-frame prediction modes. Details will not be described again in this specification to avoid repetition.
[0337] In this embodiment of the present application, the processor 2210 may be a central processing unit (CPU), or the processor 2210 may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or another programmable logic device, discrete gate or transistor logic device, discrete hardware component, or the like. The general-purpose processor may be a microprocessor or any conventional processor or the like.
[0338] The memory 2230 may include a read-only memory (ROM) device or a random access memory (RAM). Any other suitable type of storage device may also be used as the memory 2230. The memory 2230 is connected to the bus systemIt may include code and data 2231 accessed by the processor 2210 by using 2250. The memory 2230 may further include an operating system 2233 and an application program 2235. The application program 2235 includes at least one program that enables the processor 2210 to execute the video encoding or decoding method described in the present application (in particular, the image prediction method described in the present application). For example, the application program 2235 may include applications 1 to N, and further includes a video encoding or decoding application (abbreviated as a video decoding application) that executes the video encoding or decoding method described in the present application.
[0339] In addition to the data bus, the bus system 2250 may further include a power bus, a control bus, a status signal bus, and the like. However, for clarity of explanation, various types of buses in the figure are marked as the bus system 2250.
[0340] Optionally, the decoding device 2200 may further include one or more output devices, for example, a display 2270. In one example, the display 2270 may be a touch display or a touch screen that combines a display and a touch unit that operably senses touch input. The display 2270 may be connected to the processor 2210 by using the bus 2250.
[0341] Note that the description and limitation of the same step or the same term are also applicable to different embodiments. For the sake of brevity, repeated descriptions in this specification are omitted as appropriate.
[0342] One of ordinary skill in the art will appreciate that the functions described herein with reference to various exemplary logical blocks, modules, and algorithm steps can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions described with reference to exemplary logical blocks, modules, and steps can be stored or transmitted on a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium that includes any medium that facilitates transfer of a computer program from one place to another (e.g., according to a communication protocol). In this manner, the computer-readable medium generally can correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium can be any medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this application. A computer program product can include a computer-readable medium.
[0343] For example, such a computer-readable storage medium can be RAM, ROM, EEPROM, CD-ROM or another compact disc storage device, magnetic disk storage device or another magnetic storage device, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these. Further, any connection can also be referred to as a computer-readable medium. For example, when instructions are transmitted from a website, server, or another remote source through coaxial cable, fiber optic, twisted pair wire, digital subscriber line (DSL), or wireless technologies such as infrared, radio waves, and microwaves, the coaxial cable, fiber optic cable, twisted pair wire, DSL, or wireless technologies such as infrared, radio waves, and microwaves are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, and actually mean non-transient tangible storage media. As used herein, "Disk" and "Disc" (both of which are "disk" in Japanese) include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), and Blu-ray disc. A disk usually reproduces data magnetically, while a disc reproduces data optically using a laser. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0344] The corresponding functionality can be performed by one or more processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated circuits or discrete logic circuits. Thus, the term "processor" as used herein may be any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Further, in some aspects, the functionality described with reference to the exemplary logic blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured to encode, or incorporated in a combined codec. Further, the techniques may be implemented entirely in one or more circuits or logic elements. In one example, the various exemplary logic blocks, units, and modules in the video encoder 20 and the video decoder 30 can be understood as corresponding circuit devices or logic elements.
[0345] The techniques in this application can be implemented in various apparatus or devices, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., a chipset). The various components, modules, or units are described in this application to emphasize the functional aspects of an apparatus configured to execute the disclosed techniques, but are not necessarily implemented by different hardware units. In fact, as described above, the various units may be integrated within a codec hardware unit or provided by interoperable hardware units (including one or more of the processors described above) in combination with appropriate software and / or firmware.
[0346] The foregoing description is merely an example of a specific implementation form of this application and is not intended to limit the protection scope of this application. Modifications or alternative forms that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application shall be deemed to be within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Explanation of Reference Signs
[0347] 12 Source Device 14 Destination Device 16 Link 18 Video Source 20 Video Encoder 22 Output Interface 28 Input Interface 30 Video Decoder 32 Display Device 41 Prediction Module 42 Motion Estimation Unit 44 Motion Compensation Unit 46 Intra-Frame Prediction Unit 50 Adder 52 Conversion Module 54 Quantization Module 56 Entropy Encoding Module 58 Inverse Quantization Module 60 Inverse Conversion Module 62 Adder 64 Reference Image Memory 80 Entropy Decoding Module 81 Prediction Processing Module 82 Motion Compensation Unit 84 Intra-Frame Prediction Unit 86 Inverse Quantization Module 88 Inverse Conversion Module 90 Reconstruction Module, Adder 92 Reference Image Memory, Decoded Image Buffer 902 Initial Forward Reference Block 903 Initial Backward Reference Block 904 Candidate Forward Reference Block 905 Candidate backward reference block 1302 Initial forward reference block 1303 Initial backward reference block 1304 Candidate forward reference block 1305 Candidate backward reference block 1800 Predictor 1801 First acquisition unit 1802 First search unit 1803 First prediction unit 1900 Predictor 1901 Second acquisition unit 1902 Second search unit 1903 Second prediction unit 2000 Predictor 2001 Third acquisition unit 2002 Third search unit 2003 Third prediction unit 2100 Predictor 2101 Fourth acquisition unit 2102 Fourth search unit 2103 Fourth prediction unit 2200 Decoding device 2210 Processor 2230 Memory 2231 Code and data 2233 Operating system 2235 Application program 2250 Bus 2270 Display
Claims
1. An image prediction method, comprising: obtaining initial motion information of a current image block; when an early termination condition is satisfied, directly determining the positions of an initial forward reference block and an initial backward reference block of the current image block as the positions of a target forward reference block and a target backward reference block of the current image block, calculating a difference between a pixel value of the initial forward reference block and a pixel value of the initial backward reference block, and when the difference between the pixel value of the initial forward reference block and the pixel value of the initial backward reference block is smaller than a matching error threshold, directly determining the positions of the initial forward reference block and the initial backward reference block of the current image block as the positions of the target forward reference block and the target backward reference block of the current image block, wherein the positions of the initial forward and backward reference blocks of the current image block are based on the initial motion information of the current image block; obtaining a predicted value of a pixel value of the current image block based on pixel values of the target forward reference block and the target backward reference block of the current image block; An image prediction method comprising the above steps.
2. The method according to claim 1, wherein the pixel value of the target forward reference block is determined based on the position of the target forward reference block, or the pixel value of the target backward reference block is determined based on the position of the target backward reference block.
3. The initial motion information includes a first motion vector and a first reference image index corresponding to a first list (L0), and a second motion vector and a second reference image index corresponding to a second list (L1), wherein the position of the initial forward reference block in a forward reference image corresponding to the first reference image index is based on the first motion vector and the position of the current image block, and the position of the initial backward reference block in a backward reference image corresponding to the second reference image index is based on the second motion vector and the position of the current image block. The method according to any one of claims 1 to 2.
4. The step of obtaining initial motion information of a current image block includes: obtaining the initial motion information from a candidate motion information list of the current image block; The method further includes a step of encoding instruction information into a bit stream, where the instruction information indicates the initial motion information in the candidate motion information list of the current image block. The method according to any one of claims 1 to 2.
5. Before the step of obtaining initial motion information of a current image block, The method further includes a step of obtaining instruction information from a bit stream of the current image block, where the instruction information indicates the initial motion information of the current image block. The method according to any one of claims 1 to 2.
6. A memory storage including instructions, and one or more processors communicating with the memory storage, An image prediction apparatus, wherein the one or more processors execute the instructions to obtain initial motion information of a current image block, when an early termination condition is satisfied, directly determine the positions of an initial forward reference block and an initial backward reference block of the current image block as the positions of a target forward reference block and a target backward reference block of the current image block, calculate a difference between a pixel value of the initial forward reference block and a pixel value of the initial backward reference block, and when the difference between the pixel value of the initial forward reference block and the pixel value of the initial backward reference block is smaller than a matching error threshold, directly determine the positions of the initial forward reference block and the initial backward reference block of the current image block as the positions of the target forward reference block and the target backward reference block of the current image block, and the positions of the initial forward and backward reference blocks of the current image block are based on the initial motion information of the current image block; configured to obtain a predicted value of a pixel value of the current image block based on pixel values of the target forward reference block and the target backward reference block of the current image block. An image prediction apparatus.
7. The pixel value of the target forward reference block is determined based on the position of the target forward reference block, or the pixel value of the target backward reference block is determined based on the position of the target backward reference block, the apparatus according to claim 6.
8. The initial motion information includes a first motion vector and a first reference picture index corresponding to a first list (L0), and a second motion vector and a second reference picture index corresponding to a second list (L1), The position of the initial forward reference block in the forward reference picture corresponding to the first reference picture index is based on the first motion vector and the position of the current picture block, The position of the initial backward reference block in the backward reference picture corresponding to the second reference picture index is based on the second motion vector and the position of the current picture block, the apparatus according to any one of claims 6 to 7.
9. The apparatus is an encoding apparatus for encoding the current block, and the one or more processors execute the instructions to obtain the initial motion information from a candidate motion information list of the current picture block, The one or more processors further execute the instructions, Encode the indication information into a bitstream, the indication information indicating the initial motion information in the candidate motion information list of the current picture block, The apparatus according to any one of claims 6 to 7.
10. The apparatus is a decoding apparatus for decoding the current block, and the one or more processors execute the instructions to Obtain indication information from the bitstream of the current picture block, the indication information indicating the initial motion information of the current picture block, the apparatus according to any one of claims 6 to 7.
11. A non-transitory computer-readable medium storing program code, the program code, when executed by a computer device, causes the computer device to Obtain initial motion information of a current picture block, and When an early termination condition is satisfied, determining the positions of the initial forward reference block and the initial backward reference block of the current image block as the positions of the target forward reference block and the target backward reference block of the current image block as they are, calculating a difference between the pixel value of the initial forward reference block and the pixel value of the initial backward reference block, and when the difference between the pixel value of the initial forward reference block and the pixel value of the initial backward reference block is smaller than a matching error threshold, determining the positions of the initial forward reference block and the initial backward reference block of the current image block as the positions of the target forward reference block and the target backward reference block of the current image block as they are, and the positions of the initial forward and backward reference blocks of the current image block are based on the initial motion information of the current image block, the step; Obtaining a predicted value of the pixel value of the current image block based on the pixel values of the target forward reference block and the target backward reference block of the current image block; Causing a method including the above to be performed; A non-transitory computer-readable medium.
12. The non-transitory computer-readable medium according to claim 11, wherein the pixel value of the target forward reference block is determined based on the position of the target forward reference block, or the pixel value of the target backward reference block is determined based on the position of the target backward reference block.
13. The initial motion information includes a first motion vector and a first reference image index corresponding to a first list (L0), and a second motion vector and a second reference image index corresponding to a second list (L1). The position of the initial forward reference block in the forward reference image corresponding to the first reference image index is based on the first motion vector and the position of the current image block. The non-transitory computer-readable medium according to any one of claims 11 to 12, wherein the position of the initial backward reference block in the backward reference image corresponding to the second reference image index is based on the second motion vector and the position of the current image block.
14. The step of obtaining initial motion information of a current image block includes: obtaining the initial motion information from a candidate motion information list of the current image block; The method further includes a step of encoding instruction information into a bit stream, where the instruction information indicates the initial motion information in the candidate motion information list of the current image block. The non-transitory computer-readable medium according to any one of claims 11 to 12. **Claim 15** Before the step of obtaining initial motion information of a current image block, The method further includes a step of obtaining instruction information from a bit stream of the current image block, where the instruction information indicates the initial motion information of the current image block. The non-transitory computer-readable medium according to any one of claims 11 to 12.
Citation Information
Patent Citations
Moving vector detecting device
JP2001145109A
Motion information derivation mode determination in video coding
WO2016160609A1