Inter prediction method and apparatus, and corresponding encoder and decoder
By optimizing the determination of motion vector predictors and differences using candidate length and direction information, the method addresses redundancy in inter prediction, enhancing coding efficiency and video compression performance.
Patent Information
- Application Number
- JP2025043063
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-29
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2039-12-25
AI Technical Summary
Existing video coding technologies face challenges in reducing redundancy and improving coding efficiency in inter prediction processes, particularly in determining motion vector predictors and differences for image blocks.
The method involves obtaining a motion vector predictor and determining a motion vector difference based on a set of candidate length and direction information, using specific index values to reduce redundancy and enhance coding efficiency.
This approach reduces redundancy and enhances coding efficiency by optimizing the determination of motion vector predictors and differences, leading to improved video compression and decoding performance.
Smart Images

Figure 2025094073000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding technology, and in particular, to a method and apparatus for inter prediction, and corresponding encoders and decoders.
Background Art
[0002] Digital video capabilities can be incorporated into a wide variety of devices including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, mobile phones or satellite radio phones (also called "smartphones"), video conferencing devices, video streaming devices, and the like. Digital video devices implement video compression technologies described in standards including, for example, MPEG-2, MPEG-4, ITU-T H.263, and ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), video coding standard H.265 / High Efficiency Video Coding (HEVC) standard, and extended versions of these standards. Video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently by implementing video compression technologies.
[0003] Video compression techniques are used to perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove inherent redundancies in a video sequence. In block-based video coding, a video slice (i.e., a video frame or a part of a video frame) can be partitioned into image blocks, which may also be referred to as tree blocks, coding units (CUs), and / or coding nodes. Image blocks within an intra-coded (I) slice of an image are coded by spatial prediction based on reference samples in neighboring blocks within the same image. In the case of image blocks within an inter-coded (P or B) slice of an image, either spatial prediction based on reference samples in neighboring blocks within the same image or temporal prediction based on reference samples in another reference image may be used. An image may sometimes be referred to as a frame, and a reference image may sometimes be referred to as a reference frame. SUMMARY OF THE INVENTION
[0004] Embodiments of the present application provide an inter-prediction method and apparatus, and corresponding encoder and decoder, for reducing redundancy to a certain extent in a coding process and improving coding efficiency.
[0005] According to a first aspect, an embodiment of the present application provides an inter-prediction method. The method includes obtaining a motion vector predictor of a current image block; obtaining an index value of a length of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block; determining target length information from a set of candidate length information based on the index value of the length, where the set of candidate length information includes candidate length information of only N motion vector differences, and N is a positive integer greater than 1 and less than 8; obtaining the motion vector difference of the current image block based on the target length information; determining the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; and obtaining a prediction block of the current image block based on the motion vector target value of the current image block.
[0006] The set of candidate length information may be preset.
[0007] In a first possible implementation related to the first aspect, the method further includes obtaining an index value of a direction of the motion vector difference of the current image block; and determining target direction information from M pieces of candidate direction information of the motion vector difference based on the index value of the direction, where M is a positive integer greater than 1, and the step of obtaining the motion vector difference of the current image block based on the target length information includes determining the motion vector difference of the current image block based on the target direction information and the target length information.
[0008] In a second possible implementation related to the first aspect or the first possible implementation of the first aspect, N is 4.
[0009] Regarding a third possible implementation of the first aspect, in relation to a second possible implementation of the first aspect, the candidate length information of the N motion vector differences includes at least one of the following: when the length index value is a first preset value, the length indicated by the target length information is 1 / 4 of the pixel length; when the length index value is a second preset value, the length indicated by the target length information is half of the pixel length; when the length index value is a third preset value, the length indicated by the target length information is 1 pixel length; or when the length index value is a fourth preset value, the length indicated by the target length information is 2 pixel lengths. Regarding either the first aspect or any one of the aforementioned possible implementations of the first aspect, in a fourth possible implementation of the first aspect, the step of obtaining the motion vector predictor of the current image block includes: constructing a candidate motion information list of the current image block, where the candidate motion information list includes L motion vectors and L is 1, 3, 4, or 5; obtaining the index value of the prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes the motion vector predictor; and obtaining the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list.
[0010] According to a second aspect, an embodiment of the present application provides an inter-prediction method. The method includes obtaining motion vector predictors for a current image block; performing a motion search in a region at a position indicated by the motion vector predictors of the current image block to obtain a motion vector target value for the current image block; and obtaining an index value of a length of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictors of the current image block and the motion vector target value, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information, and the set of candidate length information includes only candidate length information for N motion vector differences, and N is a positive integer greater than 1 and less than 8.
[0011] In connection with the second aspect, in a first possible implementation of the second aspect, the step of obtaining an index value of a length of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block includes obtaining a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block; and determining, based on the motion vector difference of the current image block, an index value of the length of the motion vector difference of the current image block and an index value of a direction of the motion vector difference of the current image block.
[0012] In connection with the second aspect or the first possible implementation of the second aspect, in a second possible implementation of the second aspect, N is 4.
[0013] According to a third aspect, an embodiment of the present application provides an inter-prediction method. The method includes obtaining motion vector predictors of a current image block; obtaining an index value of a direction of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block; determining target direction information from a set of candidate direction information based on the index value of the direction, where the set of candidate direction information includes candidate direction information of M motion vector differences and M is a positive integer greater than 4; obtaining the motion vector difference of the current image block based on the target direction information; determining the motion vector target value of the current image block based on the motion vector difference of the current block and the motion vector predictor of the current image block; and obtaining a predicted block of the current image block based on the motion vector target value of the current image block.
[0014] In relation to the third aspect, in a first possible implementation of the third aspect, the method further includes obtaining an index value of a length of the motion vector difference of the current image block; and determining target length information from N candidate length information of the motion vector differences based on the index value of the length, where N is a positive integer greater than 1, and the step of obtaining the motion vector difference of the current image block based on the target direction information includes determining the motion vector difference of the current image block based on the target direction information and the target length information.
[0015] In relation to the third aspect or the first possible implementation of the third aspect, in a second possible implementation of the third aspect, M is 8.
[0016] Regarding a second possible implementation of the third aspect, in a third possible implementation of the third aspect, the candidate direction information of the M motion vector differences is such that when the index value of the direction is a first preset value, the direction indicated by the target direction information is exactly right; when the index value of the direction is a second preset value, the direction indicated by the target direction information is exactly left; when the index value of the direction is a third preset value, the direction indicated by the target direction information is exactly down; when the index value of the direction is a fourth preset value, the direction indicated by the target direction information is exactly up; when the index value of the direction is a fifth preset value, the direction indicated by the target direction information is bottom right; when the index value of the direction is a sixth preset value, the direction indicated by the target direction information is top right; when the index value of the direction is a seventh preset value, the direction indicated by the target direction information is bottom left; or when the index value of the direction is an eighth preset value, the direction indicated by the target direction information is top left, and includes at least one of them.
[0017] Regarding any one of the third aspect or the aforementioned possible implementations of the third aspect, in a fourth possible implementation of the third aspect, the step of obtaining the motion vector predictor of the current image block includes the step of constructing a candidate motion information list of the current image block, where the candidate motion information list includes L motion vectors, and L is 1, 3, 4, or 5; the step of obtaining the index value of the prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes the motion vector predictor; and the step of obtaining the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list.
[0018] According to the fourth aspect, an embodiment of the present application provides an inter-prediction method. The method includes obtaining motion vector predictors of a current image block, performing a motion search in a region at a position indicated by the motion vector predictors of the current image block to obtain a motion vector target value of the current image block, and obtaining an index value of a direction of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictors of the current image block and the motion vector target value, and the index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information, and the set of candidate length information includes M pieces of candidate direction information of motion vector differences, and M is a positive integer greater than 4.
[0019] In relation to the fourth aspect, in a first possible implementation of the fourth aspect, the step of obtaining an index value of a direction of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block includes obtaining a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block, and determining an index value of a length of the motion vector difference of the current image block and an index value of a direction of the motion vector difference of the current image block based on the motion vector difference of the current image block.
[0020] In relation to the fourth aspect or the first possible implementation of the fourth aspect, in a second possible implementation of the fourth aspect, M is 8.
[0021] According to a fifth aspect, an embodiment of the present application provides an inter prediction method. The method includes obtaining a first motion vector predictor of a current image block and a second motion vector predictor of the current image block, where the first motion vector predictor corresponds to a first reference frame and the second motion vector predictor corresponds to a second reference frame; obtaining a first motion vector difference of the current image block, where the first motion vector difference of the current image block is used to indicate a difference between the first motion vector predictor and a first motion vector target value of the current image block, and the first motion vector target value and the first motion vector predictor correspond to the same reference frame; determining a second motion vector difference of the current image block based on the first motion vector difference, where the second motion vector difference of the current image block is used to indicate a difference between the second motion vector predictor and a second motion vector target value of the current image block, the second motion vector target value and the second motion vector predictor correspond to the same reference frame, and when a direction of the first reference frame with respect to a current frame in which the current image block is located is the same as a direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference, or when the direction of the first reference frame with respect to the current frame in which the current image block is located is opposite to the direction of the second reference frame with respect to the current frame, a plus sign or a minus sign of the second motion vector difference is opposite to a plus sign or a minus sign of the first motion vector difference, and an absolute value of the second motion vector difference is the same as an absolute value of the first motion vector difference; determining the first motion vector target value of the current image block based on the first motion vector difference and the first motion vector predictor; determining the second motion vector target value of the current image block based on the second motion vector difference and the second motion vector predictor; and obtaining a predicted block of the current image block based on the first motion vector target value and the second motion vector target value.
[0022] According to the sixth aspect, an embodiment of the present application provides an inter-prediction device. The device includes a prediction unit configured to obtain a motion vector predictor of a current image block, and an acquisition unit configured to obtain an index value of a length of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block. The prediction unit is further configured to determine target length information from a set of candidate length information based on the index value of the length, where the set of candidate length information includes candidate length information of only N motion vector differences, and N is a positive integer greater than 1 and less than 8. The prediction unit is further configured to obtain the motion vector difference of the current image block based on the target length information, determine the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block, and obtain a prediction block of the current image block based on the motion vector target value of the current image block.
[0023] In relation to the sixth aspect, in a first conceivable implementation of the sixth aspect, the acquisition unit is further configured to obtain an index value of a direction of the motion vector difference of the current image block, and the prediction unit is further configured to determine target direction information from candidate direction information of M motion vector differences based on the index value of the direction, where M is a positive integer greater than 1. The prediction unit is configured to determine the motion vector difference of the current image block based on the target direction information and the target length information.
[0024] In relation to the sixth aspect or the first conceivable implementation of the sixth aspect, in a second conceivable implementation of the sixth aspect, N is 4.
[0025] Regarding a second possible implementation of the sixth aspect, in a third possible implementation of the sixth aspect, the candidate length information of the N motion vector differences includes at least one of the following: when the index value of the length is a first preset value, the length indicated by the target length information is 1 / 4 of the pixel length; when the index value of the length is a second preset value, the length indicated by the target length information is half of the pixel length; when the index value of the length is a third preset value, the length indicated by the target length information is 1 pixel length; or when the index value of the length is a fourth preset value, the length indicated by the target length information is 2 pixel lengths.
[0026] Regarding either the sixth aspect or any one of the aforementioned possible implementations of the sixth aspect, in a fourth possible implementation of the sixth aspect, the prediction unit is configured to: construct a candidate motion information list for the current image block, where the candidate motion information list may include L motion vectors and L is 1, 3, 4, or 5; obtain an index value of prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes a motion vector predictor; and obtain the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list.
[0027] According to the seventh aspect, an embodiment of the present application provides an inter-prediction device. The device includes an acquisition unit configured to acquire motion vector predictors of a current image block, and a prediction unit configured to perform a motion search in a region at a position indicated by the motion vector predictors of the current image block to obtain a motion vector target value of the current image block. The prediction unit is further configured to obtain an index value of the length of the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block, where the motion vector difference of the current image block is used to indicate the difference between the motion vector predictors of the current image block and the motion vector target value, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information. The set of candidate length information includes only candidate length information for N motion vector differences, and N is a positive integer greater than 1 and less than 8, and is configured to perform the acquisition.
[0028] In relation to the seventh aspect, in a first conceivable implementation of the seventh aspect, the prediction unit is configured to obtain the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictors of the current image block, and based on the motion vector difference of the current image block, determine an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block.
[0029] In relation to the seventh aspect or the first conceivable implementation of the seventh aspect, in a second conceivable implementation of the seventh aspect, N is 4.
[0030] According to the eighth aspect, an embodiment of the present application provides an inter-prediction device. The device includes a prediction unit configured to obtain a motion vector predictor of a current image block, and an acquisition unit configured to obtain an index value of a direction of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block. The prediction unit is further configured to determine target direction information from a set of candidate direction information based on the index value of the direction, where the set of candidate direction information includes candidate direction information of M motion vector differences, and M is a positive integer greater than 4. The prediction unit is further configured to obtain the motion vector difference of the current image block based on the target direction information, determine the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block, and obtain a prediction block of the current image block based on the motion vector target value of the current image block.
[0031] In relation to the eighth aspect, in a first conceivable implementation of the eighth aspect, the acquisition unit is further configured to obtain an index value of a length of the motion vector difference of the current image block, and the prediction unit is further configured to determine target length information from N candidate length information of motion vector differences based on the index value of the length, where N is a positive integer greater than 1. The prediction unit is configured to determine the motion vector difference of the current image block based on the target direction information and the target length information.
[0032] In relation to the eighth aspect or the first conceivable implementation of the eighth aspect, in a second conceivable implementation of the eighth aspect, M is 8.
[0033] Regarding a third possible implementation of the eighth aspect, in relation to a second possible implementation of the eighth aspect, the candidate direction information of the M motion vector differences may include at least one of the following: when the index value of the direction is a first preset value, the direction indicated by the target direction information is exactly right; when the index value of the direction is a second preset value, the direction indicated by the target direction information is exactly left; when the index value of the direction is a third preset value, the direction indicated by the target direction information is exactly down; when the index value of the direction is a fourth preset value, the direction indicated by the target direction information is exactly up; when the index value of the direction is a fifth preset value, the direction indicated by the target direction information is bottom - right; when the index value of the direction is a sixth preset value, the direction indicated by the target direction information is top - right; when the index value of the direction is a seventh preset value, the direction indicated by the target direction information is bottom - left; or when the index value of the direction is an eighth preset value, the direction indicated by the target direction information is top - left.
[0034] Regarding either the eighth aspect or any one of the aforementioned possible implementations of the eighth aspect, in a fourth possible implementation of the eighth aspect, the prediction unit is configured to: construct a candidate motion information list for the current image block, where the candidate motion information list includes L motion vectors and L is 1, 3, 4, or 5; obtain the index value of the prediction information of the motion information of the current image block within the candidate motion information list, where the prediction information of the motion information of the current image block includes a motion vector predictor; and obtain the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block within the candidate motion information list.
[0035] According to a ninth aspect, an embodiment of the present application provides an inter-prediction device. The device includes an acquisition unit configured to acquire a motion vector predictor of a current image block, and a prediction unit configured to perform a motion search in a region at a position indicated by the motion vector predictor of the current image block to obtain a motion vector target value of the current image block. The prediction unit is further configured to obtain an index value in the direction of the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block, where the motion vector difference of the current image block is used to indicate the difference between the motion vector predictor of the current image block and the motion vector target value, and the index value in the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information. The set of candidate direction information includes M pieces of candidate length information of the motion vector difference, and M is a positive integer greater than 4, and is configured to perform the acquisition.
[0036] In relation to the ninth aspect, in a first possible implementation of the ninth aspect, the prediction unit is configured to obtain the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block, and determine, based on the motion vector difference of the current image block, an index value of the length of the motion vector difference of the current image block and an index value in the direction of the motion vector difference of the current image block.
[0037] In relation to the ninth aspect or the first possible implementation of the ninth aspect, in a second possible implementation of the ninth aspect, M is 8.
[0038] According to the tenth aspect, an embodiment of the present application provides an inter prediction device. The device includes an acquisition unit and a prediction unit. The acquisition unit is configured to acquire a first motion vector predictor of a current image block and a second motion vector predictor of the current image block. The first motion vector predictor corresponds to a first reference frame, and the second motion vector predictor corresponds to a second reference frame. The acquisition unit is further configured to acquire a first motion vector difference of the current image block, where the first motion vector difference of the current image block is used to indicate the difference between the first motion vector predictor and a first motion vector target value of the current image block. The first motion vector target value and the first motion vector predictor correspond to the same reference frame. The prediction unit is configured to determine a second motion vector difference of the current image block based on the first motion vector difference, where the second motion vector difference of the current image block is used to indicate the difference between the second motion vector predictor and a second motion vector target value of the current image block. The second motion vector target value and the second motion vector predictor correspond to the same reference frame. When the direction of the first reference frame with respect to the current frame in which the current image block is arranged is the same as the direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference. When the direction of the first reference frame with respect to the current frame in which the current image block is arranged is opposite to the direction of the second reference frame with respect to the current frame, the plus or minus sign of the second motion vector difference is opposite to the plus or minus sign of the first motion vector difference, and the absolute value of the second motion vector difference is the same as the absolute value of the first motion vector difference. The prediction unit is further configured to determine the first motion vector target value of the current image block based on the first motion vector difference and the first motion vector predictor, determine the second motion vector target value of the current image block based on the second motion vector difference and the second motion vector predictor, and acquire a predicted block of the current image block based on the first motion vector target value and the second motion vector target value.
[0039] According to the 11th aspect, an embodiment of the present application provides a video decoder. The video decoder is configured to decode a bitstream to obtain an image block, and is an inter prediction device according to any one of the 1st aspect or a conceivable implementation of the 1st aspect, and is configured to obtain a prediction block of a current image block. An inter prediction device and a reconstruction module configured to reconstruct the current image block based on the prediction block.
[0040] According to the 12th aspect, an embodiment of the present application provides a video encoder. The video encoder is configured to encode an image block, and is an inter prediction device according to any one of the 2nd aspect or a conceivable implementation of the 2nd aspect. The inter prediction device is configured to obtain an index value of the length of the motion vector difference of the current image block based on the motion vector predictor of the current image block. The index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information. An inter prediction device and an entropy encoding module configured to encode the index value of the length of the motion vector difference of the current image block into a bitstream.
[0041] According to the 13th aspect, an embodiment of the present application provides a video decoder. The video decoder is configured to decode a bitstream to obtain an image block, and is an inter prediction device according to any one of the 3rd aspect or a conceivable implementation of the 3rd aspect, and is configured to obtain a prediction block of a current image block. An inter prediction device and a reconstruction module configured to reconstruct the current image block based on the prediction block.
[0042] According to a 14th aspect, an embodiment of the present application provides a video encoder. The video encoder is configured to encode an image block and is an inter prediction device according to any one of the 4th aspect or a conceivable implementation of the 4th aspect. The inter prediction device is configured to obtain an index value in the direction of a motion vector difference of a current image block based on a motion vector predictor of the current image block. The index value in the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information. The inter prediction device and an entropy encoding module configured to encode the index value in the direction of the motion vector difference of the current image block into a bitstream are included.
[0043] According to a 15th aspect, an embodiment of the present application provides a video data decoding device. The device includes a memory configured to store video data in the form of a bitstream and a video decoder according to any one of the 11th aspect, the 13th aspect, the 15th aspect, or any one of the 11th aspect, the 13th aspect, or the 15th aspect.
[0044] According to a 16th aspect, an embodiment of the present application provides a video data encoding device. The device includes a memory configured to store video data, where the video data includes one or more image blocks, and a video encoder according to any one of the 12th aspect, the 14th aspect, or any one of the implementations of the 12th aspect and the 14th aspect.
[0045] According to a 17th aspect, an embodiment of the present application provides an encoding device. The encoding device includes a non-volatile memory and a processor coupled to each other. The processor calls program code stored in the memory to execute some or all of the steps in any method according to any one of the 2nd aspect, the 4th aspect, or any one of the implementations of the 2nd aspect and the 4th aspect.
[0046] According to the 18th aspect, an embodiment of the present application provides a decoding device. The decoding device includes a non-volatile memory and a processor coupled to each other. The processor calls program code stored in the memory to execute some or all of the steps in any method according to any one of the 1st aspect, the 3rd aspect, the 5th aspect, or any one implementation of the 1st aspect, the 3rd aspect, and the 5th aspect.
[0047] According to the 19th aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores program code, and the program code includes instructions used to execute some or all of the steps in a method according to any one of the 1st aspect to the 5th aspect, or any one implementation of the 1st aspect to the 5th aspect.
[0048] According to the 20th aspect, an embodiment of the present application provides a computer program product. When the computer program product is executed on a computer, the computer can be made to execute some or all of the steps in a method according to any one of the 1st aspect to the 5th aspect, or any one implementation of the 1st aspect to the 5th aspect.
[0049] It should be understood that the technical solutions in the 2nd aspect to the 10th aspect of the present application are consistent with the technical solution in the 1st aspect. The beneficial effects realized by various aspects and corresponding realizable implementations are the same, and the details will not be described again.
Brief Description of the Drawings
[0050] To more clearly explain the technical solutions in the embodiments of the present invention, the attached drawings or background in the embodiments of the present invention will be described below.
[0051]
Figure 1A
[0052]
Figure 1B
[0053]
Figure 2
[0054]
Figure 3
[0055]
Figure 4
[0056]
Figure 5
[0057]
Figure 6
[0058]
Figure 7
[0059]
Figure 8
[0060]
Figure 9
[0061]
Figure 10
[0062]
Figure 11
[0063]
Figure 12
[0064]
Figure 13
[0065]
Figure 14
[0066]
Figure 15
[0067]
Figure 16
Embodiments for Carrying Out the Invention
[0068] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention. In the following description, reference is made to the accompanying drawings, which form a part hereof and which illustrate specific aspects of the embodiments of the present invention, or specific aspects in which the embodiments of the present invention can be used. It should be understood that the embodiments of the present invention may be used in other aspects and may include structural or logical changes not depicted in the accompanying drawings. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the present invention is defined by the appended claims. For example, it should be understood that the content disclosed in connection with the described method may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, when one or more specific method steps are described, a corresponding device may include one or more units (e.g., a function unit for performing one or more of the described method steps, such as one unit for performing one or more of these steps, or multiple units each performing one or more of these multiple steps), even if such one or more units are not explicitly described or illustrated in the accompanying drawings. Additionally, for example, when a specific device is described based on one or more units such as function units, a corresponding method may include steps used to perform one or more functions of the one or more units (e.g., one step used to perform one or more functions of one or more units, or multiple steps each used to perform one or more functions of one or more units within multiple units), even if one or more of such steps are not explicitly described or illustrated in the accompanying drawings. Furthermore, it should be understood that, unless otherwise specified, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.
[0069] The technical solutions in the embodiments of the present invention can be applied not only to existing video coding standards (such as standards like H.264 and HEVC), but also to future video coding standards (such as the H.266 standard). The terms used in the implementation of the present invention are merely intended to describe specific embodiments of the present invention and are not intended to limit the present invention. First, several concepts that can be used in the embodiments of the present invention will be briefly described below.
[0070] Video coding generally refers to the processing of a series of images that make up a video or video sequence. In the field of video coding, the terms "picture", "frame", and "image" may be used synonymously. The video coding used in this specification refers to video encoding or video decoding. Video encoding is performed on the source side and typically processes the original video image (e.g., by compressing it) to reduce the amount of data representing the video image, thereby optimizing storage and / or transmission. Video decoding is performed on the destination side and usually involves the reverse process compared to the encoder for reconstructing the video image. The "coding" of the video image in the embodiments should be understood as the "encoding" or "decoding" of the video sequence. The combination of the encoding component and the decoding component is also called a codec (CODEC).
[0071] A video sequence includes a series of pictures, one picture is further divided into a plurality of slices, and one slice is further divided into a plurality of blocks. Video coding is performed by blocks. In some new video coding standards, the concept of "block" has been further extended. For example, the H.264 standard has introduced macroblocks (MBs). A macroblock may further be divided into a plurality of prediction blocks (partitions) that can be used for predictive coding. In the high efficiency video coding (HEVC) standard, basic concepts such as "coding unit (CU)", "prediction unit (PU)", and "transform unit (TU)" are used. A plurality of block units are obtained by functional division and are explained by using a new tree-based structure. For example, a CU may be divided into smaller CUs by quadtree division, and this smaller CU may be further divided to generate a quadtree structure. A CU is a basic unit for dividing and encoding a coded picture. The PU and TU also have a similar tree structure. A PU may correspond to a prediction block and is a basic unit for predictive coding. A CU is further divided into a plurality of PUs in a certain partitioning pattern. A TU may correspond to a transform block and is a basic unit for transforming a prediction residual. However, all of the CU, PU, and TU are essentially the concept of a block (or image block).
[0072] For example, in HEVC, by using a quadtree structure represented as a coding tree, a CTU is partitioned into a plurality of CUs. The decision of whether to encode an image area by using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may further be partitioned into one, two, or four PUs based on the partition pattern of the PU. In one PU, the same prediction process is applied, and the relevant information is transmitted to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the partition pattern of the PU, the CU may be partitioned into a plurality of transform units (TUs) based on another quadtree structure similar to the coding tree used for the CU. In the latest development of video compression technology, a CTU is partitioned by using a quad-tree and a multi-type tree to obtain a plurality of CUs. A multi-type tree includes a binary-tree and a ternary-tree. In the partitioning structure, the CU may be square or rectangular.
[0073] In this specification, for ease of explanation and understanding, an encoded image block within a currently coded image may be referred to as the current block. For example, in encoding, the current block is the block being encoded, and in decoding, the current block is the block being decoded. A decoded image block within a reference image that is used to predict the current block is referred to as a reference block. Specifically, a reference block is a block that provides a reference signal for the current block, and the reference signal represents pixel values within the image block. A block within a reference image that provides a prediction signal for the current block may be referred to as a prediction block. The prediction signal represents pixel values, sample values, or sample signals within the prediction block. For example, after traversing multiple reference blocks, the best reference block is found. The best reference block provides a prediction for the current block, and this block is referred to as the prediction block. The current block may also be referred to as the current image block.
[0074] In the case of lossless video coding, the original video image can be reconstructed. That is, (assuming no transmission loss or other data loss occurs during storage or transmission,) the reconstructed video image has the same quality as the original video image. In the case of lossy video coding, for example, further compression is performed by quantization, reducing the amount of data required to represent the video image, and the video image cannot be completely reconstructed on the decoder side. That is, the quality of the reconstructed video image is lower or worse than the quality of the original video image.
[0075] Several H.261 video coding standards are used in the "lossy hybrid video codec" (i.e., spatial and temporal prediction in the sample area is combined with 2D transform coding for applying quantization in the transform area). Each image of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. That is, on the encoder side, video is typically processed, i.e., encoded, at the block (video block) level. For example, prediction blocks are generated by spatial (intra picture) prediction and temporal (inter picture) prediction, the prediction blocks are subtracted from the current block (the block being processed or to be processed) to obtain a residual block, the residual block is transformed and quantized in the transform area to reduce (compress) the amount of data to be transmitted. On the decoder side, the inverse processing part for the encoder is applied to the coded or compressed block to reconstruct the current block for representation. Further, the encoder repeats the decoder's processing loop so that the encoder and decoder generate the same prediction (e.g., intra prediction and inter prediction) and / or reconstruction to process, i.e., code, subsequent blocks.
[0076] Hereinafter, the system architecture to which the embodiments of the present invention are applied will be described. FIG. 1A is a schematic block diagram of an example of a video coding system 10 to which an embodiment of the present invention is applied. As shown in FIG. 1A, the video coding system 10 may include a source device 12 and a destination device 14. The source device 12 generates encoded video data. Consequently, the source device 12 may sometimes be referred to as a video encoding device. The destination device 14 may decode the encoded video data generated by the source device 12. Consequently, the destination device 14 may sometimes be referred to as a video decoding device. In various implementation solutions, the source device 12, the destination device 14, or both the source device 12 and the destination device 14 may include one or more processors and a memory coupled to the one or more processors. As described herein, the memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other computer-accessible medium that can be used to store the desired program code in the form of instructions or data structures. The source device 12 and the destination device 14 may include various devices including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, wireless communication devices, and the like.
[0077] In FIG. 1A, source device 12 and destination device 14 are depicted as separate devices, but in alternative embodiments of the device, both the source device 12 and the destination device 14, or the functions of both the source device 12 and the destination device 14, i.e., the source device 12 or corresponding function and the destination device 14 or corresponding function, may be included. In such embodiments, the source device 12 or corresponding function and the destination device 14 or corresponding function may be implemented by using the same hardware and / or software, or by using separate hardware and / or software or any combination thereof.
[0078] The communication connection between source device 12 and destination device 14 may be implemented via link 13, and destination device 14 may receive the encoded video data from source device 12 via link 13. Link 13 may include one or more media or devices capable of moving the encoded video data from source device 12 to destination device 14. In one example, link 13 may include one or more communication media that enable source device 12 to directly transmit the encoded video data to destination device 14 in real time. In this example, source device 12 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to destination device 14. The one or more communication media may include wireless communication media and / or wired communication media, such as the radio frequency (RF) spectrum, or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, which may be, for example, a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include a router, a switch, a base station, or another device that facilitates communication from source device 12 to destination device 14.
[0079] Source device 12 includes an encoder 20. Optionally, source device 12 may further include an image source 16, an image pre-processor 18, and a communication interface 22. In certain implementations, encoder 20, image source 16, image pre-processor 18, and communication interface 22 may be hardware components within source device 12, or may be software programs within source device 12. Separate explanations are as follows.
[0080] The image source 16 may include, for example, any type of image capture device configured to capture real-world images, and / or any type of device for generating images or comments (in the case of encoding screen content, some text on the screen may also be regarded as part of the image or image to be encoded), such as a computer graphics processor configured to generate computer animation images, or any type of device configured to acquire and / or provide real-world images or computer animation images (such as screen content or virtual reality (VR) images), and / or any combination thereof (such as augmented reality (AR) images). The image source 16 may be a camera for capturing images or a memory for storing images. The image source 16 may further include any type of (internal or external) interface through which previously captured or generated images are stored and / or through which images are acquired or received. When the image source 16 is a camera, the image source 16 may be, for example, a local camera or a camera integrated into the source device. When the image source 16 is a memory, the image source 16 may be, for example, a local memory or a memory integrated into the source device. When the image source 16 includes an interface, the interface may be, for example, an external interface for receiving images from an external video source. The external video source may be, for example, an external image capture device such as a camera, an external memory, or an external image generation device. The external image generation device may be, for example, an external computer graphics processor, a computer, or a server. The interface may be any type of interface according to any unique interface protocol or a standardized interface protocol, such as a wired interface or a wireless interface, or an optical interface.
[0081] An image may be regarded as a two-dimensional array or matrix of picture elements. The picture elements in the array may also be called samples. The size and / or resolution of an image is defined by the number of samples in the horizontal and vertical directions (or, the horizontal and vertical axes) of the array or image. Usually, three color components are used for color representation. For example, an image may be represented as or may include three sample arrays. For example, in the RGB format or color space, an image includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space. For example, an image in the YUV format includes a luminance component indicated by Y (L may sometimes be used instead) and two chrominance components indicated by U and V. The luminance (luma) component Y represents brightness or gray-level intensity (which is the same in a grayscale image, for example), and the two chrominance (chroma) components U and V represent chrominance or color information components. Correspondingly, an image in the YUV format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (U and V). An image in the RGB format may be converted (or transformed) into an image in the YUV format, and vice versa. Such a process is also known as color transformation (or conversion). If an image is monochrome, the image may include only a luma sample array. In this embodiment of the present invention, the image transmitted from the image source 16 to the image processor may also be called the original image data 17.
[0082] The image preprocessor 18 is configured to receive the original image data 17 and preprocess the original image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. For example, the preprocessing performed by the image preprocessor 18 may include trimming, color format conversion (e.g., from the RGB format to the YUV format), color correction, or noise removal.
[0083] An encoder 20 (also referred to as a video encoder 20) is configured to receive pre - processed image data 19 and process the pre - processed image data 19 by using a related prediction mode (such as the prediction mode in each embodiment of this specification) to provide encoded image data 21. (Hereinafter, the structural details of the encoder 20 will be further described based on FIGS. 2, 4, or 5). In some embodiments, the encoder 20 may be configured to execute various embodiments described below to implement the encoder - side application of the inter - prediction method described in the present invention.
[0084] A communication interface 22 is configured to receive the encoded image data 21 and transmit the encoded image data 21 via a link 13 to a destination device 14 or any other device (such as a memory) for storage or direct reconstruction. Any other device may be any device used for decoding or storage. The communication interface 22 may be configured to package the encoded image data 21 into an appropriate format, such as data packets, for transmission via the link 13, for example.
[0085] The destination device 14 includes a decoder 30. Optionally, the destination device 14 may further include a communication interface 28, an image post - processor 32, and a display device 34. Separate explanations are as follows.
[0086] The communication interface 28 may be configured to receive the encoded image data 21 from the source device 12 or any other source. Any other source may be, for example, a storage device, which may be, for example, an encoded image data storage device. The communication interface 28 may be configured to transmit or receive the encoded image data 21 via the link 13 between the source device 12 and the destination device 14, or via any type of network. The link 13 may be, for example, a direct wired connection or a wireless connection. Any type of network may be, for example, a wired network or a wireless network, or any combination thereof, or any type of private network or public network, or any combination thereof. The communication interface 28 may be configured to unpack data packets transmitted via the communication interface 22 to obtain the encoded image data 21.
[0087] Both the communication interface 28 and the communication interface 22 may be configured as a unidirectional communication interface or a bidirectional communication interface, for example, to send and receive messages to establish a connection, and to perform acknowledgment and exchange of any other information related to the communication link and / or data transmission, for example, encoded image data transmission.
[0088] The decoder 30 (also referred to as the decoder 30) is configured to receive the encoded image data 21 and provide the decoded image data 31 or the decoded image 31 (hereinafter, the structural details of the decoder 30 will be further described based on FIGS. 3, 4, or 5). In some embodiments, the decoder 30 may be configured to execute various embodiments described below to implement the decoder-side application of the inter-prediction method described in the present invention.
[0089] The image post-processor 32 is configured to post-process the decoded image data 31 (also referred to as the reconstructed image data) to obtain the post-processed image data 33. The post-processing performed by the image post-processor 32 may include color format conversion (e.g., from YUV format to RGB format), color correction, trimming, resampling, or any other processing. The image post-processor 32 may further be configured to transmit the post-processed image data 33 to the display device 34.
[0090] The display device 34 is configured to receive the post-processed image data 33 and display the image to, for example, a user or viewer. The display device 34 may be any type of display configured to present the reconstructed image, such as an integrated or external display or monitor, or include the same. For example, the display may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0091] Although the source device 12 and the destination device 14 are depicted as separate devices in FIG. 1A, alternative embodiments of the device may include both the source device 12 and the destination device 14, or the functions of both the source device 12 and the destination device 14, i.e., the source device 12 or corresponding function and the destination device 14 or corresponding function. In such embodiments, the source device 12 or corresponding function and the destination device 14 or corresponding function may be implemented by using the same hardware and / or software, or by using separate hardware and / or software or any combination thereof.
[0092] As will be apparent to those skilled in the art based on this specification, the functions of these different units, or the presence and (exact) partitioning of the functions of the source device 12 and / or the destination device 14, shown in FIG. 1A, may vary depending on the actual device and application. The source device 12 and the destination device 14 may each be any one of a wide variety of devices, such as any type of handheld device or fixed device, for example, a notebook computer or laptop computer, a mobile phone, a smartphone, a pad or tablet computer, a video camera, a desktop computer, a set-top box, a television, a camera, an in-vehicle device, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiving device, or a broadcast transmitting device, and may or may not use any type of operating system.
[0093] The encoder 20 and the decoder 30 may each be implemented as any one of a variety of suitable circuits, such as one or more microprocessors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), discrete logic circuitry, hardware, or any combination thereof. When these techniques are implemented partially using software, the device may store software instructions in a suitable non-transitory computer-readable storage medium and execute these instructions using hardware such as one or more processors to perform the techniques of this disclosure. Any one of the foregoing (including hardware, software, combinations of hardware and software, and the like) may be regarded as one or more processors.
[0094] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the technology of the present application may be applied to a video coding setting (e.g., video encoding or video decoding) that does not necessarily include any data communication between the encoding device and the decoding device. In other examples, the data may be obtained from local memory or streamed via a network. The video encoding device may encode the data and store it in memory, and / or the video decoding device may obtain the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data into memory and / or obtain data from memory and decode it.
[0095] FIG. 1B is an exemplary diagram of an example of a video coding system 40 that includes the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3 according to an exemplary embodiment. The video coding system 40 can implement combinations of various techniques in embodiments of the present invention. In the implementation shown, the video coding system 40 may include an imaging device 41, an encoder 20, a decoder 30 (and / or a video encoder / decoder implemented by the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0096] As shown in FIG. 1B, the imaging device 41, the antenna 42, the processing unit 46, the logic circuit 47, the encoder 20, the decoder 30, the processor 43, the memory 44, and / or the display device 45 can communicate with each other. As described, the video coding system 40 is shown with the encoder 20 and the decoder 30, but in different examples, the video coding system 40 may include only the encoder 20 or only the decoder 30.
[0097] In some examples, the antenna 42 may be configured to transmit or receive an encoded bitstream of video data. Additionally, in some examples, the display device 45 may be configured to present video data. In some examples, the logic circuit 47 may be implemented by the processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, or the like. The video coding system 40 may optionally include a processor 43. The optional processor 43 may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, or the like. In some examples, the logic circuit 47 may be implemented by hardware, such as video coding dedicated hardware, and the processor 43 may be implemented by general-purpose software, an operating system, or the like. Additionally, the memory 44 may be any type of memory, such as volatile memory (e.g., Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM)), or non-volatile memory (e.g., flash memory). By way of non-limiting example, the memory 44 may be implemented by cache memory. In some examples, the logic circuit 47 may access the memory 44 (e.g., to implement an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 may include a memory (e.g., a cache) for implementing an image buffer.
[0098] In some examples, the encoder 20 implemented by using a logic circuit may include an image buffer (e.g., implemented by the processing unit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the encoder 20 implemented by using the logic circuit 47 so as to implement various modules described in relation to the encoder system or subsystem shown in FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit may be configured to execute various operations described herein.
[0099] In some examples, the decoder 30 may be implemented in a similar manner by the logic circuit 47 so as to implement various modules described in relation to the decoder 30 shown in FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, the decoder 30 implemented by using a logic circuit may include an image buffer (implemented by the processing unit 2820 or the memory 44) and a graphics processing unit (e.g., implemented by the processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the decoder 30 implemented by using the logic circuit 47 so as to implement various modules described in relation to the decoder system or subsystem shown in FIG. 3 and / or any other decoder system or subsystem described herein.
[0100] In some examples, the antenna 42 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data, indicators, index values, mode selection data, or the like related to the video frame coding described herein, for example, data related to coding partitioning (e.g., transform coefficients or quantized transform coefficients, optional indicators (as described), and / or data defining coding partitioning). The video coding system 40 may further include a decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present video frames.
[0101] In this embodiment of the invention, it should be understood that for the examples described in relation to the encoder 20, the decoder 30 may be configured to perform the reverse process. With respect to signaling syntax elements, the decoder 30 may be configured to receive and parse such syntax elements and, correspondingly, decode the associated video data. In some examples, the encoder 20 may entropy code syntax elements into the encoded video bitstream. In such examples, the decoder 30 may parse such syntax elements and, correspondingly, decode the associated video data.
[0102] Note that the inter prediction method described in the embodiments of the present invention is mainly used in the inter prediction process, and this process exists in both the encoder 20 and the decoder 30. The encoder 20 / decoder 30 in the embodiments of the present invention may conform to video standard protocols such as H.263, H.264, HEVV, MPEG-2, MPEG-4, VP8, or VP9, or may be an encoder / decoder conforming to a next-generation video standard protocol (such as H.266).
[0103] FIG. 2 is a schematic / conceptual block diagram of an exemplary encoder 20 configured to implement an embodiment of the present invention. In the example of FIG. 2, the encoder 20 includes a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy encoding unit 270. The prediction processing unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a mode selection unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not depicted in this figure). The encoder 20 shown in FIG. 2 may also be referred to as a video encoder or a hybrid video encoder based on a hybrid video codec.
[0104] For example, the residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 20, while for example, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, and the prediction processing unit 260 form the inverse signal path of the encoder. The inverse signal path of the encoder corresponds to the signal path of a decoder (see decoder 30 in FIG. 3).
[0105] Encoder 20 receives, for example, from input 202, an image 201 or an image block 203 of image 201, for example, an image from a series of images forming a video or video sequence. Image block 203 may also be referred to as the current image block or the image block to be encoded. Image 201 may be referred to as the current image or the image to be encoded (especially in video coding to distinguish the current image from other images, where other images are, for example, previously encoded and / or decoded images within the same video sequence, i.e., including the current image).
[0106] In some embodiments, encoder 20 may include a partitioning unit (not depicted in FIG. 2) configured to partition image 201 into a plurality of blocks such as image block 203. Image 201 is typically partitioned into a plurality of non - overlapping blocks. The partitioning unit may use the same block size for all images within the video sequence and the corresponding grid defining the block size, or may vary the block size between images or subsets or groups of images and be configured to partition each image into a plurality of corresponding blocks.
[0107] In one example, prediction processing unit 260 of encoder 20 may be configured to perform any combination of the above - described partitioning techniques.
[0108] Similar to image 201, image block 203 may also be or be regarded as a two - dimensional array or matrix of samples having sample values, although of a smaller size than image 201. That is, image block 203 may include, for example, one sample array (e.g., the luma array in the case of a monochrome image 201), three sample arrays (e.g., one luma array and two chroma arrays in the case of a color image), or any other number and / or type of arrays depending on the color format applied. The size of image block 203 is defined by the number of samples in the horizontal and vertical directions (or horizontal and vertical axes) of image block 203.
[0109] The encoder 20 shown in FIG. 2 is configured to encode the image 201 block by block. For example, the encoder encodes and predicts each image block 203.
[0110] The residual calculation unit 204 is configured to calculate a residual block 205 based on the image block 203 and the prediction block 265 (more details regarding the prediction block 265 will be provided below). For example, the residual block 205 within the sample region is obtained by subtracting the sample values of the prediction block 265 from the sample values of the image block 203 sample by sample (pixel by pixel).
[0111] The transformation processing unit 206 is configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transformation coefficients 207 within the transform domain. The transformation coefficients 207 may also be referred to as transform residual coefficients and represent the residual block 205 within the transform domain.
[0112] The conversion processing unit 206 may be configured to apply integer approximations of DCT / DST such as the conversion specified by HEVC / H.265. Compared with the orthogonal DCT transform, such integer approximations are typically scaled by a certain factor. In order to comply with the norm of the residual block processed by using the forward transform and the inverse transform, an additional scale factor is applied as part of the conversion process. The scale factor is usually selected based on several constraints. For example, the scale factor is a power of 2 of the shift operation, the bit depth of the conversion coefficient, or a trade-off between accuracy and implementation cost. For example, for the inverse transform by the inverse conversion processing unit 212 on the decoder side 30 (and, for example, the corresponding inverse transform by the inverse conversion processing unit 212 on the encoder side 20), a specific scaling factor is specified, for example, for the forward transform by the conversion processing unit 206 on the encoder side 20, the corresponding scaling factor may be appropriately specified.
[0113] The quantization unit 208 is configured to quantize the transform coefficients 207, for example, by applying scalar quantization or vector quantization, to obtain quantized transform coefficients 209. The quantized transform coefficients 209 may also be referred to as quantization residual coefficients 209. In the quantization process, the bit depth associated with some or all of the transform coefficients 207 may be reduced. For example, during quantization, an n-bit transform coefficient may be truncated to an m-bit transform coefficient, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, in the case of scalar quantization, different scales may be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, and larger quantization steps correspond to coarser quantization. The applicable quantization steps may be indicated by a quantization parameter (QP). The quantization parameter may be, for example, an index of a predefined set of applicable quantization steps. For example, a smaller quantization parameter may correspond to finer quantization (smaller quantization steps), and a larger quantization parameter may correspond to coarser quantization (larger quantization steps), or vice versa. Quantization may include division by a quantization step and corresponding quantization and / or inverse quantization (e.g., performed by the inverse quantization unit 210), or may include multiplication by a quantization step. In embodiments according to some standards such as HEVC, the quantization parameter may be used to determine the quantization step. Generally, the quantization step may be calculated based on the quantization parameter using a fixed-point approximation of an equation including division. A further scaling factor may be introduced for quantization and dequantization to restore the norm of the residual block, and the norm of the residual block may be modified due to the scale used in the fixed-point approximation of the equation of the quantization step and the quantization parameter. In an exemplary implementation, the scale of the inverse transform and the scale of dequantization may be combined. Alternatively, a customized quantization table may be used and signaled, for example, in a bitstream, from the encoder to the decoder.Quantization is an irreversible operation, and the loss increases as the quantization step increases.
[0114] The inverse quantization unit 210 is configured to apply the inverse quantization of the quantization unit 208 to the quantization coefficients to obtain the dequantization coefficients 211. For example, based on or by using the same quantization step as the quantization unit 208, it is configured to apply the inverse of the quantization scheme applied by the quantization unit 208. The dequantization coefficients 211 may also be referred to as dequantized residual coefficients 211 and usually correspond to the transform coefficients 207, although they are not the same as the transform coefficients due to the loss caused by quantization.
[0115] The inverse transform processing unit 212 is configured to apply the inverse transform of the transform applied by the transform processing unit 206, such as the inverse discrete cosine transform (DCT) or the inverse discrete sine transform (DST), to obtain the inverse transform block 213 within the sample region. The inverse transform block 213 may also be referred to as the inverse transform dequantization block 213 or the inverse transform residual block 213.
[0116] The reconstruction unit 214 (e.g., the analog adder 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265, to obtain the reconstructed block 215 within the sample region.
[0117] Optionally, for example, the buffer unit 216 of the line buffer 216 (abbreviated as "buffer" 216) is configured to buffer or store the reconstructed block 215 and the corresponding sample values, for example, for intra prediction. In other embodiments, the encoder may be configured to use the unfiltered reconstructed block and / or the corresponding sample values stored in the buffer unit 216 for any type of estimation and / or prediction, for example, intra prediction.
[0118] For example, in one embodiment, the encoder 20 may be configured such that the buffer unit 216 is used not only to store the reconstructed block 215 for the intra prediction unit 254, but also for the loop filter unit 220 (not depicted in FIG. 2), and / or, for example, the buffer unit 216 and the decoded image buffer unit 230 may be configured to form one buffer. In other embodiments, the filtered block 221 and / or the block or sample from the decoded image buffer 230 (neither the block nor the sample is depicted in FIG. 2) may be used as the input or basis for the intra prediction unit 254.
[0119] The loop filter unit 220 (simply referred to as the "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221 and perform smoothing of pixel transitions or improvement of video quality. The loop filter unit 220 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter, for example, a bilateral filter, an adaptive loop filter (ALF), an edge enhancement filter or a smoothing filter, or a collaborative filter. Although the loop filter unit 220 is shown as an in-loop filter in FIG. 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221. The decoded image buffer 230 may store the reconstructed encoded block after the loop filter unit 220 performs a filtering operation on the reconstructed encoded block.
[0120] In one embodiment, the encoder 20 (and correspondingly, the loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information), for example, directly or after entropy encoding is performed by the entropy encoding unit 270 or any other entropy encoding unit, such that, for example, the decoder 30 can receive the same loop filter parameters and apply the same loop filter parameters to decoding.
[0121] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use in encoding video data by the encoder 20. The DPB 230 may be formed by any one of various memory devices such as a dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), and resistive RAM (RRAM (registered trademark))), or another type of memory device. The DPB 230 and the buffer 216 may be provided by the same memory device or separate memory devices. In one example, the decoded picture buffer (DPB) 230 is configured to store the filtered block 221. The decoded picture buffer 230 may further be configured to store the same current picture or a different picture, e.g., another previously filtered block of a previously reconstructed picture, e.g., the block 221 that has been previously reconstructed and filtered, and may provide the previously reconstructed, i.e., decoded, complete picture (and corresponding reference blocks and samples) and / or the partially reconstructed current picture (and corresponding reference blocks and samples) for, e.g., inter prediction. In one example, when the reconstructed block 215 is reconstructed without loop filtering, the decoded picture buffer (DPB) 230 is configured to store the reconstructed block 215.
[0122] The prediction processing unit 260, also referred to as the block prediction processing unit 260, is configured to receive or acquire an image block 203 (the current image block 203 of the current image 201), the reconstructed image data, e.g., reference samples of the same (current) image from the buffer 216, and / or reference image data 231 of one or more previously decoded images from the decoded image buffer 230, and to process such data for prediction, i.e., to provide a prediction block 265 which can be an inter prediction block 245 or an intra prediction block 255.
[0123] The mode selection unit 262 may be configured to select a corresponding prediction block 245 or 255 to be used as the prediction block 265 and / or a prediction mode (e.g., an intra prediction mode or an inter prediction mode) for the calculation of the residual block 205 and the reconstruction of the reconstructed block 215.
[0124] In one embodiment, the mode selection unit 262 may be configured to select a prediction mode (e.g., from among the prediction modes supported by the prediction processing unit 260). The prediction mode provides the best match or the minimum residual (the minimum residual means better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or both considerations or balancing. The mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO). That is, it may be configured to select a prediction mode that provides the minimum rate distortion optimization or a prediction mode whose associated rate distortion meets at least the selection criteria of the prediction mode.
[0125] Hereinafter, the prediction processing (e.g., executed by the prediction processing unit 260) and mode selection (e.g., executed by the mode selection unit 262) executed by the exemplary encoder 20 will be described in detail.
[0126] As described above, the encoder 20 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0127] The set of intra prediction modes may include 35 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes such as those defined in H.265, or may include 67 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes such as those defined in the yet - to - be - finalized H.266.
[0128] In a possible implementation, a set of inter prediction modes depends on available reference images (i.e., at least some of the decoded images stored in the DPB 230 as described above, for example) and other inter prediction parameters. For example, it depends on whether the entire reference image or only a part of the reference image, for example, a search window area around the area of the current block, is used for searching for the best matching reference block, and / or, for example, whether pixel interpolation such as half-pixel interpolation and / or 1 / 4-pixel interpolation is applied. A set of inter prediction modes may include, for example, an Advanced Motion Vector Prediction (AMVP) mode and a merge mode. In a specific implementation, a set of inter prediction modes may include an improved control point-based AMVP mode and an improved control point-based merge mode in embodiments of the present invention. In one example, the intra prediction unit 254 may be configured to execute any combination of the inter prediction techniques described below.
[0129] In addition to the aforementioned prediction modes, in embodiments of the present invention, a skip mode and / or a direct mode may be applied.
[0130] The prediction processing unit 260 may further divide the image block 203 into smaller block partitions or sub-blocks by repeatedly using, for example, a quad-tree (QT) partitioning, a binary-tree (BT) partitioning, a triple-tree (TT) partitioning, or any combination thereof, and may be configured to perform predictions for each of these block partitions or sub-blocks. The mode selection includes the selection of the tree structure of the divided image block 203 and the selection of the prediction mode applied to each of these block partitions or sub-blocks.
[0131] The inter prediction unit 244 may include a motion estimation (ME) unit (not depicted in FIG. 2) and a motion compensation (MC) unit (not depicted in FIG. 2). The motion estimation unit is configured to receive or acquire the image block 203 (the current image block 203 of the current image 201) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, reconstructed blocks of one or more other / different previously decoded images 231, for motion estimation. For example, the video sequence may include the current image and a previously decoded image 31. That is, the current image and the previously decoded image 31 may be part of a series of images forming the video sequence or may form a series of images.
[0132] For example, the motion estimation unit (not depicted in FIG. 2) may be configured to select a reference block from among a plurality of reference blocks of the same or different images of a plurality of other images, and provide a reference image to the motion compensation unit (not depicted in FIG. 2), and / or provide an offset (spatial offset) between the motion vector (the position (coordinates X and Y) of the reference block) and the position of the current block as an inter prediction parameter. This offset is also referred to as a motion vector (MV).
[0133] The motion compensation unit is configured to obtain an inter-prediction parameter and perform inter-prediction based on or by using the inter-prediction parameter to obtain an inter-prediction block 245. The motion compensation performed by the motion compensation unit (not depicted in FIG. 2) may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation (interpolation may be performed at a sub-sample accuracy level). By interpolation filtering, additional pixel samples may be generated from known pixel samples, thereby potentially increasing the number of candidate prediction blocks that can be used for coding an image block. When the motion compensation unit 246 receives the motion vector of the PU of the current image block, it may identify the position of the prediction block indicated by the motion vector in one of a plurality of reference image lists. The motion compensation unit 246 may further generate syntax elements associated with the block and the video slice that are used by the video decoder 30 when decoding the image block of the video slice.
[0134] Specifically, the inter-prediction unit 244 may transmit syntax elements to the entropy coding unit 270, and the syntax elements include an inter-prediction parameter (e.g., indication information of an inter-prediction mode selected for prediction of the current block after traversing a plurality of inter-prediction modes, or at least one of an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block). In a possible application scenario, when there is only one inter-prediction mode, the inter-prediction parameter may alternatively not be held in the syntax element. In this case, the decoder side 30 may directly perform decoding in the default prediction mode. It can be understood that the inter-prediction unit 244 may be configured to perform any combination of inter-prediction techniques.
[0135] The intra prediction unit 254 is configured to obtain, for example, receive, the image block 203 (current image block) and one or more previously reconstructed blocks of the same image, such as reconstructed neighboring blocks, for intra estimation. The encoder 20 may be configured to select, for example, an intra prediction mode from among a plurality of (predetermined) intra prediction modes.
[0136] In one embodiment, the encoder 20 may be configured to select an intra prediction mode according to an optimization criterion, for example, based on a minimum residual (e.g., an intra prediction mode that provides a prediction block 255 that is most similar to the current image block 203) or a minimum rate distortion.
[0137] The intra prediction unit 254 is further configured to determine an intra prediction block 255, for example, based on the intra prediction parameters of the selected intra prediction mode. In any case, after selecting the intra prediction mode of the block, the intra prediction unit 254 is further configured to provide the intra prediction parameters, that is, information indicating the selected intra prediction mode of the block, to the entropy coding unit 270. In one example, the intra prediction unit 254 may be configured to perform any combination of intra prediction techniques.
[0138] Specifically, the intra prediction unit 254 may transmit syntax elements to the entropy coding unit 270, and the syntax elements include intra prediction parameters (e.g., indication information of the intra prediction mode selected for the prediction of the current block after traversing a plurality of intra prediction modes). In a possible application scenario, when there is only one intra prediction mode, the intra prediction parameters may alternatively not be held in the syntax elements. In this case, the decoder side 30 may directly perform decoding in the default prediction mode.
[0139] The entropy encoding unit 270 applies (or bypasses) an entropy encoding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CAVLC) scheme, arithmetic coding scheme, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique) to one or all of the quantized coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters, to obtain encoded image data 21 that can be output via output 272, for example, in the form of an encoded bitstream 21. The encoded bitstream may be transmitted to the video decoder 30 or archived for later transmission or acquisition by the video decoder 30. The entropy encoding unit 270 may further be configured to entropy encode other syntax elements of the currently encoded video slice.
[0140] Other structural variations of the video encoder 20 may be configured to encode the video stream. For example, the non-transform-based encoder 20 may directly quantize the residual signal without the transform processing unit 206 for some blocks or frames. In another implementation, the encoder 20 may include a quantization unit 208 and an inverse quantization unit 210 that may be combined into a single unit.
[0141] Specifically, in this embodiment of the present invention, the encoder 20 may be configured to implement the inter prediction method described in the following embodiments.
[0142] It should be understood that other structural variations of the video encoder 20 may be configured to encode a video stream. For example, for some image blocks or image frames, the video encoder 20 may directly quantize the residual signal without being processed by the conversion processing unit 206 and, correspondingly, without being processed by the inverse conversion processing unit 212. Alternatively, for some image blocks or image frames, the video encoder 20 may not generate residual data, and correspondingly, the conversion processing unit 206, the quantization unit 208, the inverse quantization unit 210, and the inverse conversion processing unit 212 do not need to perform processing. Alternatively, the video encoder 20 may directly store the reconstructed image block as a reference block without being processed by the filter 220. Alternatively, the quantization unit 208 and the inverse quantization unit 210 in the video encoder 20 may be combined together. The loop filter 220 is optional. In the case of lossless compression encoding, the conversion processing unit 206, the quantization unit 208, the inverse quantization unit 210, and the inverse conversion processing unit 212 are optional. It should be understood that in different application scenarios, the inter prediction unit 244 and the intra prediction unit 254 may be selectively used.
[0143] FIG. 3 is a schematic / conceptual block diagram of an exemplary decoder 30 configured to implement an embodiment of the present invention. The video decoder 30 is configured to receive, for example, encoded image data (e.g., an encoded bitstream) 21 encoded by the encoder 20 and obtain a decoded image 231. In the decoding process, the video decoder 30 receives from the video encoder 20 an encoded video bitstream representing video data, e.g., an image block of an encoded video slice, and associated syntax elements.
[0144] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an analog adder 314), a buffer 316, a loop filter 320, a decoded image buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 30 may perform a decoding path generally inverse to the encoding path described in relation to the video encoder 20 of FIG. 2.
[0145] The entropy decoding unit 304 performs entropy decoding on the encoded image data 21 to obtain, for example, quantization coefficients 309 and / or decoded encoding parameters (not depicted in FIG. 3), such as any one or all of (decoded) inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 is further configured to transfer inter prediction parameters, intra prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0146] The inverse quantization unit 310 may have the same function as the inverse quantization unit 110. The inverse transform processing unit 312 may have the same function as the inverse transform processing unit 212. The reconstruction unit 314 may have the same function as the reconstruction unit 214. The buffer 316 may have the same function as the buffer 216. The loop filter 320 may have the same function as the loop filter 220. The decoded image buffer 330 may have the same function as the decoded image buffer 230.
[0147] The prediction processing unit 360 may include an inter prediction unit 344 and an intra prediction unit 354. The inter prediction unit 344 may have a function similar to that of the inter prediction unit 244, and the intra prediction unit 354 may have a function similar to that of the intra prediction unit 254. The prediction processing unit 360 is typically configured to perform block prediction and / or obtain a prediction block 365 from the encoded data 21, and to receive or obtain prediction-related parameters (e.g., at least one of an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block) and / or information regarding the selected prediction mode, for example, from the entropy decoding unit 304 (explicitly or implicitly).
[0148] When the video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 for the image blocks of the current video slice based on the signaled intra prediction mode and the data from previously decoded blocks of the current frame or image. When the video frame is coded as an inter-coded (B or P) slice, the inter prediction unit 344 (e.g., the motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 for the video blocks of the current video slice based on the motion vector received from the entropy decoding unit 304 and another syntax element (e.g., the syntax element may be at least one of an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block). In the case of inter prediction, the prediction block may be generated from one of a plurality of reference images within one reference image list. The video decoder 30 may construct reference frame lists 0 and 1 by using a default construction technique based on the reference images stored in the DPB 330.
[0149] The prediction processing unit 360 is configured to determine prediction information for video blocks of a current video slice by analyzing motion vectors and / or other syntax elements, and to generate a prediction block for the currently decoded video block using the prediction information. In an example of the present invention, the prediction processing unit 360 uses some received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for encoding video blocks within a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), one or more construction information of reference picture lists of the slice, motion vectors of each inter-coded video block within the slice, an inter prediction state of each inter-coded video block within the slice, and other information, and decodes video blocks within the current video slice. In another example of the present disclosure, the syntax elements received from the bitstream by the video decoder 30 include syntax elements in one or more of an adaptive parameter set (APS), a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header.
[0150] The inverse quantization unit 310 may be configured to inverse-quantize (i.e., dequantize) the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The inverse quantization process may include determining the degree of quantization to be applied, as well as the degree of inverse quantization to be applied, using quantization parameters calculated by the video encoder 20 for each video block within the video slice.
[0151] The inverse transform processing unit 312 is configured to apply an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to generate a residual block in the pixel domain.
[0152] The reconfiguration unit 314 (e.g., analog adder 314) is configured to add, for example, the sample values of the reconfigured residual block 313 and the sample values of the prediction block 365, so as to add the inverse transform block 313 (i.e., the reconfigured residual block 313) to the prediction block 365 to obtain the reconfigured block 315 within the sample region.
[0153] The loop filter unit 320 (during or after the coding loop) is configured to filter the reconfigured block 315 to obtain the filtered block 321, and perform smoothing of pixel transitions or improvement of video quality. In one example, the loop filter unit 320 may be configured to execute any combination of the filtering techniques described below. The loop filter unit 320 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or another filter, for example, a bilateral filter, an adaptive loop filter (ALF), an edge enhancement filter or a smoothing filter, or a collaborative filter. Although the loop filter unit 320 is shown as an in-loop filter in FIG. 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0154] Next, the decoded video block 321 in a given frame or image is stored in the decoded image buffer 330 that stores the reference image used for subsequent motion compensation.
[0155] The decoder 30 is configured to output, for example, the decoded image 31 via the output 332 for presentation to the user or browsing by the user.
[0156] Other variations of the video decoder 30 may be configured to decode a compressed bitstream. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 may directly inverse-quantize the residual signal without the inverse transform processing unit 312 for some blocks or frames. In another implementation, the video decoder 30 may include an inverse quantization unit 310 and an inverse transform processing unit 312 that may be combined into a single unit.
[0157] Specifically, in this embodiment of the present invention, the decoder 30 is configured to implement the inter prediction method described in the following embodiments.
[0158] It should be understood that other structural variations of the video decoder 30 may be configured to decode an encoded video bitstream. For example, the video decoder 30 may generate an output video stream without the processing performed by the filter 320. Alternatively, for some image blocks or image frames, the entropy decoding unit 304 of the video decoder 30 does not obtain quantization coefficients by decoding, and correspondingly, the inverse quantization unit 310 and the inverse transform processing unit 312 do not need to perform processing. The loop filter 320 is optional. In the case of lossless compression, the inverse quantization unit 310 and the inverse transform processing unit 312 are optional. It should be understood that in different application scenarios, the inter prediction unit and the intra prediction unit may be selectively used.
[0159] It should be understood that in the encoder 20 and the decoder 30 of the present application, the processing result of a certain procedure can be output to the next procedure after being further processed. For example, after procedures such as interpolation filtering, motion vector derivation, or loop filtering, operations such as clip or shift are further performed on the processing result of the corresponding procedure.
[0160] For example, the motion vectors of the sub-blocks of the current image block may be further processed from the motion vectors of the control points of the current image block or the motion vectors of the adjacent affine coding blocks. This is not limited in the present application. For example, the value of the motion vector is restricted within a range of a specific bit width. Assuming that the allowable bit width of the motion vector is bitDepth, the range of the value of the motion vector is from -2^(bitDepth-1) to 2^(bitDepth-1)-1, and the symbol "^" represents exponentiation. When bitDepth is 16, the range of values is from -32768 to 32767. When bitDepth is 18, the range of values is from -131072 to 131071. In another example, the value of the motion vector (e.g., the motion vectors MV of four 4×4 sub-blocks within an 8×8 image block) is restricted such that the maximum difference between the integer parts of the MVs of the four 4×4 sub-blocks does not exceed N pixels, for example, does not exceed 1 pixel.
[0161] To restrict the motion vector within a specific bit width, the following two methods may be used.
[0162] Method 1: The overflow most significant bit of the motion vector is removed. ux=(vx+2 bitDepth )%2 bitDepth vx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux uy=(vy+2 bitDepth )%2 bitDepth vy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy
[0163] vx represents the horizontal component of the motion vector of the image block or the sub-block of the image block. vy represents the vertical component of the motion vector of the image block or the sub-block of the image block. ux and uy are intermediate values. bitDepth represents the bit depth.
[0164] For example, the value of vx is -32769, and 32767 is derived according to the above formula. The value is stored in the computer in two's complement representation, and the two's complement representation of -32769 is 1,0111,1111,1111,1111 (17 bits), and the process executed by the computer due to overflow discards the most significant bit. Therefore, the value of vx is 0111,1111,1111,1111, that is, 32767. This value is consistent with the result derived by the process according to the formula.
[0165] Method 2: As shown in the following formula, clipping is performed on the motion vector. vx = Clip3(-2 bitDepth-1 , 2 bitDepth-1 , -1, vx) vy = Clip3(-2 bitDepth-1 , 2 bitDepth-1 , -1, vy)
[0166] vx represents the horizontal component of the motion vector of the image block or the sub-block of the image block. vy represents the vertical component of the motion vector of the image block or the sub-block of the image block. x, y, and z correspond to the three input values of the MV clamping process Clip3. Clip3 is defined to indicate that the value of z is clipped to the range [x, y].
Number
[0167] FIG. 4 is a schematic structural diagram of a video coding device 400 (e.g., a video encoding device 400 or a video decoding device 400) according to an embodiment of the present invention. The video coding device 400 is suitable for implementing the embodiments described herein. In one embodiment, the video coding device 400 may be a video decoder (e.g., decoder 30 in FIG. 1A) or a video encoder (e.g., encoder 20 in FIG. 1A). In another embodiment, the video coding device 400 may be one or more components of decoder 30 in FIG. 1A or encoder 20 in FIG. 1A.
[0168] The video coding device 400 includes an inlet port 410 and a receiving unit (Rx) 420 configured to receive data, a processor, a logic unit, or a central processing unit (CPU) 430 configured to process data, a transmitting unit (Tx) 440 and an outlet port 450 configured to transmit data, and a memory 460 configured to store data. The video coding device 400 may further include optical / electrical components and electrical / optical (EO) components coupled to the inlet port 410, the receiving unit 420, the transmitting unit 440, and the outlet port 450 for optical or electrical signals to enter and exit.
[0169] Processor 430 is implemented by using hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiving unit 420, a transmitting unit 440, an output port 450, and a memory 460. Processor 430 includes a coding module 470 (e.g., an encoding module 470 or a decoding module 470). The encoding / decoding module 470 implements the embodiments disclosed herein and implements the inter-prediction method provided in the embodiments of the present invention. For example, the encoding / decoding module 470 implements, processes, or provides various coding operations. Therefore, including the encoding / decoding module 470 provides a substantial improvement in the functions of the video coding device 400 and affects the conversion of the video coding device 400 to different states. Alternatively, the encoding / decoding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0170] Memory 460 includes one or more disks, tape drives, and solid state drives and is used as an overflow data storage device for storing programs when the programs are selectively executed and for storing instructions and data read during the execution of the programs. Memory 460 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0171] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 12 and the destination device 14 of FIG. 1A according to an exemplary embodiment. The apparatus 500 can implement the technology of the present application. That is, FIG. 5 is a schematic block diagram of the implementation of an encoding device or a decoding device (abbreviated as the coding device 500) according to an embodiment of the present application. The coding device 500 may include a processor 510, a memory 530, and a bus system 550. The processor and the memory are connected via the bus system. The memory is configured to store instructions. The processor is configured to execute the instructions stored in the memory. The memory of the coding device stores program code. The processor can call the program code stored in the memory to execute the video encoding method or the video decoding method described in the present application, particularly various new inter-prediction methods. To avoid repetition, the details are not described again here.
[0172] In this embodiment of the present application, the processor 510 may be a central processing unit (abbreviated as "CPU"), or the processor 510 may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or another programmable logic device, a discrete gate device or transistor logic device, a discrete hardware component, or the like. The general-purpose processor may be a microprocessor, or any conventional processor or the like.
[0173] Memory 530 may include a read-only memory (ROM) device or a random access memory (RAM) device. Alternatively, any other suitable type of storage device may be used as memory 530. Memory 530 may include code and data 531 that are accessed by processor 510 using bus 550. Memory 530 may further include an operating system 533 and an application program 535. Application program 535 includes at least one program that enables processor 510 to execute the video encoding method or video decoding method described in this application (in particular, the inter prediction method described in this application). For example, application program 535 may include applications 1 to N, and may further include a video encoding application or a video decoding application (abbreviated as a video coding application) that executes the video encoding method or video decoding method described in this application.
[0174] Bus system 550 may include not only a data bus, but may also include a power bus, a control bus, a status signal bus, and the like. However, for clarity of explanation, the various types of buses in this figure are represented as bus system 550.
[0175] Optionally, coding device 500 may further include one or more output devices, such as display 570. In one example, display 570 may be a touch-sensitive display that combines a display and a touch-sensitive unit operable to sense touch input. Display 570 may be connected to processor 510 via bus 550.
[0176] Hereinafter, the solutions in the embodiments of this application will be described in detail.
[0177] Inter prediction in a video encoding method or a video decoding method executed by an inter prediction unit 244, an inter prediction unit 344, an encoder 20, a decoder 30, a video coding device 400, or a coding device 500 includes determination of motion information. Specifically, the motion information may be determined by a motion estimation unit, and the motion information may include at least one type of reference picture information and motion vector information. The reference picture information may include at least one of unidirectional / bidirectional prediction information (bidirectional prediction means that two reference blocks are required to determine a prediction block of a current picture block, and in bidirectional prediction, two groups of motion information are required to determine two reference blocks), information regarding a reference picture list, and a reference picture index corresponding to the reference picture list. The motion vector information may include a motion vector, and the motion vector indicates position offsets in a horizontal direction and a vertical direction. The motion vector information may further include a motion vector difference (MVD). The determination of the motion information and the determination of the prediction block may include one of the following modes.
[0178] In the AMVP mode, first, the encoder constructs a candidate motion vector list based on the motion vectors of spatially or temporally adjacent blocks (e.g., but not limited to, encoded blocks) of the current block, and then determines the motion vector of the motion vector predictor (MVP) of the current block from the candidate motion vector list by calculating the bitrate distortion. The encoder transfers to the decoder side the index value of the selected motion vector predictor in the candidate motion vector list and the index value of the reference frame (the reference frame may also be called the reference image). Further, the encoder performs a motion search in the vicinity of the MVP center to obtain a better motion vector (also called the motion vector target value) of the current block. The encoder transfers to the decoder side the difference (Motion vector difference) between the MVP and the best motion vector. First, the decoder constructs a candidate motion vector list by using the motion vectors of spatially or temporally adjacent blocks (e.g., but not limited to, decoded blocks) of the current block, obtains the motion vector predictor based on the candidate motion vector list and the obtained index value of the motion vector predictor in the candidate motion vector list, obtains a better motion vector based on the obtained difference between the MVP and the better motion vector, and obtains the predicted block of the current block based on the better motion vector and the reference frame obtained based on the index value of the reference frame. Note that when there is one candidate in the candidate motion vector list, the index value of the selected motion vector predictor in the candidate motion vector list may not be transmitted.
[0179] The encoder side may be the source device 12, the video coding system 40, the encoder 20, the video coding device 400, or the coding device 500. The decoder side may be the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. The index value of the motion vector prediction factor in the candidate motion vector list and the index value of the reference frame (the reference frame may also be called a reference picture) may be syntax elements used for transmission in the foregoing description.
[0180] In merge mode, the encoder first constructs a candidate motion information list based on the motion information of spatially or temporally adjacent blocks (e.g., but not limited to, encoded blocks) of the current block, calculates the rate distortion, determines the best motion information from the candidate motion information list as the motion information of the current block, and transfers the index value of the position of the best motion information from the candidate motion information list (shown as a merge index also applicable to the following description) to the decoder side. FIG. 6 shows the spatial and temporal candidate motion information of the current block. The spatial candidate motion information is the motion information of five spatially adjacent blocks (A0, A1, B0, B1, and B2). If the adjacent blocks are not available or the intra coding mode is used, the spatial candidate motion information is not added to the candidate motion information list. The temporal candidate motion information of the current block is obtained by scaling the MV of the collocated blocks in the reference frame based on the reference frame and the picture order count (POC) of the current frame. First, it is determined whether the block at position T0 in the reference frame is available. If the block is not available, the block at position T1 is selected. The decoder side first constructs a candidate motion information list based on the motion information of spatially or temporally adjacent blocks (e.g., but not limited to, decoded blocks) of the current block, and the motion information in the motion information list includes a motion vector and an index value of the reference frame. The decoder side then obtains the best motion information based on the candidate motion information list and the index value of the position of the best motion information in the candidate motion information list, and obtains the predicted block of the current block based on the best motion information. Note that when there is one candidate in the candidate motion information list, the index value of the position of the best motion information in the candidate motion information list may not be transmitted.
[0181] Note that the candidate motion information list in the embodiment of the present invention is constructed based on the motion information of spatially or temporally adjacent blocks of the current block, but is not limited thereto. The candidate motion information list may be constructed or improved by using at least one of the motion information of spatially adjacent blocks, the motion information of temporally adjacent blocks, pairwise average merging candidate, history-based merging candidate, and zero motion vector merging candidate. For a detailed description of the construction process, refer to JVET-L1001-v6. The embodiment of the present invention is not limited thereto.
[0182] In the merge mode with motion vector difference (MMVD), MVD transmission based on the merge mode is added. Specifically, on the encoder side, a motion search is further performed in the proximity area centered on the best motion information in the candidate motion information list to obtain a better motion vector (which may also be called a motion vector target value) for the current block. The encoder side transmits the difference (Motion vector difference) between the better motion vector and the motion vector included in the best motion information in the candidate motion information list to the decoder side. After obtaining the best motion information, the decoder side further obtains a better motion vector based on the difference and the motion vector included in the best motion information in the candidate motion information list. Next, based on the better motion vector and the reference frame indicated by the index value of the reference frame included in the best motion information in the candidate motion information list, a predicted block of the current block is obtained.
[0183] Note that the MMVD mode is a mode in which MVD transmission based on the skip mode is added. Compared with the merge mode, the skip mode can be understood as residual information between the predicted block of the current block and the original block of the current block. Similarly, based on the skip mode, the encoder side further performs motion search in the proximity area centered on the best motion information in the candidate motion information list to obtain a better motion vector for the current block. The encoder side then transmits the difference (Motion vector difference) between the better motion vector and the motion vector included in the best motion information in the candidate motion information list to the decoder side. After obtaining the best motion information, the decoder side further obtains a better motion vector based on the difference and the motion vector included in the best motion information in the candidate motion information list. Next, based on the better motion vector and the reference frame indicated by the index value of the reference frame included in the best motion information in the candidate motion information list, the predicted block of the current block is obtained. For the description of the skip mode, please refer to the existing H.266 draft (working draft, e.g., JVET-L1001-v6). It will not be described in detail here again.
[0184] In the MMVD mode, multiple merge candidates within the VVC are used. One or more of these merge candidates are selected, and then an MV extension formula is executed based on the selected one or more candidates. The MV extension formula is implemented by using a simplified identification method. The identification method includes the steps of identifying the starting point of the MV, the motion step, and the motion direction. By using the existing merge candidate list, the selected candidate may be the MRG_TYPE_DEFAULT_N mode. Based on the selected candidate, the initial position of the MV is determined. The basic candidate IDX (Table 1) indicates a specific candidate selected as the best candidate in the candidate list.
Table 1
[0185] The basic candidate IDX represents the index value of the position of the best motion information in the candidate motion information list. The Nth MVP indicates that the Nth item in the candidate motion information list is the MVP.
[0186] When the MVD is transmitted, offset values such as x and y may be transmitted, or the length and direction of the MVD may be transmitted, or the index value of the length of the MVD (indicating a distance from 1 / 4 pixel to 32 pixels) and the index value of the direction of the MVD (up, down, left, or right) may be transmitted.
[0187] The index value of the length of the MVD is used to indicate the length of the MVD. A correspondence relationship between the index value of the length of the MVD (Distance IDX) and the length of the MVD (Pixel distance) may be preset, and the correspondence relationship may be shown in Table 2.
Table 2
[0188] The index value of the direction of the MVD is used to indicate the direction of the MVD. A correspondence relationship between the index value of the direction of the MVD (Direction IDX) and the direction of the MVD (x-axis or y-axis) may be preset, and the correspondence relationship may be shown in Table 3.
Table 3
[0189] In Table 3, when the value of the y-axis is N / A, it may indicate that the direction of the MVD is independent of the y-axis direction, and when the value of the x-axis is N / A, it may indicate that the direction of the MVD is independent of the x-axis direction.
[0190] In the decoding process, the MMVD flag (mmvd_flag, used to indicate whether the current block is decoded in MMVD mode) is parsed after the skip flag (cu_skip_flag, used to indicate whether the current block is decoded in skip mode) or the merge flag (merge_flag, used to indicate whether the current block is decoded in merge mode). If the skip flag or the merge flag is true, it is necessary to parse the value of the MMVD flag. If the MMVD flag is true, it is necessary to encode or decode other flag values corresponding to MMVD.
[0191] Furthermore, in a bidirectional inter prediction (or also called bidirectional prediction) scenario, the decoder side or the encoder side may decode or encode only one-direction MVD information, and the MVD information in the other direction may be obtained based on the one-direction MVD information. The specific process may be as follows.
[0192] Specifically, whether the reference frame corresponding to the one-way MVD and the reference frame corresponding to the other-way MVD are in the same direction or in the opposite direction may be determined based on the POC value of the frame in which the current block is placed and the POC values of the reference frames in these two directions. For example, if the plus or minus sign of the first difference obtained by subtracting the POC value of the reference frame in one direction from the POC value of the frame in which the current block is placed is the same as the plus or minus sign of the second difference obtained by subtracting the POC value of the reference frame in the other direction from the POC value of the frame in which the current block is placed, the reference frame corresponding to the one-way MVD and the reference frame corresponding to the other-way MVD are in the same direction. Otherwise, if the plus or minus sign of the first difference obtained by subtracting the POC value of the reference frame in one direction from the POC value of the frame in which the current block is placed is opposite to the plus or minus sign of the second difference obtained by subtracting the POC value of the reference frame in the other direction from the POC value of the frame in which the current block is placed, the reference frame corresponding to the one-way MVD and the reference frame corresponding to the other-way MVD are in different directions.
[0193] Note that the direction of the reference frame may be the direction of the reference frame with respect to the direction of the current frame (the frame in which the current block is placed), or the direction of the current frame with respect to the direction of the reference frame. In a specific implementation process, whether the reference frame corresponding to the one-way MVD and the reference frame corresponding to the other-way MVD are in the same direction or in the opposite direction may be determined based on the difference obtained by subtracting the POC value of the current frame from the POC value of the current frame, or may be determined based on the difference obtained by subtracting the POC value of the current frame from the POC value of the reference frame.
[0194] If the reference frame corresponding to the MVD in one direction and the reference frame corresponding to the MVD in the other direction are in the same direction, the sign of the MVD in the other direction is the same as the sign of the MVD in one direction. For example, if the MVD in one direction is (x, y), the MVD in the other direction is (x, y). Specifically, the MVD in one direction may be obtained based on the index value of the length of the MVD in one direction and the index value of the direction of the MVD in one direction.
[0195] Alternatively, if the reference frame corresponding to the MVD in one direction and the reference frame corresponding to the MVD in the other direction are in different directions, the sign of the MVD in the other direction is opposite to the sign of the MVD in one direction. For example, if the MVD in one direction is (x, y), the MVD in the other direction is (-x, -y).
[0196] Scale the MVD in the other direction based on the MVD in one direction, the first POC difference, and the second POC difference to obtain a better MVD in the other direction. The first POC difference is the difference between the POC value of the frame in which the current block is located and the POC value of the reference frame corresponding to the MVD in one direction, and the second POC difference is the difference between the POC value of the frame in which the current block is located and the POC value of the reference frame corresponding to the MVD in the other direction. For a specific description of the scaling method, please refer to the existing H.266 draft (working draft, for example, JVET-L1001-v6). Details are not described here.
[0197] The foregoing MVD solution can be further optimized. For example, index values with relatively large pixel distance values are rarely used, the value of the direction of the MVD can only indicate four directions, and in a specific process, the step (3) of scaling the MVD in the other direction based on the MVD in one direction, the first POC difference, and the second POC difference to obtain a better MVD in the other direction is complex. Therefore, the embodiments of the present invention provide a series of improved solutions.
[0198] FIG. 7 is a schematic flowchart of an inter prediction method according to an embodiment of the present invention. The method may be executed by a destination device 14, a video coding system 40, a decoder 30, a video coding device 400, or a coding device 500. Specifically, the method may be executed by the video decoder 30, or specifically, by an entropy decoding unit 304 and a prediction processing unit 360 (or, for example, an inter prediction unit 344 within the prediction processing unit 360). The method may include the following steps.
[0199] S701: Obtain a motion vector prediction factor for a current image block.
[0200] In a specific implementation process, the step of obtaining a motion vector prediction factor for a current image block may include constructing a candidate motion information list for the current image block, where the candidate motion information list includes L motion vectors, and L is 1, 3, 4, or 5; obtaining an index value of prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes a motion vector prediction factor; and obtaining a motion vector prediction factor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list. The candidate motion information list for the current image block may be a merge candidate motion information list. Correspondingly, the inter prediction method provided in this embodiment of the present invention may be applied to the MMVD mode.
[0201] The index value in the candidate motion information list may be an index value in a variable-length coding form. For example, when L is 3, 1 may be used to indicate the first item in the candidate motion information list, 01 may be used to indicate the second item in the candidate motion information list, and 00 may be used to indicate the third item in the candidate motion information list. Alternatively, when L is 4, 1 may be used to indicate the first item in the candidate motion information list, 01 may be used to indicate the second item in the candidate motion information list, 001 may be used to indicate the third item in the candidate motion information list, and 000 may be used to indicate the fourth item in the candidate motion information list. Others may be inferred by analogy.
[0202] Regarding the method of obtaining the motion vector predictor by constructing the candidate motion information list of the current image block, refer to the above description of modes such as the AMVP mode, merge mode, MMVD mode, or skip mode. Details will not be described again here.
[0203] S702: Obtain the index value of the length of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector predictor and the motion vector target value of the current image block.
[0204] The index value of the length of the motion vector difference may be used to indicate one piece of candidate length information in a set of candidate length information.
[0205] A set of candidate length information may include at least two pieces of candidate length information or may include one piece of candidate length information.
[0206] One candidate length information may be used to indicate the length of one motion vector difference, and the candidate length information may be a length value or information used to derive a length value. The length may be represented by using the Euclidean distance, or the length may include the absolute value of the x component and the absolute value of the y component of the motion vector difference. Of course, the length may alternatively be represented by using another norm. This is not limited here.
[0207] Note that since the motion vector is a two-dimensional array, the motion vector difference may be represented by using a two-dimensional array. When the motion vector is a three-dimensional array, the motion vector difference may be represented by a three-dimensional array.
[0208] S703: Determine target length information from a set of candidate length information based on the index value of the length. The set of candidate length information includes only candidate length information of N motion vector differences, where N is a positive integer greater than 1 and less than 8.
[0209] N may be 4. In some possible implementations, index values of different lengths may indicate different lengths. For example, among the candidate length information of N motion vector differences, when the index value of the length is the first preset value, the length indicated by the target length information is 1 / 4 of the pixel length; when the index value of the length is the second preset value, the length indicated by the target length information is half of the pixel length; when the index value of the length is the third preset value, the length indicated by the target length information is 1 pixel length; or when the index value of the length is the fourth preset value, the length indicated by the target length information is 2 pixel lengths. At least one of these is included. Note that the first preset value to the fourth preset value may not be sequential, may be independent of each other, and are only used to distinguish different preset values. Of course, the first preset value to the fourth preset value may alternatively be sequential or may have a sequence attribute. In some possible implementations, the correspondence between the index value of the length and the length of the MVD can be shown in Table 4.
Table 4
[0210] "Pel" has the same meaning as "pixel". For example, 1 / 4 pel indicates a length of 1 / 4 of a pixel. For a similar explanation, refer to the explanation of Table 2.
[0211] MmvdDistance represents a value for obtaining the length of the MVD. For example, the value of the length of the MVD can be obtained by shifting MmvdDistance 2 bits to the right.
[0212] S704: Obtain the motion vector difference of the current image block based on the target length information.
[0213] The method may further include obtaining an index value of the direction of the motion vector difference of the current image block, and determining target direction information from M pieces of candidate direction information of the motion vector difference based on the index value of the direction, where M is a positive integer greater than 1.
[0214] The M pieces of candidate direction information of the motion vector difference may be M pieces of candidate direction information.
[0215] The index value of the direction of the motion vector difference may be used to indicate one piece of candidate direction information among the M pieces of candidate direction information of the motion vector difference.
[0216] One piece of candidate direction information may be used to indicate the direction of one motion vector difference. Specifically, the candidate direction information may be a sign indicating a plus sign or a minus sign, and the sign may be the sign of the x component of the motion vector difference, or the sign of the y component of the motion vector difference, or the signs of the x component and the y component of the motion vector. Alternatively, the candidate direction information may be information used to derive the sign.
[0217] In a specific implementation process, the step of obtaining the motion vector difference of the current image block based on the target length information may include the step of determining the motion vector difference of the current image block based on the target direction information and the target length information.
[0218] S705: Determine the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block.
[0219] In some possible implementations, the sum of the motion vector difference of the current image block and the motion vector predictor of the current image block may be used as the motion vector target value of the current image block.
[0220] S706: Obtain the prediction block of the current image block based on the motion vector target value of the current image block.
[0221] For details of the inter prediction, refer to the foregoing description. Here, it will not be described in detail again.
[0222] FIG. 8 is a schematic flowchart of an inter prediction method according to an embodiment of the present invention. The method may be executed by a source device 12, a video coding system 40, an encoder 20, a video coding device 400, or a coding device 500. Specifically, the method may be executed by a prediction processing unit 260 (or, for example, an inter prediction unit 244 within the prediction processing unit 260) within the encoder 30. The method may include the following steps.
[0223] S801: Obtain the motion vector predictor of the current image block.
[0224] For the process, refer to the foregoing description of the AMVP mode, merge mode, MMVD mode, or skip mode. Here, it will not be described in detail again.
[0225] S802: Perform a motion search in the region at the position indicated by the motion vector predictor of the current image block to obtain the motion vector target value of the current image block.
[0226] S803: Based on the motion vector target value of the current image block and the motion vector predictor of the current image block, obtain the index value of the length of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector predictor of the current image block and the motion vector target value, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information. The set of candidate length information includes only candidate length information for N motion vector differences, and N is a positive integer greater than 1 and less than 8.
[0227] Based on the motion vector target value of the current image block and the motion vector predictor of the current image block, the step of obtaining the index value of the length of the motion vector difference of the current image block may include the step of obtaining the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block, and the step of determining, based on the motion vector difference of the current image block, the index value of the length of the motion vector difference of the current image block and the index value of the direction of the motion vector difference of the current image block.
[0228] N may be 4.
[0229] FIG. 8 illustrates an encoder-side method corresponding to the decoder-side method described in FIG. 7. For related descriptions, refer to FIG. 7 or the related descriptions above. Details will not be described again here.
[0230] FIG. 9 is a schematic flowchart of an inter prediction method according to an embodiment of the present invention. The method may be executed by the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. Specifically, the method may be executed by the video decoder 30, or specifically, by the entropy decoding unit 304 and the prediction processing unit 360 (or, for example, the inter prediction unit 344 within the prediction processing unit 360). The method may include the following steps.
[0231] S901: Obtain the motion vector predictor of the current image block.
[0232] In a specific implementation process, the step of obtaining the motion vector prediction factor of the current image block may include the step of constructing a candidate motion information list of the current image block, where the candidate motion information list includes L motion vectors, and L is 1, 3, 4, or 5; the step of obtaining the index value of the prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes the motion vector prediction factor; and the step of obtaining the motion vector prediction factor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list. The candidate motion information list of the current image block may be a merge candidate motion information list. Correspondingly, the inter prediction method provided in this embodiment of the present invention may be applied to the MMVD mode.
[0233] For the method of obtaining the motion vector prediction factor by constructing the candidate motion information list of the current image block, please refer to the above description of modes such as the AMVP mode, merge mode, MMVD mode, or skip mode. Details will not be described again here.
[0234] S902: Obtain the index value of the direction of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector prediction factor and the target value of the motion vector of the current image block.
[0235] The index value of the direction of the motion vector difference may be used to indicate one piece of candidate direction information in a set of candidate direction information.
[0236] A set of candidate direction information may include at least two pieces of candidate direction information or may include one piece of candidate direction information.
[0237] One candidate direction information may be used to indicate the direction of one motion vector difference. Specifically, the candidate direction information may be a sign indicating a plus sign or a minus sign, and the sign may be the sign of the x component of the motion vector difference, or the sign of the y component of the motion vector difference, or the signs of the x component and the y component of the motion vector. Alternatively, the candidate direction information may be information used to derive the sign.
[0238] Since the motion vector is a two-dimensional array, the motion vector difference may be represented by using a two-dimensional array. When the motion vector is a three-dimensional array, the motion vector difference may be represented by a three-dimensional array.
[0239] S903: Determine target direction information from a set of candidate direction information based on the index value of the direction. The set of candidate direction information includes candidate direction information of M motion vector differences, and M is a positive integer greater than 4.
[0240] M may be 8. In some possible implementations, index values in different directions may indicate different directions. For example, for the candidate direction information of M motion vector differences, when the index value of the direction is the first preset value, the direction indicated by the target direction information is exactly right; when the index value of the direction is the second preset value, the direction indicated by the target direction information is exactly left; when the index value of the direction is the third preset value, the direction indicated by the target direction information is exactly down; when the index value of the direction is the fourth preset value, the direction indicated by the target direction information is exactly up; when the index value of the direction is the fifth preset value, the direction indicated by the target direction information is bottom - right; when the index value of the direction is the sixth preset value, the direction indicated by the target direction information is top - right; when the index value of the direction is the seventh preset value, the direction indicated by the target direction information is bottom - left; or when the index value of the direction is the eighth preset value, the direction indicated by the target direction information is top - left, and may include at least one of these. Note that the first preset value to the eighth preset value may not be sequential, are independent of each other, and are only used to distinguish different preset values. Of course, the first preset value to the eighth preset value may alternatively be sequential or have a sequence attribute.
[0241] In some possible implementations, the correspondence between the index value of the direction and the direction of the MVD may be shown in Table 5, Table 6, or Table 7.
Table 5
Table 6
Table 7
[0242] In Table 5, when the value on the x-axis is "+", it may indicate that the direction of the MVD is the positive direction of the x-axis. When the value on the y-axis is "+", it may indicate that the direction of the MVD is the positive direction of the y-axis. When the value on the x-axis is "-", it may indicate that the direction of the MVD is the negative direction of the x-axis. When the value on the y-axis is "-", it may indicate that the direction of the MVD is the negative direction of the y-axis. When the value on the x-axis is N / A, it may indicate that the direction of the MVD is independent of the direction of the x-axis. When the value on the y-axis is N / A, it may indicate that the direction of the MVD is independent of the direction of the y-axis. When the values on both the x-axis and the y-axis are "+", it may indicate that the direction of the MVD with a projection in the x-axis direction is the positive direction, and the direction of the MVD with a projection in the y-axis direction is also the positive direction. When the values on both the x-axis and the y-axis are "-", it may indicate that the direction of the MVD with a projection in the x-axis direction is the negative direction, and the direction of the MVD with a projection in the y-axis direction is also the negative direction. When the value on the x-axis is "+" and the value on the y-axis is "-", it may indicate that the direction of the MVD with a projection in the x-axis direction is the positive direction, and the direction of the MVD with a projection in the y-axis direction is the negative direction. When the value on the x-axis is "-" and the value on the y-axis is "+", it may indicate that the direction of the MVD with a projection in the x-axis direction is the negative direction, and the direction of the MVD with a projection in the y-axis direction is the positive direction. The positive direction of the x-axis may indicate the left direction, and the positive direction of the y-axis may indicate the downward direction.
[0243] In Table 6 or Table 7, the x-axis may represent the sign coefficient of the component x of the MVD, and the product of the sign coefficient of the component x of the MVD and the absolute value of the component x is the component x of the MVD. The y-axis may represent the sign coefficient of the component y of the MVD, and the product of the sign coefficient of the component y and the absolute value of the component y is the component y of the MVD.
[0244] S904: Obtain the motion vector difference of the current image block based on the target direction information.
[0245] The method may further include obtaining an index value of the length of the motion vector difference of the current image block, and determining target length information from N pieces of candidate length information of the motion vector difference based on the index value of the length, where N is a positive integer greater than 1.
[0246] The N pieces of candidate length information of the motion vector difference may be a set of candidate length information in the embodiment of FIG. 7, specifically, a set of candidate length information provided in Table 4. For the length of the motion vector difference, the index value of the length of the motion vector difference, the candidate length information, and the description of the motion vector difference, refer to FIG. 7 or the foregoing description. Details are not described herein again.
[0247] In a specific implementation process, the step of obtaining the motion vector difference of the current image block based on the target direction information may include determining the motion vector difference of the current image block based on the target direction information and the target length information.
[0248] S905: Determine the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block.
[0249] In some possible implementations, the sum of the motion vector difference of the current image block and the motion vector predictor of the current image block may be used as the motion vector target value of the current image block.
[0250] S906: Obtain the prediction block of the current image block based on the motion vector target value of the current image block.
[0251] For details of inter prediction, refer to the foregoing description. Details are not described herein again.
[0252] For the content of the embodiment in FIG. 9, which is the same as the content in FIG. 7 and the foregoing description, please refer to FIG. 7 and the foregoing description. Details will not be described again here.
[0253] FIG. 10 is a schematic flowchart of an inter prediction method according to an embodiment of the present invention. The method may be executed by a source device 12, a video coding system 40, an encoder 20, a video coding device 400, or a coding device 500. Specifically, the method may be executed by a prediction processing unit 260 (or, for example, an inter prediction unit 244 within the prediction processing unit 260) within the encoder 30. The method may include the following steps.
[0254] S1001: Obtain the motion vector predictor of the current image block.
[0255] For the process, please refer to the foregoing description of the AMVP mode, merge mode, MMVD mode, or skip mode. Details will not be described again here.
[0256] S1002: Perform a motion search in the region at the position indicated by the motion vector predictor of the current image block to obtain the motion vector target value of the current image block.
[0257] S1003: Based on the motion vector target value of the current image block and the motion vector predictor of the current image block, obtain the index value of the direction of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector predictor of the current image block and the motion vector target value, and the index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information. The set of candidate direction information includes M pieces of candidate length information of the motion vector difference, and M is a positive integer greater than 4.
[0258] Based on the motion vector target value of the current image block and the motion vector predictor of the current image block, the step of obtaining the index value of the direction of the motion vector difference of the current image block may include the step of obtaining the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block, and the step of determining, based on the motion vector difference of the current image block, the index value of the length of the motion vector difference of the current image block and the index value of the direction of the motion vector difference of the current image block.
[0259] M may be 8.
[0260] FIG. 10 illustrates an encoder-side method corresponding to the decoder-side method described in FIG. 9. For related descriptions, refer to FIG. 9 or the related descriptions above. Details are not described again here.
[0261] FIG. 11 is a schematic flowchart of an inter prediction method according to an embodiment of the present invention. The method may be executed by the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. Specifically, the method may be executed by the video decoder 30, or specifically, by the entropy decoding unit 304 and the prediction processing unit 360 (or, for example, the inter prediction unit 344 within the prediction processing unit 360). The method may include the following steps.
[0262] S1101: Obtain a first motion vector predictor of the current image block and a second motion vector predictor of the current image block. The first motion vector predictor corresponds to the first reference frame, and the second motion vector predictor corresponds to the second reference frame.
[0263] S1102: Obtain the first motion vector difference of the current image block. The first motion vector difference of the current image block is used to indicate the difference between the first motion vector predictor and the first motion vector target value of the current image block, and the first motion vector target value and the first motion vector predictor correspond to the same reference frame.
[0264] S1103: Determine the second motion vector difference of the current image block based on the first motion vector difference. The second motion vector difference of the current image block is used to indicate the difference between the second motion vector predictor and the second motion vector target value of the current image block, and the second motion vector target value and the second motion vector predictor correspond to the same reference frame. When the direction of the first reference frame with respect to the current frame in which the current image block is located is the same as the direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference, or when the direction of the first reference frame with respect to the current frame in which the current image block is located is opposite to the direction of the second reference frame with respect to the current frame, the plus or minus sign of the second motion vector difference is opposite to the plus or minus sign of the first motion vector difference, and the absolute value of the second motion vector difference is the same as the absolute value of the first motion vector difference.
[0265] S1104: Determine the first motion vector target value of the current image block based on the first motion vector difference and the first motion vector predictor.
[0266] The first motion vector target value may be the sum of the first motion vector difference and the first motion vector predictor.
[0267] S1105: Determine the second motion vector target value of the current image block based on the second motion vector difference and the second motion vector predictor.
[0268] The second motion vector target value may be the sum of the second motion vector difference and the second motion vector predictor.
[0269] S1106: Obtain a predicted block of the current image block based on the first motion vector target value and the second motion vector target value.
[0270] For the content of the embodiment in FIG. 11, which is the same as the content in the foregoing description, please refer to the foregoing description. Details will not be described again here.
[0271] Based on the same inventive concept as the foregoing method, as shown in FIG. 12, an embodiment of the present invention further provides an inter prediction apparatus 1200. The inter prediction apparatus 1200 may be the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500, or may be a component of the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. Alternatively, the inter prediction apparatus 1200 may include an entropy decoding unit 304 and a prediction processing unit 360 (or, for example, an inter prediction unit 344 within the prediction processing unit 360). The inter prediction apparatus 1200 includes an acquisition unit 1201 and a prediction unit 1202. The acquisition unit 1201 and the prediction unit 1202 may be implemented by using software. For example, the acquisition unit 1201 and the prediction unit 1202 may be software modules, or the acquisition unit 1201 and the prediction unit 1202 may be a processor and a memory that execute instructions. Alternatively, the acquisition unit 1201 and the prediction unit 1202 may be implemented by using hardware. For example, the acquisition unit 1201 and the prediction unit 1202 may be modules within a chip.
[0272] The prediction unit 1202 may be configured to obtain a motion vector predictor of the current image block.
[0273] The acquisition unit 1201 may be configured to acquire an index value of the length of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector predictor and the target value of the motion vector of the current image block.
[0274] In some possible implementations, the acquisition unit 1201 may include an entropy decoding unit 304 configured to acquire an index value of the length of the motion vector difference of the current image block or an index value of the direction of the motion vector difference of the current image block. The prediction unit 1202 may include a prediction unit 360, specifically, may include an inter prediction unit 344.
[0275] The prediction unit 1202 is further configured to determine target length information from a set of candidate length information based on the index value of the length, where the set of candidate length information includes only candidate length information of N motion vector differences, and N is a positive integer greater than 1 and less than 8, to determine, to acquire the motion vector difference of the current image block based on the target length information, to determine the target value of the motion vector of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block, and to acquire the prediction block of the current image block based on the target value of the motion vector of the current image block.
[0276] The acquisition unit 1201 may further be configured to acquire an index value of the direction of the motion vector difference of the current image block. Correspondingly, the prediction unit 1202 may further be configured to determine target direction information from M candidate direction information of motion vector differences based on the index value of the direction, where M is a positive integer greater than 1. After the target direction information is acquired, the prediction unit 1202 may be configured to determine the motion vector difference of the current image block based on the target direction information and the target length information.
[0277] N may be, for example, 4. In some possible implementations, the candidate length information of the N motion vector differences may include at least one of the following: when the index value of the length is the first preset value, the length indicated by the target length information is 1 / 4 of the pixel length; when the index value of the length is the second preset value, the length indicated by the target length information is half of the pixel length; when the index value of the length is the third preset value, the length indicated by the target length information is 1 pixel length; or when the index value of the length is the fourth preset value, the length indicated by the target length information is 2 pixel lengths.
[0278] The prediction unit 1202 is configured to construct a candidate motion information list for the current image block, where the candidate motion information list may include L motion vectors, and L may be 1, 3, 4, or 5, and to obtain an index value of the prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes a motion vector prediction factor, and to obtain a motion vector prediction factor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list.
[0279] Specifically, it can be understood that the functions of the units of the inter prediction apparatus 1200 in this embodiment can be implemented according to the methods in the embodiments of the foregoing inter prediction method. For the specific implementation process, please refer to the relevant descriptions in the embodiments of the foregoing method. Details will not be described again here.
[0280] Based on the same inventive concept as the foregoing method, as shown in FIG. 13, an embodiment of the present invention further provides an inter prediction apparatus 1300. The inter prediction apparatus 1300 may be the source device 12, the video coding system 40, the encoder 20, the video coding device 400, or the coding device 500, or may be a component of the source device 12, the video coding system 40, the encoder 20, the video coding device 400, or the coding device 500. Alternatively, the inter prediction apparatus 1300 may include a prediction processing unit 260 (or, for example, an inter prediction unit 244 within the prediction processing unit 260). The inter prediction apparatus 1300 includes an acquisition unit 1301 and a prediction unit 1302. The acquisition unit 1301 and the prediction unit 1302 may be implemented by using software. For example, the acquisition unit 1301 and the prediction unit 1302 may be software modules, or the acquisition unit 1301 and the prediction unit 1302 may be a processor and a memory that execute instructions. Alternatively, the acquisition unit 1301 and the prediction unit 1302 may be implemented by using hardware. For example, the acquisition unit 1301 and the prediction unit 1302 may be modules within a chip.
[0281] The acquisition unit 1301 may be configured to acquire a motion vector predictor of the current image block.
[0282] The prediction unit 1302 may be configured to perform a motion search in a region at a position indicated by the motion vector predictor of the current image block to acquire a motion vector target value of the current image block.
[0283] In some possible implementations, the acquisition unit 1301 and the prediction unit 1302 may be used as an implementation of the prediction processing unit 260.
[0284] The prediction unit 1302 may be further configured to obtain an index value of the length of the motion vector difference of the current image block based on the target value of the motion vector of the current image block and the motion vector prediction factor of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector prediction factor of the current image block and the target value of the motion vector. The index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information. The set of candidate length information includes only candidate length information for N motion vector differences, and N is a positive integer greater than 1 and less than 8.
[0285] The prediction unit 1302 may be configured to obtain the motion vector difference of the current image block based on the target value of the motion vector of the current image block and the motion vector prediction factor of the current image block, and determine, based on the motion vector difference of the current image block, an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block.
[0286] N may be, for example, 4.
[0287] Specifically, it can be understood that the functions of the units of the inter prediction device 1300 in this embodiment can be implemented according to the methods in the embodiments of the foregoing method. For the specific implementation process, please refer to the relevant descriptions in the embodiments of the foregoing method. Details will not be described again here.
[0288] Based on the same inventive concept as the foregoing method, as shown in FIG. 14, an embodiment of the present invention further provides an inter prediction device 1400. The inter prediction device 1400 may be the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500, or may be a component of the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. Alternatively, the inter prediction device 1400 may include an entropy decoding unit 304 and a prediction processing unit 360 (or, for example, an inter prediction unit 344 within the prediction processing unit 360). The inter prediction device 1400 includes an acquisition unit 1401 and a prediction unit 1402. The acquisition unit 1401 and the prediction unit 1402 may be implemented by using software. For example, the acquisition unit 1401 and the prediction unit 1402 may be software modules, or the acquisition unit 1401 and the prediction unit 1402 may be a processor and a memory that execute instructions. Alternatively, the acquisition unit 1401 and the prediction unit 1402 may be implemented by using hardware. For example, the acquisition unit 1401 and the prediction unit 1402 may be modules within a chip.
[0289] The prediction unit 1402 may be configured to obtain a motion vector predictor of the current image block.
[0290] The acquisition unit 1401 may be configured to obtain an index value in the direction of the motion vector difference of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector predictor and the motion vector target value of the current image block.
[0291] In some possible implementations, the acquisition unit 1401 may include an entropy decoding unit 304 configured to acquire an index value of the length of the motion vector difference of the current image block or an index value of the direction of the motion vector difference of the current image block. The prediction unit 1402 may include a prediction unit 360, specifically, may include an inter prediction unit 344.
[0292] The prediction unit 1402 may further be configured to determine target direction information from a set of candidate direction information based on the index value of the direction, where the set of candidate direction information includes candidate direction information of M motion vector differences, and M is a positive integer greater than 4; to acquire the motion vector difference of the current image block based on the target direction information; to determine the target value of the motion vector of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; and to acquire the prediction block of the current image block based on the target value of the motion vector of the current image block.
[0293] The acquisition unit 1401 may further be configured to acquire an index value of the length of the motion vector difference of the current image block. Correspondingly, the prediction unit 1402 may further be configured to determine target length information from N pieces of candidate length information of motion vector differences based on the index value of the length, where N is a positive integer greater than 1. After the target length information is acquired, the prediction unit 1402 may be configured to determine the motion vector difference of the current image block based on the target direction information and the target length information.
[0294] M may be, for example, 8. In some possible implementations, for the candidate direction information of the M motion vector differences, when the index value of the direction is the first preset value, the direction indicated by the target direction information is exactly right; when the index value of the direction is the second preset value, the direction indicated by the target direction information is exactly left; when the index value of the direction is the third preset value, the direction indicated by the target direction information is exactly down; when the index value of the direction is the fourth preset value, the direction indicated by the target direction information is exactly up; when the index value of the direction is the fifth preset value, the direction indicated by the target direction information is bottom right; when the index value of the direction is the sixth preset value, the direction indicated by the target direction information is top right; when the index value of the direction is the seventh preset value, the direction indicated by the target direction information is bottom left; or when the index value of the direction is the eighth preset value, the direction indicated by the target direction information is top left, and it may include at least one of these cases.
[0295] The prediction unit 1402 is configured to construct a candidate motion information list for the current image block, where the candidate motion information list may include L motion vectors, and L may be 1, 3, 4, or 5; to obtain the index value of the prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes a motion vector predictor; and to obtain the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list.
[0296] Specifically, it can be understood that the functions of the units of the inter prediction device 1400 in this embodiment can be implemented according to the methods in the embodiments of the foregoing inter prediction method. For the specific implementation process, please refer to the relevant descriptions in the embodiments of the foregoing method. Details will not be described again here.
[0297] Based on the same inventive concept as the foregoing method, as shown in FIG. 15, an embodiment of the present invention further provides an inter prediction device 1500. The inter prediction device 1500 may be the source device 12, the video coding system 40, the encoder 20, the video coding device 400, or the coding device 500, or may be a component of the source device 12, the video coding system 40, the encoder 20, the video coding device 400, or the coding device 500. Alternatively, the inter prediction device 1500 may include a prediction processing unit 260 (or, for example, an inter prediction unit 244 within the prediction processing unit 260). The inter prediction device 1500 includes an acquisition unit 1501 and a prediction unit 1502. The acquisition unit 1501 and the prediction unit 1502 may be implemented by using software. For example, the acquisition unit 1501 and the prediction unit 1502 may be software modules, or the acquisition unit 1501 and the prediction unit 1502 may be a processor and a memory that execute instructions. Alternatively, the acquisition unit 1501 and the prediction unit 1502 may be implemented by using hardware. For example, the acquisition unit 1501 and the prediction unit 1502 may be modules within a chip.
[0298] The acquisition unit 1501 may be configured to acquire a motion vector predictor of a current image block.
[0299] The prediction unit 1502 may be configured to perform a motion search in a region at a position indicated by the motion vector predictor of the current image block to obtain a motion vector target value of the current image block.
[0300] In some possible implementations, the acquisition unit 1501 and the prediction unit 1502 may be used as an implementation of the prediction processing unit 260.
[0301] The prediction unit 1502 may be further configured to obtain an index value of the direction of the motion vector difference of the current image block based on the target value of the motion vector of the current image block and the motion vector prediction factor of the current image block. The motion vector difference of the current image block is used to indicate the difference between the motion vector prediction factor of the current image block and the target value of the motion vector. The index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information. The set of candidate direction information includes candidate length information of M motion vector differences, and M is a positive integer greater than 4.
[0302] The prediction unit 1502 may be configured to obtain the motion vector difference of the current image block based on the target value of the motion vector of the current image block and the motion vector prediction factor of the current image block, and determine, based on the motion vector difference of the current image block, an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block.
[0303] M may be, for example, 8.
[0304] Specifically, it can be understood that the functions of the units of the inter prediction apparatus 1500 in this embodiment can be implemented according to the methods in the embodiments of the foregoing method. For the specific implementation process, please refer to the relevant descriptions in the embodiments of the foregoing method. Details are not described herein again.
[0305] Based on the same inventive concept as the foregoing method, as shown in FIG. 16, an embodiment of the present invention further provides an inter prediction apparatus 1600. The inter prediction apparatus 1600 may be the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500, or may be a component of the destination device 14, the video coding system 40, the decoder 30, the video coding device 400, or the coding device 500. Alternatively, the inter prediction apparatus 1600 may include an entropy decoding unit 304 and a prediction processing unit 360 (or, for example, an inter prediction unit 344 within the prediction processing unit 360). The inter prediction apparatus 1600 includes an acquisition unit 1601 and a prediction unit 1602. The acquisition unit 1601 and the prediction unit 1602 may be implemented by using software. For example, the acquisition unit 1601 and the prediction unit 1602 may be software modules, or the acquisition unit 1601 and the prediction unit 1602 may be a processor and a memory that execute instructions. Alternatively, the acquisition unit 1601 and the prediction unit 1602 may be implemented by using hardware. For example, the acquisition unit 1601 and the prediction unit 1602 may be modules within a chip.
[0306] The acquisition unit 1601 may be configured to acquire a first motion vector predictor of a current image block and a second motion vector predictor of the current image block. The first motion vector predictor corresponds to a first reference frame, and the second motion vector predictor corresponds to a second reference frame.
[0307] The acquisition unit 1601 may further be configured to acquire a first motion vector difference of the current image block. The first motion vector difference of the current image block is used to indicate a difference between the first motion vector predictor and a first motion vector target value of the current image block, and the first motion vector target value and the first motion vector predictor correspond to the same reference frame.
[0308] The prediction unit 1602 determines the second motion vector difference of the current image block based on the first motion vector difference, where the second motion vector difference of the current image block is used to indicate the difference between the second motion vector predictor and the second motion vector target value of the current image block. The second motion vector target value and the second motion vector predictor correspond to the same reference frame, and when the direction of the first reference frame with respect to the current frame in which the current image block is located is the same as the direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference, or when the direction of the first reference frame with respect to the current frame in which the current image block is located is opposite to the direction of the second reference frame with respect to the current frame, the plus or minus sign of the second motion vector difference is opposite to the plus or minus sign of the first motion vector difference, and the absolute value of the second motion vector difference is the same as the absolute value of the first motion vector difference, and determines, based on the first motion vector difference and the first motion vector predictor, the first motion vector target value of the current image block, determines, based on the second motion vector difference and the second motion vector predictor, the second motion vector target value of the current image block, and obtains the predicted block of the current image block based on the first motion vector target value and the second motion vector target value, and may be configured to perform.
[0309] In some possible implementations, the acquisition unit 1601 and the prediction unit 1602 may be used as an implementation of the prediction processing unit 360.
[0310] Specifically, it can be understood that the functions of the units of the inter prediction apparatus 1600 in this embodiment can be implemented according to the methods in the embodiments of the foregoing inter prediction method. For the specific implementation process, please refer to the relevant descriptions in the embodiments of the foregoing method. Details will not be described again here.
[0311] One of ordinary skill in the art can understand that the functions described in connection with the various illustrative logical blocks, modules, and algorithm steps disclosed and described herein can be implemented by hardware, software, firmware, or any combination thereof. When the functions described in connection with these illustrative logical blocks, modules, and steps are implemented in software, they may be stored on a computer-readable medium as one or more instructions or code, or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one location to another (e.g., in accordance with a communication protocol). Thus, a computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this application. A computer program product may include a computer-readable medium.
[0312] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other compact disc storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is considered to be a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carriers, signals, or other transient media, but rather, are meant to be non-transitory tangible storage media. As used herein, disk (both disk and disc) includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-ray disc. Disk typically magnetically reproduces data, and disc optically reproduces data with a laser. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0313] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated circuits or discrete logic circuits. Accordingly, the term "processor" as used herein may be any one of the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described in connection with the exemplary logic blocks, modules, and stages described herein may be provided within dedicated hardware modules and / or software modules configured for encoding and decoding, or incorporated within a combined codec. Additionally, these techniques may be implemented entirely with one or more circuits or logic elements.
[0314] The techniques of the present application may be implemented in various devices or apparatuses including a wireless transceiver, an integrated circuit (IC), or a set of ICs (e.g., a chipset). In the present application, various components, modules, or units are described to emphasize the functional aspects of the apparatuses configured to execute the disclosed techniques, but these are not necessarily implemented by different hardware units. In fact, as described above, the various units may be combined with a codec hardware unit along with appropriate software and / or firmware, or provided by interoperable hardware units (including one or more of the foregoing processors).
[0315] In the foregoing embodiments, each embodiment has its respective focus. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0316] The foregoing description is only a specific implementation of this application and is not intended to limit the protection scope of this application. Any modifications or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application shall be included in the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims. [Other conceivable items] (Item 1) A method of inter prediction, obtaining a motion vector predictor of a current image block; obtaining an index value of the length of the motion vector difference of the current image block, wherein the motion vector difference of the current image block is used to indicate the difference between the motion vector predictor and the target value of the motion vector of the current image block; determining target length information from a set of candidate length information based on the index value of the length, wherein the set of candidate length information includes candidate length information of only N motion vector differences, and N is a positive integer greater than 1 and less than 8; obtaining the motion vector difference of the current image block based on the target length information; determining the target value of the motion vector of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; obtaining a predicted block of the current image block based on the target value of the motion vector of the current image block A method comprising the steps of. (Item 2) The method further comprises: obtaining an index value of the direction of the motion vector difference of the current image block; determining target direction information from M candidate direction information of motion vector differences based on the index value of the direction, wherein M is a positive integer greater than 1; and further comprising: The step of obtaining the motion vector difference of the current image block based on the target length information is the step of determining the motion vector difference of the current image block based on the target direction information and the target length information having the method according to item 1. (Item 3) The method according to item 1 or 2, wherein N is 4. (Item 4) The candidate length information of the N motion vector differences is when the index value of the length is a first preset value, the length indicated by the target length information is 1 / 4 of the pixel length, when the index value of the length is a second preset value, the length indicated by the target length information is half of the pixel length, when the index value of the length is a third preset value, the length indicated by the target length information is 1 pixel length, or when the index value of the length is a fourth preset value, the length indicated by the target length information is 2 pixel lengths including at least one of the above, the method according to item 3. (Item 5) The step of obtaining the motion vector predictor of the current image block is the step of constructing a candidate motion information list of the current image block, the candidate motion information list including L motion vectors, where L is 1, 3, 4, or 5, the step of constructing, the step of obtaining the index value of the prediction information of the motion information of the current image block in the candidate motion information list, the prediction information of the motion information of the current image block including the motion vector predictor, the step of obtaining, the step of obtaining the motion vector predictor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list The method according to any one of items 1 to 4, having (Item 6) A method of inter prediction, comprising: Obtaining a motion vector predictor of a current image block; Performing a motion search in a region at a position indicated by the motion vector predictor of the current image block to obtain a motion vector target value of the current image block; Obtaining an index value of a length of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block, wherein the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor of the current image block and the motion vector target value, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information, and the set of candidate length information includes candidate length information for only N motion vector differences, and N is a positive integer greater than 1 and less than 8. A method comprising the above. (Item 7) The step of obtaining an index value of a length of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block includes: Obtaining a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector predictor of the current image block; Determining an index value of the length of the motion vector difference of the current image block and an index value of a direction of the motion vector difference of the current image block based on the motion vector difference of the current image block. Having The method according to item 6. (Item 8) The method according to item 6 or 7, wherein N is 4. (Item 9) A method for inter prediction, obtaining a motion vector predictor for a current image block; obtaining an index value of a direction of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block; determining target direction information from a set of candidate direction information based on the index value of the direction, where the set of candidate direction information includes candidate direction information of M motion vector differences and M is a positive integer greater than 4; obtaining the motion vector difference of the current image block based on the target direction information; determining the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; obtaining a prediction block of the current image block based on the motion vector target value of the current image block and a method comprising the steps. (Item 10) The method further comprises: obtaining an index value of a length of the motion vector difference of the current image block; determining target length information from N pieces of candidate length information of motion vector differences based on the index value of the length, where N is a positive integer greater than 1; and the step of obtaining the motion vector difference of the current image block based on the target direction information is: determining the motion vector difference of the current image block based on the target direction information and the target length information The method according to item 9 having the above steps. The method according to item 9. The method according to item 9. (Item 11) The method according to item 9 or 10, where M is 8. (Item 12) The candidate direction information of the M motion vector differences is When the index value in the above direction is a first preset value, the direction indicated by the above target direction information is exactly right, When the index value in the above direction is a second preset value, the direction indicated by the above target direction information is exactly left, When the index value in the above direction is a third preset value, the direction indicated by the above target direction information is exactly down, When the index value in the above direction is a fourth preset value, the direction indicated by the above target direction information is exactly up, When the index value in the above direction is a fifth preset value, the direction indicated by the above target direction information is lower right, When the index value in the above direction is a sixth preset value, the direction indicated by the above target direction information is upper right, When the index value in the above direction is a seventh preset value, the direction indicated by the above target direction information is lower left, or When the index value in the above direction is an eighth preset value, the direction indicated by the above target direction information is upper left The method according to item 11, including at least one of the above. (Item 13) The step of obtaining the motion vector predictor of the current image block is The step of constructing a candidate motion information list of the current image block, the candidate motion information list includes L motion vectors, and L is 1, 3, 4, or 5, and the step of constructing; The step of obtaining the index value of the prediction information of the motion information of the current image block in the candidate motion information list, the prediction information of the motion information of the current image block includes the motion vector predictor, and the step of obtaining; Obtaining the motion vector prediction factor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list The method according to any one of items 9 to 12, having . (Item 14) A method for inter prediction, comprising: Obtaining a motion vector prediction factor of a current image block; Performing a motion search in a region at a position indicated by the motion vector prediction factor of the current image block to obtain a motion vector target value of the current image block; Obtaining an index value of a direction of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block, wherein the motion vector difference of the current image block is used to indicate a difference between the motion vector prediction factor of the current image block and the motion vector target value, and the index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information in a preset set of candidate direction information, the set of candidate direction information includes candidate length information of M motion vector differences, and M is a positive integer greater than 4 A method comprising the above steps. (Item 15) The step of obtaining the index value of the length of the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block comprises: Obtaining the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block; Determining the index value of the length of the motion vector difference of the current image block and the index value of the direction of the motion vector difference of the current image block based on the motion vector difference of the current image block Having the above steps The method according to item 14. (Item 16) The method according to item 14 or 15, wherein M is 8. (Item 17) A method of inter prediction, obtaining a first motion vector predictor of a current image block and a second motion vector predictor of the current image block, wherein the first motion vector predictor corresponds to a first reference frame and the second motion vector predictor corresponds to a second reference frame; obtaining a first motion vector difference of the current image block, wherein the first motion vector difference of the current image block is used to indicate a difference between the first motion vector predictor and a first motion vector target value of the current image block, and the first motion vector target value and the first motion vector predictor correspond to the same reference frame; determining a second motion vector difference of the current image block based on the first motion vector difference, wherein the second motion vector difference of the current image block is used to indicate a difference between the second motion vector predictor and a second motion vector target value of the current image block, the second motion vector target value and the second motion vector predictor correspond to the same reference frame, and when a direction of the first reference frame with respect to a current frame in which the current image block is arranged is the same as a direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference, or when the direction of the first reference frame with respect to the current frame in which the current image block is arranged is opposite to the direction of the second reference frame with respect to the current frame, a plus sign or a minus sign of the second motion vector difference is opposite to a plus sign or a minus sign of the first motion vector difference, and an absolute value of the second motion vector difference is the same as an absolute value of the first motion vector difference. Determining a first motion vector target value of the current image block based on the first motion vector difference and the first motion vector predictor; Determining a second motion vector target value of the current image block based on the second motion vector difference and the second motion vector predictor; Obtaining a prediction block of the current image block based on the first motion vector target value and the second motion vector target value A method comprising the steps of. (Item 18) An apparatus for inter prediction, the apparatus comprising: A prediction unit configured to obtain a motion vector predictor of a current image block; An acquisition unit configured to obtain an index value of a length of a motion vector difference of the current image block, wherein the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block; Comprising: The prediction unit further: Determining target length information from a set of candidate length information based on the index value of the length, wherein the set of candidate length information includes candidate length information of only N motion vector differences, and N is a positive integer greater than 1 and less than 8; Obtaining the motion vector difference of the current image block based on the target length information; Determining the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; Obtaining a prediction block of the current image block based on the motion vector target value of the current image block Configured to perform: Apparatus. (Item 19) The acquisition unit is further configured to acquire an index value in the direction of the motion vector difference of the current image block. The prediction unit is further configured to determine target direction information from candidate direction information of M motion vector differences based on the index value in the above direction, where M is a positive integer greater than 1. The prediction unit is configured to determine the motion vector difference of the current image block based on the target direction information and the target length information. The apparatus according to item 18. (Item 20) The apparatus according to item 18 or 19, where N is 4. (Item 21) The candidate length information of the N motion vector differences is When the index value of the length is a first preset value, the length indicated by the target length information is 1 / 4 of the pixel length. When the index value of the length is a second preset value, the length indicated by the target length information is half of the pixel length. When the index value of the length is a third preset value, the length indicated by the target length information is 1 pixel length, or When the index value of the length is a fourth preset value, the length indicated by the target length information is 2 pixel lengths. The apparatus according to item 20, including at least one of the above. (Item 22) The prediction unit is Constructing a candidate motion information list of the current image block, where the candidate motion information list includes L motion vectors, and L is 1, 3, 4, or 5. Acquiring an index value of prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes the motion vector predictor. Obtaining the motion vector prediction factor based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list The apparatus according to any one of items 18 to 21, which is configured to perform the above. (Item 23) An apparatus for inter prediction, the apparatus comprising An acquisition unit configured to obtain a motion vector prediction factor of a current image block; A prediction unit configured to perform a motion search in a region at a position indicated by the motion vector prediction factor of the current image block to obtain a motion vector target value of the current image block and comprising The prediction unit is further configured to obtain an index value of a length of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector prediction factor of the current image block and the motion vector target value, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information in a preset set of candidate length information, the set of candidate length information includes candidate length information for only N motion vector differences, and N is a positive integer greater than 1 and less than 8, and is configured to perform the obtaining. Apparatus. (Item 24) The prediction unit obtains the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block, and determines the index value of the length of the motion vector difference of the current image block and the index value of the direction of the motion vector difference of the current image block based on the motion vector difference of the current image block The apparatus according to item 23, configured as such. (Item 25) The apparatus according to item 23 or 24, where N is 4. (Item 26) An apparatus for inter prediction, the apparatus comprising: A prediction unit configured to obtain a motion vector predictor of a current image block; An acquisition unit configured to obtain an index value in a direction of a motion vector difference of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector predictor and a motion vector target value of the current image block; Comprising: The prediction unit is further configured to determine target direction information from a set of candidate direction information based on the index value in the direction, where the set of candidate direction information includes candidate direction information of M motion vector differences, and M is a positive integer greater than 4; to obtain the motion vector difference of the current image block based on the target direction information; to determine the motion vector target value of the current image block based on the motion vector difference of the current image block and the motion vector predictor of the current image block; and to obtain a prediction block of the current image block based on the motion vector target value of the current image block. Apparatus. (Item 27) The acquisition unit is further configured to obtain an index value of a length of the motion vector difference of the current image block; The prediction unit is further configured to determine target length information from N pieces of candidate length information of motion vector differences based on the index value of the length, where N is a positive integer greater than 1. The prediction unit Determines the motion vector difference of the current image block based on the target direction information and the target length information. Configured as such. The apparatus according to item 26. (Item 28) The apparatus according to item 26 or 27, where M is 8. (Item 29) The candidate direction information of the M motion vector differences is when the index value in the above direction is a first preset value, the direction indicated by the target direction information is exactly right, when the index value in the above direction is a second preset value, the direction indicated by the target direction information is exactly left, when the index value in the above direction is a third preset value, the direction indicated by the target direction information is exactly down, when the index value in the above direction is a fourth preset value, the direction indicated by the target direction information is exactly up, when the index value in the above direction is a fifth preset value, the direction indicated by the target direction information is bottom right, when the index value in the above direction is a sixth preset value, the direction indicated by the target direction information is top right, when the index value in the above direction is a seventh preset value, the direction indicated by the target direction information is bottom left, or when the index value in the above direction is an eighth preset value, the direction indicated by the target direction information is top left The apparatus according to item 28, including at least one of the above. (Item 30) The prediction unit is to construct a candidate motion information list of the current image block, where the candidate motion information list includes L motion vectors and L is 1, 3, 4, or 5, and to obtain an index value of prediction information of the motion information of the current image block in the candidate motion information list, where the prediction information of the motion information of the current image block includes the motion vector predictor. Based on the candidate motion information list and the index value of the motion information of the current image block in the candidate motion information list, obtaining the motion vector prediction factor The apparatus according to any one of items 26 to 29, which is configured to perform the above operations (Item 31) An apparatus for inter prediction, the apparatus comprising An acquisition unit configured to obtain a motion vector prediction factor of a current image block A prediction unit configured to perform a motion search in a region at a position indicated by the motion vector prediction factor of the current image block to obtain a motion vector target value of the current image block and The prediction unit is further configured to obtain an index value of a direction of a motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block, where the motion vector difference of the current image block is used to indicate a difference between the motion vector prediction factor and the motion vector target value of the current image block, the index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information in a preset set of candidate direction information, the set of candidate direction information includes candidate length information of M motion vector differences, and M is a positive integer greater than 4 Apparatus (Item 32) The prediction unit obtains the motion vector difference of the current image block based on the motion vector target value of the current image block and the motion vector prediction factor of the current image block, and determines an index value of the length of the motion vector difference of the current image block and an index value of the direction of the motion vector difference of the current image block based on the motion vector difference of the current image block The apparatus according to item 31, configured as such. (Item 33) The apparatus according to item 31 or 32, where M is 8. (Item 34) An inter-prediction apparatus comprising an acquisition unit and a prediction unit, where the acquisition unit is configured to acquire a first motion vector predictor of a current image block and a second motion vector predictor of the current image block, the first motion vector predictor corresponds to a first reference frame, the second motion vector predictor corresponds to a second reference frame, and the acquisition unit is further configured to acquire a first motion vector difference of the current image block, where the first motion vector difference of the current image block is used to indicate a difference between the first motion vector predictor and a first motion vector target value of the current image block, and the first motion vector target value and the first motion vector predictor correspond to the same reference frame. The prediction unit determines a second motion vector difference of the current image block based on the first motion vector difference, where the second motion vector difference of the current image block is used to indicate a difference between the second motion vector predictor and a second motion vector target value of the current image block. The second motion vector target value and the second motion vector predictor correspond to the same reference frame, and when a direction of the first reference frame with respect to a current frame in which the current image block is disposed is the same as a direction of the second reference frame with respect to the current frame, the second motion vector difference is the first motion vector difference, or when the direction of the first reference frame with respect to the current frame in which the current image block is disposed is opposite to the direction of the second reference frame with respect to the current frame, a plus sign or a minus sign of the second motion vector difference is opposite to a plus sign or a minus sign of the first motion vector difference, and an absolute value of the second motion vector difference is the same as an absolute value of the first motion vector difference. The prediction unit is configured to determine the second motion vector difference, determine a first motion vector target value of the current image block based on the first motion vector difference and the first motion vector predictor, determine a second motion vector target value of the current image block based on the second motion vector difference and the second motion vector predictor, and obtain a prediction block of the current image block based on the first motion vector target value and the second motion vector target value. Apparatus. (Item 35) A video decoder configured to decode a bitstream to obtain an image block, An inter prediction apparatus according to any one of Items 18 to 22, the inter prediction apparatus being configured to obtain a prediction block of a current image block, A reconstruction module configured to reconstruct the current image block based on the prediction block A video decoder including the reconstruction module. (Item 36) A video encoder configured to encode an image block, An inter prediction device according to any one of items 23 to 25, wherein the inter prediction device is configured to obtain an index value of a length of a motion vector difference of the current image block based on a motion vector predictor of the current image block, and the index value of the length of the motion vector difference of the current image block is used to indicate one piece of candidate length information within a preset set of candidate length information, the inter prediction device; An entropy encoding module configured to encode the index value of the length of the motion vector difference of the current image block into a bitstream A video encoder comprising the same. (Item 37) A video decoder configured to decode a bitstream to obtain an image block, An inter prediction device according to any one of items 26 to 30, the inter prediction device being configured to obtain a prediction block of a current image block; A reconstruction module configured to reconstruct the current image block based on the prediction block; A video decoder comprising the same. (Item 38) A video encoder configured to encode an image block, An inter prediction device according to any one of items 31 to 33, wherein the inter prediction device is configured to obtain an index value of a direction of a motion vector difference of the current image block based on a motion vector predictor of the current image block, and the index value of the direction of the motion vector difference of the current image block is used to indicate one piece of candidate direction information within a preset set of candidate direction information, the inter prediction device; An entropy encoding module configured to encode the index value of the direction of the motion vector difference of the current image block into a bitstream A video encoder comprising the same. (Item 39) A video coding device comprising a non-volatile memory and a processor coupled to each other, wherein the processor calls program code stored in the memory and executes the method according to any one of Items 1 to 17.
Claims
1. 1. An encoding method performed by an encoder, the method comprising: obtaining a first motion vector predictor for a current image block and a second motion vector predictor for the current image block, the first motion vector predictor corresponding to a first reference frame and the second motion vector predictor corresponding to a second reference frame; obtaining a first motion vector difference for the current image block, the first motion vector difference for the current image block indicating a difference between the first motion vector predictor and a first motion vector target value for the current image block, the first motion vector target value and the first motion vector predictor corresponding to a same reference frame, the first motion vector difference being used to determine a second motion vector difference for the current image block, the second motion vector difference indicating a difference between the second motion vector predictor and a second motion vector target value for the current image block, the second motion vector target value and the second motion vector predictor corresponding to a same reference frame, and the second motion vector difference is the first motion vector difference if an orientation of the first reference frame with respect to a current frame in which the current image block is located is the same as an orientation of the second reference frame with respect to the current frame; obtaining a length index value of the first motion vector difference and a direction index value of the first motion vector difference; encoding the index value of the length of the first motion vector difference and the index value of the direction of the first motion vector difference into a bitstream; A method comprising:
2. The method of claim 1 , wherein the index value of the length of the first motion vector difference of the current image block is used to indicate one candidate length information within a preset set of candidate length information.
3. The method of claim 1 , wherein the index value of the direction of the first motion vector difference of the current image block is used to indicate one candidate direction information within a preset set of candidate direction information.
4. A decoding method performed by a decoder, the method comprising: obtaining a length index value of a first motion vector difference of a current image block and a direction index value of the first motion vector difference by parsing a bitstream; obtaining a first motion vector predictor for the current image block and a second motion vector predictor for the current image block, the first motion vector predictor corresponding to a first reference frame and the second motion vector predictor corresponding to a second reference frame; obtaining the first motion vector difference according to the index value of the length of the first motion vector difference and the index value of the direction of the first motion vector difference, the first motion vector difference of the current image block indicating a difference between the first motion vector predictor and a first motion vector target value of the current image block, the first motion vector target value and the first motion vector predictor corresponding to a same reference frame; determining a second motion vector difference of the current image block based on the first motion vector difference, the second motion vector difference of the current image block indicating a difference between the second motion vector predictor and a second motion vector target value of the current image block, the second motion vector target value and the second motion vector predictor corresponding to a same reference frame, and the second motion vector difference is the first motion vector difference if an orientation of the first reference frame with respect to a current frame in which the current image block is located is the same as an orientation of the second reference frame with respect to the current frame; determining the first motion vector target value for the current image block based on the first motion vector difference and the first motion vector predictor; determining the second motion vector target value for the current image block based on the second motion vector difference and the second motion vector predictor; obtaining a prediction block of the current image block based on the first motion vector target value and the second motion vector target value; A method comprising:
5. The method of claim 4 , wherein the index value of the length of the first motion vector difference of the current image block is used to indicate one candidate length information within a preset set of candidate length information.
6. The method of claim 4 , wherein the index value of the direction of the first motion vector difference of the current image block is used to indicate one candidate direction information within a preset set of candidate direction information.
7. At least one processor; a memory coupled to said at least one processor and storing instructions which, when executed by said at least one processor, cause said at least one processor to perform the method of any one of claims 1 to 3; An encoding device comprising:
8. At least one processor; a memory coupled to said at least one processor and storing instructions which, when executed by said at least one processor, cause said at least one processor to perform the method of any one of claims 4 to 6; A decoding device comprising:
9. A computer program which, when executed on a computer or on a processor, enables said computer or said processor to carry out the method according to any one of claims 1 to 6.
10. 1. A method for storing an encoded bitstream of video data, the method comprising: receiving the bitstream, the bitstream having a length index value of a first motion vector difference and a direction index value of the first motion vector difference of a current image block, the length index value of the first motion vector difference and the direction index value of the first motion vector difference are used to determine the first motion vector difference, the first motion vector difference indicating a difference between a first motion vector predictor and a first motion vector target value of the current image block, the first motion vector target value and the first motion vector predictor correspond to a same reference frame, and the first motion vector difference is determined based on the first motion vector difference of the current image block; and receiving a second motion vector difference for determining a second motion vector difference for the current image block, the second motion vector difference indicating a difference between a second motion vector predictor and a second motion vector target value for the current image block, the second motion vector difference indicating a difference between a second motion vector predictor and a second motion vector target value for the current image block, the second motion vector predictor corresponding to a same reference frame, the first motion vector predictor corresponding to a first reference frame, the second motion vector predictor corresponding to a second reference frame, and the second motion vector difference being the first motion vector difference if an orientation of the first reference frame with respect to a current frame in which the current image block is located is the same as an orientation of the second reference frame with respect to the current frame; storing the bitstream on a storage medium; A method comprising:
11. 1. A method for transmitting an encoded bitstream of video data, the method comprising: obtaining the bitstream, the bitstream having a length index value of a first motion vector difference of a current image block and a direction index value of the first motion vector difference, the length index value of the first motion vector difference and the direction index value of the first motion vector difference are used to determine the first motion vector difference, the first motion vector difference indicating a difference between a first motion vector predictor and a first motion vector target value of the current image block, the first motion vector target value and the first motion vector predictor correspond to a same reference frame, and the first motion vector difference is determined based on the first motion vector difference of the current image block. and determining a second motion vector difference for the current image block, the second motion vector difference indicating a difference between a second motion vector predictor and a second motion vector target value for the current image block, the second motion vector difference indicating a difference between a second motion vector predictor and a second motion vector target value for the current image block, the second motion vector predictor corresponding to a same reference frame, the first motion vector predictor corresponding to a first reference frame, the second motion vector predictor corresponding to a second reference frame, and the second motion vector difference being the first motion vector difference if an orientation of the first reference frame with respect to a current frame in which the current image block is located is the same as an orientation of the second reference frame with respect to the current frame; transmitting the bitstream; A method comprising:
Citation Information
Patent Citations
Video encoding device and video decoding device using high-precision skip encoding and method thereof
US20170339425A1
Method for encoding and decoding motion information and device for encoding and decoding motion information
WO2019054736A1
Systems and methods for performing motion vector prediction for video coding using motion vector predictor origins
WO2019151093A1
Encoding method and device thereof, and decoding method and device thereof
WO2019168244A1
Coding device, decoding device, coding method, and decoding method
WO2020017367A1