Method and apparatus for decoder-side motion vector refinement in video coding and decoding

By deriveing the initial motion vector in video encoding and decoding and adjusting the generation value, and optimizing the motion vector refinement process, the problem of low efficiency of high-resolution video encoding and decoding is solved, and more efficient video image reconstruction is achieved.

CN114080808BActive Publication Date: 2025-08-01BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080048631.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-06
Filing Date
2020-07-03
Publication Date
2025-08-01
Estimated Expiration
2040-07-03

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies are used to process high-resolution videos, it is difficult to effectively improve the encoding and decoding efficiency and maintain image quality.

Method used

By derive the initial motion vector of the current codec unit, multiple motion vector candidates are generated, and the motion vector is refined by adjusting the generation value, and the motion vector refinement process is optimized, using the decoder-side motion vector refinement (DMVR) technology.

Benefits of technology

It improves the efficiency of video encoding and decoding, reduces the computational complexity in the motion vector refinement process, and improves the reconstruction quality of video images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114080808B_ABST
    Figure CN114080808B_ABST
Patent Text Reader

Abstract

A method for video coding and decoding is provided. The method includes: deriving an initial motion vector (MV) of a current coding unit (CU); deriving a plurality of MV candidates for decoder-side motion vector refinement (DMVR); determining a cost value for each of the initial MV and the MV candidates; receiving signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; obtaining updated cost values by adjusting at least one of the cost values based on the parameter; and deriving a refined MV based on the updated cost values of the initial MV and the MV candidates.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 871,130, filed on Jul. 6, 2019, entitled “Decoder - side Motion Vector Refinement for Video Coding”, the entire content of which is incorporated herein by reference for all purposes. Technical Field

[0003] This application generally relates to video encoding, decoding, and compression, and more particularly, but not limited to, methods and apparatuses for decoder - side motion vector refinement (DMVR) in video encoding and decoding. Background Art

[0004] Various electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc., support digital video. Electronic devices send, receive, encode, decode, and / or store digital video data by implementing video compression / decompression. Digital video devices implement video encoding and decoding techniques, such as those described in standards defined by Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), MPEG - 2, MPEG - 4, ITU - T H.263, ITU - T H.264 / MPEG - 4 Part 10, Advanced Video Coding (AVC), ITU - T H.265 / High Efficiency Video Coding (HEVC) and extensions of such standards.

[0005] Video encoding and decoding typically use prediction methods (e.g., inter - frame prediction, intra - frame prediction), which utilize the redundancy present in video images or sequences. An important goal of video encoding and decoding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. As evolving video services become available, there is a need for encoding techniques with better encoding efficiency.

[0006] Video compression typically includes performing spatial (intra) prediction and / or temporal (inter) prediction to reduce or remove redundancy inherent in video data. For block-based video coding and decoding, a video frame is segmented into one or more strips, each strip having a plurality of video blocks, which may also be referred to as coding tree units (CTUs). Using a quadtree with a nested multi-type tree structure, a CTU can be split into coding units (CUs), where a CU defines a pixel region sharing the same prediction mode. Each CTU can contain one coding unit (CU), or can be recursively split into smaller CUs until a predefined minimum CU size is reached. Each CU (also referred to as a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be coded in an intra mode, an inter mode, or an IBC mode. Video blocks in the intra-coded (I) strips of a video frame are coded using spatial prediction, which is with respect to reference samples in adjacent blocks within the same video frame. Video blocks in the inter-coded (P or B) strips of a video frame can use spatial prediction or temporal prediction, where spatial prediction is with respect to reference samples in adjacent blocks within the same video frame, and temporal prediction is with respect to reference samples in other previous reference video frames and / or future reference video frames.

[0007] In some examples of the present disclosure, the term "unit" defines an image region covering all components (such as luminance and chrominance); the term "block" is used to define a region covering a specific component (e.g., luminance), and when considering a chrominance sampling format (such as 4:2:0), blocks of different components (e.g., luminance and chrominance) can be different in spatial position.

[0008] A prediction block for a current video block to be coded and decoded is derived based on spatial prediction or temporal prediction of previously encoded reference blocks (e.g., adjacent blocks). The process of finding the reference blocks can be done through a block matching algorithm. Residual data, which represents the pixel difference between the current block to be coded and decoded and the prediction block, is referred to as a residual block or prediction error. Inter-coded blocks are coded based on a motion vector and the residual block, where the motion vector points to the reference block in the reference frame that forms the prediction block. The process of determining the motion vector is generally referred to as motion estimation. Intra-coded blocks are coded based on the intra-prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to the transform domain (e.g., frequency domain), resulting in residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, and then entropy-coded into a video bitstream to achieve even greater compression.

[0009] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., flash memory) for access by another electronic device having digital video capabilities or directly transmitted to the electronic device, either wired or wirelessly. The electronic device then performs video decompression (which is the reverse process of the video compression described above), e.g., by parsing the encoded video bitstream to obtain syntax elements from the bitstream, and reconstructing the digital video data from the encoded video bitstream into its original format based at least in part on the syntax elements obtained from the bitstream, and the electronic device presents the reconstructed digital video data on a display of the electronic device.

[0010] As the digital video quality changes from high definition to 4K×2K or even 8K×4K, the amount of video data to be encoded / decoded grows exponentially. There has been a long-standing challenge in how to encode / decode video data more efficiently while maintaining the image quality of the decoded video data. Summary of the Invention

[0011] Generally, the present disclosure describes examples of techniques related to decoder-side motion vector refinement (DMVR) in video coding and decoding.

[0012] According to a first aspect of the present disclosure, there is provided a method for video coding and decoding, including: deriving an initial motion vector (MV) of a current coding unit (CU); deriving a plurality of MV candidates for decoder-side motion vector refinement (DMVR); determining a cost value for each of the initial MV and the MV candidates; receiving signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; obtaining an updated cost value by adjusting at least one of the cost values based on the parameter; and deriving a refined MV based on the initial MV and the updated cost values of the MV candidates.

[0013] According to a second aspect of the present disclosure, there is provided a device for video coding and decoding, including: one or more processors; and a memory configured to store instructions executable by the one or more processors; wherein the one or more processors are configured, when executing the instructions, to: derive an initial motion vector (MV) of a current coding unit (CU); derive a plurality of MV candidates for decoder-side motion vector refinement (DMVR); determine a cost value for each of the initial MV and the MV candidates; receive signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; obtain an updated cost value by adjusting at least one of the cost values based on the parameter; and derive a refined MV based on the initial MV and the updated cost values of the MV candidates.

[0014] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, including instructions stored therein, wherein when the instructions are executed by one or more processors, the instructions cause the one or more processors to perform actions, the actions including: deriving an initial motion vector (MV) of a current coding unit (CU); deriving a plurality of MV candidates for decoder-side motion vector refinement (DMVR); determining a cost value for each of the initial MV and the MV candidates; receiving signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; obtaining updated cost values by adjusting at least one of the cost values based on the parameter; and deriving a refined MV based on the initial MV and the updated cost values of the plurality of MV candidates. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] A more specific description of the examples of the present disclosure will be presented by referring to specific examples shown in the accompanying drawings. Given that these drawings only depict some examples and are thus not considered to be a limitation of the scope, the examples will be described and explained by using the drawings with additional features and details.

[0016] Figure 1 is a block diagram showing an exemplary video encoder according to some embodiments of the present disclosure.

[0017] Figure 2 is a block diagram showing an exemplary video decoder according to some embodiments of the present disclosure.

[0018] Figure 3 is a schematic diagram showing an example of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure.

[0019] Figure 4 is a schematic diagram showing an example of a DMVR search process according to some embodiments of the present disclosure.

[0020] Figure 5 is a schematic diagram showing an example of a DMVR integer luminance sample search pattern according to some embodiments of the present invention.

[0021] Figure 6 is a block diagram showing an exemplary apparatus for video coding and decoding according to some embodiments of the present disclosure.

[0022] Figure 7 is a flowchart showing an exemplary process of decoder-side motion vector refinement (DMVR) in video coding and decoding according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to those of ordinary skill in the art that various alternatives may be used. For example, it will be apparent to those of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0024] The terms used in this disclosure are for the purpose of describing exemplary examples only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "or" and "and / or" as used herein are intended to mean and include any and all possible combinations of one or more of the associated listed items unless the context clearly dictates otherwise.

[0025] References throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Unless otherwise clearly stated, features, structures, elements, or characteristics described in connection with one or some embodiments are also applicable to other embodiments.

[0026] Throughout this disclosure, unless otherwise clearly stated, the terms "first," "second," "third," etc. are used as designations only for referring to related elements (e.g., devices, components, compositions, steps, etc.) and do not imply any spatial or temporal order. For example, "a first device" and "a second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be named arbitrarily.

[0027] As used herein, depending on the context, the terms "if" or "when" may be understood to mean "upon" or "in response to." These terms, if they appear in the claims, may not indicate that the related limitation or feature is conditional or optional.

[0028] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" may include a memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may be physically attached to each other or positioned adjacent to each other, or may not be physically attached to each other or positioned adjacent to each other.

[0029] A unit or module can be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module can include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.

[0030] Figure 1 FIG. shows a block diagram illustrating an exemplary block-based hybrid video encoder 100, which can be used in conjunction with many video coding and decoding standards that use block-based processing. In encoder 100, a video frame is segmented into a plurality of video blocks for processing. For each given video block, a prediction is formed based on an inter-frame prediction method or an intra-frame prediction method. In inter-frame prediction, one or more prediction values are formed based on pixels from a previous reconstructed frame through motion estimation and motion compensation. In intra-frame prediction, a prediction value is formed based on reconstructed pixels in the current frame. Through mode decision, the best prediction value can be selected to predict the current block.

[0031] The prediction residual, which represents the difference between the current video block and its prediction value, is sent to transformation circuitry 102. The transform coefficients are then sent from transformation circuitry 102 to quantization circuitry 104 for entropy reduction. The quantized coefficients are then fed to entropy coding / decoding circuitry 106 to generate a compressed video bitstream. As Figure 1 shown, prediction-related information 110 from the inter-frame prediction circuitry and / or the intra-frame prediction circuitry 112, such as video block segmentation information, motion vectors, reference picture indices, and intra-frame prediction modes, is also fed through entropy coding / decoding circuitry 106 and saved into the compressed video bitstream 114.

[0032] In encoder 100, decoder-related circuitry is also required for prediction purposes to reconstruct pixels. First, the prediction residual is reconstructed through inverse quantization circuitry 116 and inverse transform circuitry 118. The reconstructed prediction residual is combined with the block prediction value 120 to generate unfiltered reconstructed pixels for the current video block.

[0033] Spatial prediction (or "intra-frame prediction") uses pixels from decoded adjacent blocks (which are referred to as reference samples) in the same video frame as the current video block to predict the current video block.

[0034] Temporal prediction (also known as "inter - frame prediction") uses reconstructed pixels from already decoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. Typically, for a given coding unit (CU) or coding block, the temporal prediction signal is signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is also signaled, which is used to identify which reference picture in the reference picture buffer the temporal prediction signal is from.

[0035] After spatial and / or temporal prediction is performed, the intra / inter - frame mode decision circuitry 121 in the encoder 100 selects the best prediction mode, for example, based on a rate - distortion optimization method. The block prediction value 120 is then subtracted from the current video block; and the resulting prediction residual is decorrelated using the transform circuitry 102 and quantization circuitry 104. The resulting quantized residual coefficients are de - quantized by the inverse quantization circuitry 116 and inverse - transformed by the inverse transform circuitry 118 to form the reconstructed residual, which is then added back to the predicted block to form the reconstructed signal for the CU. Before the reconstructed CU is placed in the reference picture buffer 117 of the picture buffer and used to decode future video blocks, loop filtering 115, such as a de - block filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), may be further applied to the reconstructed CU. To form the output video bitstream 114, the coding mode (inter - frame or intra - frame), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 106 to be further compressed and packetized to form the bitstream.

[0036] For example, the de - block filter is available in the current versions of AVC, HEVC, and VVC. In HEVC, an additional loop filter called sample adaptive offset (SAO) is defined to further improve coding efficiency. Another loop filter called adaptive loop filter (ALF) is being actively researched.

[0037] These loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. They can also be turned off according to the decision made by the encoder 100 to save computational complexity.

[0038] It should be noted that intra - frame prediction is typically based on unfiltered reconstructed pixels, while inter - frame prediction is based on filtered reconstructed pixels (if the encoder 100 turns on these filter options).

[0039] Figure 2is a block diagram showing an exemplary block-based video decoder 200, which can be used in conjunction with many video coding and decoding standards. This decoder 200 is similar to the reconstruction-related part in the Figure 1 encoder 100 residing in. In decoder 200, first, the incoming video bitstream 201 is decoded by entropy decoding 202 to derive quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residuals. The block predictor value mechanism, which is implemented in the intra / inter mode selector 212, is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residuals from the inverse transform 206 and the prediction output generated by the block predictor value mechanism using an adder 214.

[0040] Before the reconstructed blocks are stored in the picture buffer 213 that serves as a reference picture repository, the reconstructed blocks can be further passed through the loop filter 209. The reconstructed video in the picture buffer 213 can be sent to drive a display device and for predicting future video blocks. When the loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0041] The video coding / decoding standards mentioned above (such as VVC, JEM, HEVC, MPEG-4 Part 10) are conceptually similar. For example, they all use block-based processing. In the Joint Video Exploration Team (JVET) meetings, JVET defined the first draft of the Versatile Video Coding (VVC) and the VVC Test Model 1 (VTM1) coding method. It was decided to include a quadtree with a nested multi-type tree using binary splitting and ternary splitting coding block structures as the initial new coding and decoding feature of VVC.

[0042] Decoder-side Motion Vector Refinement (DMVR) in VVC

[0043] Decoder-side motion vector refinement (DMVR) is a technique for blocks coded in the bi-predictive merge mode. In this mode, bilateral matching (BM) prediction can be used to further refine the two motion vectors (MVs) of the block.

[0044] Figure 3 is a schematic diagram showing an example of decoder-side motion vector refinement (DMVR). As Figure 3As shown, the bilateral matching method is used to refine the motion information of the current CU 322 by searching for the closest match between two reference blocks 302 and 312 of the current CU 322 along the motion trajectory of the current CU in two associated reference pictures of the current CU 322 (i.e., refPic in list L0 300 and refPic in list L1 310). Based on the initial motion information from the merge mode, the patterned rectangular blocks 322, 302, and 312 indicate the current CU and its two reference blocks. Based on the MV candidates used in the motion refinement search process (i.e., the motion vector refinement process), the patterned rectangular blocks 304, 314 indicate a pair of reference blocks.

[0045] The MV differences between the MV candidates (i.e., MV0’ and MV1’) and the initial MVs (also called the original MVs) (i.e., MV0 and MV1) are MVdiff and -MVdiff, respectively. Both the MV candidates and the initial MVs are bi-directional motion vectors. During DMVR, multiple such MV candidates around the initial MV can be examined. Specifically, for each given MV candidate, its two associated reference blocks can be located in the reference pictures in its list 0 and list 1, respectively, and the difference between them can be calculated.

[0046] The block difference can also be referred to as the cost value and is typically measured as the sum of absolute differences (SAD) or the SAD with line subsampling (i.e., the SAD calculated using the interlaced lines of the involved blocks). In some other examples, the mean-removed SAD or the sum of squared differences (SSD) can also be used as the cost value. The MV candidate with the lowest cost value or SAD between the two reference blocks becomes the refined MV and is used to generate the bi-directional prediction signal as the actual prediction for the current CU.

[0047] In VVC, DMVR is applied to CUs that satisfy the following conditions:

[0048] · The CU is encoded and decoded using the CU-level merge mode with bi-directional prediction MVs (not the sub-block merge mode);

[0049] · With respect to the current picture, one reference picture of the CU is in the past (i.e., having a POC less than the POC of the current picture) and the other reference picture is in the future (i.e., having a POC greater than the POC of the current picture);

[0050] · The POC distances from the two reference pictures to the current picture (i.e., the absolute POC differences) are the same; and

[0051] · The CU has more than 64 luma samples in size and more than 8 luma samples in height.

[0052] The refined MVs derived through the DMVR process are used to generate inter-prediction samples and are also used in temporal motion vector prediction for future picture coding. The original MVs are used in the deblocking process and are also used in spatial motion vector prediction for future CU coding.

[0053] Search Scheme in DMVR

[0054] As Figure 3 shown, MV candidates (or search points) surround the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point represented by a candidate MV pair (MV0’, MV1’) examined by DMVR follows the following two equations:

[0055] MV0′ = MV0 + MV diff

[0056] MV1′ = MV1 - MV diff ,

[0057] where MV diff represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. In the current VVC, the refinement search range is two integer luminance samples from the initial MV.

[0058] Figure 4 illustrates an example of the search process of DMVR. As Figure 4 shown, the search process includes an integer sample offset search stage 402 and a fractional sample refinement stage 404.

[0059] To reduce the search complexity, a fast search method with an early termination mechanism is applied in the integer sample offset search stage 402. Instead of a 25-point full search, a 2-iteration search scheme is applied to reduce the number of SAD check points. Figure 5 illustrates an example of the DMVR integer luminance sample search pattern for the integer sample offset search stage 402. Figure 5 Each rectangular box in Figure 5As shown, according to the fast search method, up to 6 SADs (SADs for the center and P1 to P5) are checked in the first iteration. In the first iteration, the initial MV is the center. First, the SADs of five points (the center and P1 to P4) are compared. If the SAD of the center (i.e., the center position) is the smallest, the integer pixel offset search stage 402 of the DMVR is terminated. Otherwise, one more position P5 is checked (determined based on the SAD distribution of P1 to P4). Then, the position with the smallest SAD among P1 to P5 is selected as the center position for the second iteration search. The process of the second iteration search is the same as that of the first iteration search. The SADs calculated in the first iteration can be reused in the second iteration, and thus only the SADs of 3 additional points may need to be calculated in the second iteration. Note that when the SAD of the center point in the first iteration is less than the number of pixels used to calculate the SAD (which is equal to w×h / 2, where w and h represent the width and height of the DMVR operation unit respectively), the entire DMVR process is terminated prematurely without further search.

[0060] After the integer pixel search 402 is the fractional pixel refinement 404. To reduce the computational complexity, the parametric error surface equation is used to derive the fractional pixel refinement 404, instead of using an additional search using SAD comparison. The fractional pixel refinement 404 is conditionally called based on the output of the integer pixel search stage. When the integer pixel search stage 402 is terminated with the center having the smallest SAD in the first iteration search or the second iteration search, the fractional pixel refinement is further applied.

[0061] In the fractional pixel refinement based on the parametric error surface, the SAD costs (or cost values) of the center position and its four adjacent positions are used to fit a two-dimensional parabolic error surface equation of the following form:

[0062] E(x, y) = A(x - x min ) 2 + B(y - y min ) 2 + C,

[0063] where (x min , y min ) corresponds to the fractional position with the smallest SAD cost, and C corresponds to the minimum cost value. By solving the above equation using the SAD cost values of five search points, (x min , y min ) can be derived by the following formula:

[0064] x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0))) (1)

[0065] y min = (E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0))) (2)

[0066] x min and y min The values of x and y are further constrained between -8 and 8, which corresponds to a half-pixel offset from the center point with 1 / 16 pixel MV accuracy. The calculated fractional offset (x min , y min ) is added to the integer distance MV refinement to obtain the sub-pixel accuracy MV refinement.

[0067] Bilinear Interpolation and Sample Padding for DMVR

[0068] In VVC, the resolution of the MV is 1 / 16 luminance samples. The samples at the fractional positions are interpolated using an 8-tap interpolation filter. In the DMVR search, when the candidate MV points to a sub-pixel position, those related fractional position samples need to be interpolated. To reduce the computational complexity, a bilinear interpolation filter is used during the search process in DMVR to generate the fractional samples.

[0069] Another effect of using the bilinear filter for interpolation is that, with a 2-sample search range, the DVMR search process does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. Again, during this 8-tap interpolation process, sample padding is used to avoid accessing more reference samples than the normal motion compensation process. More specifically, during the 8-tap interpolation process based on the refined MV, the samples beyond those required for the motion compensation based on the original MV will be filled from their adjacent available samples.

[0070] Maximum DMVR Processing Unit

[0071] When the width and / or height of the CU is greater than 16 luminance samples, the DMVR operation for the CU is performed based on DMVR processing units with a maximum width and / or height equal to 16 samples. In other words, in such a case, the original CU is divided into sub-blocks with a width and / or height equal to 16 luminance samples for the DMVR operation. The maximum processing unit size for the DMVR search process is limited to 16×16.

[0072] In the current VVC design, there is no control flag to control the enabling of DMVR. However, it is not guaranteed that the MV refined by DMVR is always better than the MV before refinement. In some cases, the DMVR refinement process may produce a refined MV that is worse than the original MV. According to some examples of the present disclosure, several methods are proposed to reduce the loss caused by such uncertainty of DMVR MV refinement.

[0073] Updated Cost Value for DMVR by Adjusting the Cost Value

[0074] Several exemplary methods are proposed to favor the original MV during the DMVR process. Note that these different methods can be applied independently or jointly.

[0075] In some examples of the present disclosure, the terms "initial MV" and "original MV" can be used interchangeably.

[0076] In some examples, during the DMVR process, the cost values for each MV candidate among the initial MV and MV candidates can be adjusted or updated to favor the initial MV. That is, after calculating the cost value (e.g., SAD) of the search point during the DMVR process, the (multiple) cost values can be adjusted to increase the probability that the initial MV has the minimum cost value among the updated cost values, i.e., to favor the initial MV.

[0077] Therefore, after obtaining the updated cost values, the initial MV is more likely to be selected as the MV with the lowest cost during the DMVR process.

[0078] Here, for illustrative purposes, the SAD value is used as an exemplary cost value. Other values (such as row subsampled SAD, mean removed SAD, or sum of squared differences (SSD)) can also be used as cost values.

[0079] In some examples, compared with the SAD values of other MV candidates, the SAD value between the reference blocks referred to by the initial MV (or original MV) is reduced by a first value Offset SAD , the first value Offset SAD is calculated by a predefined process. Therefore, the initial MV is favored over other candidate MVs because its SAD value is reduced.

[0080] In one example, the value of Offset SAD can be determined as 1 / N of the SAD value associated with the initial MV, where N is an integer (e.g., 4, 8, or 16).

[0081] In another example, the value of Offset SAD can be determined as a constant value M.

[0082] In another example, Offset SAD can be determined according to the warp decoding information in the current CU, and the warp decoding information includes at least one or a combination of the following: the codec block size, the magnitude of the motion vector, the SAD of the initial MV, and the relative position of the DMVR processing unit. For example, Offset SAD can be determined as 1 / N of the SAD value associated with the initial MV, where N is an integer value (such as 4, 8, or 16) selected based on the block size of the current CU. When the current block size is greater than or equal to a predefined size (e.g., 16×16), the value of N is set to 8; otherwise, the value of N is set to 4. For example, Offset SAD can be determined as 1 / N of the SAD value associated with the initial MV, where N is an integer value (such as 4, 8, or 16) selected based on the distance between the center position of the DMVR processing unit and the center position of the current CU. When the distance is greater than or equal to a predefined threshold, N is set to a value (e.g., 8); otherwise, N is set to another value (e.g., 4).

[0083] In these examples, it is described that the SAD value associated with the initial MV is reduced by a value Offset SAD . In practice, this concept can be implemented differently. For example, instead of reducing the SAD value associated with the initial MV, during the DMVR search process, the value of Offset SAD can be added to those SADs associated with other MV candidates, and the results in both cases are equivalent.

[0084] In some other examples, the SAD value between the reference blocks referred to by non-initial MV candidates is increased by a second value Offset SAD ’, and the second value Offset SAD ’ is calculated through a predefined process. The second value Offset SAD ’ and the first value Offset SAD can be the same or different. Therefore, the initial MV is advantageous because the SAD value of the non-initial MV is increased.

[0085] In one example, the value of Offset SAD ’ can be determined as 1 / N of the SAD value associated with the non-initial MV, where N is an integer (such as 4, 8, or 16).

[0086] In another example, the value of Offset SAD ’ can be determined as a constant value M.

[0087] In another example, Offset SADThe value of 'Offset' can be determined according to the warp decoding information in the current CU, which may include the codec block size, the magnitude of the motion vector, the SAD value of non-initial MVs, and / or the relative position of the DMVR processing unit within the current CU. For example, this value can be determined as 1 / N of the SAD value from the BM using non-initial MVs, where N is an integer selected based on the block size (e.g., 4, 8, or 16). When the current block size is greater than or equal to a predefined size (e.g., 16×16), the value of N is set to 8; otherwise, the value of N is set to 4. For example, Offset SAD The value of 'Offset' can be determined as 1 / N of the SAD value from the BM using non-initial MVs, where N is an integer value (e.g., 4, 8, or 16) selected based on the distance between the center position of the DMVR processing unit and the center position of the current CU. When the distance is greater than or equal to a predefined threshold, N is set to a value (e.g., 8); otherwise, N is set to another value (e.g., 4).

[0088] In these examples, it is described that the SAD value associated with a non-initial MV candidate is increased by a value Offset SAD '. In practice, this concept can be implemented differently. For example, instead of increasing the SAD value associated with a non-initial MV, during the DMVR search process, the value of Offset SAD ' can be subtracted from the SAD associated with the initial MV, and the result is equivalent.

[0089] In some other examples, based on an appropriate subset of samples used for SAD calculation associated with non-initial MVs, the BM SAD associated with the initial MV is calculated. That is, compared to the SAD value of the MV candidate, fewer samples are used to determine the SAD value of the initial MV. This can be similar to reducing the SAD value of the initial MV.

[0090] According to some examples of the present disclosure, a parameter can be signaled to the decoder for adjusting or updating the cost value for each MV candidate in the initial MV and / or MV candidates to favor the initial MV. The value of the parameter can be signaled in the bitstream in the sequence parameter set, picture parameter set, slice header, codec tree unit (CTU), and / or codec unit (CU).

[0091] In some examples, the parameter can be a value used when adjusting at least one of the cost values described in the above examples, such as N or M. For example, in the case of reducing the SAD value of the initial MV, the SAD value of the initial MV can be reduced by the reciprocal of the value of the signaled parameter multiplied by the cost value of the initial MV (i.e., Offset SADThe value is determined to be 1 / N of the SAD value associated with the initial MV, or the parameter value is decreased (i.e., Offset). SAD The value is determined to be the constant value M. The set of codewords can be designed for signaling the value N or M. Based on the set of codewords, a parameter value to be signaled is selected from a predefined set of values, and each codeword in the codeword set corresponds to one of the values in the predefined set. In one example, the set of values can be predefined as {4, 8, 16}. Binary codewords can be assigned to each value within the predefined set. Examples of binary codewords are shown in Table 1 below.

[0092] Table 1 Examples of codewords indicating parameter values for signaling

[0093] Value of Parameter Codeword 4 0 8 10 16 11

[0094] In some other examples, in the sequence parameter set, picture parameter set, slice header, CTU, and / or CU, a special value can be signaled into the bitstream to indicate that the initial MV has an updated cost value of zero, which is equivalent to the case of disabling DMVR. In one example, when the cost value of the initial MV is decreased by Offset SAD and the value of Offset SAD is determined to be 1 / N (where N is an integer) of the SAD value associated with the initial MV, N = 1 (i.e., the value of the parameter signaled is 1) will cause the SAD associated with the original MV to be equal to zero. In such a case, the refined MV derived through the DMVR process is always the original MV (i.e., in this case, the original MV is the refined MV), which is equivalent to disabling DMVR. In some examples, the special value one (1) can be included in the predefined set of values of the parameter, which can be, for example, {1, 4, 8, 16}.

[0095] According to the above examples, the DMVR process is modified such that in the integer-sample offset search stage, the initial MV is favored compared to other MV candidates, thereby reducing the loss caused by possible scenarios where the refined MV is worse than the original MV.

[0096] Figure 6 is a block diagram showing an exemplary apparatus for video coding and decoding according to some embodiments of the present disclosure. Apparatus 600 can be a terminal, such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.

[0097] As Figure 6As shown in [FIGURE], device 600 may include one or more of the following components: processing component 602, memory 604, power component 606, multimedia component 608, audio component 610, input / output (I / O) interface 612, sensor component 614, and communication component 616.

[0098] The processing component 602 generally controls the overall operation of the device 600, such as operations related to display, telephone calls, data communication, camera operations, and recording operations. The processing component 602 may include one or more processors 620 for executing instructions to complete all or part of the steps of the above methods. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.

[0099] The memory 604 is configured to store different types of data to support the operation of the device 600. Examples of such data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, and the memory 604 may be static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a disk, or a compact disc.

[0100] The power component 606 powers different components of the device 600. The power component 606 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 600.

[0101] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some examples, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen for receiving input signals from the user. The touch panel may include one or more touch sensors for sensing touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch actions or swipe actions but also detect the duration and pressure associated with the touch operation or swipe operation. In some examples, the multimedia component 608 may include a front camera and / or a rear camera. When the device 600 is in an operating mode (such as a shooting mode or a video mode), the front camera and / or the rear camera may receive external multimedia data.

[0102] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC). When the device 600 is in an operating mode (such as a call mode, a recording mode, and a voice recognition mode), the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some examples, the audio component 610 further includes a speaker for outputting audio signals.

[0103] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module. The above peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to a home button, a volume button, a power-on button, and a lock button.

[0104] The sensor component 614 includes one or more sensors for providing a status assessment of different aspects of the device 600. For example, the sensor component 614 can detect the on / off state of the device 600 and the relative positions of components. For example, these components are the display and the keyboard of the device 600. The sensor component 614 can also detect a change in the position of the device 600 or a change in the position of a component of the device 600, the presence or absence of user contact on the device 600, the orientation or acceleration / deceleration of the device 600, and a change in the temperature of the device 600. The sensor component 614 can include a proximity sensor, which is configured to detect the presence of nearby objects without any physical contact. The sensor component 614 can also include an optical sensor, such as a CMOS or CCD image sensor used in imaging applications. In some examples, the sensor component 614 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0105] The communication component 616 is configured to facilitate wired or wireless communication between the device 600 and other devices. The device 600 can access a wireless network based on communication standards (such as WiFi, 4G, or a combination thereof). In one example, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In one example, the communication component 616 can further include a near field communication (NFC) module for facilitating short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0106] In one example, the apparatus 600 may be implemented by one or more of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components to perform the above method.

[0107] The non-transitory computer-readable storage medium may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a hybrid drive or a solid state hybrid drive (SSHD), a read only memory (ROM), a compact disk read only memory (CD-ROM), magnetic tape, a floppy disk, etc.

[0108] Figure 7 FIG. is a flowchart illustrating an exemplary process of decoder-side motion vector refinement in video coding according to some embodiments of the present disclosure.

[0109] In step 702, the processor 620 derives an initial motion vector (MV) of a current coding unit (CU).

[0110] In step 704, the processor 620 derives a plurality of MV candidates for decoder-side motion vector refinement (DMVR).

[0111] In step 706, the processor 620 determines a cost value for each MV candidate among the initial MV and the MV candidates.

[0112] In step 708, the processor 620 receives signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV.

[0113] In step 710, the processor 620 obtains updated cost values by adjusting at least one of the cost values based on the parameter.

[0114] In step 712, the processor 620 derives a refined MV based on the initial MV and the updated cost values of the MV candidates.

[0115] The parameter may be signaled in one or a combination of the following: a sequence parameter set, a picture parameter set, a slice header, a coding tree unit (CTU), and / or a coding unit (CU).

[0116] The value of the signaled parameter may be selected from a predefined set of values based on a set of codewords, where each codeword in the set of codewords corresponds to one of the values in the predefined set.

[0117] Adjusting at least one cost value in the cost values may include: reducing the cost value of the initial MV by a first value, where the first value is determined using a predefined process; or increasing the cost value of the MV candidate by a second value, where the second value is determined using a predefined process.

[0118] In some examples, an apparatus for video encoding and decoding is provided. The apparatus includes one or more processors 620; and a memory 604 configured to store instructions executable by the one or more processors; where the one or more processors are configured to perform the method as Figure 7 shown.

[0119] In some other examples, a non-transitory computer-readable storage medium 604 is provided, having instructions stored therein. When the instructions are executed by one or more processors 620, the instructions cause the processors to perform the method as Figure 7 shown.

[0120] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art in light of the teachings presented in the foregoing description and the associated drawings.

[0121] The examples were chosen and described to explain the principles of the present disclosure and to enable others of ordinary skill in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and the various embodiments with various modifications suitable for the particular purposes contemplated. Accordingly, it will be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the present disclosure.

Claims

1. A method for video encoding and decoding, comprising: Deriving an initial motion vector MV of a current coding unit CU; Deriving a plurality of MV candidates for decoder-side motion vector refinement DMVR; Determining a cost value for each of the initial MV and the MV candidates; Receiving signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; Obtaining updated cost values by adjusting the at least one of the cost values based on the parameter; And Deriving a refined MV based on the updated cost values of the initial MV and the MV candidates, wherein the parameter is determined according to coded information in the current CU, and the coded information includes at least one or a combination of the following: coded block size, magnitude of the motion vector, SAD of the initial MV, and relative position of the DMVR processing unit.

2. The method according to claim 1, wherein the parameter is signaled in one or a combination of the following: sequence parameter set, picture parameter set, slice header, coding tree unit CTU, and / or coding unit CU.

3. The method according to claim 1, wherein the value of the signaled parameter is selected from a predefined set of values based on a set of codewords, and each of the codewords corresponds to one of the values in the predefined set.

4. The method according to claim 3, wherein adjusting the at least one cost value among the cost values includes: Reducing the cost value of the initial MV by a first value determined using a predefined process.

5. The method according to claim 4, wherein the first value is determined as the reciprocal of the value of the parameter multiplied by the cost value of the initial MV.

6. The method according to claim 5, wherein the predefined set of values includes a value one, which indicates that the initial MV has an updated cost value of zero.

7. The method according to claim 4, wherein the first value is equal to the value of the parameter.

8. The method according to claim 1, wherein adjusting the at least one cost value among the cost values includes: Increasing the cost value of the MV candidate by a second value determined using a predefined process.

9. The method according to claim 3, wherein the predefined set of values includes {4, 8, 16}.

10. The method according to claim 1, wherein the cost value includes sum of absolute differences SAD.

11. An apparatus for video encoding and decoding, comprising: One or more processors; And A memory configured to store instructions executable by the one or more processors; wherein the one or more processors are configured to: Derive an initial motion vector MV of a current coding unit CU; Derive a plurality of MV candidates for decoder-side motion vector refinement DMVR; Determine a cost value for each of the initial MV and the MV candidates; Receive signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; Obtain updated cost values by adjusting the at least one of the cost values based on the parameter; And Derive a refined MV based on the updated cost values of the initial MV and the MV candidates, The parameter is determined according to the transcoded information in the current CU, and the transcoded information includes at least one or a combination of the following: transcoding block size, magnitude of the motion vector, SAD of the initial MV, and relative position of the DMVR processing unit.

12. The apparatus according to claim 11, wherein the parameter is signaled in one or a combination of the following: sequence parameter set, picture parameter set, slice header, coding tree unit CTU, and / or coding unit CU.

13. The apparatus according to claim 11, wherein the value of the signaled parameter is selected from a predefined set of values based on a set of codewords, and each of the codewords corresponds to one of the values in the predefined set.

14. The apparatus according to claim 13, wherein adjusting at least one of the cost values comprises: Reduce the cost value of the initial MV by a first value determined using a predefined process.

15. The apparatus according to claim 14, wherein the first value is determined as the reciprocal of the value of the parameter multiplied by the cost value of the initial MV.

16. The apparatus according to claim 15, wherein the predefined set of values includes a value one, which indicates that the initial MV has an updated cost value of zero.

17. The apparatus according to claim 14, wherein the first value is equal to the value of the parameter.

18. The apparatus according to claim 11, wherein adjusting at least one of the cost values includes increasing the cost value of the MV candidate by a second value determined using a predefined process.

19. The apparatus according to claim 13, wherein the predefined set of values includes {4, 8, 16}.

20. The apparatus according to claim 11, wherein the cost value includes sum of absolute differences SAD.

21. A non-transitory computer-readable storage medium including instructions stored therein, wherein when the instructions are executed by one or more processors, the instructions cause the one or more processors to perform actions, the actions including: Derive an initial motion vector MV of a current coding unit CU; Derive a plurality of MV candidates for decoder-side motion vector refinement DMVR; Determine a cost value for each of the initial MV and the MV candidates; Receive signaling corresponding to a parameter for adjusting at least one of the cost values in favor of the initial MV; Obtain an updated cost value by adjusting at least one of the cost values based on the parameter; and Derive a refined MV based on the updated cost values of the initial MV and the MV candidates, wherein the parameter is determined according to the transcoded information in the current CU, and the transcoded information includes at least one or a combination of the following: transcoding block size, magnitude of the motion vector, SAD of the initial MV, and relative position of the DMVR processing unit.

22. The non-transitory computer-readable storage medium according to claim 21, wherein the parameter is signaled in one or a combination of the following: sequence parameter set, picture parameter set, slice header, coding tree unit CTU, and / or coding unit CU.

23. The non-transitory computer-readable storage medium according to claim 21, wherein the value of the parameter signaled is selected from a predefined set of values based on a set of codewords, each of the codewords corresponding to one of the values in the predefined set.

24. The non-transitory computer-readable storage medium according to claim 23, wherein adjusting at least one of the cost values includes: Reduce the cost value of the initial MV by a first value determined using a predefined process.

25. The non-transitory computer-readable storage medium according to claim 24, wherein the first value is determined as the reciprocal of the value of the parameter multiplied by the cost value of the initial MV.

26. The non-transitory computer-readable storage medium according to claim 25, wherein the predefined set of values includes the value one, which indicates that the initial MV has an updated cost value of zero.

27. The non-transitory computer-readable storage medium according to claim 24, wherein the first value is equal to the value of the parameter.

28. The non-transitory computer-readable storage medium according to claim 21, wherein adjusting at least one of the cost values includes: Increase the multiple cost values of the MV candidate by a second value determined using a predefined process.

29. The non-transitory computer-readable storage medium according to claim 23, wherein the predefined set of values includes {4, 8, 16}.

30. The non-transitory computer-readable storage medium according to claim 21, wherein the cost value includes the sum of absolute differences SAD.

Citation Information

Patent Citations

  • Method and apparatus for decoder-side motion vector refinement in video coding

    CN116916026A