Method and device for pattern-based motion vector derivation for video coding

Through the method of template matching and multi-reference table derivation, combined with different codeword sending and merging indexes and other optimization methods, the problem of limited improvement in style-based motion vector derivation performance in existing video encoding systems is solved, achieving more efficient encoding and decoding and reducing bandwidth and complexity.

CN114449287BActive Publication Date: 2025-05-06MEDIATEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210087199.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-03-16
Filing Date
2017-03-14
Publication Date
2025-05-06
Estimated Expiration
2037-03-14

AI Technical Summary

Technical Problem

When using style-based motion vector derivation, existing video encoding systems have limited performance improvement and high complexity. Especially in the decoder-side motion vector derivation method, there are problems with bandwidth and complexity.

Method used

A new video encoding method is proposed to deduce the best in-reference table templates for the current template through template matching, and repeat this process between multiple reference tables until the specified number is reached. At the same time, different codewords are used to send messages to merge indexes, disable MV costs, disable weighted prediction of PMVD, reduce block matching operations, and perform range derivation in multi-parameter CABAC.

Benefits of technology

Improves encoding and decoding efficiency, reduces bandwidth and complexity, and improves video encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114449287B_ABST
    Figure CN114449287B_ABST
Patent Text Reader

Abstract

The present invention discloses a video encoding method and apparatus based on bidirectional matching or template matching using decoder-derived motion information. According to one method, template matching is used to derive the best template in a first reference table (e.g., Table 0 / Table 1) of a current template. A new current template is determined based on the current template, the best template in the first reference table, or both. The best template in a second reference table (e.g., Table 0 / Table 1) is derived using the new current template. The process is repeated for the first reference table and the second reference table until a certain number of repetitions is reached. The derivation of the new current template may depend on the picture type of the picture containing the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description of the case

[0002] This application is a divisional application of the invention patent application with application number 201780018068.4, and invention name is method and device for style-based motion vector derivation for video coding. Technical Field

[0003] The present invention relates to motion compensation for video coding using decoder side derived motion information. More particularly, the present invention relates to performance improvement or complexity reduction for pattern-based motion vector derivation. Background Art

[0004] In a common video coding system using motion-compensated inter prediction, motion information is usually sent from the encoder to the decoder so that the decoder can correctly perform motion-compensated inter prediction. In such a system, motion information will occupy some coding bits. In order to improve coding efficiency, a motion vector derivation method at the decoder side is disclosed in VCEG-AZ07 (Jianle Chen et al., Further improvements to HMKTA-1.0, ITU-Telecommunications Standardization Sector, Study Group 16 Question 6, Video Coding Experts Group (VCEG), 52nd Meeting: June 19-26, 2015, Warsaw, Poland). According to VCEG-AZ07, the motion vector derivation method at the decoder side uses two frame rate up conversion (Frame Rate Up-Conversion, FRUC) modes. One of the FRUC modes is called B-slice bilateral matching and the other of the FRUC modes is called template matching of P-slice or B-slice.

[0005] Figure 1An example of FRUC bidirectional matching mode is shown, where the motion information of the current block 110 is derived based on two reference images. The motion information of the current block is derived by finding the best match between two blocks (120 and 130) in two different reference images (i.e., Ref0 and Ref1) along the motion trajectory 140 of the current block (Cur block). Under the assumption of continuous motion trajectory, the motion vector MV0 associated with ref0 and the motion vector MV1 associated with Ref1 pointing to the two reference blocks should be proportional to the temporal distances, i.e., TD0 and TD1, between the current image (i.e., Cur pic) and the two reference images.

[0006] Figure 2 An example of FRUC template matching mode is shown. Neighboring areas (220a and 220b) of the current block 210 in the current image (i.e., Cur pic) are used as a template to match the corresponding templates (230a and 230b) in the reference image (i.e., Ref0). The best match between templates 220a / 220b and templates 230a / 230b determines a decoder derived motion vector 240. Although Ref0 is Figure 2 As shown, Ref1 can also be used as a reference image.

[0007] According to VCEG-AZ07, FRUC_mrg_flag is signaled when merge_flag or skip_flag is true. If FRUC_mrg_flag is 1, then FRUC_merge_mode is signaled to indicate that bilateral matching merge mode or template matching merge mode is selected. If FRUC_mrg_flag is 0, this implies that regular merge mode is used and a mergeindex is signaled in this case. In video coding, in order to improve coding efficiency, the motion vector of a block can be predicted using motion vector prediction (MVP), in which a candidate list is generated. In merge mode, a merge candidate list can be used to encode a block. When a block is encoded using merge mode, the motion information (egmotion vector) of the block can be represented by one of the candidate MVs in the merge MV table. Therefore, instead of directly transmitting the motion information of the block, a merge index is transmitted to the decoder. The decoder maintains the same merge table and uses the merge index to retrieve the merge candidate signaled by the merge index. In general, the merge candidate list contains a small number of candidates and transmitting the merge index is more efficient than transmitting motion information. When a block is encoded in merge mode, its motion information is "merged" with the motion information of neighboring blocks by signaling a merge index, rather than actually transmitting the motion information. However, the prediction residuals are still transmitted. In the case where the prediction residual is zero or very small, the prediction residual is "skipped" (i.e. skip mode) and the block is encoded in skip mode using the merge index to identify the merge MV in the merge table.

[0008] Although the term FRUC refers to motion vectors for frame rate up conversion (Frame Rate Up-Conversion), its technology is used for the decoder to derive one or more merged MV candidates without actually transmitting motion information. As such, FRUC is also referred to as decoder-derived motion information in this application. Because the template matching method is a pattern-based MV derivation technique, the template matching method of FRUC is also referred to as pattern-based MV Derivation (PMVD) in this application.

[0009] At the decoder side, a new temporal MVP called temporal derived MVP is derived by scanning all MVs in all reference frames. To derive the LIST_0 temporal derived MVP, for each LIST_0 MV in the LIST_0 reference frame, the MV is scaled to point to the current frame. The 4x4 block in the current frame pointed to by the scaled MV is the target current block. The MV is further scaled to point to the reference image with refIdx equal to 0 in LIST_0 for the target current block. The further scaled MV is stored in the LIST_0 MV field for the target current block. Figure 3A and Figure 3B An example of MVP derived in the time domain showing the derivation of List_0 and List_1 respectively. Figure 3A and Figure 3B In the example, each small square block corresponds to a 4x4 block. The process of temporally derived MVP scans all MVs in the 4x4 blocks in all reference images to generate the temporally derived LIST_0 and LIST_1 MVPs of the current frame. For example, Figure 3A In FIG. 1 , block 310, block 312 and block 314 correspond to a 4x4 block of the current image, a List_0 reference image with index equal to 0 (i.e., refidx=0) and a List_0 reference image with index equal to 1 (i.e., refidx=1), respectively. The motion vectors 320 and 330 within the two blocks of the List_0 reference image with index equal to 1 are known. Then, the temporally derived MVPs 322 and 332 can be derived by scaling the motion vectors 320 and 330, respectively. The scaled MVP is then assigned to a corresponding block. Similarly, in Figure 3B In FIG. 1 , block 340, block 342 and block 344 correspond to a 4x4 block of the current image, a List_1 reference image with index equal to 0 (i.e., refidx=0) and a List_1 reference image with index equal to 1 (i.e., refidx=1), respectively. Motion vectors 350 and 360 within the two blocks of the List_0 reference image with index equal to 1 are known. Then, the temporally derived MVPs 352 and 362 can be derived by scaling the motion vectors 350 and 360, respectively.

[0010] For the bidirectional matching merge mode and the template matching merge mode, two-stage matching is adopted. The first stage is PU-level matching, and the second stage is sub-PU-level matching. In the PU-level matching, multiple initial MVs in LIST_0 and LIST_1 are selected respectively. These MVs include MVs from merge candidates (i.e., traditional merge candidates, such as those specified in the HEVC standard) and MVs from MVP derived from the time domain. Two different staring MV groups are generated for the two tables. For each MV in one table, an MV pair is generated by synthesizing the MV and a mirrored MV, which is derived by scaling the MV to another table. For each MV pair, two reference blocks are compensated by using the MV pair. The sum of absolutely differences (SAD) of the two blocks is calculated. The MV pair with the smallest SAD is selected as the best MV pair.

[0011] After deriving the best MV for a PU, a diamond search is performed to improve the MV pair. The improvement precision is 1 / 8-pel. The search range for improvement is limited to ±1 pixel. The final MV pair is the MV pair derived at the PU level. Diamond search is a fast block matching motion estimation algorithm well known in the video coding community. Therefore, the details of the diamond search algorithm are not repeated here.

[0012] For the second stage sub-PU-level search, the current PU is divided into sub-PUs. The depth of the sub-PU (e.g., 3) is signaled in the sequence parameter set (SPS). The minimum sub-PU size is a 4x4 block. For each sub-PU, multiple starting MVs are selected in LIST_0 and LIST_1, which include PU-level derived MVs, zero MV, HEVC collocated TMVP of the current sub-PU and the bottom-right block, time-domain derived MVP of the current sub-PU, and MVs of the left and above PUs / sub-PUs. The best M pair of sub-PUs is determined by using a similar mechanism such as PU-level search. A diamond search is performed to improve the MV pairs. Motion compensation is performed on the sub-PU to generate a predictor for the sub-PU.

[0013] For the template matching merge mode, the reconstructed pixels of the top 4 columns and the left 4 rows are used to form a template. Template matching is performed to find the best matching template and its corresponding MV. Template matching also uses two-stage matching. In PU-level matching, multiple starting MVs in LIST_0 and LIST_1 are selected respectively. These MVs include MVs from merge candidates (i.e., traditional merge candidates, such as those specified in the HEVC standard) and MVs from time-domain derived MVPs. Two different starting MVs for two tables are generated. For each MV in a table, the SAD cost of the template with the MV is calculated. The MV with the minimum cost is the best MV. A diamond search is then performed to improve the MV. The improvement accuracy is 1 / 8-pel. The improvement search range is limited to ±1 pixel. The final MV is the PU-level derived MV. The MVs in LIST_0 and LIST_1 are generated independently.

[0014] For the second stage sub-PU-level search, the current PU is split into sub-PUs. The depth of the sub-PU (e.g., 3) is signaled in the sequence parameter set SPS. The minimum sub-PU size is a 4x4 block. For each sub-PU at the left or top PU boundary, multiple starting MVs are selected in LIST_0 and LIST_1, which include the MVs of the PU-level derived MVs, the zero MV, the HEVC collocated TMVP of the current sub-PU and the bottom-right block, the time-domain derived MVP of the current sub-PU, and the MVs of the left and above PUs / sub-PUs. The best M pair of the sub-PU is determined by using a similar mechanism as for PU-level search. A diamond search is performed to improve the MV pair. Motion compensation is performed on the sub-PU to generate a predictor for the sub-PU. Because these PUs are not at the left or top PU boundary, the second stage sub-PU-level search is not used, and the corresponding MV is set equal to the MV in the first stage.

[0015] In this decoder MV derivation method, template matching is also used to generate the MVP for inter-mode coding. When a reference image is selected, template matching is performed to find an optimal template on the selected reference image. Its corresponding MV is the derived MVP. The MVP is inserted into the first position in AMVP. AMVP stands for Advanced MV Prediction, in which the current MV is coded predictively using a candidate table. The difference between the current MV and the MV candidate selected from the candidate table is encoded.

[0016] Bi-directional optical flow (BIO) was revealed in JCTVC-C204 (authors are Elena Alshina and Alexander Alshin, "Bi-directional optical flow", Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 3rd Meeting: Guangzhou, China, 7-15, October 2010) and VECG-AZ05 (authors are E. Alshina et al., Known tools performance investigation for next generation video coding, ITU-T SG 16 Question 6, Video Coding Experts Group (VCEG), 52nd Meeting: June 19-26, 2015, Warsaw, Poland, Document: VCEG-AZ05). BIO uses the assumption of optical flow and steady motion to achieve sample-level motion refinement. This is only used for true bidirectional prediction blocks, which are predicted from two reference frames, one of which is the previous frame and the other is the subsequent frame. In VCEG-AZ05, BIO uses a 5x5 window to derive the motion refinement for each sample. Therefore, for an NxN block, the motion compensation result of an (N+4)x(N+4) block and the corresponding gradient information are used to derive the sample-based motion improvement of the NxN block. According to VCEG-AZ05, a 6-tap gradient filter and a 6-tap interpolation filter are used to generate the gradient information of BIO. Therefore, the computational complexity of BIO is higher than that of traditional bidirectional prediction. In order to further improve the performance of BIO, the following method is proposed.

[0017] In a technical paper written by Marpe et al. (D. Marpe, H. Schwarz, and T. Wiegand, "Context-Based Adaptive Binary Arithmetic Coding in the H.264 / AVC Video Compression Standard", IEEE Transactions on Circuits and Systems for Video Technology, Vol. 13, No. 7, pp. 620-636, July 2003), a multi-parameter probability update (context adaptive binary coding) for HEVC CABAC is proposed. The parameter N = 1 / (1-α) is a measure of the number of previously coded bins, which has an important influence on the current update ("window size"). This value determines in a certain sense the average storage required by the system. The choice of the parameter that determines the sensitivity of the model is a difficult and important issue. A sensitive system can react quickly to real changes. On the other hand, an insensitive model will not react to noise and random errors. Both parameters are useful, but they restrict each other. By using different control signals, α can be changed during encoding. However, this method is very laborious and complicated. In this way, different values ​​of αi can be calculated at the same time:

[0018] pi_new = (1- αi) y+ α i pi_old. (1)

[0019] Weighted average is used to predict the probability of the next unit:

[0020] pnew = Σβi pi_new. (2)

[0021] In the above equation, βi is a weighting factor. In AVC CABAC, lookup tables (i.e., m_aucNextStateMPS and m_aucNextStateLPS) and exponential meshes are used for probability updates. However, probability updates can use uniform meshes and explicit calculations with multiplication free formulas.

[0022] Assuming that the probability pi is represented by an integer from 0 to 2k, the probability is determined as follows:

[0023] pi=Pi / 2k.

[0024] If 'αi is the inverse of a quadratic number (i.e., αi = 1 / 2Mi), then we get the multiplication-free equation for the probability update:

[0025] Pi = (Y >> Mi) + P – (Pi >> Mi). (3)

[0026] In the above equation, “>>Mi” represents a right shift operation of Mi. The equation predicts the probability that the next unit will be “1”, where Y=2k if the last coded unit was “1” and Y=0 if the last coded unit was “0”.

[0027] To balance increased complexity with improved performance, a linear combination of probability estimates involving only two parameters is used:

[0028] P0 = (Y>>4) + P0 – (P0>>4) (4)

[0029] P1 = (Y>>7) + P1 – (P0>>7) (5)

[0030] P = (P0 + P1+1)>>1 (6)

[0031] For probability calculation in AVC CABAC, the floating point value is always less than or equal to 1 / 2. If the probability exceeds this limit, LPS (least probable symbol) becomes MPS (most probable symbol) to keep the probability within the above interval. This concept has some obvious advantages, such as reducing the size of the lookup table.

[0032] However, direct generalization of the above method for multi-parameter update models may encounter some difficulties. In practice, one probability estimate can exceed the limit while the other is still less than 1 / 2. Therefore, either MPS / LPS switching is required for each Pi or some averages require MPS / LPS switching. In both cases, it introduces additional complexity but does not achieve significant performance improvement. Therefore, the present application proposes to increase the allowed level of probability (in floating point values) to a maximum of 1 and prohibit MPS / LPS switching. Therefore, a lookup table (LUT) storing RangeOne or RangeZero is derived. Summary of the invention

[0033] In order to solve the above technical problems, the present application provides a new method and device for video encoding using motion compensation.

[0034] In one method, template matching is used to derive the best template in a first reference table (e.g., Table 0 / Table 1) of a current template. A new current template is determined based on the current template, the best template in the first reference table, or both. The best template in a second reference table (e.g., Table 0 / Table 1) is derived using the new current template. The process is repeated for the first reference table and the second reference table until a number of repetitions is reached. The derivation of the new current template may depend on the picture type of the picture containing the current block.

[0035] The method and device for video encoding using motion compensation provided by the present invention can improve encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Shows an example of motion compensation using the bidirectional matching technique, where the current block is predicted by two reference blocks along the motion track.

[0037] Figure 2 An example of motion compensation using template matching techniques is shown, where a template of the current block is matched with a reference template in a reference image.

[0038] Figure 3A An example of the temporal motion vector prediction (MVP) derivation process for the LIST_0 reference picture is shown.

[0039] Figure 3B An example of the temporal motion vector prediction (MVP) derivation process for the LIST_1 reference picture is shown.

[0040] Figure 4 A flow chart showing a video coding system using decoder-side derived motion information according to an embodiment of the present invention, wherein merge indexes are signaled using different codewords.

[0041] Figure 5 A flowchart of a video coding system using decoder-side derived motion information according to an embodiment of the present invention is shown, wherein a first-stage MV or a first-stage MV pair is used as the only initial MV or MV pair for second-stage search or the center MV of the search window.

[0042] Figure 6 A flow chart of a video coding system using decoder-side derived motion information according to an embodiment of the present invention is shown, wherein after finding a reference template of a first reference table, a current template is modified for template search in another reference table.

[0043] Figure 7A flowchart of a video coding system using decoder-side derived motion information according to an embodiment of the present invention is shown, wherein sub-PU search is disabled in template search.

[0044] Figure 8 An example of pixels within a current block and a reference block showing the difference between the current block and the reference block.

[0045] Fig. 9 An example of pixels in a current block and a reference block for calculating differences between the current block and the reference block according to an embodiment of the present invention is shown, wherein sub-blocks with virtual pixel values ​​are used to reduce the operation of calculating the differences.

[0046] Fig.10 A flowchart of a video encoding system using decoder-side derived motion information according to an embodiment of the present invention is shown, wherein a decoder-side MV or a decoder-side MV pair is derived according to a decoder-side MV derivation process, wherein the decoder-side MV derivation process uses a block difference calculation based on a reduced bit depth in an MV derivation associated with the decoder-side MV derivation process. DETAILED DESCRIPTION

[0047] The following description is the best way to implement the invention, which is intended to illustrate the basic principles of the invention and should not be construed as limiting. The scope of the invention is best determined by reference to the appended claims.

[0048] In the present invention, multiple methods are disclosed to reduce bandwidth or complexity or improve coding efficiency for decoder-side motion vector derivation.

[0049] Signalling Merge Index with Different Codewords

[0050] In the bidirectional matching merge mode and the template matching merge mode, the LIST_0 and LIST_1 MVs in the merge candidates are used as the starting MVs. The best MV is undoubtedly derived by searching all these MVs. These merge modes will result in high memory bandwidth. Therefore, the present invention discloses a method to signal the merge index of the bidirectional matching merge mode or the template matching merge mode. If the merge index is signaled, the best starting MVs in LIST_0 and LIST_1 are known. Bidirectional matching or template matching only needs to perform an improved search around the signaled merge candidates. For bidirectional matching, if the merge match is a uni-directional MV, the corresponding MV in the other table can be generated by using a mirrored (scaled) MV.

[0051] In another embodiment, by using a predetermined MV generation method, the starting MV and / or MV pair in LIST_0, LIST_1 are known. The best starting MV or the best MV pair in LIST_0 and / or LIST_1 is explicitly signaled to reduce bandwidth requirements.

[0052] Although bidirectional matching and template matching have been commonly used in a two-stage manner, the method of signaling a combined index using different codewords according to the present invention is not limited to a two-stage manner.

[0053] In another embodiment, when a merge index is signaled, the selected MV can also be used to exclude or select some candidates in the first stage, i.e., PU-level matching. For example, some MVs in the candidate list that are far from the selected MV can be excluded. In addition, N MVs in the candidate list that are closest to the selected MV but in different reference frames can be selected.

[0054] In the above method, the codeword of the merge index can be a fixed-length (FL) code, a unary code, or a truncated unary (TU) code. The context of the merge index for the bidirectional matching and template matching merge modes can be different from that of the normal merge mode. A different context model set can be used. The codeword can be slice-type dependent or signaled in the slice header. For example, the TU code can be used in a random-access (RA) picture or for a B-picture when the picture order count (POC) of the reference frame is not always less than the current picture. The FL code can be used for low-latency B / P pictures or can be used for P-pictures or B-pictures where the POC of the reference frame is not always less than the current picture.

[0055] Figure 4A flow chart of a time coding system using information derived at the decoder side according to the embodiment is shown, wherein the merge index is signaled using different codewords. The steps shown in the following flow chart or any subsequent flow chart, as well as other flow charts in the disclosure, can be implemented as executable program code on one or more processors (e.g., one or more CPUs) at the encoder side and / or the decoder side. The steps shown in the flow chart can also be implemented based on hardware, such as one or more electronic devices or processors for executing the steps of the flow chart. According to the method, in step 410, input data related to a current block in a current image is received at the encoder side, or a video stream containing encoded data related to the current block in the current image is received at the decoder side. In step 420, a decoder-side merge candidate for the current block using bidirectional matching, template matching, or both is derived. In step 430, a merge candidate group including the merge candidate at the decoder side is generated. Methods for generating a merge candidate group are known in the industry. Generally, it includes motion information of spatial and / or temporal neighboring blocks as merge candidates. The merge candidate at the decoder side derived according to the present embodiment is included in the merge candidate group. In step 440, a current merge index of the selected current block is signaled at an encoder side or decoded at a decoder side using one of at least two different codeword groups or one of at least two contexts encoded using a context basis. The at least two different codeword groups or the at least two contexts encoded using a context basis are used to encode a merge index associated with a merge candidate of the merge candidate group.

[0056] No MV Cost, Sub-block Refined from Merge Candidate

[0057] In the bidirectional matching merge mode and the template matching merge mode, the initial MV is first derived from the neighboring blocks and the template co-located block. In the pattern-based MV search, the MV cost (i.e., the MV difference multiplied by lambda, λ) is added to the prediction distortion. The method of the present invention limits the searched MV to around the initial MV. The MV cost is usually used at the encoder end to reduce the bit overhead of the MVD (MV difference), because signaling the MVD consumes coding bits. However, the motion vector derived at the decoder end is a decoder-end process, which does not require additional end information. Therefore, in one embodiment, the MV cost can be eliminated.

[0058] In the bidirectional matching merge mode and the template matching merge mode, a two-stage MV search is adopted. The best MV in the first search stage (CU / PU-level stage) is used as one of the initial MVs in the second search stage. The search window of the second stage is centered on the initial MV of the second search stage. However, it requires memory bandwidth. In order to further reduce the bandwidth requirement, the present invention discloses a method of using the initial MV of the first search stage as the center MV of the search window of the second stage sub-block search. In this way, the search window of the first stage can be reused in the second stage. No additional bandwidth is required.

[0059] In VCEG-AZ07, for sub-PU MV search in template and bidirectional matching, the left and top MVs of the current PU are used as initial search candidates. In one embodiment, in order to reduce memory bandwidth, the second stage sub-PU search only uses the best MV of the first stage as the initial MV of the second stage.

[0060] In another embodiment, in combination with the above-mentioned method of sending combined index, a search window of an explicit sending combined index MV is used in the first-stage and second-stage searches.

[0061] Figure 5 A flowchart of a video encoding system using motion information derived at the decoder end according to an embodiment of the present invention is shown, wherein the first stage MV or the first stage MV pair is used as the only initial MV or MV pair for the second stage search or as the center MV of the search window. According to the method, at step 510, input data related to a current block in a current image is received, wherein each current block is divided into a plurality of sub-blocks. At the encoder end, the input data may correspond to pixel data to be encoded and the input data may correspond to encoded data to be decoded at the decoder end. At step 520, the first stage MV or the first stage MV pair is derived based on one or more first stage MV candidates using bidirectional matching, template matching, or both. At step 530, the second stage MVs of the plurality of sub-blocks are derived by deriving one or more second stage MVs for each sub-block using bidirectional matching, template matching, or both, wherein the first stage MV or the first stage MV pair is used as the only initial MV or MV pair for the second stage bidirectional matching, template matching, or both or as the center MV of the search window. At step 540, a final MV or a final MVP is derived from a set of MV candidates or MVP candidates including the second stage MV. In step 550 , the current block or the current MV of the current block is encoded or decoded using the final MV or the final MVP at the encoder end or the decoder end, respectively.

[0062] Disable Weighted Prediction for PMVD

[0063] In template matching merge mode and bidirectional matching merge mode, weighted prediction is disabled according to this method. If both LIST_0 and LIST_1 contain matching reference blocks, the weight is 1:1.

[0064] Matching Criterion

[0065] When the best or several best LIST_0 / LIST_1 templates are found, the templates in LIST_0 / LIST_1 can be used to search for templates in LIST_1 / LIST_0 (i.e., templates in LIST_0 are used to search for templates in LIST_1, and vice versa). For example, the current template of List_0 can be modified to "2*(current template)-LIST_0 template", where the LIST_0 template corresponds to the best LIST_0 template. The new current template is used to search for the best template in LIST_1. The notation "2*(current template)-LIST_0 template" means a pixel-wise operation between the current template and the best template found in reference table 0 (i.e., LIST_0 template). Although traditional template matching wants to achieve the best match between the current template and the reference template in reference table 0 and the best match between the current template and the reference template in reference table 1 respectively. The modified current template of another reference table can help achieve the best match together. Repeated searches can be used. For example, after finding the best LIST_1 template, the current template can be modified to "2*(current template) - LIST_1 template". The modified new current template is used to search for the best template in LIST_0 again. The number of repetitions and the first target reference table should be defined in the standard.

[0066] The proposed matching criteria of LIST_1 can be slice-type dependent. For example, when the image sequence numbers of the reference frames are not all smaller than the current image, "2*(current template) - LIST_0 template" can be used for random-access (RA) images or for B-images, and "current template" can be used for other types of images; vice versa.

[0067] Figure 6A flow chart of a video encoding system using decoder-side derived motion information according to the present invention is shown, wherein after finding a reference template in a first reference table, a current template is modified for use in template searches in other reference tables. According to the method, at step 610, input data related to a current block or subblock in a current image is received. At the encoder end, the input data may correspond to pixel data to be encoded, and the input data may correspond to encoded data to be decoded at the decoder end. At step 620, one or more first best templates of the current template of the current block or subblock are derived using template matching, which are pointed to by one or more first best MVs in the first reference table, wherein the one or more first best MVs are derived based on template matching. At step 630, after deriving the one or more first best templates, a new current template is determined based on the current template, the one or more first best templates, or both. In step 640, one or more second best templates of a new current template of the current block or subblock are derived using template matching, which are pointed to by one or more second best MVs in the second reference table, wherein the one or more second best MVs are derived according to template matching, and the first reference table and the second reference table belong to a group including Table 0 and Table 1, and the first reference table is different from the second reference table. In step 650, one or more final MVs or final MVPs are determined from a group of MV candidates or MVP candidates, and the MV candidates or MVP candidates include one or more best MVs associated with one or more first best MVs and one or more second best MVs. In step 660, the current block or subblock or the current MV of the current block or subblock is encoded or decoded using the final MV or final MVP at the encoder end or the decoder end, respectively.

[0068] Disable Sub-PU-Level Search for templateMatching

[0069] According to the method of the present invention, sub-PU search for template matching merge mode is disabled. Sub-PU search is only used for bidirectional matching merge. For template matching merge mode, since the entire PU / CU can have the same MV, BIO can be used for coding blocks in template matching merge mode. As mentioned above, BIO is used for truly bi-directional predicted blocks to improve motion vectors.

[0070] Figure 7A flow chart of a video coding system using decoder-side derived motion information according to the present method is shown, wherein a sub-PU search for template search is disabled. According to the present method, at step 710, input data related to a current block or sub-block in a current image is received. At the encoder side, the input data may correspond to pixel data to be encoded, and the input data may correspond to encoded data to be decoded at the decoder side. At step 720, a first-stage MV or a first-stage MV pair is derived based on one or more MV candidates using bidirectional matching or template matching. At step 730, it is checked whether bidirectional matching or template matching is used. If bidirectional matching is used, steps 740 and 750 are performed. If template matching is used, step 760 is performed. At step 740, second-stage MVs of multiple sub-blocks are generated by deriving one or more second-stage MVs for each sub-block based on the first-stage MV or the first-stage MV pair using bidirectional matching, wherein the current block is divided into the multiple sub-blocks. At step 750, a final MV or a final MVP is determined from a set of MV candidates or MVP candidates including the second-stage MV. In step 760, a final MV or final MVP is determined from a set of MV candidates or MVP candidates including the first stage MV. In step 770, after the final MV or final MVP is determined, the current block or the current MV of the current block is encoded or decoded using the final MV or final MVP at the encoder or decoder, respectively.

[0071] Reduce the Operations of Block Matching

[0072] For MV derivation at the decoder side, the SAD cost of the template with multiple MVs is calculated to find the best MV at the decoder side. In order to reduce the operation of SAD calculation, a method of approximate the SAD between the current block 810 and the reference block 820 is disclosed. In the SAD calculation of traditional block matching, Figure 8 As shown, the squared differences between the corresponding pixel pairs of the current block (8x8 block) and the reference block (8x8 block) are calculated and summed up as shown in equation (1) to obtain the final sum of the squared differences, where Ci,j and Ri,j represent pixels in the current block 810 and the reference block 820 respectively, where the width is equal to N and the height is equal to M.

[0073]

[0074] For speedup, the current block and the reference block are divided into sub-blocks of size KxL, where K and L can be any integers. Fig. 9As shown, the current block 910 and the reference block 920 are both 8x8 blocks and are divided into 2x2 sub-blocks. Each sub-block is then treated as a virtual pixel and a virtual pixel value is used to represent each sub-block. The virtual pixel value can be the sum of pixels within the sub-block, the average of pixels within the sub-block, the dominant pixel value within the sub-block, a pixel within the sub-block, a preset pixel value, or any other way of calculating a value using pixels within the sub-block. SAD can be calculated by the sum of absolute differences between the virtual pixels of the current block and the reference block. In addition, the sum of the squared difference (SSD) can be calculated by the sum of squared differences between the virtual pixels of the current block and the reference block. Therefore, the SAD or SSD per pixel is approximated by the SAD or SSD of the virtual pixel, which requires fewer operations (e.g., fewer multiplications).

[0075] Furthermore, in order to maintain similar search results, the method also discloses a refinement search stage after finding M best matches using the SAD or SSD of virtual pixels, where M can be any positive integer. For each of the M best candidates, a per-pixel SAD or SSD can be calculated to find the final best matching block.

[0076] In order to reduce the complexity of SAD and SSD calculation, the method of the present invention calculates the first K-bit MSB (or truncated L-bit LSB) data. For example, for a 10-bit video input, it can use the 8-bit MSB to calculate the distortion of the current block and the reference block.

[0077] Fig.10 A flow chart of a video encoding system using decoder-side derived motion information according to the present method is shown. According to the present method, at step 1010, input data associated with a current block or sub-block in a current image is received. At the encoder side, the input data may correspond to pixel data to be encoded, and the input data may correspond to encoded data to be decoded at the decoder side. At step 1020, a decoder-side MV or a decoder-side MV pair is derived according to a decoder-side MV derivation process using a block difference calculation based on a reduced bit depth in an MV search associated with a decoder-side MV derivation process. At step 1030, a final MV or a final MVP is determined from a set of MV candidates or MVP candidates including a decoder-side MV or a decoder-side MV pair. At step 1040, the current block or the current MV of the current block is encoded or decoded using the final MV or the final MVP at the encoder side or the decoder side, respectively.

[0078] Range Derivation for Multi-Parameter CABAC

[0079] In multi-parameter CABAC, the method of the present invention uses the LPS table to derive the RangeOne or RangeZero for each probability state. The average RangeOne or RangeZero can be derived by averaging the RangeOnes or RangeZeros. The encoded RangeOne (RangeOne for coding, ROFC) and the encoded RangeZero (RangeZero for coding, RZFC) can be derived using equation (8):

[0080] RangeZero_0=(MPS_0==1)? RLPS_0:(range–RLPS_0);

[0081] RangeZero_1=(MPS_1==1)? RLPS_1:(range–RLPS_1);

[0082] ROFC=(2*range–RangeZero_0–RangeZero_1)>>1; or

[0083] ROFC = (2*range – RangeZero_0 – RangeZero_1)>>1; (8)

[0084] In CABAC, the method of the present invention uses a "stand-alone" context model on some syntaxes. The probability or probability state of the stand-alone context may be different from other contexts. For example, the probability or probability conversion of the "stand-alone" context model may use a different mathematical model. In one embodiment, a context model with a fixed probability may be used for the stand-alone context. In another embodiment, a context model with a fixed probability range may be used for the stand-alone context.

[0085] The above flowchart is to show an example of video encoding according to the present invention. A person skilled in the art may modify each step, rearrange the steps, split a step, or merge steps to implement the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics are used to show examples to implement embodiments of the present invention. A person skilled in the art may implement the present invention by replacing syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0086] The description presented above is intended to enable one of ordinary skill in the art to implement the present invention in the context of the specific application and its requirements as provided. Various variations of the described embodiments will be apparent to those skilled in the art, and the general principles defined therein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but is intended to be consistent with the broadest scope of the principles and novel features disclosed herein. In the above detailed description, various specific details are described to provide a thorough understanding of the present invention. Furthermore, it should be understood by those skilled in the art that the present invention is implementable.

[0087] The embodiments described in the present invention may be implemented using various hardware, software code, or a combination of both. For example, one embodiment of the present invention may be software code integrated into one or more circuits of a video compression chip or integrated into the video compression software that performs the above-mentioned process. One embodiment of the present invention may also be software code to be run on a digital signal processor (DSP) to perform the above-mentioned process. The present invention may also be related to multiple functions performed by a computer processor, a digital signal processor, a microprocessor, a field programmable gate array (FPGA), etc. These processors can be used to perform specific tasks according to the present invention by running machine-readable software code or firmware code that defines the specific method implemented by the present invention. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles and software code languages ​​and other configuration codes to perform tasks according to the present invention do not deviate from the spirit and scope of the present invention.

[0088] The present invention may be implemented in other forms without departing from its spirit or essential characteristics. The embodiments described above need to be considered comprehensively and are only intended to illustrate the present invention rather than to limit the present invention. All changes within the meaning and scope of the equivalent of the claims are within the scope of protection of the claims.

Claims

1. A video coding and decoding method using video compensation, characterized in that: The method includes: receiving input data associated with a current block or sub-block within a current image; Using template matching to derive one or more first best templates of the current template of the current block or sub-block, which are pointed to by one or more first best motion vectors in the first reference table, wherein the one or more first best motion vectors are derived according to template matching; After deriving the one or more first best templates, determining a new current template based on the current template, the one or more first best templates, or both; deriving one or more second best templates of the new current template of the current block or subblock using template matching, which are pointed to by one or more second best motion vectors in a second reference table, wherein the one or more second best motion vectors are derived according to the template matching, and wherein the first reference table and the second reference table belong to a group including LIST_0 and LIST_1, and the first reference table is different from the second reference table; determining a final motion vector or a final motion vector predictor from a set of motion vector candidates or motion vector predictor candidates, the set of motion vector candidates or motion vector predictor candidates comprising one or more best motion vectors associated with the one or more first best motion vectors and the one or more second best motion vectors; and The final motion vector or the final motion vector predictor is used to encode or decode the current block or subblock or the current motion vector of the current block or subblock at the encoder end or the decoder end, respectively.

2. The method according to claim 1, characterized in that Wherein the new current template is equal to (2*(current template)-LIST_X template), and wherein the LIST_X template corresponds to the one or more first best templates, the one or more first best templates are pointed to by one or more first best motion vectors in the first reference table, and "X" is equal to "0" or "1".

3. The method according to claim 2, characterized in that Wherein after deriving one or more second best templates, in order to derive one or more first best templates again in the first reference table, another new current template is determined based on the current template, the one or more second best templates or both, and the best template search is repeated in the first reference table and the second reference table until the repetition reaches a certain number.

4. The method according to claim 1, characterized in that Whether the new current template is allowed depends on the picture type of the picture containing the current block.

5. A video encoding and decoding device using motion compensation, characterized in that: The device comprises one or more electronic circuits or processors for: receiving input data associated with a current block or sub-block within a current image; Using template matching to derive one or more first best templates of the current template of the current block or sub-block, which are pointed to by one or more first best motion vectors in the first reference table, wherein the one or more first best motion vectors are derived according to template matching; After deriving the one or more first best templates, determining a new current template based on the current template, the one or more first best templates, or both; deriving one or more second best templates of the new current template of the current block or subblock using template matching, which are pointed to by one or more second best motion vectors in a second reference table, wherein the one or more second best motion vectors are derived according to the template matching, and wherein the first reference table and the second reference table belong to a group including LIST_0 and LIST_1, and the first reference table is different from the second reference table; determining a final motion vector or a final motion vector predictor from a set of motion vector candidates or motion vector predictor candidates, the set of motion vector candidates or motion vector predictor candidates comprising one or more best motion vectors associated with the one or more first best motion vectors and the one or more second best motion vectors; and The final motion vector or the final motion vector predictor is used to encode or decode the current block or subblock or the current motion vector of the current block or subblock at the encoder end or the decoder end, respectively.

Citation Information

Patent Citations

  • Methods for decoder-side motion vector derivation

    CN102131091A

  • Buffering prediction data in video coding

    CN103688541A