Video decoding apparatus, video coding apparatus

By prioritizing candidates and establishing a selection mechanism for equal TM costs in the bcwCandidatesList generation, the video coding technique addresses prediction accuracy challenges, enhancing codec quality and efficiency.

WO2025134860A1PCT designated stage expired Publication Date: 2025-06-26SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/043568
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-10
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing video coding techniques, such as those using the BCW method for inter prediction, face challenges in improving prediction accuracy, especially in generating the bcwCandidatesList and comparing Template Matching Costs (TM costs).

Method used

The proposed solution enhances the generation method of the bcwCandidatesList by defining priorities for optional candidates and determining which candidates are included based on these priorities. Additionally, a selection mechanism is established for cases where TM costs are equal, without adding additional calculations.

Benefits of technology

This approach improves the quality of codecs by enhancing prediction accuracy without increasing computational complexity, thereby improving the overall efficiency of video coding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024043568_26062025_PF_FP_ABST
    Figure JP2024043568_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This invention aims to enhance precision by improving the bcwIdx generation method of the BCW method. The new generation method derives templates based on current block information and generates bcwCandidatesList based on mergeBcwIdx and equal weight. Additionally, in situations where multiple candidates have the same cost, a priority is given for choosing.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO DECODING APPARATUS, VIDEO CODING APPARATUS

[0001] The embodiments of the present invention relate to, a video decoding apparatus, a video coding apparatus.

[0002] A video coding apparatus which generates coded data by coding a video, and a video decoding apparatus which generates decoded images by decoding the coded data are used for efficient transmission or recording of videos.

[0003] For example, specific video coding schemes include H.264 / AVC, High-Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) schemes, and the like.

[0004] In such a video coding scheme, images (pictures) constituting a video are managed in a hierarchical structure including slices obtained by splitting an image, coding tree units (CTUs) obtained by splitting a slice, units of coding (coding units; which is referred to as CUs) obtained by splitting a coding tree unit, and transform units (TUs) obtained by splitting a coding unit, and are coded / decoded for each CU.

[0005] In such a video coding scheme, usually, a prediction image is generated based on a local decoded image that is obtained by coding / decoding an input image (a source image), and prediction error components (which may be referred to also as “difference images” or “residual images”) obtained by subtracting the prediction image from the input image are coded. Generation methods of prediction images include an inter-picture prediction (an inter-prediction) and an intra-picture prediction (intra prediction).

[0006] In recent video coding and decoding techniques, NPL1 introduced the BCW method for inter prediction. The BCW method is used when the currently block is coded with bi-prediction. BCW will give two weights for the two reference frame RefPicL0 and RefPicL1, the final prediction results are calculated by weighted sum of the two prediction results from these two reference frame.

[0007] NPL 1: Ru-Ling Liao, Jie Chen, Yan Ye, Xinwei Li (Alibaba), “EE2-2.2: Template matching based BCW index derivation for merge mode”, JVET-AB0079, JVET 28th Meeting, Mainz, DE, 20-28 October 2022

[0008] BCW (Bi-prediction with CU-level weight) is a method designed to enhance the results of bi-prediction and is used in inter prediction. For non-merge coded Coding Units (CUs), the BCW index is signaled. In the case of merge-coded CUs, the BCW index is inferred from neighboring blocks. NPL1 introduced a method that derives the BCW index for merge-coded CUs based on Template Matching Cost (TM cost). When the current block is a merge-coded block, BCW generates a candidate list called bcwCandidatesList, which contains weights assigned to RefPicL0. The final decision on which weight to use is determined by comparing the SAD costs. This invention aims to improve prediction accuracy by enhancing the process of generating the bcwCandidatesList and the comparison process of TM costs.

[0009] The aim of this invention is to improve prediction accuracy by modifying the generation method of the bcwCandidateslist. In this invention, we define priorities for optional candidates and determine which candidates are included in the candidate list based on these priorities. Additionally, we establish a selection mechanism for candidates in cases where their TM costs are equal.

[0010] According to an aspect of the present invention, the quality of the codecs can be improved without adding additional calculations.

[0011] FIG. 1 is a schematic diagram illustrating a configuration of an image transmission system according to the present embodiment.FIG. 2 is a diagram showing the hierarchical structure of the coded stream data.FIG. 3 is a schematic diagram of the video decoding apparatus.FIG. 4 shows the structure of the inter prediction parameter derivation unit.FIG. 5 is a schematic diagram of the inter prediction image generation unit.FIG. 6 is a diagram showing the details of the inter prediction parameter coding unit.FIG. 7 is a block diagram showing the structure of a video coding apparatus.FIG. 8 is a diagram showing the structure of BCW parameter derivation unit.FIG. 9 shows the template area of target(currently) block.FIG. 10 shows the relationship between bcwIdx and weights.

[0012] First Embodiment Hereinafter, embodiments of the present disclosure is described with reference to the drawings.

[0013] FIG. 1 is a schematic diagram illustrating a configuration of an image transmission system 1 according to the present embodiment.

[0014] The image transmission system 1 is a system in which a coding stream obtained by coding a coding target image is transmitted, the transmitted coding stream is decoded, and an image is displayed. The image transmission system 1 includes a video coding apparatus (image coding apparatus) 11, a network 21, a video decoding apparatus (image decoding apparatus) 31, and a video display apparatus (image display apparatus) 41.

[0015] An image T is input to the video coding apparatus 11.

[0016] The network 21 transmits a coding stream Te generated by the video coding apparatus 11 to the video decoding apparatus 31. The network 21 is the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or a combination thereof. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network configured to transmit broadcast waves of digital terrestrial television broadcasting, satellite broadcasting or the like. Furthermore, the network 21 may be substituted by a storage medium in which the coding stream Te is recorded, such as a Digital Versatile Disc (DVD: trademark) or a Blu-ray Disc (BD: trademark).

[0017] The video decoding apparatus 31 decodes each of the coding streams Te transmitted from the network 21 and generates one or multiple decoded images Td which are decoded.

[0018] The video display apparatus 41 displays all or part of the one or multiple decoded images Td generated by the video decoding apparatus 31. For example, the video display apparatus 41 includes a display device such as a liquid crystal display and an organic Electro-Luminescence (EL) display. Forms of the display include a stationary type, a mobile type, an HMD type, and the like. In addition, in a case that the video decoding apparatus 31 has a high processing capability, an image having high image quality is displayed, and in a case that the apparatus only has a lower processing capability, an image which does not require high processing capability and display capability is displayed.

[0019] Operator Operators and notations used in the present specification is described below.

[0020] >> is an arithmetic right bit shift, << is an arithmetic left bit shift, & is a bitwise AND, | is a bitwise OR, ^ is a bitwise XOR, |= is an OR assignment operator, and || indicates a logical sum.

[0021] x ? y : z is a ternary operator to take y in a case that x is true (other than 0) and take z in a case that x is false (0).

[0022] Clip3(x, y, z) is a function to clip z in a value equal to or greater than x and less than or equal to y, and a function to return x in a case that z is less than x (z < x), return y in a case that z is greater than y (z > y), and return z in other cases.

[0023] abs(a) is a function that returns the absolute value of a.

[0024] Int(a) is a function that returns the integer value of a.

[0025] floor(a) is a function that returns the maximum integer equal to or less than a.

[0026] ceil(a) is a function that returns the minimum integer equal to or greater than a.

[0027] a / d represents division of a by d (round down decimal places).

[0028] x = y..z represents x takes on integer values starting from y to z, inclusive, with x, y, and z being integer numbers and z being greater than or equal to y.

[0029] Structure of Coding Stream Te Prior to the detailed description of the video coding apparatus 11 and the video decoding apparatus 31 according to the present embodiment, a data structure of the coding stream Te generated by the video coding apparatus 11 and decoded by the video decoding apparatus 31 is described.

[0030] FIG. 2 is a diagram illustrating a hierarchical structure of data of the coding stream Te. The coding stream Te includes a sequence and multiple pictures constituting the sequence illustratively. (a) to (f) of FIG. 2 are diagrams illustrating a coded video sequence defining a sequence SEQ, a coded picture prescribing a picture PICT, a coding slice prescribing a slice S, a coding slice data prescribing slice data, a coding tree unit included in the coding slice data, and a coding unit (CU) included in each coding tree unit, respectively.

[0031] Coded Video Sequence In the coded video sequence (CVS, coding stream), a set of data referred to by the video decoding apparatus 31 to decode the coded sequence sequences to be processed is defined. As illustrated in FIG. 2, the CVS includes a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a picture (PICT), and Supplemental Enhancement Information (SEI).

[0032] In the video parameter set VPS, in a video including multiple layers, a set of coding parameters common to multiple videos and a set of coding parameters associated with the multiple layers and an individual layer included in the video are defined.

[0033] In the sequence parameter set SPS, a set of coding parameters referred to by the video decoding apparatus 31 to decode a target sequence is defined. For example, a width and a height of a picture are defined. Note that multiple SPSs may exist. In that case, any of multiple SPSs is selected from the PPS.

[0034] In the picture parameter set PPS, a set of coding parameters referred to by the video decoding apparatus 31 to decode each picture in a target sequence is defined. For example, a reference value (pic_init_qp_minus26) of a quantization step size used for decoding of a picture and a flag (weighted_pred_flag) indicatingan application of a weighted prediction are included. Note that multiple PPSs may exist. In that case, any of multiple PPSs is selected from each picture in a target sequence.

[0035] Coded Picture In the coded picture, a set of data referred to by the video decoding apparatus 31 to decode the picture PICT to be processed is defined. As illustrated in FIG. 2, the picture PICT includes a slice 0 to a slice NS-1 (NS is the total number of slices included in the picture PICT).

[0036] Note that in a case that it is not necessary to distinguish each of the slice 0 to the slice NS-1 below, subscripts of reference signs may be omitted. In addition, the same applies to other data with subscripts included in the coding stream Te which is described below.

[0037] Coding Slice In the coding slice, a set of data referred to by the video decoding apparatus 31 to decode the slice S to be processed is defined. As illustrated in FIG. 2, the slice includes a slice header and a slice data.

[0038] The slice header includes a coding parameter group referred to by the video decoding apparatus 31 to determine a decoding method for a target slice. Slice type specification information (slice_type) indicating a slice type is one example of a coding parameter included in the slice header.

[0039] Examples of slice types that may be specified by the slice type specification information include (1) I slice using only an intra prediction in coding, (2) P slice using a unidirectional prediction or an intra prediction in coding, and (3) B slice using a unidirectional prediction, a bidirectional prediction, or an intra prediction in coding, and the like. Note that the inter prediction is not limited to a uni-prediction and a bi-prediction, and the prediction image may be generated by using a larger number of reference pictures. Hereinafter, in a case that a slice is referred to as the P or B slice, the slice indicates a slice that includes a block in which the inter prediction may be used.

[0040] Note that, the slice header may include a reference to the picture parameter set PPS (pic_parameter_set_id).

[0041] Coding Slice Data In the coding slice data, a set of data referred to by the video decoding apparatus 31 to decode the slice data to be processed is defined. The slice data include CTUs as illustrated in FIG. 2. The CTU is a block of a fixed size (for example, 64 x 64) constituting a slice.

[0042] Coding Tree Unit In FIG. 2, a set of data referred to by the video decoding apparatus 31 to decode the CTU to be processed is defined. The CTU is split into coding units CUs, each of which is a basic unit of coding processing, by a recursive Quad Tree split (QT split), Binary Tree split (BT split), or Ternary Tree split (TT split). The BT split and the TT split are collectively referred to as a Multi Tree split (MT split). Nodes of a tree structure obtained by recursive quad tree splits are referred to as Coding Nodes. Intermediate nodes of a quad tree, a binary tree, and a ternary tree are coding nodes, and the CTU itself is also defined as the highest coding node.

[0043] Coding Unit As illustrated in FIG. 2, a set of data referred to by the video decoding apparatus 31 to decode the coding unit to be processed is defined. Specifically, the CU includes a CU header CUH, a prediction parameter, a transform parameter, a quantization transform coefficient, and the like. In the CU header, a prediction mode and the like are defined.

[0044] There are cases that the prediction processing is performed in units of CU or performed in units of sub-CU obtained by further splitting the CU. In a case that the sizes of the CU and the sub-CU are equal to each other, the number of sub-CUs in the CU is one. In a case that the CU is larger in size than the sub-CU, the CU is split into sub-CUs. For example, in a case that the CU has a size of 8 x 8, and the sub-CU has a size of 4 x 4, the CU is split into four sub-CUs which include two horizontal splits and two vertical splits.

[0045] There are two types of predictions (prediction modes), which are an intra prediction and an inter prediction. The intra prediction refers to a prediction in an identical picture, and the inter prediction refers to prediction processing performed between different pictures (for example, between pictures of different display times).

[0046] Transform and quantization processing is performed in units of CU, but the quantization transform coefficient may be subjected to entropy coding in units of subblock such as 4 x 4.

[0047] Prediction parameter A prediction image is derived by a prediction parameter accompanying a block. The prediction parameter includes prediction parameters of the intra prediction and the inter prediction.

[0048] Inter prediction parameter.

[0049] The prediction parameter of the inter prediction is described below. The inter prediction parameter includes prediction list utilization flags predFlag0 and predFlag1, reference picture indices, refIdxL0 and refIdxL1, and motion vectors, mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether the reference picture lists (L0 list, L1 list) are used, and when the value is 1, the corresponding reference picture list is utilized. In this specification, when stating "a flag indicating whether XX is true or not," if the flag is anything other than 0 (e.g., 1), XX is considered true, and if the flag is 0, XX is considered false. Logical negation, logical AND, etc., treat 1 as true and 0 as false (similarly below). However, in actual devices or methods, other values can be used as true or false.

[0050] The syntax elements used to derive inter-prediction parameters include, for example, the merge_flag (general_merge_flag), merge_idx, merge_subblock_flag, regular_merge_flag, ciip_flag, merge_gpm_partition_index, merge_gpm_idx0, merge_gpm_idx1, inter_pred_idc, reference picture indices refIdxLX, motion vector predictor index mvp_LX_idx, motion vector difference mvdLX, and motion vector precision mode amvr_mode. The merge_subblock_flag indicates whether sub-block level inter prediction is used. The regular_merge_flag is a flag indicating whether regular merge mode or MMVD is used. The ciip_flag indicates whether CIIP (combined inter-picture merge and intra-picture prediction) mode or GPM mode (Geometric Partitioning Merge mode) is used. The merge_gpm_partition_idx is an index indicating the partition shape in GPM mode. The merge_gpm_idx0 and merge_gpm_idx1 are indices indicating the merge indices in GPM mode. The inter_pred_idc is an identifier used in AMVP mode to select the reference picture. The mvp_LX_idx is the index for predicting vectors used to derive motion vectors.

[0051] Reference picture list The reference picture list is a list composed of reference pictures stored in the reference picture memory (RefPicList). For each Coding Unit (CU), the refIdxLX specifies which picture from the reference picture list RefPicListX (X=0 or 1) stored in the reference picture memory 306 is actually referenced. Here, LX is a notation used for distinguishing between L0 prediction and L1 prediction. Subsequently, LX is replaced with L0 and L1 to differentiate parameters for the L0 list and the L1 list.

[0052] Merge prediction and AMVP prediction The decoding (encoding) process for prediction parameters includes Merge mode and Advanced Motion Vector Prediction (AMVP) mode. The general_merge_flag is a flag used to identify these modes. Merge mode is a prediction mode that omits some or all of the motion vector differences. It derives prediction parameters from already processed neighboring block prediction parameters without including prediction list usage flags (predFlagLX), reference picture indices (refIdxLX), and motion vectors (mvLX) in the encoded data. AMVP mode includes inter_pred_idc, refIdxLX, and mvLX in the encoded data. Note that mvLX is encoded as the prediction vector mvpLX identifier (mvp_LX_idx) and the difference vector mvdLX. The general term for prediction modes that omit or simplify motion vector differences is called General Merge Mode, and it can be selected using the general_merge_flag along with AMVP prediction.

[0053] When general_merge_flag is 1, the regular_merge_flag can be transmitted. If regular_merge_flag is 1, normal merge mode or MMVD can be chosen; otherwise, CIIP mode or GPM mode can be selected. CIIP mode generates a prediction image by weighted summation of inter-predicted and intra-predicted images. GPM mode generates a prediction image by dividing the target CU into two non-rectangular prediction units. The inter_pred_idc is a value indicating the type and number of reference pictures; it can take one of the values PRED_L0, PRED_L1, or PRED_BI. PRED_L0 and PRED_L1 indicate single predictions using one reference picture each from the L0 list and L1 list, respectively. PRED_BI indicates a two way prediction (bi-prediction) using two reference pictures managed in both the L0 and L1 lists. The merge_idx is an index indicating which prediction parameter (merge candidate) in prediction parameter candidates should be used as the prediction parameter for the target block. The prediction parameter candidates are derived from processed blocks.

[0054] Motion vector mvLX indicates the shift amount between blocks on two different pictures. The prediction vector and difference vector for mvLX are respectively called mvpLX and mvdLX.

[0055] The relationship between inter_pred_idc and prediction list usage flags predFlagLX inter_pred_idc and predFlagL0, predFlagL1 are related as follows and are mutually convertible: inter_pred_idc = (predFlagL1<<1)+predFlagL0   predFlagL0 = inter_pred_idc & 1   predFlagL1 = inter_pred_idc >> 1 It's important to note that inter prediction parameters can be determined using prediction list usage flags (predFlagLX) or inter prediction identifiers (inter_pred_idc). The determination based on prediction list usage flags (predFlagLX) can be replaced with determination based on inter prediction identifiers (inter_pred_idc), and vice versa.

[0056] The determination of the bi-directional prediction The flag biPred, indicating whether it is a bi-directional prediction, can be derived by checking if both prediction list usage flags (predFlagL0 and predFlagL1) are set to 1. Alternatively, biPred can be determined based on whether the inter prediction identifier (inter_pred_idc) indicates the use of two prediction lists (reference pictures).

[0057] Configuration of video decoding apparatus A configuration of the video decoding apparatus 31 (FIG. 3) according to the present embodiment is described.

[0058] The video decoding apparatus 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding apparatus) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit 308, an inverse quantization and inverse transform processing unit 311, an addition unit 312, and a prediction parameter derivation unit 320. Note that a configuration in which the loop filter 305 is not included in the video decoding apparatus 31 is also used in accordance with the video coding apparatus 11 described later.

[0059] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes, from coded data, parameter set information such as the VPS, the SPS, and the PPS, and a slice header (slice information). The CT information decoding unit 3021 decodes a CT from coded data. The CU decoding unit 3022 decodes a CU from coded data. In a case that a TU includes a prediction error, the TU decoding unit 3024 decodes QP update information (quantization correction value) and a quantization prediction error (residual_coding) from coded data.

[0060] The prediction image generation unit 308 is composed of an inter prediction image generation unit 309 (FIG.5) and an intra prediction image generation unit 310. The prediction parameter derivation unit 320 is composed of an inter prediction parameter derivation unit 303 (FIG.4) and an intra prediction parameter derivation unit.

[0061] Furthermore, an example in which a CTU and a CU are used as units of processing is described below, but the processing is not limited to this example, and processing in units of sub-CU may be performed. Alternatively, by replacing the CTU and the CU by a block and replacing the sub-CU by a subblock, and processing in units of blocks or subblocks may be performed.

[0062] The entropy decoding unit 301 performs entropy decoding on the coding stream Te input from the outside and separates and decodes individual codes (syntax elements). The separated codes include prediction information to generate a prediction image, a prediction error to generate a difference image, and the like. Entropy coding has a variable length coding method for syntax elements according to the context (probability model) adaptively selected according to the type of syntax elements and the surrounding conditions, and a variable length coding method for syntax elements using a predetermined table or formula.

[0063] The entropy decoding unit 301 outputs the decoded codes to the parameter decoding unit 302. The decoded codes include, for example, prediction mode (predMode), general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, and so on. The control of which codes to decode is performed based on the instructions from the parameter decoding unit 302.

[0064] The parameter decoding unit 302 notifies the entropy decoding unit 301 of which syntax elements need be decoded. The entropy decoding unit 301 outputs the syntax element to the prediction parameter derivation unit 320.

[0065] The prediction parameter derivation unit 320 may derive the prediction parameters based on the output of the paremater decoding unit 302 and the prediction parematers which saved in the prediction parameter memory 307. The derived prediction parameters is output into the prediction image generation unit 308 and also is saved in the prediction parameter memory 307.

[0066] The loop filter 305 is a filter provided in the coding loop, and is a filter that removes block distortion and ringing distortion and improves image quality. The loop filter 305 applies a filter such as a deblocking filter, a Sample Adaptive Offset (SAO), and an Adaptive Loop Filter (ALF) on a decoded image of a CU generated by the addition unit 312.

[0067] The reference picture memory 306 stores the decoded image of the CU generated by the addition unit 312 in a predetermined position for each target picture and target CU.

[0068] The prediction parameter memory 307 stores prediction parameters in a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores a parameter derived by the prediction parameter derivation unit 320, a prediction mode predMode separated by the entropy decoding unit 301, and the like.

[0069] The prediction image generation unit 308 receives input of the prediction parameter derived by the prediction parameter deviation unit 320, and the like. In addition, the prediction image generation unit 308 reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or a subblock by using the prediction parameter and the read reference picture (reference picture block) in the prediction mode indicated by the prediction mode predMode. Here, the reference picture block refers to a set of pixels (referred to as a block because they are normally rectangular) on a reference picture and is a region that is referred to generate a prediction image.

[0070] Inter Prediction Parameter Derivation Unit Configuration As shown in FIG.4, the Inter Prediction Parameter Derivation Unit 303 derives inter prediction parameters based on syntax elements input from the parameter decoding unit 302. It references the prediction parameters stored in the prediction parameter memory 307 and outputs them to the Inter Prediction Image Generation Unit 309 and the prediction parameter memory 307. The Inter Prediction Parameter Derivation Unit 303, along with its internal elements such as the AMVP Prediction Parameter Derivation Unit 3032, Merge Prediction Parameter Derivation Unit 3036, MV Addition Unit 3038, and BCW (Bi-prediction with CU-level Weights) parameter derivation unit 3039 are common parts in video encoding and decoding devices. Therefore, they can be collectively referred to as the Motion Vector Derivation Unit (Motion Vector Derivation Device).

[0071] When general_merge_flag is 1, indicating the merge prediction mode, merge_idx is derived, and the Merge Prediction Parameter Derivation Unit 3036 outputs it. When general_merge_flag is 0, indicating the AMVP prediction mode, the AMVP Prediction Parameter Derivation Unit 3032 derives mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx. The MV Addition Unit 3038 adds the derived mvpLX and mvdLX to obtain mvLX.

[0072] Merge Prediction The Merge Prediction Parameter Derivation Unit 3036 includes the Merge Candidate Derivation Unit 30361 and the Merge Candidate Selection Unit 30362. Merge candidates, including prediction parameters (predFlagLX, mvLX, refIdxLX), are constructed and stored in the merge candidate list. The merge candidates in the merge candidate list are assigned indices according to predetermined rules.

[0073] The Merge Candidate Derivation Unit 30361 derives merge candidates using the motion vectors and refIdxLX of the decoded neighboring blocks. Additionally, the Merge Candidate Derivation Unit 30361 may apply spatial merge candidate derivation processing, temporal merge candidate derivation processing, etc., as explained later.

[0074] As part of the spatial merge candidate derivation process, the Merge Candidate Derivation Unit 30361 reads prediction parameters stored in the prediction parameter memory 307 according to predetermined rules and sets them as merge candidates. For example, it reads the prediction parameters at the neighboring positions A1, B1, B0, A0, B2 as shown in FIG.9(a).

[0075] A1: (xCb-1, yCb+cbHeight-1) B1: (xCb+cbWidth-1, yCb-1) B0: (xCb+cbWidth, yCb-1) A0: (xCb-1, yCb+cbHeight) B2: (xCb-1, yCb-1) Here, (xCb, yCb) denotes the top-left coordinates of the target block, and cbWidth and cbHeight denote the width and height of the target block.

[0076] For the temporal merge derivation process, the Merge Candidate Derivation Unit 30361 reads the prediction parameters for the block C in the reference image, which includes the right-bottom CBR or the central coordinates of the target block. These parameters are read from the prediction parameter memory 307 and stored as the merge candidate "Col" in the merge candidate list mergeCandList[].

[0077] The order of storing in mergeCandList[] can be, for example, spatial merge candidates (B1, A1, B0, A0, B2) followed by the temporal merge candidate "Col." Note that unavailable reference blocks (e.g., intra prediction reference block) are not stored in the merge candidate list.

[0078] i = 0 if(availableFlagB1)  mergeCandList[i++] = B1 if(availableFlagA1)  mergeCandList[i++] = A1 if(availableFlagB0)  mergeCandList[i++] = B0 if(availableFlagA0)  mergeCandList[i++] = A0 if(availableFlagB2)  mergeCandList[i++] = B2 if(availableFlagCol)  mergeCandList[i++] = Col Additionally, historical merge candidates HmvpCand, pairwise average candidates avgCand, and zero merge candidates zeroCandm can be added to mergeCandList[].The Merge Candidate Selection Unit 30362 selects a merge candidate N from the merge candidate list based on the merge_idx as follows: N = mergeCandList[merge_idx] Here, N is a label representing the merge candidate, taking values such as A1, B1, B0, A0, B2, Col, etc. The motion information of the merge candidate represented by the label N is given by (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN.

[0079] The Merge Candidate Selection Unit 30362 uses merge_idx to select (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN as the inter prediction parameters for the target block. The selected merge candidate's inter prediction parameters are stored in the prediction parameter memory 307 and output to the Inter Prediction Image Generation Unit 309.

[0080] MMVD Prediction Unit 30376 The MMVD Prediction Unit 30376 handles processing in the MMVD (Merge with Motion Vector Difference) mode. In merge prediction, motion vectors obtained from neighboring blocks are used as merge candidates' motion vectors. The MMVD mode refines motion vectors by adding specified distance and direction difference vectors to the merge candidates' motion vectors. In the MMVD mode, the MMVD Prediction Unit 30376 uses merge candidates and restricts the value range of difference vectors to specified distances (e.g., 6 variations, 8 variations) and directions (e.g., 4 directions, 8 directions, 16 directions) to efficiently derive motion vectors.

[0081] BCW parameter derivation unit 3039 The BCW parameter derivation unit 3039 derives / selects a bcwIdx, indicating the selection of which weight from a pre-defined set of weights to use for the current bi-prediction. The BCW parameter derivation unit 3039 is used only when the current block is subjected to bi-prediction. In such cases, the current block has two reference pictures, RefPicL0 and RefPicL1, as well as two motion vectors, MvL0 and MvL1. When the current block is subjected to uni-prediction, the bcwIdx may be derived as the default value (e.g. bcwIdx = bcwIdxEqual = 2). bcwIdxEqual is the bcwIndex where the weight of prediction image using RefPicL0 (weightPredL0) is equal to the weight of prediction image using RefPicL1 (weightPredL1).

[0082] The BCW parameter derivation unit 30391 comprises the Reference sample derivation unit 3091, Template derivation unit 30392, and BCW weights derivation device 30393. The BCW weights derivation device 30393 includes the Candidates derivation unit 303931, Weight derivation Unit 303932, Template prediction image generation unit 303933, Template cost derivation unit 303934, and Choosing unit 303935.

[0083] Template derivation unit 3092 Template derivation unit 3092 derives two templates, tempSampleTop and tempSampleLeft, for the current block. The templates, tempSampleTop and tempSampleLeft, store the already reconstructed pixel values. The width of tempSampleTop, denoted as tempSampleTopWidth, is equal to the current block's width (curBlockWidth), and the length, tempSampleTopHeight, is equal to 1. The width of tempSampleLeft, denoted as tempSampleLeftWidth, is set equal to 1, and the length, tempSampleLeftHeight, is set equal to the current block's height (curBlockHeight). FIG.9 illustrates the relationship between the templates and the current block. In FIG. 9 is just an example, and in encoding or decoding devices, the size of the current block is not limited to 8x8; it can be of other sizes and shapes.

[0084] tempSampleTopHeight and tempSampleLeftWidth may be determined based on different methods according to the information of the current block. For example: Method 1. Set tempSampleTopHeight and tempSampleLeftWidth to the same fixed value.

[0085] Method 2. Set tempSampleTopHeight and tempSampleLeftWidth to different fixed values.

[0086] Method 3. Set tempSampleTopHeight and tempSampleLeftWidth to different values based on the size of the current block.

[0087] Method 4. Set tempSampleTopHeight and tempSampleLeftWidth to different values based on the shape of the current block.

[0088] Specifically, each method may be constrcuted as follows: (Method 1: same fixed value) Set the values of tempSampleTopHeight and tempSampleLeftWidth to a fixed value n (tempSampleTopHeight = tempSampleLeftWidth = n). The fixed value n may be a positive integer [1, 2, 3, 4,...]. The value of n is not greater than the maximum of the current block's height and width.

[0089] (Method 2: different fixed value) Set the values of tempSampleTopHeight and tempSampleLeftWidth to two fixed values n1 and n2 (tempSampleTopHeight = n1, tempSampleLeftWidth = n2) (n1 != n2). n1 and n2 may be a positive integers [1, 2, 3, 4,...]. The values of n1 and n2 are not greater than the current block's height and width.

[0090] Method 3: different values based on the size of the current block) Set the values of tempSampleTopHeight and tempSampleLeftWidth based on the size of the current block. Template derivation unit 3092 may derive tempSampleTopHeight as the larger value when the current block width increases. Template derivation unit 3092 may derive tempSampleLeftWidth as the larger value when the current block height increases. For example, set a threshold list sh = [8] and a value list L = [2, 4]. When the width of the current block is greater than or equal to sh[0], tempSampleTopHeight = L[1]; otherwise, tempSampleTopHeight = L[0]. When the height of the current block is greater than or equal to sh[0], tempSampleLeftWidth = L[1]; otherwise, tempSampleLeftWidth = L[0]. The number of thresholds (elements in sh) isthe number of heights (elements in L) minus 1. For example, sh = [8, 16, 32], L = [1, 2, 3, 4]. In this case, the value of tempSampleTopHeight is determined as follows (consistent with the rule for tempSampleLeftWidth): if (curBlockWidth >= sh[0]) tempSampleTopHeight = L[1]; else if (curBlockWidth >= sh[1]) tempSampleTopHeight = L[2]; else if (curBlockWidth >= sh[2]) tempSampleTopHeight = L[3]; else tempSampleTopHeight = L[0]; Template derivation unit 3092 may derive tempSampleTopHeight and tempSampleTopLeft by arithmetic calculations, e.g. right shifting, division. For example, the following applies.

[0091] tempSampleTopHeight = Min (curBlockWidth >> 3, 4) tempSampleLeftWidth = Min (curBlockHeight >> 3, 4) More generally, the following applies.

[0092] tempSampleTopHeight = Min ((curBlockWidth + offset) >> shift, maxLen) tempSampleLeftWidth = Min ((curBlockHeight + offset)>> shiftt, maxLen) or tempSampleTopHeight = Min ((curBlockWidth + offset) / Divt, maxLen) tempSampleLeftWidth = Min ((curBlockHeight + offset) / Divt, maxLen) (Method 4: different values based on the shape of the current block) Set the values of tempSampleTopHeight and tempSampleLeftWidth based on the shape of the current block. More specifically, Template derivation unit 3092 may derive them based on the aspect ratio rate of the current block Template derivation unit 3092 may derive tempSampleTopHeight and tempSampleLeftWidth as the same value if the current block is square (i.e. curBlockHeight == curBlockWidth). Template derivation unit 3092 may derive tempSampleTopHeight and tempSampleLeftWidth as the same value if the current block is square (i.e., curBlockHeight == curBlockWidth). When the current block is vertically long (i.e., curBlockHeight > curBlockWidth), Template derivation unit 3092 may derive tempSampleTopHeight as a larger value than tempSampleLeftWidth. Similarly, when the current block is horizontally long (i.e., curBlockHeight < curBlockWidth), Template derivation unit 3092 may derive tempSampleTopHeight as a smaller value than tempSampleLeftWidth. It can also be said that Template derivation unit 3092 may derive tempSampleLeftWidth as smaller value than tempSampleTopHeight when the current block is horozonzztally long (i.e. curBlockHeight < curBlockWidth).

[0093] Template derivation unit 3092 may derive tempSampleTopHeight and tempSampleLeftWidth to satisfy the condition tempSampleTopHeight / tempSampleLeftWidth = curBlockHeight / curBlockWidth. A specific implementation method is as follows: aspectRatio may be clipped as follows.

[0094] aspectRatio = Min (aspectRatio, MAX_RATIO) MAX_RATIO may be 3, 4. A variable valBase may take other positive integer value greater than 0.

[0095] Reference sample derivation unit 3091 Reference sample derivation unit 30391 obtains refSampleTopL0 and refSampleLeftL0 from RefPicL0 based on MvL0. Similarly, it obtains refSampleTopL1 and refSampleLeftL1 from RefPicL1 based on MvL1. It is important to note that the sizes of refSampleTopL0 and refSampleTopL1 must be consistent with tempSampleTop. Similarly, the sizes of refSampleLeftL0 and refSampleLeftL1 is the same as tempSampleLeft.

[0096] Candidates derivation unit 303931 Candidates derivation unit 303931 may derives bcwCandidatesList, which stores the available weights. BCW has a predefined weight list WeightList = [1,3,4,5,7,2,6], where the selectable weights are saved. The bcwIdx represents the index of the currently selected weight, and its values are integers in the closed interval [1,7].

[0097] bcwCandidatesList stores possible values for bcwIdx, and its elements can be derived based on different methods, such as: Method C1. All elements in WeightList.

[0098] Method C2. Derive bcwCandidatesList according to mergeBcwIdx.

[0099] Method C3. Always use / add the equal weight candidate in bcwCandidatesList.

[0100] Each method may be constrcuted as follows: (Method C1: All elements in WeightList) Candidates derivation unit 303931 derives bcwCandidatesList such that bcwCandidatesList[i] = WeightList[i] for I = 0..numCand -1. In this method, the elements in bcwCandidatesList are the same as the elements in WeightList.

[0101] (Method C2: Derive bcwCandidatesList according to mergeBcwIdx) mergeBcwIdx is the bcwIdx used in the reference block selected during merge prediction. If N is the selected merge position where the candidate of merge posision may be A1, B1, B0, A0, or B2. The mergeBcwIdx is set equal to bcwIdx of N.

[0102] Candidates derivation unit 303931 may derive bcwCandidatesList such that mergeBcwIdx and its surrounding values are included in bcwCandidatesList. More specifically based on a threshold, n, the candidates with weights within the positive or negative n range of the weight corresponding to mergeBcwIdx may be added to bcwCandidatesList. For example, with n=1 and mergeBcwIdx=6 (corresponding weight=2), the candidates with weights corresponding weight=1 (=2-1) or bcwIdx=4, and corresponding weight=3 (=2+1) or bcwIdx=3, are added to bcwCandidatesList, resulting in bcwCandidatesList=[6,4,3]. When mergeBcwIdx corresponds to a weight equal to 1, candidates with weights within the positive 2n range of the weight corresponding to mergeBcwIdx may be added to bcwCandidatesList. When mergeBcwIdx corresponds to a weight equal to 7, candidates with weights within the negative 2n range of the weight corresponding to mergeBcwIdx may be added to bcwCandidatesList. For instance, with n=1 and mergeBcwIdx=0 (corresponding weight=7), the candidates with weights 6 (7-1) and 5 (7-2) are added to bcwCandidatesList, resulting in bcwCandidatesList=[0,5,1].

[0103] When generating bcwCandidatesList, there are several possible orders for adding candidates to the bcwCandidatesList. Taking the case of mergeBcwIdx = 5 as an example: 1.When generating bcwCandidatesList, mergeBcwIdx may be placed in it first, and then the candidates calculated based on the threshold may be added. The order of candidates calculated based on the threshold may be exchanged. In this case, bcwCandidatesList = [5, 0, 1] or [5, 1, 0].

[0104] 2.The candidates calculated based on the threshold may be added first, and then mergeBcwIdx may be included. The order of candidates calculated based on the threshold may be exchanged. In this case, bcwCandidatesList = [1, 0, 5] or [0, 1, 5].

[0105] 3.All candidates may be added to bcwCandidatesList based on their weights value. In this case, bcwCandidatesList = [0, 5, 1].

[0106] 4.Candidates may be added to bcwCandidatesList based on the position of their weights in WeightList. In this case, bcwCandidatesList = [0, 1, 5] or [5, 1, 0].

[0107] (Method C3: Always use / add the equal weight candidate in bcwCandidatesList) Based on the statistics that the equal weight candidate where weightPredL0 is equal to weightPredL1 (bcwIdx= bcwIdxEqual, e.g. 2, weightPredL0=weightPredL1=4) are selected more often than other cases. In this method, on top of Method C2, priority is given to the situation where bcwIdx=bcwIdxEqual. Candidates derivation unit 303931 derives bcwCandidatesList such that bcwIdxEqual is always added to bcwCandidatesList.

[0108] As one example, Candidates derivation unit 303931 may add bcwIdxEqual to the first position of bcwCandidatesList such that bcwCandidatesList[0] = bcwIdxEqual.

[0109] As other example, Candidates derivation unit 303931 may add bcwIdxEqual to the last position of bcwCandidatesList if there’s no bcwIdxEqual in bcwCandidatesList. The process may be as follows.

[0110] found = false; For i..numCand - 1, If (bcwCandidatesList[i] == bcwIdxEqual) found = true; If found = false, bcwCandidatesList[numCand - 1] = bcwIdxEqual; where numCand is the length of bcwCandidatesList or a constant value.

[0111] Template prediction image generation unit 303933 The Template Prediction Image Generation Unit 303933 iterates through the index saved in bcwCandidatesList, selecting weights based on the chosen index, and calculates refSampleTop[] and refSampleLeft[] for each index. The number of elements in bcwCandidatesList are denoted by num, and the index is represented as bcwIdx = bcwCandidatesList[i] (where i is in the range of [0,numCand -1].

[0112] bcwIdx = bcwCandidatesList[i] (i>= 0 and i < numCand) refSampleTop[bcwIdx] = ((8-WeightList[bcwIdx]) * refSampleTop0 + WeightList[bcwIdx] * refSampleTop1 + 4) >> 3 refSampleLeft[bcwIdx] = ((8-WeightList[bcwIdx]) * refSampleLeft0 + WeightList[bcwIdx] * refSampleLeft1 + 4) >> 3 Template Cost Derivation Unit 303934 The Template Cost Derivation Unit 303934 iterates through refSampleTop[] and refSampleLeft[], calculating the costs for all candidate samples based on the already reconstructed template regions tempSampleTop and tempSampleLeft. It computes the SAD costs for template areas, resulting in topCost[] and leftCost[], and finally determines the overall cost as the sum of topCost[] and leftCost[], represented as costList[].

[0113] The Template Cost Derivation Unit 303934 may set priorities by reducing their costs of specific candidates depending on bcwIdx. Specifically, let curBcwIdxCost be the cost calculated from the weight corresponding to the current bcwIdx.

[0114] (Case-1) If this bcwIdx equal to mergeBcwIdx, then curBcwIdxCost = curBcwIdxCost - curBcwIdxCost*W1 / 32 or curBcwIdxCost = (curBcwIdxCost*WS1)>>shiftVal. where W1 may be 1 .. 4. WS2 may be (1<<shiftVal) - (W1), shiftVal = 3, 4, 5 or 6 (Case-2) If bcwIdx equal to bcwIdxEqual (equal weights), then curBcwIdxCost = curBcwIdxCost - curBcwIdxCost*W1 / 32 or curBcwIdxCost = (curBcwIdxCost*WS1)>>shiftVal. bcwIdxEqual may be 2 but not limited to 2.

[0115] (Case3(Case1 && Case2)) If both conditions are satisfied, then curBcwIdxCost = curBcwIdxCost - curBcwIdxCost*W2 / 32 or curBcwIdxCost = (curBcwIdxCost*WS2)>>shiftVal.

[0116] where W2 > W1. WS2 may be (1<<shiftVal) - (W2) In practical implementation, may choose to apply only case 1, case 2, or both. Regarding the settings of W1 and W2, there are various options. W1 can take an integer value within the closed interval [0, 32]. When W1 = 0, curBcwIdxCost = curBcwIdxCost (remains unchanged). When W1 = 32, curBcwIdxCost = 0. For other values of W1, the formula curBcwIdxCost = curBcwIdxCost - curBcwIdxCost * W1 / 32 is used for calculation. W2 needs to be greater than W1, and it can also take any value in the closed interval [0, 32].This operations reduce the cost for these cases and increases their likelihood of being selected.

[0117] Choosing Unit 303935 The Choosing Unit 303935 selects the bcwIdx with the minimum cost from the obtained CostList[] and outputs it to the Inter Prediction Image Generation Unit 309. It's worth noting that when there are multiple candidates with the same cost, the priority for selection is as follows: 1.Select candidates that are both mergeBcwIdx and equal weight (bcwIdx == mergeBcwIdx && bcwIdx == 2).

[0118] 2.Select candidates with mergeBcwIdx (bcwIdx == mergeBcwIdx).

[0119] 3.Select candidates with equal weight (bcwIdx == bcwIdxEqual).

[0120] 4.If none of the above conditions are met, select based on their order in CostList[].

[0121] In specific implementations, the order of priorities 1, 2, and 3 may be altered.

[0122] Inter Prediction Image Generation Unit 309 If predMode indicates inter-prediction, the Inter Prediction Image Generation Unit 309, as part of the Prediction Image Generation Unit 308, generates the predicted image for a block or sub-block based on the inter-prediction parameters input from the Inter Prediction Parameter Derivation Unit 303 and the reference picture. FIG.5 is a schematic diagram illustrating the configuration of the Inter Prediction Image Generation Unit 309 included in the Prediction Image Generation Unit 308 according to the present embodiment. The Inter Prediction Image Generation Unit 309 comprises a motion compensation unit 3091, a generation unit 3095. The generation unit 3095 includes an Intra-Inter unit 30951, a GPM unit 30952, a BIO unit 30954, and a weight prediction unit 3094. The Inter Prediction Image Generation Unit 309 outputs the generated prediction image for the block to the addition unit 312.

[0123] Motion compensation unit 3091 The motion compensation unit 3091 generates a motion compensation image by reading the reference block from the reference picture memory 306 based on the inter-prediction parameters (predFlagLX, refIdxLX, mvLX) input from the Inter Prediction Parameter Derivation Unit 303. The reference block is a block located at the position shifted by mvLX on the reference picture RefPicLX specified by refIdxLX. If mvLX is not of integer precision, a filter called the motion compensation filter is applied to generate pixels at fractional positions and create the motion compensation image (also can be named interpolation image).

[0124] The motion compensation unit 3091 first derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the prediction block coordinates (x, y) using the following equations: xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+x xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+y yFrac = mvLX[1]&(MVPREC-1) Here, (xPb, yPb) is the upper-left coordinate of the bW*bH-sized block, and x=0…bW-1, y=0…bH-1. MVPREC represents the precision of mvLX (1 / MVPREC pixel precision), for example, MVPREC=16.

[0125] The motion compensation unit 3091 performs horizontal interpolation processing using the interpolation filter on the reference picture refImg to derive a temporary image temp[][]. The Σ denotes the sum for k=0 to NTAP-1, shift1 is the normalization parameter to adjust the value range, and offset1=1<<(shift1-1).

[0126] temp[x][y]=(ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Subsequently, the motion compensation unit 3091 performs vertical interpolation processing on the temporary image temp[][] to derive the interpolation image Pred[][]. The Σ denotes the sum for k=0 to NTAP-1, shift2 is the normalization parameter to adjust the value range, and offset2=1<<(shift2-1).

[0127] Pred[x][y]=(ΣmcFilter[yFrac][k] *temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 In the case of bi-prediction, for each L0 list and L1 list, the motion compensation unit 3091 derives interpolation images PredL0[][] and PredL1[][] from the above Pred[][] and generates an interpolation image Pred[][] from PredL0[][] and PredL1[][].

[0128] GPM unit 30952 The GPM unit 30952 generates the prediction image for the GPM mode by taking the weighted sum of multiple inter-prediction images when ciip_mode is 0.

[0129] IntraInter unit 30951 The IntraInter unit 30951 generates the prediction image for the CIIP mode when ciip_mode is 1. It does so by taking the weighted sum of the inter-prediction image and the intra-prediction image.

[0130] BIO unit 30954 The BIO unit 30954 generates the prediction image by referring to two prediction images (the first prediction image and the second prediction image) and a gradient correction term in the bi-prediction mode.

[0131] Weighted prediction unit 3094 The weighted prediction unit 3094 generates the prediction image for the block by multiplying the interpolation image PredLX by weighting coefficients. If the current block is bi-predicted, a weighted calculation is performed based on the weight indicated by bcwIdx. BCW has a predefined weight list WeightList = [1,3,4,5,7,2,6], where the weights corresponding to the prediction image PredL1 generated by MvL1 are determined by weightPredL1 = WeightList[bcwIdx]. The total weight is set to 8; therefore, the weight for the prediction image PredL0 generated by MvL0 may be calculated as weightPredL0 = 8 - weightPredL1. For a more intuitive understanding, the relationship between bcwIdx and weights is illustrated in FIG.10. The final prediction image Pred is obtained by the weighted sum of PredL0 and PredL1, given by the formula: Pred(x, y) = (weightPred0 * PredL0(x, y) + weightPredL1 * PredL1(x, y) + 4) >> 3. Pred(x, y) represents all pixels in the Pred. The sizes of Pred, PredL0, and PredL1 may be the same.

[0132] The inverse quantization and inverse transform processing unit 311 performs inverse quantization on a quantization transform coefficient input from the prediction parameter derivation unit 320 to calculate a transform coefficient. This quantization transform coefficient is a coefficient obtained by performing a frequency transform such as a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or the like on prediction errors to quantize in coding processing. The inverse quantization and inverse transform processing unit 311 performs an inverse frequency transform such as an inverse DCT, an inverse DST, or the like on the calculated transform coefficient to calculate a prediction error. The inverse quantization and inverse transform processing unit 311 outputs the prediction error to the addition unit 312.

[0133] The addition unit 312 adds the prediction image of the block input from the intra prediction image generation unit 310 and the prediction error input from the inverse quantization and inverse transform processing unit 311 for each pixel and generates a decoded image of the block. The addition unit 312 stores the decoded image of the block in the reference picture memory 306 and outputs the image to the loop filter 305.

[0134] Configuration of video coding apparatus Next, a configuration of the video coding apparatus 11 according to the present embodiment is described. FIG. 7 is a block diagram illustrating a conuration of the video coding apparatus 11 according to the present embodiment. The video coding apparatus 11 is configured to include a prediction image generation unit 101, a subtraction unit 102, a transform and quantization processing unit 103, an inverse quantization and inverse transform processing unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (a prediction parameter storage unit, a frame memory) 108, a reference picture memory (a reference image storage unit, a frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, prediction parameter derivation unit 120, and an entropy coding unit 104.

[0135] The prediction image generation unit 101 generates a prediction image for each CU that is a region obtained by splitting each picture of the image T. The operation of the prediction image generation unit 101 is the same as that of the inter prediction image generation unit 309 already described, and thus descriptions thereof is omitted.

[0136] The subtraction unit 102 subtracts a pixel value of the prediction image of the block input from the prediction image generation unit 101 from a pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform and quantization processing unit 103.

[0137] The transform and quantization processing unit 103 calculates a transform coefficient by performing a frequency transform on the prediction error input from the subtraction unit 102, and derives a quantization transform coefficient by quantization. The transform and quantization proceessing unit 103 outputs the quantization transform coefficient to the entropy coding unit 104 and the inverse quantization and inverse transform processing unit 105.

[0138] The inverse quantization and inverse transform processing unit 105 is the same as the inverse quantization and inverse transform processing unit 311 (FIG. 3) in the video decoding apparatus 31, and descriptions thereof are omitted. The calculated prediction error is output to the addition unit 106.

[0139] To the entropy coding unit 104, the quantization transform coefficient is input from the transform and quantization processing unit 103, and coding parameters are input from the parameter coding unit 111. The entropy coding unit 104 performs entropy coding on split information, the prediction parameters, the quantization transform coefficient, and the like to generate and output the coding stream Te.

[0140] The parameter coding unit 111 instructs the entropy coding unit 104 to encode the prediction parameters and quantization coefficients, derived from the prediction parameter derivation unit 120.

[0141] The parameter encoding unit 111 includes the header encoding unit 1110, CT information encoding unit 1111, and CU encoding unit 1112. The CU encoding unit 1112 further includes the TU encoding unit 1114. The approximate operation of each module is described below.

[0142] The header encoding unit 1110 performs encoding processing for header information, partition information, prediction information, quantization transform coefficients, and other parameters.

[0143] The CT information encoding unit 1111 encodes QT, MT (BT, TT) partition information, and the like.

[0144] The CU encoding unit 1112 encodes CU information, prediction information, partition information, and the like.

[0145] The TU encoding unit 1114 encodes QP update information and quantized prediction errors when prediction errors are present in the TU.

[0146] The CT information encoding unit 1111 and CU encoding unit 1112 supply syntax elements such as inter-prediction parameters (predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra-prediction parameters, and quantization transform coefficients to the parameter encoding unit 111.

[0147] The prediction parameter derivation unit 120 derives the syntax element from the parameters inputted from the coding parameter determination unit 110. Some parts of the prediction parameter derivation unit 120 have the same structure as the prediction parameter derivation unit 320.

[0148] The configuration of the inter prediction parameter encoding unit 112 The inter prediction parameter encoding unit 112 is as shown in FIG.6, comprising the parameter encoding control unit 1121 and the inter prediction parameter derivation unit 303. The inter prediction parameter derivation unit 303 shares a common configuration with the video decoding device. The parameter encoding control unit 1121 includes the merge index derivation unit 11211 and the vector candidate index derivation unit 11212.

[0149] The merge index derivation unit 11211 derives merge candidates and outputs them to the inter prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives prediction vector candidates and outputs them to both the inter prediction parameter derivation unit 303 and the parameter encoding unit 111.

[0150] The configuration of the intra prediction parameter encoding unit includes the parameter encoding control unit and the intra prediction parameter derivation unit. The intra prediction parameter derivation unit shares a common configuration with the video decoding device. In the video encoding device, the inputs to the inter prediction parameter derivation unit 303 may be determined by the encoding parameter determination unit 110, the prediction parameter memory 108, and output to the parameter encoding unit 111.

[0151] The addition unit 106 adds a pixel value of the prediction image of the block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization and inverse transform processing unit 105 to each other for each pixel, and generates a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0152] The loop filter 107 applies a deblocking filter, an SAO, and an ALF to the decoded image generated by the addition unit 106. Note that the loop filter 107 need not necessarily include the above-described three types of filters, and may have a configuration of only the deblocking filter, for example.

[0153] The prediction parameter memory 108 stores the prediction parameters generated by the prediction parameter derivation unit 120 for each target picture and CU at a predetermined position. It may stores the transform coefficients created by the transform and quantization processing unit 103.

[0154] The reference picture memory 109 stores the decoded image generated by the loop filter 107 for each target picture and CU at a predetermined position.

[0155] The coding parameter determination unit 110 selects one set among multiple sets of coding parameters. A coding parameter refers to the above-mentioned QT, BT, or TT split information, the prediction parameter, or a parameter to be coded, the parameter being generated in association therewith. The prediction image generation unit 101 generates the prediction image by using these coding parameters.

[0156] The coding parameter determination unit 110 calculates, for each of the multiple sets, an RD cost value indicating the magnitude of an amount of information and a coding error. The RD cost value is, for example, the sum of a code amount and the value obtained by multiplying a coefficient λ by a square error. The coding parameter determination unit 110 selects a set of coding parameters of which cost value calculated is a minimum value. With this configuration, the entropy coding unit 104 outputs the selected set of coding parameters as the coding stream Te. The coding parameter determination unit 110 outputs the determined coding parameters in the parameter coding unit 111, the prediction parameter derivation unit 120, the prediction image generation unit 101.

[0157] Note that, some of the video coding apparatus 11 and the video decoding apparatus 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the intra prediction image generation unit 310, the inverse quantization and inverse transform processing unit 311, the addition unit 312, the prediction parameter derivation unit 320, the prediction image generation unit 101, the subtraction unit 102, the transform and quantization processing unit 103, the entropy coding unit 104, the inverse quantization and inverse transform processing unit 105, the loop filter 107, the coding parameter determination unit 110, and the parameter coding unit 111, the prediction parameter derivation unit 120, may be realized by a computer. In that case, this configuration may be realized by recording a program for realizing such control functions on a computer-readable recording medium and causing a computer system to read the program recorded on the recording medium for execution. Note that the “computer system” mentioned here refers to a computer system built into either the video coding apparatus 11 or the video decoding apparatus 31 and is assumed to include an OS and hardware components such as a peripheral apparatus. Furthermore, a “computer-readable recording medium” refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, and the like, and a storage device such as a hard disk built into the computer system. Moreover, the “computer-readable recording medium” may include a medium that dynamically stores a program for a short period of time, such as a communication line in a case that the program is transmitted over a network such as the Internet or over a communication line such as a telephone line, and may also include a medium that stores the program for a fixed period of time, such as a volatile memory included in the computer system functioning as a server or a client in such a case. Furthermore, the above-described program may be one for realizing some of the above-described functions, and also may be one capable of realizing the above-described functions in combination with a program already recorded in a computer system.

[0158] Furthermore, a part or all of the video coding apparatus 11 and the video decoding apparatus 31 in the embodiment described above may be realized as an integrated circuit such as a Large Scale Integration (LSI). Each function block of the video coding apparatus 11 and the video decoding apparatus 31 may be individually realized as processors, or part or all may be integrated into processors. The circuit integration technique is not limited to LSI, and the integrated circuits for the functional blocks may be realized as dedicated circuits or a multi-purpose processor. In a case that with advances in semiconductor technology, a circuit integration technology with which an LSI is replaced appears, an integrated circuit based on the technology may be used.

[0159] The embodiment of the present disclosure has been described in detail above referring to the drawings, but the specific configuration is not limited to the above embodiments and various amendments may be made to a design that fall within the scope that does not depart from the gist of the present disclosure.

[0160] The embodiment of the present invention may be applied to a video decoding device that decodes encoded data of image data, and a video encoding device that generates encoded data from image data. In addition, the data structure of the encoded data is generated by the video encoding device and referenced by the video decoding device.

[0161] Reference Signs List 31 Image decoding apparatus 301 Entropy decoding unit 302 Parameter decoding unit 3020 Header decoding unit 3021 CT information decoding unit 3022 CU decording unit 3024 TU decoding unit 3026 Scaling list decoding unit 303 Inter prediction parameter derivation unit 3036 Merge prediction parameter derivation unit 30376 MMVD prediction unit 3032 AMVP prediction parameter derivation unit 3039 BCW parameter derivation unit 30391 Reference sample derivation unit 30392 Template derivation unit 30393 BCW weight derivation device 303931 Candidates derivation unit 303932 Weight derivation unit 303933 Template prediction image generation unit 303934 Template cost derivation unit 303935 Choosing unit 3038 MV addition unit 307 Prediction parameter memory 308 Prediction image generation unit 309 Inter prediction image generation unit 3091 Motion compensation unit 3095 Gneration unit 30951 IntraInter unit 30952 GPM unit 3094 Weight prediction unit 30954 BIO unit 310 Intra prediction image generation unit 320 Prediction parameter derivation unit 311 Inverse quantization and inverse transform processing unit 312 Addition unit 11 Image coding apparatus 101 Prediction image generation unit 102 Subtraction unit 103 Transform and quantization processing unit 104 Entropy coding unit 105 Inverse quantization and inverse transform processing unit 107 Loop filter 108 Prediction parameter memory 110 Coding parameter determination unit 111 Parameter coding unit 1110 Hearder coding unit 1111 CT information coding unit 1112 CU coding unit 1114 TU coding unit 120 Prediction parameter derivation unit 112 Inter prediction parameter derivation unit 1121 Parameter encoding control unit 11211 Merge index derivation unit 11212 Vector candidate index derivation unit <Cross Reference> This patent application claims priority under on US Patent Provisional Application No. 63 / 613,268 filed on December 21, 2023, the entire contents of which are hereby incorporated by reference.

Claims

1. A video decoding apparatus for generating a prediction image, the video decoding apparatus comprising a BCW prediction unit configured to derive plural of inter prediction mode using a weights candidate list with priority generated from a predefined weights list.

2. The video decoding apparatus of claim 1, wherein the video decoding apparatus further comprising a BCW parameter derivation unit configured to derive bcwCandidatesList based on indices of weights, wherein the indices of weights within plus and minus n (or plus or minus 2n) range from the indices of weights are added to the bcwCandidatesList.

3. The video decoding apparatus of claim 1, wherein the video decoding apparatus further comprising a BCW parameter derivation unit configured to derive bcwCandidatesList based on an equal wight, wherein indices of the equal weight are added to the bcwCandidatesList.

4. The video decoding apparatus of claim 1, wherein the video decoding apparatus further comprising a BCW parameter derivation unit configured to derive templates based on current block's size or shape.

Citation Information

Patent Citations

  • Index reordering of BI-prediction with CU-level weight (BCW) by using template-matching

    WO2023076948A1