Image decoding device, image decoding method, and program
Patent Information
- Application Number
- JP2024001393
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-01-09
AI Technical Summary
【0011】 本発明によれば、符号化効率の高い画像復号装置、画像復号方法及びプログラムを提供することができる。
Smart Images

Figure 0007923782000001 
Figure 0007923782000002 
Figure 0007923782000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image decoding device, an image decoding method, and a program. [Background technology]
[0002] Non-Patent Documents 1 and 2 disclose interpretation methods.
[0003] Interpretation generates predicted pixels for the block to be decoded from decoded pixels (reference pixels) in a decoded picture (reference picture) that is different from the picture to be decoded.
[0004] Furthermore, interpretation uses control information to select the motion vector of the block to be decoded, which is necessary for generating the predicted pixels of the block to be decoded, from a list of candidate motion vectors located spatially or temporally adjacent to or near the block to be decoded. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] ITU-T H.266 / VVC [Non-Patent Document 2] M.Coban et al., Algorithm description of Enhanced Compression Model 10 (ECM 10), JVET-AE2025, 2023 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In Non-Patent Documents 1 and 2, the candidate motion vectors that can be selected by interpretation are limited to motion vectors located spatially or temporally adjacent to or near the block to be decoded, which presents a problem in that there is room for improvement in encoding efficiency.
[0007] Therefore, the present invention has been made in view of the above-mentioned problems, and aims to provide an image decoding device, an image decoding method, and a program with high encoding efficiency. [Means for solving the problem]
[0008] The first feature of the present invention is an image decoding device comprising: a decoding unit that decodes code information by variable length and outputs quantization values and control information; an inverse quantization unit that inversely quantizes the quantization values and outputs conversion coefficients; an inverse conversion unit that inversely transforms the conversion coefficients and outputs predicted residual pixels; an intra-prediction unit that generates intra-prediction pixels from the control information and decoded pixels; a decoded picture buffer that stores the decoded pixels; an inter-prediction unit that generates inter-prediction pixels from the control information and the decoded pixels stored in the decoded picture buffer; and the intra-prediction pixels and the I The interface prediction unit comprises an adder that generates the decoded pixels by adding at least one of the interface prediction pixels, and the interface prediction unit generates a chain of motion vectors by recursively searching for motion vectors stored in the decoded picture buffer for motion vectors at positions that are spatially or temporally adjacent or close to the decoded block stored in a merge candidate list that stores motion vector candidates, and adding all the derived motion vectors, thereby generating a chain of motion vectors, and deriving the chain of motion vectors as a candidate (merge candidate) for the motion vector of the decoded block, thereby performing chain of motion vector prediction.
[0009] A second feature of the present invention is an image decoding method comprising: step A, variable-length decoding of coded information to output quantized values and control information; step B, inverse quantization of the quantized values to output conversion coefficients; step C, inverse transformation of the conversion coefficients to output predicted residual pixels; step D, generating intra-predicted pixels from the control information and decoded pixels; step E, storing the decoded pixels in a decoded picture buffer; step F, generating inter-predicted pixels from the control information and the decoded pixels stored in the decoded picture buffer; and the intra-predicted pixels and inter-predicted pixels for the predicted residual pixels. The gist of this method is to perform chain-link motion vector prediction, which includes step G of generating the decoded pixels by adding at least one of the measured pixels, and in step F of the above, a chain-link motion vector is generated by recursively searching for motion vectors stored in the decoded picture buffer for motion vectors at positions that are spatially or temporally adjacent or close to the decoded block stored in a merge candidate list that stores motion vector candidates for the decoded block, and adding all the motion vectors derived from this search, and deriving the chain-link motion vector as a candidate (merge candidate) for the motion vector of the decoded block.
[0010] A third feature of the present invention is a program that causes a computer to function as an image decoding device, wherein the image decoding device comprises: a decoding unit that decodes code information by variable length and outputs quantization values and control information; an inverse quantization unit that inversely quantizes the quantization values and outputs conversion coefficients; an inverse conversion unit that inversely transforms the conversion coefficients and outputs predicted residual pixels; an intra-prediction unit that generates intra-prediction pixels from the control information and decoded pixels; a decoded picture buffer that stores the decoded pixels; an inter-prediction unit that generates inter-prediction pixels from the control information and the decoded pixels stored in the decoded picture buffer; and a prior The system comprises an adder that generates the decoded pixels by adding an intra-predicted pixel and at least one of the inter-predicted pixels. The inter-prediction unit generates a chain of motion vectors by recursively searching for motion vectors stored in the decoded picture buffer for motion vectors at positions that are spatially or temporally adjacent or close to the decoded block, which are stored in a merge candidate list that stores motion vector candidates, and adding all the derived motion vectors. The gist of the system is to perform chain of motion vector prediction, deriving the chain of motion vectors as candidates (merge candidates) for the motion vector of the decoded block. [Effects of the Invention]
[0011] According to the present invention, it is possible to provide an image decoding device, an image decoding method, and a program with high encoding efficiency. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 shows an example of the functional block of an image decoding device 200 according to one embodiment. [Figure 2] Figure 2 is a conceptual diagram of interpretation. [Figure 3] Figure 3 is a diagram illustrating the reference pixel values referenced when the area outside the picture to be decoded, stored in the decoded picture buffer 207 disclosed in Non-Patent Document 1 and Non-Patent Document 2, is referenced from a future block to be decoded. [Figure 4] FIG. 4 is a conceptual diagram illustrating generation of the inter prediction padding region described with reference to FIG. 3. [Figure 5] FIG. 5 is a diagram showing an example of a merge candidate list which is a method for deriving a motion vector in an inter prediction unit 205. [Figure 6] FIG. 6 is a diagram showing an example of names of a plurality of types of merge candidates and derivation orders of these merge candidates in regular merge disclosed in Non-Patent Literature 1 and Non-Patent Literature 2. [Figure 7] FIG. 7 is a diagram showing an example of search positions of spatial merge candidates and TMVP candidates in a block to be decoded. [Figure 8] FIG. 8 is a diagram showing an example of a method for deriving a daisy-chain motion vector in daisy-chain motion vector prediction. [Figure 9] FIG. 9 is a diagram showing an example of a method for deriving a daisy-chain motion vector in daisy-chain motion vector prediction. [Figure 10] FIG. 10 is a diagram showing an example of recursive search for motion vectors including bi-prediction. [Figure 11] FIG. 11 is a diagram showing an example of a daisy-chain motion vector derived by recursive search for motion vectors including bi-prediction. [Figure 12] FIG. 12 is a diagram showing an example of searching for daisy-chain motion vector prediction candidates after searching for a plurality of types of merge candidates for regular merge disclosed in Non-Patent Literature 1 and Non-Patent Literature 2. [Figure 13] FIG. 13 is a diagram showing an example of searching for daisy-chain motion vector prediction candidates after searching for a plurality of types of merge candidates for regular merge disclosed in Non-Patent Literature 1 and Non-Patent Literature 2. [Figure 14] FIG. 14 is a diagram showing an example of merge candidates which are derivation targets of daisy-chain motion vector prediction candidates and positions (index numbers) in a merge candidate list to which the daisy-chain motion vector prediction candidates are added. [Figure 15]Figure 15 shows an example of the merge candidate that is the target of derivation of the chain-linked motion vector prediction candidate, and the position (index number) within the merge candidate list to which the chain-linked motion vector prediction candidate is added. [Figure 16] Figure 16 shows an example of the merge candidate that is the target of derivation of the chain-linked motion vector prediction candidate, and the position (index number) within the merge candidate list to which the chain-linked motion vector prediction candidate is added. [Figure 17] Figure 17 shows an example of the merge candidate that is the target of derivation of the chain-linked motion vector prediction candidate, and the position (index number) within the merge candidate list to which the chain-linked motion vector prediction candidate is added. [Figure 18] Figure 18 shows an example of pruning processing when the interpretation unit 205 of the image decoding device 200 according to one embodiment adds merge candidates corresponding to the chain motion vectors derived by chain motion vector prediction to the merge candidate list. [Figure 19] Figure 19 shows an example of a method for controlling the reference source when deriving the chained motion vector in chained motion vector prediction by the interpretation unit 205 of the image decoding device 200 according to one embodiment. [Figure 20] Figure 20 shows an example of a method for controlling the reference point when deriving the chained motion vector in chained motion vector prediction by the interpretation unit 205 of the image decoding device 200 according to one embodiment. [Figure 21] Figure 21 shows an example of a method for controlling the source and destination references when deriving chained motion vectors in chained motion vector prediction by the interpretation unit 205 of the image decoding device 200 according to one embodiment, in a case where GDR is applied. [Figure 22] Figure 22 shows an example of a method for controlling the reference source when deriving the chained motion vector in chained motion vector prediction. [Figure 23]Figure 23 shows an example of the process by which the decoding unit 201 of an image decoding device 200 according to one embodiment determines whether or not the sps_chained_mvp_enabled_flag has been decoded based on sps_temporal_mvp_enabled_flag. [Figure 24] Figure 24 shows an example of a control point motion vector cpMv derived by the interpretation unit 205 of an image decoding device 200 according to one embodiment. [Modes for carrying out the invention]
[0013] Embodiments of the present invention will be described below with reference to the drawings. Note that the components in the following embodiments can be replaced with existing components as appropriate, and various variations are possible, including combinations with other existing components. Therefore, the description of the following embodiments does not limit the content of the invention as described in the claims.
[0014] <First Embodiment> The image decoding device 200 according to this embodiment will be described below with reference to Figures 1 to 24.
[0015] The image decoding device 200 according to the first embodiment of the present invention is intended for a variety of image signals (hereinafter referred to as "images"). For example, the image decoding device 200 according to this embodiment is intended for YUV (YCbCr) images composed of luminance pixels and chrominance pixels, RGB images composed of RGB pixels, and monochrome images. Here, each pixel constituting an image has discrete values (pixel values) with a predetermined bit width.
[0016] Figure 1 is a diagram showing an example of the functional block of the image decoding device 200 according to this embodiment.
[0017] As shown in Figure 1, the image decoding device 200 includes a code input unit 210, a decoding unit 201, an inverse quantization unit 202, an inverse transform unit 203, an intra prediction unit 204, an inter prediction unit 205, an adder 206, a decoded picture buffer 207, and an image output unit 220.
[0018] Furthermore, in the following descriptions of the functions of each part, the term "pixel" may refer to a block (unit) composed of pixels, a tree block which is the largest block size, or a slice, tile, or image (picture) larger than a tree block.
[0019] Although not shown in Figure 1, the image decoding device 200 may also include an in-loop filter between the adder 206 and the decoded picture buffer 207 to correct the pixel values of the decoded pixels.
[0020] The code input unit 210 is configured to acquire code information encoded by the image encoding device.
[0021] The decoding unit 201 is configured to decode control information and quantization values from the code information input from the code input unit 210. For example, the decoding unit 201 is configured to output control information and quantization values by performing variable-length decoding on such code information.
[0022] Here, the quantized values are sent to the inverse quantization unit 202, and the control information is sent to the decoded picture buffer 207, the intra prediction unit 204, and the inter prediction unit 205. This control information (syntax) includes information necessary for controlling the decoded picture buffer 207, the intra prediction unit 204, and the inter prediction unit 205, and may also include header information such as a sequence parameter set, a picture parameter set, a picture header, and a slice header.
[0023] The inverse quantization unit 202 is configured to inversely quantize the quantized values sent from the decoding unit 201 to obtain the decoded conversion coefficients. These conversion coefficients are then sent to the inverse transformation unit 203.
[0024] The inverse transform unit 203 is configured to inversely transform the transformation coefficients sent from the inverse quantization unit 202 to obtain the decoded predicted residual. This predicted residual is then sent to the adder 206.
[0025] The intra-prediction unit 204 is configured to generate intra-prediction pixels based on the decoded pixels and control information sent from the decoding unit 201. Here, the decoded pixels are obtained via the adder 206 and stored in the decoded picture buffer 207. The intra-prediction pixels are prediction pixels that are added with the prediction residual in the adder 206. The intra-prediction pixels are sent to the adder 206.
[0026] Here, the prediction performed by the intra prediction unit 204 may be the intra prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2. Alternatively, the prediction performed by the intra prediction unit 204 may be the intra block copy or intra template matching prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0027] The decoded picture buffer 207 cumulatively stores decoded pixels sent from the adder 206 and inter prediction information corresponding to decoded pixels sent from the inter prediction unit 205.
[0028] Here, the decoded pixels stored in the decoded picture buffer 207 and the inter-prediction information corresponding to such decoded pixels are referenced by the inter-prediction unit 205. Details of the inter-prediction information will be described later.
[0029] The interprediction unit 205 is configured to generate interprediction pixels based on the decoded pixels and interprediction information obtained by referring to the decoded picture buffer 207, and control information sent from the decoding unit 201.
[0030] Furthermore, the interpretation unit 205 sends the generated interpretation prediction pixels to the addition unit 206, and sends the interpretation prediction information used to generate the interpretation prediction pixels to the decoded picture buffer 207 as corresponding to the pixels to be decoded.
[0031] Here, the prediction performed by the interpretation unit 205 may be the interpretation prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2. Alternatively, the prediction performed by the interpretation unit 205 may be the intrablock copy prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0032] The adder 206 is configured to calculate a decoded pixel by adding the predicted residual sent from the inverse transform unit 203 with at least one of the input intra-predicted pixels and inter-predicted pixels. The decoded pixel is then sent to the image output unit 220, the decoded picture buffer 207, and the intra-prediction unit 204.
[0033] The image output unit 220 is configured to output the decoded pixels transmitted from the adder 206.
[0034] In the following sections, the roles of the interpretation unit 205 and the decoded picture buffer 207, which are characteristic configurations of the image decoding device 200 according to this embodiment, will be explained using the example that the prediction performed by the interpretation unit 205 is an interpretation.
[0035] <Basic roles of the interpretation unit 205 and the decoded picture buffer 207> The basic role of the interpretation unit 205 is to predict the blocks to be decoded with high accuracy in the subsequent adder 206. This involves deriving one or more motion vectors (Mv) for the blocks to be decoded and predicting the pixels (hereinafter, reference pixels) of the decoded blocks (hereinafter, reference blocks) referenced by Mv within a decoded picture (hereinafter, reference picture) that is different from the decoded picture to which the blocks to be decoded belong, that is, performing interpretation.
[0036] Figure 2 shows a conceptual diagram of interpretation. Specifically, Figure 2 shows the MvL0 of two motion vectors derived by the interpretation unit 205 in order to generate interpretation prediction pixels of the decoded block (CurBlk) within the decoded picture (CurPic) in interpretation. k and MvL1 k Based on two different reference pictures (RerPicL0 k and RerPicL1 k ) Reference block (RerBlkL0 k and RerBlkL1 k This shows an example of how ) is referenced.
[0037] Here, L0 and L1 indicate the numbers of the reference picture list, which describes the reference pictures that can be accessed from the picture to be decoded, and k corresponds to the value of the control information (index) used to identify the motion vector, which will be explained later in the method for deriving the motion vector.
[0038] The image decoding device 200 according to this embodiment may derive two different reference picture lists, List 0 (L0) and List 1 (L1), for each picture to be decoded, as described in Non-Patent Documents 1 and 2.
[0039] Here, the decoded picture buffer 207 may derive a reference picture list for each picture to be decoded based on the control information decoded by the decoding unit 201.
[0040] For example, the control information may include information such as the number of reference pictures to be included in the reference picture list and candidate reference pictures (picture reference structure) for the picture to be decoded.
[0041] Details of how motion vectors are derived in the interpretation unit 205 will be described later. The general method for generating interpretation pixels in the interpretation unit 205 is as follows.
[0042] The interpretation unit 205 may generate interpretation pixels from a single reference block, or it may generate interpretation pixels using two reference blocks as shown in Figure 2, or it may generate interpretation pixels using two or more reference blocks.
[0043] Here, the interpretation unit 205 may apply the following method as a way to generate interpretation pixels using two reference blocks.
[0044] As an example, the interpretation unit 205 may generate interpretation predicted pixels by simply averaging the pixel values of two reference blocks.
[0045] As another example, the interpretation unit 205 may generate interpretation pixels by weighted averaging the pixel values of the two reference blocks on a block-by-block basis using predetermined weight values.
[0046] The interpretation unit 205 may apply the block-level weighted biprediction (BCW: Bi-prediction with CU-level Weights) disclosed in Non-Patent Documents 1 and 2 as a technique for such block-level weighted averaging.
[0047] Furthermore, in Non-Patent Documents 1 and 2, multiple weight values can be selected in BCW, including a simple average, i.e., a 1:1 weight value, and such weight values are specified by the value of an internal parameter (bcwIdx, the third internal parameter).
[0048] Here, the interpretation unit 205 derives bcwIdx by inheritance from control information or the decoded picture buffer 207. Details of the inheritance will be described later.
[0049] As another example, the interpretation unit 205 may divide the block to be decoded into two parts by a predetermined straight line, and generate interpretation pixels by weighting the pixel values of the reference block corresponding to each divided region using weight values corresponding to the distance from the dividing line.
[0050] The interpretation unit 205 may apply the Geometric Partitioning Mode (GPM) disclosed in Non-Patent Documents 1 and 2 as a technique for weighted averaging for each region obtained by dividing such block into two by a straight line.
[0051] Here, if the interpretation unit 205 refers to a decimal pixel-precision position within a reference block, it may generate interpretation pixels using an interpolation filter to derive the reference pixel value of such decimal pixel-precision position from the reference pixel values at integer pixel-precision positions in the up, down, left, and right directions of such decimal pixel position.
[0052] Furthermore, when generating interpretation pixels, the interpretation prediction unit 205 may adaptively apply two interpolation filters with different cutoff frequencies for interpretation, as disclosed in Non-Patent Documents 1 and 2.
[0053] Of these two interpolation filters, the one with the lower cutoff frequency is called a Switchable Interpolation Filter (SIF).
[0054] The interpretation unit 205 may apply SIF when the motion vector refers to a half-pixel precision position, and when the internal parameter (hpelIfIdx, first internal parameter) that determines whether or not to apply SIF indicates that SIF should be applied, similar to Non-Patent Documents 1 and 2.
[0055] The interpretation unit 205 derives hpelIfIdx by inheriting from control information or the decoded picture buffer 207. Details of the inheritance will be described later.
[0056] On the other hand, if the interpretation unit 205 references a position with fractional pixel precision, it may round (clip) the reference point of the motion vector to the nearest integer pixel precision position and generate interpretation pixels without applying an interpolation filter.
[0057] Furthermore, the interpretation unit 205 may apply local illumination compensation (LIC), as disclosed in Non-Patent Document 2, when generating the final interpretation predicted pixels.
[0058] LIC is a technique that corrects the reference pixel values within a reference block using linear prediction based on the decoded pixels adjacent to the decoded block and the reference block, respectively.
[0059] The interpretation unit 205 may apply LIC if an internal parameter (licFlag, second internal parameter) that controls whether or not to apply LIC indicates that LIC should be applied, similar to Non-Patent Document 1 and Non-Patent Document 2.
[0060] The interpretation unit 205 derives licFlag by inheriting control information or the decoded picture buffer 207. Details of the inheritance will be described later.
[0061] Furthermore, the interpretation unit 205 may apply the multi-hypothetical prediction (MHP) disclosed in Non-Patent Document 2 when generating the final interpretation predicted pixels.
[0062] MHP is a technique that generates the final inter-predicted pixels from the reference pixel values of at least three reference blocks using a weighted average.
[0063] The interpretation unit 205 may apply MHP if the internal parameter (mhpFlag, fourth internal parameter) that identifies whether or not to apply MHP indicates that MHP should be applied, similar to Non-Patent Document 2.
[0064] The interpretation unit 205 derives the mhpFlag by inheriting control information or the decoded picture buffer 207. Details of the inheritance will be described later.
[0065] The interpretation unit 205 sends interpretation pixels for the decoded block to the adder 206, and also sends to the decoded picture buffer 207 the motion vector used to generate the interpretation pixels, an index (reference index) for identifying the reference picture referenced by the motion vector in the reference picture list, and a counter (POC: Picture Order Count) representing the picture output order of the reference picture corresponding to the reference index. Hereinafter, the motion vector used to generate these interpretation pixels, the reference index, and the POC corresponding to the reference index will be referred to as interpretation information. Furthermore, the interpretation unit 205 may also send at least one or all of the internal parameters bcwIdx, hpelIfIdx, licFlag, or mhpFlag, which modify the method of generating the interpretation pixels described above, along with the interpretation information.
[0066] The decoded picture buffer 207 stores not only the decoded pixels of the decoded block sent from the adder 206, but also inter prediction information for the decoded block corresponding to those decoded pixels.
[0067] The decoded picture buffer 207 may store such inter prediction information in a predetermined size (number of pixels). For example, the decoded picture buffer 207 may store such inter prediction information in units of the minimum size (4 × 4 pixels) of the block to be decoded, as disclosed in Non-Patent Documents 1 and 2.
[0068] The decryption picture buffer 207 stores such interpretation information in the minimum size of the block to be decrypted, thereby improving the interpretation accuracy when such interpretation information is referenced from the block to be decrypted in the future.
[0069] The decryption picture buffer 207 accumulates decrypted pixels on a per-picture-to-decryption basis. Alternatively, the decryption picture buffer 207 may accumulate decrypted pixels on a per-slice-to-decryption basis. Hereafter, the case in which decrypted pixels are accumulated on a per-picture-to-decryption basis will be explained as an example.
[0070] The decoded picture buffer 207 may sequentially delete decoded pixels and corresponding inter prediction information for each decoded picture unit if it is determined that the decoded pixels and corresponding inter prediction information for each decoded picture unit will not be referenced as reference pictures and reference blocks by future decoded pictures or decoded blocks. This configuration reduces the amount of information stored in the decoded picture buffer 207.
[0071] The decoded picture buffer 207 may derive the number of decoded pictures to be stored and the reference picture structure between the decoded pictures based on the control information decoded by the decoding unit 201.
[0072] As an example of modification, the decoded picture buffer 207 may additionally store inter-prediction pixels if such pixels exist in a predetermined area outside the picture to be decoded, in addition to the picture to be decoded.
[0073] Figure 3 is a diagram illustrating the reference pixel values referenced when the area outside the picture to be decoded, stored in the decoded picture buffer 207 disclosed in Non-Patent Document 1 and Non-Patent Document 2, is referenced from a future block to be decoded.
[0074] As shown in Figure 3, Non-Patent Document 1 (VVC: Versatile Video Coding) generates a reference pixel value by duplicating the pixel value at the nearest boundary of the decoded picture when a future decoded block is referenced within a predetermined range outside the decoded picture.
[0075] In VVC, the specified range (padding area) is defined as the maximum width (number of pixels) of the block to be decoded plus 16 pixels.
[0076] In Non-Patent Document 2 (ECM: Enhanced Compression Model), as shown in Figure 3, only a predetermined range is stored between the area outside the picture to be decoded and the padding area as an inter-predicted pixel value of the decoded block that can be referenced from future decoded blocks (inter-predicted padding area).
[0077] This improves the accuracy of future inter-prediction of target blocks because even if areas outside the decryption target picture are referenced from future decryption target blocks, inter-predicted pixel values for the decryption target block are available.
[0078] Figure 4 is a conceptual diagram showing how the interpredictive padding region described in Figure 3 is generated.
[0079] As shown in Figure 4, if the adjacent block of the reference block corresponding to the block to be decoded has inter-prediction pixels, inter-prediction pixels can be generated outside the decoded picture even if the block to be decoded is adjacent to the decoded picture. Therefore, the decoded picture buffer 207 can accumulate the inter-prediction padding region shown in Figure 3.
[0080] <Method for deriving motion vectors in the interpretation unit> The method for deriving motion vectors in the interpretation unit 205 of the image decoding device 200 according to this embodiment will be described below with reference to Figures 5 to 17.
[0081] Figure 5 shows an example of a merge candidate list, which is a method for deriving motion vectors in the interpretation unit 205.
[0082] The merge candidate list is a list containing candidate inter-prediction information (i.e., merge candidates) constructed by the inter-prediction unit 205 to derive motion vectors using the merge (merge mode) disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0083] The interpretation unit 205 selects a merge candidate from among multiple merge candidates in the merge candidate list based on an index (merge_idx) sent from the decoding unit 201 to identify the merge candidate in the merge candidate list, and uses it as interpretation information for generating interpretation pixels.
[0084] The upper limit of the number of merge candidates stored in the merge candidate list may be a fixed value or a variable value. If it is a variable value, control information can be provided in the sequence parameter set, etc., as in Non-Patent Documents 1 and 2, and the variable value can be set based on such control information.
[0085] The interpretation unit 205 identifies the interpretation information to be used to generate interpretation pixels from the merge candidate list based on one or more merge_idx sent from the decoding unit 201.
[0086] For example, the interpretation unit 205 may derive interpretation information by decoding merge_idx once for each region obtained by dividing the GPM target block by a straight line.
[0087] As shown in Figure 5, the inter-prediction unit 205 stores inter-prediction information (merge candidates) corresponding to L0 and L1 of the merge candidate list.
[0088] In the example shown in Figure 5, the interpretation unit 205 only stores motion vectors and reference indices as interpretation information. However, in addition to these, it may also store at least one or all of the following internal parameters that change the method of generating the interpretation pixels described above: bcwIdx, hpelIfIdx, licFlag, or mhpFlag.
[0089] The interpretation unit 205 may search for and add interpretation information, i.e., interpretation information located spatially or temporally adjacent or close to the block to be decoded, to the merge candidate list as a merge candidate.
[0090] As an example, the interpretation unit 205 may combine and apply the search for multiple types of merge candidates disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0091] Figure 6 shows the names of the various merge candidates in a regular merge disclosed in Non-Patent Documents 1 and 2, and the derivation order of these merge candidates.
[0092] Specifically, in Non-Patent Document 1 (VVC), the interpretation unit 205 checks whether or not to use the following merge candidates, in the order of Spatial merge candidates, Temporal Motion Vector Prediction (TMVP) candidates, History-based Motion Vector Prediction (HMVP) candidates, and Pairwise merge candidates.
[0093] In Non-Patent Document 2 (ECM), the interpretation unit 205 checks whether non-adjacent (NA) spatial merge candidates are used between the spatial merge candidates and the TMVP candidates, in addition to the merge candidates described above in Non-Patent Document 1.
[0094] The following is an overview of each merge candidate:
[0095] Figure 7 shows the search locations for spatial merge candidates and TMVP candidates in the block to be decoded.
[0096] As shown in Figure 7, in the spatial merge candidate, interpretation information is searched at the following positions adjacent to the decoded block (CurBlk): bottom left (A0), left (A1), top right (B0), top (B1), and top left (B2).
[0097] In other words, for spatial merge candidates, interpretation information for spatially adjacent locations to the decryption target block (CurBlk) is used.
[0098] Furthermore, in the TMVP candidate, interpretation information is searched for the central lower right (Col) of the decryption target block (CurBlk) and the lower right outside the decryption target block (CurBlk) (H) in a picture different from the decryption target block (reference picture) shown in Figure 7.
[0099] In other words, TMVP candidates use interpretation information from locations temporally adjacent to the block to be decoded.
[0100] In HMVP candidates and non-adjacent space merge candidates, interpretation information from spatially distant (or adjacent) locations to the decryption target block is used.
[0101] In particular, for HMVP candidates, a table capable of storing inter-prediction information in a FIFO format is constructed, and each time new inter-prediction information is searched for, the oldest inter-prediction information is discarded from the table, and the new inter-prediction information is added to the table.
[0102] The interpretation unit 205 of the image decoding device 200 according to this embodiment can design the position for searching for interpretation information and the size of the table for HMVP candidates, similar to the HMVP candidates and non-adjacent space merge candidates disclosed in Non-Patent Documents 1 and 2.
[0103] In addition, the inter prediction unit 205 may, in the search for merge candidates in the normal merge described above, modify the inter prediction information (motion vector) by re-searching for a new reference block position from the surrounding pixels of the reference block of the inter prediction information to be searched, so as to reduce the difference (hereinafter referred to as the template matching cost) between the decoded pixels adjacent to the reference block (template of the block to be decoded) and the decoded pixels adjacent to the block to be decoded (template of the block to be decoded).
[0104] As an example, the interpretation unit 205 may apply the template matching merge disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0105] In addition, the inter-prediction unit 205 may, in selecting a merge candidate from the merge candidate list in the normal merge described above, derive the inter-prediction information of L0 from the control information, and select as the merge candidate of L1 the merge candidate that minimizes the difference between the reference block template based on the inter-prediction information of L0 derived from the control information and the reference block template based on the inter-prediction information stored in L1 of the merge candidate list.
[0106] The inter-prediction unit 205 can derive merge candidates for L1 (i.e., inter-prediction information) based on the inter-prediction information for L0 derived from such control information. Therefore, in addition to the control information necessary for deriving the inter-prediction information for L0, it does not require decoding of merge_idx to derive the inter-prediction information for L1.
[0107] As an example, the interpretation unit 205 may apply the Adaptive motion vector prediction-merge mode disclosed in Non-Patent Documents 1 and 2.
[0108] As an example of modification, the inter-prediction unit 205 may reverse L0 and L1 in the example described above. That is, the inter-prediction unit 205 may derive the inter-prediction information for L1 from the control information and the inter-prediction information for L0 from the merge candidate list.
[0109] In addition, the inter-prediction unit 205 may derive inter-prediction information by constructing a merge candidate list for merging, which derives inter-prediction information in sub-block units obtained by dividing the block to be decoded.
[0110] As an example, the inter-prediction unit 205 may derive inter-prediction information by constructing a list of merge candidates for subblock TMVP (SbTMVP) or affine merge as disclosed in Non-Patent Documents 1 and 2.
[0111] <1. Basic Concepts of Chain-Linked Motion Vector Prediction (Derivation of Chain-Linked Motion Vectors)> The methods for deriving interpretation information (hereinafter referred to as motion vectors for the sake of simplicity) by the interpretation unit described above are all limited to information stored at locations that are spatially or temporally adjacent or close to the block to be decoded.
[0112] To further improve coding efficiency, the inter prediction unit 205 may derive, as the motion vector of the block to be decoded, a motion vector generated by adding a motion vector derived by recursively searching for motion vectors accumulated in the decoded picture buffer 207 with respect to motion vectors at positions spatially or temporally adjacent or close to the block to be decoded.
[0113] Herein, a technique of generating a new motion vector (chained motion vector) by recursively (in a chained manner) searching for motion vectors accumulated in the decoded picture buffer 207 starting from motion vectors at positions spatially or temporally adjacent or close to the block to be decoded is referred to as "Chained Motion Vector Prediction (CMVP)" in the present specification. Hereinafter, the technical content of such chained motion vector prediction will be described in detail.
[0114] A method of deriving a chained motion vector in chained motion vector prediction will be described with reference to FIGS. 8 to 17.
[0115] FIG. 8 is a diagram showing an example of a method of deriving a chained motion vector in chained motion vector prediction.
[0116] As shown in FIG. 8, the inter prediction unit 205 obtains the L0 motion vector MvL0 corresponding to any merge candidate stored in the merge candidate list k (0) and a reference index as a starting point, and recursively (in a chained manner) searches for motion vectors accumulated in the decoded picture buffer 207 to generate a new motion vector.
[0117] For example, as shown in (Expression 1) and (Expression 2), the inter prediction unit 205 obtains the reference picture RefPicL0 k (m) A new motion vector generated by chained motion vector prediction for the reference block RefBlk and a reference picture can be derived.
[0118] MvL0 k / m =MvL0k (0)+MvL0 k (1) + MvL0 k (2) + ... + MvL0 k (m) (Formula 1) RefPicL0 k / m =RefPicL0 k (m) (Formula 2) In other words, the new motion vector MvL0 generated by chain-reaction motion vector prediction k / m This can be calculated by recursively (chaining together) searching for motion vectors stored in the decoded picture buffer 207 and adding up all the resulting motion vectors.
[0119] Furthermore, the new reference picture generated by chain-reaction motion vector prediction is MvL0 k / m This is the reference picture that is being referenced.
[0120] Figure 9 shows an example of a method for deriving chained motion vectors in chained motion vector prediction.
[0121] As shown in Figure 9, in chain-reaction motion vector prediction, the interpretation unit 205 may add block vectors (vectors indicating references within the same picture) to the recursive search targets in addition to the search for motion vectors.
[0122] For example, if a block vector (Bv) is added, the reference picture RefPicL0 k For the reference block RefBlk of (m), Bv is added to the derivation formula for the new motion vector generated by chain-reaction motion vector prediction, as shown in (Equation 3).
[0123] MvL0 k / m =MvL0 k (0)+Bv k (0)+MvL0 k (1) + MvL0 k (2) + ... + MvL0 k (m) (Formula 3) Figure 10 shows an example of a recursive search for motion vectors, including biprediction (i.e., interprediction using two different motion vectors for L0 and L1). Figure 11 shows an example of a chain of motion vectors derived by a recursive search for motion vectors, including biprediction.
[0124] As shown in Figures 10 and 11, in chain-linked motion vector prediction, the interpretation unit 205 uses each search list number (L0 or L1) and each search depth (or each search count) of the motion vectors stored in the decoded picture buffer 207 to newly search for and derive one or more motion vectors (up to two in the example of Figures 10 and 11) from the single reference motion vector, and adds all the already derived motion vectors to each of the one or more newly derived motion vectors to obtain the chain-linked motion vector MvL0 k / m Generates.
[0125] Figures 12 and 13 show an example of searching for chain-linked motion vector prediction candidates after searching for multiple types of merge candidates for normal merging as disclosed in Non-Patent Documents 1 and 2.
[0126] As shown in Figures 12 and 13, the interpretation unit 205 may search for chain-linked motion vector prediction candidates after searching for predetermined merge candidates.
[0127] For example, as shown in Figure 12, the interpretation unit 205 may, in VVC, search for chain-linked motion vector prediction candidates after searching for TMVP candidates.
[0128] Because chain-reaction motion vector prediction searches for new motion vectors using the motion vectors of reference pictures rather than the decoded pictures stored in the decoded picture buffer 207, it will not be effective unless TMVP is enabled for the decoded sequence, the group of decoded pictures, the decoded picture, or the decoded slice.
[0129] Therefore, searching for TMVP candidates followed by a chain of motion vector prediction candidates is a natural design method for deriving motion vectors.
[0130] As an example of a modification shown in Figure 12, the interpretation unit 205 may, as shown in Figure 13, search for chain-reaction motion vector prediction candidates in VVC and ECM after searching for HMVP candidates.
[0131] Chain motion vector prediction, due to its nature of generating new motion vectors through recursive motion vector retrieval, must be applied only when at least one merge candidate is stored in the merge candidate list.
[0132] As shown in Figure 13, by searching for HMVP candidates that can be expected to store one or more motion vector candidates in the merge candidate list, and then searching for chained motion vector prediction candidates, the likelihood of generating new motion vectors through chained motion vectors increases.
[0133] As another example of modification, the interpretation unit 205 may, in the ECM, search for chained motion vector prediction candidates after searching for spatial merge candidates, as shown in Figure 12.
[0134] Figure 7 shows the search locations for spatial merge candidates. For example, increasing the number of search locations for spatial merge candidates increases the likelihood that one or more motion vector candidates will be stored in the merge candidate list, making it easier to generate new motion vector candidates through chained merge candidates.
[0135] Figures 14 to 17 show examples of merge candidates that are the target of derivation for chain-linked motion vector prediction candidates, and the position (index number) within the merge candidate list to which the chain-linked motion vector prediction candidates are added.
[0136] Note that in Figures 14 to 17, only motion vectors are shown as interpretation information stored in the merge candidate list in order to simplify the explanation of the technical details, but the reference index mentioned above is also stored.
[0137] Furthermore, in addition to the motion vector and reference index, the interpretation unit 205 may also store internal parameters that modify the interpretation pixel generation method described above, namely bcwIdx, hpelIfIdx, licFlag, and mhpFlag.
[0138] As shown in Figure 14, the inter-prediction unit 205 may, after storing predetermined merge candidates in a merge candidate list, apply chain-reaction motion vector prediction sequentially starting from the first merge candidate stored in the merge candidate list, and sequentially store all new merge candidates generated by the chain-reaction motion vector prediction for each merge candidate in the merge candidate list.
[0139] Here, the predetermined merge candidates may be any of the multiple types of merge candidates in a normal merge as explained using Figures 12 and 13, or they may be a predetermined number of merge candidates (a predetermined number in merge_idx).
[0140] Such predetermined number may be set as a fixed value, or it may be set as a variable value for each sequence, group of pictures, picture, or slice to be decoded, using control information for each sequence, group of pictures, or slice to be decoded.
[0141] As an example of a modification shown in Figure 14, the interpretation unit 205 may, as shown in Figure 15, store predetermined merge candidates in a merge candidate list, then apply chain-link motion vector prediction sequentially starting from the first merge candidate stored in the merge candidate list, and sequentially store the new chain-link motion vectors (i.e., merge candidates) generated at each search depth of the chain-link motion vector prediction for each merge candidate into the merge candidate list.
[0142] As an example of modifications to Figures 14 and 15, the interpretation unit 205 may, as shown in Figure 16, after the first merge candidate is stored in the merge candidate list, apply chain-link motion vector predictions sequentially starting from that merge candidate, and sequentially store all the new chain-link motion vectors (i.e., merge candidates) generated by the chain-link motion vector prediction for each merge candidate into the merge candidate list.
[0143] As an example of the modifications shown in Figures 14 to 16, the interpretation unit 205 may, as shown in Figure 17, after the first merge candidate is stored in the merge candidate list, apply chain-link motion vector prediction sequentially starting from that merge candidate, and sequentially store the new chain-link motion vectors (i.e., merge candidates) generated at each search depth of the chain-link motion vector prediction for each merge candidate into the merge candidate list.
[0144] As an example of the modifications shown in Figures 14 to 17, the interpretation unit 205 may generate new merge candidates by combining the chained motion vectors generated using each search depth of the chained motion vector prediction for each merge candidate list with the chained motion vectors generated at a different search depth.
[0145] As an example of the modifications shown in Figures 14 to 17, the interpretation unit 205 may generate new merge candidates by combining the chained motion vectors generated using each search list number or search depth for the chained motion vector prediction for each merge candidate list with the chained motion vectors generated using different search list numbers or search depths.
[0146] <2. Pruning process for chain-like motion vector prediction candidates> Using Figure 18, the pruning process when the interpretation unit 205 of the image decoding device 200 according to this embodiment adds merge candidates (the aforementioned daisy-chain motion vector prediction candidates) corresponding to the daisy-chain motion vectors derived by daisy-chain motion vector prediction to the merge candidate list will be explained.
[0147] Figure 18 shows an example of pruning processing when the interpretation unit 205 of the image decoding device 200 according to this embodiment adds merge candidates corresponding to the chain motion vectors derived by chain motion vector prediction to the merge candidate list.
[0148] The interpretation unit 205 determines whether to add merge candidates corresponding to the chain motion vectors derived by chain motion vector prediction based on predetermined conditions.
[0149] Specifically, as shown in Figure 18, if the interpretation unit 205 determines in step S000 that a predetermined condition is met, it determines in step S002 not to add a merge candidate corresponding to the chained motion vector derived by chained motion vector prediction.
[0150] On the other hand, if the interpretation unit 205 determines in step S000 that the predetermined conditions are not met, it determines in step S001 to add a merge candidate corresponding to the chain motion vector derived by chain motion vector prediction.
[0151] Here, for example, the predetermined condition may be that the reference picture and motion vector associated with each of the merge candidates corresponding to at least one merge candidate stored in the merge candidate list and the derived chain motion vector match.
[0152] When the interpretation unit 205 determines that at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain of motion vectors match, it decides not to add the merge candidate corresponding to the chain of motion vectors to the merge candidate list, that is, to prune (make the merge candidate corresponding to the chain of motion vectors subject to pruning). This prevents cases where two or more identical motion vectors are stored in the merge candidate list, reduces the amount of code in merge_idx (merge index) sent to the image decoding unit 200, and improves encoding efficiency.
[0153] Alternatively, the predetermined condition may be that at least one merge candidate stored in the merge candidate list and one or more reference pictures and motion vectors associated with each of the merge candidates corresponding to the derived chain motion vector all match.
[0154] Alternatively, the predetermined condition may be that the reference pictures and motion vectors of merge candidate list number 0 (L0) and merge candidate list number 1 (L1), associated with at least one merge candidate stored in the merge candidate list and each of the merge candidates corresponding to the derived chain of motion vectors, all match.
[0155] Furthermore, the predetermined conditions may include that the reference pictures associated with at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain of motion vectors match, and that the error (Euclidean distance) of the reference position of the motion vectors associated with at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain of motion vectors is less than a predetermined threshold.
[0156] This enhances the pruning process for merge candidates corresponding to chained motion vectors, more effectively preventing cases where two or more similar motion vectors are stored in the merge candidate list, reducing the amount of code in merge_idx sent to the image decoder 200, and improving encoding efficiency.
[0157] As a further control, the interpretation unit 205 may set a predetermined threshold for the case where there is one motion vector to compare to be greater than the predetermined threshold for the case where there are two motion vectors to compare.
[0158] In this way, by setting a predetermined threshold for the case where there is one motion vector to be compared to a predetermined threshold for the case where there are two motion vectors to be compared, it becomes possible to perform pruning with higher accuracy.
[0159] Here, the interpretation unit 205 may set a predetermined threshold of 1.5 pixels when there is only one motion vector to compare.
[0160] Alternatively, the interpretation unit 205 may set a predetermined threshold of 1.0 pixels or 2.0 pixels when there is only one motion vector to compare.
[0161] Furthermore, the interpretation unit 205 may set a predetermined threshold of 1.0 pixel when there are two motion vectors to be compared.
[0162] Alternatively, the interpretation unit 205 may set a predetermined threshold of 0.5 pixels or 1.5 pixels when there are two motion vectors to compare.
[0163] Furthermore, the predetermined conditions described above may also be that the reference picture and motion vector associated with each of the merge candidates, which are linked to at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain of motion vectors, and the internal parameters that change the method for generating predetermined interprediction pixels, all match.
[0164] Alternatively, the predetermined condition may be that the reference picture, motion vector, and internal parameter (hpelIfIdx) that identifies whether or not a switching interpolation filter is applied, associated with at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector, all match.
[0165] Alternatively, the predetermined condition may be that the reference picture, motion vector, and internal parameter (licFlag) that controls whether or not local luminance compensation is applied, associated with at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector, all match.
[0166] Furthermore, the predetermined condition may be that the reference picture and motion vector associated with at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector, as well as the internal parameter (bcwIdx) that identifies the weight values of the block-level weighted biprediction, match.
[0167] The interpretation unit 205, by including the above-mentioned internal parameters in predetermined conditions, generates interpretation pixels in a different way when the internal parameters are different, even when the reference picture and motion vector to be compared are the same. This leaves room for improvement in interpretation accuracy, and thus an improvement in coding efficiency can be expected.
[0168] <3. Controlling the source and destination of references when deriving chain-reaction motion vectors> Using Figures 19 to 22, the control method of the reference source and reference destination when the interpretation unit 205 of the image decoding device 200 according to this embodiment derives a chain of motion vectors in chain of motion vector prediction will be explained.
[0169] Firstly, using Figure 19, we will explain the method for controlling the reference source when deriving the chained motion vector in chained motion vector prediction.
[0170] Figure 19 shows an example of a method for controlling the reference source when deriving the chained motion vectors in chained motion vector prediction by the interpretation unit 205.
[0171] In the chain-linked motion vector prediction, the interpretation unit 205 may set the starting point of the first motion vector constituting the chain-linked motion vector to part or all of the block to be decoded when deriving the chain-linked motion vector.
[0172] As an example of setting a part of the block to be decoded, as shown in Figure 19, the interpretation unit 205 may, when deriving the chain of motion vectors in chain of motion vector prediction, set the starting point of the first motion vector constituting the chain of motion vectors to the center (C) of the block to be decoded (CurBlk).
[0173] Alternatively, as shown in Figure 19, the interpretation unit 205, in predicting a chain of motion vectors, may, when deriving a chain of motion vectors, set the starting point of the first motion vector constituting the chain of motion vectors to a predetermined position in the decoded block (CurBlk) in addition to the center (C) of the decoded block (CurBlk), and search whether motion vectors are accumulated at the reference destination of the motion vectors in a predetermined order.
[0174] Specifically, as shown in Figure 19, in chain-link motion vector prediction, when deriving chain-link motion vectors, the interpretation unit 205 may set the starting point of the first motion vector constituting the chain-link motion vector to be the four corners of the decoded block (CurBlk) in addition to the center (C) of the decoded block (CurBlk), and search whether motion vectors are accumulated at the reference locations of the motion vectors in a predetermined order: center (C), upper left (TL), upper right (TR), lower left (BL), and lower right (BR) of the decoded block.
[0175] The interpretation unit 205 may generate a chain of motion vectors for each motion vector that is searched and derived at each point.
[0176] As an example of modification, the interpretation unit 205 may compare the difference between the adjacent decoded pixel values of the reference block and the decoded target block of each motion vector searched and derived at each point, and generate a chain of motion vectors limited to at least one motion vector with a small difference.
[0177] As another example of modification, the interpretation unit 205 may derive a new motion vector for generating a chain of motion vectors by weighting and averaging each motion vector derived from the search at each point.
[0178] Specifically, the interpretation unit 205 may, when weighting and averaging each motion vector, scale each motion vector based on the distance of each motion vector's reference picture from the block to be decoded to generate a new motion vector. The scaling target may be set to the reference picture with the minimum distance from the block to be decoded.
[0179] When the interpretation unit 205 derives a chain of motion vectors, it sets the starting point of the first motion vector constituting the chain of motion vectors to the four corners of the block to be decoded, in addition to the center of the block to be decoded. This increases the likelihood of deriving a chain of motion vectors, and as a result, an improvement in coding efficiency can be expected.
[0180] As a further control, the interpretation unit 205 may, when deriving the chain of motion vectors in chain of motion vector prediction, control the starting point of the first motion vector constituting the chain of motion vectors according to the block size or block aspect ratio of the block to be decoded.
[0181] For example, in chain-link motion vector prediction, the interpretation unit 205 may, when deriving the chain-link motion vector, set the starting point of the first motion vector constituting the chain-link motion vector to only the center of the block to be decoded when the block size of the block to be decoded is less than or equal to the first size.
[0182] If the block size of the block to be decoded is less than or equal to the first size, when deriving the chain of motion vectors, the first motion vector constituting the chain of motion vectors is likely to be the same at the center and the four corners of the block to be decoded.
[0183] Therefore, when the block size of the block to be decoded is less than or equal to the first size, the amount of processing required to search for the chain of motion vectors can be reduced by setting the starting point of the first motion vector constituting the chain of motion vectors to only the center of the block to be decoded (excluding the four corners of the block to be decoded).
[0184] Alternatively, in chain-linked motion vector prediction, when deriving chain-linked motion vectors, the interpretation unit 205 may set the starting point of the first motion vector constituting the chain-linked motion vector to the four corners of the decoded block in addition to the center of the decoded block, if the block size of the decoded block is larger than the first size, and search whether motion vectors are accumulated at the reference points of the motion vectors in a predetermined order: center of the decoded block, upper left, upper right, lower left, and lower right.
[0185] If the block size of the block to be decoded is larger than the first size, when deriving the chain of motion vectors, the first motion vector that makes up the chain of motion vectors is likely to differ at the center and the four corners of the block to be decoded.
[0186] Therefore, when the block size of the block to be decoded is larger than the first size, when deriving the chain of motion vectors, by setting the starting point of the first motion vector constituting the chain of motion vectors to the four corners of the block to be decoded in addition to the center of the block to be decoded, different chain of motion vectors can be generated, and an improvement in coding efficiency can be expected.
[0187] Here, since the blocks to be decoded may be non-square, the block size (number of pixels in the block) of the blocks to be decoded may be replaced with the width and height of the blocks to be decoded, and the above control may be implemented.
[0188] Specifically, in chain-linked motion vector prediction, when deriving chain-linked motion vectors, the interpretation unit 205 may set the starting point of the first motion vector constituting the chain-linked motion vector to be at the four corners of the decoded block in addition to the center of the decoded block, if the width or height of the decoded block is greater than a predetermined number of pixels, and search whether motion vectors are accumulated at the reference destination of the motion vector in the predetermined order of center, upper left, upper right, lower left, and lower right of the decoded block.
[0189] The interpretation unit 205 may set the predetermined number of pixels to 8 pixels, 16 pixels, or 32 pixels.
[0190] Alternatively, in chain-linked motion vector prediction, when deriving chain-linked motion vectors, the interpretation unit 205 may set the starting point of the first motion vector constituting the chain-linked motion vector to be the center of the decoded block, as well as the lower left and upper right corners of the decoded block, if the aspect ratio of the blocks to be decoded is greater than a predetermined ratio, and search whether motion vectors are accumulated at the reference points of the motion vectors in a predetermined order: center, lower left, and upper right corners of the decoded block.
[0191] The interpretation unit 205 may set the predetermined ratio to 4, 8, or 16.
[0192] Secondly, using Figures 20 and 21, we will explain an example of a method for controlling the reference when deriving the chained motion vector in chained motion vector prediction.
[0193] Figure 20 shows an example of a method for controlling the reference point when deriving the chained motion vectors in chained motion vector prediction by the interpretation unit 205.
[0194] As shown in Figure 20, when the interpretation unit 205 recursively searches for motion vectors to derive a chain of motion vectors, if the reference position of each motion vector refers to an outside reference picture, it may search for the motion vector at the nearest reference picture's accumulation location of motion vectors from that reference position.
[0195] Even if the reference position of a motion vector refers to an area outside the reference picture, a chain of motion vectors linked from that motion vector may be located within a different reference picture, allowing for the generation of new chained motion vectors.
[0196] Alternatively, as shown in Figure 20, the interpretation unit 205 may terminate the generation of the chain of motion vectors if, during the recursive motion vector search process for deriving the chain of motion vectors, the reference position of each motion vector refers to an area outside the reference picture.
[0197] If the reference position of a motion vector refers to an area outside the reference picture, there is a high probability that the motion vectors linked from that motion vector will also refer to areas outside the reference picture in different reference pictures. Therefore, the amount of decoding processing can be reduced by stopping the generation of linked motion vectors.
[0198] Alternatively, if the decoded picture buffer 207 is accumulating an interprediction padding area outside the reference picture, the interprediction unit 205 may perform the same processing by replacing the above description of "outside the reference picture" with the description of "outside the interprediction padding area".
[0199] Figure 21 shows an example of a method for controlling the source and destination references when deriving the chained motion vectors in chained motion vector prediction by the interpretation unit 205, in a case where sequential decoder update (GDR), disclosed in Non-Patent Document 1, known as an encoding and decoding technology for low-latency video transmission, is applied to a group of pictures to be decoded.
[0200] GDR is a technology that reduces the processing delay of video encoding and decoding at a constant bitrate compared to video encoding and decoding in a random access configuration that includes intra pictures, which is widely used in video transmission. This is achieved by updating only a portion of the screen using encoding and decoding with intra prediction (hereinafter referred to as intra update), and moving this portion of the screen for each picture to be decoded.
[0201] Figure 21 illustrates the intra-update region as an intra-slice, showing an example where such a slice moves from the top to the bottom edge of the picture to be decoded.
[0202] Herein, in this specification, the three divided regions in the decryption target picture by GDR are referred to as the region after intra-update, the region during intra-update, and the region before intra-update.
[0203] As shown in Figure 21, when sequential decoder updates are applied to the group of pictures to be decoded, the interpretation unit 205 may restrict the recursive search for motion vectors in chained motion vector prediction so as not to exceed the region being updated intra-.
[0204] For example, as shown in Figure 21, if decoder updates are applied sequentially to the group of pictures to be decoded, and the blocks to be decoded are included in the region after intra-update, the interpretation unit 205 may restrict the region that can be referenced in each search of the recursive motion vector in chained motion vector prediction to only the region after intra-update in the reference picture for each search target.
[0205] Alternatively, if the interpretation unit 205 sequentially applies decoder updates to the group of pictures to be decoded, and the blocks to be decoded are included in the region before the intra-update, the region that can be referenced in each search of the recursive motion vector in the chain-linked motion vector prediction may be limited to only the region before the intra-update in the reference picture for each search target.
[0206] Thus, when the interpretation unit 205 sequentially applies decoder updates to the group of pictures to be decoded, it restricts the recursive search for the sequence of motion vectors in the sequence of motion vector prediction so as not to exceed the region being updated intra-, thereby enabling the generation of sequence of motion vectors with high prediction accuracy.
[0207] Figure 22 shows an example of how the interconnect prediction unit 205 controls the reference destination when the chained motion vectors are derived in chained motion vector prediction, in a case where reference of motion vectors to a reference picture with a different resolution (sampling ratio) than the decoded picture is permitted in the decoded block, as disclosed in Non-Patent Document 1.
[0208] In RPR, the interpretation unit 205 calculates the ratio of the width and height of the reference picture to the width and height of the decoded picture, and scales the pixel position and motion vector length of the reference picture relative to the decoded picture according to that ratio to generate interpretation predicted pixels.
[0209] Therefore, for example, as shown in Figure 22, if the width and height of the reference picture are 0.5 times the width and height of the decoded picture, the pixel positions and motion vector lengths of the reference picture relative to the decoded picture are scaled by 0.5 times.
[0210] On the other hand, when the interpretation unit 205 generates a chain of motion vectors, it is necessary to connect the motion vectors at the same scale as the decoded picture.
[0211] Therefore, as shown in Figure 22, the interpretation unit 205 may, when referencing motion vectors in reference pictures with a different resolution (sampling ratio) than the picture to be decoded and recursively searching for motion vectors from such reference pictures to generate a chain of motion vectors, descale the length of the recursively derived motion vectors according to the ratio of the width and height of the picture to be decoded to the width and height of each reference picture, and then add them together to generate a chain of motion vectors.
[0212] As an example of modification, if the interpretation unit 205 allows referencing motion vectors for reference pictures with a different resolution (sampling ratio) than the picture to be decoded, it may apply chained motion vectors only to reference pictures with the same sampling ratio as the picture to be decoded.
[0213] In other words, the interpretation unit 205 may, when it is permitted to reference a motion vector to a reference picture with a different resolution (sampling ratio) than the picture to be decoded, not apply a chain of motion vectors to a reference picture that does not have the same sampling ratio as the picture to be decoded.
[0214] <4. Inheritance of Interpretation Information in Chain-Based Motion Vector Prediction> The interpretation unit 205 may control whether or not to inherit predetermined internal parameters that change the method of generating interpretation pixels when deriving merge candidates by chain-reaction motion vector prediction.
[0215] (Switching interpolation filter) As an example, when deriving merge candidates by chain-repeated motion vector prediction, the interpretation unit 205 may not inherit the internal parameter (hpelIfIdx) that identifies whether or not to apply a switching interpolation filter associated with each recursively searched motion vector, and may always set the internal parameter (hpelIfIdx) that identifies whether or not to apply a switching interpolation filter corresponding to the merge candidates derived by chain-repeated motion vector prediction to indicate that the switching interpolation filter is not applied.
[0216] Alternatively, when deriving merge candidates by chain-link motion vector prediction, the interpretation unit 205 may inherit an internal parameter (hpelIfIdx) that specifies whether or not to apply a switching interpolation filter associated with each recursively searched motion vector, specifically the internal parameter (hpelIfIdx) that specifies whether or not to apply a switching interpolation filter associated with the first or last motion vector constituting the chain-link motion vector, as an internal parameter (hpelIfIdx) that specifies whether or not to apply a switching interpolation filter corresponding to the merge candidate derived by chain-link motion vector prediction.
[0217] Alternatively, when the interpretation unit 205 derives merge candidates by chain-repeated motion vector prediction, if at least one of the internal parameters (hpelIfIdx) that specify whether or not to apply a switching interpolation filter associated with each recursively searched motion vector indicates that a switching interpolation filter should be applied, then the internal parameter (hpelIfIdx) that specifies whether or not to apply a switching interpolation filter corresponding to the merge candidate derived by chain-repeated motion vector prediction may also indicate that a switching interpolation filter should be applied.
[0218] Furthermore, the interpretation unit 205 may always set an internal parameter (hpelIfIdx) that identifies whether or not to apply a switching interpolation filter corresponding to the merge candidate derived by the chain motion vector prediction, to indicate that the switching interpolation filter should not be applied, if the chain motion vector generated by the chain motion vector prediction consists of at least one block vector.
[0219] (Local brightness compensation) As an example, when deriving merge candidates by chain-repeated motion vector prediction, the interpretation unit 205 may not inherit the internal parameter (licFlag) that controls whether or not to apply local luminance compensation associated with each recursively searched motion vector, and may always set the internal parameter (licFlag) that controls whether or not to apply local luminance compensation corresponding to the merge candidates derived by chain-repeated motion vector prediction to indicate that local luminance compensation is not applied.
[0220] Alternatively, when deriving merge candidates by chain motion vector prediction, the interpretation unit 205 may inherit an internal parameter (licFlag) that controls whether or not to apply local luminance compensation associated with each recursively searched motion vector, specifically the internal parameter (licFlag) that controls whether or not to apply local luminance compensation associated with the first or last motion vector constituting the chain motion vector, as an internal parameter (licFlag) that controls whether or not to apply local luminance compensation corresponding to the merge candidate derived by chain motion vector prediction.
[0221] Alternatively, when the interpretation unit 205 derives merge candidates by chain-reaction motion vector prediction, if at least one of the internal parameters (licFlag) that control whether or not to apply local luminance compensation associated with each recursively searched motion vector indicates that local luminance compensation should be applied, then the internal parameter (licFlag) that controls whether or not to apply local luminance compensation corresponding to the merge candidate derived by chain-reaction motion vector prediction may also indicate that local luminance compensation should be applied.
[0222] Furthermore, if the chain of motion vectors generated by chain of motion vector prediction consists of at least one block vector, the interpretation unit 205 may always set an internal parameter (licFlagg) that controls whether or not to apply local luminance compensation corresponding to the merge candidate derived by chain of motion vector prediction to indicate that local luminance compensation should not be applied.
[0223] (Block-based weighted biprediction) As an example, when deriving merge candidates by chain-repeated motion vector prediction, the interpretation unit 205 may not inherit the internal parameter (bcwIdx) that identifies the weight values of the block-unit weighted biprediction associated with each recursively searched motion vector, but may always set the internal parameter (bcwIdx) that identifies the weight values of the block-unit weighted biprediction corresponding to the merge candidates derived by chain-repeated motion vector prediction to indicate that the weight values of the block-unit weighted biprediction are a simple average.
[0224] Alternatively, when deriving merge candidates by chain-link motion vector prediction, the interpretation unit 205 may inherit, from among the internal parameters (bcwIdx) that identify the weight values of block-unit weighted bipredictions associated with each recursively searched motion vector, the internal parameter (bcwIdx) that identifies the weight values of block-unit weighted bipredictions associated with the first or last motion vector constituting the chain-link motion vector, as the internal parameter (bcwIdx) that identifies the weight values of block-unit weighted bipredictions corresponding to the merge candidates derived by chain-link motion vector prediction.
[0225] As an example of modification, when the interpretation unit 205 derives merge candidates by chain-repeated motion vector prediction, if at least one of the internal parameters (bcwIdx) that identify the weight values of the block-unit weighted biprediction associated with each recursively searched motion vector indicates that the weight values of the block-unit weighted biprediction are the simple average, then the internal parameter (bceIdx) that identifies the weight values of the block-unit weighted biprediction corresponding to the merge candidate derived by chain-repeated motion vector prediction may also indicate that the weight values of the block-unit weighted biprediction are the simple average.
[0226] (Multiple Hypothetical Predictions) As an example, when deriving merge candidates by chain-repeated motion vector prediction, the interpretation unit 205 may choose not to inherit the internal parameter (mhpFlag, fourth internal parameter) that controls whether or not to apply multiple hypothesis predictions associated with each recursively searched motion vector, and instead always set the internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions corresponding to the merge candidates derived by chain-repeated motion vector prediction to indicate that multiple hypothesis predictions should not be applied.
[0227] As an example of modification, the interpretation unit 205 may also be shown to inherit, when deriving merge candidates by chain-link motion vector prediction, the internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions associated with each recursively searched motion vector, specifically the internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions associated with the first or last motion vector constituting the chain-link motion vector, as the internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions corresponding to the merge candidates derived by chain-link motion vector prediction.
[0228] As an example of modification, when the interpretation unit 205 derives merge candidates by chain-repeated motion vector prediction, if at least one of the internal parameters (mhpFlag) that control whether or not to apply multiple hypothesis predictions associated with each recursively searched motion vector indicates that multiple hypothesis predictions should be applied, then the internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions corresponding to the merge candidates derived by chain-repeated motion vector prediction may also indicate that multiple hypothesis predictions should be applied.
[0229] As a further modification, the interpretation unit 205 may be configured to always set an internal parameter (mhpFlag) that controls whether or not to apply multiple hypothesis predictions corresponding to merge candidates derived by the daisy-chain motion vector prediction, to indicate that multiple hypothesis predictions should not be applied, if the daisy-chain motion vectors generated by the daisy-chain motion vector prediction consist of at least one block vector.
[0230] As described above, if the internal parameters that change the method of generating interprediction pixels are not always inherited in chain-reaction motion vector prediction, the amount of code required for interprediction information to be stored in the decoded picture buffer 207 can be reduced.
[0231] On the other hand, as mentioned above, if the internal parameters that change the method of generating interpretation pixels are inherited by chain-reaction motion vector prediction, it may be possible to improve the accuracy of interpretation pixels using merge candidates generated by chain-reaction motion vector prediction, and an improvement in coding efficiency can be expected.
[0232] Furthermore, when inheriting the aforementioned internal parameters, if the chain of motion vectors consists of at least one block vector, the need to change the interpretation pixel generation method is low. Therefore, by not inheriting such internal parameters, i.e., by setting the method of generating interpretation pixels to not change, an improvement in code efficiency can be expected.
[0233] <5. Rearranging merge candidates, including chain-linked motion vector prediction candidates, and correcting the vectors> (Sort merge candidates) The later a chain of motion vector prediction candidates is derived compared to other merge candidates, the less likely it is to be stored in the merge candidate list.
[0234] As a way to solve this problem, the interpretation unit 205 may calculate the template matching cost of all merge candidates stored in the merge candidate list, which includes merge candidates derived by chain-reaction motion vector prediction, and then sort the storage positions of each merge candidate in the merge candidate list in ascending order, with the template matching cost decreasing from smallest to largest.
[0235] As an example, the interpretation unit 205 may use the Adaptive Reordering Merge Candidate (ARMC) disclosed in Non-Patent Document 2 as a technique for rearranging the storage positions of each merge candidate in the merge candidate list based on the template matching cost of each merge candidate.
[0236] Furthermore, as mentioned above, chain motion vector prediction recursively generates new motion vectors. Therefore, if other merge candidates are derived after the chain motion vector prediction candidate has been derived, those other merge candidates may not be stored in the merge candidate list.
[0237] As a way to solve this problem, the interpretation unit 205 may calculate the template matching costs of all unpruned merge candidates, including the merge candidates derived by chain-reaction motion vector prediction, and store them in the merge candidate list in ascending order of increasing template matching costs.
[0238] Alternatively, the interpretation unit 205 may perform a predetermined number of operations to calculate the template matching costs of all unpruned merge candidates, including the merge candidates derived by chain-reaction motion vector prediction, and to store them in the merge candidate list in ascending order of template matching costs, in ascending order. For example, this predetermined number of operations may be two or three.
[0239] This ensures that merge candidates capable of generating highly accurate interpretation pixels are placed within the merge candidate list, leading to improved coding efficiency.
[0240] (Vector correction) The interpretation unit 205 may correct the motion vectors associated with the merge candidates derived by chain-reaction motion vector prediction in a predetermined manner.
[0241] For example, the interpretation unit 205 may correct the motion vectors associated with the merge candidates derived by chain-reaction motion vector prediction using template matching.
[0242] Here, the interpretation unit 205 may apply a technique similar to the template matching disclosed in Non-Patent Document 2 as such template matching.
[0243] For example, the interpretation unit 205 may correct the motion vectors associated with the merge candidates derived by chain-reaction motion vector prediction using bilateral matching.
[0244] Here, the interpretation unit 205 may apply techniques similar to the bilateral matching techniques disclosed in Non-Patent Documents 1 and 2 as such bilateral matching. For example, the interpretation unit 205 may apply the decoder-side motion vector correction (DMVR) disclosed in Non-Patent Document 1.
[0245] For example, the interpretation unit 205 may correct the motion vector associated with the merge candidate derived by chain-reaction motion vector prediction by adding the motion vector difference derived from the control information.
[0246] Here, the interpretation unit 205 may apply a technique similar to the merge with motion vector difference (MMVD) disclosed in Non-Patent Documents 1 and 2 as control information indicating the motion vector difference.
[0247] The interpretation unit 205 corrects the motion vectors associated with the merge candidates derived by chain-reaction motion vector prediction using a predetermined method, thereby improving the interpretation accuracy and potentially improving coding efficiency.
[0248] <6. Various limitations on chain-reaction motion vector prediction (applicability, number of pictures, number of merge candidates)> Using Figure 23, various methods for limiting chain-reaction motion vector prediction by the decoding unit 201 and the interpretation unit 205 of the image decoding device 200 according to this embodiment will be explained. (Whether or not chain-reaction motion vector prediction is applied) Firstly, regarding the applicability of chain-reaction motion vector prediction, as mentioned above, by its very nature, chain-reaction motion vector prediction is applicable when TMVP is applicable. In other words, if TMVP is not applicable, chain-reaction motion vector prediction is not applicable.
[0249] Therefore, the applicability of chain-reaction motion vector prediction can be controlled according to control information (syntax) that controls the applicability of time motion vector prediction (TMVP).
[0250] In other words, the interpretation unit 205 may control whether or not chained motion vector prediction is applicable based on control information sent from the decoding unit 201 that controls whether or not time motion vector prediction (TMVP) is applicable.
[0251] In Non-Patent Document 1, two syntaxes are defined to control whether TMVP is applicable or not: a syntax for each sequence to be decoded (sps_temporal_mvp_enabled_flag, first control information, first syntax) and a syntax for each picture to be decoded (ph_temporal_mvp_enabled_flag, second control information, second syntax).
[0252] In other words, the applicability of chain-reaction motion vector prediction can be controlled using units similar to or lower than those used for TMVP.
[0253] For example, the method for controlling whether or not to apply chain-reaction motion vector prediction in the interpretation prediction 205 section, which controls whether or not to apply chain-reaction motion vector prediction at the unit of the blocks to be decoded, is as follows.
[0254] The interpretation unit 205 may control whether or not chain-reaction motion vector prediction can be applied to the decoded block based on control information sent from the decoding unit 201 that controls whether or not time motion vector prediction (TMVP) can be applied to the decoded sequence unit.
[0255] More specifically, the interpretation unit 205 determines that chain-reaction motion vector prediction is applicable to the decoded block when the control information sent from the decoding unit 201, which controls whether or not time motion vector prediction (TMVP) can be applied to the decoded sequence unit, indicates that time motion vector prediction is applicable.
[0256] On the other hand, if the control information described above determines that time motion vector prediction is not applicable, the interpretation unit 205 determines that chain-reaction motion vector prediction is not applicable to the block to be decoded.
[0257] Alternatively, the inter prediction unit 205 may control whether to apply daisy-chain motion vector prediction to the block to be decoded, based on control information that controls whether temporal motion vector prediction (TMVP) is applicable for each decoding target picture unit sent from the decoding unit 201.
[0258] More specifically, when the control information sent from the decoding unit 201 for controlling whether temporal motion vector prediction (TMVP) is applicable for each decoding target picture unit specifies that temporal motion vector prediction can be applied, the inter prediction unit 205 specifies that daisy-chain motion vector prediction can be applied to the block to be decoded.
[0259] On the other hand, when the above control information specifies that temporal motion vector prediction cannot be applied, the inter prediction unit 205 specifies that daisy-chain motion vector prediction cannot be applied to the block to be decoded.
[0260] The inter prediction unit 205 may control whether to apply daisy-chain motion vector prediction in units of decoding target sequences or in units of decoding targets smaller than decoding target sequences, based on control information that controls whether temporal motion vector prediction (TMVP) is applicable for each decoding target sequence unit sent from the decoding unit 201.
[0261] Alternatively, the inter prediction unit 205 may control whether to apply daisy-chain motion vector prediction in units of decoding target pictures or in units of decoding targets smaller than decoding target pictures, based on control information that controls whether temporal motion vector prediction (TMVP) is applicable for each decoding target picture unit sent from the decoding unit 201.
[0262] Alternatively, before the inter prediction unit 205 controls whether to apply daisy-chain motion vector prediction in units of blocks to be decoded, the decoding unit 201 may control whether to apply daisy-chain motion vector prediction based on control information that controls whether temporal motion vector prediction (TMVP) is applicable.
[0263] FIG. 23 is a diagram illustrating an example of processing in which a decoding unit 201 determines whether to decode a syntax (sps_chained_mvp_enabled_flag) that defines whether chained motion vector prediction is applicable per decoding target sequence unit, based on a syntax (sps_temporal_mvp_enabled_flag) that controls whether temporal motion vector prediction (TMVP) is applicable per decoding target sequence unit.
[0264] As shown in FIG. 23, when the decoding unit 201 determines in step S100 that the syntax (sps_temporal_mvp_enabled_flag) that controls whether temporal motion vector prediction (TMVP) is applicable per decoding target sequence unit is 1, in step S101, the decoding unit 201 decodes the syntax (sps_chained_mvp_enabled_flag, a third syntax) that controls whether chained motion vector prediction is applicable per decoding target sequence unit.
[0265] On the other hand, when the decoding unit 201 determines in step S100 that the syntax (sps_temporal_mvp_enabled_flag) that controls whether temporal motion vector prediction (TMVP) is applicable per decoding target sequence unit is not 1, in step S102, the decoding unit 201 does not decode the syntax (sps_chained_mvp_enabled_flag) that controls whether chained motion vector prediction is applicable per decoding target sequence unit.
[0266] Here, the decoding unit 201 determines that time motion vector prediction (TMVP) can be applied to each sequence to be decoded if the value of sps_temporal_mvp_enabled_flag is 1, determines that time motion vector prediction (TMVP) cannot be applied to each sequence to be decoded if the value of sps_temporal_mvp_enabled_flag is 0, and if sps_temporal_mvp_enabled_flag has not been decoded, it estimates the value of sps_temporal_mvp_enabled_flag to be 0.
[0267] Furthermore, the decoding unit 201 determines that chained motion vector prediction can be applied to each sequence to be decoded if the value of sps_chained_mvp_enabled_flag is 1, determines that chained motion vector prediction cannot be applied to each sequence to be decoded if the value of sps_chained_mvp_enabled_flag is 0, and estimates the value of sps_chained_mvp_enabled_flag to be 0 if sps_chained_mvp_enabled_flag has not been decoded.
[0268] In this way, by determining whether or not to decode sps_chained_mvp_enabled_flag based on the value of sps_temporal_mvp_enabled_flag, the amount of code can be reduced because sps_chained_mvp_enabled_flag does not need to be decoded in cases where it is obvious that decoding sps_chained_mvp_enabled_flag is unnecessary.
[0269] Chain-reaction motion vector prediction can be used in conjunction with the aforementioned Adaptive Reordering Merge Candidate (ARMC) method, which involves comparing the template matching costs of merge candidates, to further improve encoding efficiency.
[0270] Therefore, as an example of a change in the control method for applying chain-reaction motion vector prediction based on the applicability of TMVP as described above, the interpretation unit 205 may control the applicability of chain-reaction motion vector prediction based on both the applicability of TMVP and the applicability of ARMC.
[0271] Specifically, the interpretation unit 205 may specify that chain-reaction motion vector prediction is applicable when TMVP is applicable and ARMC is applicable. In other cases, the interpretation unit 205 may specify that chain-reaction motion vector prediction is not applicable.
[0272] Furthermore, the applicability of ARMC may be determined based on the values of control information similar to that used for TMVP.
[0273] (Number of reference pictures and search depth for chain-link motion vector prediction) The following describes how to limit the number of reference pictures or the search depth that can be referenced when recursively searching for motion vectors in chain-reaction motion vector prediction.
[0274] First, we will explain how to limit the number of reference pictures that can be referenced when recursively searching for motion vectors in chain-reaction motion vector prediction.
[0275] As an example, the interpretation unit 205 may restrict the recursive search for motion vectors in chained motion vector prediction to only at least one reference picture included in the reference picture list of the block to be decoded.
[0276] The interpretation unit 205 reduces the decoding processing load compared to when no such restriction is imposed (for example, when the interpretation unit 205 searches all the reference pictures stored in the decoded picture buffer 207) by limiting the recursive search for motion vectors in chained motion vector prediction to only at least one reference picture included in the reference picture list of the block to be decoded.
[0277] Alternatively, the interpretation unit 205 may restrict the recursive search for motion vectors in the chain-reaction motion vector prediction to only the reference pictures referenced by the motion vector of the block to be decoded.
[0278] In other words, the interpretation unit 205 re-searches only the block vectors within the reference picture referenced by the motion vector of the block to be decoded, and generates a chain of motion vectors through chain of motion vector prediction.
[0279] Secondly, we will describe a method for limiting the recursive motion vector search depth in chained motion vector prediction.
[0280] As an example, the interpretation unit 205 may limit the recursive motion vector search depth for predicting chained motion vectors in the decoded block based on the control information.
[0281] Specifically, the interpretation unit 205 may limit the search depth of the recursive motion vectors in the chain-linked motion vector prediction in the decoded block based on control information (sps_max_num_my_chains) that sets an upper limit on the search depth of the recursive motion vectors in the chain-linked motion vector prediction for each sequence to be decoded sent from the decoding unit.
[0282] Here, the control information for each sequence to be decoded may be lower-level control information, such as picture-level or slice-level control information, or it may be hierarchical control information combining these.
[0283] Alternatively, the interpretation unit 205 may set a predetermined fixed value for the recursive motion vector search depth of the chained motion vector prediction. The interpretation unit 205 may set 1, 2, or 3 as such predetermined fixed value.
[0284] By limiting the search depth of daisy-chaining motion vector prediction in this way, the amount of decoding processing required for deriving daisy-chaining motion vectors can be reduced compared to the case where the amount of decoding processing is not limited.
[0285] (Number of merge candidates for daisy-chaining motion vector prediction) Hereinafter, a method for limiting the number of merge candidates derived by daisy-chaining motion vector prediction will be described.
[0286] As an example, the inter prediction unit 205 may limit the number of merge candidates derived by daisy-chaining motion vector prediction for the block to be decoded based on control information.
[0287] Specifically, the inter prediction unit 205 may be configured based on control information (sps_max_num_cmvp_cand) that sets an upper limit on the search depth of the number of merge candidates derived by daisy-chaining motion vector prediction for each sequence unit to be decoded sent from the decoding unit, the number of merge candidates derived by daisy-chaining motion vector prediction for the block to be decoded may be limited.
[0288] Here, the control information in units of sequences to be decoded may be control information in units of pictures or slices, which is lower-level control information than the units of sequences to be decoded, or may be hierarchical control information combining these.
[0289] Alternatively, the inter prediction unit 205 may set a predetermined fixed value for the number of merge candidates derived by daisy-chaining motion vector prediction.
[0290] (Limitation of block vector search in daisy-chaining motion vector prediction) It has been described above that block vector search may be included in the recursive motion vector search of daisy-chaining motion vector prediction. However, block vectors referencing the same picture are easily found in screen images, game images, etc. containing text and computer graphics (CG), but are difficult to find in images captured by a camera.
[0291] Therefore, it is desirable to be able to restrict the search for block vectors in chain motion vector prediction depending on the features of the image.
[0292] As an example, the interpretation unit 205 may restrict the block vector search based on control information that specifies whether or not a block vector search is possible in the chain-reaction motion vector prediction for the block to be decoded.
[0293] Specifically, the interpretation unit 205 may restrict the ability to search for block vectors in chain-reaction motion vector prediction for a given block based on control information (sps_cmvp_bv_search_enabled_flag) that specifies whether or not block vectors can be searched in chain-reaction motion vector prediction for each sequence to be decoded, which is sent from the decoding unit 201.
[0294] Here, the control information for each sequence to be decoded may be lower-level control information, such as picture-level or slice-level control information, or it may be hierarchical control information combining these. (Restrictions on the reference sources for chain-reaction motion vector prediction) In the recursive search for a sequence of motion vectors in chain-reaction motion vector prediction, the method described above adds the four corners of the target block to the center of the initial motion vector.
[0295] However, compared to screen images and game images, where pixel values tend to change rapidly at the pixel level, images captured by a camera are less prone to change.
[0296] Therefore, depending on the features of the image, it is desirable to be able to restrict the starting point of the first motion vector in the recursive search for a chain of motion vectors.
[0297] As an example, the interpretation unit 205 may restrict the reference source for the chain motion vector prediction based on control information that specifies whether to add at least one predetermined location within the chain motion vector prediction for the block to be decoded, in addition to the center of the block to be decoded. Such predetermined locations may be the four corners within the block to be decoded as described above.
[0298] Specifically, the interpretation unit 205 may, based on sps_cmvp_multi_position_flag, restrict whether to add at least one predetermined position on the decryption target block in addition to the center of the decryption target block by adding the starting point of the first motion vector in the chain motion vector prediction for the decryption target block.
[0299] Here, sps_cmvp_multi_position_flag is control information that specifies whether to add at least one predetermined position on the block to be decoded, in addition to the center of the block to be decoded, to the starting point of the first motion vector in the chain motion vector prediction for the block to be decoded.
[0300] Here, the control information for each sequence to be decoded may be lower-level control information, such as picture-level or slice-level control information, or it may be hierarchical control information combining these.
[0301] <7. Applicability of Chain-Based Motion Vector Prediction> This document describes the technical characteristics of applying chain-reaction motion vector prediction to merges other than the normal merges disclosed in Non-Patent Documents 1 and 2. (MMVD, GPM, TM merge) Firstly, the MMVD, GPM, and template matching merge (TM merge) disclosed in Non-Patent Documents 1 and 2 construct merge lists, similar to normal merges. However, in Non-Patent Document 2, the pruning of merge candidates during the construction of these merge candidate lists is enhanced compared to normal merges, making these merge candidate lists less likely to be identical to the merge candidates stored in the merge candidate list of a normal merge.
[0302] Therefore, in MMVD, GPM, and TM merges, if the merge candidate list from the normal merge is not reused, applying chain-reaction motion vector prediction to the merge candidates in the respective merge candidate lists for MMVD, GPM, and TM merges, in addition to applying chain-reaction motion vector prediction to the merge candidates in the normal merge candidate list, can be expected to improve encoding efficiency.
[0303] Therefore, in addition to applying chain-reaction motion vector prediction to the merge candidates in the merge candidate list for normal merge, the interpretation unit 205 may also apply chain-reaction motion vector prediction to the merge candidates in the merge candidate list for merge mode motion vector differences.
[0304] Furthermore, the interpretation unit 205 may apply chain-reaction motion vector prediction to the merge candidates in the merge candidate list for normal merge, as well as to the merge candidates in the merge candidate list for geometric partitioning mode.
[0305] Furthermore, the interpretation unit 205 may apply chain-reaction motion vector prediction to the merge candidates in the merge candidate list for normal merges, as well as to the merge candidates in the merge candidate list for template matching merges.
[0306] Here, the interpretation unit 205 may apply pruning, motion vector correction, and reordering of merge candidates (motion vectors) to the chain of motion vectors generated for merge mode motion vector difference, geometric partitioning mode, and template matching merge, similar to normal merging.
[0307] (Adaptive motion vector prediction - merge mode) In the Adaptive Motion Vector Prediction-Merge mode disclosed in Non-Patent Documents 1 and 2 mentioned above, the motion vector of either List 0 (L0) or List 1 (L1) of the reference picture list is derived by Adaptive Motion Vector Prediction (AMVP), i.e., by control information. Therefore, if the application of chained motion vector prediction is made to merge candidates in the merge candidate list of reference picture list numbers not derived by AMVP, the number of selectable merge candidates in Adaptive Motion Vector Prediction-Merge mode increases, which is expected to improve inter-prediction accuracy and coding efficiency.
[0308] Therefore, in the adaptive motion vector prediction-merge mode, the interpretation unit 205 may apply chain motion vector prediction to merge candidates for reference picture list numbers in the merge candidate list that have not been derived by adaptive motion vector prediction.
[0309] Furthermore, in the adaptive motion vector prediction-merge mode, the interpretation unit 205 may restrict the generated chain of motion vectors from referencing the reference pictures of motion vectors derived by adaptive motion vector prediction when applying chain of motion vector prediction to merge candidates for reference picture lists in the merge candidate list that have not been derived by adaptive motion vector prediction.
[0310] Here, the interpretation unit 205 may apply pruning when storing the merge candidate list, motion vector correction, and rearrangement of merge candidates (motion vectors) to the chain of motion vectors generated for the adaptive motion vector prediction-merge mode, similar to normal merging.
[0311] (Subblock merge mode) The affine merge disclosed in Non-Patent Documents 1 and 2 above derives the motion vectors for each subblock within the decoded block using an affine transform. As shown in Figure 24, in the case of single prediction (either L0 or L1), two or three control point motion vectors cpMv are derived, and in the case of double prediction (both L0 and L1), four or six control point motion vectors cpMv are derived.
[0312] Note that the L0 and L1 control point motion vectors for single-prediction and double-prediction all refer to the same reference picture.
[0313] In affine merging, these control point motion vectors are derived from merge candidates stored in a merge candidate list for affine merging, as disclosed in Non-Patent Documents 1 and 2.
[0314] The inter-prediction unit 205 may derive merge candidates to be stored in the merge candidate list for affine merge in the same manner as the method for deriving merge candidates for affine merge disclosed in Non-Patent Document 1 and Non-Patent Document 2.
[0315] For example, if two or three control point motion vectors are accumulated at positions spatially adjacent to or near the block to be decoded, the interpretation unit 205 may add them to the merge candidate list for affine merging of the block to be decoded.
[0316] Alternatively, if multiple motion vectors are accumulated at positions spatially adjacent to or near the block to be decoded, the interpretation unit 205 may generate control point motion vectors for the block to be decoded based on these vectors and add them to the merge candidate list for affine merging.
[0317] Therefore, in affine merge mode, the interpretation unit 205 may recursively search for motion vectors for at least two or more control point motion vectors stored in the merge candidate list, add the derived motion vectors to each of the initial control point motion vectors to generate a chain of control point motion vectors, and store them in the merge candidate list.
[0318] Furthermore, in affine merge mode, when the interpretation unit 205 recursively searches for motion vectors for at least two or more control point motion vectors stored in the merge candidate list, it may restrict the reference picture to be searched for each control point motion vector to the same reference picture.
[0319] Alternatively, when the interpretation unit 205 recursively searches for motion vectors for at least two or more control point motion vectors of L0 and L1 stored in the merge candidate list in affine merge mode, it may restrict the reference picture to be searched for each control point motion vector of L0 and L1 to the same reference picture.
[0320] Here, the interpretation unit 205 may apply pruning processing when storing the merge candidate list, correction of the motion vectors, and reordering of merge candidates (motion vectors) to the chain of control point motion vectors generated for affine merging, similar to normal merging.
[0321] According to the image decoding device of this embodiment, the interpretation unit 205 recursively searches for motion vectors stored in the decoding picture buffer 207 for motion vectors at positions that are spatially or temporally adjacent or close to the decoding target block, and derives a motion vector generated by adding all the derived motion vectors as the motion vector of the decoding target block. As a result, the number of motion vectors that can be selected in interpretation increases, interpretation accuracy is improved, and coding efficiency is improved.
[0322] The image decoding device 200 described above may be implemented as a program that causes a computer to execute each function (each process). [Industrial applicability]
[0323] Furthermore, according to this embodiment, for example, it is possible to achieve an overall improvement in service quality in video communication, thereby contributing to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote sustainable industrialization and foster innovation." [Explanation of Symbols]
[0324] 200…Image decoding device 201...Decoding section 202...Inverse quantization section 203...Inverse Transformation Section 204...Intra Prediction Unit 205...Interface Forecasting Department 206... Adder 207...Decrypted picture buffer 210... Code input section 220...Image output unit
Claims
1. An image decoding device, It includes an inter-prediction unit that generates inter-prediction pixels from control information and decoded pixels stored in a decoded picture buffer, The image decoding apparatus is characterized in that the interpretation unit generates a chain of motion vectors by recursively searching for motion vectors stored in the decoding picture buffer for motion vectors at positions that are spatially or temporally adjacent or close to the decoding target block stored in the merge candidate list which stores candidate motion vectors, and adding all the derived motion vectors, and derives the chain of motion vectors as a candidate for the motion vector of the decoding target block, thereby performing chain of motion vector prediction.
2. The image decoding apparatus according to claim 1, characterized in that the interpretation unit, in the chain-linked motion vector prediction, adds block vectors indicating references within the same picture to the recursive search targets in addition to the motion vectors stored in the decoded picture buffer.
3. The image decoding device according to claim 1, wherein the interpretation unit, in the chain-linked motion vector prediction, searches for one or more motion vectors from a single motion vector that is a reference source for each search list number of the merge candidate list and the search depth during the recursive search of motion vectors in the chain-linked motion vector prediction, and generates the chain-linked motion vector by adding all the already derived motion vectors to each of the one or more newly derived motion vectors.
4. The image decoding apparatus according to claim 1, characterized in that the interpretation unit searches for a chain of motion vector prediction candidates after searching for a predetermined merge candidate.
5. The image decoding apparatus according to claim 4, characterized in that the interpretation unit searches for the chain-linked motion vector prediction candidates after searching for time motion vector prediction candidates.
6. The image decoding apparatus according to claim 4, characterized in that the interpretation unit searches for the chain-linked motion vector prediction candidates after searching for the history-based motion vector prediction candidates.
7. The image decoding apparatus according to claim 4, characterized in that the interpretation unit searches for the chain-linked motion vector prediction candidates after searching for spatial merge candidates.
8. The aforementioned interpretation unit, The predetermined merge candidates are stored in the merge candidate list, The chain-reaction motion vector prediction is applied sequentially, starting with the first merge candidate stored in the merge candidate list. The image decoding device according to claim 4, characterized in that new merge candidates generated by the chain-link motion vector prediction for each of the merge candidates are sequentially stored in the merge candidate list.
9. The aforementioned interpretation unit, The predetermined merge candidates are stored in the merge candidate list, The chain-reaction motion vector prediction is applied sequentially, starting with the first merge candidate stored in the merge candidate list. The image decoding device according to claim 4, characterized in that new merge candidates generated for each of the merge candidates using each search depth in the chain-link motion vector prediction are sequentially stored in the merge candidate list.
10. The aforementioned interpretation unit, Store the first merge candidate in the merge candidate list. The chain-reaction motion vector prediction is applied sequentially to the merge candidates, The image decoding device according to any one of claims 1 to 4, characterized in that new merge candidates generated by the chain-link motion vector prediction are sequentially stored in the merge candidate list for each of the merge candidates.
11. The aforementioned interpretation unit, Store the first merge candidate in the merge candidate list. The chain-reaction motion vector prediction is applied sequentially to the merge candidates, The image decoding apparatus according to any one of claims 1 to 4, characterized in that new merge candidates generated for each of the merge candidates using the respective search depths of the chain-link motion vector prediction are sequentially stored in the merge candidate list.
12. The image decoding apparatus according to any one of claims 1 to 4, characterized in that the interpretation unit generates new merge candidates by combining the chained motion vectors generated using each search depth in the chained motion vector prediction with the chained motion vectors generated using a different search depth for each merge candidate list.
13. The image decoding apparatus according to any one of claims 1 to 4, characterized in that the interpretation unit generates new merge candidates by combining a chain of motion vectors generated for each merge candidate list using each search list number or each search depth of the chain of motion vector prediction with a chain of motion vectors generated using a different search list number or search depth than the previous search list number or each search depth.
14. An image decoding method, This includes generating interpredicted pixels from control information and decoded pixels stored in the decoded picture buffer, Image decoding method characterized by generating interpredicted pixels by generating a chain of motion vectors by recursively searching for motion vectors accumulated in the decoding picture buffer for motion vectors at positions spatially or temporally adjacent or close to the decoding target block stored in a merge candidate list that stores motion vector candidates, and adding all derived motion vectors to generate a chain of motion vectors, and deriving the chain of motion vectors as a candidate for the motion vector of the decoding target block, thereby performing chain of motion vector prediction.
15. A program that makes a computer function as an image decoding device, The aforementioned image decoding device is It includes an inter-prediction unit that generates inter-prediction pixels from control information and decoded pixels stored in a decoded picture buffer, The program is characterized by performing a chain-link motion vector prediction, in which the interpretation unit recursively searches for motion vectors stored in the decoding picture buffer for motion vectors at positions that are spatially or temporally adjacent or close to the decoding target block stored in the merge candidate list that stores candidate motion vectors for motion vectors, and adds all the derived motion vectors to generate a chain-link motion vector, and derives the chain-link motion vector as a candidate for the motion vector of the decoding target block.
Citation Information
Patent Citations
ITTH.266/
Systems and methods for performing motion vector prediction using a derived set of motion vectors
WO2019216325A1
Systems and methods for deriving a motion vector prediction in video coding
WO2020100988A1
Motion compensation considering out-of-boundary conditions in video coding
WO2023076700A1
Method, apparatus, and medium for video processing
WO2023098899A1