Image decoding device, image decoding method, and program

By recursively searching for motion vectors adjacent to or close to the decoded target block in the decoded image buffer, chained motion vector predictions are generated, which solves the problem of limited motion vector candidate selection in inter-frame prediction in the prior art and improves coding efficiency.

CN121753337APending Publication Date: 2026-03-27KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, the motion vector candidates selected for inter-frame prediction are limited to positions that are spatially or temporally adjacent to or close to the target block being decoded, leaving room for improvement in coding efficiency.

Method used

By recursively searching for motion vectors at positions adjacent to or close to the target block in the decoded image buffer, chained motion vectors are generated and chained motion vector prediction is performed as motion vector candidates for the target block, thus controlling whether temporal motion vector prediction can be applied.

Benefits of technology

It improves the encoding efficiency of image decoding and provides a more efficient encoding method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753337A_ABST
    Figure CN121753337A_ABST
Patent Text Reader

Abstract

And coding efficiency is improved. In this image decoding device (200), an inter prediction unit (205) controls whether or not chain motion vector prediction can be applied on the basis of control information that controls whether or not temporal motion vector prediction can be applied, said control information being transmitted from a decoding unit (201).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an image decoding apparatus, an image decoding method, and a program. Background Technology

[0002] Inter-frame prediction has been disclosed in Non-Patent Literature 1 and Non-Patent Literature 2.

[0003] Inter-frame prediction generates predicted pixels for the target block from decoded pixels (reference pixels) within a decoded image (reference image) that is different from the target image.

[0004] In addition, inter-frame prediction uses control information to select the motion vector of the decoded target block needed to generate the predicted pixels of the decoded target block from a pool of candidates of motion vectors that are spatially or temporally adjacent to or close to the decoded target block. Existing technical documents Non-patent literature

[0005] Non-patent literature 1: ITU-T H.266 / VVC Non-Patent Literature 2: M. Coban et al., Algorithm description of Enhanced Compression Model 10 (ECM 10), JVET-AE2025, 2023 Summary of the Invention The problem that the invention aims to solve

[0006] In Non-Patent Literature 1 and Non-Patent Literature 2, the candidates for motion vectors that can be selected using inter-frame prediction are limited to motion vectors that are spatially or temporally adjacent to or close to the target block being decoded. Therefore, there is room for improvement in coding efficiency.

[0007] Therefore, the present invention was made in view of the above-mentioned problems, and its object is to provide an image decoding device, image decoding method and program with high encoding efficiency. Solution for solving the problem

[0008] The first feature of this invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs prediction residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels to the aforementioned prediction residual pixels. The difference pixels are added together to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors located spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit controls whether the aforementioned chained motion vector prediction can be applied based on control information sent from the aforementioned decoding unit to control whether temporal motion vector prediction can be applied.

[0009] The second feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that performs an addition operation on at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels. The aforementioned predicted residual pixels are added to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors located spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned decoding unit controls whether the aforementioned chained motion vector prediction can be applied based on control information that controls whether temporal motion vector prediction can be applied.

[0010] The third feature of this invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that performs an addition operation on at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels and the aforementioned predicted residual pixels. The aforementioned decoded pixels are generated by adding motion vectors. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The aforementioned chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit restricts the recursive search for motion vectors in the aforementioned chained motion vector prediction to only at least one reference image included in the reference image list of the aforementioned decoded target block.

[0011] The fourth feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels to the aforementioned predicted residual pixels. The residual pixels are added together to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The aforementioned chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit restricts the recursive search for motion vectors in the aforementioned chained motion vector prediction to reference images that only refer to the motion vectors of the aforementioned decoded target block.

[0012] The fifth feature of this invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantization values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantization values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs prediction residual pixels; an intra-frame prediction unit that generates intra-frame prediction pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame prediction pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame prediction pixels and the aforementioned inter-frame prediction pixels to the aforementioned prediction residual pixels. The motion vectors are added together to generate the aforementioned decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all the derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The inter-frame prediction unit limits the search depth of the recursive motion vectors in the aforementioned chained motion vector prediction in the aforementioned decoded target block based on the aforementioned control information sent from the aforementioned decoding unit.

[0013] The sixth feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that performs an addition operation on at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels. The aforementioned predicted residual pixels are added to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The aforementioned chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit configures a predetermined fixed value as the search depth of the recursive motion vectors in the aforementioned chained motion vector prediction.

[0014] The seventh feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantization values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantization values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs prediction residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels to the aforementioned prediction residual. The pixels are added together to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all the derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merge candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit limits the number of merge candidates derived by the aforementioned chained motion vector prediction in the aforementioned decoded target block based on the aforementioned control information sent from the aforementioned decoding unit.

[0015] The eighth feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that performs an addition operation on the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels. One of the missing pixels is added to the aforementioned predicted residual pixel to generate the aforementioned decoded pixel. The aforementioned inter-frame prediction unit generates a chained motion vector by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all the derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vector as a candidate for the motion vector of the aforementioned decoded target block. The aforementioned merge candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit configures a predetermined fixed value as the number of merge candidates derived by the aforementioned chained motion vector prediction.

[0016] The ninth feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantization values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantization values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels to the aforementioned predicted residual pixels. The aforementioned decoded pixels are generated by summing the motion vectors. The inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors that are spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and summing all derived motion vectors. The chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit determines whether to perform a search for block vectors in the aforementioned chained motion vector prediction for the aforementioned decoded target block based on the aforementioned control information sent from the aforementioned decoding unit.

[0017] The tenth feature of the present invention is an image decoding apparatus, comprising: a decoding unit that performs variable-length decoding on code information and outputs quantization values ​​and control information; an inverse quantization unit that performs inverse quantization on the aforementioned quantization values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the aforementioned transform coefficients and outputs prediction residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the aforementioned control information and decoded pixels; a decoded image buffer that stores the aforementioned decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the aforementioned control information and the aforementioned decoded pixels stored in the aforementioned decoded image buffer; and an adder that adds at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels to the aforementioned prediction residual pixels to generate the aforementioned decoded pixels, wherein the aforementioned inter-frame prediction unit performs inverse transform on ... to generate the aforementioned decoded pixels. The motion vectors that are spatially or temporally adjacent to or close to the decoded target block in the candidate list of candidates stored in the storage of motion vectors are recursively searched for motion vectors stored in the aforementioned decoded image buffer and all derived motion vectors are added together to generate chained motion vectors. Chained motion vector prediction is then performed, and the aforementioned chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned candidate list stores the motion vector candidates. Based on the aforementioned control information sent from the aforementioned decoding unit, the aforementioned inter-frame prediction unit determines whether to add the starting point of the first motion vector in the aforementioned chained motion vector prediction for the aforementioned decoded target block to at least one predetermined position of the aforementioned decoded target block, in addition to adding it to the center of the aforementioned decoded target block.

[0018] The eleventh feature of this invention is an image decoding method, comprising: step A, performing variable-length decoding on code information and outputting quantization values ​​and control information; step B, performing inverse quantization on the aforementioned quantization values ​​and outputting transform coefficients; step C, performing inverse transform on the aforementioned transform coefficients and outputting predicted residual pixels; step D, generating intra-frame predicted pixels from the aforementioned control information and decoded pixels; step E, storing the aforementioned decoded pixels in a decoded image buffer; step F, generating inter-frame predicted pixels from the aforementioned control information and the stored aforementioned decoded pixels; and step G, combining at least one of the aforementioned intra-frame predicted pixels and the aforementioned inter-frame predicted pixels with the aforementioned predicted residual image buffer. The aforementioned decoded pixels are generated by summing the motion vectors. In step F, a chain of motion vectors is generated by recursively searching for motion vectors in the aforementioned decoded image buffer for positions that are spatially or temporally adjacent to or close to the decoded target block stored in the merge candidate list and summing all the derived motion vectors. Chain of motion vector prediction is then performed, and the aforementioned chain of motion vector prediction derives the aforementioned chain of motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merge candidate list stores the motion vector candidates. In step F, the application of the aforementioned chain of motion vector prediction is controlled based on control information that controls whether temporal motion vector prediction can be applied.

[0019] The twelfth feature of the present invention is a program that enables a computer to function as an image decoding device, the image decoding device comprising: a decoding unit that performs variable-length decoding on code information and outputs quantized values ​​and control information; an inverse quantization unit that performs inverse quantization on the quantized values ​​and outputs transform coefficients; an inverse transform unit that performs inverse transform on the transform coefficients and outputs predicted residual pixels; an intra-frame prediction unit that generates intra-frame predicted pixels from the control information and decoded pixels; a decoded image buffer that stores the decoded pixels; an inter-frame prediction unit that generates inter-frame predicted pixels from the control information and the stored decoded pixels; and an adder that performs an addition operation on at least one of the intra-frame predicted pixels and the inter-frame predicted pixels. The aforementioned predicted residual pixels are added to generate the aforementioned decoded pixels. The aforementioned inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors located spatially or temporally adjacent to or close to the decoded target block stored in the aforementioned decoded image buffer and adding all derived motion vectors. The aforementioned chained motion vector prediction derives the aforementioned chained motion vectors as candidates for motion vectors of the aforementioned decoded target block. The aforementioned merged candidate list stores the motion vector candidates. The aforementioned inter-frame prediction unit controls whether the aforementioned chained motion vector prediction can be applied based on control information sent from the aforementioned decoding unit to control whether temporal motion vector prediction can be applied. Invention Effects

[0020] According to the present invention, an image decoding apparatus, an image decoding method, and a program with high encoding efficiency can be provided. Attached Figure Description

[0021] Figure 1 This is a diagram illustrating an example of the functional blocks of an image decoding apparatus 200 according to one embodiment. Figure 2 This is a diagram illustrating the concept of inter-frame prediction. Figure 3 It is a diagram used to illustrate the reference pixel values ​​disclosed in each of Non-Patent Document 1 and Non-Patent Document 2, when the reference is outside the decoded target image stored in the decoded image buffer 207 from the future decoded target block. Figure 4 It is generated in Figure 3 The conceptual diagram of the inter-frame prediction padding region is explained in the text. Figure 5 This is a diagram showing an example of a merged candidate list as a method for deriving motion vectors in the inter-frame prediction unit 205. Figure 6 This is a diagram illustrating an example of the names of various merger candidates and the order in which these merger candidates are derived in a regular merge disclosed in Non-Patent Document 1 and Non-Patent Document 2. Figure 7 This is a diagram illustrating an example of the search positions for spatial merge candidates and TMVP candidates in the decoded target block. Figure 8 This is a diagram illustrating an example of a method for deriving chain motion vectors in chain motion vector prediction. Figure 9 This is a diagram illustrating an example of a method for deriving chain motion vectors in chain motion vector prediction. Figure 10 This is a diagram illustrating an example of a recursive search for motion vectors that include bidirectional prediction. Figure 11 This is a diagram illustrating an example of a chain of motion vectors derived through a recursive search that includes bidirectional prediction of motion vectors. Figure 12 This is a diagram illustrating an example, as disclosed in Non-Patent Document 1 and Non-Patent Document 2, of searching for chain motion vector prediction candidates after searching for multiple merger candidates in a conventional merging process. Figure 13 This is a diagram illustrating an example, as disclosed in Non-Patent Document 1 and Non-Patent Document 2, of searching for chain motion vector prediction candidates after searching for multiple merger candidates in a conventional merging process. Figure 14This is a diagram illustrating an example of the merged candidates as derived objects of chained motion vector prediction candidates, and the position (index number) of the chained motion vector prediction candidates added to the merged candidate list. Figure 15 This is a diagram illustrating an example of the merged candidates as derived objects of chained motion vector prediction candidates, and the position (index number) of the chained motion vector prediction candidates added to the merged candidate list. Figure 16 This is a diagram illustrating an example of the merged candidates as derived objects of chained motion vector prediction candidates, and the position (index number) of the chained motion vector prediction candidates added to the merged candidate list. Figure 17 This is a diagram illustrating an example of the merged candidates as derived objects of chained motion vector prediction candidates, and the position (index number) of the chained motion vector prediction candidates added to the merged candidate list. Figure 18 This is a diagram illustrating an example of a pruning process in which the inter-frame prediction unit 205 of an image decoding apparatus 200 according to one embodiment adds merging candidates corresponding to chain motion vectors derived by chain motion vector prediction to the merging candidate list. Figure 19 This is a diagram illustrating an example of a method for controlling the generation of a reference source when generating a chain motion vector in the chain motion vector prediction of an inter-frame prediction unit 205 of an image decoding apparatus 200 according to one embodiment. Figure 20 This is a diagram illustrating an example of a method for controlling the generation of a reference destination when generating a chain motion vector in the chain motion vector prediction of the inter-frame prediction unit 205 of an image decoding apparatus 200 according to one embodiment. Figure 21 This diagram illustrates an example of a method for controlling the reference source and reference destination when exporting chain motion vectors in the chain motion vector prediction of the inter-frame prediction unit 205 of an image decoding apparatus 200 according to one embodiment, in the case of applying GDR. Figure 22 This is a diagram illustrating an example of a method for controlling the reference source when deriving chain motion vectors in chain motion vector prediction. Figure 23 This is a diagram illustrating an example of a decoding unit 201 of an image decoding apparatus 200 according to an embodiment, which determines whether there is a decoding process for sps_chained_mvp_enabled_flag based on sps_temporal_mvp_enabled_flag. Figure 24This is a diagram illustrating an example of a control point motion vector cpMv derived by the inter-frame prediction unit 205 of an image decoding apparatus 200 according to one embodiment. Detailed Implementation

[0022] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Furthermore, the constituent elements in the following embodiments can be appropriately replaced with existing constituent elements, and various modifications are possible, including combinations with other existing constituent elements. Therefore, the description of the following embodiments does not limit the scope of the invention as described in the claims.

[0023] <First Implementation Method> The following is for reference Figures 1 to 24 The image decoding apparatus 200 according to this embodiment will be described.

[0024] The image decoding apparatus 200 according to the first embodiment of the present invention targets various image signals (hereinafter referred to as images). For example, the image decoding apparatus 200 according to this embodiment targets YUV (YCbCr) images including luminance pixels and chrominance pixels, RGB images including RGB pixels, and monochrome images. Here, the pixels constituting each image have discrete values ​​(pixel values) with a predetermined bit width.

[0025] Figure 1 This is a diagram illustrating an example of the functional blocks of the image decoding apparatus 200 according to this embodiment.

[0026] like Figure 1 As shown, the image decoding device 200 includes a code input unit 210, a decoding unit 201, an inverse quantization unit 202, an inverse transform unit 203, an intra-frame prediction unit 204, an inter-frame prediction unit 205, an adder 206, a decoded image buffer 207, and an image output unit 220.

[0027] In addition, the parts described as "pixels" in the functional descriptions of each part below can be blocks (units) that include pixels, tree blocks that are the largest size of blocks, slices larger than tree blocks, bricks, or images (pictures).

[0028] In addition, although in Figure 1 Although not shown in the diagram, the image decoding device 200 may include a loop filter between the adder 206 and the decoded image buffer 207 to correct the pixel values ​​of the decoded pixels.

[0029] The code input unit 210 is configured to acquire code information encoded by the image encoding device.

[0030] The decoding unit 201 is configured to decode control information and quantization values ​​from code information input from the code input unit 210. For example, the decoding unit 201 is configured to output control information and quantization values ​​by performing variable-length decoding on the code information.

[0031] Here, the quantization value is sent to the inverse quantization unit 202, and control information is sent to the decoded image buffer 207, the intra-frame prediction unit 204, and the inter-frame prediction unit 205. Furthermore, this control information (syntax) may include information required for controlling the decoded image buffer 207, the intra-frame prediction unit 204, and the inter-frame prediction unit 205, including header information such as sequence parameter sets, image parameter sets, image headers, and slice headers.

[0032] The inverse quantization unit 202 is configured to inverse quantize the quantized value sent from the decoding unit 201 and use it as the transform coefficients to be decoded. These transform coefficients are then sent to the inverse transform unit 203.

[0033] The inverse transform unit 203 is configured to perform an inverse transform on the transform coefficients sent from the inverse quantization unit 202 and use them as the decoded prediction residual. This prediction residual is sent to the adder 206.

[0034] The intra-prediction unit 204 is configured to generate intra-predicted pixels based on the decoded pixels and control information sent from the decoding unit 201. Here, the decoded pixels are pixels obtained via the adder 206 and stored in the decoded image buffer 207. Furthermore, the intra-predicted pixels are predicted pixels used to add to the prediction residual in the adder 206. Additionally, the intra-predicted pixels are sent to the adder 206.

[0035] Here, the prediction performed by the intra-prediction unit 204 may be the intra-prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2. Alternatively, the prediction performed by the intra-prediction unit 204 may be the intra-block copying or intra-template matching prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0036] The decoded image buffer 207 cumulatively stores inter-frame prediction information corresponding to the decoded pixels sent from the adder 206 and the decoded pixels sent from the inter-frame prediction unit 205.

[0037] Here, the inter-frame prediction unit 205 references the decoded pixels and the corresponding inter-frame prediction information stored in the decoded image buffer 207. Details regarding the inter-frame prediction information will be described later.

[0038] The inter-frame prediction unit 205 is configured to generate inter-frame prediction pixels based on the decoded pixels obtained from the reference decoded image buffer 207, inter-frame prediction information, and control information sent from the decoding unit 201.

[0039] Then, the inter-frame prediction unit 205 sends the generated inter-frame prediction pixel to the addition unit 206, and sends the inter-frame prediction information used to generate the inter-frame prediction pixel as inter-frame prediction information corresponding to the decoding target pixel to the decoding image buffer 207.

[0040] Here, the prediction performed by the inter-frame prediction unit 205 may be the inter-frame prediction disclosed in Non-Patent Document 1 and Non-Patent Document 2. Alternatively, the prediction performed by the inter-frame prediction unit 205 may be the intra-frame block copying disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0041] The adder 206 is configured to add the prediction residual sent from the inverse transform unit 203 to at least one of the input intra-frame prediction pixels and inter-frame prediction pixels to calculate the decoded pixel. The decoded pixel is sent to the image output unit 220, the decoded image buffer 207, and the intra-frame prediction unit 204.

[0042] The image output unit 220 is configured to output the decoded pixels sent from the adder 206.

[0043] Hereinafter, taking the prediction performed by the inter-frame prediction unit 205 as an example, the function of the inter-frame prediction unit 205 and the decoded image buffer 207, which are characteristic configurations of the image decoding apparatus 200 according to this embodiment, will be explained.

[0044] <Basic functions of inter-frame prediction unit 205 and decoded image buffer 207> The basic function of the inter-frame prediction unit 205 is to derive one or more motion vectors (Mv) for the decoded target block in order to predict the decoded target block with high accuracy in the subsequent adder 206, and to predict the pixels (hereinafter referred to as reference pixels) of the decoded target block (hereinafter referred to as reference block) referenced by Mv in a decoded image (hereinafter referred to as reference image) that is different from the decoded target image to which the decoded target block belongs. That is, to perform inter-frame prediction.

[0045] exist Figure 2 The diagram below illustrates the concept of inter-frame prediction. Specifically, Figure 2An example is shown below: In inter-frame prediction, in order to generate inter-frame prediction pixels of the decoded target block (CurBlk) located within the decoded target image (CurPic), reference blocks (RerBlkL0k and RerBlkL1k) in two different reference images (RerPicL0k and RerPicL1k) are referenced based on the two motion vectors MvL0k and MvL1k derived by the inter-frame prediction unit 205.

[0046] Here, L0 and L1 indicate the numbers of a list of reference images that can be referenced from the decoded target image (reference image list), and k corresponds to the value of the control information (index) used to determine the motion vectors described in the method for deriving motion vectors described later.

[0047] As in Non-Patent Document 1 and Non-Patent Document 2, the image decoding apparatus 200 of this embodiment can export list 0 (L0) and list 1 (L1) as two different reference image lists according to the target image to be decoded.

[0048] Here, the decoded image buffer 207 can export the target image according to the reference image list based on the control information decoded by the decoding unit 201.

[0049] For example, control information may include the number of reference images that should be included in the list of reference images, and candidate reference images (image reference structure) for the target image to be decoded.

[0050] The detailed method for deriving motion vectors in the inter-frame prediction unit 205 will be described later. An overview of the method for generating inter-frame prediction pixels in the inter-frame prediction unit 205 is as follows.

[0051] The inter-frame prediction unit 205 can generate inter-frame prediction pixels from a single reference block, as shown in the example. Figure 2 The diagram shows the generation of inter-frame predicted pixels using two reference blocks, but more than two reference blocks can also be used to generate inter-frame predicted pixels.

[0052] Here, the inter-frame prediction unit 205 can apply the following method as a way to generate inter-frame prediction pixels using two reference blocks.

[0053] As an example, the inter-frame prediction unit 205 can simply average the pixel values ​​of two reference blocks to generate inter-frame prediction pixels.

[0054] As another example, the inter-frame prediction unit 205 can use predetermined weight values ​​to generate inter-frame predicted pixels by weighting the pixel values ​​of two reference blocks on a block-by-block basis.

[0055] The inter-frame prediction unit 205 can be applied to the block-weighted bi-directional prediction (BCW: Bi-prediction with CU-level Weights) disclosed in Non-Patent Document 1 and Non-Patent Document 2 as a block-weighted averaging technique.

[0056] Furthermore, in non-patent documents 1 and 2, in BCW, it is possible to select multiple weight values, including a simple average weight value of 1:1, which is determined by the value of the internal parameter (bcwIdx, the third internal parameter).

[0057] Here, the inter-frame prediction unit 205 derives bcwIdx through control information or inheritance from the decoded image buffer 207. Details of inheritance will be described later.

[0058] As another example, the inter-frame prediction unit 205 can generate inter-frame prediction pixels by dividing the decoded target block into two parts on a predetermined straight line and using a weighted average of the pixel values ​​of the reference blocks corresponding to each of the divided regions. The aforementioned weight value corresponds to the distance from the dividing straight line.

[0059] The inter-frame prediction unit 205 can be applied to the geometric partitioning mode (GPM) disclosed in Non-Patent Document 1 and Non-Patent Document 2 as a weighted average technique for each region that linearly divides the block into two regions using a straight line.

[0060] Here, in the case of a fractional pixel precision position within the motion vector reference reference block, the inter-frame prediction unit 205 can use an interpolation filter to generate inter-frame prediction pixels. The aforementioned interpolation filter is used to derive the reference pixel value at the fractional pixel precision position from the reference pixel value at the integer pixel precision position in the up, down, left, and right directions located at the fractional pixel position.

[0061] Furthermore, when generating inter-frame prediction pixels, the inter-frame prediction unit 205 can adaptively apply two interpolation filters disclosed in Non-Patent Document 1 and Non-Patent Document 2, which have different cutoff frequencies for inter-frame prediction.

[0062] The interpolation filter with the lower cutoff frequency among the two interpolation filters is called the Switchable Interpolation Filter (SIF).

[0063] Similar to Non-Patent Document 1 and Non-Patent Document 2, when the motion vector references a half-pixel precision position and the internal parameter (hpelIfIdx, first internal parameter) indicating whether to apply SIF indicates the application of SIF, the inter-frame prediction unit 205 can apply SIF.

[0064] Furthermore, the inter-frame prediction unit 205 derives hpelIfIdx through control information or inheritance from the decoded image buffer 207. Details regarding inheritance will be described later.

[0065] On the other hand, when the motion vector references a position with fractional pixel precision, the inter-frame prediction unit 205 can round (trim) the reference destination of the motion vector to an integer pixel precision position that is closest to that fractional pixel precision position to generate inter-frame prediction pixels without applying an interpolation filter.

[0066] In addition, the inter-frame prediction unit 205 may apply the local illumination compensation (LIC) disclosed in Non-Patent Document 2 when generating the final inter-frame prediction pixel.

[0067] LIC is a technique that uses linear prediction to correct the values ​​of reference pixels in the reference block based on the decoded pixels that are adjacent to the target block and the reference block, respectively.

[0068] Similar to Non-Patent Document 1 and Non-Patent Document 2, when the internal parameter (licFlag, second internal parameter) controlling whether to apply LIC indicates the application of LIC, the inter-frame prediction unit 205 can apply LIC.

[0069] Furthermore, the inter-frame prediction unit 205 derives the licFlag through control information or inheritance from the decoded image buffer 207. Details regarding inheritance will be described later.

[0070] In addition, the inter-frame prediction unit 205 may apply the multi-hypothesis prediction (MHP) disclosed in Non-Patent Document 2 when generating the final inter-frame prediction pixels.

[0071] MHP is a technique that generates the final inter-frame predicted pixels by weighted averaging of reference pixel values ​​from at least three reference blocks.

[0072] Similar to Non-Patent Document 2, when the internal parameter (mhpFlag, fourth internal parameter) indicating whether to apply MHP indicates the application of MHP, the inter-frame prediction unit 205 may apply MHP.

[0073] Furthermore, the inter-frame prediction unit 205 derives mhpFlag through control information or inheritance from the decoded image buffer 207. Details regarding inheritance will be described later.

[0074] The inter-frame prediction unit 205 sends inter-frame prediction pixels for the decoded target block to the adder 206, and sends to the decoded image buffer 207 a motion vector for generating the inter-frame prediction pixels, an index (reference index) for determining the reference image reference within the reference image list, and a counter (POC: Picture Order Count) indicating the image output order of the reference images corresponding to the reference index. Hereinafter, the motion vector, reference index, and POC corresponding to the reference index used to generate these inter-frame prediction pixels will be referred to as inter-frame prediction information. Furthermore, the inter-frame prediction unit 205 may include at least one or all of bcwIdx, hpelIfIdx, licFlag, or mhpFlag in the inter-frame prediction information and send it, where bcwIdx, hpelIfIdx, licFlag, or mhpFlag are internal parameters that modify the aforementioned method for generating inter-frame prediction pixels.

[0075] In addition to the decoded pixels of the decoded target block sent from the adder 206, the decoded image buffer 207 also stores the inter-frame prediction information of the decoded target block corresponding to the decoded pixels.

[0076] The decoding image buffer 207 can store the inter-frame prediction information in a predetermined size (number of pixels). For example, the decoding image buffer 207 can store the inter-frame prediction information in units of the minimum size (4×4 pixels) of the decoding target block as disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0077] The decoded image buffer 207 can improve the inter-frame prediction accuracy when referencing the decoded target block to be decoded in the future by storing the inter-frame prediction information at the minimum size of the decoded target block.

[0078] The decoded image buffer 207 stores decoded pixels on a per-target-image basis. Alternatively, the decoded image buffer 207 may store decoded pixels on a per-slice-of-target-image basis. The following explanation will use the case of storing decoded pixels on a per-target-image basis as an example.

[0079] If it is determined that the stored decoded pixels, per unit of the decoded target image, and the inter-frame prediction information corresponding to those decoded pixels, are not referenced from future decoded target images or blocks as reference images or blocks, the decoded image buffer 207 can sequentially delete the decoded pixels, per unit of the decoded target image, and the inter-frame prediction information corresponding to those decoded pixels. According to this configuration, the amount of information stored in the decoded target image buffer 207 can be reduced.

[0080] The decoded image buffer 207 can derive the number of decoded target images and the reference image structure between the decoded target images based on the control information decoded by the decoding unit 201.

[0081] As a variation, if the inter-frame predicted pixel exists in a predetermined area outside the decoded target image in addition to the decoded target image, the decoded image buffer 207 may additionally store the inter-frame predicted pixel.

[0082] Figure 3 It is a diagram used to illustrate the reference pixel values ​​disclosed in each of Non-Patent Document 1 and Non-Patent Document 2, when the reference is outside the decoded target image stored in the decoded image buffer 207 from the future decoded target block.

[0083] like Figure 3 As shown in Non-Patent Document 1 (VVC: Versatile Video Coding), when referencing a future decoded target block within a predetermined range outside the decoded target image, the pixel value located at the nearest decoded target image boundary is copied to generate a reference pixel value.

[0084] In VVC, the range in which 16 pixels are added to the maximum size (number of pixels) of the width of the decoded target block is specified as the specified range (fill area).

[0085] In non-patent literature 2 (ECM: Enhanced Compression Model), such as Figure 3 As shown, only a predetermined range is stored between the outside of the decoded target image and the filling area, which serves as the region (inter-frame prediction filling area) from which the inter-frame predicted pixel values ​​of the decoded target block can be referenced from future decoded target blocks.

[0086] Therefore, even if the target image is not referenced from the future target block, the inter-frame prediction pixel values ​​of the target block still exist, thus improving the inter-frame prediction accuracy of the future target block.

[0087] Figure 4 It is generated in Figure 3 The conceptual diagram of the inter-frame prediction padding region is explained in the text.

[0088] like Figure 4 As shown, when the adjacent blocks of the reference block corresponding to the decoding target block have inter-frame prediction pixels, even if the decoding target block is adjacent to the decoding target image, inter-frame prediction pixels can be generated outside the decoding target image. Therefore, the decoding image buffer 207 can store... Figure 3 The inter-frame prediction padding region is shown.

[0089] <Methods for deriving motion vectors in the inter-frame prediction unit> The following is for reference Figures 5 to 17 The method for deriving motion vectors in the inter-frame prediction unit 205 of the image decoding apparatus 200 according to this embodiment will be described.

[0090] Figure 5 This is a diagram showing an example of a merged candidate list as a method for deriving motion vectors in the inter-frame prediction unit 205.

[0091] The merge candidate list is a list of candidates (i.e. merge candidates) disclosed in Non-Patent Document 1 and Non-Patent Document 2 that contains inter-frame prediction information constructed by the inter-frame prediction unit 205 for deriving motion vectors in merge (merge mode).

[0092] The inter-frame prediction unit 205 extracts merge candidates from a plurality of merge candidates located in the merge candidate list based on the index (merge_idx) sent from the decoding unit 201 for determining the merge candidates in the merge candidate list, and uses them as inter-frame prediction information for generating inter-frame prediction pixels.

[0093] The upper limit of the number of merging candidates stored in the merging candidate list can be a fixed value or a variable value. In the case of a variable value, for example, as in Non-Patent Document 1 and Non-Patent Document 2, control information can be set in the sequence parameter set, etc., and the variable value can be configured based on the control information.

[0094] The inter-frame prediction unit 205 determines the inter-frame prediction information for generating inter-frame prediction pixels from the merge candidate list based on one or more merge_idx values ​​sent from the decoding unit 201.

[0095] For example, for the aforementioned GPM, the inter-frame prediction unit 205 can decode merge_idx one by one for each region that divides the decoded target block into two by a straight line, so as to derive inter-frame prediction information.

[0096] like Figure 5 As shown, the inter-frame prediction unit 205 stores the inter-frame prediction information (merging candidates) corresponding to L0 and L1 in the merging candidate list.

[0097] In addition, Figure 5 In the example, only motion vectors and reference indices are recorded. However, in addition to this, the inter-frame prediction unit 205 may also store at least one or more of bcwIdx, hpelIfIdx, licFlag or mhpFlag, which are internal parameters of the method for changing the generation of the aforementioned inter-frame prediction pixels, as inter-frame prediction information.

[0098] The inter-frame prediction unit 205 can search for inter-frame prediction information, that is, inter-frame prediction information that is spatially or temporally adjacent to or close to the target block being decoded, and add it to the merging candidate list as a merging candidate.

[0099] As an example, the inter-frame prediction unit 205 can combine and apply the search of multiple merged candidates disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0100] Figure 6 The names of several merger candidates in regular merges disclosed in Non-Patent Document 1 and Non-Patent Document 2 are shown, along with the order in which these merger candidates are derived.

[0101] Specifically, in Non-Patent Document 1 (VVC), the inter-frame prediction unit 205 confirms whether or not these merging candidates have been used, in the order of Spatialmerge candidates, Temporal Motion Vector Prediction (TMVP) candidates, History-based Motion Vector Prediction (HMVP) candidates, and Pairwise merge candidates.

[0102] In Non-Patent Document 2 (ECM), in addition to the aforementioned merge candidates in Non-Patent Document 1, the inter-frame prediction unit 205 also confirms whether there are non-adjacent (NA) spatial merge candidates among the spatial merge candidates and TMVP candidates.

[0103] The following is a summary of each merger candidate.

[0104] Figure 7 This is a diagram showing the search locations of spatial merge candidates and TMVP candidates in the decoded target block.

[0105] like Figure 7As shown, in the spatial merging candidate, inter-frame prediction information is searched at the positions of the lower left (A0), left (A1), upper right (B0), upper (B1), and upper left (B2) adjacent to the decoded target block (CurBlk).

[0106] That is, in the spatial merging candidate, inter-frame prediction information from the spatially adjacent positions of the decoded target block (CurBlk) is used.

[0107] Additionally, in the TMVP candidates, a search is conducted in images different from the decoded target block (see reference image). Figure 7 Inter-frame prediction information for each position in the lower right center (Col) and the lower right center (H) outside the decoded target block (CurBlk).

[0108] That is, in the TMVP candidate, inter-frame prediction information from positions that are temporally adjacent to the target block being decoded is used.

[0109] In HMVP candidates and non-adjacent spatial merging candidates, inter-frame prediction information from locations spatially far (near) the decoded target block is used.

[0110] Specifically, in the HMVP candidate, a table is constructed that can store FIFO-type inter-frame prediction information. Each time new inter-frame prediction information is searched, the oldest inter-frame prediction information is discarded from the table, and the new inter-frame prediction information is added to the table.

[0111] Similar to the HMVP candidates and non-adjacent spatial merging candidates disclosed in Non-Patent Document 1 and Non-Patent Document 2, the inter-frame prediction unit 205 of the image decoding apparatus 200 according to this embodiment can design the location for searching inter-frame prediction information and the size of the table for HMVP candidates.

[0112] Furthermore, the inter-frame prediction unit 205 can re-search for new reference block positions from the surrounding pixels of the reference block of the inter-frame prediction information of the search target, and correct the inter-frame prediction information (motion vector) so that in the search for merging candidates in the aforementioned conventional merging, the difference (hereinafter referred to as template matching cost) between the decoded pixels (template of the reference block) adjacent to the reference block based on the inter-frame prediction information of the search target and the decoded pixels (template of the decoded target block) adjacent to the decoded target block becomes smaller.

[0113] As an example, the inter-frame prediction unit 205 can be applied to template matching and merging disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0114] Furthermore, in the selection of merge candidates from the merge candidate list in the aforementioned conventional merging, the inter-frame prediction unit 205 can derive inter-frame prediction information of L0 from the control information and select the merge candidate with the smallest difference between the template of the reference block based on the inter-frame prediction information stored in the merge candidate list L1 and the template of the reference block based on the inter-frame prediction information of L0 derived from the control information, as the merge candidate of L1.

[0115] Since the inter-frame prediction unit 205 can derive the merging candidate (i.e., inter-frame prediction information) of L1 based on the inter-frame prediction information of L0 derived from the control information, it is not necessary to decode merge_idx in order to derive the inter-frame prediction information of L1, except for the control information required to derive the inter-frame prediction information of L0.

[0116] As an example, the inter-frame prediction unit 205 can be applied to the adaptive motion vector prediction-merge mode disclosed in Non-Patent Document 1 and Non-Patent Document 2.

[0117] As a variation, the inter-frame prediction unit 205 can reverse L0 and L1 in the aforementioned example. That is, the inter-frame prediction unit 205 can derive inter-frame prediction information for L1 using control information and derive inter-frame prediction information for L0 from the merging candidate list.

[0118] Furthermore, the inter-frame prediction unit 205 can construct a merge candidate list for merging, which derives inter-frame prediction information in units of sub-blocks after the decoding target block is divided, in order to derive inter-frame prediction information.

[0119] As an example, the inter-frame prediction unit 205 can construct a merge candidate list for sub-block TMVP (SbTMVP) or affine merge disclosed in Non-Patent Document 1 and Non-Patent Document 2 to derive inter-frame prediction information.

[0120] <1. Basic Concepts of Chain Motion Vector Prediction (Derivation of Chain Motion Vectors)> The aforementioned methods for deriving inter-frame prediction information (hereinafter, for simplicity, referred to as motion vectors) through the inter-frame prediction unit are all limited to methods that store the information at a location that is spatially or temporally adjacent to or close to the target block being decoded.

[0121] To further improve coding efficiency, the inter-frame prediction unit 205 can derive the following motion vector as the motion vector of the decoding target block: that is, the motion vector is generated by recursively searching for motion vectors stored in the decoding image buffer 207 for positions that are adjacent or close to the decoding target block in space or time, and adding the derived motion vectors.

[0122] Here, the technique of recursively (chainedly) searching the motion vectors stored in the decoded image buffer 207 and generating new motion vectors (chained motion vectors) starting from motion vectors located spatially or temporally adjacent to or close to the decoded target block is referred to in this specification as "Chained Motion Vector Prediction (CMVP)". The technical content of this chained motion vector prediction is described in detail below.

[0123] use Figures 8 to 17 This section explains the method for deriving chain motion vectors in chain motion vector prediction.

[0124] Figure 8 This is a diagram illustrating an example of a method for deriving chain motion vectors in chain motion vector prediction.

[0125] like Figure 8 As shown, the inter-frame prediction unit 205 recursively (chain-likely) searches for motion vectors stored in the decoded image buffer 207, starting with the motion vector MvL0k(0) corresponding to any merge candidate stored in the merge candidate list and the reference index, and generates new motion vectors.

[0126] For example, as shown in Equation 1 and Equation 2, for the reference image RefPicL0 k The inter-frame prediction unit 205 can derive new motion vectors and reference images generated in chained motion vector prediction from the reference block RefBlk (m).

[0127] MvL0 k / m =MvL0 k (0)+MvL0 k (1)+MvL0 k (2)+…+MvL0 k (m) (Equation 1) RefPicL0 k / m =RefPicL0 k (m) (Equation 2) That is, the motion vectors stored in the decoded image buffer 207 can be recursively (chained) searched and all derived motion vectors can be summed to calculate the new motion vector MvL0 generated in the chained motion vector prediction. k / m .

[0128] Additionally, the new reference image generated in the chain motion vector prediction is MvL0. k / m Reference image used for reference.

[0129] Figure 9This is a diagram illustrating an example of a method for deriving chain motion vectors in chain motion vector prediction.

[0130] like Figure 9 As shown, in chained motion vector prediction, the inter-frame prediction unit 205 can add block vectors (vectors indicating reference destinations within the same image) to the recursive search target in addition to searching for motion vectors.

[0131] For example, when adding a block vector (Bv), for the reference image RefPicL0 k The reference block RefBlk of (m), as shown in Equation 3, adds Bv to the derivation of the new motion vector generated in the chained motion vector prediction.

[0132] MvL0 k / m =MvL0 k (0)+Bv k (0)+MvL0 k (1)+MvL0 k (2)+…+MvL0 k (m) (Equation 3) Figure 10 This is a diagram illustrating an example of a recursive search for motion vectors that includes bidirectional prediction (i.e., inter-frame prediction using two motion vectors different from L0 and L1). Figure 11 This is a diagram illustrating an example of a chain of motion vectors derived through a recursive search that includes bidirectional prediction of motion vectors.

[0133] like Figure 10 as well as Figure 11 As shown, in chained motion vector prediction, the inter-frame prediction unit 205 uses the search list numbers (L0 or L1) and search depths (or search counts) of the motion vectors stored in the decoded image buffer 207 to search for one or more new motion vectors from a reference source. Figure 10 as well as Figure 11 In the example, at most two motion vectors are derived, and all the derived motion vectors are added to each of the newly derived motion vectors to generate a chained motion vector MvL0. k / m .

[0134] Figure 12 as well as Figure 13 This is a diagram illustrating an example, as disclosed in Non-Patent Document 1 and Non-Patent Document 2, of searching for chain motion vector prediction candidates after searching for multiple merger candidates in a conventional merging process.

[0135] like Figure 12 as well as Figure 13As shown, the inter-frame prediction unit 205 can search for chain motion vector prediction candidates after searching for predetermined merging candidates.

[0136] For example, such as Figure 12 As shown, in VVC, the inter-frame prediction unit 205 can search for chained motion vector prediction candidates after searching for TMVP candidates.

[0137] Chained motion vector prediction is invalid if the TMVP is invalid in the property of using motion vectors from a reference image instead of the decoded target image stored in the decoded image buffer 207 to search for new motion vectors, or if the TMVP is invalid in the decoded target sequence or decoded target image group or decoded target image or decoded target slice.

[0138] Therefore, searching for chained motion vector prediction candidates after searching for TMVP candidates can be seen as a natural design approach as a motion vector derivation method.

[0139] As Figure 12 Examples of changes, such as Figure 13 As shown, in VVC and ECM, the inter-frame prediction unit 205 can search for chained motion vector prediction candidates after searching for HMVP candidates.

[0140] In terms of the property of generating new motion vectors by searching recursively, chained motion vector prediction must be applied when at least one merging candidate is stored in the merging candidate list.

[0141] like Figure 13 As shown, by searching for chained motion vector prediction candidates after searching for HMVP candidates, the probability of generating new motion vectors through chained motion vectors is increased. The aforementioned HMVP candidates can be expected to save more than one motion vector candidate for the merged candidate list.

[0142] As other examples of changes, such as Figure 12 As shown, in ECM, the inter-frame prediction unit 205 can search for chain motion vector prediction candidates after merging candidates in the search space.

[0143] Although Figure 7 The search positions for spatial merging candidates are shown in the figure. However, for example, if the search positions for spatial merging candidates are increased, the probability of keeping more than one motion vector candidate in the merging candidate list increases, and it is easier to generate new motion vector candidates through chained merging of candidates.

[0144] Figures 14 to 17 This is a diagram illustrating an example of the merged candidates as derived objects of chained motion vector prediction candidates, and the position (index number) of the chained motion vector prediction candidates added to the merged candidate list.

[0145] In addition, Figures 14 to 17 In order to simplify the explanation of the technical content, only motion vectors are recorded as inter-frame prediction information stored in the merge candidate list, but the aforementioned reference index is also stored.

[0146] In addition to motion vectors and reference indices, the inter-frame prediction unit 205 may also store bcwIdx, hpIfIdx, licFlag, and mhpFlag, which are internal parameters for changing the aforementioned inter-frame prediction pixel generation method.

[0147] like Figure 14 As shown, the inter-frame prediction unit 205 can sequentially apply chained motion vector prediction starting from the first merging candidate stored in the merging candidate list after storing the predetermined merging candidates in the merging candidate list, and sequentially store all new merging candidates generated in the chained motion vector prediction for each merging candidate in the merging candidate list.

[0148] Here, the pre-defined merge candidate can be used Figure 12 as well as Figure 13 The description refers to any of the multiple merge candidates in a regular merge, or it can be a predetermined number (the predetermined number of merge_idx) of merge candidates.

[0149] The predetermined number can be configured as a fixed value, or it can be configured as a variable value using control information in units of decoding target sequences, decoding target image groups, decoding target images, or decoding target slices.

[0150] As Figure 14 Examples of changes, such as Figure 15 As shown, the inter-frame prediction unit 205 can, after saving the predetermined merging candidates in the merging candidate list, sequentially apply chain motion vector prediction starting from the first merging candidate saved in the merging candidate list, and sequentially save the new chain motion vectors (i.e., merging candidates) generated at each search depth for the chain motion vector prediction of each merging candidate in the merging candidate list.

[0151] As Figure 14 as well as Figure 15 Examples of changes, such as Figure 16 As shown, after the initial merging candidate is saved in the merging candidate list, the inter-frame prediction unit 205 can sequentially apply chain motion vector prediction from the merging candidate and sequentially save all new chain motion vectors (i.e., merging candidates) generated in the chain motion vector prediction for each merging candidate in the merging candidate list.

[0152] As Figures 14 to 16 Examples of changes, such as Figure 17 As shown, after the initial merging candidate is saved in the merging candidate list, the inter-frame prediction unit 205 can sequentially apply chain motion vector prediction from the merging candidate, and sequentially save the new chain motion vectors (i.e., merging candidates) generated at each search depth for the chain motion vector prediction of each merging candidate in the merging candidate list.

[0153] As Figures 14 to 17 In a modified example, the inter-frame prediction unit 205 can combine the chain motion vectors generated at each search depth for each chain motion vector prediction of each merge candidate list, and the chain motion vectors generated at search depths different from each search depth, to generate new merge candidates.

[0154] As Figures 14 to 17 In a modified example, the inter-frame prediction unit 205 can combine chain motion vectors generated from each search list number or each search depth predicted for each merge candidate list, and chain motion vectors generated using search list numbers or search depths different from those search list numbers or search depths, to generate new merge candidates.

[0155] <2. Pruning of Candidates for Chain Motion Vector Prediction> use Figure 18 This describes the pruning process performed by the inter-frame prediction unit 205 of the image decoding apparatus 200 in this embodiment when it adds a merging candidate (the aforementioned merging motion vector prediction candidate) corresponding to the merging candidate derived by the merging motion vector prediction to the merging candidate list.

[0156] Figure 18 This is a diagram illustrating an example of a pruning process in which the inter-frame prediction unit 205 of the image decoding apparatus 200 according to this embodiment adds merging candidates corresponding to the chain motion vectors derived by chain motion vector prediction to the merging candidate list.

[0157] The inter-frame prediction unit 205 determines whether to add a merging candidate corresponding to the chain motion vectors derived from the chain motion vector prediction based on predetermined conditions.

[0158] Specifically, such as Figure 18 As shown, when the inter-frame prediction unit 205 determines in step S000 that the predetermined conditions are met, in step S002, it determines that no merge candidate corresponding to the chain motion vector derived by the chain motion vector prediction is added.

[0159] On the other hand, when the inter-frame prediction unit 205 determines in step S000 that the predetermined conditions are not met, in step S001, it determines to add a merge candidate corresponding to the chain motion vector derived by the chain motion vector prediction.

[0160] Here, for example, the predetermined conditions may be consistent with the reference image and motion vector associated with each of at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector.

[0161] When the inter-frame prediction unit 205 determines that the reference image and motion vector are consistent with each of the at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector, it determines that the merge candidate corresponding to the chain motion vector will not be added to the merge candidate list, that is, the merge candidate corresponding to the chain motion vector will be removed (as a pruning target). This prevents more than two identical motion vectors from being stored in the merge candidate list, reduces the amount of code sent to the merge_idx (merge index) of the image decoding device 200, and improves coding efficiency.

[0162] Alternatively, the predetermined condition may be that it is consistent with at least one merge candidate stored in the merge candidate list and with one or more reference images and motion vectors associated with each of the merge candidates corresponding to the derived chain of motion vectors.

[0163] Alternatively, the predetermined conditions may be that the reference image and motion vector are all identical to those associated with each of at least one merge candidate in the merge candidate list and each of the merge candidate corresponding to the derived chain motion vector in the merge candidate list 0 (L0) and merge candidate list 1 (L1).

[0164] Furthermore, the predetermined condition may be that the reference image associated with each of at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector is consistent with the reference image associated with each of the merge candidates associated with the derived chain motion vector, and the error (Euclidean distance) of the reference position of the merge candidate associated with each of the merge candidates stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector is less than a predetermined threshold.

[0165] Therefore, by strengthening the pruning process of the merge candidates corresponding to the chained motion vectors, the situation of storing more than two identical motion vectors in the merge candidate list is further prevented, the amount of code sent to merge_idx in the image decoding device 200 is reduced, and the encoding efficiency is improved.

[0166] As a further control, the inter-frame prediction unit 205 can configure a predetermined threshold for the case where there is only one motion vector being compared to a predetermined threshold that is greater than the predetermined threshold for the case where there are two motion vectors being compared.

[0167] Thus, by configuring the predetermined threshold for the case where there is only one motion vector to be compared to a predetermined threshold for the case where there are two motion vectors to be compared, more precise pruning processing can be performed.

[0168] Here, the inter-frame prediction unit 205 can configure a predetermined threshold of 1.5 pixels when the motion vector being compared is one.

[0169] Alternatively, the inter-frame prediction unit 205 can configure a predetermined threshold of 1.0 pixels or 2.0 pixels when the motion vector being compared is one.

[0170] In addition, the inter-frame prediction unit 205 can configure a predetermined threshold of 1.0 pixels when there are two motion vectors being compared.

[0171] Alternatively, the inter-frame prediction unit 205 can configure a predetermined threshold of 0.5 pixels or 1.5 pixels when there are two motion vectors being compared.

[0172] In addition, the aforementioned predetermined conditions may be consistent with the internal parameters of the generation method for each associated reference image and motion vector in at least one merge candidate stored in the merge candidate list and the merge candidate corresponding to the derived chain motion vector, as well as the generation method for changing the predetermined inter-frame prediction pixels.

[0173] Alternatively, the predetermined conditions may be consistent with at least one merge candidate stored in the merge candidate list and the associated reference image and motion vector of each merge candidate corresponding to the derived chain of motion vectors, as well as the internal parameter (hpelIfIdx) for determining whether to apply the switching interpolation filter.

[0174] Alternatively, the predetermined conditions may be consistent with at least one merge candidate stored in the merge candidate list and the associated reference image and motion vector of each merge candidate corresponding to the derived chain of motion vectors, as well as the internal parameter (licFlag) controlling whether local brightness compensation is applied.

[0175] Furthermore, the predetermined conditions can be consistent with at least one merge candidate stored in the merge candidate list and the associated reference image and motion vector of each merge candidate corresponding to the derived chain motion vector, as well as the internal parameters (bcwIdx) for determining the weight values ​​of the block-based weighted bidirectional prediction.

[0176] By including the aforementioned internal parameters in the predetermined conditions, even when the reference image being compared and the motion vector are the same, the inter-frame prediction unit 205 can generate inter-frame prediction pixels using different methods when the internal parameters are different, leaving room for improving the accuracy of inter-frame prediction, and thus an improvement in coding efficiency can be expected.

[0177] <3. Control of reference source and reference destination when deriving chained motion vectors> use Figures 19 to 22 The control method for the reference source and reference destination when the inter-frame prediction unit 205 of the image decoding apparatus 200 according to this embodiment derives the chain motion vector in the chain motion vector prediction will be described.

[0178] First, use Figure 19 The method for controlling the reference source when deriving chain motion vectors in chain motion vector prediction is explained.

[0179] Figure 19 This is a diagram illustrating an example of a method for controlling the deriving of a reference source when deriving a chain motion vector in the chain motion vector prediction by the inter-frame prediction unit 205.

[0180] When deriving chained motion vectors in chained motion vector prediction, the inter-frame prediction unit 205 can configure the starting point of the first motion vector constituting the chained motion vector in part or all of the decoded target block.

[0181] As an example of something configured as part of the decoding target block, such as Figure 19 As shown, when deriving chained motion vectors in chained motion vector prediction, the inter-frame prediction unit 205 can configure the starting point of the first motion vector constituting the chained motion vector at the center (C) of the decoded target block (CurBlk).

[0182] Or, such as Figure 19 As shown, when deriving chained motion vectors in chained motion vector prediction, in addition to being positioned at the center (C) of the decoded target block (CurBlk), the inter-frame prediction unit 205 can also position the starting point of the first motion vector constituting the chained motion vector at a predetermined position in the decoded target block (CurBlk), and search in a predetermined order whether the motion vector is stored in the reference destination of the motion vector.

[0183] Specifically, such as Figure 19 As shown, when deriving chained motion vectors in chained motion vector prediction, in addition to being positioned at the center (C) of the decoded target block (CurBlk), the inter-frame prediction unit 205 can also position the starting point of the first motion vector constituting the chained motion vector at the four corners of the decoded target block (CurBlk), and search for whether the motion vector is stored in the reference destination of the motion vector in a predetermined order of the center (C), upper left (TL), upper right (TR), lower left (BL), and lower right (BR) of the decoded target block.

[0184] The inter-frame prediction unit 205 can generate chained motion vectors for each motion vector searched and derived at each point.

[0185] As a variation, the inter-frame prediction unit 205 can compare the difference between the reference block of each motion vector searched and derived at each point and the adjacent decoded pixel values ​​of the decoding target block, and generate a chain of motion vectors by limiting at least one motion vector with a smaller difference.

[0186] As another variation, the inter-frame prediction unit 205 can perform a weighted average of the motion vectors searched and derived at each point to derive a new motion vector for generating a chain of motion vectors.

[0187] Specifically, the inter-frame prediction unit 205 can scale each motion vector based on the distance between the reference image of each motion vector and the decoded target block when weighting and averaging each motion vector to generate new motion vectors. The scaling destination can be configured as the reference image with the smallest distance to the decoded target block.

[0188] When deriving chained motion vectors, in addition to being positioned at the center of the decoded target block, the inter-frame prediction unit 205 further increases the likelihood of deriving chained motion vectors by positioning the starting point of the initial motion vector constituting the chained motion vector at the four corners of the decoded target block. As a result, an improvement in coding efficiency can be expected.

[0189] As a further control, when deriving chained motion vectors in chained motion vector prediction, the inter-frame prediction unit 205 can control the starting point of the initial motion vector constituting the chained motion vector according to the block size or block aspect ratio of the decoded target block.

[0190] For example, when deriving chained motion vectors in chained motion vector prediction, if the block size of the target block being decoded is less than or equal to a first size, the inter-frame prediction unit 205 can configure the starting point of the first motion vector constituting the chained motion vector only at the center of the target block being decoded.

[0191] When the block size of the target block being decoded is less than the first size, when deriving the chained motion vectors, there is a high probability that the initial motion vectors constituting the chained motion vectors are the same at the center of the target block and at the four corners of the target block.

[0192] Therefore, when the block size of the target block to be decoded is less than or equal to the first size, when deriving the chain motion vector, by setting the starting point of the first motion vector constituting the chain motion vector only at the center of the target block to be decoded (excluding the four corners of the target block), the amount of search processing for the chain motion vector can be reduced.

[0193] Alternatively, when deriving chained motion vectors in chained motion vector prediction, if the block size of the decoded target block is larger than the first size, in addition to being positioned at the center of the decoded target block, the inter-frame prediction unit 205 may also position the starting point of the first motion vector constituting the chained motion vector at the four corners of the decoded target block, and search for whether the motion vector is stored in the reference destination of the motion vector in a predetermined order of center, upper left, upper right, lower left, and lower right of the decoded target block.

[0194] When the block size of the target block being decoded is larger than the first size, when deriving the chained motion vectors, the initial motion vectors constituting the chained motion vectors are more likely to be different at the center of the target block and at the four corners of the target block.

[0195] Therefore, when the block size of the target block being decoded is larger than the first size, when deriving chain motion vectors, in addition to being positioned at the center of the target block, different chain motion vectors are generated by also positioning the starting point of the first motion vector constituting the chain motion vector at the four corners of the target block, which can be expected to improve coding efficiency.

[0196] Here, we also assume that the target block for decoding is not square. Therefore, the block size (number of pixels in the block) of the target block for decoding can be replaced with the width and height of the target block for decoding, and the aforementioned control can be implemented.

[0197] Specifically, when deriving chained motion vectors in chained motion vector prediction, if the width or height of the decoded target block is greater than a predetermined number of pixels, in addition to being positioned at the center of the decoded target block, the inter-frame prediction unit 205 may also position the starting point of the initial motion vector constituting the chained motion vector at the four corners of the decoded target block, and search for whether the motion vector is stored in the reference destination of the motion vector in a predetermined order of center, upper left, upper right, lower left, and lower right of the decoded target block.

[0198] The inter-frame prediction unit 205 can be configured with any one of 8 pixels, 16 pixels, or 32 pixels as the predetermined number of pixels.

[0199] Alternatively, when deriving chained motion vectors in chained motion vector prediction, if the aspect ratio of the block to be decoded is greater than a predetermined ratio, in addition to being positioned at the center of the target block, the inter-frame prediction unit 205 may also position the starting point of the initial motion vector constituting the chained motion vector at the lower left and upper right of the target block, and search for whether the motion vector is stored in the reference destination of the motion vector in the predetermined order of the center, lower left and upper right of the target block.

[0200] The inter-frame prediction unit 205 can be configured to use any one of 4, 8 or 16 as the predetermined ratio.

[0201] Secondly, use Figure 20 as well as Figure 21 An example of a control method for the reference destination when deriving chain motion vectors in chain motion vector prediction is illustrated.

[0202] Figure 20 This is a diagram illustrating an example of a method for controlling the derivation of a reference destination when generating a chain motion vector in the chain motion vector prediction by the inter-frame prediction unit 205.

[0203] like Figure 20 As shown, during the recursive motion vector search process when deriving chained motion vectors, if the reference position of each motion vector is not a reference image, the inter-frame prediction unit 205 can search for motion vectors at the storage location of motion vectors in the reference image closest to that reference position.

[0204] Even if the reference position of the motion vector is outside the reference image, new chained motion vectors can still be generated because the motion vectors linked from this motion vector may be located in different reference images.

[0205] Or, such as Figure 20 As shown, during the recursive motion vector search process when deriving chained motion vectors, if the reference position of each motion vector is not referenced by the reference image, the inter-frame prediction unit 205 can terminate the generation of chained motion vectors.

[0206] When the reference position of a motion vector is outside the reference image, since the motion vectors chained from that motion vector may also reference outside the reference image in different reference images, the amount of decoding processing can be reduced by terminating the generation of the chained motion vectors.

[0207] Alternatively, if the decoding image buffer 207 stores the inter-frame prediction padding region outside the reference image, the aforementioned record of "outside the reference image" can be replaced with the record of "outside the inter-frame prediction padding region", and the inter-frame prediction unit 205 can perform the same processing.

[0208] Figure 21 This diagram illustrates an example of a method for controlling the reference source and reference destination when a Gradual Decoding Refresh (GDR) disclosed in Non-Patent Document 1, which is known as an encoding / decoding technique for low-latency video transmission, is applied to a group of target images for decoding. The method involves controlling the generation of chain motion vectors in the chain motion vector prediction by the inter-frame prediction unit 205.

[0209] GDR is a technique that reduces the processing latency of video encoding / decoding at a certain bit rate by using intra-frame prediction encoding / decoding to update only a portion of the frame (hereinafter referred to as intra-frame update) and moving that portion of the frame according to the decoded target image. This is compared to video encoding / decoding in random access configurations that include intra-frame images, which are widely used in video transmission.

[0210] Figure 21 The intra-slice diagram illustrates the intra-update region, showing an example of the slice moving from the top to the bottom of the decoded target image.

[0211] In this specification, the three regions in the decoded target image divided by GDR are referred to as the region after intra-frame update, the region during intra-frame update, and the region before intra-frame update.

[0212] like Figure 21 As shown, when progressive decoding refresh is applied in the target image group, the inter-frame prediction unit 205 can limit the search of recursive motion vectors predicted by chained motion vectors to no more than the region in the intra-frame update.

[0213] For example, such as Figure 21 As shown, when progressive decoding refresh is applied in the group of decoded target images and the decoded target block is included in the region updated intraframe, the interframe prediction unit 205 can limit the region that can be referenced in each search of the recursive motion vectors predicted by the chained motion vectors to the region updated intraframe only in the reference image of each search target.

[0214] Alternatively, when progressive decoding refresh is applied in the group of target images and the target blocks are included in the region before the intra-frame update, the inter-frame prediction unit 205 can limit the region that can be referenced in each search of the recursive motion vectors predicted by the chained motion vectors to only the region in the reference image of each search target before the intra-frame update.

[0215] Thus, when progressive decoding refresh is applied in the target image group, the inter-frame prediction unit 205 can generate chain motion vectors with high prediction accuracy by limiting the search of recursive motion vectors predicted by chain motion vectors to no more than the region in the intra-frame update.

[0216] Figure 22This diagram illustrates an example of controlling the reference destination when deriving chain motion vectors in chain motion vector prediction via the inter-frame prediction unit 205 in the case of applying Reference Picture Resampling (RPR) disclosed in Non-Patent Document 1. The aforementioned RPR allows reference motion vectors for reference pictures with different resolutions (sampling ratios) than the decoded picture in the decoded target block.

[0217] In RPR, the inter-frame prediction unit 205 calculates the ratio of the width and height of the reference image to the width and height of the decoded image, and scales the pixel position and motion vector length of the reference image relative to the decoded image based on the ratio to generate inter-frame prediction pixels.

[0218] Therefore, for example, such as Figure 22 As shown, when the width and height of the reference image are 0.5 times the width and height of the decoded image, the pixel position and motion vector length of the reference image relative to the decoded image are scaled to 0.5 times.

[0219] On the other hand, when generating chained motion vectors, the inter-frame prediction unit 205 needs to connect the motion vectors at the same scale as the decoded image.

[0220] Therefore, as Figure 22 As shown, when the inter-frame prediction unit 205 references motion vectors from a reference image with a different resolution (sampling ratio) than the target image and recursively searches for motion vectors from the reference image to generate chained motion vectors, it can inversely scale the length of the recursively derived motion vectors according to the ratio of the width and height of the target image to the width and height of each reference image, and then add them together to generate chained motion vectors.

[0221] As a variation, when it is permissible to reference motion vectors for reference images with a different resolution (sampling ratio) than the decoded target images for a decoded target sequence or a decoded target image group, the inter-frame prediction unit 205 may apply chained motion vectors only to reference images with the same sampling ratio as the decoded target images.

[0222] In other words, when it is permissible to reference motion vectors for a reference image with a different resolution (sampling ratio) than the decoded target image for a decoded target sequence or a decoded target image group, the inter-frame prediction unit 205 can be configured not to apply chained motion vectors to a reference image with a different sampling ratio than the decoded target image.

[0223] <4. Inheritance of Inter-Frame Prediction Information in Chained Motion Vector Prediction> When merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 can control whether or not predetermined internal parameters of the inter-frame prediction pixel generation method are inherited and changed.

[0224] (Switch the interpolation filter) As an example, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 can always be configured such that the internal parameter (hpelIfIdx) corresponding to the merging candidates derived through chained motion vector prediction indicates that the switching interpolation filter should not be applied, and does not inherit the internal parameter (hpelIfIdx) associated with each motion vector in the recursive search that determines whether to apply the switching interpolation filter.

[0225] Alternatively, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may inherit the internal parameter (hpelIfIdx) associated with the first or last motion vector constituting the chained motion vector from the internal parameter (hpelIfIdx) for determining whether to apply the switching interpolation filter, which is associated with each motion vector in the recursive search, as the internal parameter (hpelIfIdx) for determining whether to apply the switching interpolation filter corresponding to the merging candidates derived through chained motion vector prediction.

[0226] Alternatively, when merging candidates are derived through chained motion vector prediction, if at least one or more of the internal parameters (hpelIfIdx) associated with each motion vector in the recursive search indicate that a switching interpolation filter should be applied, the inter-frame prediction unit 205 may be configured such that the internal parameters (hpelIfIdx) corresponding to the merging candidates derived through chained motion vector prediction indicate that a switching interpolation filter should be applied.

[0227] Furthermore, when the chain motion vectors generated by chain motion vector prediction include at least one or more block vectors, the inter-frame prediction unit 205 may be configured to always indicate that the internal parameter (hpelIfIdx) corresponding to the merge candidate derived by chain motion vector prediction, which determines whether to apply the switching interpolation filter, is not applied.

[0228] (Local brightness compensation) As an example, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 can always be configured such that the internal parameter (licFlag) corresponding to the merging candidates derived through chained motion vector prediction, which controls whether local brightness compensation is applied, indicates that local brightness compensation is not applied, and does not inherit the internal parameter (licFlag) associated with each motion vector in the recursive search that controls whether local brightness compensation is applied.

[0229] Alternatively, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may inherit the internal parameter (licFlag) associated with the first or last motion vector constituting the chained motion vector from the internal parameters (licFlag) for controlling whether to apply local brightness compensation, which is associated with each motion vector in the recursive search, and use it as the internal parameter (licFlag) for controlling whether to apply local brightness compensation corresponding to the merging candidates derived through chained motion vector prediction.

[0230] Alternatively, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may be configured such that, if at least one of the internal parameters (licFlag) associated with each motion vector in the recursive search indicates the application of local brightness compensation, the internal parameter (licFlag) corresponding to the merging candidate derived through chained motion vector prediction indicates the application of local brightness compensation.

[0231] Furthermore, when the chain motion vectors generated by chain motion vector prediction include at least one or more block vectors, the inter-frame prediction unit 205 can always configure the internal parameter (licFlagg) corresponding to the merging candidate derived by chain motion vector prediction, which controls whether local brightness compensation is applied, to indicate that local brightness compensation is not applied.

[0232] (Weighted bidirectional prediction on a block-by-block basis) As an example, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 can always be configured such that the internal parameter (bcwIdx) that determines the weight value of the weighted bidirectional prediction on a block-by-block basis, corresponding to the merging candidates derived through chained motion vector prediction, indicates that the weight value of the weighted bidirectional prediction on a block-by-block basis is a simple average value, and does not inherit the internal parameter (bcwIdx) that determines the weight value of the weighted bidirectional prediction on a block-by-block basis associated with each motion vector in the recursive search.

[0233] Alternatively, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may inherit the internal parameter (bcwIdx) associated with the weight value of the weighted bidirectional prediction on a block-by-block basis, which is associated with the first or last motion vector constituting the chained motion vector, as the internal parameter (bcwIdx) for determining the weight value of the weighted bidirectional prediction on a block-by-block basis corresponding to the merging candidate derived through chained motion vector prediction.

[0234] As a variation, when merging candidates are derived through chained motion vector prediction, if at least one of the internal parameters (bcwIdx) associated with each motion vector in the recursive search that determine the weight values ​​of the weighted bidirectional prediction on a block-by-block basis indicates that the weight values ​​of the weighted bidirectional prediction on a block-by-block basis are determined to be a simple average, the inter-frame prediction unit 205 may be configured such that the internal parameter (bceIdx) corresponding to the merging candidate derived through chained motion vector prediction determines the weight values ​​of the weighted bidirectional prediction on a block-by-block basis to be determined to be a simple average.

[0235] (Multiple Hypothesis Prediction) As an example, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 can be configured such that the internal parameter (mhpFlag) corresponding to the merging candidates derived through chained motion vector prediction, which controls whether to apply multi-hypothesis prediction, indicates that multi-hypothesis prediction is not applied, and does not inherit the internal parameter (mhpFlag, fourth internal parameter) associated with each motion vector of the recursive search, which controls whether to apply multi-hypothesis prediction.

[0236] As a variation, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may be configured to indicate whether the control associated with each motion vector of the inherited and recursive search applies the internal parameters (mhpFlag) of the multi-hypothesis prediction, which are the internal parameters (mhpFlag) of the control associated with the first or last motion vector constituting the chained motion vector, as the internal parameters (mhpFlag) of the control for whether to apply the multi-hypothesis prediction corresponding to the merging candidates derived through chained motion vector prediction.

[0237] As a variation, when merging candidates are derived through chained motion vector prediction, the inter-frame prediction unit 205 may be configured such that, if at least one of the internal parameters (mhpFlag) associated with the control whether to apply multi-hypothesis prediction indicates the application of multi-hypothesis prediction, the internal parameter (mhpFlag) corresponding to the merging candidates derived through chained motion vector prediction indicates the application of multi-hypothesis prediction.

[0238] As a further variation, when the chain motion vectors generated by chain motion vector prediction include at least one or more block vectors, the inter-frame prediction unit 205 may be configured to always indicate that multi-hypothesis prediction is not applied, corresponding to the merge candidates derived by chain motion vector prediction and the internal parameter (mhpFlag) that controls whether multi-hypothesis prediction is applied.

[0239] As mentioned above, if the internal parameters of the method for generating inter-frame predicted pixels are not inherited in chained motion vector prediction, the amount of code required to store inter-frame prediction information in the decoded image buffer 207 is reduced.

[0240] On the other hand, as mentioned above, if the internal parameters of the method for generating inter-frame predicted pixels are inherited and modified in chained motion vector prediction, the inter-frame predicted pixels generated in the chained motion vector prediction can potentially be made more accurate, and an improvement in coding efficiency can be expected.

[0241] Furthermore, since the necessity of changing the generation method of inter-frame predicted pixels is low when the chained motion vector includes at least one block vector under the aforementioned inherited internal parameters, an improvement in coding efficiency can be expected by configuring it not to inherit the internal parameters, i.e., not to change the generation method of inter-frame predicted pixels.

[0242] <5. Including candidate merging, candidate rearrangement, and vector correction for chained motion vector prediction candidates> (Rearrange and merge candidates) The later a chain of motion vector prediction candidates is derived compared to other merged candidates, the less likely the chain of motion vector prediction candidates are to be saved in the merged candidate list.

[0243] As a method to solve this problem, the inter-frame prediction unit 205 can calculate the template matching costs of all merge candidates stored in the merge candidate list, and rearrange the storage positions of each merge candidate in the merge candidate list in ascending order according to the template matching costs from low to high. The aforementioned merge candidate list includes merge candidates derived by chained motion vector prediction.

[0244] As an example, the inter-frame prediction unit 205 may use the Adaptive Reordering Merge Candidate (ARMC) disclosed in Non-Patent Document 2 as a technique related to the reordering of the storage position of each merging candidate in the merging candidate list based on the template matching cost of each merging candidate.

[0245] Furthermore, as mentioned above, since chained motion vector prediction recursively generates new motion vectors, other merged candidates may not be saved in the merged candidate list when other merged candidates are derived after the chained motion vector prediction candidates are derived.

[0246] As a method to solve this problem, the inter-frame prediction unit 205 can calculate the template matching cost of all unpruned merge candidates, including merge candidates derived from chained motion vector prediction, and store them in the merge candidate list in ascending order according to the template matching cost from low to high.

[0247] Alternatively, the inter-frame prediction unit 205 may perform a series of processes a predetermined number of times, wherein the series of processes includes calculating the template matching costs of all unpruned merge candidates, including those derived from chained motion vector predictions, and storing them in an ascending order of template matching costs in a merge candidate list. For example, the predetermined number of times may be two or three.

[0248] Therefore, merging candidates that can generate inter-frame predicted pixels with high accuracy are placed in the merging candidate list, which can be expected to improve coding efficiency.

[0249] (Correction Vector) The inter-frame prediction unit 205 can use a predetermined method to correct the motion vectors associated with the merged candidates derived from chained motion vector prediction.

[0250] As an example, the inter-frame prediction unit 205 can use template matching to correct motion vectors associated with merged candidates derived through chained motion vector prediction.

[0251] Here, the inter-frame prediction unit 205 may apply the same template matching technique as disclosed in Non-Patent Document 2 as the template matching.

[0252] As an example, the inter-frame prediction unit 205 can use bilateral matching to correct motion vectors associated with merged candidates derived through chained motion vector prediction.

[0253] Here, the inter-frame prediction unit 205 may apply the same bilateral matching technique as disclosed in Non-Patent Document 1 and Non-Patent Document 2 as the bilateral matching. For example, the inter-frame prediction unit 205 may apply the decoder-side motion vector refinement (DMVR) disclosed in Non-Patent Document 1.

[0254] As an example, the inter-frame prediction unit 205 can differentially add motion vectors derived from control information to correct motion vectors associated with merged candidates derived through chained motion vector prediction.

[0255] Here, the inter-frame prediction unit 205 can apply the same technique as the merge with motion vector difference (MMVD) disclosed in Non-Patent Document 1 and Non-Patent Document 2 as control information indicating motion vector difference.

[0256] By using a predetermined method, the inter-frame prediction unit 205 corrects the motion vectors associated with the merged candidates derived from chained motion vector prediction, thereby improving the accuracy of inter-frame prediction and improving coding efficiency.

[0257] <6. Limitations of chained motion vector prediction (whether to apply / number of images / number of candidate images to merge)> use Figure 23 Various limiting methods for chain motion vector prediction by the decoding unit 201 and the inter-frame prediction unit 205 of the image decoding apparatus 200 according to this embodiment will be described.

[0258] (Whether to apply chain motion vector prediction) First, regarding whether or not chained motion vector prediction can be applied, as mentioned above, if TMVP can be applied in its nature, then chained motion vector prediction can be applied. In other words, if TMVP cannot be applied, then chained motion vector prediction cannot be applied.

[0259] Therefore, it is possible to control whether chained motion vector prediction can be applied based on the control information (syntax) of whether time motion vector prediction (TMVP) can be applied.

[0260] That is, the inter-frame prediction unit 205 can control whether to apply chained motion vector prediction based on the control information sent from the decoding unit 201 to control whether to apply temporal motion vector prediction (TMVP).

[0261] Here, in Non-Patent Document 1, the syntax (sps_temporal_mvp_enabled_flag, first control information, first syntax) based on the decoded target sequence and the syntax (ph_temporal_mvp_enabled_flag, second control information, second syntax) based on the decoded target image are defined as the syntax for controlling whether TMVP can be applied.

[0262] That is, it is also possible to control whether chain motion vector prediction can be applied in units that are the same as or lower than TMVP.

[0263] For example, the control method for whether chained motion vector prediction can be applied in inter-frame prediction 205, which controls whether chained motion vector prediction can be applied on a unit of decoding target blocks, is as follows.

[0264] The inter-frame prediction unit 205 can control whether to apply chained motion vector prediction for the decoded target block based on control information sent from the decoding unit 201 to control whether to apply time motion vector prediction (TMVP) in units of the decoded target sequence.

[0265] More specifically, if the control information sent from the decoding unit 201 determines that time motion vector prediction (TMVP) can be applied in units of the decoded target sequence, the inter-frame prediction unit 205 determines that chained motion vector prediction for the decoded target block can be applied.

[0266] On the other hand, if the aforementioned control information determines that temporal motion vector prediction cannot be applied, the inter-frame prediction unit 205 determines that chained motion vector prediction for the decoded target block cannot be applied.

[0267] Alternatively, the inter-frame prediction unit 205 can control whether to apply chained motion vector prediction for the decoded target block based on control information sent from the decoding unit 201 to control whether to apply time motion vector prediction (TMVP) per decoded target image.

[0268] More specifically, if the control information sent from the decoding unit 201 determines that temporal motion vector prediction (TMVP) can be applied in units of decoded target images, the inter-frame prediction unit 205 determines that chained motion vector prediction for decoded target blocks can be applied.

[0269] On the other hand, if the aforementioned control information determines that temporal motion vector prediction cannot be applied, the inter-frame prediction unit 205 determines that chained motion vector prediction for the decoded target block cannot be applied.

[0270] The inter-frame prediction unit 205 can control whether to apply chained motion vector prediction based on the time motion vector prediction (TMVP) in units of the decoded target sequence or the decoded target sequence smaller than the decoded target sequence, based on the control information sent from the decoding unit 201.

[0271] Alternatively, the inter-frame prediction unit 205 can control whether to apply chained motion vector prediction based on the control information sent from the decoding unit 201 regarding whether to apply time motion vector prediction (TMVP) on a unit of decoded target image or on a unit of decoded target smaller than the decoded target image.

[0272] Alternatively, before the inter-frame prediction unit 205 controls whether chained motion vector prediction can be applied on a per-decoded target block basis, the decoding unit 201 can control whether chained motion vector prediction can be applied based on control information that controls whether temporal motion vector prediction (TMVP) can be applied.

[0273] Figure 23 This diagram illustrates an example of a process where the decoding unit 201 determines whether or not the sps_chained_mvp_enabled_flag is being decoded, based on the syntax (sps_temporal_mvp_enabled_flag) that controls whether or not the sps_chained_mvp_enabled_flag can be applied on a unit basis, in order to specify whether chained motion vector prediction can be applied on a unit basis.

[0274] like Figure 23 As shown, when it is determined in step S100 that the syntax (sps_temporal_mvp_enabled_flag) for controlling whether the temporal motion vector prediction (TMVP) can be applied on a unit of the decoded target sequence is 1, the decoding unit 201 decodes the syntax (sps_chained_mvp_enabled_flag, third syntax) for controlling whether the chained motion vector prediction can be applied on a unit of the decoded target sequence in step S101.

[0275] On the other hand, when it is determined in step S100 that the syntax (sps_temporal_mvp_enabled_flag) for controlling whether temporal motion vector prediction (TMVP) can be applied on a unit basis in the decoded target sequence is not 1, the decoding unit 201 does not decode the syntax (sps_chained_mvp_enabled_flag) for controlling whether chained motion vector prediction can be applied on a unit basis in the decoded target sequence in step S102.

[0276] Here, the decoding unit 201 determines that temporal motion vector prediction (TMVP) can be applied on a unit basis to the decoded target sequence when the value of sps_temporal_mvp_enabled_flag is 1, determines that temporal motion vector prediction (TMVP) cannot be applied on a unit basis to the decoded target sequence when the value of sps_temporal_mvp_enabled_flag is 0, and estimates the value of sps_temporal_mvp_enabled_flag to be 0 when sps_temporal_mvp_enabled_flag is not decoded.

[0277] Furthermore, when the value of sps_chained_mvp_enabled_flag is 1, the decoding unit 201 determines that chained motion vector prediction can be applied on a unit basis to the decoded target sequence; when the value of sps_chained_mvp_enabled_flag is 0, it determines that chained motion vector prediction cannot be applied on a unit basis to the decoded target sequence; and when sps_chained_mvp_enabled_flag is not decoded, the value of sps_chained_mvp_enabled_flag is estimated to be 0.

[0278] Thus, since the value of sps_temporal_mvp_enabled_flag is used to determine whether sps_chained_mvp_enabled_flag has been decoded, when it is clear that decoding of sps_chained_mvp_enabled_flag is not required, the amount of code can be reduced.

[0279] When chained motion vector prediction is combined with Adaptive Reordering Merge Candidate (ARMC) by comparing the template matching costs of the aforementioned merged candidates, improvements in coding efficiency can be more readily expected.

[0280] Therefore, as a modified example of the aforementioned control method for whether chained motion vector prediction can be applied based on whether TMVP can be applied, the inter-frame prediction unit 205 can control whether chained motion vector prediction can be applied based on both whether TMVP can be applied and whether ARMC can be applied.

[0281] Specifically, when both TMVP and ARMC can be applied, the inter-frame prediction unit 205 can determine that chained motion vector prediction can be applied. Otherwise, the inter-frame prediction unit 205 can determine that chained motion vector prediction cannot be applied.

[0282] Furthermore, the same control information as TMVP can be set to determine whether ARMC can be applied, and the determination can be made based on the value of this control information.

[0283] (Number of reference images for chain motion vector prediction / search depth) The following explains the methods for limiting the number of reference images or search depth that can be used when searching for recursive motion vectors in chained motion vector prediction.

[0284] First, we will explain the method for limiting the number of reference images that can be consulted when searching for recursive motion vectors in chained motion vector prediction.

[0285] As an example, the inter-frame prediction unit 205 may restrict the search for recursive motion vectors in chained motion vector prediction to only at least one reference image included in the list of reference images of the decoded target block.

[0286] By restricting the search for recursive motion vectors predicted by the inter-frame prediction unit 205 to only at least one reference image included in the reference image list of the decoded target block, the amount of decoding processing can be reduced compared to the case where no such restriction is imposed (e.g., the inter-frame prediction unit 205 searches all reference images stored in the decoded image buffer 207).

[0287] Alternatively, the inter-frame prediction unit 205 can restrict the search for recursive motion vectors in chained motion vector prediction to only the reference image of the motion vector reference of the target block.

[0288] In other words, the inter-frame prediction unit 205 only re-searches for block vectors within the reference image of the motion vector reference of the decoded target block, and generates chain motion vectors through chain motion vector prediction.

[0289] Secondly, the method for limiting the search depth of recursive motion vectors in chain-like motion vector prediction is explained.

[0290] As an example, the inter-frame prediction unit 205 can limit the search depth of the recursive motion vectors predicted by the chained motion vectors in the decoded target block based on control information.

[0291] Specifically, the inter-frame prediction unit 205 can limit the search depth of the recursive motion vectors in the chained motion vector prediction in the decoded target block based on the control information (sps_max_num_my_chains) sent from the decoding unit that configures the upper limit of the search depth of the recursive motion vectors for chained motion vector prediction on a per-decoded target sequence.

[0292] Here, the control information based on the decoded target sequence can be control information based on images or slices, which is lower than the control information based on the decoded target sequence, or it can be hierarchical control information that combines these control information.

[0293] Alternatively, the inter-frame prediction unit 205 can predict the search depth of recursive motion vectors using chained motion vector prediction. The inter-frame prediction unit 205 can configure 1, 2, or 3 to this predetermined fixed value.

[0294] By limiting the search depth for chained motion vector prediction in this way, the amount of decoding processing required to derive the chained motion vectors can be reduced compared to the unrestricted case.

[0295] (The number of merged candidates for chain motion vector prediction) The following explains the method for limiting the number of merged candidates derived from chain motion vector prediction.

[0296] As an example, the inter-frame prediction unit 205 can limit the number of merge candidates derived from chained motion vector prediction in the decoded target block based on control information.

[0297] Specifically, the inter-frame prediction unit 205 can limit the upper limit of the search depth of the merge candidate number derived by chain motion vector prediction in the decoded target block based on the control information (sps_max_num_cmvp_cand) sent from the decoding unit, which configures the upper limit of the search depth of the merge candidate number derived by chain motion vector prediction on a per-decode target sequence basis.

[0298] Here, the control information based on the decoded target sequence can be control information based on images or slices, which is lower than the control information based on the decoded target sequence, or it can be hierarchical control information that combines these control information.

[0299] Alternatively, the inter-frame prediction unit 205 can configure a predetermined fixed value as the number of merged candidates derived from chained motion vector prediction.

[0300] (Limitations of block vector search in chained motion vector prediction) The above describes how the search for block vectors can be included in the recursive search for motion vectors in chained motion vector prediction. However, while block vectors referencing the same image are easily found in screen images including characters, computer-generated graphics (CG), and game images, they are difficult to find in images captured by a camera.

[0301] Therefore, it is desirable to restrict the search for block vectors in chain motion vector prediction based on the features of the image.

[0302] As an example, the inter-frame prediction unit 205 can limit the block vector search based on control information, which determines whether the block vector in the chain motion vector prediction for the decoded target block can be searched.

[0303] Specifically, the inter-frame prediction unit 205 can limit whether to search for block vectors in the chain motion vector prediction for the decoded target block based on the control information (sps_cmvp_bv_search_enabled_flag) sent from the decoding unit 201, which determines whether to search for block vectors in the chain motion vector prediction on a per-decoded target sequence basis.

[0304] Here, the control information based on the decoded target sequence can be control information based on images or slices, which is lower than the control information based on the decoded target sequence, or it can be hierarchical control information that combines these control information.

[0305] (Limitations of the reference source for chain motion vector prediction) The above describes a method for adding the four corners of the target block to the starting point of the initial motion vector in the recursive motion vector search of chained motion vector prediction, in addition to decoding the center of the target block.

[0306] However, images captured by a camera are much harder to change than screen images or game images, where pixel values ​​can change rapidly on a pixel-by-pixel basis.

[0307] Therefore, it is desirable to limit the starting point of the initial motion vector in the recursive search for chained motion vector prediction based on the features of the image.

[0308] As an example, the inter-frame prediction unit 205 can limit the reference source for chained motion vector prediction based on control information. This control information determines whether, in addition to the center of the decoded target block, at least one predetermined position of the decoded target block should be added to the starting point of the initial motion vector in the chained motion vector prediction for the decoded target block. This predetermined position can be one of the four corners within the decoded target block.

[0309] Specifically, the inter-frame prediction unit 205 can, based on sps_cmvp_multi_position_flag, limit whether, in addition to the center of the decoded target block, at least one or more predetermined positions of the decoded target block are added to the starting point of the first motion vector in the chained motion vector prediction for the decoded target block.

[0310] Here, sps_cmvp_multi_position_flag is the control information that determines whether, in addition to the center of the decoded target block, at least one predetermined position of the decoded target block is added to the starting point of the first motion vector in the chained motion vector prediction for the decoded target block.

[0311] Here, the control information based on the decoded target sequence can be control information based on images or slices, which is lower than the control information based on the decoded target sequence, or it can be hierarchical control information that combines these control information.

[0312] <7. Applications of Chain-Based Motion Vector Prediction> This describes the technical features disclosed in Non-Patent Document 1 and Non-Patent Document 2, which apply chain motion vector prediction to mergers other than conventional mergers.

[0313] (MMVD, GPM, and TM merged) First, the MMVD, GPM, and template matching merge (TM merge) methods disclosed in Non-Patent Document 1 and Non-Patent Document 2 construct merge lists in the same way as conventional merges. However, in Non-Patent Document 2, the pruning (trimming) of merge candidates when constructing these merge candidate lists is enhanced compared to conventional merges, making it difficult for these merge candidate lists to be identical to the merge candidate lists stored in conventional merges.

[0314] Therefore, in MMVD, GPM, and TM merging, without using the merge candidate list of regular merging, if chained motion vector prediction is applied to the merge candidate lists of each of MMVD, GPM, and TM merging, in addition to applying chained motion vector prediction to the merge candidates of the merge candidate list of regular merging, then an improvement in coding efficiency can be expected.

[0315] Therefore, in addition to applying chained motion vector prediction to the merge candidate list for regular merging, the inter-frame prediction unit 205 can also apply chained motion vector prediction to the merge candidate list for merging mode motion vector difference.

[0316] In addition to applying chain motion vector prediction to the merge candidates in the merge candidate list for regular merging, the inter-frame prediction unit 205 can also apply chain motion vector prediction to the merge candidates in the merge candidate list for geometric segmentation mode.

[0317] In addition to applying chained motion vector prediction to the merge candidates in the merge candidate list for regular merging, the inter-frame prediction unit 205 can also apply chained motion vector prediction to the merge candidates in the merge candidate list for template matching merging.

[0318] Here, regarding the chain motion vectors generated for merging mode motion vector difference, geometric segmentation mode, and template matching merging, the inter-frame prediction unit 205 can apply the same pruning process, motion vector correction, and merging candidate (motion vector) rearrangement as in regular merging when saving the merging candidate list.

[0319] (Adaptive Motion Vector Prediction - Merging Mode) In the Adaptive Motion Vector Prediction-Merge mode disclosed in Non-Patent Documents 1 and 2, since the motion vectors of either List 0 (L0) or List 1 (L1) of the reference image list are derived through Adaptive Motion Vector Prediction (AMVP), i.e. through control information, if chained motion vector prediction is applied to the merge candidate list of the reference image list number of the application object for which chained motion vector prediction was not derived in AMVP, the number of merge candidate options that can be selected in the Adaptive Motion Vector Prediction-Merge mode increases, thereby improving the inter-frame prediction accuracy and the expected improvement in coding efficiency.

[0320] Therefore, in the adaptive motion vector prediction-merging mode, the inter-frame prediction unit 205 can apply chained motion vector prediction to the merge candidate of the reference image list number that was not derived in the adaptive motion vector prediction.

[0321] Furthermore, in the adaptive motion vector prediction-merging mode, when chained motion vector prediction is applied to a merge candidate in the merge candidate list that is a reference image list number not derived in the adaptive motion vector prediction, the inter-frame prediction unit 205 may restrict the reference image of the motion vector derived in the adaptive motion vector prediction to not refer to the generated chained motion vector.

[0322] Here, regarding the chained motion vectors generated for the adaptive motion vector prediction-merging mode, the inter-frame prediction unit 205 can apply the same pruning process, motion vector correction, and rearrangement of merge candidates (motion vectors) as in regular merging.

[0323] (Sub-block merging mode) In the affine merging disclosed in the aforementioned Non-Patent Document 1 and Non-Patent Document 2, in order to derive the motion vector of each sub-block within the decoded target block in the affine transformation, such as Figure 24 As shown, in the case of unidirectional prediction (either L0 or L1), two or three control point motion vectors cpMv are derived, and in the case of bidirectional prediction (both L0 and L1), four or six control point motion vectors cpMv are derived.

[0324] Furthermore, the control point motion vectors of each of the L0 and L1 in unidirectional or bidirectional prediction are referenced to the same reference image.

[0325] In affine merging, as disclosed in Non-Patent Document 1 and Non-Patent Document 2, these control point motion vectors are derived from merge candidates stored in a merge candidate list for affine merging.

[0326] Similar to the deriving method for affine merging candidates disclosed in Non-Patent Document 1 and Non-Patent Document 2, the inter-frame prediction unit 205 can derive merging candidates stored in the merging candidate list for affine merging.

[0327] For example, if two or three control point motion vectors are stored in locations that are spatially adjacent to or close to the decoded target block, the inter-frame prediction unit 205 can add these motion vectors to the merging candidate list for affine merging of the decoded target block.

[0328] Alternatively, if multiple motion vectors are stored in locations that are spatially adjacent to or close to the target block being decoded, the inter-frame prediction unit 205 can generate control point motion vectors for the target block being decoded based on these motion vectors and add them to the merge candidate list for affine merging.

[0329] Therefore, in affine merging mode, the inter-frame prediction unit 205 can recursively search for motion vectors for at least two control point motion vectors stored in the merging candidate list and add the derived motion vectors to each of the initial control point motion vectors to generate chained control point motion vectors, and store them in the merging candidate list.

[0330] Furthermore, when recursively searching for motion vectors for at least two control point motion vectors stored in the merging candidate list in affine merging mode, the inter-frame prediction unit 205 can restrict the reference image for the search target of each control point motion vector to the same reference image.

[0331] Alternatively, in affine merging mode, when recursively searching for motion vectors for at least two control point motion vectors stored in the merging candidate list for L0 and L1, the inter-frame prediction unit 205 may restrict the reference image of the search target for each control point motion vector in L0 and L1 to the same reference image.

[0332] Here, regarding the chain control point motion vectors generated for affine merging, the inter-frame prediction unit 205 can apply the same pruning process, motion vector correction, and rearrangement of merging candidates (motion vectors) as in regular merging.

[0333] According to the image decoding apparatus of this embodiment, since the inter-frame prediction unit 205 derives the following motion vector as the motion vector of the decoding target block, that is, it generates a motion vector by recursively searching for motion vectors stored in the decoding image buffer 207 for positions that are adjacent or close to the decoding target block in space or time and adding all the derived motion vectors, the options for motion vectors that can be selected in inter-frame prediction are increased, the accuracy of inter-frame prediction is improved, and the coding efficiency is improved.

[0334] The aforementioned image decoding device 200 can be a program that enables the computer to perform various functions (steps) and is implemented. Industrial availability

[0335] Furthermore, according to this embodiment, for example, since it is possible to improve the overall quality of service in dynamic image communication, it is possible to contribute to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs): "Building resilient infrastructure, promoting sustainable industrialization and pursuing the expansion of innovation." Figure Labels

[0336] 200 Image Decoding Device 201 Decoding Department 202 Inverse Quantization Department 203 Inverse Transformation Unit 204 Intra-frame Prediction Unit 205 Inter-frame Prediction Unit 206 Adder 207 Decoding Image Buffer 210 Code Input Section 220 Image Output Unit

Claims

1. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit controls whether the chained motion vector prediction can be applied based on control information sent from the decoding unit to control whether the temporal motion vector prediction can be applied.

2. The image decoding device according to claim 1, characterized in that, The inter-frame prediction unit controls whether to apply the chained motion vector prediction for the decoded target block based on first control information sent from the decoding unit to control whether to apply time motion vector prediction in units of the decoded target sequence.

3. The image decoding apparatus according to claim 2, characterized in that, If the first control information determines that temporal motion vector prediction can be applied, the inter-frame prediction unit determines that chained motion vector prediction can be applied for the decoded target block. If the first control information determines that temporal motion vector prediction cannot be applied, the inter-frame prediction unit determines that chained motion vector prediction for the decoded target block cannot be applied.

4. The image decoding apparatus according to claim 1, characterized in that, The inter-frame prediction unit controls whether to apply the chained motion vector prediction for the decoded target block based on second control information sent from the decoding unit to control whether to apply time motion vector prediction in units of the decoded target image.

5. The image decoding apparatus according to claim 4, characterized in that, If the second control information determines that temporal motion vector prediction can be applied, the inter-frame prediction unit determines that the chained motion vector prediction for the decoded target block can be applied. If the second control information determines that temporal motion vector prediction cannot be applied, the inter-frame prediction unit determines that the chained motion vector prediction for the decoded target block cannot be applied.

6. The image decoding apparatus according to claim 1, characterized in that, The inter-frame prediction unit controls whether to apply the chained motion vector prediction, either in units of the decoded target sequence or in units of decoded targets smaller than the decoded target sequence, based on first control information sent from the decoding unit to control whether time motion vector prediction can be applied in units of the decoded target sequence.

7. The image decoding apparatus according to claim 1, characterized in that, The inter-frame prediction unit controls whether to apply the chained motion vector prediction, either in units of the decoded target image or in units of decoded targets smaller than the decoded target image, based on second control information sent from the decoding unit to control whether time motion vector prediction can be applied in units of the decoded target image.

8. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The decoding unit controls whether the chain motion vector prediction can be applied based on control information that controls whether time motion vector prediction can be applied.

9. The image decoding apparatus according to claim 8, characterized in that, When the first syntax for controlling whether time-based motion vector prediction can be applied, based on the decoded target sequence, is set to 1, the decoding unit decodes the syntax for controlling whether the chain-like motion vector prediction can be applied, based on the decoded target sequence. If the first syntax is not 1, the decoding unit does not decode the syntax that controls whether chained motion vector prediction can be applied on a unit basis, based on the decoded target sequence.

10. The image decoding apparatus according to claim 9, characterized in that, When the value of the first syntax is 1, the decoding unit is determined to be able to apply time motion vector prediction on a unit basis, based on the decoded target sequence. When the value of the first syntax is 0, the decoding unit determines that it cannot apply time motion vector prediction on a unit basis, based on the decoded target sequence. If the first syntax is not decoded, the decoding unit estimates the value of the first syntax as 0.

11. The image decoding apparatus according to claim 8, characterized in that, When the value of the third syntax, which controls whether chained motion vector prediction can be applied on a unit basis (based on the decoded target sequence), is 1, the decoding unit determines that the chained motion vector prediction can be applied on a unit basis (based on the decoded target sequence). When the value of the third syntax is 0, the decoding unit is determined to be able to apply chained motion vector prediction on a unit basis, based on the decoded target sequence. If the third syntax is not decoded, the decoding unit estimates the value of the third syntax as 0.

12. The image decoding apparatus according to claim 1, characterized in that, The inter-frame prediction unit controls whether the chained motion vector prediction can be applied based on whether temporal motion vector prediction can be applied and whether adaptive reordering and merging of candidates can be applied.

13. The image decoding apparatus according to claim 12, characterized in that, The inter-frame prediction unit is configured to apply the chained motion vector prediction when temporal motion vector prediction can be applied and adaptive reordering and merging of candidates can be applied.

14. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction to derive these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores the motion vector candidates. The inter-frame prediction unit restricts the search for recursive motion vectors in the chained motion vector prediction to only at least one reference image included in the reference image list of the decoded target block.

15. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit restricts the recursive motion vector search in the chained motion vector prediction to a reference image that is only used for the motion vector reference of the decoded target block.

16. An image decoding apparatus, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit limits the search depth of recursive motion vectors in the chained motion vector prediction within the decoded target block based on the control information sent from the decoding unit.

17. The image decoding apparatus according to claim 16, characterized in that, The control information configures the upper limit of the search depth of the recursive motion vectors in the chained motion vector prediction on a unit basis, based on the decoded target sequence.

18. An image decoding apparatus, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit configures a predetermined fixed value for the search depth of the recursive motion vectors in the chained motion vector prediction.

19. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit limits the number of merge candidates derived from the chained motion vector prediction in the decoded target block based on the control information sent from the decoding unit.

20. The image decoding apparatus according to claim 19, characterized in that, The control information configures an upper limit on the search depth of the merged candidate number derived from the chain motion vector prediction, on a unit basis, of the decoded target sequence.

21. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients. The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching for motion vectors located spatially or temporally adjacent to or close to the decoded target block in the merged candidate list of stored motion vectors. All derived motion vectors are summed to generate chained motion vectors. Chained motion vector prediction then derives these chained motion vectors as candidates for motion vectors of the decoded target block. The inter-frame prediction unit configures a predetermined fixed value for the number of merged candidates derived from the chained motion vector prediction.

22. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit determines whether to perform a block vector search in the chain motion vector prediction for the target block of the decoded block based on the control information sent from the decoding unit.

23. The image decoding apparatus according to claim 22, characterized in that, The control information determines whether a block vector search can be performed in the chain motion vector prediction, based on the decoded target sequence.

24. An image decoding device, characterized in that, include: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the decoded pixels stored in the decoded image buffer; as well as An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. Based on the control information sent from the decoding unit, the inter-frame prediction unit determines whether to add the starting point of the first motion vector in the chained motion vector prediction for the decoded target block to at least one predetermined position in the decoded target block, in addition to adding it to the center of the decoded target block.

25. The image decoding apparatus according to claim 24, characterized in that, The control information determines, on a unit basis, whether to add the starting point of the first motion vector in the chained motion vector prediction for the decoded target block to at least one predetermined position in the decoded target block, in addition to adding it to the center of the decoded target block.

26. An image decoding method, characterized in that, include: Step A: Variable-length decoding of the code information and output of quantization values ​​and control information; Step B: Inverse quantization is performed on the quantized value and the transformation coefficients are output; Step C: Perform an inverse transformation on the transformation coefficients and output the predicted residual pixels; Step D: Generate intra-frame predicted pixels from the control information and the decoded pixels; Step E: Store the decoded pixels in the decoded image buffer; Step F: Generate inter-frame predicted pixels from the control information and the stored decoded pixels; as well as Step G involves adding at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. In step F, a chain of motion vectors is generated by recursively searching the motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merge candidate list, summing all derived motion vectors, and then performing chain motion vector prediction. This chain motion vector prediction yields the chain of motion vectors as candidates for motion vectors of the decoded target block. The merge candidate list stores these motion vector candidates. In step F, the application of the chain motion vector prediction is controlled based on control information that controls whether time motion vector prediction can be applied.

27. A program, characterized in that, To enable the computer to function as an image decoding device. The image decoding device includes: The decoding unit performs variable-length decoding on the code information and outputs quantized values ​​and control information. The inverse quantization unit performs inverse quantization on the quantized value and outputs the transformation coefficients; The inverse transform unit performs an inverse transform on the transform coefficients and outputs the predicted residual pixels; The intra-frame prediction unit generates intra-frame predicted pixels from the control information and the decoded pixels; Decode the image buffer to store the decoded pixels; The inter-frame prediction unit generates inter-frame prediction pixels from the control information and the stored decoded pixels; and An adder adds at least one of the intra-frame predicted pixels and the inter-frame predicted pixels to the predicted residual pixels to generate the decoded pixels. The inter-frame prediction unit generates chained motion vectors by recursively searching motion vectors stored in the decoded image buffer for positions spatially or temporally adjacent to or close to the decoded target block stored in the merged candidate list, summing all derived motion vectors, and then performing chained motion vector prediction. The chained motion vector prediction derives these chained motion vectors as candidates for motion vectors of the decoded target block. The merged candidate list stores these motion vector candidates. The inter-frame prediction unit controls whether the chained motion vector prediction can be applied based on control information sent from the decoding unit to control whether the temporal motion vector prediction can be applied.