Image decoding device, image decoding method, and program
Geometric block partition merging in image decoding systems addresses the issue of inappropriate block divisions by selecting optimal shapes for object boundaries, enhancing coding performance and image quality.
Patent Information
- Application Number
- JP2024166919
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2039-12-26
AI Technical Summary
Conventional image decoding techniques that divide blocks into rectangles may fail to select appropriate block divisions when object boundaries appear in directions other than the block boundaries, leading to increased prediction errors and reduced subjective image quality.
Implementing geometric block partition merging, where the merging mode identification unit determines the application of geometric block partitioning based on the block aspect ratio, allowing for appropriate block division shapes to be selected for object boundaries regardless of their direction.
This approach reduces prediction errors and improves coding performance and subjective image quality by enabling accurate block partitioning for object boundaries.
Smart Images

Figure 0007772893000001 
Figure 0007772893000002 
Figure 0007772893000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image decoding device, an image decoding method, and a program. [Background technology]
[0002] Non-Patent Documents 1 and 2 disclose a rectangular block division technique (rectangular division technique) called QTBTTT (Quad-Tree-Binary-Tree-Ternary-Tree). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Versatile Video Coding (Draft 7), JVET-N1001 [Non-patent document 2] ITU-T H.265 High Efficiency Video Coding Summary of the Invention [Problem to be solved by the invention]
[0004] However, when the above-described conventional technique only divides blocks into rectangles, there is a problem in that when an object boundary appears in any direction relative to the block boundary, it is possible that an appropriate block division for the object boundary may not be selected.
[0005] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide an image decoding device, an image decoding method, and a program that can realize an improved coding performance effect by reducing prediction errors and an improved subjective image quality effect by selecting an appropriate block division boundary for an object boundary by applying geometric block division merging to a target block that has been divided into rectangular parts, thereby enabling an appropriate block division shape to be selected for an object boundary that appears in any direction. [Means for solving the problem]
[0006] A first feature of the present invention is an image decoding device including a merging unit configured to apply geometric block partition merging to a target block that has been rectangularly partitioned, the merging unit including a merging mode identification unit configured to identify whether or not to apply the geometric block partition merging, and the merging mode identification unit determining whether or not to apply the geometric block partition merging based on a block aspect ratio of the target block.
[0007] A second feature of the present invention is an image decoding method including a step of applying geometric block division merging to a target block that has been divided into rectangular blocks, the step including a step A of determining whether or not to apply the geometric block division merging, and the step A determining whether or not to apply the geometric block division merging based on the block aspect ratio of the target block.
[0008] A third feature of the present invention is a program that causes a computer to function as an image decoding device, the image decoding device including a merging unit configured to apply geometric block partition merging to a target block that has been rectangularly partitioned, the merging unit including a merging mode identification unit configured to identify whether or not to apply the geometric block partition merging, and the merging mode identification unit determining whether or not to apply the geometric block partition merging based on a block aspect ratio of the target block. [Effects of the Invention]
[0009] According to the present invention, by applying geometric block partition merging to a target block that has been partitioned into rectangles, an appropriate block partition shape can be selected for object boundaries that appear in any direction, thereby providing an image decoding device, image decoding method and program that can achieve improved coding performance by reducing prediction errors and improved subjective image quality by selecting appropriate block partition boundaries for object boundaries. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an example of a configuration of an image processing system 1 according to an embodiment. [Figure 2] 1 is a diagram illustrating an example of functional blocks of an image encoding device 100 according to an embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of functional blocks of an inter prediction unit 111 of an image encoding device 100 according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of functional blocks of an image decoding device 200 according to an embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of functional blocks of an inter prediction unit 241 of an image decoding device 200 according to an embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of functional blocks of a merge unit 111A2 of an inter prediction unit 111 of an image encoding device 100 according to an embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of functional blocks of a merger unit 241A2 of an inter prediction unit 241 of an image decoding device 200 according to an embodiment. [Figure 8] 10 is a flowchart showing an example of a method for specifying whether or not geometric block partition merging is applied in a merge mode specifying unit 241A21 of the inter prediction unit 241 of the image decoding device 200 according to an embodiment. [Figure 9]10 is a flowchart showing an example of a method for specifying whether or not geometric block partition merging is applied in a merge mode specifying unit 241A21 of the inter prediction unit 241 of the image decoding device 200 according to an embodiment. [Figure 10] 10 is a flowchart showing an example of a method for specifying whether or not geometric block partition merging is applied in a merge mode specifying unit 241A21 of the inter prediction unit 241 of the image decoding device 200 according to an embodiment. [Figure 11] FIG. 10 is a diagram showing an example of a method for defining a geometric block division pattern according to a modified example. [Figure 12] FIG. 10 is a diagram showing an example of an elevation angle φ that defines a geometric block division pattern according to a modified example. [Figure 13] FIG. 10 is a diagram showing an example of a position (distance ρ) that defines a geometric block division pattern according to a modified example. [Figure 14] 10A and 10B are diagrams illustrating an example of a method for defining an elevation angle φ and a distance ρ according to a modified example. [Figure 15] FIG. 10 is a diagram showing an example of an index table showing combinations of elevation angles φ and distances (positions) ρ that define geometric block division patterns according to a modified example. [Figure 16] FIG. 10 is a diagram illustrating an example of control of the decoding method of "partition_idx" according to the block size or block aspect ratio of the target block according to one modified example. [Figure 17] 10 is a flowchart illustrating an example of a process of registering and pruning motion information in a merge list according to an embodiment. [Figure 18] FIG. 10 is a diagram showing an example of a merge list constructed as a result of a process of registering and pruning motion information to the merge list according to an embodiment. [Figure 19] FIG. 10 is a diagram illustrating spatial merging according to an embodiment. [Figure 20] FIG. 10 is a diagram illustrating a time merge according to an embodiment. [Figure 21] 10A and 10B are diagrams illustrating an example of a scaling process of a motion vector inherited in a temporal merge according to an embodiment. [Figure 22] FIG. 10 is a diagram illustrating history merging according to an embodiment. [Figure 23] 10A and 10B are diagrams illustrating an example of a geometric block division pattern in a target block when geometric block division merging is enabled according to a modified example, and an example of the positional relationship between spatial merge and temporal merge for the target block. [Figure 24] 10A and 10B are diagrams illustrating an example of a geometric block division pattern in a target block when geometric block division merging is enabled according to a modified example, and an example of the positional relationship between spatial merge and temporal merge for the target block. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0012] First Embodiment An image processing system 10 according to a first embodiment of the present invention will be described below with reference to Figures 1 to 10. Figure 1 is a diagram showing the image processing system 10 according to this embodiment.
[0013] As shown in FIG. 1, an image processing system 10 according to this embodiment includes an image encoding device 100 and an image decoding device 200.
[0014] The image encoding device 100 is configured to generate encoded data by encoding an input image signal. The image decoding device 200 is configured to generate an output image signal by decoding the encoded data.
[0015] The encoded data may be transmitted from the image encoding device 100 to the image decoding device 200 via a transmission path. The encoded data may be stored in a storage medium and then provided from the image encoding device 100 to the image decoding device 200.
[0016] (Image encoding device 100) The image encoding device 100 according to this embodiment will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the image encoding device 100 according to this embodiment.
[0017] As shown in FIG. 2, the image encoding device 100 includes an inter prediction unit 111, an intra prediction unit 112, a subtractor 121, an adder 122, a transform / quantization unit 131, an inverse transform / inverse quantization unit 132, an encoding unit 140, an in-loop filter processing unit 150, and a frame buffer 160.
[0018] The inter prediction unit 111 is configured to generate a prediction signal by inter prediction (inter-frame prediction).
[0019] Specifically, the inter prediction unit 111 is configured to identify a reference block included in the reference frame by comparing the frame to be coded (hereinafter referred to as the target frame) with a reference frame stored in the frame buffer 160, and to determine a motion vector (mv) for the identified reference block.
[0020] The inter prediction unit 111 is configured to generate, for each current block to be coded (hereinafter, current block) based on the reference block and the motion vector. The inter prediction unit 111 is configured to output the prediction signal to the subtractor 121 and the adder 122. Here, the reference frame is a frame different from the current frame.
[0021] The intra prediction unit 112 is configured to generate a prediction signal by intra prediction (prediction within a frame).
[0022] Specifically, the intra prediction unit 112 is configured to identify a reference block included in a target frame, and generate a prediction signal for each target block based on the identified reference block. The intra prediction unit 112 is also configured to output the prediction signal to the subtractor 121 and the adder 122.
[0023] Here, the reference block is a block that is referenced for the current block, for example, a block that is adjacent to the current block.
[0024] The subtractor 121 is configured to subtract the prediction signal from the input image signal and output the prediction residual signal to the transform / quantization unit 131. Here, the subtractor 121 is configured to generate a prediction residual signal that is the difference between the prediction signal generated by intra prediction or inter prediction and the input image signal.
[0025] The adder 122 is configured to add the prediction signal to the prediction residual signal output from the inverse transform / inverse quantization unit 132 to generate a pre-filter decoded signal, and to output the pre-filter decoded signal to the intra prediction unit 112 and the in-loop filter processing unit 150.
[0026] Here, the unfiltered decoded signal forms a reference block used by the intra prediction unit 112.
[0027] The transform / quantization unit 131 is configured to perform a transform process on the prediction residual signal and to obtain coefficient level values. Furthermore, the transform / quantization unit 131 may be configured to quantize the coefficient level values.
[0028] Here, the transform process is a process of transforming a prediction residual signal into a frequency component signal. In this transform process, a basis pattern (transform matrix) corresponding to a discrete cosine transform (DCT) or a basis pattern (transform matrix) corresponding to a discrete sine transform (DST) may be used.
[0029] The inverse transform and inverse quantization unit 132 is configured to perform inverse transform processing on the coefficient level values output from the transform and quantization unit 131. Here, the inverse transform and inverse quantization unit 132 may be configured to perform inverse quantization on the coefficient level values prior to the inverse transform processing.
[0030] Here, the inverse transform processing and inverse quantization are performed in the reverse order to the transform processing and quantization performed by the transform / quantization unit 131 .
[0031] The encoding unit 140 is configured to encode the coefficient level values output from the transform / quantization unit 131 and output encoded data.
[0032] Here, for example, the coding is entropy coding that assigns codes of different lengths based on the probability of occurrence of coefficient level values.
[0033] The encoding unit 140 is also configured to encode control data used in the decoding process in addition to the coefficient level values.
[0034] Here, the control data may include size data such as a coding block (CU: Coding Unit) size, a prediction block (PU: Prediction Unit) size, and a transform block (TU: Transform Unit) size.
[0035] Furthermore, the control data may include header information such as a sequence parameter set (SPS), a picture parameter set (PPS), and a slice header, as will be described later.
[0036] The in-loop filtering unit 150 is configured to perform filtering on the unfiltered decoded signal output from the adder 122 and to output the filtered decoded signal to the frame buffer 160 .
[0037] Here, for example, the filtering is deblocking filtering that reduces distortion occurring at the boundary portions of blocks (encoded blocks, predicted blocks, or transformed blocks).
[0038] The frame buffer 160 is configured to store reference frames used by the inter prediction unit 111.
[0039] Here, the filtered decoded signal forms a reference frame used in the inter prediction unit 111.
[0040] (Inter prediction unit 111) The inter prediction unit 111 of the image encoding device 100 according to this embodiment will be described below with reference to Fig. 3. Fig. 3 is a diagram showing an example of functional blocks of the inter prediction unit 111 of the image encoding device 100 according to this embodiment.
[0041] As shown in FIG. 3, the inter prediction unit 111 has an mv derivation unit 111A, an AMVR unit 111B, an mv refinement unit 111B, and a prediction signal generation unit 111D.
[0042] The inter prediction unit 111 is an example of a prediction unit configured to generate a prediction signal to be included in a current block based on a motion vector.
[0043] As shown in FIG. 3, the mv derivation unit 111A has an AMVP (Adaptive Motion Vector Prediction) unit 111A1 and a merge unit 111A2, and is configured to receive as input a target frame and a reference frame from the frame buffer 160 and acquire a motion vector.
[0044] The AMVP unit 111A1 is configured to identify a reference block included in the reference frame by comparing the target frame with the reference frame, and to search for a motion vector for the identified reference block.
[0045] Furthermore, the above-described search process is performed on a plurality of reference frame candidates, and the reference frame and motion vector to be used for prediction of the current block are determined and output to the subsequent predicted signal generation unit 111D.
[0046] A maximum of two reference frames and two motion vectors can be used for one block. Using only one set of reference frame and motion vector for one block is called "uni-prediction," while using two sets of reference frame and motion vector is called "bi-prediction." Hereinafter, the first set will be called "L0" and the second set will be called "L1."
[0047] Furthermore, when the AMVP unit 111A finally transmits the above-mentioned determined motion vector to the image decoding device 200, in order to reduce the amount of coding, it selects an mvp (motion vector predictor) derived from adjacent encoded motion vectors that has a small difference from the motion vector of the target block, i.e., a small motion vector difference (mvd).
[0048] The index indicating the selected mvp and mvd and the index indicating the reference frame (hereinafter, Refidx) are coded by the coding unit 140 and transmitted to the image decoding device 200. This process is generally called adaptive motion vector prediction coding (AMVP).
[0049] Note that the above-mentioned motion vector search method, reference frame and motion vector determination method, mvp selection method, and mvd calculation method can be performed using known techniques, so details thereof will be omitted.
[0050] Unlike the AMVP unit 111A1, the merge unit 111A2 does not search for and derive motion information of the target block and then transmit mvd as the difference with an adjacent block, but is configured to take a target frame and a reference frame as input, and use an adjacent block in the same frame as the target block or a block at the same position in a frame different from the target frame as the reference block, and inherit and use the motion information of the reference block as is. Such processing is generally called merge coding (hereinafter, "merge").
[0051] If the block is a merge block, a merge list for the block is first constructed. The merge list is a list of multiple combinations of reference frames and motion vectors. Each combination is assigned an index (hereinafter referred to as a merge index). Instead of individually encoding the combination information of RefiDx and the motion vector (hereinafter referred to as motion information), the image encoding device 100 encodes only the merge index and transmits it to the image decoding device 200.
[0052] By using a common merge list construction method between the image encoding device 100 and the image decoding device 200, the image decoding device 200 can decode motion information from merge index information alone. The merge list construction method and the geometric block division method of inter prediction blocks according to this embodiment will be described later.
[0053] The mv refinement unit 111B is configured to perform a refinement process to modify the motion vector output from the merge unit 111A2. For example, a known refinement process to modify the motion vector is DMVR (Decoder-side Motion Vector Refinement) described in Non-Patent Document 1. In this embodiment, the known method described in Non-Patent Document 1 can be used as the refinement process, and therefore a description thereof will be omitted.
[0054] The prediction signal generation unit 111C is configured to receive a reference frame and a motion vector as input and output an MC prediction image signal. As the processing in the prediction signal generation unit 111C can use a known method described in Non-Patent Document 1, a description thereof will be omitted.
[0055] (Image decoding device 200) The image decoding device 200 according to this embodiment will be described below with reference to Fig. 4. Fig. 4 is a diagram showing an example of functional blocks of the image decoding device 200 according to this embodiment.
[0056] As shown in FIG. 4, the image decoding device 200 includes a decoding unit 210, an inverse transform / inverse quantization unit 220, an adder 230, an inter prediction unit 241, an intra prediction unit 242, an in-loop filtering unit 250, and a frame buffer 260.
[0057] The decoding unit 210 is configured to decode the coded data generated by the image coding device 100, and to decode the coefficient level values.
[0058] Here, the decoding is, for example, entropy decoding, which is the reverse procedure of the entropy encoding performed by the encoding unit 140.
[0059] The decoding unit 210 may also be configured to obtain the control data by decoding the encoded data.
[0060] As mentioned above, the control data may include size data such as the coding block size, the prediction block size, and the transform block size.
[0061] The inverse transform / inverse quantization unit 220 is configured to perform inverse transform processing on the coefficient level values output from the decoding unit 210. Here, the inverse transform / inverse quantization unit 220 may be configured to perform inverse quantization on the coefficient level values prior to the inverse transform processing.
[0062] Here, the inverse transform processing and inverse quantization are performed in the reverse order to the transform processing and quantization performed by the transform / quantization unit 131 .
[0063] The adder 230 is configured to add the prediction signal to the prediction residual signal output from the inverse transform / inverse quantization unit 220 to generate a pre-filtered decoded signal, and output the pre-filtered decoded signal to the intra prediction unit 242 and the in-loop filter processing unit 250.
[0064] Here, the unfiltered decoded signal forms a reference block used by the intra prediction unit 242.
[0065] Like the inter prediction unit 111, the inter prediction unit 241 is configured to generate a prediction signal by inter prediction (inter-frame prediction).
[0066] Specifically, the inter prediction unit 241 is configured to generate a prediction signal for each prediction block based on a motion vector decoded from encoded data and a reference signal included in a reference frame. The inter prediction unit 241 is configured to output the prediction signal to the adder 230.
[0067] Like the intra prediction unit 112, the intra prediction unit 242 is configured to generate a prediction signal by intra prediction (intra-frame prediction).
[0068] Specifically, the intra prediction unit 242 is configured to identify a reference block included in the target frame, and generate a prediction signal for each prediction block based on the identified reference block. The intra prediction unit 242 is configured to output the prediction signal to the adder 230.
[0069] Similar to the in-loop filter processing unit 150, the in-loop filter processing unit 250 is configured to perform filtering on the unfiltered decoded signal output from the adder 230 and to output the filtered decoded signal to the frame buffer 260.
[0070] Here, for example, the filtering process is deblocking filtering process that reduces distortion occurring at the boundary portions of blocks (encoded blocks, predicted blocks, transformed blocks, or sub-blocks obtained by dividing these).
[0071] Similar to the frame buffer 160, the frame buffer 260 is configured to store reference frames used by the inter prediction unit 241.
[0072] Here, the filtered decoded signal forms a reference frame used by the inter prediction unit 241.
[0073] (Inter prediction unit 241) The inter prediction unit 241 according to this embodiment will be described below with reference to Fig. 5. Fig. 5 is a diagram showing an example of functional blocks of the inter prediction unit 241 according to this embodiment.
[0074] As shown in FIG. 5, the inter prediction unit 241 has an mv decoding unit 241A, an mv refinement unit 241B, and a prediction signal generation unit 241C.
[0075] The inter prediction unit 241 is an example of a prediction unit configured to generate a prediction signal included in a prediction block based on a motion vector.
[0076] The mv decoding unit 241A has an AMVP unit 241A1 and a merge unit 241A2, and is configured to obtain motion vectors by decoding the target frame and reference frame input from the frame buffer 260, and the control data received from the image encoding device 100.
[0077] The AMVP unit 241A1 is configured to receive a target frame, a reference frame, and an index indicating mvp and mvd, Refidx, from the image encoding device 100, and decode motion vectors. As a known method can be used for decoding motion vectors, details thereof will be omitted.
[0078] The merge unit 241A2 is configured to receive a merge index from the image encoding device 100 and decode a motion vector.
[0079] Specifically, the merging unit 241A2 is configured to construct a merge list and acquire, from the constructed merge list, a motion vector corresponding to the received merge index, in the same manner as in the image encoding device 100. The method of constructing the merge list will be described in detail later.
[0080] The mv refinement unit 241B is configured to perform a refinement process to modify a motion vector, similar to the mv refinement unit 111B.
[0081] The predicted signal generation unit 241C is configured to generate a predicted signal based on a motion vector, similar to the predicted signal generation unit 111C.
[0082] Details of the prediction signal generation process when the geometric block division merge is enabled will be described later. As for the prediction signal generation process when other merge modes are enabled, the known technology described in Non-Patent Document 1 can be used in this embodiment, and therefore the description will be omitted.
[0083] (Merge section) 6 and 7, a description will be given of the merging unit 111A2 of the inter prediction unit 111 of the image encoding device 100 according to this embodiment and the merging unit 241A2 of the inter prediction unit 241 of the image decoding device 200. Figures 6 and 7 are diagrams showing example functional blocks of the merging unit 111A2 of the inter prediction unit 111 of the image encoding device 100 according to this embodiment and the merging unit 241A2 of the inter prediction unit 241 of the image decoding device 200.
[0084] As shown in FIG. 6, the merger 111A2 includes a merge mode specification unit 111A21, a geometric block division unit 111A22, and a merge list construction unit 111A23.
[0085] As shown in FIG. 7, the merger 241A2 includes a merge mode identifier 241A21, a geometric block divider 241A22, and a merge list constructor 241A23.
[0086] The difference between merge unit 111A2 and merge unit 241A2 is that the input and output of various indexes described below are reversed in merge mode identification unit 111A21 and merge mode identification unit 241A21, geometric block division unit 111A22 and geometric block division unit 241A22, and merge list construction unit 111A23 and merge list construction unit 241A23.
[0087] That is, the various indexes output by the merge unit 111A2 are input by the merge unit 241A2. Other than that, the functions of the merge unit 111A2 and the merge unit 241A2 are the same, and therefore, for simplicity of explanation, the functions of the merge unit 241A2 will be representatively explained below.
[0088] The merge mode specifying unit 241A21 is configured to specify whether or not the geometric block merge mode is to be applied to the target block that has been divided into rectangular blocks.
[0089] In addition to the geometric block division merge, the merge modes include, for example, the normal merge adopted in Non-Patent Document 1, sub-block merge, MMVD (Merge mode with MVD), CIIP (Combined inter and intra prediction), and IBC (Intra Block Copy).
[0090] In this embodiment, in addition to these merge modes, a new geometric block division merge is applied, which allows an appropriate block division shape to be selected using geometric block division, even if the object boundary appears in any direction for a target block that has been divided into rectangular blocks.As a result, prediction errors are reduced, which can be expected to improve coding performance and subjective image quality near object boundaries.
[0091] The geometric block division unit 241A22 is configured to identify a geometric block division pattern of the target block that has been divided into rectangular blocks, and divide the target block into geometric blocks using the identified division pattern. The method for identifying the geometric block division pattern will be described in detail later.
[0092] The merge list construction unit 241A23 is configured to construct a merge list for the current block and decode the motion information.
[0093] The process of constructing a merge list consists of three stages: a process of checking whether motion information is available, a process of registering and pruning motion information, and a process of decoding motion information. Details of each stage will be described later.
[0094] (Method of Determining Whether or Not to Apply Geometric Block Division Merge in the Merge Mode Determining Unit 241A21) A method for specifying whether or not to apply geometric block division merging in the merge mode specifying unit 241A21 will be described below with reference to Figures 8 to 10. Figures 8 to 10 are flowcharts showing an example of a method for specifying whether or not to apply geometric block division merging in the merge mode specifying unit 241A21.
[0095] As shown in FIG. 8, the merge mode specification unit 241A21 is configured to specify that the geometric block division merge is to be applied (the merge mode of the target block is the geometric block division merge) when the normal merge is not applied and the CIIP is not applied.
[0096] Specifically, as shown in FIG. 8, in step S7-1, the merge mode determination unit 241A21 determines whether or not to apply normal merging, and if it determines that normal merging is to be applied, it proceeds to step S7-5, and if it determines that normal merging is not to be applied, it proceeds to step S7-2.
[0097] In step S7-2, merge mode specification section 241A21 specifies whether or not CIIP is to be applied, and if it is determined that CIIP is to be applied, the process proceeds to step S7-3, and if it is determined that CIIP is not to be applied, the process proceeds to step S7-4.
[0098] The merge mode specifying unit 241A21 may skip step S7-2 and proceed directly to step S7-4, which means that if normal merging is not applicable, the merge mode of the target block is specified as geometric block division merging.
[0099] In step S7-3, the merge mode specifying unit 241A21 specifies that CIIP is to be applied to the target block (the merge mode of the target block is CIIP), and ends this process.
[0100] The image encoding device 100 and the image decoding device 200 store, as an internal parameter, the result of specifying whether or not geometric block division merging is to be applied to the current block.
[0101] In step S7-4, the merge mode specifying unit 241A21 specifies that geometric block division merge is to be applied to the target block (the merge mode of the target block is geometric block division merge), and ends this process.
[0102] In step S7-5, the merge mode specifying unit 241A21 specifies that normal merge is to be applied to the target block (the merge mode of the target block is normal merge), and ends this process.
[0103] In the flowchart of FIG. 8, even if step S7-2 is replaced with a merge other than CIIP in the future, the merge mode specification unit 241A21 may specify whether or not to apply the geometric block division merge depending on the result of determining whether or not the replaced merge can be applied.
[0104] Furthermore, in the flowchart of FIG. 8A, even if a merge other than CIIP is added between step S7-2 and step S7-4 in the future, the merge mode specification unit 241A21 may specify whether or not to apply the geometric block division merge depending on the result of determining whether or not the added merge can be applied.
[0105] Next, the conditions for determining whether or not to apply normal merging in step S7-1 will be described with reference to FIG.
[0106] As shown in FIG. 9, in step S7-1-1, the merge mode specifying unit 241A21 determines whether or not the normal merge flag needs to be decoded.
[0107] If the merge mode identification unit 241A21 determines that the judgment condition of step S7-1-1 is met, i.e., that normal merge decoding is necessary, it proceeds to step S7-1-2, and if it determines that the condition of step S7-1-1 is not met, i.e., that normal merge decoding is not necessary, it proceeds to step S7-1-3.
[0108] Here, if the merge mode identification unit 241A21 determines that the normal merge flag does not need to be decoded, it can estimate the value of the normal merge flag based on the ``general_merge_flag'' which indicates whether the target block is inter-predicted by merging and the sub-block merge flag which indicates whether sub-block merging is applied, similar to the method described in non-patent document 1.
[0109] Regarding such an estimation method, the same method as that described in Non-Patent Document 1 can be used in this embodiment, and therefore a description thereof will be omitted.
[0110] The condition for determining whether or not decoding of the normal merge flag in step S7-1-1 is necessary is made up of a condition for determining whether or not CIIP is applicable and a condition for determining whether or not geometric block merging is applicable. - Criteria for determining whether CIIP is applicable: (1) The area of the target block is 64 pixels or more.
[0111] (2) Indicates that the CIIP enable flag at the SPS level is enabled (value is 1).
[0112] (3) The skip mode flag of the target block is disabled (its value is 0). (4) The width of the target block is less than 128 pixels.
[0113] (5) The height of the target block is less than 128 pixels. - Geometry block merging applicability criteria: (1) The area of the target block is 64 pixels or more.
[0114] (6) The SPS level geometric block division merge enable flag is enabled (value is 1).
[0115] (7) The maximum number of merge indexes that can be registered in the merge list for the geometric block division merge (hereinafter referred to as the maximum number of geometric block division merge candidates) is greater than 1.
[0116] (8) The width of the target block is 8 pixels or more.
[0117] (9) The height of the target block is 8 pixels or more.
[0118] (10) The slice type including the target block is a B slice (bi-predictive slice).
[0119] The merge mode identification unit 241A21 determines that the CIIP is applicable when all of the above conditional expressions (1) to (5) are satisfied in the CIIP applicability determination conditions, and otherwise determines that the CIIP is not applicable.
[0120] Note that the same conditional expressions described in Non-Patent Document 1 can be used in this embodiment for determining whether or not the CIIP is applicable, and therefore a description thereof will be omitted.
[0121] Furthermore, the merge mode specification unit 241A21 determines that the geometric block division merge is applicable when all of the above-mentioned conditional expressions (1) and (6) to (10) are satisfied in the conditions for determining whether the geometric block merge is applicable, and otherwise determines that the geometric block division merge is not applicable.
[0122] The conditions (1) and (6) to (10) will be described in detail later.
[0123] If the merge mode determination unit 241A21 determines whether or not the above-mentioned CIIP applicability determination condition or the geometric block division merge applicability determination condition is satisfied among the decoding necessity determination conditions of the normal merge flag in step S7-1-1, it proceeds to step S7-1-2, and if neither is satisfied, it proceeds to step S7-1-3.
[0124] In addition, the CIIP applicability determination condition may be omitted from the decoding necessity determination condition for the normal merge flag in step S7-1-1, which means that only the geometric block division merge applicability determination condition is added to the decoding necessity determination condition for the normal merge flag.
[0125] In step S7-1-2, the merge mode identification unit 241A21 decodes the normal merge flag, and the process proceeds to step S7-1-3.
[0126] In step S7-1-3, the merge mode identification unit 241A21 determines whether the value of the normal merge flag is 1, and if it is 1, proceeds to step S7-5, and if it is not 1, proceeds to step S7-2.
[0127] Next, the conditions for determining whether or not CIIP is applied in step S7-2 will be described with reference to FIG.
[0128] As shown in FIG. 10, in step S7-2-1, the merge mode identification unit 241A21 determines whether or not the CIIP flag needs to be decoded.
[0129] If the merge mode identification unit 241A21 determines that step S7-2-1 is satisfied, i.e., that decoding of the CIIP flag is necessary, it proceeds to step S7-2-2; if the merge mode identification unit 241A21 determines that step S7-2-1 is not satisfied, i.e., that decoding of the CIIP flag is not necessary, it proceeds to step S7-2-3.
[0130] Here, if the merge mode identification unit 241A21 determines that the CIIP flag does not need to be decrypted, it estimates the value of the CIIP flag as follows.
[0131] If all of the following conditions are met, the merge mode identification unit 241A21 treats the CIIP as valid, i.e., the value of the CIIP flag as 1; otherwise, it treats the CIIP as invalid, i.e., the value of the CIIP flag as 0.
[0132] (1) The area of the target block is 64 pixels or more.
[0133] (2) The normal merge flag is 0.
[0134] (3) Indicates that the CIIP enable flag at the SPS level is enabled (value is 1).
[0135] (4) The skip mode flag of the target block is enabled (its value is 0).
[0136] (5) The width of the target block is less than 128 pixels.
[0137] (6) The height of the target block is less than 128 pixels.
[0138] (12) “general_merge_flag” is 1.
[0139] (13) The subblock merge flag is 0.
[0140] The conditions for determining whether the CIIP flag needs to be decoded in step S7-2-1 are made up of the above-mentioned CIIP applicability determination condition and geometric block division / merging availability determination condition.
[0141] If both the CIIP applicability condition and the geometric block division mergeability determination condition are met, the merge mode identification unit 241A21 proceeds to step S7-2-2, and if either one is not met, the merge mode identification unit 241A21 proceeds to step S7-2-3.
[0142] Here, conditional expression (1) may be omitted from step S7-2-1 because it has already been determined that the condition is satisfied in the normal merge flag decoding necessity determination condition before entering the CIIP flag decoding necessity determination condition.
[0143] In step S7-2-2, the merge mode identification unit 241A21 decodes the CIIP flag, and the process proceeds to step S7-1-3.
[0144] In step S7-2-3, the merge mode identification unit 241A21 determines whether the value of the CIIP flag is 1, and if it is 1, proceeds to step S7-3, and if it is not 1, proceeds to step S7-4.
[0145] (Explanation of the conditions for determining whether or not geometric block division and merging are possible) The following describes in detail (1) and (6) to (10) relating to the geometric block division merge possibility determination conditions. (1) To limit the target blocks to which geometric block division merging can be applied to relatively large blocks, the area of the target block is set to 64 pixels or more. (6) A new flag is added to indicate whether or not geometric block division merging is applicable at the SPS level. If the flag is invalid, it can be determined that geometric block merging is not applicable to the target block. Therefore, the flag is added to the geometric block division merging applicability determination formula. (7) In order to avoid exceeding the worst-case memory bandwidth required for motion compensation prediction of inter-prediction blocks in Non-Patent Document 1, the lower limit of the width of the target block is set to 8 pixels or more. (8) In order to avoid exceeding the worst-case memory bandwidth required for motion compensation prediction of inter-prediction blocks in Non-Patent Document 1, the lower limit of the height of the target block is set to 8 pixels or more. (9) A target block to which geometric block partition merging is applicable has two different motion information across a partition boundary. Therefore, if the target block is included in a B slice, it can be determined that geometric block partition merging is applicable to the target block. On the other hand, if the target block is included in a slice other than a B slice, that is, if the target block cannot have two different motion vectors, it can be determined that geometric block partition merging is not applicable to the target block. (10) As described above, geometric block partition merge requires two different pieces of motion information. The merge list construction unit identifies (decodes) these two different pieces of motion information from the motion information associated with two different merge indexes registered in the merge list. Therefore, if the maximum number of geometric block partition merge candidates is designed or identified to be one or less, it can be determined that geometric block partition merge is inapplicable to the target block. Therefore, a parameter for calculating the maximum number of geometric block partition merge candidates is stored within the geometric block partitioning unit. This maximum number of geometric block partition merge candidates may use the same value as the maximum number of merge indexes that can be registered in the merge list for normal merge (hereinafter referred to as the maximum number of merge candidates). Alternatively, to use a different value, for example, a flag specifying how many candidates to reduce from the maximum number of merge candidates may be transmitted from the image encoding device 100 to the image decoding device 200 and calculated by decoding the flag.
[0146] [Modification example 1: Adding a condition based on block aspect ratio to the conditions for determining whether or not geometric block splitting and merging is possible] Hereinafter, a first modified example of the present invention will be described with reference to FIG. 11, focusing on differences from the first embodiment described above.
[0147] In this modified example 1, in order to further restrict the application of geometric block division merging to the target block, a judgment based on the block size (upper limit) or block aspect ratio of the target block may be added to the above-mentioned geometric block division merging feasibility determination conditions.
[0148] First, it is considered to add a conditional expression that sets the upper limit of each of the height and width of a target block identified as being applicable to geometric block division merging to, for example, 64 pixels or less.
[0149] The reason for setting the upper limit of the height and width to 64 pixels or less is to avoid violating the constraints imposed by maintaining the pipeline processing units of the image decoding device 200, called Virtual Pipeline Data Units (VPDUs), which are adopted in Non-Patent Document 1.
[0150] In Non-Patent Document 1, the size of the VPDU is set to 64×64 pixels, so the upper limit of the range to which geometric block division and merging can be applied is set to 64 pixels.
[0151] On the other hand, the designer may set an even smaller upper limit, for example, 32 pixels or less, at his / her discretion.
[0152] In geometric block division merging, when generating an MC predicted image signal, the MC predicted image signal is generated based on two different motion vectors that cross a geometric block division boundary, and is generated using a blending mask table that is weighted by the distance from the geometric block division boundary.
[0153] By reducing the upper limits of the width and height of the target block, the size of the blending mask table can be reduced, and it is desirable from an implementation standpoint to reduce the size of the blending mask table that needs to be stored in the memory of the image encoding device 100 and the image decoding device 200, thereby reducing the memory storage capacity.
[0154] Secondly, it is considered to add a conditional expression that specifies that the aspect ratio of a target block to which geometric block division and merging is applicable must be 4 or less, for example.
[0155] Such elongated rectangular blocks with an aspect ratio of 8 or more are unlikely to occur in natural images. Therefore, if the application of geometric block division merging to such rectangular blocks is prohibited, it becomes possible to reduce the number of variations in the division shapes of geometric block division (the number of variations in parameters that define the positions and directions of geometric block division boundaries within a rectangular block), which has the effect of reducing the storage capacity of parameters for defining these variations in the image encoding device 100 and the image decoding device 200.
[0156] (Method of defining geometric block division patterns in the geometric block division section) A method for defining a geometric block division pattern in the geometric block division unit 241A22 according to this embodiment will be described below with reference to Fig. 11. Fig. 11 is a diagram showing an example of a method for defining a geometric block division pattern according to this embodiment.
[0157] The geometric block division pattern may be defined, for example, as shown in Figure 11, by two parameters: the position of the geometric block division boundary line, i.e., the distance ρ from the center point (hereinafter referred to as the center point) of the rectangularly divided target block to the geometric block division boundary line (hereinafter referred to as the division boundary line), and the elevation angle φ from the horizontal direction of the perpendicular line from the center point to the division boundary line.
[0158] Furthermore, the combination of the distance ρ and the elevation angle φ may be transmitted from the image encoding device 100 (the merging unit 111A2) to the image decoding device 200 (the merging unit 241A2) using an index.
[0159] The more variations in the combinations of the distance ρ and the angle φ that define the geometric block division pattern, the greater the reduction in prediction error that can be expected. However, there is a trade-off between the increased processing time required to identify combinations of the distance ρ and the elevation angle φ and the increased code length of the index to indicate such combinations. Therefore, the designer may use quantized distance ρ and elevation angle φ as intended.
[0160] Here, the quantized distance ρ may be designed, for example, in units of a predetermined number of pixels, and the quantized elevation angle φ may be designed, for example, as an angle obtained by equally dividing 360 degrees.
[0161] [Change 2: Method of defining elevation angle ρ] Hereinafter, Modification 2 of the present invention will be described with reference to Fig. 12, focusing on differences from the above-mentioned first embodiment and Modification 1. Fig. 12 is a diagram showing a modification of the elevation angle φ that defines the above-mentioned geometric block division pattern.
[0162] In the above-mentioned modified example 1, as an example of a method for determining the elevation angle φ, a method for determining the elevation angle φ using angles obtained by equally dividing 360 degrees was shown. However, in this modified example, the elevation angle φ may also be determined using the aspect ratio and horizontal and vertical directions of the possible blocks of the target block, as shown in Figure 12(b).
[0163] For example, as shown in Fig. 12(c), the elevation angle φ can be expressed using an arctangent function. In this modified example, as shown in Fig. 12(a), an example is shown in which a total of 24 different elevation angles φ are defined.
[0164] [Modification 3: How to define the position of the dividing boundary line (distance ρ)] Hereinafter, Modification 3 of the present invention will be described with reference to Fig. 13, focusing on differences from the above-described first embodiment and Modifications 1 and 2. Fig. 13 is a diagram showing a modification of the position (distance ρ) that defines the above-described geometric block division pattern.
[0165] In the above-mentioned modified examples 1 and 2, the distance ρ was defined as the perpendicular distance from the center point of the target block to the division boundary line, but in this modified example, the distance ρ may also be defined as a predetermined distance (predetermined position) in the horizontal or vertical direction from the center point, including the center point, as shown in Figure 13.
[0166] In the example of FIG. 13, two distances ρ are defined depending on the block aspect ratio of the target block.
[0167] First, as shown in FIG. 13(a), the distance ρ may be defined only in the horizontal direction for a horizontally long block.
[0168] Secondly, as shown in FIG. 13(b), the distance ρ may be defined only in the vertical direction for a vertically elongated block.
[0169] However, the vertical and horizontal distances ρ may be defined as long as the elevation angle φ is 0 degrees (180 degrees) for horizontally long blocks and as long as the elevation angle is 90 degrees (270 degrees) for vertically long blocks.
[0170] For square blocks, the distance ρ may be defined in either the horizontal direction or the vertical direction, or in both directions. Also, the distance ρ may vary with respect to a predetermined distance (predetermined position) obtained by dividing the width or height of the target block into eight equal parts, as shown in Fig. 13, for example.
[0171] In the example of Figure 13, the distance ρ is defined as the distance (position) from the center point in the horizontal (left-right) or vertical (up-down) direction that is 0 / 8, 1 / 8, 2 / 8, or 3 / 8 times the width or height of the target block.
[0172] Here, the reason why the vertical and horizontal distances (distances in the short side direction) are not specified for elevation angles other than 0 degrees (180 degrees) for horizontal blocks and other than 90 degrees (270 degrees) for vertical blocks is that the effect of increasing the number of variations in the geometric block division pattern by specifying the distance ρ in the short side direction is smaller than the effect of increasing the number of variations in the geometric block division pattern by specifying the distance ρ in the long side direction.
[0173] As shown in FIG. 13, the division boundary line based on the geometric block division can be designed as a perpendicular line to the elevation angle φ that passes through the division boundary line and a point at the predetermined distance (predetermined position) ρ defined as described above.
[0174] Note that this modified example 3 can be combined with the above modified example 2, and for angles that are 180 degrees apart in elevation angle φ, the perpendicular to the above distance ρ is the same, that is, the division boundary line is the same.
[0175] Therefore, for example, the use of elevation angles φ in the range of 0 degrees or more and less than 180 degrees and elevation angles φ in the range of 180 degrees or more and less than 360 degrees can be restricted with respect to the distance ρ in the left-right direction or the up-down direction, with the center point being line-symmetrical, as shown in Figure 13, thereby making it possible to avoid duplication of geometric block division patterns.
[0176] The combination of the range of the elevation angle φ and the distance ρ in the horizontal (left-right) and vertical (up-down) directions will still show variations of the same geometric block division pattern even if the combination is reversed, so the designer may freely change the implementation as desired.
[0177] [Change Example 4: Change in the method of defining the elevation angle φ and position (distance) ρ, reducing variations] 14 and 15, a fourth modified example of the present invention will be described below, focusing on differences from the first embodiment and the first to third modified examples. Fig. 14 is a diagram showing a modified example relating to the method of defining the elevation angle φ and the distance ρ.
[0178] In the above-mentioned modified example 2, a method for determining the elevation angle φ using the block aspect ratio was explained, and in the above-mentioned modified example 3, a method for determining the distance (position) ρ using a predetermined distance (predetermined position) in the horizontal and vertical directions from the center point was explained.
[0179] In this modified example, in order to simplify the implementation of the elevation angle φ and distance (position) ρ, as shown in Figures 14(a) to (c), the line that passes through a point at a predetermined distance (position) ρ and is indicated by the elevation angle φ may be defined as the division boundary line itself.
[0180] Furthermore, in order to reduce the geometric block division patterns, the variations in distance ρ and elevation angle φ may be reduced as shown in FIGS. 14(a) to 14(c).
[0181] Specifically, in FIGS. 14(a) to 14(c), the variations of the distance ρ shown in the above-mentioned third modification example are limited to two types: 0 / 8 and 2 / 8 times the width or height of the target block.
[0182] 14(a) to 14(c), the possible values of the elevation angle φ are limited according to the position where the division boundary line passes and the block aspect ratio of the target block. As a result, in the examples of Figures 14(a) to 14(c), the variations of the geometric block division pattern are limited to a total of 16 ways.
[0183] (Method for identifying geometric block division patterns) A method for identifying a geometric block division pattern will be described below with reference to Fig. 15. Fig. 15 is a diagram showing an example of an index table showing combinations of elevation angles φ and distances (positions) ρ that define the geometric block division patterns described above.
[0184] In FIG. 15, the elevation angle φ and distance ρ that define the geometric block division pattern are associated with each other by "angle_idx" and "distance_idx", and further, the combination of these is defined by "partition_idx".
[0185] For example, the elevation angle φ and distance ρ shown in the above-mentioned modified example 2 and modified example 3 can be defined as integers from 0 to 23 and 0 to 3, respectively, using "angle_idx" and "distance_idx" as shown in the table of Figure 12, and by decoding "partition_idx" which indicates this combination, the elevation angle φ and distance ρ can be uniquely determined, thereby making it possible to uniquely identify the geometric block partition pattern.
[0186] The image encoding device 100 and the image decoding device 200 each have an index table indicating such geometric block partitioning patterns, and the image encoding device 100 transmits to the image decoding device 200 a "partition_idx" corresponding to the geometric block partitioning pattern that has the smallest encoding cost when geometric block partitioning is enabled, and the geometric block partitioning unit 241A22 of the image decoding device 200 decodes this "partition_idx".
[0187] [Modification example 5: Controlling the decoding method of "partition_idx" according to the size or aspect ratio of the target block] Hereinafter, Modification 5 of the present invention will be described with reference to Fig. 16, focusing on differences from the above-described first embodiment and Modifications 1 to 4. Fig. 16 is a diagram for explaining an example of control of the decoding method of "partition_idx" according to the block size or block aspect ratio of the target block.
[0188] In the above example, as an example of a method for identifying a geometric block partitioning pattern, a decoding method of "partition_idx" that does not depend on the block size or aspect ratio of the target block (a method for identifying a geometric block partitioning pattern using a fixed index table) was shown.
[0189] On the other hand, in this modification, the decoding method of "partition_idx" can be controlled according to the block size or aspect ratio of the current block, and the coding efficiency can be further improved as follows.
[0190] When focusing on the block size or aspect ratio of the target block, as shown in Figure 16, even if the number of variations in the dividing boundary lines is the same for a relatively small block such as an 8x8 pixel block and a relatively large block such as a 32x32 pixel block, the density of the dividing boundary lines relative to the area of the block may differ.
[0191] Similarly, as shown in FIG. 16, there may be cases where the density of the division boundary lines differs between horizontal and vertical directions of horizontally long blocks and vertically long blocks.
[0192] In such a case, for example, the index table used to identify the elevation angle φ and the distance ρ from “partition_idx” may be changed depending on the block size or block aspect ratio of the target block.
[0193] For example, for target blocks with small block sizes, an index table with a small number of geometric block partition patterns, i.e., a small maximum value of "partition_idx", is used, and for target blocks with large block sizes, an index table with a large number of geometric block partition patterns, i.e., a large maximum value of "partition_idx", is used, thereby improving coding efficiency compared to the above example using a fixed index table.
[0194] Furthermore, the correlation between the size of the block and the number of geometric block division patterns (proportional relationship between the block size and the number of geometric block division patterns) may be reversed.
[0195] In other words, for target blocks with small block sizes, an index table with a large number of geometric block division patterns, i.e., a large maximum value of "partition_idx", may be used, and for target blocks with large block sizes, an index table with a small number of geometric block division patterns, i.e., a small maximum value of "partition_idx", may be used.
[0196] When the block size of the target block and the number of geometric block division patterns are proportional to each other, the density of the division boundary lines can be made uniform even if the block sizes of the target blocks are different.
[0197] Therefore, it is expected that the probability of aligning the division boundary line generated by the geometric block division with the object boundary occurring within the target block can be made uniform for each block size.
[0198] On the other hand, if the block size of the target block and the number of geometric block division patterns are inversely proportional, the above-mentioned effect cannot be expected, but it is possible to expect an effect of further increasing the probability of aligning the division boundary line with the object boundary, limited to small block sizes.
[0199] Generally, in natural images, areas where object boundaries run in complex directions (appearing in multiple arbitrary directions) are often coded using small blocks, while areas where object boundaries run nearly horizontally or vertically are often coded using large blocks because the error with rectangular block division is small.
[0200] Therefore, based on this idea, by increasing the number of geometric block division patterns for small blocks, it is expected that the probability of aligning the division boundary line with the object boundary will be increased, and as a result, it is expected that the prediction error will be reduced.
[0201] Similarly, for horizontally long blocks, the number of geometric block division patterns may be increased in the horizontal direction, and conversely, the number of geometric block division patterns may be decreased in the vertical direction.
[0202] On the other hand, for vertically long blocks, the number of geometric block division patterns may be increased in the horizontal direction, and conversely, the number of geometric block division patterns may be decreased in the vertical direction.
[0203] Furthermore, the correlation between the size of the block aspect ratio and the number of geometric block division patterns may be reversed, as explained above in relation to the size of the block size and the number of geometric block division patterns.
[0204] In other words, for a horizontally long block, the number of geometric block division patterns may be reduced in the horizontal direction, and conversely, the number of geometric block division patterns may be increased in the vertical direction.
[0205] On the other hand, for vertically long blocks, the number of geometric block division patterns may be reduced in the horizontal direction, and conversely, the number of geometric block division patterns may be increased in the vertical direction.
[0206] The reason for reversing the correlation in this way is the same as the idea shown in the explanation of the correlation between the block size and the number of geometric block divisions, and therefore the explanation will be omitted.
[0207] Note that this method of controlling the number of geometric block division patterns according to the block aspect ratio can be realized by changing the index table to be used, as in the above case.
[0208] Another possible implementation example is to limit the range of "partition_idx" values that can be decoded on the index table depending on the block size or aspect ratio of the target block, while the index table itself is fixed.
[0209] This does not contribute to improving the coding efficiency because the code length state is not shortened, but from the perspective of the image coding device 100, it can contribute to speeding up the coding process because it allows the process up to cost calculation for unnecessary geometric block division patterns to be omitted.
[0210] Second Embodiment Hereinafter, the second embodiment of the present invention will be described with reference to FIGS. 17 to 22, focusing on the differences from the first embodiment described above.
[0211] (Motion information availability check process) The motion information availability confirmation process in this embodiment will be described below.
[0212] As described above, the motion information availability confirmation process is the first processing step constituting the merge list construction process.
[0213] Specifically, the motion information availability confirmation process checks whether motion information is available in a reference block that is spatially or temporally adjacent to the target block. Here, the method for checking motion information is omitted here because the known method described in Non-Patent Document 1 can be used in this embodiment.
[0214] (Registering movement information to the merge list and pruning processing) Hereinafter, the process of registering and pruning motion information in the merge list according to this embodiment will be described with reference to FIG.
[0215] FIG. 17 is a flowchart showing an example of the process of registering and pruning motion information in the merge list according to this embodiment.
[0216] As shown in FIG. 17, the process of registering and pruning motion information in a merge list according to this embodiment may be configured by a total of five processes of registering and pruning motion information, similar to Non-Patent Document 1.
[0217] Specifically, the motion information registration and pruning process may include spatial merging in step S14-1, temporal merging in step S14-2, history merging in step S14-3, pairwise average merging in step S14-4, and zero merging in step S14-5. Details of each process will be described later.
[0218] 18 is a diagram showing an example of a merge list constructed as a result of the process of registering and pruning motion information to the merge list. As described above, a merge list is a list in which motion information corresponding to a merge index is registered.
[0219] Here, the maximum number of merge indexes is set to 5 in Non-Patent Document 1, but may be set freely according to the designer's intention.
[0220] Also, mvL0, mvL1, RefIdxL0, and RefIdxL1 in FIG. 18 indicate the motion vectors and reference image indexes of the reference image lists L0 and L1, respectively.
[0221] Here, the reference image lists L0 and L1 indicate lists in which reference frames are registered, and the reference frames are identified by RefIdx.
[0222] Note that although the merge list shown in FIG. 18 indicates motion vectors and reference image indices for both L0 and L1, some reference blocks may be unidirectionally predicted.
[0223] In this case, one (unipredictive) motion vector and one reference image index are registered in the list. Note that, before each of the processes in steps S15-1 to S15-4, if it is confirmed by the above-mentioned motion information availability confirmation process that there is no motion information in the reference block, each process is skipped.
[0224] (Spatial Merge) FIG. 19 is a diagram illustrating the spatial merge.
[0225] Spatial merging is a technique for inheriting mv, RefIdx, and hpelIfIdx from adjacent blocks that exist in the same frame as the target block.
[0226] Specifically, the merge list construction unit 241A23 is configured to inherit the above-mentioned mv, Refidx, and hpelfIdx from adjacent blocks that are in the positional relationship shown in Fig. 15 and register them in the merge list. The processing order may be B1->A1->B0->A0->B2, as in Non-Patent Document 1.
[0227] (Motion information pruning process) The process of registering motion information to the merge list may be performed in the order described above, but a motion information pruning process may also be implemented, as in non-patent document 1, to prevent motion information identical to motion information already registered in the merge list from being registered.
[0228] The purpose of implementing the motion information pruning process is to increase the variety of motion information registered in the merge list, and from the perspective of the image encoding device 100, the purpose is to be able to select motion information that has the lowest specified cost according to the image characteristics.
[0229] On the other hand, the image decoding device 200 can generate a prediction signal with high prediction accuracy using the motion information selected by the image encoding device 100, and as a result, an improvement in encoding performance can be expected.
[0230] In the motion information pruning process for spatial merge, for example, when motion information corresponding to B1 is registered in the merge list, the identity of the registered motion information corresponding to B1 is confirmed when the motion information corresponding to A1 in the next processing order is registered. Here, confirming the identity of the motion information means comparing whether mv and RefIdx are the same. Here, if the identity is confirmed, the motion information corresponding to A1 is not registered in the merge list.
[0231] The motion information corresponding to B0 is compared with the motion information corresponding to B1, the motion information corresponding to A0 is compared with the motion information corresponding to A1, and the motion information corresponding to B2 is compared with the motion information corresponding to B1 and A1. If there is no motion information corresponding to the position to be compared, the confirmation of identity may be skipped and the corresponding motion information may be registered.
[0232] Furthermore, in Non-Patent Document 1, the maximum number of merge indices that can be registered by spatial merging is set to 4, and if four pieces of motion information have already been registered in the merge list for adjacent block B2 by previous spatial merging processes, the processing of adjacent block B2 is skipped. In this embodiment, similar to Non-Patent Document 1, the processing of adjacent block B2 may be determined based on the number of existing registered motion information.
[0233] (Time Merge) FIG. 20 is a diagram illustrating the time merge.
[0234] Temporal merging is a technique in which a block that exists in a different frame from the target block but is located at the same position (C0 in Figure 17) or a block located at the same position (C1 in Figure 17) is identified as a reference block, and the motion vector and reference image index are inherited.
[0235] In Non-Patent Document 1, the maximum number of motion information pieces that can be registered in a merge list using a temporal merge list is 1, and in this embodiment, the same value may be used, or may be changed as desired by the designer.
[0236] Furthermore, the motion vectors inherited in the temporal merge are scaled, as shown in Figure 21.
[0237] Specifically, as shown in FIG. 21, mv of the reference block is scaled as follows based on the distance tb between the reference frame of the target block and the frame in which the target frame exists and the distance td between the reference frame of the reference block and the reference frame of the reference block.
[0238] mv'=(td / tb)×mv In the temporal merge, this scaled mv' is registered in the merge list as the motion vector corresponding to the merge index.
[0239] (History Merge) FIG. 22 is a diagram illustrating history merging.
[0240] History merge is a technique in which motion information held by inter-prediction blocks that have been coded earlier than the target block is separately recorded in a recording area called a history merge table, and if the number of motion information registered in the merge list has not reached the maximum number at the end of the spatial merge and temporal merge processes, the motion information registered in the history merge table is sequentially registered in the merge list.
[0241] FIG. 22 is a diagram showing an example of a process in which motion information is registered in a merge list using a history merge table.
[0242] The movement information registered in this history merge table is managed by a history merge index, and the maximum number of registrations is set to 6 in Non-Patent Document 1, but a similar value may be used or may be changed according to the designer's intention.
[0243] Furthermore, in Non-Patent Document 1, the registration process of motion information in this history merge table employs FIFO processing. In other words, when the number of motion information registered in the history merge table reaches the maximum number of registrations, the motion information linked to the last registered history merge index is deleted, and motion information linked to new history merge indexes is registered sequentially.
[0244] As in Non-Patent Document 1, the motion information registered in the history merge table may be initialized (all motion information may be deleted from the history merge table) when the current block crosses CTUs.
[0245] (pairwise average merge) Pairwise average merging is a technique for generating new motion information using motion information associated with two pairs of merge indexes already registered in a merge list, and registering the new motion information in the merge list.
[0246] As for the two sets of merge indexes used for pairwise average merging, as in Non-Patent Document 1, the 0th and 1st merge indexes registered in the merge list may be used as fixed, or the designer may change them to a different combination of two sets.
[0247] In pairwise average merging, new motion information is generated by averaging the motion information associated with two pairs of merge indexes already registered in the merge list.
[0248] Specifically, for example, if there are two motion vectors corresponding to two sets of merge indexes (i.e., in the case of bi-prediction), the pairwise average merge motion vectors mvL0Avg / mvL1Avg are calculated as follows using the motion vectors mvL0P0 / mvL0P1 and mvL1P0 / mvL1P1 in the L0 and L1 directions, respectively:
[0249] mvL0Avg=(mvL0P0+mvL0P1) / 2 mvL1Avg=(mvL1P0+mvL1P1) / 2 Here, if either mvL0P0 and mvL1P0 or mvL0P1 and mvL1P1 does not exist, the non-existent vector is treated as a zero vector and calculated as described above.
[0250] In this case, Non-Patent Document 1 specifies that the reference image index RefIdxL0P0 / RefIdxL0P1 associated with the merge index P0 should always be used as the reference image index associated with the pairwise average merge index.
[0251] (Zero Merge) Zero merge is a process of adding a zero vector to the merge list if the number of registered motion information in the merge list has not reached the maximum number at the time when the pairwise average merge process described above is completed. Regarding the registration method, a known method described in Non-Patent Document 1 can be used in this embodiment, so a description thereof will be omitted.
[0252] (Motion information decoding process from merge list) The merge list construction unit 241A23 is configured to decode motion information from the merge list constructed after the above-described motion information registration and pruning process is completed.
[0253] For example, the merge list constructor 241A23 is configured to select, from the merge list, motion information corresponding to the merge index transmitted from the merge list constructor 111A23 and decode it.
[0254] On the other hand, although not shown, the merge index selection method in merge list constructor 111A23 is configured to transmit motion information that results in the smallest coding cost to merge list constructor 241A23.
[0255] Here, the merge index that is registered in the merge list earlier, i.e., the merge index with a smaller index number, has a shorter code length (lower coding cost), so the merge list construction unit 111A23 tends to select the merge index with a smaller index number.
[0256] Here, in the merge list construction unit 241A23, when geometric block division merging is disabled, one merge index is decoded for the target block, but when geometric block merging is enabled, two regions m / n exist for the target block that straddle the geometric block division boundary, so two different merge indexes m / n are decoded.
[0257] For example, by decoding such a merge index m / n as follows, even if the merge index for m / n transmitted from the merge list construction unit 111A2 is the same, it is possible to assign a different merge index to m / n.
[0258] m=merge_idx0[xCb][yCb] n=merge_idx1[xCb][yCb]+(merge_idx1[xCb][yCb]>=m)?1:0 Here, xCb and yCb are position information of the pixel value located at the top left corner of the target block.
[0259] Note that merge_idx1 does not need to be decoded if the maximum number of merge candidates for a geometric block partition merge is 2 or less. This is because, when the maximum number of merge candidates for a geometric block partition merge is 2 or less, the merge index in the merge list corresponding to merge_idx1 is identified as another merge index different from the merge index in the merge list selected by merge_idx0.
[0260] [Modification Example 6: Controlling merge list construction when geometric block division merge is enabled] Hereinafter, Modification 6 of the present invention will be described with reference to Figures 23 and 24, focusing on differences from the above-described second embodiment. Specifically, with reference to Figures 23 and 24, control of merge list construction when geometric block division merging is enabled according to this modification will be described.
[0261] 23 and 24 are diagrams showing an example of a geometric block division pattern in a target block when geometric block division merge is enabled, and an example of the positional relationship between spatial merge and temporal merge for the target block.
[0262] If the target block has the geometric block division pattern shown in Figure 23, the adjacent blocks with the closest motion information for the two regions m / n across the block division boundary line are B2 for m and B1 / A1 / A0 / B0 / C1 for n.
[0263] Furthermore, if the target block has the geometric block division pattern shown in Figure 21, the adjacent blocks with the closest motion information for the two regions m / n across the division boundary are C1 for m and B1 / A0 / A1 / B0 / B2 for n.
[0264] When geometric block division is effective, it is expected that prediction accuracy will be improved by making it easier to register motion information of the adjacent block that is closest to the above m / n in the merge list.
[0265] Furthermore, from the perspective of improving prediction accuracy, lowering the registration priority of motion information at similar spatial positions, such as B1 / B0 and A1 / A0, in the merge list and adding motion information at different spatial positions, such as B2 and C1, to the merge list can be expected to further improve prediction accuracy.
[0266] In addition, as described above, if the motion information of the nearest neighboring block is registered in the merge list as a merge index with a small index number, the coding efficiency can be improved.
[0267] In the above configuration, for example, if the motion information for which the registration priority in the merge list is to be changed is motion information for B0 and A0, even if the motion information availability confirmation process confirms that motion information exists at these positions, the motion information is treated as unavailable, thereby increasing the probability of registering motion information corresponding to B2 and C1 following B0 and A0 in the merge list.
[0268] [Modification example 7: How to change the registration priority of other movement information] Modification 7 of the present invention will be described below, focusing on the differences from the above-mentioned second embodiment and Modification 6. In the above-mentioned Modification 6, the process of changing the registration priority of motion information when geometric block division merging is enabled is performed by treating motion information at a predetermined position in the spatial merge process, for example, motion information of B0 and A0, as unavailable even if motion information exists or if it is confirmed that the information is not identical to already registered motion information, thereby indirectly increasing the probability (priority) of registering subsequent spatially or temporally adjacent motion information, for example, motion information of B2 or C1 (C0).
[0269] Meanwhile, a method for changing the registration priority of other movement information will be described below. For example, the order of registering motion information may be changed directly by decomposing the processing order of the spatial merge in step S14-1 and the temporal merge in step S14-2 in the merge list shown in Fig. 17. For example, when the geometric block division merge is applied, the motion information C1 or C0 (hereinafter referred to as Col) registered in the spatial merge and the temporal merge is executed before B0 and A0 during the spatial merge process.
[0270] For example, the following two implementation examples are possible.
[0271] Implementation example 1: A method of branching the merge list between normal merge and merge list depending on whether or not geometric block split merge is applied. if ( !merge_geo_flag ) i = 0 if( availableFlagB1 ) mergeCandList[ i++ ] = B1 if( availableFlagA1 ) mergeCandList[ i++ ] = A1 if( availableFlagB0 ) mergeCandList[ i++ ] = B0 if( availableFlagA0 ) mergeCandList[ i++ ] = A0 if( availableFlagB2 ) mergeCandList[ i++ ] = B2 if( availableFlagC1 ) mergeCandList[ i++ ] = Col else i = 0 if( availableFlagB1 ) mergeCandList[ i++ ] = B1 if( availableFlagA1 ) mergeCandList[ i++ ] = A1 if( availableFlagB2 ) mergeCandList[ i++ ] = B2 if( availableFlagC1 mergeCandList[ i++ ] = Col if( availableFlagB0 ) mergeCandList[ i++ ] = B0 if( availableFlagA0 ) mergeCandList[ i++ ] = A0 Here, merge_geo_flag is an internal parameter that holds the result of determining whether or not geometric block division merging is applied. When this parameter is 0, geometric block division merging is not applied, and when this parameter is 1, it is applied.
[0272] The availableFlag is an internal parameter that holds the result of the motion information availability check process for each spatial and temporal merge candidate. A value of 0 indicates that motion information is not available, and a value of 1 indicates that motion information is available.
[0273] Also, mergeCandList indicates the registration process of motion information at each position. The first if statement determines whether or not geometric block division merging is applied using merge_geo_flag. If it is determined not to be applied, the process enters the same merge list construction process as normal merging. If it is determined to be applied, the process enters the merge list construction process for geometric block division merging, which replaces the normal merge with the motion information availability check process and motion information registration / pruning process.
[0274] Implementation example 2: A method for partially sharing the merge list construction between normal merge and geometric block split merge i = 0 if( availableFlagB1 ) mergeCandList[ i++ ] = B1 if( availableFlagA1 ) mergeCandList[ i++ ] = A1 if( availableFlagB0 && !merge_geo_flag ) mergeCandList[ i++ ] = B0 if( availableFlagA0 && !merge_geo_flag) mergeCandList[ i++ ] = A0 if( availableFlagB2 ) mergeCandList[ i++ ] = B2 if( availableFlagC1 ) mergeCandList[ i++ ] = Col if( availableFlagB0 && merge_geo_flag ) mergeCandList[ i++ ] = B0 if( availableFlagA0 && merge_geo_flag) mergeCandList[ i++ ] = A0 Although this Implementation Example 2 increases the total number of processing stages compared to Implementation Example 1, it has the advantage that, because the merge list for geometric block division merging is partially shared with the normal merge list, it does not need to have resources for a completely independent merge list construction processing circuit, as in Implementation Example 1.
[0275] In the above configuration, the registration priority of B2 and Col is set higher than that of B0 and A0, but the registration priority of B2 and Col may be set higher than that of B1 and A1. As another configuration example, for example, either B2 or Col may be set higher in priority than that of B0 and A0, or higher in priority than that of B1 and A1.
[0276] Note that if B2 is registered in the merge list before B1, A1, B0, and A0, the above-described confirmation of identity (comparison condition) of B2 with the motion information corresponding to B1, A1, B0, and A may be eliminated. On the other hand, confirmation of identity with the motion information corresponding to B2 may be added when registering the motion information corresponding to B1 and A1.
[0277] Furthermore, although the above describes a method for changing the registration priority of motion information while the merge list is being constructed, the registration priority may also be changed after the merge list is constructed using the method described below.
[0278] Specifically, when motion information is registered in the merge list, the merge process and the position corresponding to the registered motion information are stored as internal parameters until the motion vector is decoded (selected). After the merge list is constructed, the merge index numbers of motion information corresponding to a specific merge process and a specific position, for example, the motion information registered by spatial merges B0 and A0, can be reordered with the merge index numbers of the temporal merge Col, thereby giving the temporal merge Col a higher priority than the spatial merges B0 and A0 (association with a smaller merge index number is possible). A similar method may be used to realize the above-described example of reordering the registration priority of indirect or direct motion information.
[0279] As described above, as a secondary effect of changing the registration priority of motion information after the merge list construction process, the merge list construction process when geometric block division merging is enabled can be made common with the merge list construction process for normal merging.
[0280] Specifically, the processing order of each merge process in the motion availability confirmation process within the merge list construction process in normal merge and geometric block division merge, and the motion information registration and pruning process can be standardized.
[0281] [Modification Example 8: Changing the merge list construction order according to block size] The eighth modified example of the present invention will be described below, focusing on the differences from the second embodiment and the sixth and seventh modified examples.
[0282] In the above-described sixth and seventh modifications, the configuration has been shown in which the criterion for determining whether to change the registration priority of motion information is determined based on whether or not geometric block division merging is applied.
[0283] On the other hand, in this modification, the determination criterion may be based on the geometric block division pattern, or may be based on the block size and aspect ratio of the target block.
[0284] According to this configuration, by checking even the geometric block division pattern and changing the registration priority of the motion information in the merge list, it is possible to expect a further improvement in prediction accuracy.
[0285] (Method of selecting (decoding) two different motion information) A method for selecting (decoding) two different pieces of motion information when geometric block division merging is enabled according to this embodiment will be described below.
[0286] As described above, when geometric block partition merge is enabled, the current block has two different pieces of motion information across a partition boundary. Furthermore, as described above, the encoding device 100 selects the two different pieces of motion information with the lowest coding cost from the merge list, and the merge index numbers of the merge list are specified using two merge indices, merge_geo_idx0 and merge_geo_idx1, for each partition region m / n. The merge index numbers are then transmitted to the image decoding device 200, and the corresponding motion information is decoded.
[0287] On the other hand, if two pieces of motion information are registered for a specified merge index number, specifically if the adjacent blocks registered in the above-mentioned merge process are bi-predictive (have motion information in both L0 and L1), it is possible that the two pieces of motion information are registered for one merge index, so one of them must be selected.
[0288] An example of the selection method will be described below.
[0289] For example, one configuration example is a method of setting the priority of motion information to be decoded in advance in the order of merge index numbers in the merge list. Specifically, in the order of processing merge index numbers, even numbers 0, 2, and 4 prioritize decoding of motion information registered in merge index L0, and odd numbers 1, 3, and 5 prioritize decoding of motion information registered in L1.
[0290] In the above configuration example, if there is no motion information for L0 or L1 corresponding to each numerical order, the existing motion information may be decoded. Also, L0 and L1 for even and odd numbers may be reversed.
[0291] Other configuration examples are also shown below. For example, for a target frame including a target block, the distances of reference frames indicated by reference indexes corresponding to L0 and L1 may be compared, and motion information including reference frames with closer distances may be preferentially decoded.
[0292] As described above, when geometric block partition merge is enabled, if two different motion information are registered in the merge list for the merge index of the partition area, the motion information including the reference frame that is close to the target frame is decoded preferentially, which is expected to reduce prediction errors.
[0293] In addition, if the distance difference between the reference frame and the target frame for two different motion information is the same, as described above, a predetermined priority can be set for the merge index numbers in the merge list, i.e., even numbers 0, 2, and 4 prioritize motion information registered in L0, and odd numbers 1, 3, and 5 prioritize motion information registered in L1.
[0294] Depending on the picture order count (POC) of the target frame, the distance difference between the target frame and the reference frames included in L0 and L1 may be the same, and therefore the priority may be determined as described above.
[0295] Alternatively, the merge list may be selected using the merge index number opposite to the list selected using the previous index number. For example, if L0 is selected at merge index 0, the opposite list, L1, may be selected for merge index 1 in the merge list. Note that if the frame distance difference is the same at merge index 0 in the merge list, L0 may be referenced.
[0296] In the above example, priority is given to decoding of motion information for which the distance between the target frame and the reference frame is short, but conversely, motion information for which the distance between the target frame and the reference frame is farther may be decoded.
[0297] Furthermore, the decoding priority of the motion information may be changed for each of the two merge indexes merge_geo_idx0 and merge_geo_idx1 for each divided region m / n, such that the motion information for the closer distance and the motion information for the farther distance are decoded.
[0298] (Prediction signal generation process for geometric block division merging) A prediction signal generation processing method when geometric block division merging is enabled according to this embodiment will be described below.
[0299] As mentioned above, when geometric block partition merging is enabled, the target block has two different motion information across the partition boundary. In this case, by weighting (blending) the motion compensation prediction signals generated for the target block based on these two different motion vectors with a weight that depends on the distance from the partition boundary, a smoothing effect of pixel values along the partition boundary can be expected.
[0300] For example, the image encoding device 100 and the image decoding device 200 have one weight table (blending table) for one geometric block partitioning pattern, and if the geometric block partitioning pattern can be identified, blending processing suitable for the geometric block partitioning pattern can be realized.
[0301] According to the above-described embodiment, by applying geometric block division merging to a target block that has been divided into rectangular segments, an appropriate block division shape can be selected for object boundaries that appear in any direction, thereby achieving an improved coding performance by reducing prediction errors and an improved subjective image quality by selecting appropriate block division boundaries for object boundaries.
[0302] In addition, the merge list construction unit 111A2 may be configured to treat the above-mentioned motion information as unusable during the motion information usability confirmation process, regardless of whether motion information exists for the specified merge process, depending on whether geometric block division merging is applied or not.
[0303] In addition, the merge list construction unit 111A2 may be configured to treat motion information for a specified merge process as a target for pruning during the motion information registration / pruning process, regardless of its identity with motion information already registered in the merge list, depending on whether geometric block division merging is applied or not, i.e., not to newly register it in the merge list.
[0304] The image encoding device 100 and the image decoding device 200 described above may be realized as a program that causes a computer to execute each function (each step).
[0305] In each of the above-described embodiments, the present invention has been described as being applied to the image encoding device 100 and the image decoding device 200, but the present invention is not limited to this and can be similarly applied to image encoding systems and image decoding systems that have the functions of the image encoding device 100 and the image decoding device 200. [Explanation of symbols]
[0306] 10...Image processing system 100...Image encoding device 111, 241...Inter prediction section 111A…mv derivation part 111A1, 241A1…AMVP section 111A2, 241A2...Merge section 111B, 241B...mv refinement department 111C, 241C...Prediction signal generation unit 111A21, 241A21...Merge mode specification section 111A22, 241A22...Geometric block division section 111A23, 241A23...Merge list construction section 112, 242...Intra prediction section 121...Subtractor 122, 230...adder 131...Transformation and quantization unit 132, 220...Inverse transform and inverse quantization units 140...encoding section 150, 250...In-loop filter processing section 160, 260...frame buffer 200...Image decoding device 210...Decoding unit
Claims
1. An image decoding device, a merging unit configured to apply geometric block division merging to a target block divided into rectangles, the merging unit includes a merge mode specifying unit configured to specify whether or not the geometric block partition merging is applied; The merge mode identification unit prohibiting application of the geometric block division merge to the target block if the block aspect ratio of the target block is 8 or more; An image decoding device, characterized in that, when the block aspect ratio of the target block is less than 8, application of the geometric block division and merging to the target block is not prohibited.
2. 1. An image decoding method comprising a step of applying geometric block division merging to a target block that has been divided into rectangles, The steps include a step A of specifying whether or not the geometric block division merge is applied, The step A comprises: prohibiting application of the geometric block division merge to the target block if the block aspect ratio of the target block is 8 or more; 10. An image decoding method, comprising: when a block aspect ratio of the target block is less than 8, not prohibiting application of the geometric block division and merging to the target block.
3. A program that causes a computer to function as an image decoding device, the image decoding device includes a merging unit configured to apply geometric block partition merging to a target block that has been rectangularly partitioned; the merging unit includes a merge mode specifying unit configured to specify whether or not the geometric block partition merging is applied; The merge mode identification unit prohibiting application of the geometric block division merge to the target block if the block aspect ratio of the target block is 8 or more; A program that does not prohibit application of the geometric block division and merging to the target block when the block aspect ratio of the target block is less than 8.
Citation Information
Patent Citations
Simplified inter prediction with geometric partitioning
WO2021104433A1