Image decoding device, image decoding method, and image decoding program

By incorporating a motion information derivation unit and a code string decoding unit in the image decoding apparatus, the HEVC image encoding technology's encoding efficiency is improved through effective motion vector correction in the merge mode.

JP2025083536AActive Publication Date: 2025-05-30JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025042224
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-13
Filing Date
2025-03-17
Publication Date
2025-05-30
Estimated Expiration
2039-09-20

AI Technical Summary

Technical Problem

The existing HEVC image encoding technology has limitations in encoding efficiency due to the reliance on the merge mode for inter prediction, where the motion vector correction is not effectively utilized.

Method used

An image decoding apparatus that includes a motion information derivation unit for scaling motion vectors from adjacent blocks, a merge candidate list generation unit, a merge candidate selection unit, and a code string decoding unit to correct the motion vector of the merge mode based on specific conditions, such as the use of long-term reference pictures.

Benefits of technology

This approach enhances the encoding efficiency of the inter prediction mode by allowing for more precise motion vector correction, thereby improving the overall performance of image decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083536000001_ABST
    Figure 2025083536000001_ABST
Patent Text Reader

Abstract

To provide an inter prediction mode with higher efficiency by correcting a motion vector in a merge mode.SOLUTION: In an image decoding method, a merge candidate list is generated; a merge candidate is selected from the merge candidate list as a selected merge candidate; a code string is decoded from an encoded stream to derive a correction vector; the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling; and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the correction merge candidate.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image decoding technology.

Background Art

[0002] There is image encoding technology such as HEVC (H.265). In HEVC, the merge mode is used as an inter prediction mode.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In HEVC, there are a merge mode and a differential motion vector mode as inter prediction modes, but the inventor has recognized that there is room to further improve the encoding efficiency by correcting the motion vector of the merge mode.

[0005] The present invention has been made in view of such a situation, and an object thereof is to provide a new inter prediction mode with higher efficiency by correcting the motion vector of the merge mode.

Means for Solving the Problems

[0006] To solve the above problems, an image decoding apparatus according to an aspect of the present invention includes motion information including a motion vector derived by scaling the motion information of a plurality of blocks adjacent to a prediction target block and the motion vector of a block at the same position on the decoded image as the prediction target block. ​​​​​​​A merge candidate list generation unit that generates a merge candidate list including merge candidates, and the merge A merge candidate selection unit that selects a merge candidate from the merge candidate list as a selected merge candidate, and a code A code string decoding unit that decodes a code string from a coded stream and derives a merge correction flag indicating whether to correct the merge candidate and a correction vector, and when the merge correction flag indicates that the merge candidate is to be corrected If either one or both of the reference pictures of the first prediction of the selected merge candidate or the reference pictures of the second prediction of the selected merge candidate are long-term reference pictures, and the reference pictures of the first prediction of the selected merge candidate and the reference pictures of the second prediction of the selected merge candidate Are in opposite directions with respect to the decoded picture including the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive a corrected merge candidate for dual prediction And a merge candidate correction unit. When the merge correction flag indicates that the merge candidate is to be corrected, the number of the selected merge candidates that can be selected is smaller than when it does not indicate that the merge candidate is to be corrected.

[0007] In addition, any combination of the above components, and those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc. are also effective as aspects of the present invention.

[0007]

[0008]

Advantages of the Invention

[0008] According to the present invention, it is possible to provide a new inter prediction mode with higher efficiency.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Embodiments for Carrying Out the Invention

[0010] [First Embodiment] Hereinafter, an image encoding apparatus, an image encoding method , and an image encoding program, as well as an image decoding apparatus, an image decoding method, and an image decoding program according to a first embodiment of the present invention will be described in detail with reference to the drawings.

[0011] FIG. 1(a) is a diagram for explaining the configuration of an image encoding apparatus 100 according to the first embodiment , and FIG. 1(b) is a diagram for explaining the configuration of an image decoding apparatus 200 according to the first embodiment.

[0012] The image encoding apparatus 100 of the present embodiment includes a block size determination unit 110, an inter prediction unit 120, a conversion unit 130, a code string generation unit 140, a partial decoding unit 150, and a frame memory 160. The image encoding apparatus 100 receives an input of an input image, performs intra prediction and inter prediction, and outputs an encoded stream. Hereinafter, an image and a picture are used with the same meaning.

[0013] The image decoding apparatus 200 includes a code string decoding unit 210, an inter prediction unit 220, an inverse conversion unit 230 , and a frame memory 240. The image decoding apparatus 200 is different from the image encoding apparatus 100 in that ​​Receives the input of the output encoded stream, performs intra prediction and inter prediction, and decodes and outputs the coded image.

[0014] The image encoding device 100 and the image decoding device 200 are implemented by hardware such as an information processing device including a CPU (Central Proce ssing Unit), a memory, etc. It is realized.

[0015] First, the functions and operations of each part of the image encoding device 100 will be described. Intra prediction is Assumed to be implemented in the same way as HEVC, and hereinafter, inter prediction will be described.

[0016] The block size determination unit 110 determines the block size for inter prediction based on the input image, and supplies the determined block size, the position of the block, and the input pixels (input values) corresponding to the block size to the inter prediction unit 120. Regarding the method for determining the block size, The RDO (Rate-Distortion Optimization) method used in the reference software of HEVC, etc., is used. Regarding the method for determining the block size, The RDO (Rate-Distortion Optimization) method used in the reference software of HEVC, etc., is

[0017] Here, the block size will be described. FIG. 2 shows an example in which a partial region of the image input to the image encoding device 100 is divided into blocks based on the block size determined by the block size determination unit 110. The block sizes are 4×4, 8×4, 4×8, 8×8, 16×8, 8×16, 32×32, ···, 128×64, 64×128, 12 8×128 exist, and the input image is divided into any of the above block sizes so that each block does not overlap. 8×8, 16×8, 8×16, 32×32, ···, 128×64, 64×128, 12 8×128 exist, and the input image is divided into any of the above block sizes so that each block does not overlap. It is divided.

[0018] The inter prediction unit 120 uses the information input from the block size determination unit 110 and the frame Using the reference picture input from the memory 160, determine the inter-prediction parameters for inter-prediction. The inter-prediction unit 120 performs inter-prediction based on the inter-prediction parameters to derive a predicted value, and supplies the block size, the position of the block, the input value, the inter-prediction parameters, and the predicted value to the conversion unit 130. Regarding the method for determining the inter-prediction parameters, an RDO (Rate-Distortion Optimization) method or the like used in the reference software of HEVC is used. The details of the inter-prediction parameters and the operation of the inter-prediction unit 120 will be described later. The conversion unit 130 subtracts the predicted value from the input value to calculate a difference value, performs processes such as orthogonal transformation and quantization on the calculated difference value to calculate prediction error data, and supplies the block size, the position of the block, the inter-prediction parameters, and the prediction error data to the bitstream generation unit 140 and the local decoding unit 150. The bitstream generation unit 140 encodes the SPS (Sequence Parameter Set), PPS (Picture Parameter Set), and other information as necessary, encodes the bitstream for determining the block size supplied from the conversion unit 130, encodes the inter-prediction parameters as a bitstream, encodes the prediction error data as a bitstream, and outputs an encoded stream. The details of the encoding of the inter-prediction parameters will be described later. The local decoding unit 150 performs processes such as inverse orthogonal transformation and inverse quantization on the prediction error data to restore the difference value, adds the difference value and the predicted value to generate a decoded image, and supplies the decoded image and the inter-prediction parameters to the frame memory 160. Using the reference picture input from the memory 160, determine the inter-prediction parameters for inter-prediction. The inter-prediction unit 120 performs inter-prediction based on the inter-prediction parameters to derive a predicted value, and supplies the block size, the position of the block, the input value, the inter-prediction parameters, and the predicted value to the conversion unit 130. Regarding the method for determining the inter-prediction parameters, an RDO (Rate-Distortion Optimization) method or the like used in the reference software of HEVC is used.

[0019] The details of the inter-prediction parameters and the operation of the inter-prediction unit 120 will be described later. The conversion unit 130 subtracts the predicted value from the input value to calculate a difference value, performs processes such as orthogonal transformation and quantization on the calculated difference value to calculate prediction error data, and supplies the block size, the position of the block, the inter-prediction parameters, and the prediction error data to the bitstream generation unit 140 and the local decoding unit 150. The bitstream generation unit 140 encodes the SPS (Sequence Parameter Set), PPS (Picture Parameter Set), and other information as necessary, encodes the bitstream for determining the block size supplied from the conversion unit 130, encodes the inter-prediction parameters as a bitstream, encodes the prediction error data as a bitstream, and outputs an encoded stream. The details of the encoding of the inter-prediction parameters will be described later.

[0020] The local decoding unit 150 performs processes such as inverse orthogonal transformation and inverse quantization on the prediction error data to restore the difference value, adds the difference value and the predicted value to generate a decoded image, and supplies the decoded image and the inter-prediction parameters to the frame memory 160. Using the reference picture input from the memory 160, determine the inter-prediction parameters for inter-prediction. The inter-prediction unit 120 performs inter-prediction based on the inter-prediction parameters to derive a predicted value, and supplies the block size, the position of the block, the input value, the inter-prediction parameters, and the predicted value to the conversion unit 130. Regarding the method for determining the inter-prediction parameters, an RDO (Rate-Distortion Optimization) method or the like used in the reference software of HEVC is used. The details of the inter-prediction parameters and the operation of the inter-prediction unit 120 will be described later. The conversion unit 130 subtracts the predicted value from the input value to calculate a difference value, performs processes such as orthogonal transformation and quantization on the calculated difference value to calculate prediction error data, and supplies the block size, the position of the block, the inter-prediction parameters, and the prediction error data to the bitstream generation unit 140 and the local decoding unit 150.

[0021] The bitstream generation unit 140 encodes the SPS (Sequence Parameter Set), PPS (Picture Parameter Set), and other information as necessary, encodes the bitstream for determining the block size supplied from the conversion unit 130, encodes the inter-prediction parameters as a bitstream, encodes the prediction error data as a bitstream, and outputs an encoded stream. The details of the encoding of the inter-prediction parameters will be described later. The local decoding unit 150 performs processes such as inverse orthogonal transformation and inverse quantization on the prediction error data to restore the difference value, adds the difference value and the predicted value to generate a decoded image, and supplies the decoded image and the inter-prediction parameters to the frame memory 160.

[0022] The frame memory 160 stores the decoded images and the inter-prediction parameters for a plurality of images. The decoded image and the inter prediction parameters are supplied to the inter prediction unit 120 .

[0023] Next, the function and operation of each unit of the image decoding device 200 will be described. It is assumed that inter prediction is performed in the same manner as EVC, and inter prediction will be described below.

[0024] The code string decoding unit 210 extracts SPS, PPS headers, and other information from the coded stream as necessary. and other information such as block size, block position, inter prediction parameters, and The prediction error data is decoded from the encoded stream, and the block size, block position, and The inter prediction unit 220 supplies the inter prediction parameters and the prediction error data to the inter prediction unit 220.

[0025] The inter prediction unit 220 receives the information input from the code stream decoding unit 210 and the frame memory 2 Using the reference picture input from 40, inter prediction is performed to derive a predicted value. The inter-prediction unit 220 determines the block size, the block position, the inter-prediction parameters, and the prediction The error data and the predicted value are supplied to an inverse transform unit 230 .

[0026] The inverse transform unit 230 performs an inverse orthogonal transform on the prediction error data provided by the inter prediction unit 220. The difference value is calculated by performing processes such as inverse quantization and dequantization, and the difference value and the predicted value are added to generate a decoded image. The decoded image and the inter-prediction parameters are supplied to the frame memory 240, and the decoded image is Output.

[0027] The frame memory 240 stores the decoded images and the inter-prediction parameters for a plurality of images. Supply the decoded image and the inter prediction parameters to the inter prediction unit 220.

[0028] Note that the inter prediction performed by the inter prediction unit 120 and the inter prediction unit 220 is the same operation, and the decoded images and the inter prediction parameters stored in the frame memory 160 and the frame memory 240 are also the same.

[0029] Subsequently, the inter prediction parameters will be described. The inter prediction parameters include a merge flag, a merge index, a valid flag for prediction LX, a motion vector for prediction LX, a reference picture index for prediction LX, a merge correction flag, and a differential motion vector for prediction LX. LX is L0 and L1. The merge flag is a flag indicating whether to use the merge mode or the differential motion vector mode as the inter prediction mode. If the merge flag is 1, the merge mode is used; if the merge flag is 0, the differential motion vector mode is used. The merge index is an index indicating the position of the selected merge candidate in the merge candidate list. The valid flag for prediction LX is a flag indicating whether prediction LX is valid or invalid. If both L0 prediction and L1 prediction are valid, it is a dual prediction; if L0 prediction is valid and L1 prediction is invalid, it is L0 prediction; if L1 prediction is valid and L0 prediction is invalid, it is L1 prediction. The merge correction flag is a flag indicating whether to correct the motion information of the merge candidate. If the merge correction flag is 1, the merge candidate is corrected; if the merge correction flag is 0, the merge candidate is not corrected. Here, the code string generation unit 140 does not encode the valid flag for prediction LX as a code string in the encoded stream. Also, the code string decoding unit 210 does not Do not decode it as a symbol string from the coded stream. The reference picture index is an index for identifying the decoded picture in Frame Memory 160. Also, the combination of the valid flag of L0 prediction, the valid flag of L1 prediction, the motion vector of L0 prediction, the motion vector of L1 prediction, the reference picture index of L0 prediction, and the reference picture index of L1 prediction is used as motion information. In addition, if the block is an intra-coded mode block or a block outside the image area, both the valid flag of L0 prediction and the valid flag of L1 prediction are set to invalid. Hereinafter, the picture type will be described as a B picture in which all of the unidirectional prediction of L0 prediction, the unidirectional prediction of L1 prediction, and the bidirectional prediction are available, but it may be a P picture in which only the unidirectional prediction is available for the picture type. In the case of a P picture, the inter prediction parameters are only for L0 prediction, and L1 prediction is processed as non-existent. Generally, when improving the coding efficiency in a B picture, the reference picture of L0 prediction is a picture in the past with respect to the picture to be predicted, and the reference picture of L1 prediction is a picture in the future with respect to the picture to be predicted. This is because the coding efficiency is improved by interpolation prediction when the reference picture of L0 prediction and the reference picture of L1 prediction are in the opposite direction with respect to the picture to be predicted. Whether the reference picture of L0 prediction and the reference picture of L1 prediction are in the opposite direction with respect to the picture to be predicted can be determined by comparing the POC (Picture Order Count) of the reference pictures. Hereinafter, it will be described that the reference picture of L0 prediction and the reference picture of L1 prediction are temporally in the opposite direction with respect to the picture to be predicted in which the block is located.

[0030]

[0031] ​​​​​​​​​​​​​​​​

[0032] Next, the details of the inter prediction unit 120 will be described. Unless otherwise specified, the configurations and operations of the inter prediction unit 120 of the video encoding apparatus 100 and the inter prediction unit 220 of the video decoding apparatus 200 are the same.

[0033] FIG. 3 is a diagram for explaining the configuration of the inter prediction unit 120. The inter prediction unit 120 includes a motion mode determination unit 121, a merge candidate list generation unit 122, a merge candidate selection unit 123, a merge candidate correction determination unit 124, a merge candidate correction unit 125, a differential motion vector

[0034] mode execution unit 126, and a predicted value derivation unit 127. The inter prediction unit 120 switches between the merge mode and the differential motion vector mode as the inter prediction mode for each block. The differential motion vector mode executed by the differential motion vector mode execution unit 126 is executed in the

[0035] same manner as in HEVC. Hereinafter, mainly the merge mode will be described. The merge mode determination unit 121 determines whether to use the merge mode as the inter prediction mode for each block. If the merge flag is 1, the merge mode is used; if

[0036] the merge flag is 0, the differential motion vector mode is used. In the inter prediction unit 120, the determination of whether to set the merge flag to 1 is performed using the RDO (rate distortion optimization) method or the like used in the reference software of HEVC. In the inter prediction unit 220, the syntax decoding unit

[0037] When the merge flag is 0, the differential motion vector mode is executed in the differential motion vector mode execution unit 126, and the differential motion vector mode is executed, and the inter-prediction parameter of the differential motion vector mode is supplied to the predicted value derivation unit 127.

[0038] When the merge flag is 1, the merge mode is executed by the merge candidate list generation unit 122, the merge candidate selection unit 12 3, the merge candidate correction determination unit 124, and the merge candidate correction unit 125, and the inter-prediction parameter of the merge mode is supplied to the predicted value derivation unit 127.

[0039] The processing when the merge flag is 1 will be described in detail hereinafter.

[0040] FIG. 4 is a flowchart for explaining the operation of the merge mode. Hereinafter, the merge mode will be described in detail with reference to FIGS. 3 and 4.

[0041] First, the merge candidate list generation unit 122 generates a merge candidate list from the motion information of the blocks adjacent to the block to be processed and the motion information of the blocks in the decoded image (S100), and supplies the generated merge candidate list to the merge candidate selection unit 123. Hereinafter, the block to be processed and the block to be predicted are used in the same meaning.

[0042] Here, the generation of the merge candidate list will be described. FIG. 5 is a diagram for explaining the configuration of the merge candidate list generation unit 122. The merge candidate list generation unit 122 includes a spatial merge candidate generation unit 201, a temporal merge candidate generation unit 202, and a merge candidate supplementation unit 203.

[0043] FIG. 6 is a diagram for explaining the blocks adjacent to the block to be processed. Here, the processing target The blocks adjacent to the elephant block are block A, block B, block C, block D, block E, block F, block G, but it is not limited to this if multiple blocks adjacent to the block to be processed are used.

[0044] FIG. 7 is a diagram for explaining the blocks on the decoded image at the same position as the block to be processed and its periphery. Here, the blocks on the decoded image at the same position as the block to be processed and its periphery are block CO1, block CO2, block CO3, but it is not limited to this if multiple blocks on the decoded image at the same position as the block to be processed and its periphery are used. Hereinafter, block CO1, CO2, CO3 are called the same-position blocks, and the decoded image including the same-position blocks is called the same-position picture.

[0045] Hereinafter, the generation of the merge candidate list will be described in detail with reference to FIGS. 5, 6, and 7.

[0046] First, the spatial merge candidate generation unit 201 sequentially checks block A, block B, block C, block D, block E, block F, block G. If either or both of the valid flags of L0 prediction and the valid flag of L1 prediction are valid, the motion information of the block is sequentially added to the merge candidate list as a merge candidate. The merge candidates generated by the spatial merge candidate generation unit 201 are called spatial merge candidates.

[0047] Next, the temporal merge candidate generation unit 202 sequentially checks block C01, block CO2, block C O3. First, for the block where either or both of the valid flag of L0 prediction and the valid flag of L1 prediction become valid, after performing processing such as scaling on the motion information, it is used as a merge candidate. ​​​​​​​​​​They are sequentially added to the motion merge candidate list. The motion merge candidates generated by the motion merge candidate generation unit 202 are called motion merge candidates.

[0048] Here, the scaling of the motion merge candidates will be described. The scaling of the motion merge candidates is the same as that of HEVC. The motion vectors of the motion merge candidates are the motion vectors of the same-position blocks scaled based on the distance between the reference picture referred to by the same-position picture and the same-position block and the picture referred to by the motion merge candidate for the block to be predicted. The distance between the picture where the block to be predicted is located and the picture referred to by the motion merge candidate is used for scaling. The motion vectors are derived by scaling.

[0049] Note that the pictures referred to by the motion merge candidates are reference pictures with a reference picture index of 0 for both L0 prediction and L1 prediction. Also, as the same-position block, it is determined whether the same-position block of L0 prediction or the same-position block of L1 prediction is used by encoding (decoding) the same-position derivation flag. As described above, the motion merge candidates are the motion vectors of either the L0 prediction or the L1 prediction of the same-position block scaled for L0 prediction and L1 prediction to derive new motion vectors for L0 prediction and L1 prediction, which are used as the motion vectors of the L0 prediction and L1 prediction of the motion merge candidates. Whether the same-position block of L0 prediction or the same-position block of L1 prediction is used is determined by encoding (decoding) the same-position derivation flag. As described above, the motion merge candidates are the motion vectors of either the L0 prediction or the L1 prediction of the same-position block scaled for L0 prediction and L1 prediction to derive new motion vectors for L0 prediction and L1 prediction, which are used as the motion vectors of the L0 prediction and L1 prediction of the motion merge candidates. One of the motion vectors of either the L0 prediction or the L1 prediction of the same-position block is scaled for L0 prediction and L1 prediction to derive new motion vectors for L0 prediction and L1 prediction, which are used as the motion vectors of the L0 prediction and L1 prediction of the motion merge candidates. scaled for L0 prediction and L1 prediction to derive new motion vectors for L0 prediction and L1 prediction, which are used as the motion vectors of the L0 prediction and L1 prediction of the motion merge candidates. to be the motion vectors of the L0 prediction and L1 prediction of the motion merge candidates.

[0050] Next, if the same motion information is included in the merge candidate list multiple times, one piece of motion information is kept and the other motion information is deleted.

[0051] Next, if the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates the merge candidate supplementation unit 203 supplements the merge candidate list until the number of merge candidates included in the merge candidate list reaches the maximum number of merge candidates Until the number of merge candidates reaches the maximum number of merge candidates, supplementary merge candidates are added to the merge candidate list to make the number of merge candidates included in the merge candidate list equal to the maximum number of merge candidates. Here, the supplementary merge candidate is motion information where both the motion vectors of the L0 prediction and the L1 prediction are (0, 0) and both the reference picture indices of the L0 prediction and the L1 prediction are 0. Here, the maximum number of merge candidates is set to 6, but it may be 1 or more. Subsequently, the merge candidate selection unit 123 selects one merge candidate from the merge candidate list (S101), supplies the selected merge candidate (referred to as the "selected merge candidate") and the merge index to the merge candidate correction determination unit 124, and sets the selected merge candidate as the motion information of the block to be processed. In the inter prediction unit 120 of the image encoding apparatus 100, one merge candidate is selected from the merge candidates included in the merge candidate list using the RDO (rate-distortion optimization) method or the like used in the reference software of HEVC to determine the merge index. In the inter prediction unit 220 of the image decoding apparatus 200, the code sequence decoding unit 210 acquires the merge index decoded from the encoded stream, and selects one merge candidate from the merge candidates included in the merge candidate list based on the merge index as the selected merge candidate. Subsequently, the merge candidate correction determination unit 124 checks whether the width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate are valid (S102). The width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and the L0 prediction of the selected merge candidate

[0052] Here, the maximum number of merge candidates is 6, but it may be 1 or more.

[0053] Subsequently, the merge candidate selection unit 123 selects one merge candidate from the merge candidate list (S101), supplies the selected merge candidate (referred to as the "selected merge candidate") and the merge index to the merge candidate correction determination unit 124, and sets the selected merge candidate as the motion information of the block to be processed. (S101), the selected merge candidate (referred to as the "selected merge candidate") and the merge index are supplied to the merge candidate correction determination unit 124, and the selected merge candidate is used as the motion information of the block to be processed. In the inter prediction unit 120 of the image encoding apparatus 100, one merge candidate is selected from the merge candidates included in the merge candidate list using the RDO (rate-distortion optimization) method or the like used in the reference software of HEVC to determine the merge index. In the inter prediction unit 120 of the image encoding apparatus 100, one merge candidate is selected from the merge candidates included in the merge candidate list using the RDO (rate-distortion optimization) method or the like used in the reference software of HEVC to determine the merge index. In the inter prediction unit 120 of the image encoding apparatus 100, one merge candidate is selected from the merge candidates included in the merge candidate list using the RDO (rate-distortion optimization) method or the like used in the reference software of HEVC to determine the merge index. In the inter prediction unit 120 of the image encoding apparatus 100, one merge candidate is selected from the merge candidates included in the merge candidate list using the RDO (rate-distortion optimization) method or the like used in the reference software of HEVC to determine the merge index. In the inter prediction unit 220 of the image decoding apparatus 200, the code sequence decoding unit 210 acquires the merge index decoded from the encoded stream, and selects one merge candidate from the merge candidates included in the merge candidate list based on the merge index as the selected merge candidate. In the inter prediction unit 220 of the image decoding apparatus 200, the code sequence decoding unit 210 acquires the merge index decoded from the encoded stream, and selects one merge candidate from the merge candidates included in the merge candidate list based on the merge index as the selected merge candidate. In the inter prediction unit 220 of the image decoding apparatus 200, the code sequence decoding unit 210 acquires the merge index decoded from the encoded stream, and selects one merge candidate from the merge candidates included in the merge candidate list based on the merge index as the selected merge candidate.

[0054] Subsequently, the merge candidate correction determination unit 124 checks whether the width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate are valid (S102). The width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate are valid. The width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate are valid. The width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, and the L0 prediction of the selected merge candidate If the condition that both or at least one of the L1 prediction and the L0 prediction is valid is not satisfied (S1 02 NO), without correcting the selected merge candidate as the motion information of the processing target block, proceed to step S111. Here, since the merge candidate list always includes a merge candidate in which at least one of the L0 prediction and the L1 prediction is valid, it is obvious that both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate is valid. Therefore, the "Is both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate valid?" in S102 is omitted and S102 may be inspected to determine whether the width of the processing target block is equal to or greater than a predetermined width and the height of the processing target block is equal to or greater than a predetermined height.

[0055] If the width of the processing target block is equal to or greater than a predetermined width, the height of the processing target block is equal to or greater than a predetermined height, and both or at least one of the L0 prediction and the L1 prediction of the selected merge candidate is valid (S 102 YES), the merge candidate correction determination unit 124 sets the merge correction flag (S10 3) and supplies the merge correction flag to the merge candidate correction unit 125. In the inter prediction unit 120 of the image encoding apparatus 100, if the prediction error when inter predicting with the merge candidate is equal to or greater than a predetermined prediction error, the merge correction flag is set to 1, and if the prediction error when inter predicting with the selected merge candidate is less than the predetermined prediction error, the merge correction flag is set to 0. In the inter prediction unit 220 of the image decoding apparatus 200, the syntax-based code sequence decoding unit 210 acquires the merge correction flag decoded from the encoded stream. Subsequently, the merge candidate correction unit 125 checks whether the merge correction flag is 1 (S10 417)

[0056] ​4). If the merge correction flag is not 1 (NO in S104), the selected merge candidate is the processing target Proceed to step S111 without correcting the motion information of the block.

[0057] If the merge correction flag is 1 (YES in S104), check whether the L0 prediction of the selected merge candidate is valid (S105). If the L0 prediction of the selected merge candidate is not valid (NO in S10 5), proceed to step S108. If the L0 prediction of the selected merge candidate is valid (YES in S1 05), determine the differential motion vector of the L0 prediction (S106). As described above, if the merge correction flag is 1, correct the motion information of the selected merge candidate, and if the merge correction flag is 0, do not correct the motion information of the selected merge candidate.

[0058] In the inter prediction unit 120 of the image encoding device 100, the differential motion vector of the L0 prediction is obtained by motion vector search. Here, the search range of the motion vector is both ±16 in the horizontal and vertical directions but may be a multiple of 2 such as ±64. In the inter prediction unit 220 of the image decoding device 200, the differential motion vector of the L0 prediction decoded by the syntax-based encoding stream by the syntax decoding unit 210 is obtained.

[0059] Subsequently, the merge candidate correction unit 125 calculates the corrected motion vector of the L0 prediction and sets the corrected motion vector of the L0 prediction as the motion vector of the L0 prediction of the motion information of the processing target block (S107).

[0060] Here, the relationship between the corrected motion vector of the L0 prediction (mvL0), the motion vector of the L0 prediction of the selected merge candidate (mmvL0), and the differential motion vector of the L0 prediction (mvdL0) is described ​​​​​is clarified. The corrected motion vector (mvL0) of the L0 prediction is the sum of the motion vector (mmvL0) of the L0 prediction of the selection merge candidate and the differential motion vector (mvdL0) of the L0 prediction, and is expressed by the following formula. Here, [0] represents the horizontal component of the motion vector, and [1] represents the vertical component of the motion vector.

[0061] mvL0[0] = mmvL0[0] + mvdL0[0] mvL0[1] = mmvL0[1] + mvdL0[1] Subsequently, it is inspected whether the L1 prediction of the selection merge candidate is valid (S108). If the L1 prediction of the selection merge candidate is not valid (NO in S108), the process proceeds to step S111. If the L1 prediction of the selection merge candidate is valid (YES in S108), the differential motion vector of the L1 prediction is determined (S109).

[0062] In the inter prediction unit 120 of the image encoding apparatus 100, the differential motion vector of the L1 prediction is obtained by motion vector search. Here, the search range of the motion vector is set to ±16 both in the horizontal direction and the vertical direction, but it can be set to a power of 2 such as ±64. In the inter prediction unit 220 of the image decoding apparatus 200, the syntax-based code sequence decoding unit 210 acquires the differential motion vector of the L1 prediction decoded from the encoded stream.

[0063] Subsequently, the merge candidate correction unit 125 calculates the corrected motion vector of the L1 prediction, and sets the corrected motion vector of the L1 prediction as the motion vector of the L1 prediction of the motion information of the processing target block (S110).

[0064] Here, the corrected motion vector (mvL1) of the L1 prediction, the motion ​​​​​​​​​​​Explain the relationship between the vector (mmvL1) and the differential motion vector (mvdL1) of the L1 prediction. The corrected motion vector (mvL1) of the L1 prediction is the sum of the motion vector (mmvL1) of the L1 prediction of the selected merge candidate and the differential motion vector (mvdL1) of the L1 prediction, and is expressed by the following formula. Here, [0] represents the horizontal component of the motion vector, and [1] represents the vertical component of the motion vector. The corrected motion vector (mvL1) of the L1 prediction is the sum of the motion vector (mmvL1) of the L1 prediction of the selected merge candidate and the differential motion vector (mvdL1) of the L1 prediction, and is expressed by the following formula. Here, [0] represents the horizontal component of the motion vector, and [1] represents the vertical component of the motion vector. That is, the following equation holds. Note that [0] represents the horizontal component of the motion vector, and [1] represents the vertical component of the motion vector. Indicates the vertical component of the motion vector.

[0065] mvL1[0] = mmvL1[0] + mvdL1[0] mvL1[1] = mmvL1[1] + mvdL1[1] Subsequently, the prediction value derivation unit 127 performs inter prediction of either L0 prediction, L1 prediction, or dual prediction based on the motion information of the processing target block, and derives a prediction value (S111). As described above, if the merge correction flag is 1, the motion vector of the selected merge candidate is corrected, and if the merge correction flag is 0, the motion vector of the selected merge candidate is not corrected. As described above, if the merge correction flag is 1, the motion vector of the selected merge candidate is corrected. If the merge correction flag is 0, the motion vector of the selected merge candidate is not corrected.

[0066] The encoding of the inter prediction parameter will be described in detail. FIG. 8 is a diagram showing a part of the syntax of a block in the merge mode. Table 1 shows the relationship between the inter prediction parameter and the syntax. In FIG. 8, cbWidth is the width of the processing target block, and cbHeight is the height of the processing target block. Both the predetermined width and the predetermined height are set to 8. By setting the predetermined width and the predetermined height, the processing amount can be reduced by not correcting the merge candidate in small block units. Here, cu_skip_flag is 1 if the block is in the skip mode, and 0 if it is not in the skip mode. The syntax of the skip mode is the same as the syntax of the merge mode. merge_idx is the merge candidate. FIG. 8 is a diagram showing a part of the syntax of a block in the merge mode. Table 1 shows the relationship between the inter prediction parameter and the syntax. In FIG. 8, cbWidth is the width of the processing target block, and cbHeight is the height of the processing target block. The predetermined width and the predetermined height are both set to 8. By setting the predetermined width and the predetermined height, the processing amount can be reduced by not correcting the merge candidate in small block units. Here, cu_skip_flag is 1 if the block is in the skip mode, and 0 if it is not in the skip mode. The syntax of the skip mode is the same as the syntax of the merge mode. merge_idx is the merge candidate. It is a merge index for selecting a merge candidate from a list.

[0067] Encode (decode) merge_idx before merge_mod_flag to determine the merge index, and then judge the encoding (decoding) of merge_mod_flag. By sharing merge_idx with the merge index in the merge mode, the encoding efficiency is improved while suppressing the increase in syntax complexity and context.

[0068] [Table 1]

[0069] Figure 9 is a diagram showing the syntax of differential motion vectors. mvd_codin in Figure 9 g(N) has the same syntax as the syntax used in the differential motion vector mode . Here, N is 0 or 1. N = 0 indicates L0 prediction, and N = 1 indicates L1 prediction.

[0070] The syntax of the differential motion vector includes a flag abs_mvd_greater0_flag[d] indicating whether the component of the differential motion vector is greater than 0, the differential motion vector flag abs_mvd_greater1_flag[d] indicating whether the component is greater than 1, the mvd_sign_fl ag[d] indicating the sign (±) of the component of the differential motion vector, and the abs_m vd_minus2[d] indicating the absolute value of the vector obtained by subtracting 2 from the component of the differential motion vector. Here, d is 0 or 1. d = 0 indicates the horizontal direction component, and d = 1 indicates the vertical direction component.

[0071] In HEVC, there are a merge mode and a differential motion vector mode as inter prediction modes. In the merge mode, since motion information can be restored with only one merge flag, it is a mode with very high encoding efficiency. However, since the motion information in the merge mode depends on the processed block, the prediction efficiency is limited in some cases, and it is necessary to improve the utilization efficiency.

[0072] On the other hand, in the differential motion vector mode, L0 prediction and L1 prediction are prepared as separate syntaxes, and for each of the prediction type (L0 prediction, L1 prediction, or dual prediction) and L0 prediction and L1 prediction, a prediction motion vector flag, a differential motion vector, and a reference picture index are required. Therefore, the differential motion vector mode has lower encoding efficiency than the merge mode, but it has a stable and high prediction efficiency for sudden motions with little correlation with the motions of spatially adjacent blocks or temporally adjacent blocks that cannot be derived in the merge mode.

[0073] In the present embodiment, by enabling correction of the motion vector in the merge mode while fixing the prediction type and the reference picture index in the merge mode, the encoding efficiency can be improved more than that in the differential motion vector mode, and the utilization efficiency can be improved more than that in the merge mode.

[0074] Also, by making the differential motion vector the difference from the motion vector of the selected merge candidate, the magnitude of the differential motion vector can be suppressed to be small, and the encoding efficiency can be suppressed.

[0075] In addition, the syntax of the differential motion vector in the merge mode is made the same as the difference in the differential motion vector mode. By making it the same as the syntax of the fractional motion vector, the differential motion vector can be added in the merge mode with few configuration changes. Even if added, the configuration changes can be minimized.

[0076] Also, by defining a predetermined width and a predetermined height, if the predicted block width is greater than or equal to the predetermined width, and the predicted block height is greater than or equal to the predetermined height, and both or at least one of the L0 prediction or L1 prediction of the selected merge candidate is valid, the process of correcting the motion vector in the merge mode is omitted, and the processing amount of correcting the motion vector in the merge mode can be suppressed. Note that when there is no need to suppress the processing amount of correcting the motion vector, there is no need to limit the correction of the motion vector by the predetermined width and the predetermined height. Next, a modification example of the present embodiment will be described. Unless otherwise specified, the modification examples can be combined with each other. Note that when there is no need to suppress the processing amount of correcting the motion vector, there is no need to limit the correction of the motion vector by the predetermined width and the predetermined height.

[0077] Hereinafter, modification examples of the present embodiment will be described. Unless otherwise specified, the modification examples can be combined with each other. can be combined with each other. [Modification Example 1] In the present embodiment, the differential motion vector is used as the syntax of the block in the merge mode. In this modification example, the differential motion vector is defined to be encoded (or decoded) as a differential unit motion vector. The differential unit motion vector is a motion vector when the picture interval is the minimum interval. In HEVC or the like, the minimum picture interval is encoded as a code sequence in the encoding stream. The differential unit motion vector is scaled according to the interval between the picture to be encoded and the reference picture in the merge mode and used as the differential motion vector. Let the POC (Picture Order Count) of the picture to be encoded be POC(Cur), and the L0 prediction in the merge mode is encoded as a code sequence in the encoding stream.

[0078] The differential unit motion vector is scaled according to the interval between the picture to be encoded and the reference picture in the merge mode and used as the differential motion vector. Let the POC (Picture Order Count) of the picture to be encoded be POC(Cur), and the L0 prediction in the merge mode Picture Order Count) be POC(Cur), and the L0 prediction in the merge mode Let the POC of the reference picture for measurement be POC(L0), and the POC of the reference picture for L1 prediction in merge mode be POC(L1). Then, the motion vector is calculated as follows: umvdL 0 represents the differential unit motion vector for L0 prediction, and umvdL1 represents the differential unit motion vector for L1 prediction.

[0079] mvL0[0] = mmvL0[0] + umvdL0[0]*(POC(Cur ) - POC(L0)) mvL0[1] = mmvL0[1] + umvdL0[1]*(POC(Cur ) - POC(L0)) mvL1[0] = mmvL1[0] + umvdL1[0]*(POC(Cur ) - POC(L1)) mvL1[1] = mmvL1[1] + umvdL1[1]*(POC(Cur ) - POC(L1)) As described above, in this modification example, by using the differential unit motion vector as the differential motion vector and reducing the signed amount of the differential motion vector, the encoding efficiency can be improved. Note that when the differential motion vector is large and the distance between the prediction target picture and the reference picture is large, the encoding efficiency can be particularly improved. Also, when there is a proportional relationship between the interval between the processing target picture and the reference picture and the speed of the object moving within the screen, both the prediction efficiency and the encoding efficiency can be improved.

[0080] In addition, it is not necessary to scale and derive based on the picture - to - picture distance like time merge candidates. In the decoder, since scaling can be achieved only by a multiplier, a divider is not required, and the circuit scale and the processing amount can be reduced. [Modification Example 2] In this embodiment, it is possible to encode (or decode) 0 as a component of the differential motion vector. For example, only L0 prediction can be changed. In this modification, 0 cannot be encoded (or decoded) as a component of the differential motion vector.

[0081] FIG. 10 is a diagram showing the syntax of the differential motion vector of Modification 2. The syntax of the differential motion vector includes a flag abs_mvd_greater1_flag[d] indicating whether the component of the differential motion vector is greater than 1, a flag abs_mvd_greater2_flag[d] indicating whether the component of the differential motion vector is greater than 2, the absolute value of the vector obtained by subtracting 3 from the component of the differential motion vector, abs_mvd_minus3[d], and a flag mvd_sign_flag[d] indicating the sign (±) of the component of the differential motion vector.

[0082] As described above, by making it impossible to encode (or decode) 0 as a component of the differential motion vector, the encoding efficiency when the component of the differential motion vector is 1 or more can be improved. [Modification 3] In this embodiment, the components of the differential motion vector are integers, and in Modification 2, they are integers excluding 0. In this modification, the components excluding the ± sign of the differential motion vector are limited to powers of 2.

[0083] Instead of abs_mvd_minus2[d] which is the syntax of this embodiment, abs_mvd_pow_plus1[d] is used. The differential motion vector mvd[d] is calculated from mvd_sign_flag[d] and abs_mvd_pow_plus1[d] as follows. ​​​​​​​​​​​

[0084] mvd[d] = mvd_sign_flag[d] * 2 ^ (abs_mvd_po w_plus1[d] + 1) Also, instead of abs_mvd_minus3[d] which is the syntax of Modification Example 2, abs_mvd_pow_plus2[d] is used. The differential motion vector mvd[d] is calculated from mvd_sign_flag[d] and abs_mvd_pow_plus2[d] as follows.

[0085] mvd[d] = mvd_sign_flag[d] * 2 ^ (abs_mvd_po w_plus2[d] + 2) By limiting the components of the differential motion vector to powers of 2, the processing amount of the encoding device can be significantly reduced while improving the prediction efficiency in the case of large motion vectors. [Modification Example 4] In this embodiment, mvd_coding(N) includes the differential motion vector, but in this modification example, mvd_coding(N) includes the motion vector magnification.

[0086] In the syntax of mvd_coding(N) of this modification example, abs_mvd_gre ater0_flag[d], abs_mvd_greater1_flag[d], and m vd_sign_flag[d] do not exist. Instead, it has a configuration including abs_mvr_plus 2[d] and mvr_sign_flag[d].

[0087] The corrected motion vector (mvLN) for LN prediction is the product of the motion vector of LN prediction of the selection merge candidate (mmvLN) and the motion vector magnification (mvrLN), and is calculated by the following formula.

[0088] ​​ mvLN[d] = mmvLN[d] * mvrLN[d] By limiting the components of the differential motion vector to powers of 2, the processing amount of the encoding device can be significantly reduced while improving the prediction efficiency in the case of large motion vectors.

[0089] Note that this modification example cannot be combined with Modification Example 1, Modification Example 2, Modification Example 3, and Modification Example 6. [Modification Example 5] In the syntax of FIG. 8 of the present embodiment, when cu_skip_flag is 1 (in the skip mode), it is assumed that there is a possibility that merge_mod_flag exists. However, in the case of the skip mode, merge_mod_flag may not exist.

[0090] In this way, by omitting merge_mod_flag, the encoding efficiency in the skip mode can be improved, and the determination of the skip mode can be simplified. [Modification Example 6] In the present embodiment, it is checked whether the LN prediction (N = 0 or 1) of the selected merge candidate is valid. If the LN prediction of the selected merge candidate is not valid, the differential motion vector is not made valid. Accordingly, without checking whether the LN prediction of the selected merge candidate is valid, the differential motion vector may be made valid regardless of whether the LN prediction of the selected merge candidate is valid. In this case, when the LN prediction of the selected merge candidate is invalid, the motion vector of the LN prediction of the selected merge candidate is set to (0, 0), and the reference picture index of the LN prediction of the selected merge candidate is set to 0.

[0091] In this way, in this modification example, regardless of whether the LN prediction of the selected merge candidate is valid ​​​By enabling the differential motion vector, the opportunity to utilize bidirectional prediction is increased, and the encoding efficiency can be improved. [Modification Example 7] In this embodiment, it is determined whether the L0 prediction and the L1 prediction of the selection merge candidate are individually valid, and it is controlled whether to encode (or decode) the differential motion vector. However, when both the L0 prediction and the L1 prediction of the selection merge candidate are valid, the differential motion vector is encoded (or decoded), and when both the L0 prediction and the L1 prediction of the selection merge candidate are not valid, the differential motion ve ctor may not be encoded (or decoded). In the case of this modification example, step S102 is as follows. The merge candidate correction determination unit 124 checks whether the width of the processing target block is equal to or greater than a predetermined width, and the height of the processing target bl

[0092] ock is equal to or greater than a predetermined height, and both the L0 prediction and the L1 prediction of the selection merge candidate are valid (S102). Moreover, step S105 and step S108 are unnecessary in this modification example.

[0093] Also, step S105 and step S108 are unnecessary in this modification example.

[0094] FIG. 11 is a diagram showing a part of the syntax of a block that is the merge mode of Modification Example 7. The syntax related to steps S102, S105, and S108 is different.

[0095] In this way, in this modification example, by enabling the differential motion vector when both the L0 prediction and the L1 prediction of the selection merge candidate are valid, the motion vector of the selection merge candidate of the frequently used bidirectional prediction is corrected, and the prediction efficiency can be improved efficiently. [Modification Example 8] In this embodiment, as the syntax of a block in the merge mode, the difference in the L0 prediction ​​Two differential motion vectors, i.e., the motion vector and the differential motion vector of the L1 prediction, were used. In this modification example, only one differential motion vector is encoded (or decoded), and one differential motion vector is shared as the correction motion vector of the L0 prediction and the correction motion vector of the L1 prediction, and the motion vector mmvLN (N = 0, 1) of the selection merge candidate and the differential motion vector m vd are used to calculate the motion vector mvLN (N = 0, 1) of the correction merge candidate as follows.

[0096] When the L0 prediction of the selection merge candidate is valid, the motion vector of the L0 prediction is calculated from the following formula.

[0097] mvL0[0] = mmvL0[0] + mvd[0] mvL0[1] = mmvL0[1] + mvd[1] When the L1 prediction of the selection merge candidate is valid, the motion vector of the L1 prediction is calculated from the following formula. The differential motion vector in the direction opposite to the L0 prediction is added. The differential motion vector may be subtracted from the motion vector of the L1 prediction of the selection merge candidate.

[0098] mvL1[0] = mmvL1[0] + mvd[0]*-1 mvL1[1] = mmvL1[1] + mvd[1]*-1 FIG. 12 is a diagram showing a part of the syntax of a block that is the merge mode of Modification 8. The validity of the L0 prediction and the validity of the L1 prediction are checked but deleted, and the absence of mvd_ coding(1) is different from this embodiment. mvd_coding(0) corresponds to one differential motion vector.

[0099] Thus, in this modification example, only one differential motion vector is used for the L0 prediction and the L1 prediction. ​​​​​​By defining, in the case of dual prediction, the number of differential motion vectors is halved and shared between L0 prediction and L 1 prediction, it is possible to improve the coding efficiency while suppressing a decrease in prediction efficiency .

[0100] Also, when the reference picture referred to by the L0 prediction of the selection merge candidate and the reference picture referred to by the L1 prediction are in the opposite direction (not in the same direction) with respect to the picture to be predicted, by adding the differential motion vector in the opposite direction, it is possible to improve the coding efficiency for a motion in a certain direction .

[0101] The effects of this modification example will be described in detail. FIG. 16 is a diagram for explaining the effects of Modification Example 8 . FIG. 16 is an image showing the state of a sphere (the area filled with diagonal lines) moving horizontally within a moving rectangular area (the area surrounded by a broken line). In such a case, the motion of the sphere with respect to the screen is the sum of the motion of the rectangular area and the motion of the sphere moving horizontally. Let picture B be the picture to be predicted, picture A be the reference picture for L0 prediction, and picture C be the reference picture for L1 prediction. Assume that picture A and picture C are reference pictures in a reverse relationship with respect to the picture to be predicted . .

[0102] When the sphere moves in a certain direction at a constant speed, if pictures A, B, and C are equally spaced from each other, the amount of movement of the sphere that cannot be obtained from adjacent blocks is added to the L0 prediction and subtracted from the L1 prediction, so that the motion of the sphere can be accurately reproduced .

[0103] When the sphere moves at a speed that is not constant in a certain direction, pictures A, B, and C are not equally spaced from each other, but if the amount of movement corresponding to the rectangular area of the sphere is equally spaced ​ Add the movement amount of the sphere that cannot be obtained from the adjacent block to the L0 prediction and subtract it from the L1 prediction By doing so, the movement of the sphere can be accurately reproduced.

[0104] Also, when the sphere moves at a constant speed in a certain direction for a certain period of time, Picture A, Picture B , Picture C are not equally spaced, but the movement amounts corresponding to the rectangular area of the sphere may be equally spaced There is a case. FIG. 17 is a diagram for explaining the effect when the picture intervals in the eighth modification are not equally spaced. This example will be described in detail with reference to FIG. 17. Pictures F0, F1 in FIG. 17 , ···, F8 each show pictures at fixed intervals. From Picture F0 to Picture F4 the sphere is stationary, and it is assumed that the sphere moves at a constant speed in a certain direction after Picture F5 . When Picture F0 and Picture F6 are reference pictures and Picture F5 is the prediction target picture , Picture F0, Picture F5, Picture F6 are not equally spaced, but the movement amounts corresponding to the rectangular area of the sphere are equally spaced. When Picture F5 is the prediction target picture , generally Picture F4 with a shorter distance is selected as the reference picture, but Picture F4 is not used and Picture F0 is selected as the reference picture because Picture F0 is a higher-quality picture with less distortion than Picture F4 . The reference pictures are usually managed in the FIFO (First-In First-Out) method, but there is a long-term reference picture as a mechanism for using a high-quality picture with less distortion as the reference picture. In this way, this modification example can improve the prediction efficiency and coding efficiency by applying it when either or both of the L0 prediction and the L1 prediction are long-term reference pictures . Also, this modification example is for L0 ​Applicable when either or both of prediction or L1 prediction are intra-picture This can improve prediction efficiency and coding efficiency.

[0105] Also, by not scaling the differential motion vector based on the inter-picture distance like the temporal merge candidate, the circuit scale and power consumption can be reduced. For example, if the differential motion vector is scaled, when the temporal merge candidate is selected as the selection merge candidate, both scaling of the temporal merge candidate and scaling of the differential motion vector are required. When the differential motion vector is scaled, when the temporal merge candidate is selected as the selection merge candidate, both scaling of the temporal merge candidate and scaling of the differential motion vector are necessary. The scaling of the temporal merge candidate and the scaling of the differential motion vector have different motion vectors as the scaling reference, so both scalings cannot be executed together and need to be executed separately.

[0106] Also, when the temporal merge candidate is included in the merge candidate list as in this embodiment, the temporal merge candidate is scaled, and if the differential motion vector is smaller than the motion vector of the temporal merge candidate, the coding efficiency can be improved without scaling the differential motion vector. Also, when the differential motion vector is large, by selecting the differential motion vector mode, a decrease in coding efficiency can be suppressed. [Modification Example 9] In this embodiment, the maximum number of merge candidates when the merge correction flag is 0 and 1 is the same. In this modification example, the maximum number of merge candidates when the merge correction flag is 1 is made less than the maximum number of merge candidates when the merge correction flag is 0. For example, the maximum number of merge candidates when the merge correction flag is 1 is set to 2. Here, when the merge correction flag is 1, ​​​​​Set the maximum number of merge candidates for the combination as the maximum corrected merge candidate number. Also, when the merge index is less than the maximum corrected merge candidate number, encode (decode) the merge correction flag, and when the merge index is greater than or equal to the maximum corrected merge candidate number, do not encode (decode) the merge correction flag. Here, the maximum number of merge candidates and the maximum corrected merge candidate number when the merge correction flag is 0 may be predetermined values, or may be obtained by encoding (decoding) into the SPS or PPS in the encoded stream.

[0107] Thus, in this modified example, by making the maximum number of merge candidates when the merge correction flag is 1 less than the maximum number of merge candidates when the merge correction flag is 0, it is possible to determine whether to correct the merge candidate only for merge candidates with a higher selection probability, thereby reducing the processing of the coding replacement while suppressing a decrease in coding efficiency. Also, when the merge index is greater than or equal to the maximum corrected merge candidate number, there is no need to encode (decode) the merge correction flag, so the coding efficiency is improved. [Second Embodiment] The configurations of the image encoding device 100 and the image decoding device 200 in the second embodiment are the same as those of the image encoding device 100 and the image decoding device 200 in the first embodiment. The operation and syntax of the merge mode in this embodiment are different from those in the first embodiment. Hereinafter, the differences between this embodiment and the first embodiment will be described.

[0108] FIG. 13 is a flowchart for explaining the operation of the merge mode in the second embodiment. FIG. 14 shows a part of the syntax of the block that is the merge mode in the second embodiment. ​​​​​​​​It is a figure. FIG. 15 is a diagram showing the syntax of differential motion vectors in the second embodiment. is.

[0109] Hereinafter, the differences from the first embodiment will be described with reference to FIGS. 13, 14, and 15. FIG. 13 is different from FIG. 4 and steps S205 to S207, and steps S209 to step S211.

[0110] If the merge correction flag is 1 (YES in S104), it is checked whether the L0 prediction of the selected merge candidate is invalid (S205). If the L0 prediction of the selected merge candidate is not invalid (NO in S20 5), the process proceeds to step S208. If the L0 prediction of the selected merge candidate is invalid (YES in S2 05), the corrected motion vector of the L0 prediction is determined (S206). In the inter prediction unit 120 of the image encoding device 100, the corrected motion vector of the L0 prediction is obtained by motion vector search. Here, the search range of the motion vector is set to ±1 in both the horizontal and vertical directions.

[0111] In the inter prediction unit 220 of the image decoding device 200, the corrected motion vector of the L0 prediction is obtained from the encoded stream. Here, the search range of the motion vector is set to ±1 in both the horizontal and vertical directions. In the inter prediction unit 220 of the image decoding device 200, the corrected motion vector of the L0 prediction is obtained from the encoded stream. Subsequently, the reference picture index of the L0 prediction is determined (S207). Here, the reference picture index of the L0 prediction is set to 0.

[0112] Subsequently, it is checked whether the slice type is B and the L1 prediction of the selected merge candidate is invalid (S208). If the slice type is not B or the L1 prediction of the selected merge candidate is not invalid (NO in S208), the process proceeds to S111. If the slice type is B and the selected merge candidate

[0113] Subsequently, it is checked whether the slice type is B and the L1 prediction of the selected merge candidate is invalid (S208). If the slice type is not B or the L1 prediction of the selected merge candidate is not invalid (NO in S208), the process proceeds to S111. If the slice type is B and the selected merge candidate is not B or the L1 prediction of the selected merge candidate is not invalid (NO in S208), the process proceeds to S111. If the slice type is B and the selected merge candidate If the supplementary L1 prediction is invalid (YES in S208), determine the correction motion vector for the L1 prediction (S209).

[0114] In the inter prediction unit 120 of the image encoding device 100, the correction motion vector for the L1 prediction is obtained by motion vector search. Here, the search range of the motion vector is set to ±1 both in the horizontal direction and the vertical direction In the inter prediction unit 220 of the image decoding device 200, the correction motion vector for the L1 prediction is obtained from the encoded stream.

[0115] Subsequently, determine the reference picture index for the L1 prediction (S110). Here, the reference picture index for the L1 prediction is set to 0.

[0116] As described above, in this embodiment, in the case of a slice type that permits dual prediction (i.e., slice type B), the merge candidate for L0 prediction or L1 prediction is converted into a merge candidate for dual prediction. By using the merge candidate for dual prediction, it is possible to expect an improvement in prediction efficiency due to the filtering effect. Also, by using the most recently decoded image as the reference picture, the search range of the motion vector can be minimized. [Modification Example] In this embodiment, the reference picture index in step S207 and step S210 is set to 0. In this modification example, when the L0 prediction of the selected merge candidate is invalid, the reference picture index of the L0 prediction is set to the reference picture index of the L1 prediction, and when the L1 prediction of the selected merge candidate is invalid, the reference picture index of the L1 prediction is set to the reference picture index of the L0 prediction.

[0117] Thus, by slightly shifting and filtering the predicted value based on the motion vector of the L0 prediction or L1 prediction of the selection merge candidate, a minute motion can be reproduced, and the prediction efficiency can be improved.

[0118] In all of the embodiments described above, the encoded bitstream output by the image encoding device is specified to have a data format that can be decoded according to the encoding method used in the embodiment. The encoded bitstream has a data format that can be read by a computer such as an HDD, SSD, flash memory, optical disk, etc., and can be recorded and provided on a recording medium, or can be provided from a server through a wired or wireless network. Accordingly, the image decoding device corresponding to this image encoding device can decode the encoded bitstream of this specific data format regardless of the providing means. When a wired or wireless network is used to exchange the encoded bitstream between the image encoding device and the image decoding device, the encoded bitstream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission device that converts the encoded bitstream output by the image encoding device into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a receiving device that receives the encoded data from the network, restores it to an encoded bitstream, and supplies it to the image decoding device are provided. The transmission device includes a memory that buffers the encoded bitstream output by the image encoding device, a packet processing unit that packets the encoded bitstream, and a transmission unit that transmits the packet through the network.

[0119] To exchange the encoded bitstream between the image encoding device and the image decoding device, when a wired or wireless network is used, the encoded bitstream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, the encoded bitstream output by the image encoding device is converted into encoded data in a data format suitable for the transmission form of the communication path and transmitted to the network. A transmission device is provided, and a receiving device that receives the encoded data from the network, restores it to an encoded bitstream, and supplies it to the image decoding device is provided. The transmission device includes a memory that buffers the encoded bitstream output by the image encoding device, a packet processing unit that packets the encoded bitstream, and a transmission unit that transmits the packet through the network. The receiving device includes a reception unit that receives the encoded data from the network, a restoration unit that restores the encoded data to an encoded bitstream, and a supply unit that supplies the encoded bitstream to the image decoding device. The transmission device includes a memory that buffers the encoded bitstream output by the image encoding device, a packet processing unit that packets the encoded bitstream, and a transmission unit that transmits the packet through the network. and a transmitting unit for transmitting the coded data. a receiving unit for receiving packetized coded data and a buffer for buffering the received coded data; The memory for storing the encoded data is a memory for processing packets to generate an encoded bit stream and for displaying the image. and a packet processing unit for providing the packet to an image decoding device.

[0120] In order to exchange encoded bitstreams between an image encoding device and an image decoding device, When a wired or wireless network is used, in addition to a transmitting device and a receiving device, Even if a relay device is provided to receive the encoded data transmitted by the transmitting device and supply it to the receiving device, The relay device includes a receiving section for receiving packetized encoded data transmitted from the transmitting device. a memory for buffering the received coded data; and a transmitting unit for transmitting the packetized coded data to the network. a reception packet processing unit that processes the received data into packets to generate an encoded bit stream; A recording medium for storing the encoded bit stream and a transmission device for packetizing the encoded bit stream. The communication packet processing unit may include a communication packet processing unit.

[0121] In addition, by adding a display unit for displaying the image decoded by the image decoding device to the configuration, Also, an imaging unit may be added to the configuration, and the captured image may be input to the image encoding device. By doing so, it may also be used as an imaging device.

[0122] FIG. 18 shows an example of a hardware configuration of the encoding / decoding device of the present application. The present invention includes the configurations of an image encoding device and an image decoding device according to the embodiments of the present invention. The encoding / decoding device 9000 includes a CPU 9001, a codec IC 9002, an I / O interface 9003, a memory 9004, an optical disk drive 9005, a network interface 9006, and a video interface 9009, and each component is connected by a bus 9010 .

[0123] The image encoding unit 9007 and the image decoding unit 9008 are typically implemented as the codec IC 9002 . The image encoding process of the image encoding device according to the embodiment of the present invention is executed by the image encoding unit 9007, and the image decoding process in the image decoding device according to the embodiment of the present invention is executed by the image encoding unit 9007. The I / O interface 9003 is realized by, for example, a USB interface and is connected to an external keyboard 9104, a mouse 9 105, etc. The CPU 9001 controls the encoding / decoding device 9 000 to execute the operation desired by the user based on the user operation input via the I / O interface 9003. Examples of user operations using the keyboard 9104, the mouse 9105, etc. include selection of which function of encoding or decoding to execute, setting of encoding quality, input / output destination of the encoding stream , input / output destination of the image, etc. When the user desires to perform an operation to play an image recorded on the disk recording medium 9100, the optical disk drive 9005 reads an encoded bit stream from the inserted disk recording medium 9100 and sends the read encoded stream to the image decoding unit 9008 of the codec IC 9002 via the bus 9010. The image decoding unit 9008 performs the image decoding process in the image decoding device according to the embodiment of the present invention on the input encoded bit stream .

[0124] . Execute and send the decoded image to an external monitor 9103 via the video interface 9009. Also, the encoding / decoding device 9000 has a network interface 9006 and can be connected to an external distribution server 9106 or a mobile terminal 9107 via the network 9101. If the user desires to play back an image recorded on a distribution server 9106 or a mobile terminal 9107 instead of the image recorded on the disk recording medium 9100, the network interface 9006 obtains an encoded stream from the network 9101 instead of reading an encoded bitstream from the input disk recording medium 9100. Also, if the user desires to play back an image recorded in the memory 9004, image decoding processing in the image decoding device according to the embodiment of the present invention is executed on the encoded stream recorded in the memory 9004.

[0125] If the user desires to encode an image captured by an external camera 9102 and record it in the memory 9004, the video interface 9009 receives the image from the camera 9102 and sends it via the bus 9010 to the image encoding unit 9007 of the codec IC 9002. The image encoding unit 9007 executes image encoding processing in the image encoding device according to the embodiment of the present invention on the image input via the video interface 9009 and creates an encoded bitstream. Then, the encoded bitstream is sent via the bus 9010 to the memory 9004. If the user desires to record the encoded stream on the disk recording medium 9100 instead of the memory 9004, the optical disk drive 9005 inserts The encoded stream is written to the inserted disk recording medium 9100.

[0126] It is also possible to realize a hardware configuration that has an image encoding device and does not have an image decoding device, or a hardware configuration that has an image decoding device and does not have an image encoding device. Such a hardware configuration is realized, for example, by replacing the codec IC 9002 with the image encoding unit 9007 or the image decoding unit 9008, respectively.

[0127] The above processes related to encoding and decoding can of course be realized as a transmission device, a storage device, and a receiving device using hardware such as an ASIC. They can also be realized by firmware stored in a ROM (Read Only Memory), a flash memory, etc., or by software of a computer such as a CPU or a Soc (System on a chip). It is also possible to record the firmware program and the software program on a recording medium readable by a computer and provide them, or to provide them from a server through a wired or wireless network work, or to provide them as data broadcast of terrestrial or satellite digital broadcast.

[0128] As described above, the present invention has been described based on the embodiments. The embodiments are examples, and various modifications are possible for each combination of these components and each processing process. It is understood by those skilled in the art that such modifications are also within the scope of the present invention.

Explanation of Signs

[0129] 100 Image encoding device, 110 Block size determination unit, 120 Inter prediction section, 121 merge mode determination section, 122 merge candidate list generation section, 123 m erge candidate selection section, 124 merge candidate correction determination section, 125 merge candidate correction section, 1 26 differential motion vector mode execution section, 127 predicted value derivation section, 130 conversion section, 140 code sequence generation section, 150 local decoding section, 160 frame memory, 200 image decoding device, 201 spatial merge candidate generation section, 202 temporal merge candidate generation section, 203 merge candidate supplementation section, 210 code sequence decoding section, 220 inter prediction section, 2 30 inverse conversion section, 240 frame memory.

Claims

1. a merge candidate list generation unit that generates a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to a block to be predicted and a motion vector of a block in an encoded image that is at the same position as the block to be predicted; and a merge candidate supplementation unit that adds a supplementary merge candidate having a motion vector of (0, 0) to the merge candidate list; a merging candidate selection unit for selecting a merging candidate from the merging candidate list as a selected merging candidate; a merge correction determination unit that sets a merge correction flag indicating whether or not to correct the merge candidate; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction unit that, when both a reference picture of a first prediction of the selected merge candidate and a reference picture of a second prediction of the selected merge candidate are not long-term reference pictures, subtracts the correction vector scaled according to an interval between a prediction target picture and a reference picture from a motion vector of a second prediction of the selected merge candidate, to derive a bi-predictive corrected merge candidate; a code string encoding unit that encodes the merge correction flag and the correction vector into a coded stream; An image encoding device comprising:

2. a merging candidate list generating step of generating a merging candidate list including, as merging candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and a motion vector of a block in an encoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplementary merge candidate having a motion vector (0, 0) to the merge candidate list; selecting a merge candidate from the merge candidate list as a selected merge candidate; a merge correction determination step of setting a merge correction flag indicating whether or not to correct the merge candidate; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the prediction target picture and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, to derive a bi-predictive corrected merge candidate; a code string encoding step of encoding the merge correction flag and the correction vector into a coded stream; 13. An image encoding method comprising the steps of:

3. a merging candidate list generating step of generating a merging candidate list including, as merging candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and a motion vector of a block in an encoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplementary merge candidate having a motion vector (0, 0) to the merge candidate list; selecting a merge candidate from the merge candidate list as a selected merge candidate; a merge correction determination step of setting a merge correction flag indicating whether or not to correct the merge candidate; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the prediction target picture and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, to derive a bi-predictive corrected merge candidate; a code string encoding step of encoding the merge correction flag and the correction vector into a coded stream; An image encoding program characterized by causing a computer to execute the above steps.

4. a merge candidate list generating unit configured to generate a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to a block to be predicted and a motion vector of a block in a decoded image that is located at the same position as the block to be predicted; a merge candidate supplementation unit that adds a supplementary merge candidate having a motion vector of (0, 0) to the merge candidate list; a merging candidate selection unit for selecting a merging candidate from the merging candidate list as a selected merging candidate; a code string decoding unit that decodes a code string from the encoded stream and derives a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction unit that, when both a reference picture of a first prediction of the selected merge candidate and a reference picture of a second prediction of the selected merge candidate are not long-term reference pictures, subtracts the correction vector scaled according to an interval between a prediction target picture and a reference picture from a motion vector of a second prediction of the selected merge candidate, to derive a bi-predictive corrected merge candidate; An image decoding device comprising:

5. a merging candidate list generating step of generating a merging candidate list including, as merging candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the prediction target block and a motion vector of a block in a decoded image that is at the same position as the prediction target block; a merge candidate supplementation step of adding a supplementary merge candidate having a motion vector (0, 0) to the merge candidate list; selecting a merge candidate from the merge candidate list as a selected merge candidate; a code string decoding step of decoding a code string from the encoded stream to derive a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the prediction target picture and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, to derive a bi-predictive corrected merge candidate; 13. An image decoding method comprising:

6. a merging candidate list generating step of generating a merging candidate list including, as merging candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the prediction target block and a motion vector of a block in a decoded image that is at the same position as the prediction target block; a merge candidate supplementation step of adding a supplementary merge candidate having a motion vector (0, 0) to the merge candidate list; selecting a merge candidate from the merge candidate list as a selected merge candidate; a code string decoding step of decoding a code string from the encoded stream to derive a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; when the merge correction flag indicates that a merge candidate is to be corrected, and either one or both of a first prediction reference picture of the selected merge candidate or a second prediction reference picture of the selected merge candidate are long-term reference pictures, and the first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate are in opposite directions with respect to a prediction target picture including the prediction target block, adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling, thereby deriving a bi-predictive corrected merge candidate; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the prediction target picture and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, to derive a bi-predictive corrected merge candidate; An image decoding program characterized by causing a computer to execute the above steps.

7. 3. A method for generating a bit stream in accordance with the image coding method according to claim 2, and storing the bit stream on a recording medium.

8. A transmission method for generating a bit stream according to the image coding method according to claim 2 and transmitting the bit stream.

Citation Information

Patent Citations

  • Moving image encoding and decoding device using motion compensation inter-frame prediction system capable of area integration

    JP1998276439A