Image decoding device, image decoding method, and image decoding program

The image decoding device enhances coding efficiency by adjusting motion vectors in HEVC's merge mode through a merge candidate list and correction mechanism, addressing limitations in existing technologies.

JP7810300B2Active Publication Date: 2026-02-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025042224
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-13
Filing Date
2025-03-17
Publication Date
2026-02-03
Estimated Expiration
2039-09-20

AI Technical Summary

Technical Problem

Existing image coding technologies like HEVC have limitations in coding efficiency, particularly in the merge mode of inter prediction, where motion vectors can be improved for better performance.

Method used

An image decoding device and method that includes a merge candidate list generation unit, a merge candidate selection unit, and a merge correction unit to adjust motion vectors based on adjacent block information and decoded image positions, using a merge correction flag to enhance coding efficiency.

Benefits of technology

The proposed solution provides a new inter prediction mode with higher efficiency by reducing the number of selectable merge candidates and improving coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007810300000002
    Figure 0007810300000002
  • Figure 0007810300000003
    Figure 0007810300000003
  • Figure 0007810300000004
    Figure 0007810300000004
Patent Text Reader

Abstract

To provide an inter prediction mode with higher efficiency by correcting a motion vector in a merge mode.SOLUTION: In an image decoding method, a merge candidate list is generated; a merge candidate is selected from the merge candidate list as a selected merge candidate; a code string is decoded from an encoded stream to derive a correction vector; the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling; and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the correction merge candidate.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image decoding technique. [Background technology]

[0002] There are image coding technologies such as HEVC (H.265). HEVC uses inter-prediction mode. Merge mode is used as the mode. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 10-276439 Summary of the Invention [Problem to be solved by the invention]

[0004] In HEVC, there are two inter prediction modes: merge mode and differential motion vector mode. However, there is room for further improvement in coding efficiency by correcting the motion vectors in merge mode. In particular, the inventor has come to realize this.

[0005] The present invention has been made in view of the above situation, and its purpose is to By correcting vectors, we provide a new inter prediction mode that is more efficient. The reason is that. [Means for solving the problem]

[0006] In order to solve the above problem, an image decoding device according to one aspect of the present invention provides a prediction target block. The motion information of the adjacent blocks and the decoded image at the same position as the block to be predicted are motion information including a motion vector derived by scaling the motion vector of the block a merge candidate list generation unit that generates a merge candidate list including the merge candidates; a merge candidate selection unit for selecting a merge candidate from the merge candidate list as a merge candidate; The code string is decoded from the encoded stream and a merge correction flag is generated to indicate whether or not to correct the merge candidate. a code string decoding unit for deriving a merge correction flag and a correction vector; If the selected merge candidate indicates that the first prediction of the selected merge candidate is to be a reference picture of the first prediction of the selected merge candidate or the selected merge candidate, one or both of the reference pictures of the second prediction of the candidate image are long-term reference pictures; The first prediction reference picture of the selected merge candidate and the second prediction reference picture of the selected merge candidate If the picture to be decoded is in the opposite direction to the picture containing the block to be predicted, The correction vector is not scaled to the motion vector of the first prediction of the selected merge candidate. The correction vector is added as it is, and the motion vector of the second prediction of the selected merge candidate is shifted. Merge candidate correction that derives bi-predictive correction merge candidates by subtracting without scaling and if the merge correction flag indicates that the merge candidate is to be corrected, Compared to the case where correction is not indicated, the number of selectable merge candidates is reduced. stomach.

[0007] Any combination of the above components, and the expression of the present invention may be used as a method, an apparatus, a system, a recording medium, Conversions between the body, computer program, etc. are also valid aspects of the present invention. be. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a new inter prediction mode with higher efficiency. can. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1(a) is a diagram illustrating the configuration of an image encoding device 100 according to the first embodiment, and FIG. 1(b) is a diagram illustrating the configuration of an image decoding device 200 according to the first embodiment. [Figure 2] FIG. 10 is a diagram showing an example in which an input image is divided into blocks based on the block size. [Figure 3] FIG. 2 is a diagram illustrating the configuration of an inter-prediction unit of the image encoding device of FIG. [Figure 4] 10 is a flowchart illustrating an operation in a merge mode according to the first embodiment. [Figure 5] 1(a) is a diagram illustrating the configuration of a merge candidate list generation unit of the image encoding device of FIG. [Figure 6] FIG. 10 is a diagram illustrating blocks adjacent to a processing target block. [Figure 7] FIG. 10 is a diagram illustrating blocks in a decoded image that are located at the same position as the target block and in its periphery. [Figure 8] FIG. 10 is a diagram illustrating a part of the syntax of a block in a merge mode according to the first embodiment. [Figure 9] FIG. 3 is a diagram illustrating the syntax of a differential motion vector according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating the syntax of a differential motion vector according to a modification of the first embodiment. [Figure 11] FIG. 10 is a diagram illustrating a part of the syntax of a block in a merge mode according to another modified example of the first embodiment. [Figure 12] FIG. 10 is a diagram illustrating a part of the syntax of a block in a merge mode according to yet another modified example of the first embodiment. [Figure 13] 10 is a flowchart illustrating an operation in a merge mode according to the second embodiment. [Figure 14] FIG. 10 is a diagram illustrating a part of the syntax of a block in a merge mode according to the second embodiment. [Figure 15]FIG. 10 is a diagram illustrating the syntax of a differential motion vector according to the second embodiment. [Figure 16] FIG. 13 is a diagram illustrating the effect of the eighth modification of the first embodiment. [Figure 17] FIG. 13 is a diagram illustrating the effect of the eighth modification of the first embodiment when picture intervals are not equal. [Figure 18] FIG. 2 is a diagram illustrating an example of a hardware configuration of a coding / decoding device according to a first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] [First embodiment] Hereinafter, an image coding apparatus and an image coding method according to a first embodiment of the present invention will be described with reference to the drawings. , and image encoding program, and image decoding device, image decoding method, and image decoding program The details of RAM will be explained.

[0011] FIG. 1(a) is a diagram illustrating the configuration of an image coding device 100 according to the first embodiment. FIG. 1(b) is a diagram illustrating the configuration of an image decoding device 200 according to the first embodiment. do.

[0012] The image coding device 100 of this embodiment includes a block size determination unit 110, an inter-prediction unit, and a a conversion unit 130, a code string generation unit 140, a local decoding unit 150, and a frame memory The image encoding device 100 receives an input image and performs intra-prediction and intra-prediction. Then, center prediction is performed and the coded stream is output. Hereafter, the terms image and picture have the same meaning. Use.

[0013] The image decoding device 200 includes a code string decoding unit 210, an inter prediction unit 220, and an inverse transform unit 230. and a frame memory 240. The image decoding device 200 is different from the image encoding device 100 in that The encoded stream is input, and intra prediction and inter prediction are performed. Output the image.

[0014] The image encoding device 100 and the image decoding device 200 are implemented by a CPU (Central Processor). It is implemented by hardware such as an information processing unit (ISU) equipped with memory, etc. It will be revealed.

[0015] First, the functions and operations of each unit of the image encoding device 100 will be described. It is assumed that inter prediction is performed in the same manner as in HEVC, and inter prediction will be described below.

[0016] The block size determination unit 110 determines the block size for inter prediction based on the input image. The input image corresponding to the determined block size, block position and block size is The block size is determined by the inter prediction unit 120. For example, the RDO (rate distortion optimization) method used in the HEVC reference software is used. Use.

[0017] Here, the block size will be explained. A part of the image to be displayed is divided into blocks based on the block size determined by the block size determination unit 110. The following shows an example of a block divided into blocks. The block sizes are 4x4, 8x4, 4x8, 8×8, 16×8, 8×16, 32×32, , 128×64, 64×128, 12 The input image is one of the above so that each block does not overlap. It is divided into blocks of size.

[0018] The inter prediction unit 120 performs a frame prediction based on the information input from the block size determination unit 110. The inter prediction is performed by using a reference picture input from the memory 160. The inter prediction unit 120 determines the inter prediction parameters. Inter prediction is performed to derive the predicted value, and the block size, block position, input value, and inter prediction are used. The prediction parameters and the predicted values ​​are supplied to the transform unit 130. The method is based on the RDO (Rate Distortion Optimization) method used in the HEVC reference software. The inter prediction parameters and the operation of the inter prediction unit 120 will be described in detail later. do.

[0019] The conversion unit 130 calculates a difference value by subtracting the predicted value from the input value, and performs an orthogonal Transformation and quantization are performed to calculate the prediction error data, and the block size and block The position, inter-prediction parameters, and prediction error data are transmitted to the codestream generator 140 and locally decoded. The signal is supplied to the section 150.

[0020] The code string generation unit 140 generates an SPS (Sequence Parameter Set) as needed. Set), PPS (Picture Parameter Set) and other information and encodes the code string for determining the block size supplied from the conversion unit 130. The inter-prediction parameters are coded as a code string, and the prediction error data is coded as a code string. For details on inter prediction parameter coding, see More details will be provided later.

[0021] The local decoding unit 150 performs processes such as inverse orthogonal transform and inverse quantization on the prediction error data to obtain a differential The difference value and the predicted value are added to generate a decoded image, and the decoded image and the inter-prediction parameters are The meters are fed to a frame memory 160 .

[0022] The frame memory 160 stores decoded images and inter-prediction parameters for multiple images. The decoded image and the inter prediction parameters are supplied to the inter prediction unit 120 .

[0023] Next, the function and operation of each unit of the image decoding device 200 will be described. It is assumed that inter prediction is performed in the same manner as EVC, and inter prediction will be described below.

[0024] The code stream decoding unit 210 extracts SPS, PPS headers, and other information from the coded stream as needed. and other information such as block size, block position, inter prediction parameters, and The prediction error data is decoded from the coded stream, and the block size, block position, and The inter-prediction unit 220 supplies the inter-prediction parameters and prediction error data.

[0025] The inter-prediction unit 220 receives the information input from the code stream decoding unit 210 and the frame memory 2 The prediction value is derived by performing inter prediction using the reference picture input from 40. The inter-prediction unit 220 determines the block size, block position, inter-prediction parameters, and prediction The error data and the predicted value are supplied to the inverse transform unit 230 .

[0026] The inverse transform unit 230 performs an inverse orthogonal transform on the prediction error data supplied from the inter prediction unit 220. The difference value is calculated by performing processes such as inverse quantization and dequantization, and the difference value and the predicted value are added to generate a decoded image. The decoded image and the inter-prediction parameters are supplied to the frame memory 240, and the decoded image is Output.

[0027] The frame memory 240 stores decoded images and inter-prediction parameters for multiple images. The decoded image and the inter prediction parameters are supplied to the inter prediction unit 220 .

[0028] Note that the inter prediction performed by the inter prediction unit 120 and the inter prediction unit 220 is the same. The decoded image and the image stored in the frame memory 160 and the frame memory 240 are The center prediction parameters are also the same.

[0029] Next, the inter-prediction parameters will be described. page flag, merge index, valid flag of predicted LX, motion vector of predicted LX, The reference picture index of the predicted LX, the merge correction flag, and the differential motion vector of the predicted LX. LX is L0 and L1. Merge flag is used to mark as inter prediction mode. This flag indicates whether to use the motion vector mode or the differential motion vector mode. If the merge flag is 1, the merge mode is used, and if the merge flag is 0, the differential motion vector is used. The merge index is the position of the selected merge candidate in the merge candidate list. The valid flag of the predicted LX is a flag indicating whether the predicted LX is valid or invalid. If both L0 prediction and L1 prediction are enabled, it is bi-predictive; if L0 prediction is enabled and L1 prediction is enabled, it is bi-predictive. If L1 prediction is enabled and L0 prediction is disabled, L1 prediction is used. The motion correction flag indicates whether or not to correct the motion information of the merge candidate. If the correction flag is 1, the merge candidate is corrected. If the correction flag is 0, the merge candidate is Here, the codestream generator 140 sets the validity flag of the predicted LX in the coded stream. The code string decoder 210 does not encode the predicted LX as a code string. The reference picture index is not decoded as a code string from the encoded stream. It is an index for identifying a decoded image in the L0 prediction area 160. L1 prediction valid flag, L0 prediction motion vector, L1 prediction motion vector, L0 prediction combination of the reference picture index of the L1 prediction and the reference picture index of the L2 prediction The motion information is

[0030] If the block is in intra-coding mode or outside the image area, L0 The prediction valid flag and the L1 prediction valid flag are both invalid.

[0031] Hereafter, the picture types are L0 unidirectional prediction, L1 unidirectional prediction and bidirectional prediction. Although all the predictions are available for B-pictures, the picture type is only unidirectional prediction. It may be an available P picture. In the case of a P picture, the inter prediction parameters are L Only L0 prediction is considered, and L1 prediction is treated as non-existent. To improve coding efficiency in the L0 prediction, the reference picture for L0 prediction is set to a picture earlier than the picture to be predicted. The reference picture for L1 prediction is a picture in the future than the picture to be predicted. This is because the reference picture for L0 prediction and the reference picture for L1 prediction are the same as the picture to be predicted. This is because interpolative prediction improves coding efficiency when the image is in the opposite direction from the L0 image. Whether the reference picture for L1 prediction and the reference picture for L2 prediction are in the opposite direction from the picture to be predicted The difference is that the POC (Picture Order Count) of the reference picture is compared. From now on, the reference picture for L0 prediction and the reference picture for L1 prediction will be used as the prediction target. The block will be described as being in the reverse direction in time relative to the picture to be predicted.

[0032] Next, the details of the inter prediction unit 120 will be described. The inter-prediction unit 120 of the image decoding device 100 and the inter-prediction unit 220 of the image decoding device 200 are The structure and operation are identical.

[0033] FIG. 3 is a diagram illustrating the configuration of the inter prediction unit 120. The inter prediction unit 120 a merge mode determination unit 121, a merge candidate list generation unit 122, a merge candidate selection unit 123, and a merge A merge candidate correction determination unit 124, a merge candidate correction unit 125, a differential motion vector mode implementation unit 1 26 and a predicted value derivation unit 127.

[0034] The inter prediction unit 120 selects a merge mode as the inter prediction mode and a differential motion vector The differential motion vector mode execution unit 126 switches the mode for each block. The merged motion vector mode is assumed to be implemented in the same way as HEVC. This section explains the do.

[0035] The merge mode determination unit 121 determines the merge mode as the inter prediction mode for each block. If the merge flag is 1, the merge mode is used, and if the merge flag is 2, the merge mode is used. If the lag is 0, the differential motion vector mode is used.

[0036] In the inter prediction unit 120, the determination of whether to set the merge flag to 1 is made based on the HEVC We use the RDO (rate distortion optimization) method used in the reference software. In the prediction unit 220, the code stream decoding unit 210 extracts the Gets the decoded merge flag. The syntax is explained in detail below.

[0037] If the merge flag is 0, the differential motion vector mode execution unit 126 executes the differential motion vector. The inter prediction parameter of the differential motion vector mode is It is supplied to 127.

[0038] If the merge flag is 1, the merge candidate list generator 122 and the merge candidate selector 12 3. The merge mode is performed by the merge candidate correction determination unit 124 and the merge candidate correction unit 125. The merge mode inter prediction parameters are supplied to the prediction value derivation unit 127.

[0039] The process when the merge flag is 1 will be explained in detail below.

[0040] FIG. 4 is a flowchart explaining the operation of the merge mode. Hereinafter, FIGS. 3 and 4 will be used. The merge mode will be explained in detail below.

[0041] First, the merge candidate list generation unit 122 calculates the motion of blocks adjacent to the target block. A merge candidate list is generated from the block information and the motion information of the decoded image (S100). The generated merge candidate list is supplied to the merge candidate selection unit 123. The terms and the block to be predicted are used interchangeably.

[0042] Here, the generation of the merge candidate list will be explained. The merge candidate list generation unit 122 generates a spatial merge candidate list. The merge candidate generation unit 201 includes a temporal merge candidate generation unit 202 and a merge candidate supplementation unit 203 .

[0043] FIG. 6 is a diagram illustrating blocks adjacent to a processing target block. The blocks adjacent to the elephant block are Block A, Block B, Block C, Block D, Block Block E, block F, and block G are blocks adjacent to the target block. If a lock is used, the present invention is not limited to this.

[0044] Figure 7 illustrates the blocks in the decoded image that are located at the same position as the target block and its surroundings. Here, the decoded image at the same position as the target block and its periphery is The blocks are CO1, CO2, and CO3. This is not limited to using multiple blocks in the decoded image that are in the same position as the block and its surroundings. Thereafter, blocks CO1, CO2, and CO3 are the same position blocks, and the same position blocks are the same position blocks. A decoded picture containing the same is called a co-located picture.

[0045] Hereinafter, the generation of the merge candidate list will be described in detail with reference to FIGS.

[0046] First, the spatial merge candidate generation unit 201 generates a spatial merge candidate from blocks A, B, C, and D. Blocks D, E, F, and G are checked in order and the valid flags of L0 prediction are checked. If either or both of the L1 prediction valid flags are valid, the motion information of the block is The spatial merge candidate generator 201 generates the merge candidates and adds them to the merge candidate list. The merge candidates are called spatial merge candidates.

[0047] Next, the temporal merge candidate generation unit 202 generates a merge candidate from blocks CO1, CO2, and C O3 is checked sequentially and first either the L0 prediction valid flag or the L1 prediction valid flag is checked. The motion information of the blocks where both are valid is scaled and then selected as merge candidates. The merge candidates generated by the temporal merge candidate generator 202 are added to the merge candidate list in sequence. are called temporal merge candidates.

[0048] Here, we will explain the scaling of temporal merge candidates. The motion vector of the temporal merge candidate is the motion vector of the co-located block. The distance vector between the co-located picture and the reference picture that the co-located block refers to is calculated as follows: Based on this, the distance between the picture containing the block to be predicted and the picture referenced by the temporal merge candidate is calculated. It is derived by scaling with the distance.

[0049] Note that the pictures referenced by the temporal merge candidates are reference pictures for both L0 prediction and L1 prediction. The reference picture has a channel index of 0. Also, the L0 block is used as the co-located block. Whether the co-located block of prediction or the co-located block of L1 prediction is used depends on the co-located block. The position derivation flag is coded (decoded) and judged. Either the L0 prediction or the L1 prediction of the position block. The motion vectors of the L0 and L1 predictions are derived by scaling the predictions. are the motion vectors of the L0 prediction and L1 prediction of the image candidate.

[0050] Next, if the merge candidate list contains multiple identical motion information, one motion information is merged. The remaining movement information is deleted.

[0051] Next, if the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, In this case, the merge candidate supplementation unit 203 selects the merge candidate with the smallest number of merge candidates in the merge candidate list. The merge candidates are added to the merge candidate list until the maximum number of merge candidates is reached. The number of merge candidates in the supplementary list is the maximum number of merge candidates. A candidate is a frame whose motion vectors for L0 prediction and L1 prediction are both (0,0) and whose motion vectors for L0 prediction and L1 prediction are both (0,0). is motion information in which both reference picture indexes are 0.

[0052] Here, the maximum number of merge candidates is set to 6, but it may be any number greater than or equal to 1.

[0053] Next, the merge candidate selection unit 123 selects one merge candidate from the merge candidate list. (S101), the selected merge candidate (called the “selected merge candidate”) and the merge index to the merge candidate correction determination unit 124, and the selected merge candidate is determined based on the motion information of the target block. The inter prediction unit 120 of the image encoding device 100 uses the HEVC reference software. The merge candidate list is then compiled using the RDO (Rate Distortion Optimization) method used in the One merge candidate is selected from the merge candidates and a merge index is determined. In the inter prediction unit 220 of the device 200, the code stream decoding unit 210 decodes the coded stream. The merge index is obtained and a list of merge candidates is created based on the merge index. Select one merge candidate from the included merge candidates as the selected merge candidate.

[0054] Next, the merge candidate correction determination unit 124 determines whether the width of the processing target block is equal to or greater than a predetermined width and whether the processing target block is equal to or smaller than a predetermined width. The height of the block to be processed is greater than or equal to a predetermined height, and both the L0 prediction and the L1 prediction of the selected merge candidate are Or, it is checked whether at least one of them is valid (S102). The width is greater than or equal to a certain value, the height of the block to be processed is greater than or equal to a certain height, and the L0 prediction of the selected merge candidate is If the condition that both or at least one of the L1 predictions is valid is not met (S1 02 NO), without correcting the selected merge candidate as the motion information of the processing target block, The process proceeds to step S111. Here, the merge candidate list always includes at least L0 prediction and L1 prediction. Since there are merge candidates with at least one valid L0 prediction and L1 prediction of the selected merge candidate, It is obvious that both or at least one of the measures is valid. "Are both or at least one of the L0 and L1 predictions of the merge candidate valid?" omitted Then, in step S102, the width of the processing target block is equal to or greater than a predetermined width and the height of the processing target block is equal to or greater than a predetermined width. It may also be possible to check whether the height is equal to or greater than the height.

[0055] The width of the processing target block is equal to or greater than a predetermined width, and the height of the processing target block is equal to or greater than a predetermined height, and If both or at least one of the L0 and L1 predictions of the selected merge candidate is valid (S 102 (YES), the merge candidate correction determination unit 124 sets the merge correction flag (S10 3) The merge correction flag is supplied to the merge candidate correction unit 125. The inter prediction unit 120 calculates a prediction error when performing inter prediction using merge candidates by using a predetermined prediction error. If the error is greater than or equal to the error, the merge correction flag is set to 1 and inter prediction is performed using the selected merge candidate. If the prediction error is not equal to or greater than the predetermined prediction error, the merge correction flag is set to 0. In the inter-prediction unit 220 of the image decoding device 200, the code stream decoding unit 210 performs the following operations based on the syntax: The merge correction flag is obtained by decoded from the coded stream based on the above.

[0056] Next, the merge candidate correction unit 125 checks whether the merge correction flag is 1 (S10 4) If the merge correction flag is not 1 (NO in S104), the selected merge candidate is processed. The process proceeds to step S111 without correcting the motion information of the block.

[0057] If the merge correction flag is 1 (YES in S104), the L0 prediction of the selected merge candidate is valid. If the L0 prediction of the selected merge candidate is not valid (S105), If the L0 prediction of the selected merge candidate is valid (NO in S15), proceed to step S108. 05 YES), the differential motion vector for L0 prediction is determined (S106). If the merge correction flag is 1, the motion information of the selected merge candidate is corrected, and the merge correction flag If is 0, the motion information of the selected merge candidate is not corrected.

[0058] In the inter prediction unit 120 of the image encoding device 100, the motion vector difference of the L0 prediction is Here, the search range for the motion vector is horizontal and vertical. Both are ±16, but any multiple of 2 such as ±64 may be used. In the prediction unit 220, the code stream decoding unit 210 extracts the encoded data from the coded stream based on the syntax. The differential motion vector of the L0 prediction decoded from the L0 prediction is obtained.

[0059] Subsequently, the merge candidate correction unit 125 calculates a corrected motion vector for L0 prediction, The corrected motion vector of the prediction is set as the motion vector of the L0 prediction of the motion information of the block to be processed. (S107).

[0060] Here, the corrected motion vector of the L0 prediction (mvL0), the motion vector of the L0 prediction of the selected merge candidate, The relationship between the L0 prediction differential motion vector (mvdL0) and the L0 prediction differential motion vector (mmvL0) is explained. The corrected motion vector for L0 prediction (mvL0) is the motion vector for L0 prediction of the selected merge candidate. The L0 prediction differential motion vector (mmvL0) is added to the L0 prediction differential motion vector (mvdL0). This results in the following formula: [0] is the horizontal component of the motion vector, and [1] is the horizontal component of the motion vector. 1 shows the vertical component of torque.

[0061] mvL0[0] = mmvL0[0] + mvdL0[0] mvL0[1] = mmvL0[1] + mvdL0[1] Subsequently, it is checked whether the L1 prediction of the selected merge candidate is valid (S108). If the L1 prediction of the page candidate is not valid (NO in S108), the process proceeds to step S111. If the L1 prediction of the selected merge candidate is valid (YES in S108), the differential motion vector of the L1 prediction is The torque is determined (S109).

[0062] In the inter prediction unit 120 of the image encoding device 100, the motion vector difference of the L1 prediction is Here, the search range for the motion vector is horizontal and vertical. Both are ±16, but are powers of 2 such as ±64. In the measurement unit 220, the code stream decoding unit 210 decodes the coded stream based on the syntax. The differential motion vector of the coded L1 prediction is obtained.

[0063] Subsequently, the merge candidate correction unit 125 calculates a corrected motion vector for L1 prediction, The corrected prediction motion vector is set as the L1 prediction motion vector of the motion information of the target block. (S110).

[0064] Here, the corrected motion vector of the L1 prediction (mvL1), the motion vector of the L1 prediction of the selected merge candidate, The relationship between the L1 prediction differential motion vector (mvdL1) and the L1 prediction differential motion vector (mmvL1) is explained. The corrected motion vector of L1 prediction (mvL1) is the motion vector of L1 prediction of the selected merge candidate. The L1 prediction differential motion vector (mvdL1) is added to the L2 prediction differential motion vector (mmvL1). This results in the following formula: [0] is the horizontal component of the motion vector, and [1] is the horizontal component of the motion vector. 1 shows the vertical component of torque.

[0065] mvL1[0] = mmvL1[0] + mvdL1[0] mvL1[1] = mmvL1[1] + mvdL1[1] Next, the prediction value derivation unit 127 performs L0 prediction, Either L1 prediction or bi-prediction is performed to derive a predicted value (S111). As described above, if the merge correction flag is 1, the motion vector of the selected merge candidate is corrected. If the merge correction flag is 0, the motion vector of the selected merge candidate is not corrected.

[0066] The coding of inter-prediction parameters will be described in detail. Table 1 shows a part of the syntax of the block to be coded. The relationship between the cbWidth and cbH in Fig. 8 is the width of the block to be processed. Eight is the height of the block to be processed. The specified width and height are both 8. By setting a predetermined height, the merge candidate is not corrected in small block units. This reduces the amount of processing. Here, cu_skip_flag is the flag for whether a block is skipped. It is 1 if the mode is skip mode, and 0 if not. The syntax is the same as for merge mode. merge_idx is the merge candidate. This is the merge index that selects the selected merge candidate from the list.

[0067] merge_idx is encoded (decoded) before merge_mod_flag. After the page index is confirmed, the merge_mod_flag is encoded (decoded) By determining that, merge_idx is used as the merge index for the merge mode, To improve coding efficiency while suppressing the complexity of syntax and the increase in context.

[0068] [Table 1]

[0069] FIG. 9 is a diagram showing the syntax of the differential motion vector. g(N) has the same syntax as that used in the differential motion vector mode. Here, N is either 0 or 1. N=0 indicates L0 prediction, and N=1 indicates L1 prediction.

[0070] The syntax of the differential motion vector specifies whether the components of the differential motion vector are greater than 0. abs_mvd_greater0_flag[d] is a flag indicating the difference motion vector. abs_mvd_greater1 is a flag indicating whether the torque component is greater than 1. _flag[d], mvd_sign_fl indicating the sign (±) of the differential motion vector component ag[d], abs_m indicates the absolute value of the vector obtained by subtracting 2 from the component of the differential motion vector vd_minus2[d] is included, where d is either 0 or 1. d=0 is horizontal directional component, and d=1 indicates the vertical component.

[0071] In HEVC, there are two inter prediction modes: merge mode and differential motion vector mode. In merge mode, motion information can be restored with just one merge flag, resulting in very efficient coding. However, the motion information in merge mode depends on the processed blocks. Therefore, the cases where the prediction efficiency is high are limited, and it is necessary to further improve the utilization efficiency. Ta.

[0072] On the other hand, the differential motion vector mode provides separate syntax for L0 prediction and L1 prediction. The prediction type (L0 prediction, L1 prediction, or bi-prediction) and the L0 and L1 predictions The predicted motion vector flag, the differential motion vector, and the reference picture index are Therefore, the differential motion vector mode has better coding efficiency than the merge mode. However, the merge mode cannot derive motions of spatially adjacent blocks or temporally adjacent blocks. The prediction efficiency is high and stable for sudden movements that have little correlation with the movements of adjacent blocks. It was a mode that became more and more popular.

[0073] In this embodiment, the prediction type and reference picture index of the merge mode are fixed. By enabling the correction of the merge mode motion vectors as they are, the coding efficiency is improved by using the differential motion vectors. This can improve the utilization efficiency more than the merge mode.

[0074] In addition, by making the difference motion vector the difference between the motion vector of the selected merge candidate, The magnitude of the motion vector can be reduced, and the coding efficiency can be reduced. .

[0075] Also, the syntax of the differential motion vector in merge mode is changed to the differential motion vector in differential motion vector mode. By making the syntax the same as that of the differential motion vector, the differential motion vector can be used in merge mode. Even if it is added, the change in the configuration can be kept to a minimum.

[0076] Also, a predetermined width and a predetermined height are defined, and the predicted block width is equal to or greater than the predetermined width and the predicted block height is equal to or greater than the predetermined width. is greater than or equal to a predetermined value, and both or at least one of the L0 prediction and L1 prediction of the selected merge candidate is If the condition that the merge mode is effective is not met, the process of correcting the motion vectors of the merge mode is By eliminating this processing, the amount of processing required to correct the motion vectors in merge mode can be reduced. In addition, when it is not necessary to suppress the amount of processing for correcting the motion vector, the predetermined width and the predetermined height are There is no need to limit the correction of the motion vectors by length.

[0077] The following describes variations of this embodiment. Unless otherwise specified, the variations are not combined with each other. It can be combined. [Variation 1] In this embodiment, the differential motion vector is used as the syntax for a block in merge mode. In this modification, the differential motion vector is coded (or encoded) as a differential unit motion vector. A differential unit motion vector is defined as a motion vector that is generated at the minimum picture interval. In HEVC, the minimum picture interval is the length of the coded stream. It is encoded as a code string in the program.

[0078] The differential motion vector is calculated according to the interval between the picture to be coded and the reference picture in merge mode. The POC( Picture Order Count) in POC(Cur), merge mode L0 preview The POC of the predicted reference picture is POC(L0), and the POC of the reference picture of L1 prediction in merge mode is POC(L1). If POC is POC(L1), the motion vector is calculated as follows: umvdL 0 is the differential unit motion vector of L0 prediction, and umvdL1 is the differential unit motion vector of L1 prediction. Indicates the rule.

[0079] mvL0[0] = mmvL0[0] + umvdL0[0]*(POC(Cur )-POC(L0)) mvL0[1] = mmvL0[1] + umvdL0[1]*(POC(Cur )-POC(L0)) mvL1[0] = mmvL1[0] + umvdL1[0]*(POC(Cur )-POC(L1)) mvL1[1] = mmvL1[1] + umvdL1[1]*(POC(Cur )-POC(L1)) As described above, in this modification, the differential unit motion vector is used as the differential motion vector. By doing so, the amount of code for the differential motion vector is reduced, thereby improving the coding efficiency. In addition, when the differential motion vector is large and the distance between the prediction target picture and the reference picture is small, In particular, the coding efficiency can be improved when the picture to be processed is large. Even when the distance between the image and the reference picture is proportional to the speed of the object moving on the screen, the prediction efficiency is improved. This can improve the coding efficiency.

[0080] In addition, there is no need to derive the temporal merge candidates by scaling them with the inter-picture distance. In the decoding device, scaling can be performed using only a multiplier, so a divider is not required, and the circuit size and processing This can reduce the amount of processing required. [Variation 2] In this embodiment, 0 can be coded (or decoded) as a component of the differential motion vector. In this modification, only the L0 prediction can be changed. It is assumed that 0 cannot be encoded (or decoded) as a component of .

[0081] FIG. 10 is a diagram showing the syntax of the differential motion vector in the second modification. The syntax of the TOR has a flag that indicates whether the component of the differential motion vector is greater than 1. abs_mvd_greater1_flag[d], the component of the differential motion vector is 2 abs_mvd_greater2_flag[d ], abs_mvd_m, which indicates the absolute value of the vector obtained by subtracting 3 from the component of the differential motion vector inus3[d], mvd_sign_fl indicating the sign (±) of the differential motion vector component ag[d] is included.

[0082] As mentioned above, 0 cannot be coded (or decoded) as a component of the differential motion vector. By setting the value to 1, it is possible to improve the coding efficiency when the component of the differential motion vector is 1 or more. This can be done. [Variation 3] In this embodiment, the components of the differential motion vector are integers, and in the second modification, they are integers excluding 0. In this modification, the components of the differential motion vector, excluding the ± signs, are limited to powers of two.

[0083] Instead of the syntax abs_mvd_minus2[d] in this embodiment, bs_mvd_pow_plus1[d] is used. The differential motion vector mvd[d] is m From vd_sign_flag[d] and abs_mvd_pow_plus1[d], It is calculated as follows.

[0084] mvd[d]=mvd_sign_flag[d]*2^(abs_mvd_po w_plus1[d]+1) Also, instead of the syntax abs_mvd_minus3[d] in variant 2, abs_mvd_pow_plus2[d] is used. The differential motion vector mvd[d] is Below mvd_sign_flag[d] and abs_mvd_pow_plus2[d] It is calculated as follows:

[0085] mvd[d]=mvd_sign_flag[d]*2^(abs_mvd_po w_plus2[d]+2) By limiting the components of the differential motion vector to powers of 2, the processing load of the encoding device is significantly reduced. While reducing the number of frames, it is possible to improve prediction efficiency in the case of large motion vectors. [Variation 4] In this embodiment, mvd_coding(N) includes a differential motion vector. However, in this modification, mvd_coding(N) includes a motion vector magnification.

[0086] In the syntax of mvd_coding(N) in this modification, abs_mvd_gre ater0_flag[d], abs_mvd_greater1_flag[d], m vd_sign_flag[d] does not exist, instead abs_mvr_plus 2[d] and mvr_sign_flag[d].

[0087] The corrected motion vector of the LN prediction (mvLN) is the motion vector of the LN prediction of the selected merge candidate. The motion vector magnification (mvrLN) is multiplied by the motion vector magnification (mmvLN) using the following formula: is calculated.

[0088] mvLN[d] = mmvLN[d] * mvrLN[d] By limiting the components of the differential motion vector to powers of 2, the processing load of the encoding device is significantly reduced. While reducing the number of frames, it is possible to improve prediction efficiency in the case of large motion vectors.

[0089] This modification example cannot be combined with modification examples 1, 2, 3, and 6. I can't do that. [Variation 5] In the syntax of FIG. 8 of this embodiment, cu_skip_flag is 1 (skip merge_mod_flag may be present if However, if you are in skip mode, merge_mod_flag does not exist. You may do so.

[0090] In this way, by omitting merge_mod_flag, the skip mode code This improves the coding efficiency and simplifies the skip mode decision. [Variation 6] In this embodiment, the LN prediction (N=0 or 1) of the selected merge candidate is checked to see if it is valid. However, if the LN prediction of the selected merge candidate is not valid, the differential motion vector is not valid. However, the LN prediction of the selected merge candidate is not checked to see if it is valid. The differential motion vector may be valid regardless of whether prediction is valid or not. If the LN prediction of the selected merge candidate is invalid, the motion vector of the LN prediction of the selected merge candidate is The reference picture index of the selected merge candidate is set to (0,0), and the reference picture index of the LN prediction of the selected merge candidate is set to 0.

[0091] In this way, in this modification, regardless of whether the LN prediction of the selected merge candidate is valid or not, By enabling differential motion vectors, the opportunities to use bidirectional prediction are increased, and coding Efficiency can be improved. [Variation 7] In this embodiment, whether the L0 prediction and the L1 prediction of the selected merge candidate are valid or not is determined individually. Whether to encode (or decode) the differential motion vector is determined by the decision, but the selective merge If both L0 and L1 predictions of a candidate are valid, the differential motion vector is coded (or If both L0 and L1 predictions of the selected merge candidate are invalid, the differential motion vector is used. In this modification, the vector may not be encoded (or decoded). Step S102 is as follows:

[0092] The merge candidate correction determination unit 124 determines whether the width of the processing target block is equal to or greater than a predetermined width and whether the processing target block is The block height is greater than or equal to a predetermined height, and both the L0 prediction and the L1 prediction of the selected merge candidate are valid. It is checked whether it exists (S102).

[0093] Furthermore, steps S105 and S108 are not necessary in this modification.

[0094] FIG. 11 is a diagram showing a part of the syntax of a block in the merge mode of the seventh modification. The syntax for steps S102, S105, and S108 is different.

[0095] In this way, in this modification, both L0 prediction and L1 prediction of the selected merge candidate are valid. By enabling differential motion vectors in this case, the selection and merging of frequently used bidirectional prediction candidates is improved. By correcting the complementary motion vector, prediction efficiency can be improved efficiently. [Variation 8] In this embodiment, the syntax of the block in merge mode is the difference of L0 prediction. Two differential motion vectors were used: the motion vector and the differential motion vector of L1 prediction. In this example, only one differential motion vector is coded (or decoded) to obtain the corrective motion vector for L0 prediction. A single differential motion vector is used as the correction motion vector for the vector and L1 prediction, and is calculated using the following formula: The motion vectors of the selected merge candidates mmvLN (N=0,1) and the differential motion vector m The motion vector mvLN (N=0, 1) of the corrected merge candidate is calculated from vd.

[0096] If the L0 prediction of the selected merge candidate is valid, the motion vector of the L0 prediction is calculated using the following formula: Calculate.

[0097] mvL0[0] = mmvL0[0] + mvd[0] mvL0[1] = mmvL0[1] + mvd[1] If the L1 prediction of the selected merge candidate is valid, the motion vector of the L1 prediction is calculated using the following formula: The L1 prediction of the selected merge candidate is calculated by adding the differential motion vector in the opposite direction to the L0 prediction. The differential motion vector may be subtracted from the measured motion vector.

[0098] mvL1[0] = mmvL1[0] + mvd[0]*-1 mvL1[1] = mmvL1[1] + mvd[1]*-1 FIG. 12 is a diagram showing a part of the syntax of a block in the merge mode of the eighth modification. The validity of L0 prediction and validity of L1 prediction have been removed, and mvd_ The difference from this embodiment is that there is no coding(1). It corresponds to one differential motion vector.

[0099] In this way, in this modification, only one differential motion vector is used for L0 prediction and L1 prediction. By defining By sharing it with 1 prediction, the coding efficiency can be improved while suppressing the decrease in prediction efficiency. can be done.

[0100] In addition, the reference picture referred to in the L0 prediction of the selected merge candidate and the reference picture referred to in the L1 prediction are If the picture is in the opposite direction (not the same direction) to the picture being predicted, the differential motion is By adding the vector in the opposite direction, the coding efficiency for motion in a certain direction is improved. It is possible.

[0101] The effects of this modification will be described in detail. Figure 16 is a diagram for explaining the effects of modification 8. Figure 16 shows a sphere (diagonal) moving horizontally within a moving rectangular area (area surrounded by a dashed line). In this case, the movement of the sphere relative to the screen The motion of the rectangular area is the sum of the motion of the sphere moving horizontally. Picture A is the reference picture for L0 prediction, and picture C is the reference picture for L1 prediction. Picture A and picture C are in the opposite direction from the picture to be predicted. is the reference picture in

[0102] If the sphere moves in a certain direction at a certain speed, then picture A, picture B, and picture C will be If they are equally spaced, the amount of movement of the sphere that cannot be obtained from the adjacent blocks is added to the L0 prediction. By subtracting this from the L1 prediction, the sphere's movement can be accurately reproduced.

[0103] If the sphere moves in a certain direction but at a non-constant speed, then Picture A, Picture B, Picture C, Although the distances between the C and C are not equal, the distances between the C and C corresponding to the rectangular areas of the sphere are equal. The amount of movement of the sphere that cannot be obtained from adjacent blocks is added to the L0 prediction and subtracted from the L1 prediction. This allows the movement of the sphere to be accurately reproduced.

[0104] Also, if a sphere moves at a constant speed in a certain direction for a certain period of time, Picture A and Picture B , picture C are not equally spaced, but the movement amounts corresponding to the rectangular areas of the sphere are equally spaced. FIG. 17 illustrates the effect of the eighth modification when the picture intervals are not equal. This example will be described in detail with reference to FIG. 17. , F8 indicate pictures at fixed intervals. The sphere is stationary in the picture F5 and moves at a constant speed in a constant direction from picture F6 onwards. Pictures F0 and F6 are reference pictures, and picture F5 is the picture to be predicted. In this case, pictures F0, F5, and F6 are not equally spaced, but are arranged in the same order as the sphere. The amount of movement corresponding to the rectangular area is equal. When picture F5 is the picture to be predicted, Generally, picture F4, which is close to the target, is selected as the reference picture. The reason for selecting picture F0 as the reference picture instead of picture F4 is that picture F0 is more likely to be a reference picture than picture F4. The reference picture is usually stored in a FIFO (First In First Out) It is managed in the First-In First-Out (First-In First-Out) method, but it is a high-quality picture with little distortion. There is a long-term reference picture as a mechanism for using a reference picture. Applies when either L0 prediction or L1 prediction or both are long-term reference pictures This makes it possible to improve prediction efficiency and coding efficiency. Apply when either prediction or L1 prediction or both are intra pictures. This makes it possible to improve prediction efficiency and coding efficiency.

[0105] Also, like temporal merge candidates, the differential motion vectors are scaled based on the inter-picture distance. By not using the differential motion vector filter, the circuit size and power consumption can be reduced. When scaling vectors, if a time merge candidate is selected as the selection merge candidate, Both scaling of the temporal merge candidates and scaling of the differential motion vectors are required. The scaling of temporal merge candidates and the scaling of differential motion vectors are the scaling criteria. Since the motion vectors are different, both scalings cannot be performed together and are performed separately. It is necessary to carry out

[0106] Furthermore, when the merge candidate list includes temporal merge candidates as in this embodiment, The temporal merge candidate is scaled and has a differential motion vector that is smaller than the motion vector of the temporal merge candidate. If the difference motion vector is small, the coding efficiency can be improved without scaling the difference motion vector. In addition, if the difference motion vector is large, the difference motion vector motion can be By selecting the code, it is possible to suppress a decrease in coding efficiency. [Variation 9] In this embodiment, the maximum number of merge candidates is the same when the merge correction flag is 0 and when it is 1. In this modification, the maximum number of merge candidates when the merge correction flag is 1 is set as the merge correction flag. The number of merge candidates is set to be less than the maximum number when the positive flag is 0. For example, the merge correction flag The maximum number of merge candidates when the merge correction flag is 1 is 2. The maximum number of merge candidates when the merge index is the maximum corrected number of merge candidates. If the number of merge correction candidates is smaller than the number of merge correction candidates, the merge correction flag is encoded (decoded) and the merge correction flag is If the index is equal to or greater than the maximum number of correction merge candidates, the merge correction flag is coded (decoded). Here, the maximum number of merge candidates when the merge correction flag is 0 is The maximum number of corrective merge candidates may be a predetermined value or may be a value that is It may also be obtained by encoding (decoding) into SPS or PPS.

[0107] In this way, in this modification, the maximum number of merge candidates when the merge correction flag is 1 is By reducing the number of merge candidates below the maximum number when the page correction flag is 0, the selection probability is increased. Only for merge candidates with high quality, the coding device determines whether to correct the merge candidate. This makes it possible to reduce the processing time for the merge index and suppress the decrease in coding efficiency. If the number of merge correction candidates is equal to or greater than the maximum number of merge correction candidates, the merge correction flag is coded (decoded). This eliminates the need for a multi-bit coding scheme, thereby improving coding efficiency. [Second embodiment] The configurations of the image encoding device 100 and the image decoding device 200 of the second embodiment are the same as those of the first embodiment. The image encoding device 100 and the image decoding device 200 are the same as those of the first embodiment. The operation and syntax of the merge mode differ from the first embodiment. The differences from the first embodiment will be described.

[0108] FIG. 13 is a flowchart illustrating the operation of the merge mode according to the second embodiment. FIG. 14 shows part of the syntax of a block in merge mode according to the second embodiment. FIG. 15 is a diagram showing the syntax of the differential motion vector according to the second embodiment. do.

[0109] Hereinafter, differences from the first embodiment will be described with reference to FIGS. 13, 14, and 15. FIG. 13 shows the same process as FIG. 4, from step S205 to step S207, and from step S209 to step S208. Step S211 is different.

[0110] If the merge correction flag is 1 (YES in S104), the L0 prediction of the selected merge candidate is invalid. If the L0 prediction of the selected merge candidate is not invalid (S205), If the L0 prediction of the selected merge candidate is invalid (NO in S2 If the result of step 05 is YES, a corrected motion vector for L0 prediction is determined (S206).

[0111] In the inter prediction unit 120 of the image encoding device 100, the corrected motion vector for L0 prediction is Here, the search range for the motion vector is horizontal and vertical. In the inter prediction unit 220 of the image decoding device 200, the correction motion of the L0 prediction is The vector is obtained from the coded stream.

[0112] Next, the reference picture index for L0 prediction is determined (S207). , the reference picture index for L0 prediction is set to 0.

[0113] Next, check whether the slice type is B and whether the L1 prediction of the selected merge candidate is invalid. (S208) If the slice type is not B or the L1 prediction of the selected merge candidate is invalid, If there is no slice (NO in S208), proceed to S111. If the complementary L1 prediction is invalid (YES in S208), a correction motion vector for the L1 prediction is determined. (S209).

[0114] In the inter prediction unit 120 of the image encoding device 100, the corrected motion vector for L1 prediction is Here, the search range for the motion vector is horizontal and vertical. In the inter prediction unit 220 of the image decoding device 200, the correction motion of the L1 prediction is The vector is obtained from the coded stream.

[0115] Next, the reference picture index for L1 prediction is determined (S110). , the reference picture index for L1 prediction is set to 0.

[0116] As described above, in this embodiment, the slice type is a slice type that allows bi-prediction. For slice type B, L0 or L1 prediction merge candidates are bi-predicted. By using it as a bi-predictive merge candidate, the filtering effect is reduced. This can be expected to improve prediction efficiency. By doing so, the search range for the motion vector can be minimized. [Variations] In this embodiment, the reference picture indexes in steps S207 and S210 In this modification, when the L0 prediction of the selected merge candidate is invalid, the reference of the L0 prediction is The reference picture index of the selected merge candidate is used as the reference picture index of L1 prediction. If L1 prediction is disabled, the reference picture index of L1 prediction is set to the reference picture index of L0 prediction. The index is the Cha index.

[0117] In this way, the predicted value of the motion vector of the L0 prediction or L1 prediction of the selected merge candidate is By filtering with a small shift, it is possible to reproduce small movements, improving prediction efficiency. can be improved.

[0118] In all the embodiments described above, the coded bitstream output by the image coding device The stream is specified so that it can be decoded according to the encoding method used in the embodiment. The encoded bitstream can be stored on HDD, SSD, flash drive, Provided by recording it on a computer-readable recording medium such as flash memory or optical disk. It may be provided by a server via a wired or wireless network. Therefore, an image decoding device corresponding to this image encoding device can decode this specific data regardless of the providing means. It is possible to decode coded bitstreams in the data format.

[0119] In order to exchange coded bitstreams between an image coding device and an image decoding device, When a wired or wireless network is used, the data format appropriate for the transmission mode of the communication path In this case, the image coding device may convert the output bitstream into The coded bit stream is converted into coded data in a data format suitable for the transmission mode of the communication channel. a transmitting device that converts the encoded data into a digital signal and transmits it to a network; A receiving device is provided for restoring the encoded bit stream to a coded bit stream and supplying it to the image decoding device. The receiving device includes a memory for buffering the coded bit stream output by the image coding device; A packet processing unit that packetizes the coded bit stream and a packet processing unit that transmits the packet through the network. and a transmitting unit for transmitting the coded data. a receiving unit for receiving packetized coded data and a buffer for buffering the received coded data; The memory for storing the encoded data is also configured to process the encoded data into packets to generate an encoded bit stream. and a packet processing unit for providing the packet to an image decoding device.

[0120] In order to exchange coded bitstreams between an image coding device and an image decoding device, When a wired or wireless network is used, in addition to the transmitting device and the receiving device, Even if a relay device is provided to receive the coded data transmitted by the transmitting device and supply it to the receiving device, The relay device has a receiving section that receives packetized coded data transmitted from the transmitting device. a memory for buffering the received coded data; and a transmitting unit for transmitting the packetized coded data to the network. a reception packet processing unit that processes the received data into packets to generate an encoded bit stream; a recording medium for storing the encoded bitstream and a transmission medium for packetizing the encoded bitstream. The communication packet processing unit may also include a communication packet processing unit.

[0121] In addition, by adding a display unit for displaying the image decoded by the image decoding device to the configuration, Alternatively, an imaging unit may be added to the configuration, and the captured image may be input to the image encoding device. By doing so, it may also be used as an imaging device.

[0122] FIG. 18 shows an example of the hardware configuration of the encoding / decoding device of the present application. The present invention includes the configurations of an image encoding device and an image decoding device according to the embodiments of the present invention. The encoding / decoding device 9000 includes a CPU 9001, a codec IC 9002, an I / O interface interface 9003, memory 9004, optical disk drive 9005, network interface The interface 9006 and the video interface 9009 are connected to each other via a bus 9010. are connected by

[0123] The image encoding unit 9007 and the image decoding unit 9008 are typically implemented by a codec IC 9002 and The image encoding process of the image encoding device according to the embodiment of the present invention is implemented as an image encoding The image decoding process is executed by the decoding unit 9007 in the image decoding device according to the embodiment of the present invention. The encoding process is performed by the image encoding unit 9007. The I / O interface 9003 For example, a USB interface can be used to connect an external keyboard 9104 and mouse 9 The CPU 9001 receives input via the I / O interface 9003. The encoding / decoding device 9 executes the operation desired by the user based on the user's operation. 000. As a user operation using the keyboard 9104, mouse 9105, etc. The CODEC allows you to select whether to perform encoding or decoding functions, set the encoding quality, and There are input / output destinations for frames, input / output destinations for images, etc.

[0124] When the user desires to play back images recorded on the disk recording medium 9100 The optical disc drive 9005 reads the encoded video from the inserted disc recording medium 9100. The read encoded stream is sent to the decoder via bus 9010. The image decoding unit 9008 of the block IC 9002 receives the coded image data and sends it to the image decoding unit 9008. The image decoding process in the image decoding device according to the embodiment of the present invention is performed on the bit stream. The decoded image is then displayed on an external monitor 9103 via a video interface 9009. The encoding / decoding device 9000 also includes a network interface 9006. , and connects to an external distribution server 9106 and a mobile terminal 9107 via a network 9101. The user can change the image recorded on the disk recording medium 9100 and use it on the distribution server. If you want to play back the images recorded on the server 9106 or the mobile terminal 9107, The network interface 9006 receives the code from the input disk recording medium 9100. Instead of reading the encoded bitstream, the network 9101 Also, when the user desires to play back the images recorded in the memory 9004, In this case, the embodiment of the present invention is applied to the coded stream recorded in the memory 9004. The image decoding device according to the present invention performs image decoding processing.

[0125] The user captures an image using an external camera 9102 and encodes it into memory 9004. When an operation is desired, the video interface 9009 receives an image from the camera 9102. The image data is input and sent to the image encoding unit 9007 of the codec IC 9002 via the bus 9010. The image encoding unit 9007 encodes the image input via the video interface 9009. The image encoding process is performed in the image encoding device according to the embodiment of the present invention, and the encoded bits are The encoded bitstream is then sent to the memory via bus 9010. The user changes the memory 9004 and codes it on the disk recording medium 9100. If it is desired to record the encrypted stream, the optical disc drive 9005 may be The encoded stream is written to the inserted disc recording medium 9100.

[0126] A hardware configuration that has an image encoding device but does not have an image decoding device, or a hardware configuration that has an image decoding device However, it is also possible to realize a hardware configuration that does not include an image coding device. The hardware configuration is, for example, a codec IC 9002, an image encoding unit 9007, or This is realized by replacing the image decoding unit 9008 with the image decoding unit 9009.

[0127] The above encoding and decoding processes are carried out by a transmission device using hardware such as an ASIC, It can be realized as a storage device, a receiving device, and also as a ROM (Read-On Firmware stored in the internal memory (internal memory) or flash memory, and the firmware stored in the CPU or SoC This can also be achieved by computer software such as (System on a chip) The firmware and software programs can be The information may be provided by recording it on a computer-readable recording medium or by providing it over a wired or wireless network. It can also be provided from a server via a network, or it can be used to provide terrestrial or satellite digital broadcasting data. It can also be provided as a broadcast.

[0128] The present invention has been described above based on the embodiments. The embodiments are merely examples, and the respective structures thereof are not intended to be limiting. The fact that various variations are possible in the combination of components and each treatment process, and that such variations It will be understood by those skilled in the art that such modifications are also within the scope of the present invention. [Explanation of symbols]

[0129] 100 Image encoding device, 110 Block size determination unit, 120 Inter prediction 121 merge mode determination unit, 122 merge candidate list generation unit, 123 merge 124 merge candidate correction determination unit; 125 merge candidate correction unit; 26 differential motion vector mode implementation unit, 127 predicted value derivation unit, 130 conversion unit, 140 code string generation unit, 150 local decoding unit, 160 frame memory, 200 Image decoding device, 201 spatial merge candidate generation unit, 202 temporal merge candidate generation unit, 203 merge candidate supplementation unit, 210 code string decoding unit, 220 inter prediction unit, 2 30 inverse transform units, 240 frame memories.

Claims

1. a merge candidate list generation unit that generates a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in an encoded image that is at the same position as the block to be predicted; and a merge candidate supplementation unit that adds a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merge candidate selector that selects a merge candidate from the merge candidate list as a selected merge candidate; a merge correction determination unit that sets a merge correction flag indicating whether or not to correct the merge candidate; If the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction unit that, when both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, subtracts the correction vector scaled according to an interval between the picture to be predicted and the reference picture from a motion vector of the second prediction of the selected merge candidate, to derive a bi-predictive corrected merge candidate; a code string encoding unit that encodes the merge correction flag and the correction vector into a coded stream; An image encoding device comprising:

2. a merge candidate list generation step of generating a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in the encoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merging candidate selection step of selecting a merging candidate from the merging candidate list as a selected merging candidate; a merge correction determination step of setting a merge correction flag indicating whether or not to correct the merge candidate; If the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the picture to be predicted and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures; and a code string encoding step of encoding the merge correction flag and the correction vector into a coded stream; An image encoding method comprising:

3. a merge candidate list generation step of generating a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in the encoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merging candidate selection step of selecting a merging candidate from the merging candidate list as a selected merging candidate; a merge correction determination step of setting a merge correction flag indicating whether or not to correct the merge candidate; If the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding a correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the picture to be predicted and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures; and a code string encoding step of encoding the merge correction flag and the correction vector into a coded stream; An image encoding program characterized by causing a computer to execute the above.

4. a merge candidate list generation unit that generates a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in a decoded image that is at the same position as the block to be predicted; and a merge candidate supplementation unit that adds a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merge candidate selector that selects a merge candidate from the merge candidate list as a selected merge candidate; a code string decoding unit that decodes a code string from the coded stream and derives a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; if the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction unit that, when both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures, subtracts the correction vector scaled according to an interval between the picture to be predicted and the reference picture from a motion vector of the second prediction of the selected merge candidate, to derive a bi-predictive corrected merge candidate; An image decoding device comprising:

5. a merge candidate list generation step of generating a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in the decoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merging candidate selection step of selecting a merging candidate from the merging candidate list as a selected merging candidate; a code string decoding step of decoding a code string from the coded stream to derive a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; if the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the picture to be predicted and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures; and An image decoding method comprising:

6. a merge candidate list generation step of generating a merge candidate list including, as merge candidates, motion information including motion vectors derived by scaling motion information of a plurality of blocks adjacent to the block to be predicted and motion vectors of a block in the decoded image that is at the same position as the block to be predicted; a merge candidate supplementation step of adding a supplemental merge candidate having a motion vector (0, 0) to the merge candidate list; a merging candidate selection step of selecting a merging candidate from the merging candidate list as a selected merging candidate; a code string decoding step of decoding a code string from the coded stream to derive a merge correction flag indicating whether or not to correct a merge candidate and a correction vector; if the merge correction flag indicates that a merge candidate is to be corrected, and one or both of the reference picture of a first prediction of the selected merge candidate and the reference picture of a second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to a picture to be predicted including the block to be predicted, deriving a bi-predictive corrected merge candidate by adding the correction vector to a motion vector of the first prediction of the selected merge candidate without scaling and subtracting the correction vector from a motion vector of the second prediction of the selected merge candidate without scaling; a merge candidate correction step of subtracting the correction vector scaled according to an interval between the picture to be predicted and a reference picture from a motion vector of the second prediction of the selected merge candidate, if both the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are not long-term reference pictures; and An image decoding program characterized by causing a computer to execute the above.

7. A storage method for generating a bitstream according to the image coding method of claim 2 and storing the bitstream on a recording medium.

8. A transmission method for generating a bitstream according to the image coding method of claim 2 and transmitting the bitstream.

Citation Information

Patent Citations

  • Moving image encoding and decoding device using motion compensation inter-frame prediction system capable of area integration

    JP1998276439A