Image decoding device, image decoding method, and image decoding program

The image decoding apparatus enhances encoding efficiency by correcting the motion vector in the merge mode through a merge candidate list and correction flag, addressing limitations in existing technologies like HEVC.

JP2026063275APending Publication Date: 2026-04-10JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
JVC KENWOOD CORP
Filing Date
2026-01-20
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing image encoding technologies, such as HEVC, have limitations in improving the encoding efficiency of the merge mode by correcting the motion vector, which affects the overall performance of inter prediction.

Method used

An image decoding apparatus that includes a merge candidate list generation unit, a merge candidate selection unit, and a merge correction flag to correct the motion vector of the merge mode, utilizing motion information from adjacent blocks and reference pictures to enhance encoding efficiency.

Benefits of technology

The proposed solution provides a more efficient inter prediction mode by correcting the motion vector, resulting in improved encoding efficiency and reduced processing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063275000001_ABST
    Figure 2026063275000001_ABST
Patent Text Reader

Abstract

By correcting the motion vector in merge mode, the interface becomes more efficient. Provides a prediction mode. [Solution] A merge candidate list is generated, and merge candidates are selected from the merge candidate list. Select as a complement, decode the code sequence from the encoded stream to derive a correction vector, and select the corrected vector. The first predicted motion vector of the candidate is added without scaling to the correction vector, and then selected. The correction vector is subtracted from the second predicted motion vector of the merge candidate without scaling to compensate. Derive the positive merge candidates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image decoding technology.

Background Art

[0002] There is image encoding technology such as HEVC (H.265). In HEVC, the merge mode is used as an inter prediction mode. and the differential motion vector mode.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In HEVC, there are a merge mode and a differential motion vector mode as inter prediction modes, but the inventor has recognized that there is room to further improve the encoding efficiency by correcting the motion vector of the merge mode. The present invention has been made in view of such a situation, and an object thereof is to provide a new inter prediction mode with higher efficiency by correcting the motion vector of the merge mode.

[0005]

Means for Solving the Problems

[0006] In order to solve the above problems, an image decoding apparatus according to an aspect of the present invention includes motion information including a motion vector derived by scaling the motion information of a plurality of blocks adjacent to a prediction target block and the motion vector of a block at the same position on a decoded image as the prediction target block. ​​​​​​A merge candidate list generation unit generates a list of merge candidates to be included as merge candidates, and the A merge candidate selection unit that selects a merge candidate from a list of merge candidates, and a code A merge correction flag indicates whether or not to decode the code sequence from the modified stream and correct the merge candidates. A code sequence decoding unit that derives a correction vector, and the merge correction flag corrects the merge candidate. If it indicates that, the reference picture of the first prediction of the selected merge candidate or the selected merge If either or both of the reference pictures for the second prediction of candidate Z are long-term reference pictures, The reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate If the Kucha is in the opposite direction to the decryption target picture containing the predicted block, The correction vector is applied to the first predicted motion vector of the selected merge candidate without scaling. The correction vector is then added to the second predicted motion vector of the selected merge candidate. Merge candidate correction: Derive the corrected merge candidate for the dual prediction by subtracting directly without kaling. The system includes a section, and if the merge correction flag indicates that the merge candidate will be corrected, then the merge candidate Compared to the case where correction of the supplement is not indicated, the number of selectable merge candidates is smaller. stomach.

[0007] Furthermore, any combination of the above components, or any expression of the present invention, may be used to describe a method, apparatus, system, or recording medium. The conversion between the human body, computer programs, etc., is also valid as an embodiment of the present invention. be. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a new interprediction mode that is more efficient. can. [Brief explanation of the drawing]

[0009] [Figure 1] FIG. 1(a) is a diagram for explaining the configuration of an image encoding apparatus 100 according to the first embodiment, and FIG. 1(b) is a diagram for explaining the configuration of an image decoding apparatus 200 according to the first embodiment. [Figure 2] It is a diagram showing an example in which an input image is divided into blocks based on a block size. [Figure 3] It is a diagram for explaining the configuration of the inter prediction unit of the image encoding apparatus in FIG. 1(a). [Figure 4] It is a flowchart for explaining the operation of the merge mode according to the first embodiment. [Figure 5] It is a diagram for explaining the configuration of the merge candidate list generation unit of the image encoding apparatus in FIG. 1(a). [Figure 6] It is a diagram for explaining a block adjacent to a processing target block. [Figure 7] It is a diagram for explaining blocks on a decoded image at the same position as a processing target block and its periphery. [Figure 8] It is a diagram showing a part of the syntax of a block that is the merge mode according to the first embodiment. [Figure 9] It is a diagram showing the syntax of the differential motion vector according to the first embodiment. [Figure 10] It is a diagram showing the syntax of the differential motion vector of a certain modification example according to the first embodiment. [Figure 11] It is a diagram showing a part of the syntax of a block that is the merge mode of another modification example according to the first embodiment. [Figure 12] It is a diagram showing a part of the syntax of a block that is the merge mode of yet another modification example according to the first embodiment. [Figure 13] It is a flowchart for explaining the operation of the merge mode according to the second embodiment. [Figure 14] It is a diagram showing a part of the syntax of a block that is the merge mode according to the second embodiment. [Figure 15]This figure shows the syntax of the difference motion vector in the second embodiment. [Figure 16] This figure illustrates the effects of Modification 8 of the First Embodiment. [Figure 17] This figure illustrates the effect of the case where the picture spacing is not equal in Modification 8 of the First Embodiment. [Figure 18] This figure illustrates an example of the hardware configuration of the encoding / decoding device according to the first embodiment. [Modes for carrying out the invention]

[0010] [First Embodiment] Below, along with the drawings, are an image encoding apparatus and image encoding method according to the first embodiment of the present invention. , and image coding program, and image decoding device, image decoding method, and image decoding program Let me explain the details of the lamb.

[0011] Figure 1(a) is a diagram illustrating the configuration of the image encoding device 100 according to the first embodiment. Figure 1(b) is a diagram illustrating the configuration of the image decoding device 200 according to the first embodiment. ru.

[0012] The image encoding device 100 of this embodiment includes a block size determination unit 110 and an interpretation unit. Section 120, conversion section 130, code sequence generation section 140, local decoding section 150, and frame memo Includes Ri 160. The image encoding device 100 receives input of an input image and performs intra prediction and i Performs interface prediction and outputs an encoded stream. Hereafter, "image" and "picture" will be used interchangeably. Use.

[0013] The image decoding device 200 comprises a code sequence decoding unit 210, an interpretation unit 220, and an inverse transformation unit 230. , and frame memory 240. The image decoding device 200 is connected to the image encoding device 100. The encoded stream output is received as input, intra-prediction and inter-prediction are performed, and the result is recovered. Output the image.

[0014] The image encoding device 100 and the image decoding device 200 are CPUs (Central Processor). The system is implemented using hardware such as an information processing device equipped with a ssing unit, memory, etc. It will be revealed.

[0015] First, the functions and operations of each part of the image coding device 100 will be explained. Intra prediction is The process will be similar to that of HEVC, and the interpretation will be explained below.

[0016] The block size determination unit 110 predicts the block size based on the input image. Determine the block size, block position, and the input image corresponding to the block size. The raw values ​​(input values) are supplied to the interpretation unit 120. Regarding the method for determining the block size... For example, the RDO (Rate Distortion Optimization) method used in HEVC reference software, etc. Use.

[0017] Now, let's explain the block size. Figure 2 shows the input to the image encoding device 100. A portion of the image is determined based on the block size determined by the block size determination unit 110. Here are examples of how it is divided into blocks. The block sizes are 4x4, 8x4, 4x8, 8x8, 16x8, 8x16, 32x32, ..., 128x64, 64x128, 12 An 8x128 grid exists, and the input image must be one of the above, ensuring that each block does not overlap. It will be divided into block sizes.

[0018] The interpretation unit 120 uses the information input from the block size determination unit 110 and the frame Using the reference picture input from memory 160, the inter prediction used for inter prediction The measurement parameters are determined. The interpretation unit 120 determines the interpretation parameters based on the interpretation parameters. Interpret and derive predicted values, block size, block position, input value, inter The prediction parameters and predicted values ​​are supplied to the conversion unit 130. The interpretation parameters are determined. Regarding the method, RDO (Rate Distortion Optimization) is used in the HEVC reference software. Methods such as the ) method are used. Details of the interpretation parameters and the operation of the interpretation unit 120 will be described later. do.

[0019] The conversion unit 130 calculates a difference value by subtracting the predicted value from the input value, and then applies an orthogonal value to the calculated difference value. The prediction error data is calculated by performing processes such as transformation and quantization, and the block size and block size The position, interpretation parameters, and prediction error data are processed by the code sequence generation unit 140 and local decoding unit 140. It is supplied to section 150.

[0020] The code sequence generation unit 140 generates SPS (Sequence Parameter) as needed. Set), PPS (Picture Parameter Set), and other information Encode the code sequence for determining the block size supplied from the conversion unit 130. The prediction parameters are encoded as a code sequence, and the prediction error data is encoded as a code sequence. The code is encoded and the encoded stream is output. For details on encoding the interpretation parameters, see below. This will be explained later.

[0021] The local decoding unit 150 performs processing such as inverse orthogonal transformation and inverse quantization on the prediction error data to obtain the difference. The values ​​are restored, the difference values ​​and predicted values ​​are added together to generate a decoded image, and the decoded image and the predicted parameters are interpreted. The meter is supplied to the frame memory 160.

[0022] The frame memory 160 stores the decoded image and interpretation parameters for multiple images. The decoded image and interpretation parameters are supplied to the interpretation prediction unit 120.

[0023] Next, the functions and operations of each part of the image decoding device 200 will be explained. Intra prediction is H This will be implemented in the same way as EVC, and the inter-prediction will be explained below.

[0024] The code sequence decoding unit 210 processes the encoded stream as needed, including the SPS, PPS header, and other necessary information. Decode other information, including block size, block position, interpretation parameters, and Decode the prediction error data from the encoded stream, and input the block size, block position, and The predictor parameters and prediction error data are supplied to the interpretation unit 220.

[0025] The interpretation unit 220 receives information from the code sequence decoding unit 210 and the frame memory 2 Using the reference picture input from 40, the system performs interpretation and derives the predicted value. The predictor unit 220 receives the block size, block position, predictor parameters, and prediction. Error data and predicted values ​​are supplied to the inverse conversion unit 230.

[0026] The inverse transformation unit 230 performs an inverse orthogonal transformation on the prediction error data supplied from the interpretation unit 220. Then, processes such as inverse quantization are performed to calculate the difference value, and the difference value and the predicted value are added to produce the decoded image. The decoded image and interpretation parameters are then supplied to the frame memory 240, and the decoded image Output.

[0027] The frame memory 240 stores the decoded image and interpretation parameters for multiple images. The decoded image and interpretation parameters are supplied to the interpretation prediction unit 220.

[0028] Note that the inter-prediction performed by the inter-prediction unit 120 and the inter-prediction unit 220 are the same. This is the operation, and the decoded image and the image stored in frame memory 160 and frame memory 240 are The center prediction parameters will also be the same.

[0029] Next, we will explain the interpretation prediction parameters. The interpretation prediction parameters include M Page flag, merge index, valid flag for predicted LX, motion vector for predicted LX, The reference picture index, merge correction flag, and differential motion vector of the measured LX. It includes LX, which is L0 and L1. The merge flag is the interprediction mode. This flag indicates whether to use the gyromode or differential motion vector mode. If the `Jiflag` is 1, use merge mode; if the merge flag is 0, use differential motion vector. Use the merge mode. The merge index is the position of the selected merge candidate in the merge candidate list. This is an index that indicates the status. The Prediction LX enabled flag indicates whether Prediction LX is enabled or disabled. It is a dual prediction if both L0 and L1 predictions are valid, and a dual prediction if L0 prediction is valid and L1 prediction is valid. If it is invalid, L0 prediction is used; if L1 prediction is valid but L0 prediction is invalid, L1 prediction is used. The merge correction flag indicates whether or not to correct the movement information of the merge candidate. If the correction flag is 1, the merge candidates will be corrected; if the merge correction flag is 0, the merge candidates It does not correct. Here, the code sequence generation unit 140 encodes the valid flag of the predicted LX into the encoding stream. It is not encoded as a code sequence. Also, the code sequence decoding unit 210 encodes the valid flag of the predicted LX. Do not decode the code stream as a code sequence. The reference picture index is the frame memo. This is an index for identifying the decoded image within L160. It also includes the effective color of L0 prediction. G, L1 prediction valid flag, L0 prediction motion vector, L1 prediction motion vector, L0 prediction A combination of the reference picture index for measurement and the reference picture index for L1 prediction. Use "se" as information about movement.

[0030] Note that if the block is an intra-encoded mode block or a block outside the image area, L0 Both the "Enable Prediction" flag and the "Enable L1 Prediction" flag will be disabled.

[0031] From here on, the picture types are L0 prediction unidirectional prediction, L1 prediction unidirectional prediction and bidirectional prediction. All measurements are described as available B-pictures, but the picture type is unidirectional prediction only. A P-picture may also be available. In the case of a P-picture, the interpretation parameter is L Only 0 predictions are considered, and L1 predictions are treated as nonexistent. Generally, B-pictograms When improving coding efficiency, the reference picture for L0 prediction is more past than the picture being predicted. The picture becomes the reference picture for L1 prediction, and the reference picture is a picture that is in the future compared to the picture being predicted. Yes. Because the reference picture for L0 prediction and the reference picture for L1 prediction are the pictures to be predicted. This is because interpolation prediction improves coding efficiency when the direction is in the opposite direction from the perspective of the source. L0 prediction Are the reference picture and the L1 prediction reference picture in opposite directions from the prediction target picture? This involves comparing the POC (Picture Order Count) of the reference picture. This can be used to make a determination. From here on, the reference picture for L0 prediction and the reference picture for L1 prediction will be the targets of the prediction. We explain this by assuming that the predicted picture containing the block is in the reverse direction in time.

[0032] Next, we will explain the details of the interpretation unit 120. Unless otherwise specified, image coding The structure of the interpretation unit 120 of the image processing device 100 and the interpretation unit 220 of the image decoding device 200 The act and the result are identical.

[0033] Figure 3 is a diagram illustrating the configuration of the interpretation unit 120. The interpretation unit 120 is Merge mode determination unit 121, merge candidate list generation unit 122, merge candidate selection unit 123, merge candidate correction determination unit 124, merge candidate correction unit 125, difference motion vector mode implementation unit 1 It includes 26 and a predicted value derivation unit 127.

[0034] The interpretation unit 120 has merge mode and differential motion vector as interpretation modes. The mode is switched for each block. The difference is performed by the differential motion vector mode implementation unit 126. The minute-motion vector mode will be implemented similarly to HEVC, and from here on, the focus will be mainly on the merge mode. Let me explain about Do.

[0035] The merge mode determination unit 121 determines the merge mode as the interpretation mode for each block. Determine whether to use it or not. If the merge flag is 1, use merge mode and merge function If the lag is 0, the differential motion vector mode is used.

[0036] In the interpretation unit 120, the determination of whether or not to set the merge flag to 1 is made based on HEVC The software uses methods such as RDO (Rate Distortion Optimization), which are employed in the reference software. In the prediction unit 220, the code sequence decoding unit 210 uses the syntax to determine the encoded stream from Retrieve the decrypted merge flag. Details of the syntax will be described later.

[0037] If the merge flag is 0, the differential motion vector mode implementation unit 126 implements the differential motion vector The differential motion vector mode is implemented, and the prediction parameters of the differential motion vector mode are used in the prediction value derivation section. It will be supplied to 127.

[0038] If the merge flag is 1, the merge candidate list generation unit 122 and the merge candidate selection unit 12 3. The merge mode is performed by the merge candidate correction determination unit 124 and the merge candidate correction unit 125. The merge mode interpretation parameters are then supplied to the prediction value derivation unit 127.

[0039] The process when the merge flag is set to 1 will be explained in detail below.

[0040] Figure 4 is a flowchart illustrating the operation of merge mode. From here on, Figures 3 and 4 will be used. Next, I will explain merge mode in detail.

[0041] First, the merge candidate list generation unit 122 generates a list of merge candidates for blocks adjacent to the block to be processed. A merge candidate list is generated from the information and the block movement information of the decoded image (S100), The resulting list of merge candidates is supplied to the merge candidate selection unit 123. The blocks to be processed thereafter... The terms "prediction target block" and "prediction target block" are used interchangeably.

[0042] Here, we will explain how to generate the merge candidate list. Figure 5 shows the merge candidate list generation unit. This is a diagram illustrating the configuration of 122. The merge candidate list generation unit 122 generates spatial merge candidates It includes a component unit 201, a time merge candidate generation unit 202, and a merge candidate supplementation unit 203.

[0043] Figure 6 illustrates the blocks adjacent to the block to be processed. Here, the processing pair Blocks adjacent to the elephant block are Block A, Block B, Block C, Block D, and Let's call them Lock E, Block F, and Block G, but there are multiple blocks adjacent to the block to be processed. This is not limited to this if locks are used.

[0044] Figure 7 illustrates blocks in the decoded image located at the same position as the block being processed and in its vicinity. This is a diagram showing the process. Here, the decoded image is located at the same position as the block to be processed and in its vicinity. The blocks are designated as block CO1, block CO2, and block CO3, but the blocks to be processed are This is not limited to using multiple blocks on the decoded image located at the same position as and around the block. No. From now on, blocks CO1, CO2, and CO3 will be placed in the same position block, and blocks in the same position block The decoded images containing these images are called same-location pictures.

[0045] The generation of the merge candidate list will be explained in detail below using Figures 5, 6, and 7.

[0046] First, the spatial merge candidate generation unit 201 generates blocks A, B, C, and B Blocks D, E, F, and G are checked sequentially to determine the valid flag for L0 prediction. If either or both of the L1 prediction valid flags are enabled, the movement information of that block will be displayed. The merge candidates are sequentially added to the merge candidate list. The spatial merge candidate generation unit 201 generates the candidates. The merge candidates that result in this are called spatial merge candidates.

[0047] Next, the time merge candidate generation unit 202 generates block C01, block CO2, and block C O3 is checked sequentially, and first either the valid flag for L0 prediction or the valid flag for L1 prediction is checked. Alternatively, the movement information of blocks where both are valid is processed, such as scaling, to create merge candidates. The merge candidates are added sequentially to the merge candidate list. The merge candidates generated by the time merge candidate generation unit 202 These are called time merge candidates.

[0048] Here, we will discuss scaling of time merge candidates. The merging is similar to HEVC. The motion vector of the time merge candidate is the motion of the same position block. The vector is used to determine the distance between the reference picture that the co-located picture and the co-located block refer to. Based on this, the distance between the picture containing the block to be predicted and the picture referenced by the time merge candidate is calculated. It is derived by scaling by distance.

[0049] Note that the picture referenced by the time merge candidate is the same for both L0 and L1 predictions. It is a reference picture with a character index of 0. Also, as a co-located block, L0 Whether the same-position blocks of the prediction or the same-position blocks of the L1 prediction are used depends on the position. The derivation flag is encoded (decoded) and determined. As described above, the time merge candidates are identical. The motion vector of either the L0 prediction or L1 prediction of the position block is L0 prediction and L1 We scale the prediction to derive new L0 and L1 prediction motion vectors, and then time-marked the time. These are the movement vectors for the L0 and L1 predictions of the candidate.

[0050] Next, if the merge candidate list contains multiple pieces of the same motion information, then select one piece of motion information. Keep all other movement information and delete it.

[0051] Next, the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates. In this case, the merge candidate replenishment unit 203 determines the number of merge candidates included in the merge candidate list to be the most Add supplementary merge candidates to the merge candidate list until the number of major merge candidates is reached. The number of merge candidates included in the supplemental list is set to the maximum number of merge candidates. Here, supplemental merge A candidate is one in which the motion vectors of both the L0 prediction and the L1 prediction are both (0,0). This is motion information where both reference picture indices are 0.

[0052] Here, the maximum number of merge candidates is set to 6, but any number greater than or equal to 1 is acceptable.

[0053] Next, the merge candidate selection unit 123 selects one merge candidate from the merge candidate list. (S101) Selected merge candidates (referred to as "selected merge candidates") and merge index The merge candidate correction determination unit 124 is supplied with the selected merge candidate and the movement information of the block to be processed. The interpretation unit 120 of the image encoding device 100 uses HEVC reference software. The merge candidate list is included using methods such as RDO (Rate Distortion Optimization) used in A. Select one merge candidate from the available merge candidates and determine the merge index. Image decoding In the interpretation unit 220 of the device 200, the code sequence decoding unit 210 decodes from the encoded stream. The merge index is retrieved, and a list of merge candidates is created based on the merge index. Select one merge candidate from the included merge candidates and choose it as the merge candidate.

[0054] Next, the merge candidate correction determination unit 124 determines if the width of the block to be processed is greater than or equal to a predetermined width, and The height of the block to be analyzed is greater than or equal to a predetermined height, and both the L0 prediction and L1 prediction of the selected merge candidate are met. Or check if at least one of them is valid (S102). The width of the block to be processed is The width must be greater than or equal to a certain value, and the height of the block to be processed must be greater than or equal to a predetermined height, and the L0 prediction of the selected merge candidate must be correct. The condition that both or at least one of the S1 predictions are valid must be met. 02 NO), without correcting the selected merge candidates as movement information of the block to be processed, Proceed to step S111. Here, the merge candidate list must include the smallest L0 and L1 predictions. Since merge candidates include at least one that is valid, the L0 prediction and L1 prediction of the selected merge candidates It is obvious that both or at least one of the measurements is valid. Therefore, the "selection" in S102 The phrase "Is both or at least one of the L0 and L1 predictions for the alternative merge candidate valid?" is omitted. Then, S102 is performed when the width of the block to be processed is greater than or equal to a predetermined width and the height of the block to be processed is predetermined It would also be a good idea to check if the height is above the required level.

[0055] The width of the block to be processed is greater than or equal to a predetermined width, and the height of the block to be processed is greater than or equal to a predetermined height, If both or at least one of the L0 and L1 predictions of the selected merge candidate is valid (S If YES is 102, the merge candidate correction determination unit 124 sets the merge correction flag (S10 3) The merge correction flag is supplied to the merge candidate correction unit 125. In the interpretation unit 120, the prediction error when interpreting with merge candidates is a predetermined prediction. If the error is greater than the specified margin, the merge correction flag is set to 1, and the inter-prediction is performed using the selected merge candidates. If the prediction error in the case is not greater than or equal to a predetermined prediction error, the merge correction flag is set to 0. In the interpretation unit 220 of the image decoding device 200, the code sequence decoding unit 210 is based on syntax Next, obtain the merge correction flag decoded from the encoded stream.

[0056] Next, the merge candidate correction unit 125 checks whether the merge correction flag is 1 (S10 4) If the merge correction flag is not 1 (NO in S104), the selected merge candidate will be processed. Without correcting the block's movement information, the process proceeds to step S111.

[0057] If the merge correction flag is 1 (YES in S104), there is an L0 prediction for the selected merge candidate. Check if it is effective (S105). If the L0 prediction of the selected merge candidate is not effective (S10 If NO. 5, proceed to step S108. If the L0 prediction of the selected merge candidate is valid (S1 YES to 05), determine the difference motion vector of the L0 prediction (S106). As described above, If the merge correction flag is 1, the movement information of the selected merge candidate will be corrected, and the merge correction flag If the value is 0, the movement information of the selected merge candidates will not be corrected.

[0058] In the interpretation unit 120 of the image coding device 100, the differential motion vector of the L0 prediction is The motion vector is determined by vector search. Here, the search range for the motion vector is horizontal and vertical. Both are set to ±16, but any multiple of 2, such as ±64, is acceptable. The interface of the image decoding device 200 - In the prediction unit 220, the code sequence decoding unit 210 determines whether the encoded stream is based on the syntax. Then, obtain the difference motion vector of the decoded L0 prediction.

[0059] Subsequently, the merge candidate correction unit 125 calculates the corrected motion vector of the L0 prediction, and L0 The corrected motion vector of the prediction is used as the motion vector of the L0 prediction of the motion information of the block to be processed. (S107).

[0060] Here, we have the corrected motion vector (mvL0) of the L0 prediction and the motion of the L0 prediction of the selected merge candidate. This explains the relationship between the vector (mmvL0) and the difference motion vector (mvdL0) of the L0 prediction. To clarify, the corrected motion vector (mvL0) of the L0 prediction is the motion vector of the L0 prediction of the selected merge candidate. This is the sum of the motion vector (mmvL0) and the difference motion vector (mvdL0) of the L0 prediction. Therefore, the following equation is obtained. Note that [0] is the horizontal component of the motion vector, and [1] is the horizontal component of the motion vector. This shows the vertical component of the torque.

[0061] mvL0[0] = mmvL0[0] + mvdL0[0] mvL0[1] = mmvL0[1] + mvdL0[1] Next, we check whether the L1 prediction of the selected merge candidates is valid (S108). If the L1 prediction for the page candidate is not valid (NO in S108), proceed to step S111. If the L1 prediction of the merge candidate is valid (YES in S108), the difference movement vector of the L1 prediction Determine the torque (S109).

[0062] In the interpretation unit 120 of the image coding device 100, the differential motion vector of the L1 prediction is The motion vector is determined by vector search. Here, the search range for the motion vector is horizontal and vertical. Both are set to ±16, but powers of 2 such as ±64 are used. Interpretation of the image decoding device 200 In the measurement unit 220, the code sequence decoding unit 210 decodes the encoded stream based on the syntax. Obtain the difference motion vector of the L1 prediction that was generated.

[0063] Subsequently, the merge candidate correction unit 125 calculates the corrected motion vector of the L1 prediction, and L1 The corrected motion vector of the prediction will be used as the motion vector of the L1 prediction of the motion information of the block to be processed. (S110).

[0064] Here, we have the corrected motion vector (mvL1) of the L1 prediction and the motion of the L1 prediction of the selected merge candidate. This explains the relationship between the vector (mmvL1) and the difference motion vector (mvdL1) of the L1 prediction. To clarify, the corrected motion vector (mvL1) of the L1 prediction is the motion vector of the L1 prediction of the selected merge candidate. This is the sum of the motion vector (mmvL1) and the difference motion vector (mvdL1) of the L1 prediction. Therefore, the following equation is obtained. Note that [0] is the horizontal component of the motion vector, and [1] is the horizontal component of the motion vector. This shows the vertical component of the torque.

[0065] mvL1[0] = mmvL1[0] + mvdL1[0] mvL1[1] = mmvL1[1] + mvdL1[1] Next, the predicted value derivation unit 127 predicts L0 based on the movement information of the block to be processed. Interpretation is performed using either L1 prediction or biprediction, and the predicted value is derived (S111). As described above, if the merge correction flag is set to 1, the motion vectors of the selected merge candidates will be corrected. Furthermore, if the merge correction flag is 0, the motion vectors of the selected merge candidates will not be corrected.

[0066] The coding of interpretation parameters will be explained in detail. Figure 8 shows the merge mode. This figure shows a portion of the block syntax. Table 1 shows the interpretation parameters and This shows the relationship between the numbers. In Figure 8, cbWidth is the width of the block to be processed, and cbH eight is the height of the block to be processed. The predetermined width and predetermined height are both set to 8. By setting a predetermined height, the merge candidates are not corrected in small block units. Processing load can be reduced. Here, cu_skip_flag is used to indicate that a block is skipped. It is 1 in skip mode and 0 otherwise. The tax is the same as the syntax in merge mode. merge_idx is the merge candidate. This is a merge index that selects merge candidates from a list.

[0067] Encode (decode) merge_idx before merge_mod_flag After determining the page index, the encoding (decoding) of merge_mod_flag is performed. Determine that merge_idx is shared with the merge index of the merge mode, This improves coding efficiency while suppressing the increasing complexity of syntax and the growing amount of context.

[0068] [Table 1]

[0069] Figure 9 shows the syntax for the difference motion vector. (mvd_codin in Figure 9) g(N) uses the same syntax as the syntax used in differential motion vector mode. Here, N is either 0 or 1. N=0 indicates an L0 prediction, and N=1 indicates an L1 prediction.

[0070] The syntax for a differential motion vector includes whether the components of the differential motion vector are greater than 0 or not. abs_mvd_greater0_flag[d] is a flag indicating the difference motion vector abs_mvd_greater1 is a flag that indicates whether the component of Tor is greater than 1 or not. _flag[d] indicates the sign (±) of the components of the difference motion vector. mvd_sign_fl ag[d] and abs_m represent the absolute value of the vector obtained by subtracting 2 from the components of the difference motion vector. The variable vd_minus2[d] is included, where d is either 0 or 1. d=0 means horizontal. The directional component is shown, and d=1 indicates the vertical component.

[0071] HEVC has two interpretation modes: merge mode and differential motion vector mode. In merge mode, motion information can be restored with just one merge flag, making it very efficient to encode. It was a high-rate mode. However, the movement information in merge mode depends on the processed blocks. Therefore, the cases in which prediction efficiency is high are limited, and there is a need to further improve utilization efficiency. Ta.

[0072] On the other hand, the differential motion vector mode provides separate syntax for L0 and L1 predictions. Furthermore, the prediction type (L0 prediction, L1 prediction, or dual prediction), and each of the L0 and L1 predictions Regarding the predicted motion vector flag, differential motion vector, and reference picture index It was necessary. Therefore, the differential motion vector mode is more efficient at encoding than the merge mode. However, this is a movement that cannot be derived in merge mode, such as spatially adjacent blocks or temporally adjacent blocks. It has high predictive efficiency and stability for sudden movements that have little correlation with the movements of adjacent blocks. I was in the mood to get excited.

[0073] In this embodiment, the merge mode prediction type and the reference picture index are fixed. By enabling motion vector correction in merge mode, encoding efficiency is improved by the difference motion vector. It can be improved over Tor Mode and more efficient than Merge Mode.

[0074] Furthermore, by making the difference motion vector the difference between the motion vector of the selected merge candidate, The magnitude of the motion vector can be reduced, thereby suppressing the encoding efficiency. .

[0075] Furthermore, the syntax for the difference motion vector in merge mode is the difference in difference motion vector mode. By making the syntax identical to that of minute motion vectors, difference motion vectors can be used in merge mode. Adding it can minimize configuration changes.

[0076] Furthermore, a predetermined width and height are defined, and the predicted block width is greater than or equal to the predetermined width, and the predicted block height is greater than or equal to the predetermined width. If the value is above a predetermined level, and both or at least one of the L0 predictions or L1 predictions of the selected merge candidate are met, If the condition that the method is effective is not met, the motion vector of the merge mode will be corrected. By eliminating the logic, the amount of processing required to correct the motion vector in merge mode can be reduced. Furthermore, if it is not necessary to suppress the amount of processing required to correct the motion vector, the predetermined width and predetermined height Therefore, there is no need to restrict the correction of the motion vector.

[0077] Modifications of this embodiment will be described below. Unless otherwise specified, the modifications are not combined with each other. They can be combined. [Example 1] In this embodiment, the syntax for a block in merge mode is a differential motion vector. We used this method. In this modified example, the difference motion vector is encoded as the difference unit motion vector ( It is defined to decode. The difference unit motion vector is defined as the picture interval at the minimum interval. This is a motion vector in a given case. In HEVC, the minimum picture interval is the encoded stream. It is encoded as a sequence of codes within the program.

[0078] The difference unit motion vector depends on the interval between the picture to be encoded and the reference picture in merge mode. The vector is scaled and used as a differential motion vector. POC of the picture to be encoded ( Picture Order Count (POC) is set to POC (Cur), and L0 is set to merge mode. The reference picture for measurement is POC(L0), and the reference picture for merge mode L1 prediction is If POC is denoted as POC(L1), the motion vector is calculated as shown in the following equation: umvdL 0 is the difference unit motion vector of the L0 prediction, and umvdL1 is the difference unit motion vector of the L1 prediction. It indicates a letter.

[0079] mvL0[0] = mmvL0[0] + umvdL0[0]*(POC(Cur )-POC(L0)) mvL0[1] = mmvL0[1] + umvdL0[1]*(POC(Cur )-POC(L0)) mvL1[0] = mmvL1[0] + umvdL1[0]*(POC(Cur )-POC(L1)) mvL1[1] = mmvL1[1] + umvdL1[1]*(POC(Cur )-POC(L1)) As described above, in this modification, the differential unit motion vector is used as the differential motion vector. By doing so, the coding efficiency is improved by reducing the code size of the differential motion vector. This is possible. Note that this is possible when the difference motion vector is large and the distance between the picture to be predicted and the reference picture is large. When the size increases, encoding efficiency can be improved in particular. Also, the picture to be processed Even when the distance between the reference picture and the speed of the object moving within the screen are proportional, the prediction efficiency is still high. This can improve coding efficiency.

[0080] Furthermore, unlike time merge candidates, there is no need to derive them by scaling with the distance between pictures. Furthermore, in a decoding device, scaling can be done using only a multiplier, eliminating the need for a divider, thus reducing circuit size and processing. It can reduce the amount of information needed. [Differentiation 2] In this embodiment, a component of the difference motion vector can encode (or decode) 0. For example, we made it possible to change only the L0 prediction. In this modified example, the differential motion vector The element 0 cannot be encoded (or decoded) as a component of the expression.

[0081] Figure 10 shows the syntax of the difference motion vector for Modification 2. The syntax of Toll includes a flag that indicates whether the components of the difference motion vector are greater than 1. A certain abs_mvd_greater1_flag[d] has two components in the difference motion vector. abs_mvd_greater2_flag[d ], abs_mvd_m represents the absolute value of the vector obtained by subtracting 3 from the components of the difference motion vector. inus3[d], mvd_sign_fl indicates the sign (±) of the components of the difference motion vector. ag[d] is included.

[0082] As described above, zero cannot be encoded (or decoded) as a component of the difference motion vector. This improves the encoding efficiency when the components of the difference motion vector are 1 or greater. It is possible. [Difference 3] In this embodiment, the components of the difference motion vector are integers, while in Modification 2, integers excluding 0 are used. In this modified example, the components of the difference motion vector, excluding the ± sign, are limited to powers of 2.

[0083] Instead of the syntax abs_mvd_minus2[d] in this embodiment, use a Use bs_mvd_pow_plus1[d]. The difference motion vector mvd[d] is m From vd_sign_flag[d] and abs_mvd_pow_plus1[d], the following formula It is calculated as follows.

[0084] mvd[d]=mvd_sign_flag[d]*2^(abs_mvd_po w_plus1[d]+1) Also, instead of the syntax abs_mvd_minus3[d] in variation 2 Use abs_mvd_pow_plus2[d]. The difference motion vector mvd[d] is From mvd_sign_flag[d] and abs_mvd_pow_plus2[d] downwards It is calculated as shown in the formula.

[0085] mvd[d]=mvd_sign_flag[d]*2^(abs_mvd_po w_plus2[d]+2) By limiting the components of the differential motion vector to powers of 2, the processing load of the encoding device can be significantly reduced. While reducing the number of motion vectors, it is possible to improve prediction efficiency in the case of large motion vectors. [Differentiation Example 4] In this embodiment, mvd_coding(N) is assumed to include the differential motion vector. However, in this modified example, mvd_coding(N) is assumed to include the motion vector scaling factor.

[0086] In the syntax of mvd_coding(N) in this modified example, abs_mvd_gre ater0_flag[d], abs_mvd_greater1_flag[d], m vd_sign_flag[d] does not exist, and instead abs_mvr_plus The configuration will include 2[d] and mvr_sign_flag[d].

[0087] The corrected motion vector (mvLN) for LN prediction is the motion vector of the LN prediction of the selected merge candidate. It is the product of the length (mmvLN) and the motion vector multiplier (mvrLN), and is given by the following formula: It is calculated.

[0088] mvLN[d] = mmvLN[d] * mvrLN[d] By limiting the components of the differential motion vector to powers of 2, the processing load of the encoding device can be significantly reduced. While reducing the number of motion vectors, it is possible to improve prediction efficiency in the case of large motion vectors.

[0089] Note that this modified example cannot be combined with Modified Example 1, Modified Example 2, Modified Example 3, and Modified Example 6. It is not possible. [Difference 5] In the syntax of Figure 8 of this embodiment, cu_skip_flag is 1 (skip If it is in mod mode, there is a possibility that merge_mod_flag exists. However, in skip mode, merge_mod_flag does not exist. You may do so.

[0090] Thus, by omitting merge_mod_flag, the sign of skip mode is changed. This improves processing efficiency and simplifies the determination of skip mode. [Modification 6] In this embodiment, we check whether the LN prediction (N=0 or 1) of the selected merge candidate is valid. Furthermore, if LN prediction for selected merge candidates is not enabled, the differential motion vector will not be enabled. However, without checking whether the LN prediction of the selected merge candidate is valid, the LN of the selected merge candidate The differential motion vector may be enabled regardless of whether the prediction is valid or not. In this case, select If LN prediction for the selected merge candidate is invalid, the movement vector of the LN prediction for the selected merge candidate The value is (0,0), and the reference picture index for the LN prediction of the selected merge candidate is set to 0.

[0091] Thus, in this modified example, regardless of whether LN prediction of the selected merge candidate is effective or not By enabling differential motion vectors, opportunities to utilize bidirectional prediction are increased, and encoding Efficiency can be improved. [Difference 7] In this embodiment, whether the L0 prediction and L1 prediction of the selected merge candidate are enabled or not is determined individually. We controlled whether to encode (or decode) the differential motion vector based on the determination, but selected merge If both the candidate's L0 and L1 predictions are valid, encode the difference motion vector (or (Decryption) If neither the L0 prediction nor the L1 prediction of the selected merge candidate is valid, then the differential motion... It is also possible to avoid encoding (or decoding) the metric. In this modified example, step The S102 is as follows:

[0092] The merge candidate correction determination unit 124 determines if the width of the block to be processed is greater than or equal to a predetermined width, and the width of the block to be processed is greater than or equal to a predetermined width. The lock height is greater than or equal to a predetermined height, and both L0 and L1 predictions for the selected merge candidate are valid. Check if it exists (S102).

[0093] Furthermore, steps S105 and S108 are unnecessary in this modified example.

[0094] Figure 11 shows a portion of the block syntax, which is the merge mode of Modification 7. The syntax for steps S102, S105, and S108 is different.

[0095] Thus, in this modified example, both L0 and L1 predictions of the selection merge candidate are valid. In some cases, enabling differential motion vectors allows for the selection and merging of frequently used bidirectional prediction candidates. By correcting the complementary motion vector, prediction efficiency can be improved efficiently. [Differentiation 8] In this embodiment, the syntax for the block in merge mode is the difference in L0 predictions. This transformation uses two difference motion vectors: the motion vector and the difference motion vector of the L1 prediction. In the example, only one differential motion vector is encoded (or decoded) to correct the L0 prediction motion. By sharing one difference motion vector as the correction motion vector for the vector and L1 prediction, the following equation is obtained: As shown, the motion vector mmvLN(N=0,1) of the selected merge candidate and the difference motion vector m The motion vector mvLN(N=0,1) of the correction merge candidate is calculated from vd.

[0096] If L0 prediction for selected merge candidates is valid, the motion vector of the L0 prediction can be obtained from the following formula. Calculate.

[0097] mvL0[0] = mmvL0[0] + mvd[0] mvL0[1] = mmvL0[1] + mvd[1] If L1 prediction is valid for the selected merge candidate, the movement vector of the L1 prediction can be obtained from the following formula. Calculate. Add the difference motion vector in the opposite direction to the L0 prediction. Select the L1 prediction of the merge candidate. You can also subtract the difference motion vector from the measured motion vector.

[0098] mvL1[0] = mmvL1[0] + mvd[0]*-1 mvL1[1] = mmvL1[1] + mvd[1]*-1 Figure 12 shows a portion of the block syntax, which is the merge mode of Modification Example 8. The function to check the effectiveness of L0 prediction and L1 prediction has been removed, and mvd_ This embodiment differs in that it does not have coding(1). mvd_coding(0) is This corresponds to a single difference motion vector.

[0099] Thus, in this modified example, only one difference motion vector is used for the L0 and L1 predictions. By defining this, in the case of biprediction, the number of difference motion vectors is halved, and L0 prediction and L By sharing it with the first prediction, the coding efficiency is improved while suppressing the decrease in prediction efficiency. It is possible.

[0100] Also, the reference picture referenced by the L0 prediction of the selected merge candidate and the reference referenced by the L1 prediction When a picture is in the opposite direction to the predicted picture (not in the same direction), differential motion occurs. By adding the vector in the opposite direction, encoding efficiency is improved for motion in a fixed direction. It is possible.

[0101] The effects of this modified example will be explained in detail. Figure 16 is a diagram illustrating the effects of modified example 8. Figure 16 shows a sphere moving horizontally within a moving rectangular area (area enclosed by dashed lines). This is an image showing the state of the area (colored with lines). In this case, the movement of the sphere relative to the screen. This movement is the sum of the movement of the rectangular area and the movement of the sphere moving horizontally. Picture B is The target picture, picture A is the reference picture for L0 prediction, and picture C is the reference picture for L1 prediction. Let's assume it's a kucha. Picture A and Picture C are in opposite directions relative to the predicted picture. This is a reference picture located at [location].

[0102] If a sphere moves in a constant direction at a constant speed, then picture A, picture B, and picture C will be... If they are at equal intervals, the amount of movement of spheres that cannot be obtained from adjacent blocks is added to the L0 prediction. By subtracting this from the L1 prediction, the movement of the sphere can be accurately reproduced.

[0103] When a sphere moves in a constant direction at a non-constant speed, picture A, picture B, pic Although the C segments are not equally spaced, the amount of movement corresponding to the rectangular area of ​​the sphere is equal if they are equally spaced. The amount of movement of spheres that cannot be obtained from adjacent blocks is added to the L0 prediction and subtracted from the L1 prediction. This allows for the accurate reproduction of the sphere's movement.

[0104] Furthermore, if a sphere moves in a constant direction at a constant speed over a certain period of time, Picture A, Picture B Although the pictures C are not equally spaced, the amount of movement corresponding to the rectangular area of ​​the sphere is equally spaced. This can happen. Figure 17 illustrates the effect of the case where the picture spacing is not equal in Modification Example 8. This is a diagram. This example will be explained in detail using Figure 17. Pictures F0 and F1 in Figure 17 , ..., F8 each indicate pictures at fixed intervals. Picture F0 to Picture F4 Assume the sphere is stationary and moves in a constant direction at a constant speed after picture F5. Pictures F0 and F6 are the reference pictures, and picture F5 is the picture to be predicted. In this case, picture F0, picture F5, and picture F6 are not equally spaced, but the sphere The movement amounts corresponding to the rectangular area will be equally spaced. This applies when picture F5 is the picture to be predicted. Generally, the closest picture, F4, is selected as the reference picture, but with picture F4... Instead of selecting Picture F0 as the reference picture, Picture F0 is Picture F4. This is the case when the picture is of high quality with minimal distortion. The reference picture is usually FIFO (Fi It is managed using the RST-In (First-Out) method, but produces high-quality pictures with minimal distortion. There is a mechanism called a long-term reference picture that uses a reference picture. Thus, this modified example is This applies when either or both of the L0 or L1 forecasts are long-range reference pictures. This improves prediction efficiency and coding efficiency. Furthermore, this modified example is L0 This applies when either the prediction or the L1 prediction, or both, are intrapictures. This allows for improvements in both prediction efficiency and coding efficiency.

[0105] Furthermore, the difference motion vector, like the time merge candidate, scales based on the distance between pictures. By not using differential motion, circuit size and power consumption can be reduced. For example, if differential motion When scaling a cult, if a time merge candidate is selected as the selected merge candidate, Both scaling of the time merge candidates and scaling of the differential motion vectors are required. The scaling of time merge candidates and the scaling of differential motion vectors are the scaling criteria. Because the resulting motion vectors are different, both scaling processes cannot be performed together and must be performed separately. It needs to be done.

[0106] Furthermore, if the merge candidate list includes time merge candidates, as in this embodiment, Time merge candidates are scaled, and the difference motion is greater than the motion vector of the time merge candidate. If the motion vector is small, encoding efficiency can be improved without scaling the difference motion vector. It can be made to do this. Also, if the difference motion vector is large, the difference motion vector motion By selecting option D, the decrease in encoding efficiency can be suppressed. [Modification 9] In this embodiment, the maximum number of merge candidates is the same when the merge correction flag is 0 or 1. In this modified example, the maximum number of merge candidates when the merge correction flag is 1 is calculated as follows: The number of merge candidates should be less than the maximum number when the positive flag is 0. For example, the merge correction flag. The maximum number of merge candidates when is 1 is set to 2. Here, when the merge correction flag is 1 The maximum number of merge candidates is set to the maximum number of corrected merge candidates. Also, the merge index is set to the maximum If the number of correction merge candidates is smaller, the merge correction flag is encoded (decoded), and the merge is If the index is greater than or equal to the maximum number of corrected merge candidates, encode the merge correction flag (recycle (Number) Avoid doing so. Here, the maximum number of merge candidates when the merge correction flag is 0 and The maximum number of correction merge candidates may be a predetermined value, or it may be in the encoded stream. The data can also be obtained by encoding (decoding) it into SPS or PPS.

[0107] Thus, in this modified example, the maximum number of merge candidates when the merge correction flag is 1 is... By making the number of merge candidates less than the maximum number when the merge correction flag is 0, the selection probability is lowered. Only for merge candidates with a high success rate, a decision is made as to whether or not to correct the merge candidate. This reduces the processing required for placement while suppressing a decrease in encoding efficiency. Furthermore, merge indexing is also possible. If the number of merge candidates exceeds the maximum number of corrected merge candidates, the merge correction flag is encoded (decoded). This eliminates the need for encoding, thus improving encoding efficiency. [Second Embodiment] The configuration of the image encoding device 100 and image decoding device 200 in the second embodiment is the same as the first embodiment. The image encoding device 100 and image decoding device 200 of the embodiment are identical. This embodiment is the The operation and syntax of the merge mode differ from that of Embodiment 1. Hereafter, this embodiment and The differences from the first embodiment will now be explained.

[0108] Figure 13 is a flowchart illustrating the operation of the merge mode in the second embodiment. Figure 14 shows a portion of the block syntax, which is the merge mode of the second embodiment. This is a diagram. Figure 15 is a diagram showing the syntax of the difference motion vector in the second embodiment. ru.

[0109] Hereafter, the differences from the first embodiment will be explained using Figures 13, 14, and 15. Figure 13 shows the same steps as in Figure 4, from step S205 to step S207, and from step S209 to step S207. Step S211 is different.

[0110] If the merge correction flag is 1 (YES in S104), there is no L0 prediction for the selected merge candidate. Check if it is effective (S205). If the L0 prediction of the selected merge candidate is not invalid (S20 If NO. 5, proceed to step S208. If the L0 prediction for the selected merge candidate is invalid (S2 YES for 05), determine the corrected motion vector for L0 prediction (S206).

[0111] In the interpretation unit 120 of the image coding device 100, the corrected motion vector of the L0 prediction is The motion vector is determined by vector search. Here, the search range for the motion vector is horizontal and vertical. Both are set to ±1. The interpretation unit 220 of the image decoding device 200 corrects the L0 prediction. The vector is obtained from the encoded stream.

[0112] Next, the reference picture index for L0 prediction is determined (S207). The reference picture index for L0 prediction is set to 0.

[0113] Next, check if the slice type is B and if the L1 prediction for the selected merge candidate is invalid. (S208). If the slice type is not B or the L1 prediction for the selected merge candidate is invalid. If not found (NO. S208), proceed to S111. Slice type is B and select merge candidate If the supplementary L1 prediction is invalid (YES in S208), determine the corrected motion vector for the L1 prediction. (S209).

[0114] In the interpretation unit 120 of the image coding device 100, the corrected motion vector of the L1 prediction is The motion vector is determined by vector search. Here, the search range for the motion vector is horizontal and vertical. Both are set to ±1. The interpretation unit 220 of the image decoding device 200 performs correction of L1 prediction. The vector is obtained from the encoded stream.

[0115] Next, the reference picture index for L1 prediction is determined (S110). The reference picture index for L1 prediction is set to 0.

[0116] As described above, in this embodiment, the slice type is a slice that allows biprediction. For type B (i.e., slice type B), merge candidates for L0 prediction or L1 prediction are bipredicted. Convert to merge candidates for measurement. By making them merge candidates for dual prediction, the filtering effect is This allows us to expect improved prediction efficiency. Also, by using the most recent decoded image as the reference picture... This minimizes the search range for motion vectors. [Differentiation] In this embodiment, the reference picture index in steps S207 and S210 This was set to 0. In this modified example, if the L0 prediction of the selected merge candidate is invalid, the L0 prediction is used. The reference picture index is used as the reference picture index for L1 prediction, and the selection merge candidates are If L1 prediction is disabled, the reference picture index for L1 prediction will be used as the reference picture for L0 prediction. Let's call it a 'cha index'.

[0117] Thus, by slightly shifting and filtering the predicted value based on the motion vector of the L0 prediction or L1 prediction of the selection merge candidate, a minute motion can be reproduced, and the prediction efficiency can be improved.

[0118] In all the embodiments described above, the coded bitstream output by the image coding apparatus is specified to have a data format that can be decoded according to the coding method used in the embodiment. The coded bitstream has a data format that can be read by a computer such as an HDD, SSD, flash memory, optical disk, etc. and recorded on a recording medium for providing or may be provided from a server through a wired or wireless network. Accordingly, the image decoding apparatus corresponding to this image coding apparatus can decode the coded bitstream of this specific data format regardless of the providing means. When a wired or wireless network is used to exchange the coded bitstream between the image coding apparatus and the image decoding apparatus, the coded bitstream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the coded bitstream output by the image coding apparatus into coded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a receiving apparatus that receives the coded data from the network, restores it to a coded bitstream, and supplies it to the image decoding apparatus are provided. The transmission apparatus includes a memory that buffers the coded bitstream output by the image coding apparatus, a packet processing unit that packetizes the coded bitstream, and a packet that is transmitted via the network

[0119] [[ID=2 (corrected)]] When a wired or wireless network is used to exchange the coded bitstream between the image coding apparatus and the image decoding apparatus, the coded bitstream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the coded bitstream output by the image coding apparatus into coded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a receiving apparatus that receives the coded data from the network, restores it to a coded bitstream, and supplies it to the image decoding apparatus are provided. The transmission apparatus includes a memory that buffers the coded bitstream output by the image coding apparatus, a packet processing unit that packetizes the coded bitstream, and a packet that is transmitted via the network When a wired or wireless network is used to exchange the coded bitstream between the image coding apparatus and the image decoding apparatus, the coded bitstream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the coded bitstream output by the image coding apparatus into coded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a receiving apparatus that receives the coded data from the network, restores it to a coded bitstream, and supplies it to the image decoding apparatus are provided. The transmission apparatus includes a memory that buffers the coded bitstream output by the image coding apparatus, a packet processing unit that packetizes the coded bitstream, and a packet that is transmitted via the network to the network. The transmission apparatus includes a memory that buffers the coded bitstream output by the image coding apparatus, a packet processing unit that packetizes the coded bitstream, and a packet that is transmitted via the network ​It includes a transmitting unit that transmits coded data. The receiving device transmits data via a network. A receiving unit that receives the encoded data that has been packetized, and a buffer that receives the encoded data It uses memory to process the encoded data, and generates an encoded bitstream by packet processing the encoded data. It includes a packet processing unit that provides packets to the image decoding device.

[0120] In order to exchange encoded bitstreams between the image encoding device and the image decoding device, When a wired or wireless network is used, in addition to the transmitting device and receiving device, Even if a relay device is provided that receives encoded data transmitted by a transmitting device and supplies it to a receiving device, Good. The relay device is a receiving unit that receives packetized encoded data transmitted by the transmitting device. And, a memory for buffering the received encoded data, and the packetized encoded data and It includes a transmitting unit that transmits to the network. Furthermore, the relay device includes a packetized encoded data A receiving packet processing unit that processes the data into packets to generate an encoded bitstream, and a coding A recording medium for storing encoded bitstreams and a transmission medium for packetizing encoded bitstreams. It may include a signal packet processing unit.

[0121] Furthermore, by adding a display unit to the configuration that displays the image decoded by the image decoding device, the display It can also function as a device. Furthermore, an imaging unit can be added to the configuration, and the captured images can be input to an image encoding device. Therefore, it can also be used as an imaging device.

[0122] Figure 18 shows an example of the hardware configuration of the encoding / decoding device of the present invention. The encoding / decoding device This includes the configuration of an image encoding device and an image decoding device according to embodiments of the present invention. The encoding / decoding device 9000 consists of a CPU 9001, a codec IC 9002, and an I / O interface. -Face 9003, Memory 9004, Optical Disc Drive 9005, Network It has an interface 9006 and a video interface 9009, and each part is connected to bus 9010 It is connected by.

[0123] The image encoding unit 9007 and the image decoding unit 9008 are typically connected to the codec IC 9002. It is implemented as follows. The image coding process of the image coding device according to an embodiment of the present invention is an image code This is performed by the numbering unit 9007, and is an image decoding in an image decoding device according to an embodiment of the present invention. The processing is performed by the image encoding unit 9007. The I / O interface 9003 is For example, this is achieved via a USB interface, and external keyboard 9104, mouse 9 Connects to 105, etc. CPU9001 receives input via I / O interface 9003. Based on the user's input, the encoding / decoding device 9 performs the action desired by the user. Control 000. User operation via keyboard 9104, mouse 9105, etc. This includes selecting whether to perform encoding or decoding, setting the encoding quality, and encoding stream. This includes input / output destinations for software, image input / output destinations, etc.

[0124] When the user wishes to play back images recorded on the disk recording medium 9100 The optical disc drive 9005 encodes video from the inserted disc recording medium 9100. The readout stream is read, and the readout encoded stream is coded via bus 9010. The image is sent to the image decoding unit 9008 of IC9002. The image decoding unit 9008 processes the input encoded video. The image decoding process in the image decoding apparatus according to the embodiment of the present invention is applied to the stream. Execute and send the decoded image to an external monitor 9103 via the video interface 9009. Also, the encoding / decoding device 9000 has a network interface 9006 and can be connected to an external distribution server 9106 and a mobile terminal 9107 via the network 9101. If the user desires to play back an image recorded on a distribution server 9106 or a mobile terminal 9107 instead of the image recorded on the disk recording medium 9100, the network interface 9006 obtains an encoded stream from the network 9101 instead of reading an encoded bitstream from the input disk recording medium 9100. Also, if the user desires to play back an image recorded in the memory 9004, image decoding processing in the image decoding device according to the embodiment of the present invention is performed on the encoded stream recorded in the memory 9004.

[0125] If the user desires to encode an image captured by an external camera 9102 and record it in the memory 9004, the video interface 9009 receives an image from the camera 9102 and sends it to the image encoding unit 9007 of the codec IC 9002 via the bus 9010. The image encoding unit 9007 performs image encoding processing in the image encoding device according to the embodiment of the present invention on the image input via the video interface 9009 and creates an encoded bitstream. Then, the encoded bitstream is sent to the memory 9004 via the bus 9010.

[0126] If the user desires to record the encoded stream on the disk recording medium 9100 instead of the memory 9004, the optical disk drive 9005 inserts <000098:The encoded stream is written to the inserted disk recording medium 9100.

[0126] Hardware configurations that have an image encoding device but no image decoding device, or hardware configurations that have an image decoding device It is also possible to realize a hardware configuration that does not include an image encoding device. The hardware configuration is, for example, that the codec IC9002 is the image encoding unit 9007, or This is achieved by replacing each of the image decoding units 9008 with their respective components.

[0127] The above encoding and decoding processes are performed using a transmission device with hardware such as an ASIC. It can, of course, be implemented as a storage device and a receiving device, as well as ROM (Read-On Firmware stored in memory (such as flash memory), CPU, and SoC This can also be achieved through computer software such as (System on a chip). It is possible to compute the firmware program and software program. It can also be recorded and provided on a recording medium that can be read by a data reader, or via a wired or wireless network. Data can be provided from the server through the network, or from terrestrial or satellite digital broadcasting. It can also be provided as a broadcast.

[0128] The present invention has been described above based on embodiments. The embodiments are illustrative and their respective components The fact that various variations are possible in the combination of constituent elements and each processing process, and such variations Those skilled in the art will understand that this also falls within the scope of the present invention. [Explanation of Symbols]

[0129] 100 Image encoding unit, 110 Block size determination unit, 120 Interpretation unit 121 Merge mode determination unit, 122 Merge candidate list generation unit, 123 M merge candidate selection unit, 124 merge candidate correction determination unit, 125 merge candidate correction unit, 1 26 Differential motion vector mode implementation unit, 127 Predicted value derivation unit, 130 Transformation unit, 140 Code sequence generation unit, 150 Local decoding unit, 160 Frame memory, 200 Image decoding device, 201 Spatial merge candidate generation unit, 202 Temporal merge candidate generation unit, 203 Merge candidate supplementation unit, 210 Code sequence decoding unit, 220 Interpretation unit, 2 30 inverse conversion units, 240 frame memory.

Claims

1. A merge candidate list generation unit generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on an encoded image at the same position as the block to be predicted, A merge candidate supplement unit adds a supplement merge candidate to the merge candidate list whose motion vector is (0,0), A merge candidate selection unit that selects a merge candidate from the aforementioned merge candidate list, A merge correction determination unit sets a merge correction flag indicating whether or not to correct the merge candidates, If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. If neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, the merge candidate correction unit derives a corrected merge candidate for the dual prediction by subtracting the correction vector, scaled according to the interval between the prediction target picture and the reference picture, from the motion vector of the second prediction of the selected merge candidate. The system includes a code sequence encoding unit that encodes the merge correction flag and the correction vector into an encoded stream, An image encoding device characterized in that, when the selected merge candidates are corrected, the number of selectable selected merge candidates is smaller compared to when the selected merge candidates are not corrected.

2. A merge candidate list generation step generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on an encoded image at the same position as the block to be predicted, A merge candidate replenishment step, which adds a supplemental merge candidate whose motion vector is (0,0) to the merge candidate list, A merge candidate selection step in which a merge candidate is selected from the aforementioned merge candidate list, A merge correction determination step that sets a merge correction flag indicating whether or not to correct the merge candidates, If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. If neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, then a merge candidate correction step is performed to derive a corrected merge candidate for the dual prediction by subtracting the correction vector, scaled according to the interval between the prediction target picture and the reference picture, from the motion vector of the second prediction of the selected merge candidate. The method includes a code sequence encoding step of encoding the merge correction flag and the correction vector into an encoded stream, An image encoding method characterized in that, when the selected merge candidates are corrected, the number of selectable selected merge candidates is smaller compared to when the selected merge candidates are not corrected.

3. A merge candidate list generation step generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on an encoded image at the same position as the block to be predicted, A merge candidate replenishment step, which adds a supplemental merge candidate whose motion vector is (0,0) to the merge candidate list, A merge candidate selection step in which a merge candidate is selected from the aforementioned merge candidate list, A merge correction determination step that sets a merge correction flag indicating whether or not to correct the merge candidates, If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. If neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, then a merge candidate correction step is performed to derive a corrected merge candidate for the dual prediction by subtracting the correction vector, scaled according to the interval between the prediction target picture and the reference picture, from the motion vector of the second prediction of the selected merge candidate. The computer is made to perform a code sequence encoding step that encodes the merge correction flag and the correction vector into an encoded stream, An image encoding program characterized in that, when the selected merge candidates are corrected, the number of selectable merge candidates is smaller compared to when the selected merge candidates are not corrected.

4. A merge candidate list generation unit generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on the decoded image at the same position as the block to be predicted, A merge candidate supplement unit adds a supplement merge candidate to the merge candidate list whose motion vector is (0,0), A merge candidate selection unit that selects a merge candidate from the aforementioned merge candidate list, A code sequence decoding unit that decodes a code sequence from an encoded stream and derives a merge correction flag and correction vector indicating whether or not to correct merge candidates, If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. If neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, the merge candidate correction unit derives a corrected merge candidate for the dual prediction by subtracting the correction vector, scaled according to the interval between the prediction target picture and the reference picture, from the motion vector of the second prediction of the selected merge candidate. An image decoding device characterized in that, when the selected merge candidates are corrected, the number of selectable selected merge candidates is smaller compared to when the selected merge candidates are not corrected.

5. A merge candidate list generation step generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on the decoded image at the same position as the block to be predicted, A merge candidate replenishment step, which adds a supplemental merge candidate whose motion vector is (0,0) to the merge candidate list, A merge candidate selection step in which a merge candidate is selected from the aforementioned merge candidate list, A code sequence decoding step involves decoding a code sequence from an encoded stream to derive a merge correction flag and correction vector indicating whether or not to correct merge candidates, and If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. The merge candidate correction step includes, if neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, then subtracting the correction vector scaled according to the interval between the prediction target picture and the reference picture from the motion vector of the second prediction of the selected merge candidate to derive a corrected merge candidate for the dual prediction, An image decoding method characterized in that, when the selected merge candidates are corrected, the number of selectable selected merge candidates is smaller compared to when the selected merge candidates are not corrected.

6. A merge candidate list generation step generates a merge candidate list that includes, as merge candidates, motion information including motion information of multiple blocks adjacent to the block to be predicted and motion vectors derived by scaling the motion vector of a block on the decoded image at the same position as the block to be predicted, A merge candidate replenishment step, which adds a supplemental merge candidate whose motion vector is (0,0) to the merge candidate list, A merge candidate selection step in which a merge candidate is selected from the aforementioned merge candidate list, A code sequence decoding step involves decoding a code sequence from an encoded stream to derive a merge correction flag and correction vector indicating whether or not to correct merge candidates, and If the merge correction flag indicates that the merge candidate should be corrected, and either or both of the reference picture of the first prediction of the selected merge candidate or the reference picture of the second prediction of the selected merge candidate are long-term reference pictures, and the reference picture of the first prediction of the selected merge candidate and the reference picture of the second prediction of the selected merge candidate are in opposite directions with respect to the prediction target picture that includes the prediction target block, then the correction vector is added to the motion vector of the first prediction of the selected merge candidate without scaling, and the correction vector is subtracted from the motion vector of the second prediction of the selected merge candidate without scaling to derive the corrected merge candidate for both predictions. If neither the reference picture of the first prediction of the selected merge candidate nor the reference picture of the second prediction of the selected merge candidate is a long-term reference picture, then the computer is made to perform a merge candidate correction step, which involves subtracting the correction vector, scaled according to the interval between the predicted picture and the reference picture, from the motion vector of the second prediction of the selected merge candidate to derive a corrected merge candidate for the dual prediction. An image decoding program characterized in that, when the selected merge candidates are corrected, the number of selectable merge candidates is smaller compared to when the selected merge candidates are not corrected.

7. A storage method for generating a bitstream according to the image encoding method described in claim 2 and storing the bitstream in a recording medium.

8. A transmission method for generating a bitstream according to the image encoding method described in claim 2, and transmitting the bitstream.

Citation Information

Patent Citations

  • Moving image encoding and decoding device using motion compensation inter-frame prediction system capable of area integration

    JP1998276439A