Video encoding device, video encoding method, and video encoding program
By constructing a merge candidate list with spatial and triangular merge candidates and encoding indices, the technique addresses high processing loads in image encoding, achieving efficient encoding and decoding of moving images with deformations.
Patent Information
- Application Number
- JP2025073444
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2039-03-08
AI Technical Summary
Existing image encoding techniques, such as those described in Patent Document 1, suffer from high processing loads due to image conversion processes, particularly when dealing with deformations like enlargement, reduction, and rotation in moving images.
The technique involves constructing a merge candidate list with spatial and triangular merge candidates, selecting appropriate merge candidates, and encoding indices to determine motion information efficiently, reducing processing load while maintaining high efficiency.
This approach enables highly efficient image encoding and decoding with reduced processing load, effectively handling deformations in moving images.
Smart Images

Figure 2025107241000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image encoding and decoding technique for dividing an image into blocks and performing prediction.
Background Art
[0002] In image encoding and decoding, an image to be processed is divided into blocks that are sets of a predetermined number of pixels, and processing is performed in block units. By appropriately dividing into blocks and appropriately setting intra prediction (intra prediction) and inter prediction (inter prediction), the encoding efficiency is improved. In the encoding and decoding of moving images, the encoding efficiency is improved by inter prediction that predicts from an encoded and decoded picture. Patent Document 1 describes a technique of applying an affine transformation during inter prediction. In moving images, it is not uncommon for an object to be accompanied by deformations such as enlargement, reduction, and rotation. By applying the technique of Patent Document 1, efficient encoding becomes possible.
[0003] In the encoding and decoding of moving images, the encoding efficiency is improved more by inter prediction that predicts from an encoded and decoded picture. Patent Document 1 describes a technique of applying an affine transformation during inter prediction. In moving images, it is not uncommon for an object to be accompanied by deformations such as enlargement, reduction, and rotation. By applying the technique of Patent Document 1, efficient encoding becomes possible.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, since the technique of Patent Document 1 involves image conversion, there is a problem that the processing load is extremely large. In view of the above problems, the present invention provides an encoding technique with low load and high efficiency.
Means for Solving the Problems
[0006] In one aspect of the present invention for solving the above problems, there are provided a merge candidate list construction unit that constructs a merge candidate list including spatial merge candidates, a normal merge candidate selection unit that selects normal merge candidates that are single prediction or dual prediction from the merge candidate list, a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list, a triangular merge candidate selection unit that selects a second index for identifying the second triangular merge candidate, and an encoding unit that encodes the first index for identifying the first triangular merge candidate and the second index for identifying the second triangular merge candidate. The triangular merge candidate selection unit uses the first index and the second index to first determine whether L0 motion information is included in the candidates in the merge candidate list. If L0 motion information is included, the L0 motion information is used as a triangular merge candidate. Next, it is determined whether L1 motion information is included in the candidates in the merge candidate list. If L1 motion information is included, the L1 motion information is used as a triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction unit adds two or more merge candidates to the merge candidate list.
Advantages of the Invention
[0007] According to the present invention, highly efficient image encoding and decoding processing can be realized with a low load.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64
Figure 65
Figure 66
Figure 67
Figure 68
Embodiments for Carrying Out the Invention
[0009] Define the technologies and technical terms used in this embodiment.
[0010] <Tree block> In the embodiment, the image to be encoded / decoded is evenly divided into units of a predetermined size. This unit is defined as a tree block. As shown in FIG. 4, in this embodiment, the size of the tree block is set to 128x128 pixels, but the size of the tree block is not limited to this and any size may be set. The tree block of the processing target (corresponding to the encoding target in the encoding process and the decoding target in the decoding process.) switches in raster scan order, that is, in the order from left to right and from top to bottom. The inside of each tree block can be further recursively divided. The block to be encoded / decoded after the tree block division is defined as an encoding block. Also, the tree block and the encoding block are collectively referred to as a block cluster. By performing appropriate block division, efficient encoding becomes possible. The size of the tree block can be a fixed value determined in advance by the encoding device and the decoding device, or a configuration can be adopted in which the size of the tree block determined by the encoding device is transmitted to the decoding device.
[0011] <Prediction mode> For each processing target coded block, perform intra prediction (MODE_INTRA) that makes a prediction from the image signals around the processed (in the case of encoding, the image obtained by decoding the signal for which encoding has been completed, the image signal, the tree block, the block, the coded block, etc.), and perform inter prediction (MODE_INTER) that makes a prediction from the image signal of the processed image. Switch between them. Define the mode for identifying this intra prediction (MODE_INTRA) and inter prediction (MODE_INTER) as the prediction mode (PredMode). The prediction mode (PredMode) has either intra prediction (MODE_INTRA) or inter prediction (MODE_INTER) as its value. It is used for the decoded image, image signal, tree block, block, coded block, etc. for which encoding has been completed in the encoding process, and in the decoding process, it is used for the image, image signal, tree block, block, coded block, etc. for which decoding has been completed. ) Perform intra prediction (MODE_INTRA) that makes a prediction from the image signals around the processed (in the case of encoding, the image obtained by decoding the signal for which encoding has been completed, the image signal, the tree block, the block, the coded block, etc.), and perform inter prediction (MODE_INTER) that makes a prediction from the image signal of the processed image. In the intra prediction (MODE_INTRA) that makes a prediction from the image signals around the processed (in the case of encoding, the image obtained by decoding the signal for which encoding has been completed, the image signal, the tree block, the block, the coded block, etc.), and the inter prediction (MODE_INTER) that makes a prediction from the image signal of the processed image, switch between them. Define the mode for identifying this intra prediction (MODE_INTRA) and inter prediction (MODE_INTER) as the prediction mode (PredMode). The prediction mode (PredMode) has either intra prediction (MODE_INTRA) or inter prediction (MODE_INTER) as its value. <Inter Prediction> In the inter prediction that makes a prediction from the image signal of the processed image, multiple processed images can be used as reference pictures. To manage multiple reference pictures, two types of reference lists, L0 (reference list 0) and L1 (reference list 1), are defined, and the reference pictures are specified using the respective reference indices. In the P slice, L0 prediction (Pred_L0) is available. In the B slice, L0 prediction (Pred_L0), L1 prediction (Pred_L1), and bi-prediction (Pred_BI) are available.
[0012] <Inter Prediction> In the inter prediction that makes a prediction from the image signal of the processed image, multiple processed images can be used as reference pictures. To manage multiple reference pictures, two types of reference lists, L0 (reference list 0) and L1 (reference list 1), are defined, and the reference pictures are specified using the respective reference indices. In the P slice, L0 prediction (Pred_L0) is available. In the B slice, L0 prediction (Pred_L0), L1 prediction (Pred_L1), and bi-prediction (Pred_BI) are available. L0 prediction (Pred_L0) is an inter prediction that refers to the reference pictures managed by L0. L1 prediction (Pred_L1) is an inter prediction that refers to the reference pictures managed by L1. Bi-prediction (Pred_BI) is an inter prediction in which both L0 prediction and L1 prediction are performed, and one reference picture managed by each of L0 and L1 is referred to. Bi-prediction (Pred_BI) is an inter prediction in which both L0 prediction and L1 prediction are performed, and one reference picture managed by each of L0 and L1 is referred to. Bi-prediction (Pred_BI) is an inter prediction in which both L0 prediction and L1 prediction are performed, and one reference picture managed by each of L0 and L1 is referred to. It is. Information for specifying L0 prediction, L1 prediction, and dual prediction is defined as an inter-prediction mode. . For constants and variables with subscript LX attached to the output in subsequent processing, it is assumed that the processing is performed for each of L0 and L1.
[0013] <Predicted motion vector mode> The predicted motion vector mode is a mode that transmits an index for specifying a predicted motion vector, a differential motion vector, an inter-prediction mode, and a reference index, and determines the inter-prediction information of the processing target block. The predicted motion vector is derived from a predicted motion vector candidate derived from a processed block adjacent to the processing target block or a block located at the same position or in its vicinity (neighborhood) in a block belonging to the processed image, and an index for specifying the predicted motion vector.
[0014] <Merge mode> The merge mode is a mode that does not transmit a differential motion vector and a reference index, and derives the inter-prediction information of the processing target block from the inter-prediction information of a processed block adjacent to the processing target block or a block located at the same position or in its vicinity (neighborhood) in a block belonging to the processed image.
[0015] A processed block adjacent to the processing target block and its inter- prediction information are defined as spatial merge candidates. A block located at the same position or in its vicinity (neighborhood) in a block belonging to the processed image as the processing target block, and the inter-prediction information derived from the inter-prediction information of that block are defined as temporal merge candidates. Each merge candidate The prediction of the processing target block is performed by the merge index, and the prediction candidate is registered in the merge candidate list. The merge candidate to be used is specified.
[0016] <Adjacent block> FIG. 11 is a diagram for explaining a reference block referred to for deriving inter prediction information in the predicted motion vector mode and the merge mode. A0, A1, A2, B0, B1, B2, B3 are processed blocks adjacent to the processing target block. T0 is a block belonging to the processed image and is located at the same position as or near the processing target coded block of the processing target image. A0, A1, A2 are located on the left side of the processing target coded block and are adjacent to the processing target coded block. B1, B3 are located above the processing target coded block and are adjacent to the processing target coded block. A0, B0, B2 are located at the lower left, upper right, and upper left of the processing target coded block, respectively. The details of how to handle adjacent blocks in the predicted motion vector mode and the merge mode will be described later.
[0017]
[0018]
[0019] <Affine transform motion compensation> Affine transform motion compensation divides a coded block into sub-blocks of a predetermined unit, and performs motion compensation by setting motion vectors for each sub-block individually. The motion vector of each sub-block is derived based on one or more control points derived from the inter prediction information of a block that is a processed block adjacent to the processing target block or a block belonging to the processed image and located at the same position as or near the processing target block. In the present embodiment, In the form, the size of the sub-block is 4×4 pixels, but the size of the sub-block is not limited to this, and motion vectors may be derived in pixel units. It is not limited thereto, and motion vectors may be derived in pixel units.
[0020] Fig. 14 shows an example of affine transform motion compensation when there are two control points. In this case, since the two control points have two parameters, namely the horizontal component and the vertical component, the affine transform in the case of two control points is called a four-parameter affine transform. CP1 and CP2 in Fig. 14 are control points. Fig. 15 shows an example of affine transform motion compensation when there are three control points. In this case, since the three control points have two parameters, namely the horizontal component and the vertical component, the affine transform in the case of three control points is called a six-parameter affine transform. CP1, CP2, and CP3 in Fig. 15 are control points.
[0021] Affine transform motion compensation is available in both the predictive motion vector mode and the merge mode. The mode in which affine transform motion compensation is applied in the predictive motion vector mode is defined as the sub-block predictive motion vector mode, and the mode in which affine transform motion compensation is applied in the merge mode is defined as the sub-block merge mode. motion compensation is applied in the merge mode is defined as the sub-block merge mode. motion compensation is applied in the merge mode is defined as the sub-block merge mode.
[0022] <Syntax of coded block> Using Fig. 12(a), Fig. 12(b), and Fig. 13, the syntax for representing the prediction mode of the coded block is described. The pred_mode_flag in Fig. 12(a) is a flag indicating whether it is an inter prediction or not. If pred_mode_flag is 0, it is an inter prediction, and if pred_mo de_flag is 1, it is an intra prediction. In the case of intra prediction, the intra prediction information i de_flag is 1, it is an intra prediction. In the case of intra prediction, the intra prediction information i de_flag is 1, it is an intra prediction. In the case of intra prediction, the intra prediction information i Send the intra_pred_mode, and send the merge_flag in the case of inter prediction. The merge_flag is a flag indicating whether to use the merge mode or the predicted motion vector mode. In the case of the predicted motion vector mode (merge_flag = 0), send the inter_affine_flag, which is a flag indicating whether to apply the sub-block predicted motion vector mode. When applying the sub-block predicted motion vector mode (inter_affine_flag = 1), send the cu_affine_type_flag. The cu_affine_type_fla g is a flag for determining the number of control points in the sub-block predicted motion vector mode.
[0023] On the other hand, in the case of the merge mode (merge_flag = 1), send the merge_subblock_flag in Fig. 12(b). The merge_subblock_flag is a flag indicating whether to apply the sub-block merge mode. In the case of the sub-block merge mode (merge_subblock_flag = 1), send the merge sub-block index merge_subblock_idx. On the other hand, when not in the sub-block merge mode (merge_ subblock_flag = 0), send the merge_triangle_flag, which is a flag indicating whether to apply the triangular merge mode. When applying the triangular merge mode (merge_triangle_flag = 1), send the direction merge_triangle_split_dir in which the block is divided, and the merge triangular indices merge_triangle_idx0, merge_triangle_idx1 for each of the two partitions after division. On the other hand, in the case of the triangular merge mode When the single mode is not applied (merge_triangle_flag = 0), the merge index merge_idx is sent.
[0024] Figure 13 shows the values of each syntax element and the corresponding prediction mode. merge_ flag = 0, inter_affine_flag = 0 corresponds to the normal prediction motion vector mode (Inter Pred Mode). merge_flag = 0, inter_affine_flag = 1 corresponds to the sub-block prediction motion vector mode ( Inter Affine Mode). merge_flag = 1, merge_subblock_flag = 0, merge_trianlge _flag = 0 corresponds to the normal merge mode (Merge Mode). merge_flag = 1, merge_subblock _flag = 0, merge_trianlge_flag = 1 corresponds to the triangle merge mode (Triangle Merge Mode). merge_flag = 1, merge_subblock_flag = 1 corresponds to the sub-block merge mode (Affine Merge Mode).
[0025] <poc> POC (Picture Order Count) is a variable associated with the picture to be encoded, and a value that increases by 1 in the output order of the picture is set. Depending on the value of the POC, it is possible to determine whether it is the same picture, determine the order relationship between pictures in the output order, or derive the distance between pictures. For example, if the POCs of two pictures have the same value, it can be determined that they are the same picture. If the POCs of two pictures have different values, it can be determined that the picture with the smaller POC value is the picture that is output first, and the difference between the POCs of the two pictures indicates the distance between the pictures in the time axis direction.
[0026] (First Embodiment) The image encoding device 100 and the image decoding device 200 according to the first embodiment of the present invention will be described.
[0027] FIG. 1 is a block diagram of the image encoding device 100 according to the first embodiment. The moving image encoding device of the embodiment includes an image encoding device 100, a block division unit 101, an inter prediction unit 102, an intra prediction unit 103, a decoded image memory 104, a prediction method determination unit 105, a residual signal generation unit 106, an orthogonal transform / quantization unit 107, a bit string encoding unit 108, an inverse quantization / inverse orthogonal transform unit 109, a decoded image signal superposition unit 110, and an encoded information storage memory 111.
[0028] The block division unit 101 recursively divides the input image to generate encoding blocks. The block division unit 101 divides the block to be divided into four parts in the horizontal and vertical directions respectively, and divides the block to be divided into either the horizontal or vertical direction. It includes 2-3 divided parts to be cut. The generated image signal of the processing target coded block is supplied to the inter prediction unit 102, the intra prediction unit 103, and the residual signal generation unit 106. Also, the information indicating the determined recursive division structure is supplied to the bit sequence coding unit 108. The detailed operation of the block division unit 101 will be described later. The inter prediction unit 102 performs inter prediction on the processing target coded block. It derives a plurality of candidate inter prediction information from the inter prediction information stored in the coding information storage memory and the decoded image signal stored in the decoded image memory 104, selects a suitable inter prediction mode from among the plurality of candidates, and supplies the selected inter prediction mode and the prediction image signal corresponding to the selected inter prediction mode to the prediction method determination unit 105. The detailed configuration and operation of the inter prediction unit 102 will be described later. The intra prediction unit 103 performs intra prediction on the processing target coded block. It generates a prediction image signal by intra prediction from the decoded image signal stored in the decoded image memory 104, selects a suitable intra prediction mode from among the plurality of intra prediction modes, and supplies the selected intra prediction mode and the prediction image signal corresponding to the selected intra prediction mode to the prediction method determination unit 105. An example of intra prediction is shown in FIG. 10. FIG. 10(a) shows the correspondence between the prediction direction of intra prediction and the intra prediction mode number. For example, the intra prediction mode 50 generates an intra prediction image by copying pixels in the vertical direction. The intra prediction mode 1 is the DC mode, in which the values of all pixels of the processing target block are set to the average value of the reference pixels. The intra prediction mode 0 is the Planar mode. The detailed operation of the block division unit 101 will be described later.
[0029] The inter prediction unit 102 performs inter prediction on the processing target coded block. Derive a plurality of candidate inter prediction information from the inter prediction information stored in the coding information storage memory and the decoded image signal stored in the decoded image memory 104. Select a suitable inter prediction mode from among the plurality of candidates. And supply the selected inter prediction mode and the prediction image signal corresponding to the selected inter prediction mode to the prediction method determination unit 105. The detailed configuration and operation of the inter prediction unit 102 will be described later. The detailed operation of the block division unit 101 will be described later.
[0030] The intra prediction unit 103 performs intra prediction on the processing target coded block. Generate a prediction image signal by intra prediction from the decoded image signal stored in the decoded image memory 104. Select a suitable intra prediction mode from among the plurality of intra prediction modes. And supply the selected intra prediction mode and the prediction image signal corresponding to the selected intra prediction mode to the prediction method determination unit 105. An example of intra prediction is shown in FIG. 10. FIG. 10(a) shows the correspondence between the prediction direction of intra prediction and the intra prediction mode number. For example, the intra prediction mode 50 generates an intra prediction image by copying pixels in the vertical direction. The intra prediction mode 1 is the DC mode, in which the values of all pixels of the processing target block are set to the average value of the reference pixels. The intra prediction mode 0 is the Planar mode. and is a mode for creating a two-dimensional intra prediction image from reference pixels in the vertical and horizontal directions. Figure 10(b) is an example of generating an intra prediction image in the case of intra prediction mode 40. For each pixel of the block to be processed, the value of the reference pixel in the direction indicated by the intra prediction mode is copied. When the reference pixel of the intra prediction mode is not at an integer position, the reference pixel value is determined by interpolation from the reference pixel values at the surrounding integer positions.
[0031] The decoded image memory 104 stores the decoded image generated in the decoded image signal superposition unit 110. The decoded image stored in the decoded image memory is supplied to the inter prediction unit 102 and the intra prediction unit 1 03.
[0032] The prediction method determination unit 105 evaluates each prediction using the encoded information, the coded amount of the residual signal, the distortion amount between the predicted image signal and the image signal to be processed, etc., and determines the optimal prediction mode (inter prediction or intra prediction). In the case of the merge mode of inter prediction, the encoded information of the merge index and the information indicating whether it is a sub-block merge mode (sub-block merge flag) is supplied to the bit string encoding unit 108, and the predicted motion vector mode of inter prediction, the encoded information such as the inter prediction mode, the predicted motion vector index, the reference indexes of L0 and L 1, the differential motion vector, and the information indicating whether it is a sub-block mode (sub-block predicted motion vector flag) is supplied to the bit string encoding unit 108. The determined encoded information is supplied to the encoded information storage memory 111.
[0033] The residual signal generation unit 106 subtracts the predicted image signal from the image signal to be processed to obtain the residual Generate a difference signal and supply it to the orthogonal transformation / quantization unit 107.
[0034] The orthogonal transformation / quantization unit 107 performs orthogonal transformation and quantization on the residual signal according to the quantization parameter, generates an orthogonally transformed / quantized residual signal, and supplies it to the bit sequence encoding unit 108 and the inverse quantization / inverse orthogonal transformation unit 109. and inverse orthogonal transformation unit 109.
[0035] The bit sequence encoding unit 108 encodes the encoding information corresponding to the prediction method determined by the prediction method determination unit 105 for each encoding block in addition to the information in units of sequence, picture, slice, and encoding block. Specifically, for each encoding block, the prediction mode PredMode, partition mode PartMode, in the case of inter prediction (PRED_INTER), a flag for determining whether it is a merge mode, a sub-block merge flag, a merge index in the case of merge mode, an inter prediction mode in the case of non-merge mode, a prediction motion vector index, information regarding the differential motion vector, encoding information such as a sub-block prediction motion vector flag, etc. is encoded according to the following prescribed syntax rules to generate a first encoded bit sequence. Also, the bit sequence encoding unit 108 entropy-encodes the orthogonally transformed and quantized residual signal according to the prescribed syntax rules to generate a second encoded bit sequence. The first encoded bit sequence and the second encoded bit sequence are multiplexed according to the prescribed syntax rules, and a bit stream is output.
[0036] The inverse quantization / inverse orthogonal transformation unit 109 performs inverse quantization and inverse orthogonal transformation on the orthogonally transformed / quantized residual signal supplied from the orthogonal transformation / quantization unit 107 to calculate a residual signal, and supplies it to the decoded image signal superposition unit 110.
[0037] The decoded image signal superposition unit 110 superimposes the predicted image signal according to the determination by the prediction method determination unit 105 and the residual signal that has been inverse quantized and inverse orthogonally transformed by the inverse quantization and inverse orthogonal transformation unit 109 to generate a decoded image, which is stored in the decoded image memory 104. Note that, after performing filtering processing to reduce distortion such as block distortion caused by encoding on the decoded image, it may be stored in the decoded image memory 104.
[0038] The encoding information storage memory 111 stores encoding information such as the prediction mode (inter prediction or intra prediction) determined by the prediction method determination unit 105. The encoding information stored in the encoding information storage memory 111 in the case of inter prediction, in addition to the determined motion vector, reference list, reference index, in the case of the merge mode of inter prediction, the merge index, encoding information of information (sub-block merge flag) indicating whether it is a sub-block merge mode, in the case of the predicted motion vector mode of inter prediction, the inter prediction mode, predicted motion vector index, reference indices of L0 and L1, differential motion vector, information (sub-block predicted motion vector flag) indicating whether it is a sub-block mode, in the case of intra prediction, it is the determined intra prediction mode, etc. The construction of the history candidate compensation list managed by the encoding information storage memory 111 will be described later.
[0039] FIG. 2 is a block diagram showing the configuration of a moving image decoding apparatus according to an embodiment of the present invention corresponding to the moving image encoding apparatus of FIG. 1. The moving image decoding apparatus of the embodiment includes a bit string decoding unit 201, a block division unit 202, an inter prediction unit 203, an intra prediction unit 204, an encoding information storage memory It includes a memory 205, an inverse quantization and inverse orthogonal transformation unit 206, a decoded image signal superposition unit 207, and a decoded image memory 208.
[0040] The decoding process of the moving image decoder in FIG. 2 corresponds to the decoding process provided inside the moving image encoder in FIG. 1. Therefore, each component of the encoding information storage memory 205, inverse quantization and inverse orthogonal transformation unit 206, decoded image signal superposition unit 207, and decoded image memory 208 in FIG. 2 corresponds to each component of the inverse quantization and inverse orthogonal transformation unit 109, decoded image signal superposition unit 110 in the moving image encoder in FIG. 1, the encoding information storage memory 111, and the decoded image memory 104, and has corresponding functions.
[0041] The bit stream supplied to the bit sequence decoder 201 is separated according to the rules of a specified syntax. The separated first encoded bit sequence is decoded to obtain information such as sequence, picture, slice, encoding information in units of encoding blocks, and encoding information in units of encoding blocks. Specifically, a prediction mode PredMode for determining whether it is an inter prediction (PRED_INTER) or an intra prediction (PRED_INTRA) in units of encoding blocks, a split mode PartMode, a flag for determining whether it is a merge mode in the case of inter prediction (PRED_INTER), a merge index in the case of the merge mode, a sub-block merge flag, an inter prediction mode in the case of a prediction motion vector mode, a prediction motion vector index, a differential motion vector, a sub-block prediction motion vector flag, etc. The encoding information is decoded according to the rules of the specified syntax described later, and the encoding information is sent to the inter prediction unit 203 or the intra prediction unit 204, and the encoding information storage memory ) or the like, and the encoding information is sent to the inter prediction unit 203 or the intra prediction unit 204, and the encoding information storage memory In the case of, a flag for determining whether it is a merge mode, a merge index in the case of the merge mode, a sub-block merge flag, an inter prediction mode in the case of a prediction motion vector mode, a prediction motion vector index, a differential motion vector, a sub-block prediction motion vector flag, etc. The encoding information is decoded according to the rules of the specified syntax described later, and the encoding information is sent to the inter prediction unit 203 or the intra prediction unit 204, and the encoding information storage memory According to the rules of the specified syntax described later, the encoding information related to the flag, etc. is decoded, and the encoding information is sent to the inter prediction unit 203 or the intra prediction unit 204, and the encoding information storage memory Supply it to 205. Decode the separated second encoded bit sequence to calculate the orthogonally transformed and quantized residual difference signal, and supply the orthogonally transformed and quantized residual difference signal to the inverse quantization and inverse orthogonal transformation unit 206 .
[0042] When the prediction mode PredMode of the encoding block to be processed is the inter prediction (PRED_INTER) and the prediction motion vector mode, the inter prediction unit 203 uses the encoding information of the already decoded image signal stored in the encoding information storage memory 205 to derive a plurality of candidate prediction motion vectors, register them in the prediction motion vector candidate list described later, and select a prediction motion vector corresponding to the prediction motion vector index decoded and supplied by the bit sequence decoding unit 201 from among the plurality of candidate prediction motion vectors registered in the prediction motion vector candidate list . Calculate the motion vector from the difference vector decoded by the bit sequence decoding unit 201 and the selected prediction motion vector, and store it in the encoding information storage memory 205 together with other encoding information. Here . The encoding information of the encoding block supplied and stored here includes the prediction mode PredMode, the partition mode Pa rtMode, flags predFlagL0[xP][yP], pr edFlagL1[xP][yP] indicating whether to use L0 prediction and L1 prediction, the reference indexes refIdxL0[xP][yP], refIdxL1[xP][yP] of L0 and L1 , the motion vectors mvL0[xP][yP], mvL1[xP][yP] of L0 and L1, etc. Here, xP and yP are indexes indicating the position of the upper left pixel of the encoding block within the picture element . When the prediction mode PredMode is inter prediction (MODE_INTER) and the inter prediction mode is L0 prediction (Pred_L0) . . . . . . In this case, the flag predFlagL0 indicating whether to use L0 prediction is 1, and whether to use L1 prediction The flag predFlagL1 indicating whether or not is 0. When the inter prediction mode is L1 prediction (Pred_L1) In this case, the flag predFlagL0 indicating whether to use L0 prediction is 0, and whether to use L1 prediction The flag predFlagL1 indicating whether or not is 1. When the inter prediction mode is dual prediction (Pred_BI) In this case, both the flag predFlagL0 indicating whether to use L0 prediction and the flag predFlagL1 indicating whether to use L1 prediction are both 1. Further, when the prediction mode Pred Mode of the encoding block to be processed is inter prediction (PRED_INTER) and in merge mode, merge candidates are derived. The encoding information of the already decoded encoding block stored in the encoding information storage memory 205 is used to derive a plurality of merge candidates, which are registered in the merge candidate list described later. Among the plurality of merge candidates registered in the merge candidate list, the merge candidate corresponding to the merge index decoded and supplied by the bit string decoder 201 is selected, and the L0 prediction of the selected merge candidate and the flags predFlagL0[xP][yP], predFlagL1[xP][yP] indicating whether to use L1 prediction the reference indices refIdxL0[xP][yP], refIdxL1[xP][yP] of L0 and L1, and the motion vectors mvL0[xP][yP], mvL1[xP][yP] of L0 and L1, etc. of the inter prediction information are stored in the encoding information storage memory 2 05. Here, xP and yP are indices indicating the position of the upper left pixel of the encoding block within the picture The detailed configuration and operation of the inter prediction unit will be described later
[0043] The intra prediction unit 204 performs intra prediction when the prediction mode PredMode of the encoding block to be processed is intra prediction (PRED_INTRA). The encoded information decoded by the bit string decoding unit 201 includes an intra prediction mode. According to the intra prediction mode, an intra prediction is performed from the decoded image signal stored in the decoded image memory 208 to generate a predicted image signal, and the predicted image signal is supplied to the decoded image signal superposition unit 207. Since the intra prediction unit 20 4 corresponds to the intra prediction unit 103 of the image encoding apparatus 100, it performs the same processing as the intra prediction unit 103. like image signals. The inverse quantization and inverse orthogonal transformation unit 206 performs inverse orthogonal transformation and inverse quantization on the orthogonal transformation and quantization residual signal decoded by the bit string decoding unit 201 to obtain an inverse orthogonal transformation and inverse quantization residual signal.
[0044] The decoded image signal superposition unit 207 superimposes the predicted image signal inter-predicted by the inter prediction unit 203 or the predicted image signal intra-predicted by the intra prediction unit 204 and the residual signal inverse orthogonally transformed and inverse quantized by the inverse quantization and inverse orthogonal transformation unit 206 to decode the encoded image signal and store it in the decoded image memory 208. When storing in the decoded image memory 208, filtering processing for reducing block distortion due to encoding may be performed on the decoded image, and then stored in the decoded image memory 208.
[0045] Next, the operation of the block division unit 101 in the image encoding apparatus 100 will be described.
[0046] Next, the operation of the block division unit 101 in the image encoding apparatus 100 will be described. FIG. 3 shows the operation of dividing an image into tree blocks and further dividing each tree block. It is a flowchart. First, the input image is divided into tree blocks of a predetermined size (step S1001). For each tree block, it is scanned in a predetermined order, that is, in raster scan order (step S1002), and the inside of the tree block to be processed is divided (step S1003).
[0047] FIG. 7 is a flowchart showing the detailed operation of the splitting process in step S1003. First , it is determined whether to divide the block to be processed into four parts (step S1101).
[0048] If it is determined to divide the block to be processed into four parts, the block to be processed is divided into four parts (step S1102). For each block obtained by dividing the block to be processed, it is scanned in the Z-scan order, that is, in the order of upper left, upper right, lower left, and lower right (step S1103). FIG. 5 is an example of the Z-scan order, and 601 in FIG. 6 is an example of dividing the block to be processed into four parts. The numbers 0 to 3 in 601 of FIG. 6 indicate the order of processing. Then, for each block divided in step S1101, the flowchart of FIG. 7 is recursively called. If it is determined not to divide the block to be processed into four parts, a 2-3 split is performed (step S1
[0049] 105).
[0050] FIG. 8 is a flowchart showing the detailed operation of the 2-3 split process in step S1105 . First, it is determined whether to perform a 2-3 split on the block to be processed, that is, whether to perform either a 2-way split or a 3-way split (step S1201).
[0051] If it is determined not to perform a 2-3 split on the block to be processed, that is, if it is determined not to split If so, end the division (step S1211) and return to the block in the upper hierarchy.
[0052] If it is determined to divide the block to be processed into two or three parts, further determine whether to divide the block to be processed into two parts (step S1202).
[0053] If it is determined to divide the block to be processed into two parts, determine whether to divide the block to be processed vertically (step S1203). Based on the result, divide the block to be processed vertically (step S1204) or divide the block to be processed horizontally ( step S1205). As a result of step S1204, the block to be processed is divided into two vertical parts as shown in FIG. 602, and as a result of step S1205, the block to be processed is divided into two horizontal parts as shown in FIG. 604.
[0054] If it is not determined to divide the block to be processed into two parts in step S1202, that is, if it is determined to divide it into three parts, determine whether to divide the block to be processed vertically (step S1206). Based on the result, divide the block to be processed vertically (step S1207) or divide the block to be processed horizontally (step S 1208). As a result of step S1207, the block to be processed is divided into three vertical parts as shown in FIG. 603, and as a result of step S1208, the block to be processed is divided into three horizontal parts as shown in FIG. 605.
[0055] After executing either step S1204 or step S1205, scan each divided block of the block to be processed in the order from left to right and from top to bottom (step S1209 ). )。The numbers 0 to 3 from 602 to 605 in FIG. 6 indicate the order of processing. For each of the divided blocks, the flowchart of FIG. 8 is recursively called.
[0056] The recursive block division described here may limit the necessity of division according to the number of divisions or the size of the block to be processed, etc. The information for limiting the necessity of division may be realized in a configuration where information transmission is not performed by making a prior agreement between the encoding device and the decoding device, or may be realized in a configuration where the encoding device determines the information for limiting the necessity of division and records it in the encoded bit sequence for transmission to the decoding device.
[0057] Here, when a certain block is divided, the block before division is called the parent block, and each block after division is called a child block.
[0058] Next, the operation of the block division unit 202 in the image decoding device 200 will be described. The block division unit 202 divides the tree block by the same processing procedure as the block division unit 101 of the image encoding device 100. However, in the block division unit 101 of the image encoding device 100, optimization methods such as estimation of the optimal shape by image recognition and distortion rate optimization are applied to determine the optimal block division shape, while the block division unit 202 in the image decoding device 200 determines the block division shape by decoding the block division information recorded in the encoded bit sequence, which is different.
[0059] The syntax (syntax rules of the encoded bit sequence) regarding the block division of the first embodiment is shown in FIG. 9. coding_quadtree() represents the syntax related to the 4-division processing of the block , multi_type_tree() represents the syntax related to the two-way or three-way block splitting process . qt_split is a flag indicating whether to split the block into four parts. If the block is split into four parts , then qt_split = 1; if not, then qt_split = 0. When splitting into four parts (qt_split = 1), for each of the four split blocks, recursively perform the four-way splitting process (coding_quadtree(0), c oding_quadtree(1), coding_quadtree(2), coding_quadtree(3)). When not splitting into four parts (qt_ split = 0), subsequent splitting is determined according to multi_type_tree(). mtt_split is a flag indicating whether to further split. When further splitting (mtt_split = 1), refer to mtt_split_vertical, which is a flag indicating whether to split vertically or horizontally, and mtt_split_binary, which is a flag for determining whether to split into two parts or three parts. mtt_split_vertic al = 1 indicates splitting vertically, and mtt_split_vertical = 0 indicates splitting horizontally . mtt_split_binary = 1 indicates splitting into two parts, and mtt_split_binary = 0 indicates splitting into three parts . Hierarchical block splitting is performed by recursively calling multi_type_tree until mtt_split = 0 .
[0060] <Inter prediction> The inter prediction method according to the embodiment is implemented in the inter prediction unit 10 2 of the moving image encoding device in FIG. 1 and the inter prediction unit 203 of the moving image decoding device in FIG. 2.
[0061] The inter prediction method according to the embodiment will be described with reference to the drawings. The inter prediction method is performed in either the encoding or decoding process in units of encoding blocks.
[0062] <Explanation of the inter prediction unit 102 on the encoding side> FIG. 16 is a diagram showing the detailed configuration of the inter prediction unit 102 of the moving image encoding apparatus of FIG. 1. The normal prediction motion vector mode derivation unit 301 derives a plurality of normal prediction motion vector candidates to select a prediction motion vector and calculates a difference vector from the detected motion vector. The detected inter prediction mode, reference index, motion vector, and calculated difference vector become the inter prediction information in the normal prediction motion vector mode. This inter prediction information is supplied to the inter prediction mode determination unit 305. The detailed configuration and processing of the normal prediction motion vector mode derivation unit 301 will be described later.
[0063] The normal merge mode derivation unit 302 derives a plurality of normal merge candidates and selects a normal merge candidate to obtain the inter prediction information in the normal merge mode. This inter prediction information is supplied to the inter prediction mode determination unit 305. The detailed configuration and processing of the normal merge mode derivation unit 302 will be described later.
[0064] The sub-block prediction motion vector mode derivation unit 303 derives a plurality of sub-block prediction motion vector candidates to select a sub-block prediction motion vector and calculates a difference vector from the detected motion vector. The detected inter prediction mode, reference index, motion vector, and calculated difference vector become the inter prediction information in the normal prediction motion vector mode. This inter prediction information is supplied to the inter prediction mode determination unit 305. The sub-block The detailed configuration and processing of the predicted motion vector mode derivation unit 303 will be described later.
[0065] The sub-block merge mode derivation unit 304 derives a plurality of sub-block merge candidates selects a sub-block merge candidate, and obtains the inter-prediction information of the sub-block merge mode . This inter-prediction information is supplied to the inter-prediction mode determination unit 305. The sub-block The detailed configuration and processing of the merge mode derivation unit 304 will be described later.
[0066] The inter-prediction mode determination unit 305 is based on the inter-prediction information supplied from the normal prediction motion vector mode derivation unit 301, the normal merge mode derivation unit 302, the sub-block prediction motion vector mode derivation unit 303, and the sub-block merge mode derivation unit 304, and determines the inter- prediction mode. The inter-prediction information corresponding to the determination result is supplied from the 305 of the inter-prediction mode determination unit to the motion compensation prediction unit 306.
[0067] Based on the determined inter-prediction information, the motion compensation prediction unit 306 performs inter-prediction on the reference image signal stored in the decoded image memory 1 04. The detailed configuration and processing will be described later.
[0068] <Explanation of the inter-prediction unit 203 on the decoding side> FIG. 22 is a diagram showing the detailed configuration of the inter-prediction unit 203 of the moving image decoding apparatus in FIG. 2.
[0069] The normal prediction motion vector mode derivation unit 401 derives a plurality of normal prediction motion vector candidates selects a prediction motion vector, and calculates a difference vector from the detected motion vector. The detected inter-prediction mode, reference index, motion vector, and difference vector are normal prediction It becomes the inter prediction information in the motion vector mode. This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 The detailed configuration and processing of the normal prediction motion vector mode derivation unit 40 1 will be described later.
[0070] The normal merge mode derivation unit 402 derives a plurality of normal merge candidates and selects a normal merge candidate to obtain the inter prediction information in the normal merge mode. This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408. The detailed configuration and processing of the normal merge mode derivation unit 402 will be described later. The detailed configuration and processing of the normal merge mode derivation unit 402 will be described later. The detailed configuration and processing of the normal merge mode derivation unit 402 will be described later.
[0071] The sub-block prediction motion vector mode derivation unit 403 derives a plurality of sub-block prediction motion vector candidates, selects a sub-block prediction motion vector, and calculates the difference vector from the detected motion vector and The detected inter prediction mode, reference index, motion vector, and calculated difference vector become the inter prediction information in the normal prediction motion vector mode. This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 The detected inter prediction mode, reference index, motion vector, and calculated difference vector become the inter prediction information in the normal prediction motion vector mode. This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 The detected inter prediction mode, reference index, motion vector, and calculated difference vector become the inter prediction information in the normal prediction motion vector mode. This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 The detailed configuration and processing of the sub-block prediction motion vector mode derivation unit 403 will be described later. will be described later.
[0072] The sub-block merge mode derivation unit 404 derives a plurality of sub-block merge candidates and selects a sub-block merge candidate to obtain the inter prediction information in the sub-block merge mode The sub-block merge mode derivation unit 404 derives a plurality of sub-block merge candidates and selects a sub-block merge candidate to obtain the inter prediction information in the sub-block merge mode This inter prediction information is supplied to the motion compensation prediction unit 406 via the switch 408 The detailed configuration and processing of the sub-block merge mode derivation unit 404 will be described later.
[0073] Based on the inter-prediction information determined by the motion compensation prediction unit 406, an inter-prediction is performed on the reference image signal stored in the decoded image memory 2 08. The detailed configuration and processing are the same as those on the encoding side.
[0074] <Normal prediction motion vector mode derivation unit (Normal AMVP)> The normal prediction motion vector mode derivation unit 301 in FIG. 17 includes a spatial prediction motion vector candidate derivation unit 321, a temporal prediction motion vector candidate derivation unit 322, a history prediction motion vector candidate derivation unit 3 23, a prediction motion vector candidate supplementation unit 325, a normal motion vector detection unit 326, a prediction motion ve ctor candidate selection unit 327, and a motion vector subtraction unit 328.
[0075] The normal prediction motion vector mode derivation unit 401 in FIG. 23 includes a spatial prediction motion vector candidate derivation unit 421, a temporal prediction motion vector candidate derivation unit 422, a history prediction motion vector candidate derivation unit 4 23, a prediction motion vector candidate supplementation unit 425, a prediction motion vector candidate selection unit 426, and a motion ve ctor addition unit 427.
[0076] Regarding the processing procedures of the normal prediction motion vector mode derivation unit 301 on the encoding side and the normal prediction motion ve ctor mode derivation unit 401 on the decoding side, the flowcharts in FIGS. 19 and 25 will be used to explain them respectively. FIG. 19 is a flowchart showing the normal prediction motion vector mode derivation processing procedure by the normal motion vector mode derivation unit 301 on the encoding side, and FIG. 25 is the normal flowchart showing the normal prediction motion vector mode derivation processing procedure by the normal motion vector mode derivation unit 401 on the decoding side is shown.
[0077] <Normal prediction motion vector mode derivation unit (Normal AMVP): Explanation on the encoding side> Referring to FIG. 19, the normal prediction motion vector mode derivation processing procedure on the encoding side will be described. FIG In the description of the processing procedure of FIG. 19, the term motion vector in the specification and the normal motion vector term in FIG. 19 shall correspond. First, the normal motion vector detection unit 326 detects the normal motion vector for each inter prediction mode and reference index (step S100 in FIG. 19).
[0078] Subsequently, the spatial prediction motion vector candidate derivation unit 321, the temporal prediction motion vector candidate derivation unit 3 22, the history prediction motion vector candidate derivation unit 323, the prediction motion vector candidate supplementation unit 325, the prediction motion vector candidate selection unit 327, and the motion vector subtraction unit 328 calculate the differential motion vectors of the motion vectors used in the inter prediction of the normal prediction motion vector mode for each of L0 and L1 (steps S101 to S106 in FIG. 19). Specifically, when the prediction mode PredMode of the processing target block is inter prediction (MODE_INTER) and the inter prediction mode is L0 prediction (Pr ed_L0), the prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 is selected, and the differential motion vector mvdL0 of the motion vector mvL0 for L0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion of L0 vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi prediction (Pred_BI), both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion Calculate the differential motion vector mvdL0 of the vector mvL0, and calculate the prediction motion vector candidate complementary list mvpListL1 of L1, calculate the prediction motion vector mvpL1 of L1, and calculate the differential motion vector mvdL1 of the motion vector mvL1 of L1 respectively.
[0079] For each of L0 and L1, perform differential motion vector calculation processing. However, for both L0 and L1 it is a common process. Therefore, in the following description, L0 and L1 are represented as a common LX for simplicity. In the process of calculating the differential motion vector of L0, X is 0, and in the process of calculating the differential motion vector of L1, X is 1. Also, during the process of calculating the differential motion vector of LX, when referring to the information of the other list instead of LX, the other list is represented as LY .
[0080] When using the motion vector mvLX of LX (step S102 in FIG. 19: YES), calculate the candidate of the prediction motion vector of LX and construct the prediction motion vector candidate list mvpListLX of LX (step S103 in FIG. 19). In the spatial prediction motion vector candidate derivation unit 321, the temporal prediction motion vector candidate derivation unit 322, the history prediction motion vector candidate derivation unit 323, and the prediction motion vector candidate supplementation unit 325 in the normal prediction motion vector mode derivation unit 301, derive a plurality of prediction motion vector candidates and construct the prediction motion vector candidate list mvpListLX. The detailed processing procedure of step S103 in FIG. 19 will be described later using the flowchart of FIG. 20 . Then, the prediction motion vector candidate selection unit 327 selects from the prediction motion vector candidate list of LX the prediction motion vector candidate of LX, and .
[0081] Subsequently, the prediction motion vector candidate selection unit 327 selects from the prediction motion vector candidate list of LX Select the predicted motion vector mvpLX of LX from the mvpListLX of LX (step S104 in FIG. 19). Calculate each differential motion vector, which is the difference between the motion vector mvLX and each candidate predicted motion vector mvpListLX[i] stored in the predicted motion vector candidate list mvpListLX. Calculate the amount of code when encoding these differential motion vectors for each element of the predicted motion vector candidate list mvpListLX. Then, among the elements registered in the predicted motion vector candidate list mvpListLX, select the candidate predicted motion vector mvpListLX[i] for which the amount of code for each candidate predicted motion vector is minimized as the predicted motion vector mvpLX, and obtain its index i. If there are multiple candidates for the predicted motion vector that results in the minimum amount of generated code in the predicted motion vector candidate list mvpListLX, select the candidate predicted motion vector mvpListLX[i] represented by the smaller index i in the predicted motion vector candidate list mvpListLX as the optimal predicted motion vector mvpLX, and obtain its index i.
[0082] Subsequently, in the motion vector subtraction unit 328, subtract the predicted motion vector mvpLX of LX selected from the motion vector mvLX of LX, and calculate the differential motion vector mvdLX of LX as mvdLX = mvLX - mvpLX (step S105 in FIG. 19).
[0083] <Normal Predicted Motion Vector Mode Derivation Unit (Normal AMVP): Explanation on the Decoding Side> Next, with reference to FIG. 25, the normal predicted motion vector mode processing procedure on the decoding side will be described. On the decoding side, a spatial predicted motion vector candidate derivation unit 421 and a temporal predicted motion vector candidate derivation unit 4 22. In the history prediction motion vector candidate derivation unit 423 and the prediction motion vector candidate supplementation unit 425, The motion vectors used in the inter prediction of the normal prediction motion vector mode are calculated for each of L0 and L1 respectively (Steps S201 to S206 in FIG. 25). Specifically, for the processing target block When the prediction mode PredMode is inter prediction (MODE_INTER) and the inter prediction mode of the processing target block is L0 prediction (Pred_L0), the prediction motion vector candidate list mvpListL0 for L0 is calculated and the prediction motion vector mvpL0 is selected, and the motion vector mvL0 for L0 is calculated. When the inter prediction mode of the processing target block is L1 prediction (Pred_L1), the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 is selected, and the motion vector mvL1 for L1 is calculated. When the inter prediction mode of the processing target block is bi-prediction (Pred_BI) , both L0 prediction and L1 prediction are performed. The prediction motion vector candidate list mvpListL0 for L0 is calculated, the prediction motion vector mvpL0 for L0 is selected, and the motion vector mvL0 for L0 is calculated. At the same time, the prediction motion vector candidate list mvpListL1 for L1 is calculated, the prediction motion vector mvpL1 for L1 is calculated, and the motion vector mvL1 for L1 is calculated respectively. Similar to the encoding side, on the decoding side, for each of L0 and L1, the motion vector calculation process is performed, but the process is common to both L0 and L1. Therefore, in the following description, L0 and L1 are represented by a common LX. LX represents the inter prediction mode used for the inter prediction of the encoding block to be processed. In the process of calculating the motion vector for L0, X is 0, and for L1
[0084] X is 1. The same process is performed for both L0 and L1. Therefore, in the following description, L0 and L1 are represented by a common LX. LX represents the inter prediction mode used for the inter prediction of the encoding block to be processed. In the process of calculating the motion vector for L0, X is 0, and for L1 X is 1. In the process of calculating the motion vector of, X is 1. Also, when calculating the motion vector of LX During the process, instead of using the same reference list as the LX to be calculated, the information of the other reference list is referenced. In this case, the other reference list is represented as LY.
[0085] When using the motion vector mvLX of LX (step S202 in FIG. 25: YES), calculate the candidate of the predicted motion vector of LX and construct the predicted motion vector candidate list mvpListLX of LX (step S203 in FIG. 25). In the normal predicted motion vector mode derivation unit 401, the spatial predicted motion vector candidate derivation unit 421, the temporal predicted motion vector candidate derivation unit 422, the history predicted motion vector candidate derivation unit 423, and the predicted motion vector candidate supplement unit 425 calculate candidates for a plurality of predicted motion vectors and construct the predicted motion vector candidate list mvpListLX. The detailed processing procedure of step S203 in FIG. 25 will be described later using the flowchart of FIG. 20 (to be continued). (to be continued). (to be continued). (to be continued).
[0086] Subsequently, the predicted motion vector candidate selection unit 426 selects the candidate mvpListLX[mvpIdxLX] of the predicted motion vector corresponding to the index mv pIdxLX of the predicted motion vector decoded and supplied by the bit sequence decoding unit 201 from the predicted motion vector candidate list mvpListLX and takes it out as the selected predicted motion vector mvpLX (step S204 in FIG. 25). (to be continued).
[0087] Subsequently, the motion vector addition unit 427 adds the differential motion vector mvdLX of LX decoded and supplied by the bit sequence decoding unit 201 and the predicted motion vector mvpLX of LX, and mvLX = mvpLX + mvdLX Calculate the motion vector mvLX of LX as (step S205 in FIG. 25).
[0088] <Normal prediction motion vector mode derivation unit (normal AMVP): Motion vector prediction method> FIG. 20 shows the normal prediction motion vector mode derivation common to the normal prediction motion vector mode derivation unit 301 of the moving image encoding apparatus and the normal prediction motion vector mode derivation unit 401 of the moving image decoding apparatus It is a flowchart showing the processing procedure of the normal prediction motion vector mode derivation process having the function. is.
[0089] The normal prediction motion vector mode derivation unit 301 and the normal prediction motion vector mode derivation unit 40 1 is provided with a prediction motion vector candidate list mvpListLX. The prediction motion vector candidate list mvpListLX has a list structure, and a prediction motion vector index indicating the location inside the prediction motion vector candidate list and a prediction motion vector candidate corresponding to the index are used as elements A storage area for storing is provided. The numbers of the prediction motion vector indexes start from 0, and the prediction motion vector candidates are stored in the storage area of the prediction motion vector candidate list mvpListLX. In the present embodiment, it is assumed that the prediction motion vector candidate list mvpListLX can register at least two prediction motion vector candidates (inter prediction information). Further, a variable numCurrMvpCand indicating the number of prediction motion vector candidates registered in the prediction motion vector candidate list mvpListLX is set to 0. In the present embodiment, the prediction motion vector candidate list mvpListLX is at least It is assumed that two prediction motion vector candidates (inter prediction information) can be registered. Furthermore, a variable numCurrMvpCand indicating the number of prediction motion vector candidates registered in the prediction motion vector candidate list mvpListLX is set to 0. candidates is set to 0.
[0090] The spatial prediction motion vector candidate derivation units 321 and 421 derive candidates for the prediction motion vector from the block adjacent to the left side. In this process, the block adjacent to the left side (A0 or also A flag availableFlagLXA indicating whether the predicted motion vector candidate of A1) can be used, and derive the motion vector mvLXA and the reference index refIdxA, and add mvLXA to the predicted motion vector candidate list mvpListLX (step S301 in FIG. 20). When it is L0, X is 0 , and when it is L1, X is 1 (the same applies hereinafter). Subsequently, the spatial predicted motion vector candidate derivation unit 32 1 and 421 derive candidates for the predicted motion vector from the blocks (B0, B1, or B2) adjacent above. In this process, the predicted motion vector candidate of the block adjacent above is used. A flag availableFlagLXB indicating whether it can be used, and the motion vector mvLXB and the reference index refIdxB are derived. If mvLXA and mvLXB are not equal, mvLXB is added to the predicted motion vector candidate list mvpListLX (step S302 in FIG. 20). The processes of step S30 1 and S302 in FIG. 20 are common except that the positions and numbers of the adjacent blocks to be referred to are different. A flag availableFlagLXN indicating whether the predicted motion vector candidate of the coded block can be used, and the motion vector mvLXN and the reference index refIdxN (N is A or B, the same hereinafter) are derived.
[0091] Subsequently, the temporal predicted motion vector candidate derivation units 322 and 422 derive candidates for the predicted motion vector from the coded blocks in a picture with a different time from the current picture to be processed frame. In this process, a flag availableFlagLXCol indicating whether the predicted motion vector candidate of the coded block in a picture with a different time can be used, and the motion vector mvLXCol , derive the reference index refIdxCol and the reference list listCol, and predict the motion vector mvLXCol and add it to the motion vector candidate list mvpListLX for the LX candidate (step S303 in FIG. 20). The derivation processing procedure of this step S30 will be described in detail later.
[0092] Note that the processing of the temporal prediction motion vector candidate derivation units 322 and 422 can be omitted for each sequence (SPS), picture (PPS), or slice unit. It is assumed that the processing of the temporal prediction motion vector candidate derivation units 322 and 422 can be omitted.
[0093] Subsequently, the history prediction motion vector candidate derivation units 323 and 423 add the history prediction motion vector candidates registered in the history prediction motion vector candidate list HmvpCandList to the motion vector candidate list mvpListLX for the LX candidate. (Step S304 in FIG. 20). The registration processing procedure of this step S304 will be described in detail later using the flowchart of FIG. 41. Subsequently, the history prediction motion vector candidate derivation units 323 and 423 add the history prediction motion vector candidates registered in the history prediction motion vector candidate list HmvpCandList to the motion vector candidate list mvpListLX for the LX candidate. (Step S304 in FIG. 20). The registration processing procedure of this step S304 will be described in detail later using the flowchart of FIG. 41.
[0094] Subsequently, the motion vector candidate supplementing units 325 and 425 add a motion vector of a predetermined value such as (0, 0) until the motion vector candidate list mv pListLX is satisfied (S3 05 in FIG. 20).
[0095] <Normal merge mode derivation unit (normal merge)> The normal merge mode derivation unit 302 in FIG. 18 includes a spatial merge candidate derivation unit 341, a temporal merge candidate derivation unit 342, an average merge candidate derivation unit 344, a history merge candidate derivation unit 345, a merge candidate supplementing unit 346, and a merge candidate selection unit 347.
[0096] The normal merge mode derivation unit 402 in FIG. 24 includes a spatial merge candidate derivation unit 441, a temporal merge candidate derivation unit 442, an average merge candidate derivation unit 444, a history merge candidate derivation unit 445, a merge It includes a candidate supplement part 446 and a merge candidate selection part 447.
[0097] FIG. 21 is a flowchart for explaining the procedure of a normal merge mode derivation process having a common function in the normal merge mode derivation part 302 of the moving image encoding apparatus and the normal merge mode derivation part 402 of the moving image decoding apparatus according to an embodiment of the present invention. and the normal merge mode derivation process.
[0098] The following describes various processes in sequence. In the following description, unless otherwise specified, the case where the slice type slice_type is a B slice will be described, but it is also applicable to the case of a P slice. However, when the slice type slice_type is a P slice, there is only L0 prediction (Pred_L0) as the inter prediction mode, and there is no L1 prediction (Pred_L1) or bi-prediction (Pred_BI). Therefore, the processes related to L1 can be omitted.
[0099] The normal merge mode derivation part 302 and the normal merge mode derivation part 402 are provided with a merge candidate list mergeCandList. The merge candidate list mergeCandList has a list structure, and a storage area is provided that stores a merge index indicating the location inside the merge candidate list and a merge candidate corresponding to the index as elements. The number of the merge index starts from 0, and the merge candidate is stored in the storage area of the merge candidate list mergeCandList. In the following processes, the merge candidate at the merge index i registered in the merge candidate list mergeCandList will be represented by mergeCandList[i]. In the present embodiment, the merge candidate list mergeCandList registers at least six merge candidates (inter prediction information). It shall be possible. Further, set 0 to the variable numCurrMergeCand indicating the number of merge candidates registered in the merge candidate list mergeCandList.
[0100] In the spatial merge candidate derivation unit 341 and the spatial merge candidate derivation unit 441, from the encoding information stored in the encoding information storage memory 111 of the moving image encoding device or the encoding information storage memory 205 of the moving image decoding device, derive the spatial merge candidates A and B from the blocks adjacent to the left and upper sides of the processing target block, and register the derived spatial merge candidates in the merge candidate list mergeC andList (step S401 in FIG. 21). Here, define N which indicates either the spatial merge candidate A, B or the temporal merge candidate Col. A flag availableFlagN indicating whether the inter prediction information of block N can be used as the spatial merge candidate N, the reference index refIdxL0N of L0 and the reference index refIdxL1N of L1 of the spatial merge candidate N, the L0 prediction flag predFlagL0N indicating whether L0 prediction is performed and the L1 prediction flag predFlagL1N indicating whether L1 prediction is performed, the motion vector mvL0N of L0, and the motion vector mvL1 N of are derived. However, in the present embodiment, since the merge candidates are derived without referring to other encoded blocks included in the block including the encoded block to be processed, the spatial merge candidates included in the block including the encoded block to be processed are not derived.
[0101] Subsequently, in the temporal merge candidate derivation unit 342 and the temporal merge candidate derivation unit 442, derive the temporal merge candidates from different pictures, and use the derived temporal merge candidates as the merge candidates Register it in the merge candidate list mergeCandList (step S402 in FIG. 21). Whether the time merge candidate can be used is indicated by the available flag availableFlagCol, whether L0 prediction of the time merge candidate is performed or not is indicated by the L0 prediction flag predFlagL0Col and whether L1 prediction is performed or not is indicated by the L1 prediction flag predFlagL1Col, and the motion vector mvL0Col of L0 and the motion vector mvL1Col of L1 are derived. For the detailed processing procedure of step S402, refer to FIG. 56 later for details explain.
[0102] Note that the processing of the time merge candidate derivation units 342 and 442 in units of sequence (SPS), picture (PPS), or slice can be omitted. It is assumed that it can be omitted.
[0103] Subsequently, in the history merge candidate derivation units 345 and 445, the history prediction motion vector candidates registered in the history prediction motion vector candidate list HmvpCandList are merged and added to the candidate list mergeCandList (step S403 in FIG. 21). For the detailed processing procedure of step S40 3, it will be described in detail later using the flowchart of FIG. 62.
[0104] Subsequently, in the average merge candidate derivation units 344 and 444, average merge candidates are derived from the merge candidate list mergeCandList, and the derived average merge candidates are merged and registered in the candidate list mergeCandList (step S404 in FIG. 21). For the detailed processing procedure of step S4 04, it will be described in detail later using the flowchart of FIG. 41. .
[0105] Next, the merge candidate supplementation unit 346 and the merge candidate supplementation unit 446 create a merge candidate list The number of merge candidates registered in mergeCandList, numCurrMergeCand, is less than the maximum number of merge candidates, M If it is smaller than axNumMergeCand, the merge candidate list mergeCandList The number of merge candidates numCurrMergeCand is the maximum number of merge candidates MaxNumMergeCand. The merge candidates are derived and registered in the merge candidate list mergeCandList (step S4 in FIG. 21). 05). In the P slice, the maximum number of merge candidates is limited to MaxNumMergeCand. The prediction mode with the motion vector value (0,0) in the index is L0 prediction (Pred_L0) Add merge candidates for B slices. For B slices, the motion vectors are ( 0,0) is added as a merge candidate whose prediction mode is bi-predictive (Pred_BI).
[0106] Next, the merging candidate selection unit 347 and the merging candidate selection unit 447 select the merging candidate list Select a merge candidate from the merge candidates registered in mergeCandList. The merge candidate selection unit 347 selects merge candidates by calculating the code amount and the distortion amount. A merge index indicating the selected merge candidate is generated. On the other hand, the decoding side merge candidate selection unit 447 supplies the decoded Based on the merge index, a merge candidate is selected, and the selected merge candidate is motion-compensated. The result is supplied to the compensation prediction unit 406.
[0107] The normal merge mode derivation unit 302 and the normal merge mode derivation unit 402 are If the size of the block (the product of the width and the height) is less than 32, merge candidates are derived in the parent block of the coded block. And in all child blocks, the merge candidates derived in the parent block are used, provided that the size of the parent block is 32 or more and it fits within the screen. And in all child blocks, the merge candidates derived in the parent block are used, provided that the size of the parent block is 32 or more and it fits within the screen. And in all child blocks, the merge candidates derived in the parent block are used, provided that the size of the parent block is 32 or more and it fits within the screen. And in all child blocks, the merge candidates derived in the parent block are used, provided that the size of the parent block is 32 or more and it fits within the screen.
[0108] The derivation of the sub - block predictive motion vector mode will be described.
[0109] FIG. 26 is a block diagram of the sub - block predictive motion vector mode derivation unit 303 in the encoding apparatus of the present embodiment. FIG. 26 is a block diagram of the sub - block predictive motion vector mode derivation unit 303 in the encoding apparatus of the present embodiment.
[0110] First, in the affine inheritance predictive motion vector candidate derivation unit 361, affine inheritance predictive motion vector candidates are derived. Details of the derivation of affine inheritance predictive motion vector candidates will be described later. First, in the affine inheritance predictive motion vector candidate derivation unit 361, affine inheritance predictive motion vector candidates are derived. Details of the derivation of affine inheritance predictive motion vector candidates will be described later. First, in the affine inheritance predictive motion vector candidate derivation unit 361, affine inheritance predictive motion vector candidates are derived. Details of the derivation of affine inheritance predictive motion vector candidates will be described later.
[0111] Subsequently, in the affine construction predictive motion vector candidate derivation unit 362, affine construction predictive motion vector candidates are derived. Details of the derivation of affine construction predictive motion vector candidates will be described later. Subsequently, in the affine construction predictive motion vector candidate derivation unit 362, affine construction predictive motion vector candidates are derived. Details of the derivation of affine construction predictive motion vector candidates will be described later. Subsequently, in the affine construction predictive motion vector candidate derivation unit 362, affine construction predictive motion vector candidates are derived. Details of the derivation of affine construction predictive motion vector candidates will be described later.
[0112] Subsequently, in the affine identical predictive motion vector candidate derivation unit 363, affine identical predictive motion vector candidates are derived. Details of the derivation of affine identical predictive motion vector candidates will be described later. Subsequently, in the affine identical predictive motion vector candidate derivation unit 363, affine identical predictive motion vector candidates are derived. Details of the derivation of affine identical predictive motion vector candidates will be described later. Subsequently, in the affine identical predictive motion vector candidate derivation unit 363, affine identical predictive motion vector candidates are derived. Details of the derivation of affine identical predictive motion vector candidates will be described later.
[0113] The sub - block motion vector detection unit 366 detects a sub - block motion vector suitable for the sub - block predictive motion vector mode, and supplies the detected vector to the sub - block predictive motion vector candidate selection unit 367 and the difference calculation unit 368. The sub - block motion vector detection unit 366 detects a sub - block motion vector suitable for the sub - block predictive motion vector mode, and supplies the detected vector to the sub - block predictive motion vector candidate selection unit 367 and the difference calculation unit 368. The sub - block motion vector detection unit 366 detects a sub - block motion vector suitable for the sub - block predictive motion vector mode, and supplies the detected vector to the sub - block predictive motion vector candidate selection unit 367 and the difference calculation unit 368.
[0114] The sub-block predicted motion vector candidate selection unit 367 selects the affine inheritance predicted motion vector candidate Affine construction predicted motion vector candidate derivation unit 361, an affine construction predicted motion vector candidate derivation unit 362, an affine same predicted motion vector candidate derivation unit 363, an affine construction predicted motion vector candidate derivation unit 364, an affine same predicted motion vector candidate derivation unit 365, an affine construction predicted motion vector candidate derivation unit 366, an affine same predicted motion vector candidate derivation unit 367, an Among the sub-block predicted motion vector candidates derived in the motion vector candidate derivation unit 363, Based on the motion vector supplied from the sub-block motion vector detection unit 366, A sub-block predicted motion vector candidate is selected, and the selected sub-block predicted motion vector Information about the candidates is supplied to the inter prediction mode determination unit 305 and the difference calculation unit 368 .
[0115] The difference calculation unit 368 calculates the motion vector supplied from the sub-block motion vector detection unit 366. The sub-block predicted motion vector candidate selection unit 367 selects the sub-block predicted motion vector candidate from the motion vector. The inter prediction mode determination unit 102 determines the difference prediction motion vector obtained by subtracting the inter prediction motion vector from the inter prediction mode prediction motion vector. Supply to 305.
[0116] FIG. 27 shows a sub-block prediction motion vector mode derivation method in the decoding device according to the present embodiment. 4 is a block diagram of unit 403.
[0117] First, the affine inheritance predicted motion vector candidate derivation unit 461 derives an affine inheritance predicted motion vector candidate. The process of the affine inheritance predicted motion vector candidate derivation unit 461 is as follows: Processing of the affine inheritance predicted motion vector candidate derivation unit 361 in the encoding device according to the embodiment is equivalent to.
[0118] Next, the affine construction prediction motion vector candidate derivation unit 462 performs affine construction prediction The affine construction prediction motion vector candidate derivation unit 462 performs the following process: The processing of the affine construction prediction motion vector candidate derivation unit 362 in the encoding device of this embodiment is the same. It is the same.
[0119] Subsequently, the affine identical prediction motion vector candidate derivation unit 463 derives affine identical prediction motion vector candidates. The processing of the affine identical prediction motion vector candidate derivation unit 463 is the same as that of the affine identical prediction motion vector candidate derivation unit 363 in the encoding device of this embodiment. It is the same. The processing of the affine identical prediction motion vector candidate derivation unit 363 in the encoding device of this embodiment is the same. It is the same.
[0120] The sub-block prediction motion vector candidate selection unit 466 selects sub-block prediction motion vector candidates from among the sub-block prediction motion vector candidates derived by the affine inheritance prediction motion vector candidate derivation unit 461, the affine construction prediction motion vector candidate derivation unit 462, and the affine identical prediction motion vector candidate derivation unit 463, based on the prediction motion vector index transmitted from the encoding device and decoded. The sub-block prediction motion vector candidates selected are supplied to the motion compensation prediction unit 406 and the addition operation unit 467 together with information regarding the selected sub-block prediction motion vector candidates. It selects sub-block prediction motion vector candidates from among the sub-block prediction motion vector candidates derived by the affine inheritance prediction motion vector candidate derivation unit 461, the affine construction prediction motion vector candidate derivation unit 462, and the affine identical prediction motion vector candidate derivation unit 463, based on the prediction motion vector index transmitted from the encoding device and decoded. The sub-block prediction motion vector candidates selected are supplied to the motion compensation prediction unit 406 and the addition operation unit 467 together with information regarding the selected sub-block prediction motion vector candidates. from among the sub-block prediction motion vector candidates derived by the affine inheritance prediction motion vector candidate derivation unit 461, the affine construction prediction motion vector candidate derivation unit 462, and the affine identical prediction motion vector candidate derivation unit 463, based on the prediction motion vector index transmitted from the encoding device and decoded. The sub-block prediction motion vector candidates selected are supplied to the motion compensation prediction unit 406 and the addition operation unit 467 together with information regarding the selected sub-block prediction motion vector candidates. It selects sub-block prediction motion vector candidates from among the sub-block prediction motion vector candidates derived by the affine inheritance prediction motion vector candidate derivation unit 461, the affine construction prediction motion vector candidate derivation unit 462, and the affine identical prediction motion vector candidate derivation unit 463, based on the prediction motion vector index transmitted from the encoding device and decoded.
[0121] The addition operation unit 467 adds the differential motion vector transmitted from the encoding device and decoded to the sub-block prediction motion vector selected by the sub-block prediction motion vector candidate selection unit 466, and supplies the generated motion vector to the motion compensation prediction unit 406. The addition operation unit 467 adds the differential motion vector transmitted from the encoding device and decoded to the sub-block prediction motion vector selected by the sub-block prediction motion vector candidate selection unit 466, and supplies the generated motion vector to the motion compensation prediction unit 406. The addition operation unit 467 adds the differential motion vector transmitted from the encoding device and decoded to the sub-block prediction motion vector selected by the sub-block prediction motion vector candidate selection unit 466, and supplies the generated motion vector to the motion compensation prediction unit 406.
[0122] <Affine inheritance prediction motion vector candidate derivation> The affine inheritance prediction motion vector candidate derivation unit 361 will be described. The affine inheritance prediction motion vector candidate derivation unit 461 is the same as the affine inheritance prediction motion vector candidate derivation unit 361. The affine inheritance prediction motion vector candidate derivation unit 461 is the same as the affine inheritance prediction motion vector candidate derivation unit 361. It is the same.
[0123] The affine inheritance prediction motion vector candidate inherits the motion vector information of the control points.
[0124] FIG. 30 is a diagram for explaining the derivation of the affine inheritance prediction motion vector candidate.
[0125] The affine inheritance prediction motion vector candidate is obtained by searching for the motion vectors of the control points possessed by the spatially adjacent encoded / decoded blocks.
[0126] Specifically, from the blocks (A0, A1) adjacent to the left side of the block to be processed and the blocks (B0, B1, B2) adjacent to the upper side of the block to be processed, search for at most one affine mode respectively, and use it as the affine inheritance prediction motion vector.
[0127] FIG. 34 is a flowchart for deriving the affine inheritance prediction motion vector candidate.
[0128] First, regard the blocks (A0, A1) adjacent to the left side of the block to be processed as the left group (step S3101), and determine whether the block containing A0 is a block using affine transformation motion compensation (affine mode) (step S3102). If A0 is in the affine mode (step S3102: YES), obtain the affine mode used by A0 (step S3103), and move to the processing of the adjacent block above. If A0 is not in the affine mode (step S3102: NO), set the target of deriving the affine inheritance prediction motion vector candidate to A0->A1, and try to obtain the affine mode from the block containing A1.
[0129] Subsequently, regard the blocks (B0, B1, B2) adjacent to the upper side of the block to be processed as the upper group and determine whether the block containing B0 is in the affine mode. Judge (step S3105). If B0 is in the affine mode (step S310 5: YES), obtain the affine mode used by B0 (step S3106), and end the process. If B0 is not in the affine mode (step S3105: NO), set the target for deriving the affine inheritance prediction motion vector candidate as B0->B1, and attempt to obtain the affine mode from the block including B1. Further, if B1 is not in the affine mode (step S310 5: NO), set the target for deriving the affine inheritance prediction motion vector candidate as B1->B2, and attempt to obtain the affine mode from the block including B2.
[0130] In this way, by dividing into groups of left blocks and upper blocks, for the left blocks, search for the affine mode in the order of blocks from bottom left to top left, and for the upper blocks, search for the affine mode in the order of blocks from top right to top left, so that two affine modes that are as different as possible can be obtained, and an affine prediction motion vector candidate with a smaller difference motion vector than any of the affine prediction motion vectors can be derived.
[0131] <Derivation of Affine Construction Prediction Motion Vector Candidate> The affine construction prediction motion vector candidate derivation unit 362 will be described. The affine construction prediction motion vector candidate derivation unit 462 is the same as the affine construction prediction motion vector candidate derivation unit 36 2.
[0132] The affine construction prediction motion vector candidate constructs the motion vector information of the control point from the motion information of spatially adjacent blocks.
[0133] FIG. 31 is a diagram for explaining the derivation of an affine construction prediction motion vector candidate.
[0134] The affine construction prediction motion vector candidate is obtained by constructing a new affine mode by combining the motion vectors of spatially adjacent encoded / decoded blocks. Specifically, the motion vectors of the control points CP0, CP1, and CP2 are derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed or
[0135] from the blocks (B1, B0) adjacent to the upper right side of the block to be processed, and from the blocks (A1, A0) adjacent to the lower left side of the block to be processed. Specifically, the motion vector of the upper left control point CP0 is derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed, the motion vector of the upper right control point CP1 is derived from the blocks (B1, B0) adjacent to the upper right side of the block to be processed, and the motion vector of the lower left control point CP2 is derived from the blocks (A1, A0) adjacent to the lower left side of the block to be processed. Specifically, the motion vector of the upper left control point CP0 is derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed, the motion vector of the upper right control point CP1 is derived from the blocks (B1, B0) adjacent to the upper right side of the block to be processed, and the motion vector of the lower left control point CP2 is derived from the blocks (A1, A0) adjacent to the lower left side of the block to be processed. Specifically, the motion vector of the upper left control point CP0 is derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed, the motion vector of the upper right control point CP1 is derived from the blocks (B1, B0) adjacent to the upper right side of the block to be processed, and the motion vector of the lower left control point CP2 is derived from the blocks (A1, A0) adjacent to the lower left side of the block to be processed. Specifically, the motion vector of the upper left control point CP0 is derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed, the motion vector of the upper right control point CP1 is derived from the blocks (B1, B0) adjacent to the upper right side of the block to be processed, and the motion vector of the lower left control point CP2 is derived from the blocks (A1, A0) adjacent to the lower left side of the block to be processed.
[0136] FIG. 35 is a flowchart for the derivation of an affine construction prediction motion vector candidate.
[0137] First, the upper left control point CP0, the upper right control point CP1, and the lower left control point CP2 are derived (step S S3201). The upper left control point CP0 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the B2, B3, A2 reference blocks. The upper right control point CP1 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the B1, B0 reference blocks. The lower left control point CP2 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the A1, A0 reference blocks. The upper right control point CP1 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the B1, B0 reference blocks. The lower left control point CP2 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the A1, A0 reference blocks. The lower left control point CP2 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the A1, A0 reference blocks. reference blocks. The lower left control point CP2 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the A1, A0 reference blocks. reference blocks. The lower left control point CP2 is calculated by searching for a reference block having the same reference image as the block to be processed in the order of priority of the A1, A0 reference blocks.
[0138] When selecting the three control point mode as the affine construction prediction motion vector (step S S3202: YES), it is determined whether all three control points (CP0, CP1, CP2) have been derived. Determine whether or not (step S3203). When all three control points (CP0, CP1, CP2) have been derived (step S3203: YES), use the affine model using the three control points (CP0, CP1, CP2) as the affine construction prediction motion vector (step S32 04). When the two control point mode is selected instead of the three control point mode (step S3 202: NO), determine whether or not all two control points (CP0, CP1) have been derived (step S3205). When all two control points (CP0, CP1) have been derived (step S3205: YES), use the affine model using the two control points (CP0, CP1) as the affine construction prediction motion vector (step S3206). (step S3205). When all two control points (CP0, CP1) have been derived (step S3205: YES), use the affine model using the two control points (CP0, CP1) as the affine construction prediction motion vector (step S3206).
[0139] <Derivation of Affine Identical Prediction Motion Vector Candidate> The affine identical prediction motion vector candidate derivation unit 363 will be described. The affine identical prediction motion vector candidate derivation unit 463 is the same as the affine identical prediction motion vector candidate derivation unit 36 3.
[0140] The affine identical prediction motion vector candidate can be obtained by deriving the same motion vector at each control point .
[0141] Specifically, similar to the affine construction prediction motion vector candidate derivation units 362 and 462, each control point information is derived, and all control points are identically set to any of CP0 to CP2 to obtain it. It can also be obtained by setting the time motion vector derived in the same manner as the normal prediction motion vector mode for all control points.
[0142] <Derivation of Sub-Block Merge Mode> The derivation of the sub-block merge mode will be described.
[0143] FIG. 28 is a block diagram of a sub-block merge mode derivation unit 304 in the encoding apparatus of the present embodiment. The sub-block merge mode derivation unit 304 includes a sub-block merge candidate list subblockMergeCandList. This has the same list structure as the merge candidate list mergeCandList in the normal merge mode derivation unit 302, and has a merge index indicating the location inside the sub-block merge candidate list and a storage area for storing sub-block merge candidates corresponding to the index as elements. In the present embodiment it is assumed that the sub-block merge candidate list subblockMergeCandList can register at least five merge candidates (inter prediction information). However, each merge candidate further has motion vector information in units of sub-blocks or has motion vector information of control points. First, in the sub-block temporal merge candidate derivation unit 381, sub-block temporal merge candidates are derived. Details of the derivation of sub-block temporal merge candidates will be described later. Subsequently, in the affine inheritance merge candidate derivation unit 382, affine inheritance merge candidates are derived. Details of the derivation of affine inheritance merge candidates will be described later. Subsequently, in the affine construction merge candidate derivation unit 383, affine construction merge candidates are derived. Details of the derivation of affine construction merge candidates will be described later. Subsequently, in the affine fixed merge candidate derivation unit 385, affine fixed merge candidates are derived.
[0144] Details of the derivation of affine fixed merge candidates will be described later.
[0145]
[0146]
[0147] Output. Details of affine-fixed merge candidate derivation will be described later.
[0148] The sub-block merge candidate selection unit 386 selects sub-block merge candidates from among the sub-block merge candidates derived by the sub-block time merge candidate derivation unit 381, the affine inheritance merge candidate derivation unit 382, the affine construction merge candidate derivation unit 383, and the affine fixed merge candidate derivation unit 385, selects sub-block merge candidates, and supplies information regarding the selected sub-block merge candidates to the inter- prediction mode determination unit 305.
[0149] FIG. 29 is a block diagram of a sub-block merge mode derivation unit 404 in the decoder according to the present embodiment. The sub-block merge mode derivation unit 404 includes a sub-block merge candidate list subblockMergeCandList. This is the same as that of the sub-block merge mode derivation unit 304.
[0150] First, the sub-block time merge candidate derivation unit 481 derives sub-block time merge candidates. The processing of the sub-block time merge candidate derivation unit 481 is the same as the processing of the sub-block time merge candidate derivation unit 381.
[0151] Subsequently, the affine inheritance merge candidate derivation unit 482 derives affine inheritance merge candidates. The processing of the affine inheritance merge candidate derivation unit 482 is the same as the processing of the affine inheritance merge candidate derivation unit 3 82.
[0152] Subsequently, the affine construction merge candidate derivation unit 483 derives affine construction merge candidates. The processing of the affine construction merge candidate derivation unit 483 is the same as the processing of the affine construction merge candidate derivation unit 3 83.
[0153] Subsequently, in the affine fixed merge candidate derivation unit 485, affine fixed merge candidates are derived. The processing of the affine fixed merge candidate derivation unit 485 is the same as that of the affine fixed merge candidate derivation unit 485. The sub-block merge candidate selection unit 486 selects sub-block merge candidates from among the sub-block merge candidates derived in the sub-block time merge candidate derivation unit 481, the affine inheritance merge candidate derivation unit 482, the affine construction merge candidate derivation unit 483, and the affine fixed merge candidate derivation unit 485, based on the index transmitted from the encoder and decoded, and supplies information regarding the selected sub-block merge candidates to the motion compensation prediction unit 406.
[0154] When the size (product of width and height) of a certain coded block is less than 32, sub-block merge candidates are derived in the parent block of that coded block. And in all child blocks, the sub-block merge candidates derived in the parent block are used. However, this is only the case when the size of the parent block is 32 or more and it fits within the screen.
[0155]
[0156] <Sub-block time merge candidate derivation> The operation of the sub-block time merge candidate derivation unit 381 will be described later.
[0157] <Affine inheritance merge candidate derivation> The affine inheritance merge candidate derivation unit 382 will be described. The affine inheritance merge candidate derivation unit 482 is the same as the affine inheritance merge candidate derivation unit 382.
[0158] The affine inheritance merge candidate inherits the control point affine model of spatially adjacent blocks from the affine models of the spatially adjacent blocks.
[0159] Figure 32 is a diagram for explaining the derivation of an affine inheritance merge candidate. The derivation of the affine merge inheritance mode candidate is obtained by searching for the motion vectors of the control points of the spatially adjacent encoded / decoded blocks, similar to the derivation of the affine inheritance prediction motion vector.
[0160] Specifically, from the blocks (A0, A1) adjacent to the left side of the block to be processed and the blocks (B0, B1, B2) adjacent to the upper side of the block to be processed, up to one affine mode is searched respectively and used for the affine merge mode.
[0161] Figure 36 is a flowchart for the derivation of an affine inheritance merge candidate.
[0162] First, the blocks (A0, A1) adjacent to the left side of the block to be processed are used as the left group (step S3301), and it is determined whether the block containing A0 is in the affine mode (step S3302). If A0 is in the affine mode (step S3102: YES), the affine model used by A0 is obtained (step S3303), and the process moves to the blocks adjacent to the upper side. If A0 is not in the affine mode (step S3302: NO), the target of the affine inheritance merge candidate derivation is set to A0->A1, and an attempt is made to obtain the affine mode from the blocks containing A1.
[0163] Subsequently, the blocks (B0, B1, B2) adjacent to the upper side of the block to be processed are used as the upper group - As a step (step S3304), it is determined whether the block including B0 is in the affine mode (step S3305). If B0 is in the affine mode (step S330 5: YES), the affine model used by B0 is obtained (step S3306), and the process ends. If B0 is not in the affine mode (step S3305: NO), the target for affine inheritance merge candidate derivation is set to B0->B1, and an attempt is made to obtain the affine mode from the block including B1. Further, if B1 is not in the affine mode (step S3305: NO) , the target for affine inheritance merge candidate derivation is set to B1->B2, and an attempt is made to obtain the affine mode from the block including B2.
[0164] <Affine Construction Merge Candidate Derivation> The affine construction merge candidate derivation unit 383 will be described. The affine construction merge candidate derivation unit 483 is the same as the affine construction merge candidate derivation unit 383.
[0165] FIG. 33 is a diagram for explaining affine construction merge candidate derivation. The affine construction merge candidate constructs an affine model of control points from the motion information and time-coded blocks of spatially adjacent blocks.
[0166] Specifically, the motion vector of the upper left control point CP0 is derived from the blocks (B2, B3, A2) adjacent to the upper left side of the block to be processed, and the motion vector of the upper right control point CP1 is derived from the blocks (B1, B0) adjacent to the upper right side of the block to be processed. The motion vector of the lower left control point CP2 is derived from the blocks (A1, A0) adjacent to the lower left side of the block to be processed, and the lower right control point is derived from the coded block (T0) adjacent to the lower right side of the block to be processed. Derive the motion vector of CP3.
[0167] Figure 37 is a flowchart for deriving affine construction merge candidates.
[0168] First, derive the upper left control point CP0, upper right control point CP1, lower left control point CP2, and lower right control point CP3 (step S3401). The upper left control point CP0 is calculated by searching for the block with motion information in the order of priority of blocks B2, B3, and A2. The upper right control point CP1 is calculated by searching for the block with motion information in the order of priority of blocks B1 and B0 . The lower left control point CP2 is calculated by searching for the block with motion information in the order of priority of blocks A1 and A0 . The lower right control point CP3 is calculated by searching the motion information of the temporal block .
[0169] Subsequently, determine whether an affine model with three control points can be constructed using the derived CP0, CP1, and CP2 (step S3402). If it can be constructed (step S3402: YES), use the three-control-point affine model with CP0, CP1, and CP2 as the affine merge candidate (step S3403).
[0170] Subsequently, determine whether an affine model with three control points can be constructed using the derived CP0, CP1, and CP3 (step S3404). If it can be constructed (step S3404: YES), use the three-control-point affine model with CP0, CP1, and CP3 as the affine merge candidate (step S3405).
[0171] Subsequently, determine whether an affine model with three control points can be constructed using the derived CP0, CP2, and CP3 It is determined whether construction is possible (step S3406), and if construction is possible (step S3406:YES), 3-control-point affine model using CP0, CP2, and CP3 The selected object is determined as a merge candidate (step S3407).
[0172] Next, an affine model is created using the three control points CP1, CP2, and CP3. It is determined whether construction is possible (step S3408), and if construction is possible (step S3408:YES), 3-control-point affine model using CP1, CP2, and CP3 The selected object is determined as a merge candidate (step S3409).
[0173] Next, an affine model can be constructed using the derived CP0 and CP1. If it is possible to construct the system (step S3410), 0: YES), the two-control-point affine model by CP0 and CP1 is considered as an affine merge candidate. (step S3411).
[0174] Next, an affine model can be constructed using the derived CP0 and CP2 control points. If it is possible to construct the system (step S341), 2: YES), the two-control-point affine model by CP0 and CP2 is selected as an affine merging candidate. (step S3413).
[0175] Here, whether or not an affine model can be constructed depends on whether or not all the control points are referenced. The condition is that the reference images are the same (affine transformation is possible). Three-control-point affine model using CP2, two-control-point affine model using CP0 and CP1 Affine models other than Ru, for the three-control-point affine model, are converted to a three-control-point affine model by CP0, CP1, CP 2, and for the two-control-point affine model, they are converted to a two-control-point affine model by CP0, C P1.
[0176] <Derivation of Affine Fixed Merge Candidates> The affine fixed merge candidate derivation unit 385 will be described. The affine fixed merge candidate derivation unit 485 is the same as the affine fixed merge candidate derivation unit 385.
[0177] The affine fixed merge candidates fix the movement information of the control points with the fixed movement information.
[0178] Specifically, the movement vector of each control point is fixed to (0, 0).
[0179] <Derivation of Temporal Prediction Movement Vectors> Prior to the description of the temporal prediction movement vectors, the temporal context of the pictures will be described with reference to FIG. 49 FIG. 49(a) shows the relationship between the coding block to be processed and the coded pictures that are temporally different from the picture to be processed. In the coding of the picture to be processed a specific coded picture that is referred to is defined as ColPic. ColPic is specified by the syntax .
[0180] Also, FIG. 49(b) shows the coded coding blocks that are at the same position as and in the vicinity of the coding block to be processed in ColPic. However, the coding blocks of T0 and T1 shown in FIG. 49(b) are schematic, and the actual positions and sizes are not limited to this . Now, for the coding block to be processed, let the position be (xCb, yCb), the width be cbWidt h, and the height be cbHeight. And xColBr = xCb + cbWidth yColBr = yCb + cbHeight Calculate. The encoded block on ColPic including the position ((xColBr >> 3) << 3, (yColBr >> 3) << 3) becomes T0. Also, xColCtr = xCb + (cbWidth >> 1) yColCtr = yCb + (cbHeight >> 1) Calculate. The encoded block on ColPic including the position ((xColCtr >> 3) << 3, (yColCtr >> 3) << 3) becomes T1.
[0181] The above description of the temporal context of the picture is for encoding, but it is the same during decoding. That is, during decoding, replace encoding in the above description with decoding and explain in the same way. It is explained.
[0182] The operation of the temporal prediction motion vector candidate derivation unit 322 in the normal prediction motion vector mode derivation unit 301 in FIG. 17 will be described with reference to FIG. 50.
[0183] First, derive ColPic (step S4201). The derivation of ColPic will be described with reference to FIG. 51. It will be described.
[0184] When the slice type slice_type is a B slice and the flag collocated_from_l0_flag is 0 (step S4211: YES, step S4212: YES), the picture ColPic at a different time becomes the picture RefPicList1[0] whose reference index in the reference list L1 is 0 ( step S4213). Otherwise, that is, when the slice type slice_type is a B slice If the flag collocated_from_l0_flag described above is 1 in S (step S4211: YES, step S4212: NO), or if the slice type slice_type is a P slice (step S4211: NO, step S4214: YES), the picture ColPic at a different time becomes the picture RefPicList0[0] with the reference index of the reference list L0 being 0 (step S4215). If the slice type is not a P slice (step S4214: NO), the process ends.
[0185] Referring to FIG. 50 again. After deriving ColPic, derive the coding block colCb and obtain the coding information (step S4202). This process will be described with reference to FIG. 52.
[0186] First, within the picture ColPic at a different time, the coding block including the lower right position at the same position as the coding block to be coded is set as the coding block colCb at a different time (step S 4221). An example of this coding block is shown as the coding block T0 in FIG. 49.
[0187] Next, obtain the coding information of the coding block colCb at a different time (step S422 2). If the PredMode of the coding block colCb at a different time is not available, or if the prediction mode PredMode of the coding block colCb at a different time is intra prediction (MODE_INTRA) (step S4223: NO, step S4224: YES), the coding block including the lower right position at the center at the same position as the coding block to be processed within the picture ColPic at a different time is different Set it as the encoding block colCb of time (step S4225). An example of this encoding block is shown in the encoding block T1 in FIG. 49.
[0188] Refer to FIG. 50 again. Next, for each reference list, derive the inter-prediction information (steps S4203, S4204). Here, for the encoding block colCb, derive the motion vector mvLXCol and the flag availableFlagLXCo l indicating whether the encoding information is valid for each reference list. LX indicates the reference list. In the derivation of reference list 0, LX is L0, and in the derivation of reference list 1, LX is L1. The derivation of the inter-prediction information will be described with reference to FIG. 53.
[0189] If different time encoding blocks colCb are not available (step S4231: NO) , or when the prediction mode PredMode is intra prediction (MODE_INTRA) (step S4232 : NO), set both the flag availableFlagLXCol and the flag predFlagLXCol to 0 (step S 4233), set the motion vector mvLXCol to (0, 0) (step S4234), and end the process.
[0190] If the encoding block colCb is available (step S4231: YES) and the prediction mode PredMod e is not intra prediction (MODE_INTRA) (step S4232: YES), calculate mvCol, refIdxCol, and availableFlagCol in the following order.
[0191] The flag PredFlagL0[xPCo indicating whether the L0 prediction of the encoding block colCb is used When [yPCol] is 0 (step S4235: YES), since the prediction mode of the encoding block colCb is Pred_L1, the motion vector mvCol is set to the same value as the motion vector MvL1[xPCol][yPCol] of L1 of the encoding block colCb (step S4236), and the reference index refIdxCol is set to the same value as the reference index RefIdxL1[xPCol][yPCol] of L1 (step S4237), and the reference list listCol is set to L1 (step S4238). Here, xPCol and yPCol are indices indicating the positions of the pixels above the left of the encoding block colCb within different-time pictures ColPic. On the other hand, when the L0 prediction flag PredFlagL0[xPCol][yPCol] of the encoding block colCb is not 0 (step S4235: NO), it is determined whether the L1 prediction flag PredFlagL1[xPCol][yPCol] of the encoding block colCb is 0. When the L1 prediction flag PredFlagL1[xPCol][yPCol] of the encoding block colCb is 0 (step S4239: YES), the motion vector mvCol is set to the same value as the motion vector MvL0[xPCol][yPCol] which is the motion vector of L0 of the encoding block colCb (step S4240), the reference index refIdxCol is set to the same value as the reference index RefIdxL0[xPCol][yPCol] of L0 (step S4241), and the reference list listCol is set to L0 (step S4242).
[0192]
[0193] The L0 prediction flag PredFlagL0[xPCol][yPCol] of the encoding block colCb and the encoding block col When both of the L1 prediction flags PredFlagL1[xPCol][yPCol] of Cb are not 0 (step S4235 :NO, and S4239:NO), the inter prediction mode of the coding block colCb is bi-predictive Since the motion vector is (Pred_BI), one of the two motion vectors L0 and L1 is selected (step Top S4243).
[0194] FIG. 54 shows the code when the inter prediction mode of the coding block colCb is bi-prediction (Pred_BI). 13 is a flowchart showing a procedure for deriving inter prediction information of a coded block.
[0195] First, the POCs of all pictures registered in all reference lists are checked against the current processing target. It is determined whether the POC is smaller than the POC of the picture (step S4251). P of all pictures registered in L0 and L1, which are all reference lists of colCb If OC is smaller than the POC of the current picture to be processed (step S4251: YES), ), LX is L0, that is, the predicted vector candidate of the motion vector of L0 of the coding block to be processed. If the complement has been derived (step S4252: YES), The inter prediction information of the L1 block is selected, and LX is the motion vector of the L1 block of the coding block to be processed. If a predicted vector candidate for the vector is derived (step S4252: NO), On the other hand, the L1 inter prediction information of the coding block colCb is selected. At least one POC of a picture registered in all reference lists L0 and L1 of If the POC is greater than the POC of the current picture being processed (step S4251: NO), If the flag collocated_from_l0_flag is 0 (step S4253: YES), Select the inter prediction information for L0 of colCb, and when the flag collocated_from_l0_flag is 1 (step S4253: NO), select the inter prediction information for L1 of the coded block colCb .
[0196] When selecting the inter prediction information for L0 of the coded block colCb (step S42 52: YES, or step S4253: YES), the motion vector mvCol is set to the same value as MvL0[xPCol] [yPCol] (step S4254), the reference index refIdxCol is set to the same value as RefI dxL0[xPCol][yPCol] (step S4255), and the list listCol is set to L0 (step S4256).
[0197] When selecting the inter prediction information for L1 of the coded block colCb (step S42 52: NO, or step S4253: NO), the motion vector mvCol is set to the same value as MvL1[xPCol][yPC ol] (step S4257), the reference index refIdxCol is set to the same value as RefIdxL1 [xPCol][yPCol] (step S4258), and the list listCol is set to L1 (step S4259).
[0198] Return to FIG. 53. If inter prediction information can be obtained from the coded block colCb, set both the flag ava ilableFlagLXCol and the flag predFlagLXCol to 1 (step S4244).
[0199] Subsequently, scale the motion vector mvCol to obtain the motion vector mvLXCol (step PS4245). The scaling operation processing procedure of this motion vector mvLXCol will be described with reference to FIG. 55. Hereinafter.
[0200] Subtract the POC of the reference picture corresponding to the reference index refIdxCol referred to in the list listCol of the coding block colCb from the POC of the picture ColPic at different times to obtain the picture interval td as follows. td = [POC of picture ColPic at different times] - [POC of the reference picture referred to in the list listCol of the coding block colCb] (step S4261). Note that when the POC of the reference picture referred to in the list listCol of the coding block colCb is earlier in the display order than the picture ColPic at different times the picture interval td is a positive value, and when the POC of the reference picture referred to in the list listCol of the coding block colCb is later in the display order than the picture ColPic at different times the picture interval td is a negative value.
[0201] Next, subtract the POC of the reference picture referred to in the list LX of the current processing target picture from the POC of the current processing target picture to obtain the picture interval tb as follows. tb = [POC of the current processing target picture] - [POC of the reference picture corresponding to the reference index of LX of the time merge candidate] (step S4262). Note that when the reference picture referred to in the list LX of the current processing target picture is earlier in the display order than the current processing target picture the picture interval tb is a positive value, and when the reference picture referred to in the list LX of the current processing target picture is later in the display order than the current processing target picture the picture interval tb is a negative value.
[0202] Subsequently, the inter-picture distances td and tb are compared (step S4263), and when the inter-picture distance td and tb are equal (step S4263: YES), the motion vector mvLXCol is set to mvLXCol = mvCol and calculated (step S4264), and this scaling operation process is terminated.
[0203] On the other hand, when the inter-picture distances td and tb are not equal (step S4263: NO), the variable tx is set to tx = (16384 + Abs(td) >> 1) / td and calculated (step S4265). Subsequently, the scaling coefficient distScaleFactor is set to distScaleFactor = Clip3(-4096, 4095, (tb * tx + 32) >> 6) and calculated (step S4266). Here, Clip3(x, y, z) is a function that limits the value z to a minimum value of x and a maximum value of y. Subsequently, the motion vector mvLXCol is set to mvLXCol = Clip3(-32768, 32767, Sign(distScaleFactor * mvLXCol) * ((Abs(distScaleFactor * mvLXCol) + 127) >> 8)) and calculated (step S4267), and this scaling operation process is terminated. Here, Sign (x) is a function that returns the sign of the value x, and Abs(x) is a function that returns the absolute value of the value x. (x) is a function that returns the sign of the value x, and Abs(x) is a function that returns the absolute value of the value x.
[0204] Referring to FIG. 50 again. Then, the motion vector mvL0Col of L0 and the predicted motion vector candidate list mvpListLX in the above-described normal prediction motion vector mode derivation unit 301 are used as candidates and add it (step S4205). However, this addition is only performed when the flag availableFlagL0Col indicating whether the coded block colCb in reference list 0 is valid is 1. Also, the motion vector mvL1Col of L1 is added as a candidate to the motion vector candidate list mvpListLX in the normal prediction motion vector mode derivation unit 301 described above (step S4205). However, this addition is only performed when the flag availabl eFlagL1Col indicating whether the coded block colCb in reference list 1 is valid is 1. Through the above, the processing of the temporal prediction motion vector candidate derivation unit 322 is terminated. The above description of the normal prediction motion vector mode derivation unit 301 is for the encoding process, but the same applies to the decoding process. That is, the operation of the temporal prediction motion vector candidate derivation unit 422 in the normal prediction motion vector mode derivation unit 401 in FIG. 23 is explained in the same way by replacing the encoding in the above description with decoding.
[0205] The above description of the normal prediction motion vector mode derivation unit 301 is for the encoding process, but the same applies to the decoding process. That is, the operation of the temporal prediction motion vector candidate derivation unit 422 in the normal prediction motion vector mode derivation unit 401 in FIG. 23 is explained in the same way by replacing the encoding in the above description with decoding. The operation of the temporal prediction motion vector candidate derivation unit 422 in the normal prediction motion vector mode derivation unit 401 in FIG. 23 is explained in the same way by replacing the encoding in the above description with decoding. is terminated.
[0206] <Temporal Merge Candidate Derivation> The operation of the temporal merge candidate derivation unit 342 in the normal merge mode derivation unit 302 in FIG. 18 will be described with reference to FIG. 56. First, ColPic is derived (step S4301). Next, the coded block colCb is derived
[0207] and the coding information is obtained (step S4302). Furthermore, for each reference list, the inter -prediction information is derived (steps S4303, S4304). The above processing is the same as S4201 to S4204 in the temporal prediction motion vector candidate derivation unit 322, so the description is omitted. is omitted.
[0208] Next, a flag availableFlagCol indicating whether the encoded block colCb is valid is calculated ( step S4305). When the flag availableFlagL0Col or the flag availableFlagL1Col is 1, availableFlagCol becomes 1. Otherwise, availableFlagCol becomes 0.
[0209] Then, the motion vector mvL0Col of L0 and the motion vector mvL1Col of L1 are added as candidates to the merge candidate list mergeCandList in the normal merge mode derivation unit 302 described above (step S4306). However, this addition is only performed when the flag availableFlagCol indicating whether the encoded block colCb is valid is 1. Thus, the processing of the temporal merge candidate derivation unit 34 (step S4306). However, this addition is only performed when the flag availableFlagCol indicating whether the encoded block colCb is valid is 1. Thus, the processing of the temporal merge candidate derivation unit 34 2 ends. 2 ends.
[0210] The above description of the temporal merge candidate derivation unit 342 is for the encoding process, but the same applies to the decoding process. That is, the operation of the temporal merge candidate derivation unit 442 in the normal merge mode derivation unit 402 in FIG. 24 is described in the same way by replacing the encoding in the above description with decoding. unit 442 in the normal merge mode derivation unit 402 in FIG. 24 is described in the same way by replacing the encoding in the above description with decoding.
[0211] <Update of the history predicted motion vector candidate list> Next, the initialization and update method of the history predicted motion vector candidate list HmvpCandList provided in the encoding information storage memory 111 on the encoding side and the encoding information storage memory 20 5 on the decoding side will be described in detail. FIG. 38 is a flowchart for explaining the history predicted motion vector candidate list initialization / update processing procedure. Next, the initialization and update method of the history predicted motion vector candidate list HmvpCandList provided in the encoding information storage memory 111 on the encoding side and the encoding information storage memory 20 5 on the decoding side will be described in detail. FIG. 38 is a flowchart for explaining the history predicted motion vector candidate list initialization / update processing procedure.
[0212] In this embodiment, the history motion vector predictor candidate list HmvpCandList is updated based on the coding information. The information storage memory 111 and the encoded information storage memory 205 are assumed to be implemented. A history candidate list update unit is provided in the prediction unit 102 and the inter-prediction unit 203 to perform history prediction. An update of the motion vector candidate list HmvpCandList may be performed.
[0213] At the beginning of the slice, the historical motion vector prediction candidate list HmvpCandList is initialized. On the encoding side, the prediction method decision unit 105 selects the normal prediction vector mode or the normal merge mode. When the selected candidate is selected, the history motion vector predictor candidate list HmvpCandList is updated. The inter prediction mode decoded by the bit sequence decoding unit 201 is a normal prediction vector mode or In the normal merge mode, the historical motion vector predictor candidate list HmvpCandList is updated.
[0214] Inter prediction mode used for inter prediction in normal prediction vector mode or normal merge mode. The inter prediction information is used as the inter prediction information candidate hMvpCand, and the historical predicted motion vector candidate list is Register it in HmvpCandList. The inter-prediction information candidate hMvpCand contains the reference index of L0. refIdxL0 and L1 reference index refIdxL1, L0 prediction indicating whether L0 prediction is performed a prediction flag predFlagL0 indicating whether or not L1 prediction is performed; and a L1 prediction flag predFlagL1 indicating whether or not L1 prediction is performed. The motion vector mvL0 of L0 and the motion vector mvL1 of L1 are included. The motion vector history prediction memory 111 and the coding information storage memory 205 on the decoding side Among the elements (i.e., inter prediction information) registered in the candidate list HmvpCandList, If there is inter prediction information having the same value as the inter prediction information candidate hMvpCand, then history prediction Delete the element from the motion vector candidate list HmvpCandList. On the other hand, for the inter prediction information If there is no inter prediction information having the same value as the candidate hMvpCand, then delete the first element of the history prediction motion vector candidate list HmvpCandList, and add the inter prediction information candidate hMvpCand to the end of the history prediction motion vector candidate list HmvpCandList List
[0215] The encoding information storage memory 111 on the encoding side and the encoding information storage memory 2 on the decoding side of the present invention 05 is provided such that the number of elements in the history prediction motion vector candidate list HmvpCandList is 6
[0216] First, initialize the history prediction motion vector candidate list HmvpCandList in units of slices (Step S2101 in FIG. 38). At the start of the slice, empty all elements of the history prediction motion vector candidate list Hm vpCandList, and set the value of the number NumHmvpCand of the history prediction motion vector candidates recorded in the history prediction motion vector candidate list HmvpCandList to 0
[0217] Note that the initialization of the history prediction motion vector candidate list HmvpCandList is performed in units of slices (the first encoded block of the slice), but it may also be performed in units of pictures, tiles, or tree blocks rows
[0218] Subsequently, for each encoded block in the slice, repeat the following update process for the history prediction motion vector candidate list Hmvp CandList (Steps S2102 to S2107 in FIG. 38).
[0219] First, perform initial settings in units of encoding blocks. Set the flag identicalCandExist indicating whether there is an identical candidate to FALSE (false), and set 0 to the deletion target index removeIdx (step S2103 in FIG. 38).
[0220] Determine whether there is an inter-prediction information candidate hMvp Cand to be registered in the history prediction motion vector candidate list HmvpCandList (step S2104 in FIG. 38). The prediction method on the encoding side is determined to be the normal prediction motion vector mode or the normal merge mode by the prediction method determination unit 105. Or when it is decoded as the normal prediction motion vector mode or the normal merge mode by the bit string decoder 201 on the decoding side, set that inter-prediction mode as hMvpCand. When the prediction method determination unit 105 on the encoding side determines the intra-prediction mode, the sub-block prediction motion vector mode, or the sub-block merge mode, or when it is decoded as the intra-prediction mode, the sub-block prediction motion vector mode, or the sub-block merge mode by the bit string decoder 201 on the decoding side, no update process is performed on the history prediction motion vector candidate list HmvpCandList, and there is no inter-prediction information candidate hMvpCand to be registered. If there is no inter-prediction information candidate hMvpCand to be registered, skip steps S2105 to S2106 (step S2104 in FIG. 38: NO). If there is an inter-prediction information candidate hMvpCand to be registered, perform the following processing from step S2105 (step S2104 in FIG. 38: YES)
[0221] Subsequently, among the elements of the history prediction motion vector candidate list HmvpCandList, the inter-prediction information candidate to be registered Determine whether there is an element identical to the candidate hMvpCand of the movement prediction information (step in FIG. 38 S2105). FIG. 39 is a flowchart of this identical element confirmation processing procedure. The historical prediction movement If the value of the number NumHmvpCand of the movement vector candidates is 0 (step S2121 in FIG. 39: NO), The historical prediction movement vector candidate list HmvpCandList is empty and there is no identical candidate, so steps S2122 to S2125 in FIG. 39 are skipped and this identical element confirmation processing procedure is terminated. If the value of the number NumHmvpCand of the historical prediction movement vector candidates is greater than 0 (step S2 121 in FIG. 39: YES), the historical prediction movement vector index hMvpIdx ranges from 0 to NumHmvpCand - 1 and the process of step S2123 is repeated (steps S2122 to S2125 in FIG. 39) . First, compare whether the element HmvpCandList[hMvpIdx] at the hMvpIdx-th position counted from 0 in the historical prediction movement vector candidate list is identical to the candidate hMvpCand of the movement prediction information (step S2123 in FIG. 39). If they are identical (step S2123 in FIG. 39: YES), set the flag identicalCandExist indicating whether there is an identical candidate to TRUE (true), set the value of hMVpIndex to the deletion target index removeIdx, and terminate this identical element confirmation process. If they are not identical (step S2123 in FIG. 39: NO), increment hMvpIdx by 1, and if the historical prediction movement vector index hMvpIdx is less than or equal to NumHmvpCand - 1, perform the processes after step S2123 (steps S2122 to S2125 in FIG. 39).
[0222] processing is performed (steps S2122 to S2125 in FIG. 39).
[0222] Return to the flowchart of FIG. 38 again, and for the historical prediction movement vector candidate list HmvpCandList Perform element shift and addition processing (step S2106 in FIG. 38). FIG. 40 shows the element shift / addition processing procedure of the history prediction motion vector candidate list HmvpCandList in step S2106 of FIG. 38. First, determine whether to add new elements after removing the elements stored in the history prediction motion vector candidate list HmvpCandList, or to add new elements without removing the elements. Specifically, compare whether the flag identicalC andExist indicating whether the same candidate exists is TRUE (true) or whether NumHmvpCand is 6 (step S21 41 in FIG. 40). If either the flag identicalCandExist indicating whether the same candidate exists is TRUE (true) or NumHmvpCand is 6 is satisfied (step S2141 in FIG. 40: YES) , add new elements after removing the elements stored in the history prediction motion vector candidate list HmvpCandList. Set the initial value of the index i to the value of removeIdx + 1. From this initial value to NumHmvpCand, repeat the element shift processing in step S2143. (Steps S2142~S2144 in FIG. 40). Copy the element of HMVPCandList[i] to HMVPCandList[i - 1] to shift the element forward (step S2143 in FIG. 40), and increment i by 1 . (Steps S2142~S2144 in FIG. 40). When the index i becomes NumHmvp Cand + 1 and the element shift processing in step S2143 is completed, add the inter-prediction information candidate hMvpCand to the end of the history prediction motion vector candidate list (step S2 145 in FIG. 40). Here, the end of the history prediction motion vector candidate list means, counted from 0, (NumHmvp Copy the element of HMVPCandList[i] to HMVPCandList[i - 1] to shift the element forward (step S2143 in FIG. 40), and increment i by 1 . (Steps S2142~S2144 in FIG. 40). When the index i becomes NumHmvp Cand + 1 and the element shift processing in step S2143 is completed, add the inter-prediction information candidate hMvpCand to the end of the history prediction motion vector candidate list (step S2 145 in FIG. 40). Here, the end of the history prediction motion vector candidate list means, counted from 0, (NumHmvp It is the (Cand - 1) - th HMVPCandList[NumHmvpCand - 1]. Thus, the historical prediction motion vector candidate ends the element shift / addition process of the supplementary list HMVPCandList. On the other hand, whether there is an identical candidate or not, the flag identicalCandExist indicating this is TRUE (true) and when neither of the conditions that NumHmvpCand is 6 is satisfied (step S2141 in FIG. 40: NO), without removing the elements stored in the historical prediction motion vector candidate list HmvpCandList, the last inter - prediction information candidate hMvpCand is added to the historical prediction motion vector candidate list (step S2146 in FIG. 40). Here, the last of the historical prediction motion vector candidate list is the HMVPC andList[NumHmvpCand] counted from 0. Also, NumHmvpCand is incremented by 1, and the element shift / addition process of the historical prediction motion vector candidate list HMVPCandList ends.
[0223] FIG. 43 is a diagram for explaining an example of the update process of the historical prediction motion vector list. When six elements (inter - prediction information) are registered in the historical prediction motion vector candidate list HMVPCandList and new inter - prediction information is to be added, the new inter - prediction information is compared with each element of the historical prediction motion vector candidate list HMVPCandList from the front (FIG. 43(a)). If the new inter - prediction information has the same value as the third element HMVP2 from the top of the historical prediction motion vector candidate list HMVPCandList, the element HMVP2 is deleted from the historical prediction motion vector candidate list HMVPCandList, and the subsequent elements HMVP3 - HMVP5 are shifted (copied) forward by one position each, and the historical prediction motion vector. Add new inter-prediction information at the end of the history motion vector candidate list HMVPCandList (Figure 43 (b)), and complete the update of the history motion vector candidate list HMVPCandList (Figure 43(c ))
[0224] <History Motion Vector Candidate Derivation Process> Next, the history motion vector candidate derivation method of the history motion vector candidate derivation unit 323 of the normal prediction motion vector mode derivation unit 301 on the encoding side and the history motion vector candidate derivation unit 423 of the normal prediction motion vector mode derivation unit 401 on the decoding side, which is a common process in step S304 of FIG. 20 will be described in detail. FIG. 41 is a flowchart for explaining the history motion vector candidate derivation processing procedure
[0225] If the number of current prediction motion vector candidates numCurrMvpCand is greater than or equal to the maximum number of elements of the prediction motion vector candidate list mvpListLX (here, 2) or the number of history prediction motion vector candidates NumHmvpCand is 0 (step S2201: NO in FIG. 41), the processing from step S220 2 to S2209 is omitted, and the history motion vector candidate derivation processing procedure ends If the number of current prediction motion vector candidates numCurrMvpCand is less than 2, which is the maximum number of elements of the prediction motion vector candidate list mvpListL X, and the number of history prediction motion vector candidates NumHmvpCand is greater than 0 (step S2201: YES in FIG. 41), the processing from step S 2202 to S2209 is performed
[0226] Subsequently, the process from step S2203 to S2208 in FIG. 41 is repeated with the index i ranging from 0 to a smaller value among 3 and the number NumHmvpCand of history prediction motion vector candidates ( FIG. 41 steps S2202 to S2209). When the number numCurrM vpCand of the current prediction motion vector candidates is 2 or more, which is the maximum number of elements in the prediction motion vector candidate list mvpListLX (FIG. 4 1 step S2203: NO), the process from step S2204 to S2209 in FIG. 41 is omitted, and this history prediction motion vector candidate derivation processing procedure is terminated. When the number numCurrMvpCand of the current prediction motion vector candidates is less than 2, which is the maximum number of elements in the prediction motion vector candidate list mvpListLX (FIG. 41 step S2203: YES), the process from step S2204 onward in FIG. 41 is performed.
[0227] Subsequently, the process from step S2205 to S2207 is performed for Y being 0 and 1 (L0 and L1) respectively (FIG. 41 steps S2204 to S2208). When the number numCurrMvpCand of the current prediction motion vector candidates is 2 or more, which is the maximum number of elements in the prediction motion vector candidate list mvpListLX (FIG. 41 step S2205: NO), the process from step S2206 or S2209 in FIG. 41 is omitted, and this history prediction motion vector candidate derivation processing procedure is terminated. When the number numCurrMvpCand of the current prediction motion vector candidates is less than 2, which is the maximum number of elements in the prediction motion vector candidate list mvpListLX (FIG. 41 step S2205: YES), the process from step S2206 onward in FIG. 41 is performed.
[0228] Subsequently, the reference index of LY in the history prediction motion vector candidate list HmvpCandList[i] is When it is the same as the reference index refIdxLX of the symbolized / decoded target motion vector (step S2206 in FIG. 41: YES), as the last element of the predicted motion vector candidate list, the predicted motion vector candidate list is counted from 0. The motion vector of LY of the element mvpListLX[numCurrMvpCand] at the numCurrMvpCand-th position is added to it (step S2207 in FIG. 41), and the number numCurrMvpCand of the current predicted motion vector candidates is incremented by 1. When the reference index of LY in the history predicted motion vector candidate list HmvpCandList[i] is not the same as the reference index refIdxLX of the symbolized / decoded target motion vector (step S2206 in FIG. 41: NO), the addition process in step S2207 is skipped.
[0229] The processes from step S2205 to S2207 in FIG. 41 above are performed both in L0 and L1 (steps S2204 to S2208 in FIG. 41).
[0230] Increment the index i by 1. When the index i is less than or equal to either 3 or the number NumHmvpCand of the history predicted motion vector candidates, the processes after step S2203 are performed again (steps S2202 to S2209 in FIG. 41).
[0231] <History merge candidate derivation process> Next, it is the processing procedure of step S404 in FIG. 21, which is a common process in the history merge candidate derivation unit 345 of the normal merge mode derivation unit 302 on the encoding side and the history merge candidate derivation unit 445 of the normal merge mode derivation unit 402 on the decoding side. The history from the history merge candidate list HmvpCandList The method for deriving merge candidates will be described in detail. FIG. 42 shows the procedure of the history merge candidate derivation process is a flowchart for explanation.
[0232] First, initialization processing is performed (step S2301 in FIG. 42). Set the value of each element from 0 to (numCu rrMergeCand - 1) of isPruned[i] to FALSE, and set the variable numOrigMergeCand to the number numCurrMergeCand of elements registered in the current merge candidate list.
[0233] Subsequently, set the initial value of the index hMvpIdx to 1, and repeat the additional processing from step S2303 to step S2310 in FIG. 42 from this initial value to NumHmvpCand (steps S2302~S2311 in FIG. 4 2). If the number numCurrMergeCand of elements registered in the current merge candidate list is less than or equal to (the maximum number of merge candidates MaxNumMergeCand - 1), since merge candidates have been added to all elements of the merge candidate list, this history merge candidate derivation process is ended (step S2303 in FIG. 42: NO). If the number numCurrMergeCand of elements registered in the current merge candidate list is less than or equal to (the maximum number of merge candidates MaxNumMergeCand - 1) (step S2303 in FIG. 42: YES), perform the processing after step S2304. First, set the value of sameMotion to FALSE (false) (step S2304 in FIG. 42). Subsequently, set the initial value of the index i to 0, and perform the processing of steps S2 306 and S2307 in FIG. 42 from this initial value to 1 (S2305~S2308 in FIG. 42).
[0234]
[0235] Next, compare the (NumHmvpCand - hMvpIdx)-th element HmvpCandList[NumHmvpCand - hMvpIdx] of the historical motion vector prediction candidate list, counted from 0, with the i-th element mergeCandList[i] of the merge candidate list, counted from 0 (step S2306 in FIG. 42). Here the merge candidate having the same value means that the values of all the constituent elements (inter prediction mode, reference index, motion vector) of the merge candidate are the same. However, the process of this step S2306 is limited to the case where hMvpIdx is greater than NumHmvpCand - 2, and mergeCandList[i] is a spatial merge candidate and isPruned[i] is FALSE (false). In the case of the same value (step S2306: YES in FIG. 39), both sameMotion and isPruned[i] are set to TRUE (true) (step S2307 in FIG. 42). In the case of not having the same value (step S2306: NO in FIG. 39) skip the process of step S2307. After the iterative process from step S2305 to step S2308 in FIG. 42 is completed, compare whether sameMotion is FALSE (false) (step S2309 in FIG. 42). If sameMotion is FALSE (false) (step S2309: YES in FIG. 42), add the (NumHmvpCand - hMvp Idx)-th element HmvpCandList[NumHmvpCand - hMvpIdx] of the historical predicted motion vector candidate list, counted from 0, to the numCurrMergeCand-th mergeCandList[numCurrMergeCand] of the merge candidate list, and increment numCurrMergeCand by 1 (step S2310 in FIG. 42). Increment the index hMvpIdx by 1 (step S2310 in FIG. 42). After the iterative process from step S2305 to step S2308 in FIG. 42 is completed, compare whether sameMotion is FALSE (false) (step S2309 in FIG. 42). If sameMotion is FALSE (false) (step S2309: YES in FIG. 42), add the (NumHmvpCand - hMvp Idx)-th element HmvpCandList[NumHmvpCand - hMvpIdx] of the historical predicted motion vector candidate list, counted from 0, to the numCurrMergeCand-th mergeCandList[numCurrMergeCand] of the merge candidate list, and increment numCurrMergeCand by 1 mCurrMergeCand] of the merge candidate list, and increment numCurrMergeCand by 1 (step S2310 in FIG. 42). Increment the index hMvpIdx by 1 (step S2310 in FIG. 42). Increment the index hMvpIdx by 1 Mentor (step S2302 in FIG. 42), and perform the repetitive process of steps S2302 to S2311 in FIG. 42.
[0236] When the confirmation of all elements in the history prediction motion vector candidate list is completed, or when merge candidates are added to all elements of the merge candidate list, the derivation process of the present history merge candidate is completed .
[0237] <Average merge candidate derivation process> Next, the average merge candidate derivation unit 344 of the normal merge mode derivation unit 302 on the encoding side and the average merge candidate derivation unit 444 of the normal merge mode derivation unit 402 on the decoding side will be described in detail. The derivation method of the average merge candidate, which is the common process in step S403 of FIG. 21, will be described in detail. FIG. 62 is a flowchart for explaining the average merge candidate derivation process. First, perform initialization processing (step S1301 in FIG. 62). Set the number numCurrMergeCand of elements registered in the current merge candidate list to the variable numOrigMergeCand.
[0238]
[0239] Subsequently, scan the merge candidate list in order from the head, and determine two pieces of motion information. Let the index i = 0 indicating the first piece of motion information and the index j = 1 indicating the second piece of motion information. ( Steps S1302 to S1303 in FIG. 62). If the number numCurrMergeCand of elements registered in the current merge candidate list is not less than (the maximum number of merge candidates MaxNumMergeCand - 1), then since merge candidates have been added to all elements of the merge candidate list, the present history merge candidate derivation process is terminated (step S1304 in FIG. 62). Currently registered in the merge candidate list If the number numCurrMergeCand of elements is less than or equal to (the maximum number of merge candidates MaxNumMergeCand - 1), the process after step S1305 is performed. Perform the processing after step S1305.
[0240] Determine whether both the motion information mergeCandList[i] of the i-th element in the merge candidate list and the motion information mergeCandList[j] of the j-th element in the merge candidate list are invalid (step S1305 in FIG. 62). If both are invalid, instead of deriving the average merge candidate of mergeCandList[i] and mergeCandList[j], move to the next element. If mergeCandList[i] and mergeCandList[j] are not both invalid, repeat the following processing with X set to 0 and 1 (steps S1306 to S1314 in FIG. 62). Determine whether the LX prediction of mergeCandList[i] is valid (step S1307 in FIG. 62). If the LX prediction of mergeCandList[i] is valid, determine whether the LX prediction of mergeCandList[j] is valid (step S1308 in FIG. 62). If the LX prediction of mergeCandList[j] is valid, that is, if both the LX prediction of mergeCandList[i] and the LX prediction of mergeCandList[j] are valid, derive the average merge candidate of the LX prediction with the motion vector obtained by averaging the motion vectors of the LX predictions of mergeCandList[i] and mergeCandList[j] and the reference index of the LX prediction of mergeCandList[i], set it to the LX prediction of averageCand, and make the LX prediction of averageCand valid (step S1309 in FIG. 62). In step S13 of FIG. 62
[0241] 08. If the LX prediction of mergeCandList[j] is not valid, that is, the LX prediction of mergeCandList[i] is valid and the LX prediction of mergeCandList[j] is invalid, then the LX average merge candidate of the LX prediction with the motion vector and reference index of the prediction is derived and set as the LX prediction of averageCand, and the LX prediction of averageCand is made valid (step S13 10 in FIG. 62). In step S1307 of FIG. 62, if the LX prediction of mergeCandList[i] is not valid, it is determined whether the LX prediction of mergeCandList[j] is valid (step S1311 in FIG. 62 ). If the LX prediction of mergeCandList[j] is valid, that is, the LX prediction of mergeCandList[i] is invalid and the LX prediction of mergeCandList[j] is valid, then the average merge candidate of the LX prediction with the motion vector and reference index of the LX prediction of mergeCandList[j] is derived and set as the LX prediction of averageC and, and the LX prediction of averageCand is made valid (step S1312 in FIG. 62 ). In step S1311 of FIG. 62, if the LX prediction of mergeCandList[j] is not valid, that is to say, both the LX prediction of mergeCandList[i] and the LX prediction of mergeCandList[j] are invalid, then the LX prediction of averageCand is made invalid (step S1312 in FIG. 62).
[0242] Here, the LX prediction is valid when the reference index refIdxLX is 0 or more and, when the LX prediction is invalid, that is, does not exist, the reference index refIdxLX is set to -1.
[0243] The average merge candidate averageCand of the L0 prediction, L1 prediction, or BI prediction generated as described above is added to the numCurrMergeCand-th mergeCandList[numCurrMergeCand] of the merge candidate list and numCurrMergeCand is incremented by 1 (step S1315 in FIG. 62). Thus, the process of deriving the average merge candidate is completed.
[0244] Note that the average merge candidate is averaged for each of the horizontal component and the vertical component of the motion vector. It is averaged.
[0245] <Sub-block time merge candidate derivation> The operation of the sub-block time merge candidate derivation unit 381 in the sub-block merge mode derivation unit 304 of FIG. 16 will be described with reference to FIG. 44. First, it is determined whether the coded block is less than 8×8 pixels (step S4002).
[0246] First, it is determined whether the coded block is less than 8×8 pixels (step S4002).
[0247] If the coded block is less than 8×8 pixels (step S4002: YES), the flag availableFlagSbCol indicating the existence of the sub-block time merge candidate is set to 0 (step S400 3), and the process of the sub-block time merge candidate derivation unit is terminated. Here, if temporal motion vector prediction is prohibited by syntax, or if sub-block time merge is prohibited, when the coded block is less than 8×8 pixels (step S4002: YES), the same processing is performed. YES), the same processing is performed.
[0248] On the other hand, if the coded block is 8×8 pixels or more (step S4002: NO), the coded pi Derive the adjacent motion information of the coded block in the chunk (step S4004).
[0249] The process of deriving the adjacent motion information of the coded block will be described with reference to FIG. 45. The process of deriving the adjacent motion information is similar to the process of the aforementioned spatial prediction motion vector candidate derivation unit 321. However, the order of searching for adjacent blocks is A0, B0, B1, A1, and B2 is not searched. First, set the adjacent block n = A0 and obtain the coding information (step S4052). The coding information includes a flag availableFlagN indicating whether the adjacent block can be used, a reference index refIdxLXN for each reference list, and a motion vector mvLXN.
[0250] Next, determine whether the adjacent block n is valid or invalid (step S4054). If the flag availableFlagN indicating whether the adjacent block can be used is 1, it is valid; otherwise, it is invalid.
[0251] If the adjacent block n is valid (step S4054: YES), set the reference index refIdxLXN to the reference index refIdxLXn of the adjacent block n (step S4056). Also, set the motion vector mvLXN to the motion vector mvLXn of the adjacent block n (step S4056), and end the process of deriving the adjacent motion information of the block.
[0252] On the other hand, if the adjacent block n is invalid (step S4054: NO), set the adjacent block n = B0, obtain the coding information (step S4052), and determine whether the adjacent block n is valid or invalid (step S4054). Then, loop in the order of B1, A1 in the same way. 。The process of deriving adjacent motion information loops until the adjacent blocks become valid, and all adjacent If blocks A0, B0, B1, and A1 are invalid, the process of deriving the adjacent motion information of the blocks ends and proceeds.
[0253] Refer to FIG. 44 again. After deriving the adjacent motion information (step S4004), the temporal motion vector is derived (step S4006).
[0254] The process of deriving the temporal motion vector will be described with reference to FIG. 46. First, the temporal motion vector tempMv is initialized to (0, 0) (step S4062).
[0255] Next, it is determined whether the adjacent motion information is valid or invalid (step S4064). If the flag availableFlagN indicating whether the adjacent blocks can be used is 1, it is valid; otherwise, it is invalid. If the adjacent motion information is invalid (step S4064: NO), the process of deriving the temporal motion vector ends. and proceeds.
[0256] On the other hand, if the adjacent motion information is valid (step S4064: YES), it is determined whether the flag predFlagL1N indicating whether L1 prediction is used in the adjacent block N is 1 or not (step S 4066). If predFlagL1N = 0 (step S4066: NO), the process proceeds to the next process ( step S4078). If predFlagL1N = 1 (step S4066: YES), it is determined whether the POCs of all the pictures registered in all the reference lists are less than or equal to the POC of the current picture being processed (step S4068). If this determination is true (step S4068: YES), the process proceeds to the next process (step S4070). and proceeds. and proceeds.
[0257] When the slice type slice_type is a B slice and the flag collocated_from_l0_flag is 0 (step S4070: YES, step and S4072: YES), check whether ColPic and the reference picture RefPicList1[refIdxL1N] (the picture with the reference index refIdxL1N in the reference list L1) are the same (step S4074). If this determination is true (step S 4074: YES), set the temporal motion vector tempMv = mvL1N (step S4076 ). If this determination is false (step S4074: NO), proceed to the next process (step S4078 ). If the slice type slice_type is not a B slice and the flag collocated_from_l0_f lag is not 0 (step S4070: NO, or step S4072: NO), proceed to the next process (step S4078).
[0258] Then, determine whether the flag predFlagL0N indicating whether L0 prediction is used in the adjacent block N is 1 (step S4078). If predFlagL0N = 1 (step S40 78: YES), check whether ColPic and the reference picture RefPicList0[refIdxL0N] (the picture with the reference index refIdxL0N in the reference list L0) are the same (step S4080). If this determination is true (step S4080: YES), set the temporal motion vector tempMv = m vL0N (step S4082). If this determination is false (step S4080: NO) , end the process of deriving the temporal motion vector.
[0259] Again, refer to FIG. 44. Next, ColPic is derived (step S4016). This process is the same as S4201 in the temporal prediction motion vector candidate derivation unit 322, so the description will be omitted.
[0260] Then, coded blocks colCb at different times are set (step S4017). This sets the coded block located at the same position as the coded block to be processed in the picture ColPic at different times at the lower right center as colCb. This coded block corresponds to the coded block T1 in FIG. 49.
[0261] Next, the position obtained by adding the temporal motion vector tempMv to the coded block colCb is set as the new colCb (step S4018). Now, let the upper left position of the coded block colCb be (xCo lCb, yColCb), and the temporal motion vector tempMv be (tempMv[0], tempMv[1]) with 1 / 16 pixel accuracy . Then, xColCb = Clip3( xCtb, xCtb + CtbSizeY + 3, xColCb + ( tempMv[0] >> 4 ) ) yColCb = Clip3( yCtb, yCtb + CtbSizeY - 1, yColCb + ( tempMv[1] >> 4 ) ) is calculated. Here, the upper left position of the tree block is (xCtb, yCtb), and the size of the tree block is CtbSizeY. The coded block on ColPic including the position ((xColCb >> 3) << 3, (yColCb >> 3) << 3) becomes the new colCb. As shown in the above formula, the position after adding tempMv is corrected within the range of the size of the tree block so that it does not deviate significantly compared to before adding tempMv . . It is done. If this position is outside the screen, it is corrected to be within the screen.
[0262] Then, it is determined whether the prediction mode PredMode of this encoded block colCb is inter prediction (MODE_INTER ) (step S4020). If the prediction mode of colCb is not inter prediction (step S4020: NO), the flag av indicating the existence of sub-block temporal merge candidates ailableFlagSbCol = 0 is set (step S4003), and the process of the sub-block temporal merge candidate derivation section is terminated.
[0263] On the other hand, when the prediction mode of colCb is inter prediction (step S4020: YES), inter prediction information is derived for each reference list (steps S4022, S4023). Here for colCb, the central motion vector ctrMvLX for each reference list and the flag ctrPredFlagLX indicating whether LX prediction is used are derived. LX indicates the reference list, and in the derivation of reference list 0, LX becomes L0, and in the derivation of reference list 1, LX becomes L1. The derivation of inter prediction information will be described with reference to FIG. 47. When encoded blocks colCb at different times are not available (step S4112: NO) or when the prediction mode PredMode is intra prediction (MODE_INTRA) (step S4114 : NO), both the flag availableFlagLXCol and the flag predFlagLXCol are set to 0 (step S 4116), the motion vector mvCol is set to (0, 0) (step S4118), and the process of deriving inter prediction information is terminated.
[0264]
[0265] Symbolic block colCb is available (step S4112: YES), prediction mode PredMod When e is not intra prediction (MODE_INTRA) (step S4114: YES), the following steps Calculate mvCol, refIdxCol, and availableFlagCol in the following order.
[0266] Flag PredFlagLX[xPCo l][yPCol] indicating whether LX prediction of symbolic block colCb is used is 1 (step S4120: YES), motion vector mvCol is set to the same value as the motion vector MvLX[xPCol][yPCol] of LX of symbolic block colCb (step S4122), reference index refIdxCol is set to the same value as the reference index RefIdxLX[xPCol] [yPCol] (step S4124), and list listCol is set to LX (step S4126). Here, xPCol and yPCol are indices indicating the positions of the top - left pixels of symbolic block colCb in different - time pictures ColPic. (step S4126). Here, xPCol, yPCol are indices indicating the positions of the top - left pixels of the symbolic block colCb in different - time pictures ColPic. is an index indicating the position of the top - left pixel of the symbolic block colCb in different - time pictures ColPic.
[0267] On the other hand, when the flag PredFlagL X[xPCol][yPCol] indicating whether LX prediction of symbolic block colCb is used is 0 (step S4120: NO), perform the following processing. First, judge whether the POC of all pictures registered in all reference lists is less than or equal to the POC of the current picture to be processed (step S4128). And judge whether the flag PredFlagLY[xPCol][yPCol] indicating whether LY prediction of colCb is used is 1 (step S4128). Here, LY prediction is defined as a different reference list from LX prediction. That is judge whether the flag PredFlagLY[xPCol][yPCol] indicating whether LY prediction of colCb is used is 1 (step S4128). Here, LY prediction is defined as a different reference list from LX prediction. That is a different reference list from LX prediction. That is When LX = L0, LY = L1; when LX = L1, LY = L0.
[0268] If this determination is true (step S4128: YES), the motion vector mvCol is set to the same value as MvLY[xPCol][yPCol], which is the motion vector of LY in the coding block colCb (step S4130), the reference index refIdxCol is set to the same value as RefIdxLY[xPCol][yPCol], which is the reference index of LY (step S4132), and the list listCol is set to LX (step S4134). On the other hand, if this determination is false (step S4128: NO), both the flag availableFlagLXCol and the flag predFlagLXCol are set to 0 (step S4116), the motion vector mvCol is set to (0, 0) (step S4118), and the process of deriving the inter-prediction information ends.
[0269] If inter-prediction information can be obtained from the coding block colCb, both the flag availableFlagLXCol and the flag predFlagLXCol are set to 1 (step S4136).
[0270]
[0271] Subsequently, the motion vector mvCol is scaled to obtain the motion vector mvLXCol (step S4138). Since this process is the same as S4245 in the temporal prediction motion vector candidate derivation unit 322, the description thereof is omitted.
[0272] Referring to FIG. 44 again, after inter-prediction information is derived for each reference list, the calculated motion vector mvLXCol is used as the central motion vector ctrMvLX, and the calculated flag predFlagLXCol is Set it as flag ctrPredFlagLX (Steps S4022, S4023).
[0273] Then, determine whether the central motion vector is valid (Step S4024). ctrPre If dFlagL0 = 0 and ctrPredFlagL1 = 0, it is determined to be invalid; otherwise, it is determined to be valid. If the central motion vector is invalid (Step S4024: NO), set the flag availableFlagSbCol indicating the existence of sub-block time merge candidates to 0 (Step S4003), and end the process of the sub-block time merge candidate derivation unit. On the other hand, if the central motion vector is valid (Step S4024: YES), set the flag availableFlagSbCol indicating the existence of sub-block time merge candidates to 1 (Step S4025), and derive sub-block motion information (Step S4026). This process will be described with reference to FIG. 48.
[0274] 5), and derive sub-block motion information (Step S4026). This process will be described with reference to FIG. 48.
[0275] First, calculate the number of sub-blocks numSbX in the width direction and the number of sub-blocks numSbY in the height direction from the width cbWidth and height cBheight of the coded block colCb (Step S4152). Also, set refIdxLXSbCol = 0 (Step S4152). After this process, the following repetitive process is performed in units of the predicted sub-block colSb. This repetition is performed while changing the index ySbIdx in the height direction from 0 to numSbY and the index xSbIdx in the width direction from 0 to numSbX. If the upper left position of the coded block colCb is (xCb, yCb), then the left
[0276] The upper position (xSb, ySb) is xSb = xCb + xSbIdx * sbWidth ySb = yCb + ySbIdx * sbHeight and calculated as follows. Next, the temporal motion vector tempMv is added to the predicted sub-block colSb and the resulting position is set as the new colSb (step S4154). Let the upper left position of the predicted sub-block colSb be (xColSb, yColSb), and the temporal motion vector tempMv be at 1 / 16 pixel precision (tempMv[0], tempMv[1]). Then, the upper left position of the new colSb is xColSb = Clip3( xCtb, xCtb + CtbSizeY + 3, xSb + ( tempMv[0] >> 4 ) ) yColSb = Clip3( yCtb, yCtb + CtbSizeY - 1, ySb + ( tempMv[1] >> 4 ) ) Here, the upper left position of the tree block is (xCtb, yCtb), and the size of the tree block is CtbSizeY. As shown in the above formula, the position after adding tempMv is corrected within a range of about the size of the tree block so that it does not deviate significantly compared to before adding tempMv. If this position is outside the screen, it is corrected to be within the screen.
[0277] And then, inter-prediction information is derived for each reference list (steps S4156, S41 58). Here, for the predicted sub-block colSb, the motion vector mvLXSbCol with respect to each reference list in sub-block units and the flag availableFl agLXSbCol indicating whether the predicted sub-block is valid are derived. LX indicates the reference list. For the derivation of reference list 0, LX is L0, and for the derivation of reference list 1, LX is L1. In the derivation of Reference List 1, LX becomes L1. Since the derivation of the inter prediction information is the same as that in S4022 and S4023 of FIG. 47, the description thereof is omitted.
[0278] After deriving the inter prediction information (steps S4156 and S4158), it is determined whether the prediction sub-block col Sb is valid (step S4160). When availableFlagL0SbCol = 0 and availa bleFlagL1SbCol = 0, colSb is determined to be invalid; otherwise, it is determined to be valid. When colSb is invalid (step S4160: NO), the motion vector mvLXSbCol is set to the central motion vector ctrMvLX and done (step S4162). Furthermore, the flag predFl agLXSbCol indicating whether LX prediction is used is set to the flag ctrPredFlagLX in the central motion vector (step S416 2). Thus, the derivation of the sub-block motion information is completed.
[0279] Referring to FIG. 44 again. Then, the motion vector mvL0SbCol of L0 and the motion vector mvL1SbCol of L1 are added as candidates to the sub-block merge candidate list subblockMergeCandList in the sub-block merge mode derivation unit 304 described above (step S4028) . However, this addition is only performed when the flag availableSbCol indicating the existence of the sub-block temporal merge candidate = 1. Thus, the processing of the temporal merge candidate derivation unit 342 is completed.
[0280] The above description of the sub-block temporal merge candidate derivation unit 381 is for the encoding process, but the same applies to the decoding process. That is, in the sub-block merge mode derivation unit 404 of FIG. 22 The operation of the sub-block time merge candidate derivation unit 481 is described in the same manner by replacing the encoding in the above description with decoding. and replacing it, and is described in the same way.
[0281] <Motion compensation prediction processing> The motion compensation prediction unit 306 acquires the position and size of the block that is currently the target of the prediction process in the encoding. Further, the motion compensation prediction unit 306 acquires the inter prediction information from the inter prediction mode determination unit 305. From the acquired inter prediction information, reference indexes and motion vectors are derived, and the reference picture specified by the reference index in the decoded image memory is moved from the same position as the image signal of the prediction block by the amount of the motion vector, and then the prediction signal is generated after acquiring the image signal at the moved position. and acquires the reference indexes and motion vectors from the acquired inter prediction information, and moves the reference picture specified by the reference index in the decoded image memory from the same position as the image signal of the prediction block by the amount of the motion vector, and then generates the prediction signal after acquiring the image signal at the moved position. and then generates a prediction signal.
[0282] In the case of prediction from a single reference picture, such as L0 prediction or L1 prediction, in inter prediction, the prediction signal obtained from one reference picture is used as the motion compensation prediction signal. In the case of prediction from two reference pictures, such as BI prediction, in inter prediction, the prediction signal obtained by weighted averaging the prediction signals obtained from two reference pictures is used as the motion compensation prediction signal, and the motion compensation prediction signal is supplied to the prediction method determination unit. Here, the ratio of weighted averaging in double prediction is set to 1:1, but weighted averaging may be performed using other ratios. and in the case of prediction from two reference pictures, such as BI prediction, in inter prediction, the prediction signal obtained by weighted averaging the prediction signals obtained from two reference pictures is used as the motion compensation prediction signal, and the motion compensation prediction signal is supplied to the prediction method determination unit. Here, the ratio of weighted averaging in double prediction is set to 1:1, but weighted averaging may be performed using other ratios. and supplies the motion compensation prediction signal to the prediction method determination unit. Here, the ratio of weighted averaging in double prediction is set to 1:1, but weighted averaging may be performed using other ratios. For example, the closer the picture interval between the picture to be predicted and the reference picture is, the larger the weighting ratio may be. Also, the calculation of the weighting ratio may be performed using a correspondence table between the combination of picture intervals and the weighting ratio.
[0283] The motion compensation prediction unit 406 has the same function as the motion compensation prediction unit 306 on the encoding side. The motion The compensation prediction unit 406 acquires the inter prediction information from the normal prediction motion vector mode derivation unit 401, the normal merge mode derivation unit 402, the sub-block prediction motion vector mode derivation unit 403, and the sub-block merge mode derivation unit 404 via the switch 408.
[0284] The motion compensation prediction unit 406 supplies the obtained motion compensation prediction signal to the decoded image signal superposition unit 207. Supply.
[0285] <Regarding the inter prediction mode> The process of performing prediction from a single reference picture is defined as single prediction. In the case of single prediction, L0 prediction or L1 prediction, that is, prediction using either one of the two reference pictures registered in the reference lists L0 and L1 is performed. L0 prediction and L1 prediction can be either forward prediction (prediction referring to a forward reference picture) or backward prediction (prediction referring to a backward reference picture). FIGS. 57 to 58 are diagrams for explaining motion compensation prediction in L0 prediction (single prediction). FIG. 57 shows a case where the inter prediction mode is L0 prediction and the reference picture of L0 (RefL0Pi
[0286] c) is at a time earlier than the picture to be processed (CurPic). FIG. 58 shows a case where it is L0 prediction and the reference picture of L0 is at a time later than the picture to be processed. FIG. 58 shows a case where it is L0 prediction and the reference picture of L0 is at a time later than the picture to be processed. Similarly, single prediction can also be performed by replacing the reference picture of L0 prediction in FIGS. 57 and 58 with the reference picture of L1 prediction (RefL1Pic). .
[0287] The process of performing prediction from two reference pictures is defined as double prediction. In the case of double prediction, it is expressed as double prediction using both L0 prediction and L1 prediction. FIGS. 59 to 61 are diagrams for explaining motion compensation prediction in double prediction. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed.
[0288] Thus, the relationship between the prediction types of L0 / L1 and time is not limited to L0 being forward prediction (prediction referring to a forward reference picture) and L1 being backward prediction (prediction referring to a backward reference picture), and they can be used in various ways. Also, in the case of dual prediction, the same reference picture can be used for each of L0 prediction and L1 prediction. Note that the determination of whether to perform motion compensation prediction using single prediction or dual prediction is made based on information (e.g., a flag) indicating whether to use L0 prediction and whether to use L1 prediction. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed. Thus, the relationship between the prediction types of L0 / L1 and time is not limited to L0 being forward prediction (prediction referring to a forward reference picture) and L1 being backward prediction (prediction referring to a backward reference picture), and they can be used in various ways. Also, in the case of dual prediction, the same reference picture can be used for each of L0 prediction and L1 prediction. Note that the determination of whether to perform motion compensation prediction using single prediction or dual prediction is made based on information (e.g., a flag) indicating whether to use L0 prediction and whether to use L1 prediction. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed.
[0289] <Regarding the reference index> In an embodiment of the present invention, in order to improve the accuracy of motion compensation prediction, it is possible to select an optimal reference picture from a plurality of reference pictures in motion compensation prediction. Therefore, the reference picture used in motion compensation prediction is used as a reference index, and the index is encoded into the encoded stream together with the encoding vector. These are diagrams for explaining prediction. FIG. 59 shows a dual prediction where the reference picture for L0 prediction is at a time before the picture to be processed, and the reference picture for L1 prediction is at a time after the picture to be processed. FIG. 60 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time before the picture to be processed. FIG. 61 shows a dual prediction where the reference pictures for both L0 prediction and L1 prediction are at a time after the picture to be processed. In an embodiment of the present invention, in order to improve the accuracy of motion compensation prediction, it is possible to select an optimal reference picture from a plurality of reference pictures in motion compensation prediction. Therefore, the reference picture used in motion compensation prediction is used as a reference index, and the index is encoded into the encoded stream together with the encoding vector.
[0290] <Motion compensation processing based on the normal prediction motion vector mode> The motion compensation prediction unit 306 is also shown in the inter prediction unit 102 on the encoding side in FIG. 16. As such, in the inter prediction mode determination unit 305, when the inter prediction information by the normal prediction motion vector mode derivation unit 301 is selected, this inter prediction information is acquired from the inter prediction mode determination unit 305, and the inter prediction mode, reference index, and motion vector of the currently processed block are derived, and a motion compensation prediction signal is generated. The generated motion compensation prediction signal is supplied to the prediction method determination unit 105. Similarly, as shown in the inter prediction unit 203 on the decoding side of FIG. 22, when the switch 408 is connected to the normal prediction motion vector mode derivation unit 401 during the decoding process, the motion compensation prediction unit 406 acquires the inter prediction information by the normal prediction motion vector mode derivation unit 401, and the inter prediction mode, reference index, and motion vector of the currently processed block are derived, and a motion compensation prediction signal is generated. The generated motion compensation prediction signal is supplied to the decoded image signal superimposing unit 207.
[0291]
[0292] <Motion compensation processing based on the normal merge mode> As shown in the inter prediction unit 102 on the encoding side of FIG. 16, when the inter prediction information by the normal merge mode derivation unit 302 is selected in the inter prediction mode determination unit 305, the motion compensation prediction unit 306 acquires this inter prediction information from the inter prediction mode determination unit 305, and the inter prediction mode, reference index, and motion vector of the currently processed block are derived, and a motion compensation prediction signal is generated. The generated motion compensation prediction signal is supplied to the prediction method determination unit 105.
[0293] Similarly, the motion compensation prediction unit 406 also functions like the inter prediction unit 203 on the decoding side in FIG. 22. As shown, when the switch 408 is connected to the normal merge mode derivation unit 402 during the decoding process, the motion compensation prediction unit 406 obtains the inter prediction information from the normal merge mode derivation unit 402, and derives the inter prediction mode, reference index, and motion vector of the block currently being processed, generates a motion compensation prediction signal. The generated motion compensation prediction signal is supplied to the decoded image signal superposition unit 207.
[0294] <Motion Compensation Processing Based on Sub-Block Prediction Motion Vector Mode> As shown by the inter prediction unit 102 on the encoding side in FIG. 16, when the inter prediction information from the sub-block prediction motion vector mode derivation unit 303 is selected by the inter prediction mode determination unit 305, the motion compensation prediction unit 306 obtains this inter prediction information from the inter prediction mode determination unit 305, and derives the inter prediction mode, reference index, and motion vector of the block currently being processed, generates a motion compensation prediction signal. The generated motion compensation prediction signal is supplied to the prediction method determination unit 105.
[0295] Similarly, the motion compensation prediction unit 406 also functions like the inter prediction unit 203 on the decoding side in FIG. 22. As shown, when the switch 408 is connected to the sub-block prediction motion vector mode derivation unit 403 during the decoding process, the motion compensation prediction unit 406 obtains the inter prediction information from the sub-block prediction motion vector mode derivation unit 403, and derives the inter prediction mode, reference index, and motion vector of the block currently being processed, generates a motion compensation prediction signal. The generated motion compensation prediction signal is supplied to the decoded image signal
[0296] <Motion compensation processing based on sub-block merge mode> When the motion compensation prediction unit 306 obtains the inter prediction information derived by the sub-block merge mode derivation unit 304 as shown in the inter prediction unit 102 on the encoding side of FIG. 16 in the inter prediction mode determination unit 305, it acquires this inter prediction information from the inter prediction mode determination unit 305, derives the inter prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is supplied to the prediction method determination unit 105.
[0297] Similarly, when the switch 408 is connected to the sub-block merge mode derivation unit 404 during decoding as shown in the inter prediction unit 203 on the decoding side of FIG. 22, the motion compensation prediction unit 406 acquires the inter prediction information derived by the sub-block merge mode derivation unit 404, derives the inter prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is supplied to the decoded image signal superposition unit 207.
[0298] <Motion compensation processing based on affine mode> In the present embodiment, motion compensation using an affine model can be used. Motion compensation using an affine model uses two to four corners of the encoded block as control points, derives the motion vector of the sub-block from the motion vectors of the control points, and performs motion compensation in sub-block units.
[0299] The following flags are determined by the inter prediction mode determination unit 305 in the encoding process It is reflected in the following flags based on the conditions of inter prediction and is encoded in the encoded stream. In the decoding process, it is determined whether to perform motion compensation by an affine model based on the following flags in the encoded stream.
[0300] The sps_affine_enabled_flag indicates whether motion compensation by an affine model can be used in inter prediction. If the sps_affine_enabled_flag is 0, it is suppressed so that motion compensation by an affine model is not performed at the sequence level. Also, the inter_affine_flag and the cu_affine_type_flag are not transmitted in the CU (Coded Unit) syntax of the coded video sequence. If the sps_affine_enabled_flag is 1, motion compensation by an affine model can be used in the coded video sequence.
[0301] The sps_affine_type_flag indicates whether motion compensation by a 6-parameter affine mode can be used in inter prediction. The 6-parameter affine model is a mode that derives the motion vector of a sub-block from six parameters of the horizontal and vertical components of the motion vectors of three control points respectively and performs motion compensation in sub-block units. It derives the motion vector in sub-block units, but derives a common reference index in coded block units.
[0302] If the sps_affine_type_flag is 0, it is suppressed so that motion compensation by a 6-parameter affine model is not performed. Also, the cu_affine_type_flag is for the CU of the coded video sequence. It is not transmitted in the syntax. If sps_affine_type_flag is 1, motion compensation using a 6-parameter affine model can be used in the coded video sequence.
[0303] If sps_affine_type_flag does not exist, it shall be 0.
[0304] When decoding a P or B slice, in the currently processed CU, if inte r_affine_flag is 1, motion compensation using an affine model is used to generate the motion compensation prediction signal for the currently processed CU.
[0305] If inter_affine_flag is 0, the affine model is not used for the currently processed CU.
[0306] If inter_affine_flag does not exist, it shall be 0.
[0307] When decoding a P or B slice, in the currently processed CU, if cu_a ffine_type_flag is 1, motion compensation using a 6-parameter affine model is used to generate the motion compensation prediction signal for the currently processed CU.
[0308] If cu_affine_type_flag is 0, motion compensation using a 4-parameter affine model is used to generate the motion compensation prediction signal for the currently processed CU. The 4-parameter affine model derives the motion vector of the sub-block from 4 parameters of the horizontal and vertical components of the motion vectors of each of the 2 control points, and is a mode of performing motion compensation in sub-block units.
[0309] <Merge differential motion vector (MMVD)> For the motion vectors of the top two merge candidates (the merge candidates with merge indices 0 and 1 in the merge candidate list), a differential motion vector can be added. This differential motion vector is called the merge differential motion vector. When adding the merge differential motion vector in the merge candidate selection unit 347 on the encoding side,
[0310] the motion vector to which the merge differential motion vector is added is supplied to the motion compensation prediction unit 306 via the inter prediction mode determination unit 305. Also, the bit string encoding unit 108 encodes information regarding the merge differential motion vector. Information regarding the merge differential motion vector refers to an index mmvd_distance_idx indicating the distance to be added to the motion vector and an index mmvd_direction_idx indicating the direction in which the motion vector is to be added. These indices are defined as in the tables shown in FIGS. 63(a) and 63(b). And, when representing the x and y components of the merge differential motion vector offset MmvdOffset as MmvdOffset[0] and MmvdOffset[1] respectively, it is as follows: MmvdOffset[0] = (MmvdDistance << 2) * MmvdSign[0] MmvdOffset[1] = (MmvdDistance << 2) * MmvdSign[1] The merge differential motion vector is derived from the merge differential motion vector offset MmvdOffset in the above formula. Details of deriving the merge differential motion vector will be described in the following case of the decoding side.
[0311] On the decoding side, when there is a merge differential motion vector, information regarding the merge differential motion vector is separated from the bit stream supplied to the bit string decoder 201, and a merge differential motion vector offset MmvdOffset is derived. Also, the merge candidate selection unit 447 derives a merge differential motion vector from the decoded merge differential motion vector offset. After adding this merge differential motion vector to the motion vector, the motion vector is supplied to the motion compensation prediction unit 406. The derivation of the merge differential motion vector mMvdLX in the merge candidate selection unit 447 will be described with reference to the flowchart of FIG. 4(a). First, it is determined whether the inter prediction mode of the coded block is bi-prediction (PRED_BI) (S4402). If it is not bi-prediction ( S4402:No), it is determined whether it is L0 prediction (PRED_L0) (S4404). If it is L0 prediction (S4404:Yes), mMvdL0 = MmvdOffset mMvdL1 = 0
[0312] and the process of deriving the merge differential motion vector ends (S4406). If it is L1 prediction (S4404:No), mMvdL0 = 0 mMvdL1 = MmvdOffset and the process of deriving the merge differential motion vector ends (S4408). On the other hand, if it is bi-prediction (S4402:Yes), the differences in POC between the processing target picture currPic and the reference pictures are calculated for each reference list, and are respectively set as currPocDiffL0 and currPocDiffL1 ( mMvdL0 = MmvdOffset mMvdL1 = 0 and the process of deriving the merge differential motion vector ends (S4406). If it is L1 prediction mMvdL0 = 0 mMvdL1 = MmvdOffset and the process of deriving the merge differential motion vector ends (S4408).
[0313] On the other hand, in the case of bi-prediction (S4402:Yes), the differences in POC between the processing target picture currPic and the reference pictures are calculated for each reference list, and are respectively set as currPocDiffL0 and currPocDiffL1 ( S4404:No), S4410). Here, the difference DiffPicOrderCnt(picA, picB) in the POC of picA and picB is DiffPicOrderCnt(picA, picB) = [POC of picA] - [POC of picB] as shown. Also, the reference picture RefPicList0[refIdxL0] is the picture indicated by the reference index refIdxL0 in the reference list L0. Similarly, the reference picture RefPicList1[refIdxL1] is the picture indicated by the reference index refIdxL1 in the reference list L1.
[0314] Next, it is determined whether -currPocDiffL0 * currPocDiffL1 >= 0 (step S4412 ). If this determination is true (step S4412: Yes), mMvdL0 = MmvdOffset mMvdL1 = -MmvdOffset and (step S4414), the process of deriving the merge differential motion vector ends. On the other hand, if this determination is false (step S4412: No), mMvdL0 = MmvdOffset mMvdL1 = MmvdOffset is set (step S4416). Next, it is determined whether the absolute value of the difference in POC with the reference list L0 is greater than or equal to the absolute value of the difference in POC with the reference list L1 (step S4418). If this determination is true (step S4418: Yes), X = 0, Y = 1 is set (step S4420), and the merge differential motion vector mMvdL1 of L1 is scaled (step S4424). Here, mMvdLY indicates mMvdL0 when Y = 0 and mMvdL1 when Y = 1. On the other hand, if this determination is false Combine (step S4418: No), set X = 1, Y = 0 (step S4422), and the merge difference of L0 Scale the motion vector mMvdL0 (step S4424). The merge difference motion vector Scaling of mMvdLY is as shown in Fig. 64(b), td = Clip3( -128, 127, currPocDiffLX ) tb = Clip3( -128, 127, currPocDiffLY ) tx = ( 16384 + Abs( td ) >> 1 ) / td distScaleFactor = Clip3( -4096, 4095, ( tb * tx + 32 ) >> 6 ) mMvdLY = Clip3( -32768, 32767, Sign( distScaleFactor * mMvdLY ) * ( (Abs( distScaleFactor * mMvdLY ) + 127 ) >> 8 ) ) is derived as such. Here, currPocDiffLX indicates currPocDiffL0 when X = 0 and cu rrPocDiffL1 when X = 1. Similarly, currPocDiffLY indicates currPocDiffL0 when Y = 0, and currPocDiffL1 when Y = 1. Also, Clip3(x, y, z) is a function that limits the value z to a minimum value of x and a maximum value of y. Sign(x) is a function that returns the sign of the value x, and Abs(x) is a function that returns the absolute value of the value x. Thus, the process of deriving the merge difference motion vector ends.
[0315] The merge difference motion vector may be added to the top two motion vectors of the sub-block merge candidates. In this case, the index mmvd_dis indicating the distance to be added to the motion vector The tance_idx is defined as in the table shown in FIG. 63(c). Sub-block merge candidate selection Since the operation of section 386 is the same as that of the merge candidate selection section 347, the description thereof is omitted. Also, the operation of the sub-block merge candidate selection section 486 is the same as that of the merge candidate selection section 447, so the description thereof is omitted.
[0316] As described above, MmvdDistance is defined as in the tables shown in FIGS. 63(a) and 63(c). Since these tables are defined with 1 / 4 pixel accuracy, the generated merge differential motion vectors may include fractional pixel accuracy. However, by encoding / decoding the flag indicating that the pixel accuracy of these tables is 1 on a slice-by-slice basis, the generated merge differential motion vectors can be changed so as not to include fractional pixel accuracy.
[0317] <Adaptive Motion Vector Resolution (AMVR)> The resolution of the differential motion vector can be adaptively changed in units of coded blocks. This resolution is referred to as the adaptive motion vector resolution.
[0318] The case of using the adaptive motion vector resolution for the normal prediction motion vector mode will be described. In this case, in the spatial prediction motion vector candidate derivation sections 321 and 421, the temporal prediction motion vector candidate derivation sections 322 and 422, and the history prediction motion vector candidate derivation sections 323 and 423, the derived candidate motion vectors are rounded according to the resolution. The resolution can be selected from 1 / 4, 1, and 4 pixel accuracies, and when the resolution is not changed, it is 1 / 4 pixel accuracy. The rounding process is performed in accordance with the resolution of the motion vector in the coded block to be processed. That is, the derived candidate motion vector mvX rightShift = leftShift = MvShift + 2 offset = 1 << ( rightShift - 1 ) mvX = ( mvX >= 0? ( mvX + offset ) >> rightShift : - ( ( - mvX + offset ) >> rightShift ) ) << leftShift and rounded. Here, when the resolution of the motion vector in the coding block to be processed is 1 / 4 pixel accuracy, MvShift = 0. Similarly, when the resolution of the motion vector is 1 pixel accuracy MvShift = 2, and when the resolution of the motion vector is 4 pixel accuracy, MvShift = 4. By the above formula, each of the x and y components of mvX is processed.
[0319] The adaptive motion vector resolution can also be used for the sub-block prediction motion vector mode. In this case, only the resolution is different from the above normal prediction motion vector mode. That is, in the affine inheritance prediction motion vector candidate derivation units 361 and 461, the affine construction prediction motion vector candidate derivation units 362 and 462, and the affine identical prediction motion vector candidate derivation units 363 and 463, the derived candidate motion vectors are rounded according to the resolution. The resolution can be selected from 1 / 16, 1 / 4, and 1 pixel accuracy, and when the resolution is not changed, it is 1 / 16 pixel accuracy. The rounding process is performed according to the resolution of the motion vector in the coding block to be processed. That is, the derived candidate motion vector mvX is rounded by the above formula. Here, when the resolution of the motion vector in the coding block to be processed is 1 / 4 pixel accuracy, MvShift = 0. Similarly, when the resolution of the motion vector is 1 pixel accuracy MvShift = 2, and when the resolution of the motion vector is 4 pixel accuracy, MvShift = 4. By the above It is MvShift = 2. According to the above formula, the x and y components of mvX are each processed.
[0320] <Triangle merge mode> The triangle merge mode is a type of merge mode, and it is a mode that performs motion compensation prediction by dividing the inside of the encoding / decoding block into diagonal partitions.
[0321] The triangle merge mode will be described with reference to FIG. 65. FIG. 65 shows the prediction state of an encoding / decoding block in the 16x16 triangle merge mode. The encoding / decoding block in the triangle merge mode is divided into 4x4 sub-blocks, and each sub-block is a single prediction partition 0 (UNI0), a single prediction partition 1 (UNI1), and a dual prediction partition 2 (BI) are assigned to three partitions. Here, the sub-blocks above the diagonal are assigned to partition 0, the sub-blocks below the diagonal are assigned to partition 1 , and the sub-blocks on the diagonal are respectively assigned to partition 2. If merge_triangle_s plit_dir is 0, the partitions are assigned as shown in FIG. 65(a), and if merge_tr iangle_split_dir is 1, the partitions are assigned as shown in FIG. 65(b) . .
[0322] For the motion compensation prediction of partition 0, the motion information of the single prediction specified by the merge triangle index 0 is used. For the motion compensation prediction of partition 1, the motion information of the single prediction specified by the merge triangle index 1 is used. For the motion compensation prediction of partition 2, the motion information of the single prediction specified by the merge triangle index 0 and the merge triangle index 1 are used. Motion information of double prediction that combines the motion information of single prediction specified by
[0323] Here, the motion information of single prediction is a pair of a motion vector and a reference index, and the motion information of double prediction is composed of two pairs of a motion vector and a reference index. Also, the motion information refers to the motion information of single prediction or double prediction.
[0324] The merge candidate selection units 347 and 447 use the derived merge candidate list mergeCandList as the triangular merge candidate list triangleMergeCandList.
[0325] The flowchart of FIG. 66 regarding the derivation of triangular merge candidates will be described.
[0326] First, use the merge candidate list mergeCandList as the triangular merge candidate list triangleMergeCandList (step S3501). The number of candidates numTriangleMerge Cand in the triangular merge candidate list is set to the same value as the number of merge candidates numCurrMergeCand.
[0327] Next, derive the motion information of single prediction for the merge triangular partition (step S3502) .
[0328] FIG. 67 is a flowchart for explaining the derivation of the motion information of single prediction for the merge triangular partition of the present embodiment.
[0329] In the present embodiment, the derivation of the motion information of single prediction with the same priority order in merge triangular partition 0 and merge triangular partition 1 is derived to reduce the processing load.
[0330] First, for the M-th candidate in the derived merge candidate list mergeCandList, candidate M is determined whether it has the motion information in the motion information list L0 (step S3601). Candidate If M has the motion information in the motion information list L0, the motion information of the motion information list L0 of candidate M is used as a triangular merge candidate (step S3602).
[0331] Subsequently, for the M-th candidate in the derived merge candidate list mergeCandList, candidate M is determined whether it has the motion information in the motion information list L1 (step S3603). Can If didate M has the motion information in the motion information list L1, the motion of the motion information list L1 of candidate M information is used as a triangular merge candidate (step S3604). For candidate M (M = numMergeCand - 1,... , 1, 0), steps S3601, S3602, step S36 03, and step S3604 are performed in descending order to additionally derive triangular merge candidates.
[0332] FIG. 68 is a diagram for explaining an example of the motion information of the triangular merge candidate of the present embodiment.
[0333] FIG. 68(a) shows an example of a merge candidate list. The merge candidate with merge index 0 has an inter prediction mode of dual prediction (Pred - BI), the motion information of the motion information list L0 is MV0 _L0, and the motion information of the motion information list L1 is MV0_L1. The merge candidate with merge index 1 has an inter prediction mode of single prediction (Pred - L0), the motion information of the motion information list L0 is MV1_L0, and it does not have the motion information of the motion information list L1. The merge candidate with merge index 2 has an inter prediction mode of single prediction (Pred - L1), the motion information of the motion information list L0 There is none, and the motion information in the motion information list L1 is MV2_L1. The merge candidate for merge index 3 is that the inter prediction mode is dual prediction (Pred - BI), and the motion information in the motion information list L0 is MV3_L0, and the motion information in the motion information list L1 is MV3_L1. The merge candidate for merge index 4 is that the inter prediction mode is single prediction (Pred - L0), the motion information in the motion information list L0 is MV4_L0, and there is no motion information in the motion information list L1. The merge candidate for merge index 3 is that the inter prediction mode is dual prediction (Pred - BI), and the motion information in the motion information list L0 is MV3_L0, and the motion information in the motion information list L1 is MV3_L1. The merge candidate for merge index 4 is that the inter prediction mode is single prediction (Pred - L0), the motion information in the motion information list L0 is MV4_L0, and there is no motion information in the motion information list L1. The merge candidate for merge index 4 is that the inter prediction mode is single prediction (Pred - L0), the motion information in the motion information list L0 is MV4_L0, and there is no motion information in the motion information list L1. There is none, and the motion information in the motion information list L1 is MV2_L1. The merge candidate for merge index 3 is that the inter prediction mode is dual prediction (Pred - BI), and the motion information in the motion information list L0 is MV3_L0, and the motion information in the motion information list L1 is MV3_L1. The merge candidate for merge index 4 is that the inter prediction mode is single prediction (Pred - L0), the motion information in the motion information list L0 is MV4_L0, and there is no motion information in the motion information list L1.
[0334] Figure 68(b) is a diagram for explaining the relationship between the merge triangle index and the merge triangle partition in the case of the example of the merge candidate list in Figure 68(a). Figure 68(b) is a diagram for explaining the relationship between the merge triangle index and the merge triangle partition in the case of the example of the merge candidate list in Figure 68(a).
[0335] The single - prediction motion information candidates for merge triangle partition 0 are composed of motion information in the order of MV0_L0, MV0_L1, MV1_L0, MV2_L1, MV3_L0. The single - prediction motion information candidates for merge triangle partition 0 are composed of motion information in the order of MV0_L0, MV0_L1, MV1_L0, MV2_L1, MV3_L0.
[0336] Similarly, the single - prediction motion information candidates for merge triangle partition 1 are also composed of motion information in the order of MV0_L0, MV0_L1, MV1_L0, MV2_L1, MV3_L0, (MV3_L1). However, in order to prevent the motion information of merge triangle partition 0 from being the same as that of merge triangle partition 1, the single - prediction motion information selected by merge triangle index 0 (merge_triangle_idx0) is excluded, Similarly, the single - prediction motion information candidates for merge triangle partition 1 are also composed of motion information in the order of MV0_L0, MV0_L1, MV1_L0, MV2_L1, MV3_L0, (MV3_L1). However, in order to prevent the motion information of merge triangle partition 0 from being the same as that of merge triangle partition 1, the single - prediction motion information selected by merge triangle index 0 (merge_triangle_idx0) is excluded, and merge triangle index 1 (merge_triangle_idx1) is derived. and merge triangle index 1 (merge_triangle_idx1) is derived. and merge triangle index 1 (merge_triangle_idx1) is derived.
[0337] In this way, by making the single - prediction motion information candidates of merge triangle partition 0 and merge triangle partition 1 the same as the priority order of the merge list candidate list, an efficient triangle In this way, by making the single - prediction motion information candidates of merge triangle partition 0 and merge triangle partition 1 the same as the priority order of the merge list candidate list, an efficient triangle The merge mode can be transmitted with a small amount of code. Also, the motion information of merge triangle partition 0 and the motion information of merge triangle partition 1 are not made the same, so that the merge triangle By transmitting index 0 and merge triangle index 1, the motion information of merge triangle partition 0 and the motion information of merge triangle partition 1 that do not need to be in the triangular merge mode and are the same are eliminated, and the triangular merge mode can be transmitted with a small amount of code is achieved. is achieved. .
[0338] In all of the embodiments described above, the encoded bit stream output by the image encoding apparatus is specified so that it can be decoded according to the encoding method used in the embodiment to have a data format. The encoded bit stream may be recorded and provided on a recording medium readable by a computer such as an HDD, SSD, flash memory, optical disk, etc., or may be provided from a server through a wired or wireless network. Accordingly, an image decoding apparatus corresponding to this image encoding apparatus can decode the encoded bit stream of this specific data format regardless of the providing means. or may be provided from a server through a wired or wireless network. Accordingly, an image decoding apparatus corresponding to this image encoding apparatus can decode the encoded bit stream of this specific data format regardless of the providing means. When a wired or wireless network is used to exchange the encoded bit stream between the image encoding apparatus and the image decoding apparatus, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. When a wired or wireless network is used to exchange the encoded bit stream between the image encoding apparatus and the image decoding apparatus, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. is achieved.
[0339] In order to exchange the encoded bit stream between the image encoding apparatus and the image decoding apparatus, when a wired or wireless network is used, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. when a wired or wireless network is used, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. when a wired or wireless network is used, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. when a wired or wireless network is used, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. when a wired or wireless network is used, the encoded bit stream may be converted into a data format suitable for the transmission form of the communication path and transmitted. In that case, a transmission apparatus that converts the encoded bit stream output by the image encoding apparatus into encoded data in a data format suitable for the transmission form of the communication path and transmits it to the network, and a reception apparatus that receives the encoded data from the network and restores it to the encoded bit stream and supplies it to the image decoding apparatus are provided. The communication device includes a memory that buffers the encoded bitstream output by the image encoding device, a packet processing unit that packetizes the encoded bitstream, and a transmission unit that transmits the packetized encoded data via a network. The receiving device includes a receiving unit that receives the packetized encoded data via a network, a memory that buffers the received encoded data, and a packet processing unit that performs packet processing on the encoded data to generate an encoded bitstream and provides it to the image decoding device. When a wired or wireless network is used to exchange the encoded bitstream between the image encoding device and the image decoding device, in addition to the transmitting device and the receiving device, a relay device that receives the encoded data transmitted by the transmitting device and supplies it to the receiving device may be provided. The relay device includes a receiving unit that receives the packetized encoded data transmitted by the transmitting device, a memory that buffers the received encoded data, and a transmission unit that transmits the packetized encoded data to the network. Furthermore, the relay device may include a receiving packet processing unit that performs packet processing on the packetized encoded data to generate an encoded bitstream, a recording medium that stores the encoded bitstream, and a transmitting packet processing unit that packetizes the encoded bitstream.
[0340] By adding a display unit that displays the image decoded by the image decoding device to the configuration, it can also be used as a display device. In that case, the display unit reads out the decoded image signal generated by the decoded image signal superimposing unit 207 and stored in the decoded image memory 208 and displays it on the screen.
[0341]
[0342] In addition, by adding an imaging unit to the configuration and inputting the captured image into the image encoding device, it can also be used as an imaging device. In that case, the imaging unit inputs the captured image signal to the block division unit 101. Apply force.
[0343] The above-mentioned processing related to encoding and decoding may of course be realized by using transmission, storage, and reception devices using hardware, and may also be realized by firmware stored in a ROM (Read-Only Memory), flash memory, etc., or by software such as a computer. The firmware program and software program may be recorded on a computer readable recording medium and provided, or may be provided from a server through a wired or wireless network , or may be provided as data broadcast of terrestrial or satellite digital broadcast.
[0344] As described above, the present invention has been described based on the embodiments. The embodiments are examples, and various modifications are possible for each combination of their constituent elements and each processing process, and it is understood by those skilled in the art that such modifications are also within the scope of the present invention.
Explanation of Signs
[0345] 100 Image encoding device, 101 Block division unit, 102 Inter prediction unit, 103 Intra prediction unit, 104 Decoded image memory, 105 Prediction method determination unit, 10 6 Residual signal generation unit, 107 Orthogonal transformation / quantization unit, 108 Bit sequence encoding unit, 109 Inverse quantization / inverse orthogonal transformation unit, 110 Decoded image signal superposition unit, 111 Encoded information Storage memory, 200 Image decoding device, 201 Bit sequence decoding unit, 202 Block CU splitting unit, 203 Inter prediction unit, 204 Intra prediction unit, 205 Encoding information Storage memory, 206 Inverse quantization and inverse orthogonal transform unit, 207 Decoded image signal superposition unit, 2 08 Decoded image memory.< / poc>
Claims
1. A moving image encoding apparatus using a merge mode, comprising: a merge candidate list construction unit that constructs a merge candidate list including spatial merge candidates; a normal merge candidate selection unit that selects a normal merge candidate that is single prediction or dual prediction from the merge candidate list; a triangular merge candidate selection unit that selects a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list; an encoding unit that encodes a first index for specifying the first triangular merge candidate and a second index for specifying the second triangular merge candidate; The triangular merge candidate selection unit derives the first triangular merge candidate from the merge candidate list including dual-prediction merge candidates using the first index, and derives the second triangular merge candidate from the merge candidate list including dual-prediction merge candidates using the second index. The triangular merge candidate selection unit uses the first index and the second index to first determine whether L0 motion information is included in the candidates in the merge candidate list. If L0 motion information is included, the L0 motion information is used as a triangular merge candidate. Next, it is determined whether L1 motion information is included in the candidates in the merge candidate list. If L1 motion information is included, the L1 motion information is used as a triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction unit adds two or more merge candidates to the merge candidate list. A moving image encoding apparatus characterized by the above.
2. A moving image encoding method using a merge mode, comprising: a merge candidate list construction step of constructing a merge candidate list including spatial merge candidates; a normal merge candidate selection step of selecting a normal merge candidate that is single prediction or dual prediction from the merge candidate list; a triangular merge candidate selection step of selecting a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list; an encoding step of encoding a first index for specifying the first triangular merge candidate and a second index for specifying the second triangular merge candidate; The triangular merge candidate selection step derives the first triangular merge candidate from the merge candidate list including dual-prediction merge candidates using the first index, and derives the second triangular merge candidate from the merge candidate list including dual-prediction merge candidates using the second index. The triangular merge candidate selection step uses the first index and the second index to first determine whether L0 motion information is included in the candidates in the merge candidate list. If L0 motion information is included, the L0 motion information is used as a triangular merge candidate. Next, it is determined whether L1 motion information is included in the candidates in the merge candidate list. If L1 motion information is included, the L1 motion information is used as a triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The triangular merge candidate selection step derives the first triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the first index, and derives the second triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the second index. The triangular merge candidate selection step uses the first index and the second index to first determine whether the candidate in the merge candidate list includes L0 motion information. If the L0 motion information is included, the L0 motion information is used as the triangular merge candidate. Next, it is determined whether the candidate in the merge candidate list includes L1 motion information. If the L1 motion information is included, the L1 motion information is used as the triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction step adds two or more merge candidates to the merge candidate list. A moving image encoding method characterized by the above.
3. A moving image encoding program using a merge mode, comprising: a merge candidate list construction step of constructing a merge candidate list including spatial merge candidates; a normal merge candidate selection step of selecting normal merge candidates that are single prediction or dual prediction from the merge candidate list; a triangular merge candidate selection step of selecting a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list; an encoding step of encoding a first index for specifying the first triangular merge candidate and a second index for specifying the second triangular merge candidate; and The triangular merge candidate selection step derives the first triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the first index, and derives the second triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the second index. The triangular merge candidate selection step uses the first index and the second index to first determine whether the candidate in the merge candidate list contains L0 motion information. If the L0 motion information is included, the L0 motion information is used as the triangular merge candidate. Next, it is determined whether the candidate in the merge candidate list contains L1 motion information. If the L1 motion information is included, the L1 motion information is used as the triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction step adds two or more merge candidates to the merge candidate list. A moving image encoding program characterized by the above.
4. A moving image decoding apparatus using a merge mode, a decoding unit that decodes a first index for specifying a first triangular merge candidate and a second index for specifying a second triangular merge candidate, a merge candidate list construction unit that constructs a merge candidate list including spatial merge candidates, a normal merge candidate selection unit that selects a normal merge candidate that is single prediction or double prediction from the merge candidate list, a triangular merge candidate selection unit that selects the first triangular merge candidate that is single prediction and the second triangular merge candidate that is single prediction from the merge candidate list, comprising: The triangular merge candidate selection unit derives the first triangular merge candidate using the first index from the merge candidate list including double-prediction merge candidates, and derives the second triangular merge candidate using the second index from the merge candidate list including double-prediction merge candidates. The triangular merge candidate selection unit uses the first index and the second index to first determine whether the candidate in the merge candidate list contains L0 motion information. If the L0 motion information is included, the L0 motion information is used as the triangular merge candidate. Next, it is determined whether the candidate in the merge candidate list contains L1 motion information. If the L1 motion information is included, the L1 motion information is used as the triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction unit adds two or more merge candidates to the merge candidate list. A moving image decoding apparatus characterized by the above.
5. A moving image decoding method using a merge mode, A decoding step of decoding a first index that identifies a first triangular merge candidate and a second index that identifies a second triangular merge candidate; A merge candidate list construction step of constructing a merge candidate list including spatial merge candidates; A normal merge candidate selection step of selecting a normal merge candidate that is single prediction or double prediction from the merge candidate list; A triangular merge candidate selection step of selecting a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list; Comprising: In the triangular merge candidate selection step, the first triangular merge candidate is derived from the merge candidate list including double prediction merge candidates using the first index, and the second triangular merge candidate is derived from the merge candidate list including double prediction merge candidates using the second index; In the triangular merge candidate selection step, using the first index and the second index, first it is determined whether L0 motion information is included in the candidates in the merge candidate list, and if L0 motion information is included, the L0 motion information is used as a triangular merge candidate, then it is determined whether L1 motion information is included in the candidates in the merge candidate list, and if L1 motion information is included, the L1 motion information is used as a triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate; The merge candidate list construction step adds two or more merge candidates to the merge candidate list A moving image decoding method characterized by the above.
6. A moving image decoding program using a merge mode, A decoding step of decoding a first index that identifies a first triangular merge candidate and a second index that identifies a second triangular merge candidate; A merge candidate list construction step of constructing a merge candidate list including spatial merge candidates; A normal merge candidate selection step of selecting a normal merge candidate that is single prediction or double prediction from the merge candidate list; A triangular merge candidate selection step of selecting a first triangular merge candidate that is single prediction and a second triangular merge candidate that is single prediction from the merge candidate list; Comprising: The triangular merge candidate selection step derives the first triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the first index, and derives the second triangular merge candidate from the merge candidate list including the merge candidates of dual prediction using the second index. The triangular merge candidate selection step uses the first index and the second index to first determine whether L0 motion information is included in the candidates in the merge candidate list. If L0 motion information is included, the L0 motion information is used as a triangular merge candidate. Next, it is determined whether L1 motion information is included in the candidates in the merge candidate list. If L1 motion information is included, the L1 motion information is used as a triangular merge candidate, thereby selecting the first triangular merge candidate and the second triangular merge candidate. The merge candidate list construction step adds two or more merge candidates to the merge candidate list. A moving image decoding program characterized by the above.
7. A storage method for storing a bitstream generated by the moving image encoding method according to claim 2 in a recording medium.
8. A transmission method for transmitting a bitstream generated by the moving image encoding method according to claim 2.
Citation Information
Patent Citations
Video encoding device, video encoding method, and video encoding program
JP7679899B2
Position dependent storage of motion information
WO2020094078A1
An encoder, a decoder and corresponding methods for merge mode
WO2020106189A1
Video coding with triangular shape prediction units
WO2020139903A1
Method and apparatus for video coding
WO2020142378A1