Image encoding apparatus and method, image decoding apparatus and method, storage medium
Patent Information
- Application Number
- CN202311401664.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2019-12-19
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2039-12-19
AI Technical Summary
[0016] According to the present invention, high-efficiency image encoding/decoding processing can be achieved with low overhead.
Smart Images

Figure CN117221605B_ABST
Abstract
Description
[0001] This divisional application is a divisional application of the invention patent with application number 202110324793.2, application date of December 19, 2019, and applicant being Intellectual Property Bridge No. 1 Co., Ltd. The invention patent application is entitled "Image Encoding Apparatus and Method, Image Decoding Apparatus and Method, Storage Medium". Technical Field
[0002] This invention relates to image encoding and decoding techniques for segmenting images into blocks and making predictions. Background Technology
[0003] In image encoding and decoding, the image to be processed is divided into blocks, which are sets of a predetermined number of pixels, and processed on a block-by-block basis. By dividing the image into appropriate blocks and appropriately setting intra-frame prediction and inter-frame prediction, encoding efficiency is improved.
[0004] In the encoding / decoding of moving images, encoding efficiency is improved by performing inter-frame prediction based on already encoded / decoded images. Patent Document 1 discloses a technique that applies affine transformation during inter-frame prediction. In moving images, objects often undergo deformations such as magnification, reduction, and rotation; by applying the technique in Patent Document 1, efficient encoding can be achieved.
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent document 1: Japanese Patent Application Publication No. 9-172644. Summary of the Invention
[0008] However, the technology in Patent Document 1 involves image transformations, resulting in a high processing load. In view of the above problems, this invention provides a low-load and efficient encoding technology.
[0009] To address the aforementioned problems, the image encoding apparatus of the first aspect of the present invention includes: an encoding information storage unit that stores inter-frame prediction information used in inter-frame prediction of encoded blocks into a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation unit that derives spatial inter-frame prediction information candidates from inter-frame prediction information of blocks spatially adjacent to the encoding target block and uses them as inter-frame prediction information candidates for the encoding target block; and a historical inter-frame prediction information candidate derivation unit that derives historical inter-frame prediction information candidates from inter-frame prediction information stored in the historical prediction motion vector candidate list and uses them as inter-frame prediction information candidates for the encoding target block. The historical inter-frame prediction information candidate derivation unit compares a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, with the spatial inter-frame prediction information candidates, and when the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0010] The second aspect of the image coding method of the present invention includes: a coding information saving step, which saves the inter-frame prediction information used in the inter-frame prediction of the coded block to a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation step, which derives spatial inter-frame prediction information candidates from the inter-frame prediction information of blocks spatially adjacent to the coding target block, and uses them as inter-frame prediction information candidates for the coding target block; and a historical inter-frame prediction information candidate derivation step, which derives historical inter-frame prediction information candidates from the inter-frame prediction information saved in the historical prediction motion vector candidate list, and uses them as inter-frame prediction information candidates for the coding target block. In the historical inter-frame prediction information candidate derivation step, a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, are compared with the spatial inter-frame prediction information candidates. When the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0011] The third aspect of the image encoding procedure of the present invention causes a computer to perform the following steps: an encoding information saving step, which saves the inter-frame prediction information used in the inter-frame prediction of the encoded block to a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation step, which derives spatial inter-frame prediction information candidates from the inter-frame prediction information of blocks spatially adjacent to the coded object block, and uses them as inter-frame prediction information candidates for the coded object block; and a historical inter-frame prediction information candidate derivation step, which derives historical inter-frame prediction information candidates from the inter-frame prediction information saved in the historical prediction motion vector candidate list, and uses them as inter-frame prediction information candidates for the coded object block. In the historical inter-frame prediction information candidate derivation step, a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, are compared with the spatial inter-frame prediction information candidates. When the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0012] The image decoding apparatus of the fourth aspect of the present invention includes: a decoding information storage unit that stores inter-frame prediction information used in inter-frame prediction of decoded blocks into a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation unit that derives spatial inter-frame prediction information candidates from inter-frame prediction information of blocks spatially adjacent to the target block and uses them as inter-frame prediction information candidates for the target block; and a historical inter-frame prediction information candidate derivation unit that derives historical inter-frame prediction information candidates from inter-frame prediction information stored in the historical prediction motion vector candidate list and uses them as inter-frame prediction information candidates for the target block. The historical inter-frame prediction information candidate derivation unit compares a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, with the spatial inter-frame prediction information candidates, and when the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0013] The fifth aspect of the image decoding method of the present invention includes: a decoding information saving step, which saves the inter-frame prediction information used in the inter-frame prediction of the decoded block to a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation step, which derives spatial inter-frame prediction information candidates from the inter-frame prediction information of blocks spatially adjacent to the decoded object block and uses them as inter-frame prediction information candidates for the decoded object block; and a historical inter-frame prediction information candidate derivation step, which derives historical inter-frame prediction information candidates from the inter-frame prediction information saved in the historical prediction motion vector candidate list and uses them as inter-frame prediction information candidates for the decoded object block. In the historical inter-frame prediction information candidate derivation step, a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, are compared with the spatial inter-frame prediction information candidates. When the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0014] The image decoding program of the sixth aspect of the present invention causes a computer to perform the following steps: a decoding information saving step, which saves the inter-frame prediction information used in the inter-frame prediction of the decoded block to a historical prediction motion vector candidate list; a spatial inter-frame prediction information candidate derivation step, which derives spatial inter-frame prediction information candidates from the inter-frame prediction information of blocks spatially adjacent to the decoded object block and uses them as inter-frame prediction information candidates for the decoded object block; and a historical inter-frame prediction information candidate derivation step, which derives historical inter-frame prediction information candidates from the inter-frame prediction information saved in the historical prediction motion vector candidate list and uses them as inter-frame prediction information candidates for the decoded object block. In the historical inter-frame prediction information candidate derivation step, a predetermined number of inter-frame prediction information stored in the historical prediction motion vector candidate list, starting from the latest inter-frame prediction information, are compared with the spatial inter-frame prediction information candidates. When the values of the inter-frame prediction information are different, they are used as historical inter-frame prediction information candidates.
[0015] Invention Effects
[0016] According to the present invention, high-efficiency image encoding / decoding processing can be achieved with low overhead. Attached Figure Description
[0017] Figure 1 This is a block diagram of an image encoding apparatus according to an embodiment of the present invention;
[0018] Figure 2 This is a block diagram of an image decoding apparatus according to an embodiment of the present invention;
[0019] Figure 3 This is a flowchart illustrating the actions of splitting tree blocks;
[0020] Figure 4This is a diagram illustrating the process of segmenting an input image into tree blocks;
[0021] Figure 5 This is a diagram illustrating the z-scan.
[0022] Figure 6A This is a diagram showing the segmented shape of the block;
[0023] Figure 6B This is a diagram showing the segmented shape of the block;
[0024] Figure 6C This is a diagram showing the segmented shape of the block;
[0025] Figure 6D This is a diagram showing the segmented shape of the block;
[0026] Figure 6E This is a diagram showing the segmented shape of the block;
[0027] Figure 7 This is a flowchart illustrating the action of dividing a block into four parts;
[0028] Figure 8 This is a flowchart used to illustrate the action of dividing a block into 2 or 3 parts;
[0029] Figure 9 It is a syntax used to describe the shape of block segmentation;
[0030] Figure 10A This is a diagram used to illustrate intra-frame prediction;
[0031] Figure 10B This is a diagram used to illustrate intra-frame prediction;
[0032] Figure 11 This is a diagram used to illustrate the reference block for inter-frame prediction;
[0033] Figure 12 Syntax used to describe the prediction pattern of coded blocks;
[0034] Figure 13 It is a graph showing the correspondence between syntactic elements and patterns related to inter-frame prediction;
[0035] Figure 14 This is a diagram used to illustrate affine transformation motion compensation with two control points;
[0036] Figure 15 This is a diagram used to illustrate affine transformation motion compensation with three control points;
[0037] Figure 16 yes Figure 1 A block diagram showing the detailed structure of the inter-frame prediction unit 102;
[0038] Figure 17 yes Figure 16 A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 301;
[0039] Figure 18 yes Figure 16 A block diagram showing the detailed structure of the typical merging mode derivation section 302;
[0040] Figure 19 It is used for explanation Figure 16 A flowchart of the normal prediction motion vector pattern export process of the normal prediction motion vector pattern export unit 301.
[0041] Figure 20 This is a flowchart illustrating the typical processing steps for deriving motion vector patterns from predicted data.
[0042] Figure 21 This is a flowchart illustrating the typical steps involved in exporting data using a merge pattern.
[0043] Figure 22 yes Figure 2 A block diagram showing the detailed structure of the inter-frame prediction unit 203;
[0044] Figure 23 yes Figure 22 A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 401;
[0045] Figure 24 yes Figure 22 A block diagram showing the detailed structure of the typical merging mode derivation section 402;
[0046] Figure 25 It is used for explanation Figure 22 A flowchart of the normal prediction motion vector pattern export process of the normal prediction motion vector pattern export unit 401.
[0047] Figure 26 This is a diagram illustrating the initialization / update process of the historical predicted motion vector candidate list;
[0048] Figure 27 This is a flowchart of the same feature confirmation process in the initialization / update process of the historical predicted motion vector candidate list.
[0049] Figure 28 This is a flowchart of the feature shifting process in the initialization / update process of the historical predicted motion vector candidate list.
[0050] Figure 29 This is a flowchart illustrating the steps involved in deriving candidate motion vectors from historical predictions.
[0051] Figure 30This is a flowchart illustrating the steps involved in exporting historical merge candidates;
[0052] Figure 31A This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0053] Figure 31B This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0054] Figure 31C This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0055] Figure 32 This is a graph used to illustrate motion compensation prediction when the reference image (RefL0Pic) of L0 is at a time before the processing object image (CurPic) in L0 prediction;
[0056] Figure 33 This is a diagram used to illustrate motion compensation prediction when the reference image for L0 prediction is at a time after the image of the object being processed in L0 prediction.
[0057] Figure 34 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction is before the object image being processed, and the reference image for L1 prediction is after the object image being processed.
[0058] Figure 35 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction and the reference image for L1 prediction are at a time before the object image being processed in dual prediction.
[0059] Figure 36 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction and the reference image for L1 prediction are at a time after the object image being processed in dual prediction.
[0060] Figure 37 This is a diagram illustrating an example of the hardware structure of the encoding / decoding apparatus according to an embodiment of the present invention;
[0061] Figure 38A This is a diagram illustrating an example of the elements of the historical predicted motion vector candidate list when the encoded block of the encoding / decoding object is the upper right block during 4-segmentation of the block;
[0062] Figure 38B This is a diagram illustrating an example of the elements of the historical predicted motion vector candidate list when the encoded block of the encoding / decoding object is the lower left block during 4-segmentation of the block;
[0063] Figure 38C This is a diagram illustrating an example of the elements of the historical predicted motion vector candidate list when the encoded block of the encoding / decoding object is the lower right block during 4-segmentation of the block;
[0064] Figure 38D It is a diagram used to illustrate the confirmation and comparison of elements for a candidate list of historical predicted motion vectors;
[0065] Figure 39 This is a flowchart illustrating the historical merging candidate derivation processing steps of the second embodiment of the present invention;
[0066] Figure 40 This is a flowchart of the same element confirmation process step in the historical predicted motion vector candidate list initialization / update processing step of the third embodiment of the present invention;
[0067] Figure 41 This is a flowchart of the same element confirmation process step in the historical predicted motion vector candidate list initialization / update processing step of the fourth embodiment of the present invention. Detailed Implementation
[0068] Define the technologies and technical terms used in this embodiment.
[0069] <Tree Block>
[0070] In this implementation, the image to be encoded / decoded is divided equally into segments of a predetermined size. This unit is defined as a tree block. Figure 4 In this design, the size of the tree block is set to 128×128 pixels, but the size is not limited to this and can be set to any size. The tree blocks, which are processed (corresponding to encoding objects in encoding processing and decoding objects in decoding processing), are switched according to the raster scan order, i.e., from left to right and from top to bottom. Each tree block can be further recursively divided internally. The block that becomes the encoding / decoding object after recursively dividing the tree block is defined as an encoding block. Furthermore, tree blocks and encoding blocks are collectively referred to as blocks. By performing appropriate block division, efficient encoding can be achieved. The size of the tree block can be a predetermined fixed value set in the encoding and decoding devices, or a structure can be adopted where the size of the tree block determined by the encoding device is transmitted to the decoding device. Here, the maximum size of the tree block is set to 128×128 pixels, and the minimum size is set to 16×16 pixels. Similarly, the maximum size of the encoding block is set to 64×64 pixels, and the minimum size is set to 4×4 pixels.
[0071] <Prediction Pattern>
[0072] The intra-frame prediction (MODE_INTRA) and inter-frame prediction (MODE_INTER) are switched on a per-processing-object-coded-block basis. The intra-frame prediction (MODE_INTRA) predicts based on the processed image signal of the object image, while the inter-frame prediction predicts based on the image signal of the processed image.
[0073] The processed image is used to decode the signal that has been encoded in the encoding process to obtain the image, image signal, tree block, block, coded block, etc., and is used to decode the image, image signal, tree block, block, coded block, etc. in the decoding process.
[0074] The mode that identifies the intra-frame prediction (MODE_INTRA) and inter-frame prediction (MODE_INTER) is defined as the prediction mode (PredMode). The prediction mode (PredMode) is represented by a value for either intra-frame prediction (MODE_INTRA) or inter-frame prediction (MODE_INTER).
[0075] Inter-frame prediction
[0076] In inter-frame prediction, which predicts based on image signals from processed images, multiple processed images can be used as reference pictures. To manage these multiple reference pictures, two reference lists, L0 (reference list 0) and L1 (reference list 1), are defined, each using a reference index to determine the reference picture. In a P-slice, L0 prediction (Pred_L0) can be used. In a B-slice, L0 prediction (Pred_L0), L1 prediction (Pred_L1), and double prediction (Pred_BI) can be used. L0 prediction (Pred_L0) is an inter-frame prediction that references the reference pictures managed by L0, and L1 prediction (Pred_L1) is an inter-frame prediction that references the reference pictures managed by L1. Double prediction (Pred_BI) is an inter-frame prediction that simultaneously performs L0 and L1 predictions and references individual reference pictures managed by each of L0 and L1. The information determining L0 prediction, L1 prediction, and double prediction is defined as the inter-frame prediction mode. Regarding the constants and variables with the subscript LX appended to the output in subsequent processing, the premise is that they will be processed as L0 and L1.
[0077] <Predicting motion vector patterns>
[0078] The predicted motion vector mode is a mode that transmits the index, differential motion vector, inter-frame prediction mode, reference index, and inter-frame prediction information for determining the predicted motion vector and deciding on the target block. The predicted motion vector is derived from the predicted motion vector candidate and the index used to determine the predicted motion vector. The predicted motion vector candidate is derived from the processed blocks adjacent to the target block or from blocks in the processed image that are located at the same position as or near the target block.
[0079] <Merge Mode>
[0080] The merging mode is as follows: without transmitting differential motion vectors or reference indices, the inter-frame prediction information of the processing object block is derived based on the inter-frame prediction information of processed blocks adjacent to the processing object block or blocks in the processed image that are located at the same position or nearby (neighboring) to the processing object block.
[0081] Processed blocks adjacent to the target block, along with their inter-frame prediction information, are defined as spatial merge candidates. Blocks belonging to the processed image that are located at the same position as or near the target block, along with their inter-frame prediction information derived from that block, are defined as temporal merge candidates. Each merge candidate is registered in a merge candidate list, and the merge candidate used in the prediction of the target block is determined by a merge index.
[0082] <Adjacent blocks>
[0083] Figure 11 This diagram illustrates the reference blocks used to derive inter-frame prediction information in both the predicted motion vector mode and the merge mode. A0, A1, A2, B0, B1, B2, and B3 are processed blocks adjacent to the object block. T0 is a block within the processed image blocks that is located at the same position or in the vicinity (nearby) of the object block in the object image.
[0084] A1 and A2 are blocks located to the left of the processing object's encoding block and adjacent to it. B1 and B3 are blocks located above the processing object's encoding block and adjacent to it. A0, B0, and B2 are blocks located to the lower left, upper right, and upper left of the processing object's encoding block, respectively.
[0085] The details of how adjacent blocks are handled in the prediction motion vector mode and the merging mode are described later.
[0086] <Affine Transformation Motion Compensation>
[0087] Affine transformation motion compensation involves dividing the coded block into sub-blocks of predetermined units and determining motion vectors for each sub-block individually. The motion vectors of each sub-block are derived based on one or more control points. These control points are derived from inter-frame prediction information of processed blocks adjacent to the target block or blocks within the processed image that are located at the same position or nearby (neighboring) to the target block. In this embodiment, the size of the sub-block is set to 4×4 pixels, but the size of the sub-block is not limited to this; motion vectors can also be derived in pixels.
[0088] Figure 14 An example of affine transformation motion compensation with two control points is shown. In this case, the two control points have two parameters: a horizontal component and a vertical component. Therefore, the affine transformation with two control points is called a four-parameter affine transformation. Figure 14 CP1 and CP2 are control points.
[0089] Figure 15 An example of affine transformation motion compensation with three control points is shown. In this case, the three control points have two parameters: a horizontal component and a vertical component. Therefore, the affine transformation with three control points is called a six-parameter affine transformation. Figure 15 CP1, CP2, and CP3 are control points.
[0090] Affine transformation motion compensation can be used in either the predictive motion vector mode or the merging mode. A mode in which affine transformation motion compensation is applied in the predictive motion vector mode is defined as the sub-block predictive motion vector mode, and a mode in which affine transformation motion compensation is applied in the merging mode is defined as the sub-block merging mode.
[0091] <Syntax of Inter-Frame Prediction>
[0092] use Figure 12 and Figure 13 The syntax related to inter-frame prediction is explained.
[0093] Figure 12 `merge_flag` indicates whether the processed object coding block is set to merge mode or predictive motion vector mode. `merge_affine_flag` indicates whether to apply sub-block merging mode in the processed object coding block in merge mode. `inter_affine_flag` indicates whether to apply sub-block predictive motion vector mode in the processed object coding block in predictive motion vector mode. `cu_affine_type_flag` is a flag used to determine the number of control points in sub-block predictive motion vector mode.
[0094] Figure 13The values of each syntactic element and their corresponding prediction methods are shown. `merge_flag = 1` and `merge_affine_flag = 0` correspond to the normal merge mode. The normal merge mode is a merge mode that is not a sub-block merge. `merge_flag = 1` and `merge_affine_flag = 1` correspond to the sub-block merge mode. `merge_flag = 0` and `inter_affine_flag = 0` correspond to the normal predicted motion vector mode. The normal predicted motion vector mode is a predicted motion vector merge that is not a sub-block predicted motion vector mode. `merge_flag = 0` and `inter_affine_flag = 1` correspond to the sub-block predicted motion vector mode. When `merge_flag = 0` and `inter_affine_flag = 1`, `cu_affine_type_flag` is further passed to determine the number of control points.
[0095] <poc>
[0096] The Picture Order Count (POC) is a variable associated with the picture to be encoded, and it is set to an incrementing value of 1 corresponding to the output order of the pictures. Based on the POC value, it is possible to determine whether the pictures are the same, the order of the pictures in the output sequence, and the distance between the exported pictures. For example, if two pictures have the same POC value, they can be considered the same picture. If two pictures have different POC values, the picture with the smaller POC value is determined to be the first image output, and the difference between the POC values of the two pictures represents the distance between them along the timeline.
[0097] (First Implementation)
[0098] The image encoding apparatus 100 and the image decoding apparatus 200 according to the first embodiment of the present invention will be described.
[0099] Figure 1 This is a block diagram of the image encoding apparatus 100 according to the first embodiment. The image encoding apparatus 100 of the embodiment includes a block segmentation unit 101, an inter-frame prediction unit 102, an intra-frame prediction unit 103, a decoded image memory 104, a prediction method determination unit 105, a residual generation unit 106, an orthogonal transform / quantization unit 107, a bit string encoding unit 108, an inverse quantization / inverse orthogonal transform unit 109, a decoded image signal overlap unit 110, and an encoding information storage memory 111.
[0100] The block segmentation unit 101 recursively segments the input image to generate coded blocks. The block segmentation unit 101 includes a 4-segmentation unit and a 2-3 segmentation unit. The 4-segmentation unit segments the blocks to be segmented in both the horizontal and vertical directions, while the 2-3 segmentation unit segments the blocks to be segmented in either the horizontal or vertical direction. The block segmentation unit 101 sets the generated coded blocks as processing target coded blocks and provides the image signal of the processing target coded blocks to the inter-frame prediction unit 102, the intra-frame prediction unit 103, and the residual generation unit 106. Furthermore, the block segmentation unit 101 provides information representing the determined recursive segmentation structure to the bit string encoding unit 108. Detailed operation of the block segmentation unit 101 will be described later.
[0101] The inter-frame prediction unit 102 performs inter-frame prediction for the target coded block. Based on the inter-frame prediction information stored in the encoded information storage memory 111 and the decoded image signal stored in the decoded image memory 104, the inter-frame prediction unit 102 derives multiple candidate inter-frame prediction information, selects a suitable inter-frame prediction mode from the derived candidates, and provides the selected inter-frame prediction mode and the corresponding prediction image signal to the prediction method determination unit 105. The detailed structure and operation of the inter-frame prediction unit 102 will be described later.
[0102] The intra-prediction unit 103 performs intra-prediction of the target coded block. The intra-prediction unit 103 uses the decoded image signal stored in the decoded image memory 104 as a reference pixel, and generates a predicted image signal based on intra-prediction information such as the intra-prediction mode stored in the encoding information storage memory 111. In intra-prediction, the intra-prediction unit 103 selects a suitable intra-prediction mode from multiple intra-prediction modes and provides the selected intra-prediction mode and the predicted image signal corresponding to the selected intra-prediction mode to the prediction method determination unit 105.
[0103] Figure 10A and Figure 10B An example of intra-frame prediction is shown. Figure 10A This diagram illustrates the correspondence between the prediction direction of intra-prediction and the intra-prediction mode number. For example, intra-prediction mode 50 generates an intra-predicted image by copying reference pixels in the vertical direction. Intra-prediction mode 1 is a DC mode, which sets the pixel values of all pixels in the processed object block to the average of the reference pixels. Intra-prediction mode 0 is a Planar mode (two-dimensional mode), which generates a two-dimensional intra-prediction image based on reference pixels in both the vertical and horizontal directions. Figure 10B This is an example of an intra-prediction image generated in intra-prediction mode 40. The intra-prediction unit 103 copies the value of a reference pixel in the direction indicated by the intra-prediction mode to each pixel of the processing object block. If the reference pixel in the intra-prediction mode is not at an integer position, the intra-prediction unit 103 determines the reference pixel value by interpolation based on the reference pixel values at surrounding integer positions.
[0104] The decoded image memory 104 stores the decoded image generated by the decoded image signal overlap portion 110. The decoded image memory 104 provides the stored decoded image to the inter-frame prediction unit 102 and the intra-frame prediction unit 103.
[0105] The prediction method determination unit 105 evaluates each of intra-frame prediction and inter-frame prediction by using coding information and the amount of coding of the residual, the amount of distortion between the predicted image signal and the image signal to be processed, etc., to determine the optimal prediction mode. In the case of intra-frame prediction, the prediction method determination unit 105 provides intra-frame prediction information such as the intra-frame prediction mode as coding information to the bit string coding unit 108. In the case of inter-frame prediction merging mode, the prediction method determination unit 105 provides inter-frame prediction information such as merging index and information indicating whether it is a sub-block merging mode (sub-block merging flag) as coding information to the bit string coding unit 108. In the case of inter-frame prediction motion vector prediction mode, the prediction method determination unit 105 provides inter-frame prediction information such as inter-frame prediction mode, prediction motion vector index, L0 and L1 reference indices, differential motion vector, and information indicating whether it is a sub-block prediction motion vector mode (sub-block prediction motion vector flag) as coding information to the bit string coding unit 108. In addition, the prediction method determination unit 105 provides the determined coding information to the coding information storage memory 111. The prediction method determination unit 105 provides the predicted image signal to the residual generation unit 106 and the decoded image signal overlap unit 110.
[0106] The residual generation unit 106 generates a residual by subtracting the predicted image signal from the image signal of the object being processed, and provides it to the orthogonal transformation / quantization unit 107.
[0107] The orthogonal transformation / quantization unit 107 performs orthogonal transformation and quantization on the residual according to the quantization parameters to generate the orthogonal transformation / quantization residual, and provides the generated residual to the bit string encoding unit 108 and the inverse quantization / inverse orthogonal transformation unit 109.
[0108] In addition to information about sequences, images, stripes, and coding block units, the bit string encoding unit 108 encodes coding information corresponding to the prediction method determined by the prediction method determination unit 105 for each coding block. Specifically, the bit string encoding unit 108 encodes the prediction mode PredMode for each coding block. When the prediction mode is inter-frame prediction (MODE_INTER), the bit string encoding unit 108 encodes coding information (inter-frame prediction information) such as the flag indicating whether it is a merging mode, the sub-block merging flag, the merging index when it is a merging mode, the inter-frame prediction mode when it is not a merging mode, the predicted motion vector index, information related to the differential motion vector, and the sub-block predicted motion vector flag, according to the prescribed syntax (bit string syntax rules), to generate a first bit string. When the prediction mode is intra-frame prediction (MODE_INTRA), the bit string encoding unit encodes coding information (intra-frame prediction information) such as the intra-frame prediction mode, according to the prescribed syntax (bit string syntax rules), to generate a first bit string. Furthermore, the bit string encoding unit 108 performs entropy encoding on the orthogonal transform and quantized residuals according to a prescribed syntax to generate a second bit string. The bit string encoding unit 108 then multiplexes the first and second bit strings according to a prescribed syntax to output a bit stream.
[0109] The inverse quantization / inverse quadrature transformation unit 109 performs inverse quantization and inverse quadrature transformation on the residual provided by the quadrature transformation / quantization unit 107 to calculate the residual, and provides the calculated residual to the decoded image signal overlap unit 110.
[0110] The decoded image signal overlay unit 110 overlays the predicted image signal corresponding to the decision of the prediction method determination unit 105 with the residual obtained by inverse quantization and inverse quadrature transformation performed by the inverse quantization / inverse quadrature transformation unit 109 to generate a decoded image, which is then stored in the decoded image memory 104. Alternatively, the decoded image signal overlay unit 110 may also perform filtering processing on the decoded image to reduce block distortion and other distortions caused by encoding before storing it in the decoded image memory 104.
[0111] The encoding information storage memory 111 stores encoding information such as the prediction mode (inter-frame prediction or intra-frame prediction) determined by the prediction method determination unit 105. In the case of inter-frame prediction, the encoding information stored in the encoding information storage memory 111 includes the determined motion vector, the reference index of reference lists L0 and L1, and the historical predicted motion vector candidate list. Furthermore, in the case of inter-frame prediction merging mode, the encoding information stored in the encoding information storage memory 111 includes, in addition to the above information, inter-frame prediction information such as a merging index and information indicating whether it is a sub-block merging mode (sub-block merging flag). Furthermore, in the case of inter-frame prediction predicting motion vector mode, the encoding information stored in the encoding information storage memory 111 includes, in addition to the above information, inter-frame prediction information such as the inter-frame prediction mode, predicted motion vector index, differential motion vector, and information indicating whether it is a sub-block predicting motion vector mode (sub-block predicting motion vector flag). In the case of intra-frame prediction, the encoding information stored in the encoding information storage memory 111 includes intra-frame prediction information such as the determined intra-frame prediction mode.
[0112] Figure 2 It means and Figure 1 The image encoding apparatus is shown in the block diagram of the image decoding apparatus according to the embodiments of the present invention. The image decoding apparatus of the embodiments includes a bit string decoding unit 201, a block segmentation unit 202, an inter-frame prediction unit 203, an intra-frame prediction unit 204, an encoded information storage memory 205, an inverse quantization / inverse quadrature transformation unit 206, a decoded image signal overlap unit 207, and a decoded image memory 208.
[0113] Figure 2 Image decoding device decoding processing and Figure 1 The decoding processing is internally configured within the image encoding device, therefore Figure 2 The encoded information storage memory 205, the inverse quantization / inverse quadrature transform unit 206, the decoded image signal overlap unit 207, and the decoded image memory 208 each have structures that are consistent with... Figure 1 The functions of each structure of the image encoding device, including the encoding information storage memory 111, the inverse quantization / inverse quadrature transformation unit 109, the decoded image signal overlap unit 110, and the decoded image memory 104, are respectively defined.
[0114] The bitstream provided to the bitstream decoding unit 201 is separated according to prescribed syntax rules. The bitstream decoding unit 201 decodes the separated first bitstream to obtain information on the sequence, image, stripe, coded block unit, and coded block unit encoding information. Specifically, the bitstream decoding unit 201 decodes the prediction mode PredMode on a coded block basis. The prediction mode PredMode is determined to be either inter-frame prediction (MODE_INTER) or intra-frame prediction (MODE_INTRA). When the prediction mode is inter-frame prediction (MODE_INTER), the bitstream decoding unit 201 decodes the encoding information (inter-frame prediction information) related to the flag indicating whether it is a merge mode, the merge index in the case of merge mode, the sub-block merge flag, the inter-frame prediction mode in the case of prediction motion vector mode, the prediction motion vector index, the differential motion vector, and the sub-block prediction motion vector flag, according to the prescribed syntax. The encoding information (inter-frame prediction information) is then provided to the encoding information storage memory 205 via the inter-frame prediction unit 203 and the block segmentation unit 202. When the prediction mode is intra-prediction (MODE_INTRA), the coded information (intra-prediction information) such as the intra-prediction mode is decoded according to the prescribed syntax, and the coded information (intra-prediction information) is provided to the coded information storage memory 205 via the inter-prediction unit 203 or the intra-prediction unit 204 and the block segmentation unit 202. The bit string decoding unit 201 decodes the separated second bit string, calculates the residual after orthogonal transformation / quantization, and provides the residual after orthogonal transformation / quantization to the inverse quantization / inverse orthogonal transformation unit 206.
[0115] When the prediction mode PredMode of the coded block being processed is the prediction motion vector mode in inter-frame prediction (MODE_INTER), the inter-frame prediction unit 203 uses the coded information of the decoded image signal stored in the coded information storage memory 205 to derive multiple candidate prediction motion vectors, and registers the derived candidate prediction motion vectors in the candidate prediction motion vector list described later. The inter-frame prediction unit 203 selects the prediction motion vector corresponding to the prediction motion vector index provided by the bit string decoding unit 201 from the multiple candidate prediction motion vectors registered in the candidate prediction motion vector list, calculates the motion vector based on the differential motion vector decoded by the bit string decoding unit 201 and the selected prediction motion vector, and stores the calculated motion vector along with other coded information in the coded information storage memory 205. Here, the encoding information of the coded blocks to be provided / saved includes the prediction mode PredMode, flags indicating whether L0 and L1 prediction are used (predFlagL0[xP][yP], predFlagL1[xP][yP]), reference indices for L0 and L1 (refIdxL0[xP][yP], refIdxL1[xP][yP]), motion vectors for L0 and L1 (mvL0[xP][yP], mvL1[xP][yP]), etc. Here, xP and yP represent the indices of the top-left pixel position of the coded block within the image. When the prediction mode PredMode is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode is L0 prediction (Pred_L0), the flag predFlagL0 indicating whether L0 prediction is used is 1, and the flag predFlagL1 indicating whether L1 prediction is used is 0. When the inter-frame prediction mode is L1 prediction (Pred_L1), the flag predFlagL0 indicating whether to use L0 prediction is 0, and the flag predFlagL1 indicating whether to use L1 prediction is 1. When the inter-frame prediction mode is double prediction (Pred_BI), both the flag predFlagL0 indicating whether to use L0 prediction and the flag predFlagL1 indicating whether to use L1 prediction are 1. Furthermore, when the prediction mode PredMode of the coded block being processed is the merge mode in inter-frame prediction (MODE_INTER), merge candidates are derived.Using the encoding information of the decoded coded blocks stored in the encoding information storage memory 205, multiple merging candidates are derived and registered in the merging candidate list (described later). From the multiple merging candidates registered in the merging candidate list, a merging candidate corresponding to the merging index provided by the bit string decoding unit 201 is selected. Inter-frame prediction information, including the flags predFlagL0[xP][yP], predFlagL1[xP][yP] indicating whether to utilize the selected merging candidate for L0 and L1 prediction, the reference indices refIdxL0[xP][yP], refIdxL1[xP][yP] for L0 and L1, and the motion vectors mvL0[xP][yP], mvL1[xP][yP] for L0 and L1, are stored in the encoding information storage memory 205. Here, xP and yP are indices representing the position of the top-left pixel of the coded block within the image. The detailed structure and operation of the inter-frame prediction unit 203 will be described later.
[0116] When the prediction mode PredMode of the encoded block being processed is intra-prediction (MODE_INTRA), the intra-prediction unit 204 performs intra-prediction. The intra-prediction mode is included in the encoded information decoded by the bit string decoding unit 201. Based on the intra-prediction mode included in the decoded information decoded by the bit string decoding unit 201, the intra-prediction unit 204 generates a predicted image signal based on the decoded image signal stored in the decoded image memory 208, and provides the generated predicted image signal to the decoded image signal overlay unit 207. Since the intra-prediction unit 204 corresponds to the intra-prediction unit 103 of the image encoding apparatus 100, it performs the same processing as the intra-prediction unit 103.
[0117] The inverse quantization / inverse quadrature transformation unit 206 performs inverse quadrature transformation and inverse quantization on the residual after quadrature transformation / quantization decoded by the bit string decoding unit 201 to obtain the residual after inverse quadrature transformation / inverse quantization.
[0118] The decoded image signal overlay unit 207 decodes the image signal by overlaying the predicted image signal obtained by inter-frame prediction by inter-frame prediction unit 203 or intra-frame prediction by intra-frame prediction unit 204, and the residual after inverse quadrature transformation / inverse quantization by inverse quantization / inverse quadrature transformation unit 206, and stores the decoded image signal in the decoded image memory 208. When storing the image signal in the decoded image memory 208, the decoded image signal overlay unit 207 may also perform filtering processing on the decoded image to reduce block distortion caused by encoding before storing it in the decoded image memory 208.
[0119] Next, the operation of the block segmentation unit 101 in the image encoding device 100 will be explained. Figure 3 This is a flowchart illustrating the process of segmenting an image into tree blocks and further segmenting each tree block. First, the input image is segmented into tree blocks of a predetermined size (step S1001). For each tree block, it is scanned in a predetermined order, i.e., raster scan order (step S1002), and the interior of the tree block to be processed is segmented (step S1003).
[0120] Figure 7 This is a flowchart illustrating the detailed actions of the segmentation process in step S1003. First, it is determined whether to divide the block of the object to be processed into 4 parts (step S1101).
[0121] If it is determined that the processing object block 4 should be divided, the processing object block 4 is divided (step S1102). For each block obtained by dividing the processing object block, it is scanned in the Z-scan order, that is, the order of top left, top right, bottom left, and bottom right (step S1103). Figure 5 This is an example of Z-scan order. Figure 6A 601 is an example of processing object block 4 after it has been split. Figure 6A The numbers 0 to 3 in step 601 indicate the processing order. Then, for each block divided in step S1101, the process is executed recursively. Figure 7 The segmentation process (step S1104).
[0122] If it is determined that the object block to be processed will not be divided into 4 parts, then 2-3 parts will be performed (step S1105).
[0123] Figure 8 This is a flowchart showing the detailed actions of the 2-3 segmentation process in step S1105. First, it is determined whether to perform 2-3 segmentation on the block to be processed, that is, whether to perform either 2-segmentation or 3-segmentation (step S1201).
[0124] If it is determined that the block to be processed will not be split into 2-3 segments, that is, if it is determined that no splitting will be performed, the splitting process ends (step S1211). That is, for blocks obtained by recursive splitting, no further recursive splitting is performed.
[0125] If it is determined that the block of the object to be processed should be divided into 2-3, then it is determined whether to further divide the block of the object to be processed into 2 (step S1202).
[0126] If it is determined that the processing object block should be divided into two parts, it is determined whether to divide the processing object block vertically (step S1203). Based on the result, the processing object block is divided into two parts vertically (step S1204), or the processing object block is divided into two parts horizontally (step S1205). As a result of step S1204, the processing object block is as follows: Figure 6B As shown in 602, it is divided into two parts, upper and lower (vertical direction). As a result of step S1205, the processed object block is as follows: Figure 6D As shown in 604, it is divided into two parts, left and right (horizontal direction).
[0127] In step S1202, if it is not determined that the processing object block is to be divided into two parts, that is, if it is determined that it is to be divided into three parts, then it is determined whether to divide the processing object block into top, middle, and bottom (vertical direction) (step S1206). Based on this result, the processing object block is divided into three parts in the top, middle, and bottom (vertical direction) (step S1207), or the processing object block is divided into three parts in the left, middle, and right (horizontal direction) (step S1208). In the result of step S1207, the processing object block is as follows: Figure 6C As shown in 603, it is divided into three parts: upper, middle, and lower (vertical direction). In the result of step S1208, the processed object block is as follows: Figure 6E As shown in 605, it is divided into three parts: left, middle, and right (horizontal direction).
[0128] After executing any one of steps S1204, S1205, S1207, or S1208, the blocks into which the processing object block is divided are scanned in order from left to right and from top to bottom (step S1209). Figures 6B to 6E The numbers 0 to 2, from 602 to 605, indicate the processing order. For each segmented block, the process is executed recursively. Figure 8 The 2-3 segmentation process (step S1210).
[0129] The recursive block partitioning described here can also limit whether partitioning is necessary based on the number of partitions or the size of the block being processed. The information limiting whether partitioning is necessary can be implemented by pre-agreeing between the encoding and decoding devices without transmitting the information, or by the encoding device determining whether partitioning is necessary and recording the information in a bit string before transmitting it to the decoding device.
[0130] When a block is divided, the block before the division is called the parent block, and the blocks after the division are called child blocks.
[0131] Next, the operation of the block segmentation unit 202 in the image decoding apparatus 200 will be described. The block segmentation unit 202 segments tree blocks according to the same processing steps as the block segmentation unit 101 in the image encoding apparatus 100. However, the difference is that in the block segmentation unit 101 of the image encoding apparatus 100, the optimal block segmentation shape is determined by applying optimization methods such as optimal shape estimation or distortion rate optimization based on image recognition. In contrast, the block segmentation unit 202 in the image decoding apparatus 200 determines the block segmentation shape by decoding the block segmentation information recorded in the bit string.
[0132] Figure 9 The syntax (bit string syntax rules) related to block partitioning in the first embodiment is shown. `coding_quadtree()` represents the syntax involved in the 4-partitioning of the block. `multi_type_tree()` represents the syntax involved in the 2-partitioning or 3-partitioning of the block. `qt_split` is a flag indicating whether the block is 4-partitioned. When the block is 4-partitioned, `qt_split` = 1; otherwise, `qt_split` = 0. In the case of 4-partitioning (`qt_split` = 1), the 4-partitioned blocks are recursively partitioned into 4 parts (`coding_quadtree(0), coding_quadtree(1), coding_quadtree(2), coding_quadtree(3)`, where 0 to 3 correspond to... Figure 6A (601). Without a 4-splitting operation (qt_split = 0), subsequent splits are determined according to multi_type_tree(). mtt_split is a flag indicating whether further splitting is required. Furthermore, if splitting is required (mtt_split = 1), the flag indicating whether the split is vertical or horizontal is transmitted, namely mtt_split_vertical, and the flag indicating whether to perform a 2-splitting or 3-splitting operation, namely mtt_split_binary, is transmitted. mtt_split_vertical = 1 indicates a vertical split, and mtt_split_vertical = 0 indicates a horizontal split. mtt_split_binary = 1 indicates a 2-splitting operation, and mtt_split_binary = 0 indicates a 3-splitting operation. In the case of a 2-splitting operation (mtt_split_binary = 1), the blocks after the 2-splitting are recursively split (multi_type_tree(0), multi_type_tree(1), where the 0 to 1 of the independent variable corresponds to Figures 6B to 6D (Numbers 602 or 604). In the case of 3-partition (mtt_split_binary = 0), the 3-partitioned blocks are recursively split (multi_type_tree(0), multi_type_tree(1), multi_type_tree(2), 0 to 2 correspond to...). Figure 6B 603 or Figure 6E (Number 605). Hierarchical block splitting is performed by recursively calling multi_type_tree until mtt_split = 0.
[0133] Inter-frame prediction
[0134] The inter-frame prediction method in the implementation method Figure 1 The inter-frame prediction unit 102 of the image coding apparatus and Figure 2 It is implemented in the inter-frame prediction unit 203 of the image decoding device.
[0135] The inter-frame prediction method according to the implementation method is described with reference to the accompanying drawings. The inter-frame prediction method is implemented in either the encoding or decoding process on a block-by-block basis.
[0136] <Explanation of the inter-frame prediction unit 102 on the coding side>
[0137] Figure 16 It is shown Figure 1 A diagram showing the detailed structure of the inter-frame prediction unit 102 of the image coding apparatus. The normal prediction motion vector pattern derivation unit 301 derives multiple normal prediction motion vector candidates to select a prediction motion vector and calculates the difference motion vector between the selected prediction motion vector and the detected motion vector. The detected inter-frame prediction pattern, reference index, motion vector, and calculated difference motion vector constitute the inter-frame prediction information for the normal prediction motion vector pattern. This inter-frame prediction information is provided to the inter-frame prediction pattern determination unit 305. The detailed structure and processing of the normal prediction motion vector pattern derivation unit 301 will be described later.
[0138] In the normal merging mode derivation unit 302, multiple normal merging candidates are derived, and a normal merging candidate is selected to obtain inter-frame prediction information for the normal merging mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305. The detailed structure and processing of the normal merging mode derivation unit 302 will be described later.
[0139] In the sub-block prediction motion vector mode derivation unit 303, multiple sub-block prediction motion vector candidates are derived to select a sub-block prediction motion vector, and the difference motion vector between the selected sub-block prediction motion vector and the detected motion vector is calculated. The detected inter-frame prediction mode, reference index, motion vector, and calculated difference motion vector constitute the inter-frame prediction information of the sub-block prediction motion vector mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305.
[0140] In the sub-block merging mode derivation unit 304, multiple sub-block merging candidates are derived, and a sub-block merging candidate is selected to obtain inter-frame prediction information for the sub-block merging mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305.
[0141] The inter-frame prediction mode determination unit 305 determines inter-frame prediction information based on the inter-frame prediction information provided by the normal prediction motion vector mode derivation unit 301, the normal merging mode derivation unit 302, the sub-block prediction motion vector mode derivation unit 303, and the sub-block merging mode derivation unit 304. The inter-frame prediction information corresponding to the determination result is provided from the inter-frame prediction mode determination unit 305 to the motion compensation prediction unit 306.
[0142] The motion compensation prediction unit 306 performs inter-frame prediction on the reference image signal stored in the decoded image memory 104 based on the determined inter-frame prediction information. The detailed structure and processing of the motion compensation prediction unit 306 will be described later.
[0143] <Explanation of the inter-frame prediction unit 203 on the decoding side>
[0144] Figure 22 It is shown Figure 2 A diagram showing the detailed structure of the inter-frame prediction unit 203 of the image decoding device.
[0145] Normally, the predictive motion vector mode derivation unit 401 derives multiple normally predicted motion vector candidates to select a predicted motion vector, and calculates the sum of the selected predicted motion vector and the decoded differential motion vector as the motion vector. The decoded inter-frame prediction mode, reference index, and motion vector constitute the inter-frame prediction information of the normally predicted motion vector mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408. The detailed structure and processing of the normally predicted motion vector mode derivation unit 401 will be described later.
[0146] In the normal merging mode derivation unit 402, multiple normal merging candidates are derived to select a normal merging candidate, thereby obtaining inter-frame prediction information for the normal merging mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408. The detailed structure and processing of the normal merging mode derivation unit 402 will be described later.
[0147] In the sub-block prediction motion vector mode derivation unit 403, multiple sub-block prediction motion vector candidates are derived to select a sub-block prediction motion vector. The selected sub-block prediction motion vector is calculated as the sum of the selected sub-block prediction motion vector and the decoded differential motion vector, which is then used as the motion vector. The decoded inter-frame prediction mode, reference index, and motion vector constitute the inter-frame prediction information of the sub-block prediction motion vector mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408.
[0148] In the sub-block merging mode derivation unit 404, multiple sub-block merging candidates are derived to select a sub-block merging candidate, thereby obtaining inter-frame prediction information for the sub-block merging mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408.
[0149] In the motion compensation prediction unit 406, inter-frame prediction is performed on the reference image signal stored in the decoded image memory 208 based on the determined inter-frame prediction information. The detailed structure and processing of the motion compensation prediction unit 406 are the same as those of the motion compensation prediction unit 306 on the encoding side.
[0150] <Typically predicted motion vector pattern derivation part (usually AMVP)>
[0151] Figure 17 The normal prediction motion vector pattern derivation unit 301 includes a spatial prediction motion vector candidate derivation unit 321, a time prediction motion vector candidate derivation unit 322, a historical prediction motion vector candidate derivation unit 323, a prediction motion vector candidate supplementation unit 325, a normal motion vector detection unit 326, a prediction motion vector candidate selection unit 327, and a motion vector subtraction unit 328.
[0152] Figure 23 The general prediction motion vector pattern derivation unit 401 includes a spatial prediction motion vector candidate derivation unit 421, a time prediction motion vector candidate derivation unit 422, a historical prediction motion vector candidate derivation unit 423, a prediction motion vector candidate supplementation unit 425, a prediction motion vector candidate selection unit 426, and a motion vector addition unit 427.
[0153] Use respectively Figure 19 , Figure 25 The flowchart describes the processing steps of the normal prediction motion vector pattern derivation unit 301 on the encoding side and the normal prediction motion vector pattern derivation unit 401 on the decoding side. Figure 19 This is a flowchart illustrating the normal predicted motion vector pattern derivation processing steps of the normal motion vector pattern derivation unit 301 based on the encoding side. Figure 25 This is a flowchart illustrating the normal predicted motion vector pattern export processing steps of the normal motion vector pattern export unit 401 based on the decoding side.
[0154] <Commonly Predictive Motion Vector Pattern Derivation Unit (typically AMVP): Explanation of the Encoding Side>
[0155] refer to Figure 19 The typical steps for deriving predicted motion vector patterns on the encoding side are explained. Figure 19 The instructions for the processing steps sometimes omit... Figure 19 The word "usually" is shown.
[0156] First, the typical motion vector detection unit 326 detects the typical motion vector for each inter-frame prediction mode and reference index. Figure 19 Step S100).
[0157] Next, the spatial prediction motion vector candidate derivation unit 321, the temporal prediction motion vector candidate derivation unit 322, the historical prediction motion vector candidate derivation unit 323, the prediction motion vector candidate supplementation unit 325, the prediction motion vector candidate selection unit 327, and the motion vector subtraction unit 328 calculate, for each L0 and L1, the differential motion vector of the motion vector used in the inter-frame prediction of the normal prediction motion vector mode. Figure 19 Steps S101 to S106). Specifically, when the prediction mode PredMode of the processed object block is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode is L0 prediction (Pred_L0), the candidate list of predicted motion vectors for L0, mvpListL0, is calculated, the predicted motion vector mvpL0 is selected, and the differential motion vector mvdL0 of the motion vector mvL0 of L0 is calculated. When the inter-frame prediction mode of the processed object block is L1 prediction (Pred_L1), the candidate list of predicted motion vectors for L1, mvpListL1, is calculated, the predicted motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 of L1 is calculated. When the inter-frame prediction mode for processing object blocks is dual prediction (Pred_BI), L0 prediction and L1 prediction are performed simultaneously. The candidate list of predicted motion vectors for L0, mvpListL0, is calculated. The predicted motion vector of L0, mvpL0, is selected. The differential motion vector of L0, mvL0, is calculated. The candidate list of predicted motion vectors for L1, mvpListL1, is calculated. The predicted motion vector of L1, mvpL1, is calculated. The differential motion vector of L1, mvL1, is calculated.
[0158] Differential motion vector calculations are performed separately for L0 and L1, but the process is common to both. Therefore, in the following explanation, L0 and L1 will be represented as a common LX. In the calculation of the differential motion vector for L0, X of LX is 0, and in the calculation of the differential motion vector for L1, X of LX is 1. Furthermore, in the calculation of the differential motion vector for LX, if information from another list is referenced instead of LX, this other list will be represented as LY.
[0159] When using the motion vector mvLX of LX ( Figure 19 Step S102: Yes), calculate the candidate predicted motion vectors of LX, and construct the candidate predicted motion vector list mvpListLX( Figure 19 Step S103). Multiple candidate predicted motion vectors are derived from the spatial predicted motion vector candidate derivation unit 321, the temporal predicted motion vector candidate derivation unit 322, the historical predicted motion vector candidate derivation unit 323, and the predicted motion vector candidate supplementation unit 325 in the normal predicted motion vector pattern derivation unit 301, constructing a predicted motion vector candidate list mvpListLX. (About...) Figure 19 The detailed processing steps of step S103 are as follows: Figure 20 The flowchart is described later.
[0160] Next, the predicted motion vector candidate selection unit 327 selects the predicted motion vector mvpLX of LX from the predicted motion vector candidate list mvpListLX of LX. Figure 19 Step S104). Here, in the candidate list of predicted motion vectors mvpListLX, a certain element (the i-th element counting from 0) is represented as mvpListLX[i]. Calculate each differential motion vector, which is the difference between the motion vector mvLX and the candidate mvpListLX[i] of each predicted motion vector stored in the candidate list of predicted motion vectors mvpListLX. For each element (predicted motion vector candidate) in the candidate list of predicted motion vectors mvpListLX, calculate the coding amount when encoding these differential motion vectors. Then, among the elements registered in the candidate list of predicted motion vectors mvpListLX, select the candidate mvpListLX[i] of the predicted motion vector with the smallest coding amount as the predicted motion vector mvpLX, and obtain the index i. If there are multiple candidates for the predicted motion vector that will become the smallest generated code amount in the candidate list of predicted motion vectors mvpListLX, the candidate mvpListLX[i] represented by the smallest index i in the candidate list of predicted motion vectors mvpListLX is selected as the best predicted motion vector mvpLX, and that index i is obtained.
[0161] Next, the motion vector subtraction unit 328 subtracts the selected predicted motion vector mvpLX of LX from the motion vector mvLX of LX, denoted as mvdLX = mvLX - mvpLX, to calculate the differential motion vector mvdLX of LX. Figure 19 Step S105).
[0162] <Typical Predictive Motion Vector Pattern Derivation Unit (Typical AMVP): Explanation of the Decoding Side>
[0163] Next, refer to Figure 25 The typical prediction motion vector mode processing steps on the decoding side are explained. On the decoding side, the spatial prediction motion vector candidate derivation unit 421, the temporal prediction motion vector candidate derivation unit 422, the historical prediction motion vector candidate derivation unit 423, and the prediction motion vector candidate supplementation unit 425 calculate, for each L0 and L1, the motion vector used in the inter-frame prediction of the typical prediction motion vector mode. Figure 25 Steps S201 to S206). Specifically, when the prediction mode PredMode of the processing object block is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode of the processing object block is L0 prediction (Pred_L0), the candidate list of predicted motion vectors for L0, mvpListL0, is calculated, the predicted motion vector mvpL0 is selected, and the motion vector mvL0 of L0 is calculated. When the inter-frame prediction mode of the processing object block is L1 prediction (Pred_L1), the candidate list of predicted motion vectors for L1, mvpListL1, is calculated, the predicted motion vector mvpL1 is selected, and the motion vector mvL1 of L1 is calculated. When the inter-frame prediction mode for processing object blocks is dual prediction (Pred_BI), L0 prediction and L1 prediction are performed simultaneously. The candidate list of predicted motion vectors for L0, mvpListL0, is calculated. The predicted motion vector of L0, mvpL0, is selected, and the motion vector of L0, mvL0, is calculated. The candidate list of predicted motion vectors for L1, mvpListL1, is calculated, and the predicted motion vector of L1, mvpL1, is calculated. The motion vector of L1, mvL1, is calculated separately.
[0164] Similar to the encoding side, motion vector calculations are also performed on L0 and L1 separately on the decoding side, but L0 and L1 are processed in a common manner. Therefore, in the following description, L0 and L1 are denoted as a common LX. LX represents the inter-prediction mode used for inter-frame prediction of the coded block being processed. In the process of calculating the motion vector of L0, X is 0, and in the process of calculating the motion vector of L1, X is 1. In addition, in the process of calculating the motion vector of LX, if information from another reference list is referenced instead of the same reference list as the LX being calculated, this other reference list is denoted as LY.
[0165] When using the motion vector mvLX of LX ( Figure 25 Step S202: Yes), calculate the candidates for predicted motion vectors of LX, and construct the candidate list of predicted motion vectors of LX, mvpListLX( Figure 25 Step S203). Multiple candidates for predicted motion vectors are calculated from the spatial predicted motion vector candidate derivation unit 421, the temporal predicted motion vector candidate derivation unit 422, the historical predicted motion vector candidate derivation unit 423, and the predicted motion vector candidate supplementation unit 425 in the normal predicted motion vector pattern derivation unit 401, and a predicted motion vector candidate list mvpListLX is constructed. (About...) Figure 25 The detailed processing steps of step S203, using Figure 20 The flowchart is described later.
[0166] Next, the prediction motion vector candidate selection unit 426 retrieves the candidate mvpListLX[mvpIdxLX] of the prediction motion vector that corresponds to the index mvpIdxLX of the prediction motion vector provided by the bit string decoding unit 201 from the prediction motion vector candidate list mvpListLX, and selects it as the selected prediction motion vector mvpLX. Figure 25 Step S204).
[0167] Next, the motion vector addition unit 427 performs an addition operation on the differential motion vector mvdLX and the predicted motion vector mvpLX of LX provided by the bit string decoding unit 201, denoted as mvLX = mvpLX + mvdLX, to calculate the motion vector mvLX of LX. Figure 25 Step S205).
[0168] <Commonly Predictive Motion Vector Pattern Derivative (AMVP): Motion Vector Prediction Methods>
[0169] Figure 20 This is a flowchart illustrating the processing steps of the general prediction motion vector pattern derivation process, which has a common function in the general prediction motion vector pattern derivation unit 301 of the image encoding apparatus and the general prediction motion vector pattern derivation unit 401 of the image decoding apparatus according to embodiments of the present invention.
[0170] Both the normal prediction motion vector pattern derivation unit 301 and the normal prediction motion vector pattern derivation unit 401 have a prediction motion vector candidate list mvpListLX. The prediction motion vector candidate list mvpListLX forms a list structure and is provided with a storage area that stores prediction motion vector indices representing positions within the prediction motion vector candidate list and the prediction motion vector candidates corresponding to those indices as elements. The prediction motion vector indices start from 0, and the prediction motion vector candidates are stored in the storage area of the prediction motion vector candidate list mvpListLX. In this embodiment, it is assumed that the prediction motion vector candidate list mvpListLX can register at least two prediction motion vector candidates (inter-frame prediction information). Furthermore, the variable numCurrMvpCand, representing the number of prediction motion vector candidates registered in the prediction motion vector candidate list mvpListLX, is set to 0.
[0171] Spatial prediction motion vector candidate derivation units 321 and 421 derive candidates for predicted motion vectors from the block adjacent to the left. In this process, the block adjacent to the left ( Figure 11 The inter-frame prediction information (A0 or A1), including flags indicating whether the predicted motion vector candidate can be used, motion vectors, reference indices, etc., is used to derive the predicted motion vector mvLXA, and then the exported mvLXA is added to the predicted motion vector candidate list mvpListLX. Figure 20 Step S301). Additionally, X is 0 during L0 prediction and 1 during L1 prediction (the same applies below). Next, the spatial prediction motion vector candidate derivation units 321 and 421 derive candidates for predicted motion vectors from the block adjacent to the upper side. In this process, reference is made to the block adjacent to the upper side ( Figure 11 The inter-frame prediction information (B0, B1, or B2) indicates whether the predicted motion vector candidate can be used, along with the motion vector, reference index, etc., to derive the predicted motion vector mvLXB. If the derived mvLXA and mvLXB are not equal, then mvLXB is added to the predicted motion vector candidate list mvpListLX. Figure 20 Step S302). Figure 20 The processing steps S301 and S302 are the same except that the positions and number of the referenced adjacent blocks are different. They derive the flag availableFlagLXN indicating whether the predicted motion vector candidate of the coded block can be utilized, the motion vector mvLXN, and the reference index refIdxN (N represents A or B, the same below).
[0172] Next, the temporal prediction motion vector candidate derivation units 322 and 422 derive candidates for predicted motion vectors from blocks in images whose times differ from the current processing object image. In this process, the flag `availableFlagLXCol` indicating whether the coded blocks of images from different times can be utilized, along with the motion vector `mvLXCol`, reference index `refIdxCol`, and reference list `listCol`, are derived, and `mvLXCol` is added to the predicted motion vector candidate list `mvpListLX`. Figure 20 Step S303).
[0173] Furthermore, it is assumed that the processing of time-predicted motion vector candidate deriving units 322 and 422, which are in units of sequence (SPS), image (PPS), or strip, can be omitted.
[0174] Next, the historical predicted motion vector candidate derivation units 323 and 423 add the historical predicted motion vector candidates registered in the historical predicted motion vector candidate list HmvpCandList to the predicted motion vector candidate list mvpListLX. Figure 20 Step S304). Details of the registration processing steps for step S304 are available via [link / reference]. Figure 29 The flowchart is described later.
[0175] Next, the predicted motion vector candidate supplementary units 325 and 425 add predicted motion vector candidates with predetermined values such as (0, 0) until the predicted motion vector candidate list mvpListLX is satisfied. Figure 20 (S305).
[0176] <Normal Merge Pattern Export Section (Normal Merge)>
[0177] Figure 18 The normal merging mode derivation section 302 includes a spatial merging candidate derivation section 341, a temporal merging candidate derivation section 342, an average merging candidate derivation section 344, a historical merging candidate derivation section 345, a merging candidate supplementation section 346, and a merging candidate selection section 347.
[0178] Figure 24 The normal merging mode derivation section 402 includes a spatial merging candidate derivation section 441, a temporal merging candidate derivation section 442, an average merging candidate derivation section 444, a historical merging candidate derivation section 445, a merging candidate supplementation section 446, and a merging candidate selection section 447.
[0179] Figure 21 This is a flowchart illustrating the steps of a normal merging mode export process that have a common function in the normal merging mode export unit 302 of the image encoding apparatus and the normal merging mode export unit 402 of the image decoding apparatus according to embodiments of the present invention.
[0180] The following sections will explain each process in turn. Furthermore, unless otherwise specified, the following explanations will focus on the case where the slice_type is B-slice, but this also applies to the case of P-slice. However, when the slice_type is P-slice, since only L0 prediction (Pred_L0) exists as the inter-frame prediction mode, and L1 prediction (Pred_L1) and double prediction (Pred_BI) do not exist, the processing surrounding L1 can be omitted.
[0181] Both the normal merge mode derivation unit 302 and the normal merge mode derivation unit 402 have a merge candidate list, mergeCandList. The merge candidate list mergeCandList forms a list structure and includes a merge index indicating the position within the merge candidate list, and a storage area for storing the merge candidates corresponding to the index as elements. The merge index number starts from 0, and merge candidates are stored in the storage area of the merge candidate list mergeCandList. In subsequent processing, it is assumed that the merge candidate registered at merge index i in the merge candidate list mergeCandList is represented by mergeCandList[i]. In this embodiment, it is assumed that the merge candidate list mergeCandList can register at least six merge candidates (inter-frame prediction information). Furthermore, the variable numCurrMergeCand, representing the number of merge candidates registered in the merge candidate list mergeCandList, is set to 0.
[0182] In the spatial merging candidate derivation unit 341 and the spatial merging candidate derivation unit 441, based on the encoding information stored in the encoding information storage memory 111 of the image encoding device or the encoding information storage memory 205 of the image decoding device, the blocks adjacent to the processing target block are derived in the order of B1, A1, B0, A0, B2. Figure 11 The space merge candidates (B1, A1, B0, A0, B2) are identified, and the exported space merge candidates are registered in the merge candidate list mergeCandList. Figure 21 (Step S401). Here, N is defined to represent any one of B1, A1, B0, A0, B2, or temporal merge candidate Col. The following are derived: availableFlagN indicating whether the inter-frame prediction information of block N can be used as a spatial merge candidate; refIdxL0N and refIdxL1N, the reference indices of L0 and L1 of spatial merge candidate N; predFlagL0N, indicating whether L0 prediction is performed; predFlagL1N, indicating whether L1 prediction is performed; motion vector mvL0N for L0; and motion vector mvL1N for L1. However, in this embodiment, since the merge candidates are derived without referring to the inter-frame prediction information of the blocks contained in the coded block to be processed, spatial merge candidates using the inter-frame prediction information of the blocks contained in the coded block to be processed are not derived.
[0183] Next, in the time merging candidate export unit 342 and the time merging candidate export unit 442, time merging candidates from images at different times are exported, and the exported time merging candidates are registered in the merge candidate list mergeCandList. Figure 21 Step S402). Derive the availableFlagCol, which indicates whether time merging candidates can be used; the predFlagL0Col, which indicates whether time merging candidates should be predicted for L0; the predFlagL1Col, which indicates whether L1 prediction should be made; and the motion vectors mvL0Col for L0 and mvL1Col for L1.
[0184] Furthermore, the processing of the time merging candidate export unit 342 and the time merging candidate export unit 442, which are based on sequences (SPS), images (PPS), or stripes, can be omitted.
[0185] Next, in the historical merge candidate derivation unit 345 and the historical merge candidate derivation unit 445, the historical predicted motion vector candidates registered in the historical predicted motion vector candidate list HmvpCandList are registered in the merge candidate list mergeCandList. Figure 21 Step S403).
[0186] Furthermore, if the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is less than the maximum number of merge candidates, MaxNumMergeCand, then the number of merge candidates registered in the merge candidate list mergeCandList numCurrMergeCand is used as the upper limit of the maximum number of merge candidates, MaxNumMergeCand, to derive historical merge candidates, and these candidates are then registered in the merge candidate list mergeCandList.
[0187] Next, in the average merge candidate derivation section 344 and the average merge candidate derivation section 444, average merge candidates are derived from the merge candidate list mergeCandList, and the exported average merge candidates are added to the merge candidate list mergeCandList. Figure 21 Step S404).
[0188] Furthermore, if the number of merge candidates registered in the merge candidate list numCurrMergeCand is less than the maximum number of merge candidates MaxNumMergeCand, the number of merge candidates registered in the merge candidate list numCurrMergeCand is capped at the maximum number of merge candidates MaxNumMergeCand, and an average number of merge candidates is derived and registered in the merge candidate list mergeCandList.
[0189] Here, the average merge candidate is a new merge candidate that has a motion vector obtained by averaging the motion vectors of the first and second merge candidates registered in the merge candidate list mergeCandList according to each L0 prediction and L1 prediction.
[0190] Next, in merge candidate supplementation section 346 and merge candidate supplementation section 446, when the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is less than the maximum number of merge candidates, MaxNumMergeCand, the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is used as the upper limit of the maximum number of merge candidates, MaxNumMergeCand, to derive additional merge candidates, and these candidates are registered in the merge candidate list mergeCandList. Figure 21 Step S405). Using the maximum number of merge candidates (MaxNumMergeCand) as the upper limit, in the P-strip, add merge candidates with a prediction mode of L0 prediction (Pred_L0) where the motion vector has a value of (0,0). In the B-strip, add merge candidates with a prediction mode of double prediction (Pred_BI) where the motion vector has a value of (0,0). The reference index when adding merge candidates is different from the reference index already added.
[0191] Next, in the merge candidate selection unit 347 and merge candidate selection unit 447, merge candidates are selected from those registered in the merge candidate list mergeCandList. On the encoding side, the merge candidate selection unit 347 selects merge candidates by calculating the amount of code and the amount of distortion, and provides the merge index representing the selected merge candidate and the inter-frame prediction information of the merge candidate to the motion compensation prediction unit 306 via the inter-frame prediction mode determination unit 305. On the other hand, on the decoding side, the merge candidate selection unit 447 selects merge candidates based on the decoded merge index and provides the selected merge candidates to the motion compensation prediction unit 406.
[0192] Normally, when the size (product of width and height) of a certain coding block is less than 32, the normal merging mode export unit 302 and the normal merging mode export unit 402 export merge candidates in the parent block of that coding block. Then, the merge candidates exported in the parent block are used in all child blocks. However, this is limited to the case where the size of the parent block is 32 or more and is contained within the screen.
[0193] <Update historical prediction motion vector candidate list>
[0194] Next, the initialization and update methods of the historical predicted motion vector candidate list HmvpCandList possessed by the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side are described in detail. Figure 26 This is a flowchart illustrating the steps involved in initializing / updating the candidate list of historical predicted motion vectors.
[0195] In this embodiment, it is assumed that the update of the historical predicted motion vector candidate list HmvpCandList is performed in the encoding information storage memory 111 and the encoding information storage memory 205. Alternatively, a historical predicted motion vector candidate list update unit may be provided in the inter-frame prediction unit 102 and the inter-frame prediction unit 203 to perform the update of the historical predicted motion vector candidate list HmvpCandList.
[0196] At the beginning of the strip, the historical predicted motion vector candidate list HmvpCandList is initially set. On the encoding side, if the prediction method determination unit 105 selects the normal predicted motion vector mode or the normal merging mode, the historical predicted motion vector candidate list HmvpCandList is updated. On the decoding side, if the prediction information decoded by the bit string decoding unit 201 is the normal predicted motion vector mode or the normal merging mode, the historical predicted motion vector candidate list HmvpCandList is updated.
[0197] Inter-frame prediction information used during inter-frame prediction in either the normal motion vector prediction mode or the normal merging mode is registered in the historical prediction motion vector candidate list hmvpCandList as inter-frame prediction information candidate hMvpCand. Inter-frame prediction information candidate hMvpCand includes the reference index refIdxL0 for L0 and the reference index refIdxL1 for L1, the L0 prediction flag predFlagL0 indicating whether L0 prediction is performed, the L1 prediction flag predFlagL1 indicating whether L1 prediction is performed, the motion vector mvL0 for L0, and the motion vector mvL1 for L1.
[0198] For each element (i.e., inter-frame prediction information) registered in the historical prediction motion vector candidate list HmvpCandList stored in the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side, it is checked sequentially from the beginning of the historical prediction motion vector candidate list HmvpCandList. If inter-frame prediction information with the same value as the inter-frame prediction information candidate hMvpCand exists, the element is deleted from the historical prediction motion vector candidate list HmvpCandList. On the other hand, if inter-frame prediction information with the same value as the inter-frame prediction information candidate hMvpCand does not exist, the element at the beginning of the historical prediction motion vector candidate list HmvpCandList is deleted, and the inter-frame prediction information candidate hMvpCand is added to the end of the historical prediction motion vector candidate list HmvpCandList.
[0199] The size of the maximum historical predicted motion vector candidate list, i.e., the maximum number of elements (maximum candidate number) in the historical predicted motion vector candidate list HmvpCandList, MaxNumHmvpCand, is set to 6. The size of the maximum historical predicted motion vector candidate list is the maximum number available in the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side of this invention. In addition, MaxNumHmvpCand can be the same value as the maximum merged candidate number MaxNumMergeCand-1, or the same value as the maximum merged candidate number MaxNumMergeCand, or a predetermined fixed value such as 5 or 6.
[0200] First, initialize the historical predicted motion vector candidate list HmvpCandList( on a strip basis). Figure 26 (Step S2101). At the beginning of the strip, clear all elements in the historical predicted motion vector candidate list HmvpCandList, and set the value of the historical predicted motion vector candidate number (current candidate number) NumHmvpCand registered in the historical predicted motion vector candidate list HmvpCandList to 0.
[0201] Furthermore, the offset value hMvpIdxOffset is set to a predetermined value. The value of the offset value hMvpIdxOffset is set to any predetermined value from 0 to (the size of the historical predicted motion vector candidate list, MaxNumHmvpCand-1). By setting the offset value hMvpIdxOffset to a value smaller than the maximum number of features in the historical predicted motion vector candidate list HmvpCandList, the number of comparisons between features described later can be reduced. The offset value hMvpIdxOffset is a predetermined value; however, the value of the offset value hMvpIdxOffset can be set by encoding / decoding on a sequence-by-sequence basis or by encoding / decoding on a strip-by-strip basis. The offset value hMvpIdxOffset will be described in detail later.
[0202] Furthermore, while it is assumed that the initialization of the historical predicted motion vector candidate list HmvpCandList is performed on a strip-by-strip basis (the initial encoded block of the strip), it can also be performed on an image-by-image basis, a tile-by-tile basis, or a tree-block-by-row basis.
[0203] Next, the historical predicted motion vector candidate list HmvpCandList is repeatedly updated for each coded block within the strip. Figure 26 Steps S2102 to S2107).
[0204] First, initial settings are performed on a block-by-block basis. The flag `identicalCandExist`, indicating the existence of identical candidates, is set to `FALSE`, and the index of the candidate to be deleted, `removeIdx`, is set to "0". Figure 26 Step S2103).
[0205] Determine whether there exists inter-frame prediction information with the same value as the registered object's inter-frame prediction information candidate hMvpCand in the historical predicted motion vector candidate list HmvpCandList. Figure 26 Step S2104). If the prediction method determination unit 105 on the encoding side determines it to be a normal prediction motion vector mode or a normal merging mode, or if the bit string decoding unit 201 on the decoding side decodes it to be a normal prediction motion vector mode or a normal merging mode, the inter-frame prediction information is set as the inter-frame prediction information candidate hMvpCand for the registered object. If the prediction method determination unit 105 on the encoding side determines it to be an intra-frame prediction mode, a sub-block prediction motion vector mode, or a sub-block merging mode, or if the bit string decoding unit 201 on the decoding side decodes it to be an intra-frame prediction mode, a sub-block prediction motion vector mode, or a sub-block merging mode, the historical prediction motion vector candidate list HmvpCandList is not updated, and there is no inter-frame prediction information candidate hMvpCand for the registered object. If there is no inter-frame prediction information candidate hMvpCand for the registered object, steps S2105 to S2106 are skipped. Figure 26 Step S2104: No). If there is a candidate hMvpCand for inter-frame prediction information of the registered object, proceed with the processing after step S2105. Figure 26 Step S2104: Yes).
[0206] Next, it is determined whether there exists an element in each element of the historical predicted motion vector candidate list HmvpCandList that has the same value as the inter-frame prediction information candidate hMvpCand of the registered object (inter-frame prediction information), that is, whether there is a matching element. Figure 26 Step S2105). Figure 27 This is a flowchart of the same element confirmation process. When the value of the historical predicted motion vector candidate number NumHmvpCand is 0 ( Figure 27 Step S2121: No), the historical predicted motion vector candidate list HmvpCandList is empty. Since there are no identical candidates, this step is skipped. Figure 27 Steps S2122 to S2125 conclude the same feature confirmation process. This applies when the value of the historical predicted motion vector candidate number NumHmvpCand is greater than 0 ( Figure 27 Step S2121: Yes), starting from the beginning of the historical predicted motion vector candidate list HmvpCandList and proceeding backwards, check sequentially whether there is inter-frame prediction information with the same value as the registered object's inter-frame prediction information candidate hMvpCand. The historical predicted motion vector index hMvpIdx ranges from 0 to NumHmvpCand-1, and the processing of step S2123 is repeated. Figure 27 Steps S2122 to S2125). First, compare whether the hMvpIdx-th element HmvpCandList[hMvpIdx] in the historical predicted motion vector candidate list (starting from 0) is the same as the inter-frame prediction information candidate hMvpCand. Figure 27 Step S2123). Under the same circumstances ( Figure 27 Step S2123: If yes, set the flag `identicalCandExist`, indicating whether there are identical candidates, to TRUE; set the current historical predicted motion vector index `hMvpIdx`, representing the location of the deleted object's index `removeIdx`, to the value of the index `hMvpIdx`; and end the identical feature confirmation process. In the case of different features ( Figure 27 Step S2123: No), indicating whether there is a candidate identical CandExist, keeps the flag FALSE, increases hMvpIdx by 1, if the historical predicted motion vector index hMvpIdx is below NumHmvpCand-1, then proceed with the processing after step S2123. Figure 27 Steps S2122 to S2125).
[0207] Return to Figure 26 The flowchart describes the process of shifting and adding elements to the historical predicted motion vector candidate list HmvpCandList. Figure 26 Step S2106). Figure 28 yes Figure 26 The flowchart for step S2106, the feature shifting / addition process in the historical predicted motion vector candidate list HmvpCandList, is as follows: First, it is determined whether to add new features after removing features stored in the historical predicted motion vector candidate list HmvpCandList, or to add new features without removing existing features. Specifically, it compares whether the flag identicalCandExist, indicating the existence of identical candidates, is TRUE, or whether the current number of candidates NumHmvpCand has reached the maximum number of candidates MaxNumHmvpCand. Figure 28 Step S2141). If the current candidate number NumHmvpCand is the same as the maximum candidate number MaxNumHmvpCand, it means that the maximum number of elements has been added to the historical predicted motion vector candidate list HmvpCandList. If either the flag identicalCandExist indicating the existence of identical candidates is TRUE or NumHmvpCand is the same as MaxNumHmvpCand, then... Figure 28 Step S2141: Yes), after deleting features stored in the historical predicted motion vector candidate list HmvpCandList, new features are added. Specifically, when the flag identicalCandExist, indicating whether there is a duplicate candidate, is TRUE, the duplicate candidate is deleted from the historical predicted motion vector candidate list HmvpCandList. When NumHmvpCand is the same value as MaxNumHmvpCand, the first candidate (feature) is deleted from the historical predicted motion vector candidate list HmvpCandList. The initial value of index i is set to the value of removeIdx+1. removeIdx is the index of the candidate to be deleted. This index i is changed from the initial value removeIdx+1 to NumHmvpCand-1, and the feature shifting process of step S2143 is repeated. Figure 28 Steps S2142 to S2144). By copying the elements of HmvpCandList[i] to HmvpCandList[i-1], the elements are shifted forward ( Figure 28 Step S2143), and increment i by 1 ( Figure 28 Steps S2142 to S2144). After index i becomes NumHmvpCand and the feature shifting process in step S2143 is completed, inter-frame prediction information candidate hmvpCand is added to the end of the historical prediction motion vector candidate list. Figure 28 Step S2145). Here, the last HmvpCandList[NumHmvpCand-1]th, counting from 0, is the (NumHmvpCand-1)th HmvpCandList[NumHmvpCand-1]. This concludes the feature shifting / addition process for this historical prediction motion vector candidate list HmvpCandList. On the other hand, if neither of the following conditions is met: the flag identicalCandExist indicating the existence of identical candidates is TRUE, or NumHmvpCand is the same value as MaxNumHmvpCand (…), then… Figure 28 Step S2141: No), that is, if the flag indicating whether there is an identical candidate is FALSE and NumHmvpCand is less than MaxNumHmvpCand, the feature stored in the historical predicted motion vector candidate list HmvpCandList is not removed, but the inter-frame prediction information candidate hMvpCand is added to the position after the last feature in the historical predicted motion vector candidate list. Figure 28 Step S2146). Here, the position after the last element in the historical predicted motion vector candidate list refers to the NumHmvpCandth HmvpCandList[NumHmvpCand] counting from 0. Before any elements are added to the historical predicted motion vector candidate list, this becomes position 0. Additionally, NumHmvpCand is incremented by 1, ending the element shifting and addition process for this historical predicted motion vector candidate list HmvpCandList.
[0208] Figures 31A to 31C This diagram illustrates an example of the update process for the historical predicted motion vector candidate list. When the historical predicted motion vector candidate list HmvpCandList contains six elements (inter-frame prediction information) of size MaxNumHmvpCand, if a new element is added, the new inter-frame prediction information is compared sequentially, starting with the elements preceding them in the historical predicted motion vector candidate list HmvpCandList. Figure 31A If the new feature has the same value as the third feature (HMVP2) from the beginning of the historical predicted motion vector candidate list (HmvpCandList), then remove feature HMVP2 from the historical predicted motion vector candidate list (HmvpCandList), shift (copy) the subsequent features HMVP3 to HMVP5 one by one forward, and add the new feature to the end of the historical predicted motion vector candidate list (HmvpCandList). Figure 31B Complete the update of the historical predicted motion vector candidate list HmvpCandList. Figure 31C ).
[0209] <Historical Predicted Motion Vector Candidate Export Processing>
[0210] Next, a detailed explanation will be provided as... Figure 20 Step S304 is a method for deriving historical predicted motion vector candidates from the historical predicted motion vector candidate list HmvpCandList. Figure 20 The processing step S304 is a common process in the historical prediction motion vector candidate derivation unit 323 of the normal prediction motion vector pattern derivation unit 301 on the encoding side and the historical prediction motion vector candidate derivation unit 423 of the normal prediction motion vector pattern derivation unit 401 on the decoding side. Figure 29 This is a flowchart illustrating the steps involved in deriving candidate motion vectors from historical predictions.
[0211] If the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (here, 2), or if the value of the historical number of predicted motion vector candidates (the number of features registered in the historical predicted motion vector list), NumHmvpCand, is 0. Figure 29 Step S2201: No), omitted Figure 29 The processing from steps S2202 to S2210 ends the historical predicted motion vector candidate export processing step. When the current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX (2), and the value of the historical predicted motion vector candidate number, NumHmvpCand, is greater than 0 (…), the process continues. Figure 29 Step S2201: Yes), execute Figure 29 The processing steps S2202 to S2210.
[0212] Next, repeat Figure 29 The processing of steps S2203 to S2209 continues until the index i reaches a smaller value from 1 to 4 (a predetermined upper limit) and the historical predicted motion vector candidate number NumHmvpCand. Figure 29 Steps S2202 to S2210). When the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (2 or more). Figure 29 Step S2203: No), omitted Figure 29 The processing steps S2204 to S2210 conclude the current historical predicted motion vector candidate export processing steps. The current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX, which is 2 ( Figure 29 If step S2203 is true, execute the following steps: Figure 29 The processing after step S2204.
[0213] Next, for the cases where the reference list LY of each element in the historical predicted motion vector candidate list HmvpCandList is L0 and L1, the processing from steps S2205 to S2208 is performed respectively. Figure 29 Steps S2204 to S2209). This indicates that for L0 and L1 of the historical predicted motion vector candidate list HmvpCandList, respectively... Figure 29 The processing of steps S2205 to S2208. When the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (more than 2). Figure 29 Step S2205: No), omitted Figure 29 The processing steps S2206 to S2210 conclude the historical predicted motion vector candidate export processing steps. This applies when the current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX (2). Figure 29 Step S2205 in the process: Yes, execute Figure 29 The processing after step S2206 in the process.
[0214] Next, the motion vectors of the features in the historical predicted motion vector candidate list are added to the predicted motion vector candidate list as predicted motion vector candidates. At this point, for features numbered from the end of the historical predicted motion vector candidate list by the offset value hMvpIdxOffset, they are checked in descending order to see if they are features not included in the predicted motion vector candidate list. If so, the features not included in the predicted motion vector candidate list are added to the predicted motion vector candidate list. Afterward, features not checked in descending order are added to the predicted motion vector candidate list, but only those from the historical predicted motion vector candidate list are added. By setting the offset value hMvpIdxOffset to a value smaller than the maximum number of features in the historical predicted motion vector candidate list HmvpCandList, the number of feature comparisons described later can be reduced. Figures 38A to 38D This explains why, when confirming the historical predicted motion vector candidate list, only the number of elements specified by the offset value hMvpIdxOffset from the end of the historical predicted motion vector candidate list are compared.
[0215] Figures 38A to 38D The diagram illustrates the relationship between three examples of 4-segmentation of a block and the historical predicted motion vector candidate list. It also explains the cases where each coded block is encoded using either the usual predicted motion vector pattern or the usual merge pattern. Figure 38A This is the graph when the encoded block of the encoding / decoding object is the upper right block. In this case, the inter-frame prediction information of the block to the left of the encoded block of the encoding / decoding object is highly likely to be the last element HMVP5 in the historical predicted motion vector candidate list. Figure 38B This is the diagram when the encoded block of the encoding / decoding object is the lower left block. In this case, the inter-frame prediction information of the upper right block of the encoded block of the encoding / decoding object is highly likely to be the last element HMVP5 in the historical predicted motion vector candidate list, and the inter-frame prediction information of the block above the encoded block of the encoding / decoding object is highly likely to be the second to last element HMVP4 in the historical predicted motion vector candidate list.
[0216] Figure 38C This diagram shows the situation when the encoded block of the encoding / decoding object is the bottom right block. In this case, the inter-frame prediction information of the left block of the encoded block is highly likely to be the last element (HMVP5) of the historical predicted motion vector candidate list; the inter-frame prediction information of the block above the encoded block is highly likely to be the second-to-last element (HMVP4) of the historical predicted motion vector candidate list; and the inter-frame prediction information of the top left block is highly likely to be the third-to-last element (HMVP3) of the historical predicted motion vector candidate list. In other words, the last element of the historical predicted motion vector candidate list has the highest probability of being derived as a spatial predicted motion vector candidate.
[0217] like Figure 38D As shown, the offset value hMvpIdxOffset is set to 1 so that only the last element HMVP5 in the historical predicted motion vector candidate list, which has the highest probability of being derived as a spatial predicted motion vector candidate, is compared. Alternatively, the offset value hMvpIdxOffset can be set to 2 so that the second-to-last element in the historical predicted motion vector candidate list, which has the second-highest probability of being derived as a spatial predicted motion vector candidate, is also compared. Furthermore, the offset value hMvpIdxOffset can be set to 3 so that the third-to-last element in the historical predicted motion vector candidate list, which has the third-highest probability of being derived as a spatial predicted motion vector candidate, is also compared. By setting the offset value hMvpIdxOffset to a value of 1 or higher in this way, the maximum number of comparisons of elements in the historical predicted motion vector candidate list can be reduced, thereby reducing the maximum processing load. Conversely, by setting the offset value hMvpIdxOffset to 0, the number of comparisons of elements in the historical predicted motion vector candidate list is 0, omitting the comparison process.
[0218] When index i is less than offset value hMvpIdxOffset, that is, when confirming whether it is a feature not included in the candidate list of predicted motion vectors ( Figure 29 Step S2206: Yes), if the reference index of the LY of the historical predicted motion vector candidate list HmvpCandList[NumHmvpCand-i] is the same as the reference index refIdxLX of the encoded / decoded object motion vector, and the LY of the feature HmvpCandList[NumHmvpCand-i] of the historical predicted motion vector candidate list is a feature that is different from any feature of the predicted motion vector candidate list mvpListLX ( Figure 29 Step S2207: YES), then as the last element of the predicted motion vector candidate list, in the element mvpListLX[numCurrMvpCand] of the predicted motion vector candidate list starting from 0, the motion vector of LY from the historical predicted motion vector candidate HmvpCandList[NumHmvpCand-i] is added to the predicted motion vector candidate list mvpListLX. Figure 29 Step S2208), and increment the current number of predicted motion vector candidates numCurrMvpCand by 1. If there are no features in the historical predicted motion vector candidate list HmvpCandList with the same reference index refIdxLX as the reference index of the encoded / decoded object motion vector, and no features that are different from any features in the predicted motion vector list mvpListLX ( Figure 29 If step S2207: No), then skip the addition process in step S2208.
[0219] On the other hand, when index i is not less than offset value hMvpIdxOffset, that is, when it is uncertain whether it is a feature not included in the candidate list of predicted motion vectors ( Figure 29 In step S2206 (No), as the last element of the predicted motion vector candidate list, add the motion vector of LY from the historical predicted motion vector candidate HmvpCandList[NumHmvpCand-i] to the element mvpListLX[numCurrMvpCand] which is the numCurrMvpCand-i element starting from 0 in the predicted motion vector candidate list. Figure 29 Step S2208), and increment the current number of predicted motion vector candidates numCurrMvpCand by 1.
[0220] Both L0 and L1 are performed as above Figure 29 Processing steps S2205 to S2208 ( Figure 29 Steps S2204 to S2209).
[0221] Increment index i by 1 ( Figure 29 In steps S2202 and S2210), if index i is less than or equal to either the predetermined upper limit value of 4 or the historical predicted motion vector candidate number NumHmvpCand, then the processing after step S2203 is performed again. Figure 29 Steps S2202 to S2210).
[0222] <Historical Merge Candidate Export Processing>
[0223] Next, a detailed explanation will be provided as... Figure 21 Step S404 is a method for deriving historical merge candidates from the historical merge candidate list HmvpCandList. Figure 21 The processing step S404 is a common process in the historical merge candidate derivation section 345 of the normal merge mode derivation section 302 on the encoding side and the historical merge candidate derivation section 445 of the normal merge mode derivation section 402 on the decoding side. Figure 30 This is a flowchart illustrating the steps involved in the historical merge candidate export process.
[0224] First, perform initialization processing ( Figure 30 Step S2301). Set the value of FALSE for each element from 0 to (numCurrMergeCand-1)th element of the flag isPruned[i], and set the variable numOrigMergeCand to the number of elements registered in the current merge candidate list numCurrMergeCand.
[0225] Next, add features from the historical predicted motion vector list that are not included in the merge candidate list to the merge candidate list. At this point, confirm and add features in descending order from the end of the historical predicted motion vector list. Set the initial value of index hMvpIdx to 1, and repeat the process from this initial value to NumHmvpCand. Figure 30 The addition process from step S2303 to step S2311 ( Figure 30 Steps S2302 to S2312). If the number of features registered in the current merge candidate list, numCurrMergeCand, is not below (maximum number of merge candidates, MaxNumMergeCand-1), that is, if the number of features registered in the current merge candidate list, numCurrMergeCand, reaches the maximum number of merge candidates, MaxNumMergeCand, then the historical merge candidate export process ends because all features have been added to the merge candidate list as merge candidates. Figure 30 Step S2303: No). If the number of features registered in the current merge candidate list, numCurrMergeCand, is less than (MaxNumMergeCand-1) ( Figure 30 If step S2303 is true, then proceed with the processing after step S2304.
[0226] Set the value of FALSE to the variable sameMotion, which represents the same motion information. Figure 30 (Step S2304). Next, the inter-frame prediction information of the elements in the historical predicted motion vector candidate list is added as merging candidates to the merging candidate list. At this time, for elements numbered by the offset value hMvpIdxOffset starting from the elements following in the historical predicted motion vector candidate list (the most recently added elements), it is checked in descending order whether they are elements not included in the merging candidate list. Elements not included in the merging candidate list are added to the merging candidate list. Then, elements from the historical predicted motion vector candidate list are added to the merging candidate list without checking whether they are elements not included in the merging candidate list in descending order. By setting the offset value hMvpIdxOffset to a value smaller than the maximum number of elements in the historical predicted motion vector candidate list HmvpCandList, the number of comparisons between elements described later can be reduced. Figures 38A to 38D This explains why, when confirming the historical predicted motion vector candidate list, only the number of features specified by the offset value hMvpIdxOffset is compared, starting from the feature following the last feature in the historical predicted motion vector candidate list (the most recently added feature). Figures 38A to 38D The diagram illustrates the relationship between three examples of 4-segmentation of a block and the historical predicted motion vector candidate list. It also explains the cases where each coded block is encoded using either the usual predicted motion vector pattern or the usual merge pattern. Figure 38A This is the graph when the encoded block of the encoding / decoding object is the upper right block. In this case, the inter-frame prediction information of the block to the left of the encoded block of the encoding / decoding object is highly likely to be the last element HMVP5 in the historical predicted motion vector candidate list. Figure 38B This is the diagram when the encoded block of the encoding / decoding object is the lower left block. In this case, the inter-frame prediction information of the upper right block of the encoded block of the encoding / decoding object is highly likely to be the last element HMVP5 in the historical predicted motion vector candidate list, and the inter-frame prediction information of the block above the encoded block of the encoding / decoding object is highly likely to be the second to last element HMVP4 in the historical predicted motion vector candidate list. Figure 38C This is the diagram when the encoded block of the encoding / decoding object is the bottom right block. In this case, the inter-frame prediction information of the block to the left of the encoded block of the encoding / decoding object is highly likely to be the last element (HMVP5) of the historical predicted motion vector candidate list; the inter-frame prediction information of the block above the encoded block of the encoding / decoding object is highly likely to be the second-to-last element (HMVP4) of the historical predicted motion vector candidate list; and the inter-frame prediction information of the block to the top left of the encoded block of the encoding / decoding object is highly likely to be the third-to-last element (HMVP3) of the historical predicted motion vector candidate list. That is, the last element of the historical predicted motion vector candidate list has the highest probability of being derived as a spatial merging candidate. For example... Figure 38D As shown, the offset value hMvpIdxOffset is set to 1 so that only the last element HMVP5 in the historical predicted motion vector candidate list with the highest probability of being exported as a spatial merge candidate is compared. Furthermore, the offset value hMvpIdxOffset can also be set to 2 so that the second-to-last element in the historical predicted motion vector candidate list with the second-highest probability of being exported as a spatial merge candidate is also compared. Additionally, the offset value hMvpIdxOffset can also be set to 3 so that the third-to-last element in the historical predicted motion vector candidate list with the third-highest probability of being exported as a spatial merge candidate is also compared. By setting the offset value hMvpIdxOffset to values 1 to 3 in this way, only elements from the last element in the historical predicted motion vector candidate list (the most recently added element) specified by the offset value hMvpIdxOffset, and spatial merge candidates, or elements stored in the merge candidate list, are compared. This reduces the maximum number of times elements in the historical predicted motion vector candidate list are compared, thereby reducing the maximum processing load.
[0227] When the index hMvpIdx is below the predetermined offset value hMvpIdxOffset, that is, when confirming whether it is a feature not included in the merge candidate list ( Figure 30 Step S2305: If yes, set the initial value of index i to 0, and proceed from this initial value to numOrigMergeCand-1. Figure 30 Processing steps S2307 and S2308 ( Figure 30 (S2306~S2309). Compare whether the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] in the historical predicted motion vector candidate list, starting from 0, is the same as the i-th element mergeCandList[i] in the merged candidate list. Figure 30 Step S2307). The same value for the merging candidates means that the values of all the components of the inter-frame prediction information (inter-frame prediction mode, L0 and L1 reference indices, L0 and L1 motion vectors) possessed by the merging candidates are the same. Therefore, in the comparison process of step S2307, when isPruned[i] is FALSE, the values of all the components (inter-frame prediction mode, L0 and L1 reference indices, L0 and L1 motion vectors) possessed by mergeCandList[i] and HmvpCandList[NumHmvpCand-hMvpIdx] are compared to see if they are the same. In the case of the same value ( Figure 30 Step S2307: Yes), set both sameMotion and isPruned[i] to TRUE. Figure 30 Step S2308). Additionally, the flag isPruned[i] indicates that the i-th element, counting from 0 in the merged candidate list, has the same value as an element in the historical predicted motion vector candidate list. In the case where they are not the same value ( Figure 30 If step S2307 is not specified, skip step S2308. Figure 30 After the repeated processing of steps S2306 to S2309 is completed, compare whether sameMotion is FALSE. Figure 30 Step S2310), when sameMotion is FALSE (false) Figure 30 Step S2310: Yes), that is, the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] from the historical predicted motion vector candidate list (starting from 0) does not exist in mergeCandCandList. Therefore, as the last element of the merge candidate list, the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] from the historical predicted motion vector candidate list (starting from 0) is added to the mergeCandList[numCurrMergeCand] of the merge candidate list (starting from 0), and numCurrMergeCand is incremented by 1. Figure 30 Step S2311). On the other hand, if the index hMvpIdx is not less than the offset value hMvpIdxOffset, that is, if it is not confirmed whether it is a feature not included in the merge candidate list ( Figure 30 If step S2305: No), then as the final element of the merge candidate list, add the (NumHmvpCand – hMvpIdx)th element HmvpCandList[NumHmvpCand – hMvpIdx] from the historical predicted motion vector candidate list, starting from 0, to the mergeCandList[numCurrMergeCand] of the merge candidate list, and increment numCurrMergeCand by 1. Figure 30 Step S2312).
[0228] Furthermore, increase the index hMvpIdx by 1. Figure 30 Step S2302) is performed. Figure 30 Repeat steps S2302 to S2312.
[0229] After confirming all elements in the historical predicted motion vector candidate list or adding all elements in the merge candidate list to the merge candidate list, the export process of this historical merge candidate is completed.
[0230] <Motion Compensation Predictive Processing>
[0231] The motion compensation prediction unit 306 acquires the position and size of the block of the object currently being predicted during encoding. Additionally, the motion compensation prediction unit 306 acquires inter-frame prediction information from the inter-frame prediction mode determination unit 305. Based on the acquired inter-frame prediction information, a reference index and motion vector are derived. After acquiring an image signal that shifts the reference image determined by the reference index in the decoded image memory 104 from the same position as the image signal of the predicted block by the amount of motion vector, a prediction signal is generated.
[0232] In inter-frame prediction, when the inter-frame prediction mode is L0 or L1 prediction, which involves prediction from a single reference image, the prediction signal obtained from one reference image is set as the motion-compensated prediction signal. When the inter-frame prediction mode is BI prediction, which involves prediction from two reference images, the signal obtained by weighted averaging of the prediction signals obtained from the two reference images is set as the motion-compensated prediction signal, and this motion-compensated prediction signal is provided to the prediction method determination unit 105. Here, the weighted average ratio of the two predictions is set to 1:1, but other ratios can also be used for weighted averaging. For example, the closer the image to be predicted is to the reference image, the larger the weighting ratio. Alternatively, a table corresponding to combinations of image intervals and weighting ratios can be used to calculate the weighting ratio.
[0233] The motion compensation prediction unit 406 has the same function as the motion compensation prediction unit 306 on the encoding side. The motion compensation prediction unit 406 acquires inter-frame prediction information from the normal prediction motion vector pattern derivation unit 401, the normal merging pattern derivation unit 402, the sub-block prediction motion vector pattern derivation unit 403, and the sub-block merging pattern derivation unit 404 via switch 408. The motion compensation prediction unit 406 provides the acquired motion compensation prediction signal to the decoded image signal overlay unit 207.
[0234] <About Inter-Frame Prediction Mode>
[0235] The process of making a prediction based on a single reference image is defined as single prediction. In the case of single prediction, a prediction is made using either an L0 prediction or an L1 prediction, which uses either of the two reference images registered in the reference list L0 or L1.
[0236] Figure 32 This shows the case where the reference image (RefL0Pic) of L0 in a single prediction is at a moment before the processing object image (CurPic). Figure 33 This illustrates the case where the reference image for the L0 prediction in a single prediction is at a time after the processed object image. Similarly, it is also possible to... Figure 32 and Figure 33 The reference image for L0 prediction is replaced with the reference image for L1 prediction (RefL1Pic) for single prediction.
[0237] The process of making predictions based on two reference images is defined as dual prediction. In the case of dual prediction, L0 prediction and L1 prediction are used to describe BI prediction. Figure 34 This illustrates the case where the reference image for L0 prediction is at a time before the object image is processed, and the reference image for L1 prediction is at a time after the object image is processed. Figure 35 This shows the situation where the reference images for L0 prediction and L1 prediction in a dual prediction are at a time before the object image is processed. Figure 36 This shows the case where the reference images for L0 prediction and L1 prediction in a dual prediction scenario are at a time after the processed object image.
[0238] Thus, the relationship between the prediction category and time of L0 / L1 can be used when L0 is not limited to the past direction and L1 is not limited to the future direction. Furthermore, in the case of dual prediction, the same reference image can be used to perform both L0 and L1 predictions. Additionally, it is determined whether motion compensation prediction is performed using single or dual prediction based on information indicating whether L0 or L1 prediction is used (e.g., flags).
[0239] <About the Reference Index>
[0240] In embodiments of the present invention, to improve the accuracy of motion compensation prediction, the optimal reference image can be selected from multiple reference images during motion compensation prediction. Therefore, the reference image used in motion compensation prediction is used as a reference index, and the reference index is encoded into the bitstream along with the differential motion vector.
[0241] <Motion compensation processing based on typical predicted motion vector patterns>
[0242] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when inter-frame prediction information based on the normal prediction motion vector mode derivation unit 301 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0243] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to the normal prediction motion vector mode derivation unit 401 during decoding, motion compensation prediction unit 406 acquires the inter-frame prediction information based on the normal prediction motion vector mode derivation unit 401, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0244] <Motion compensation processing based on the usual merging pattern>
[0245] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when inter-frame prediction information based on the normal merging mode derivation unit 302 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0246] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to the normal merging mode derivation unit 402 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on the normal merging mode derivation unit 402, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0247] <Motion Compensation Processing Based on Sub-Block Predicted Motion Vector Patterns>
[0248] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when the inter-frame prediction information based on the sub-block prediction motion vector mode derivation unit 303 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0249] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to sub-block prediction motion vector mode derivation unit 403 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on sub-block prediction motion vector mode derivation unit 403, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to decoded image signal overlay unit 207.
[0250] <Motion Compensation Processing Based on Sub-Block Merging Pattern>
[0251] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when the inter-frame prediction mode determination unit 305 selects the inter-frame prediction information based on the sub-block merging mode derivation unit 304, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0252] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to sub-block merging mode derivation unit 404 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on sub-block merging mode derivation unit 404, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0253] <Motion Compensation Processing Based on Affine Transform Prediction>
[0254] In both the normal prediction motion vector mode and the normal merging mode, affine model-based motion compensation can be used based on the following flags. These flags are reflected in the following flags based on the inter-frame prediction conditions determined by the inter-frame prediction mode determination unit 305 during the encoding process, and are encoded into the bitstream. During the decoding process, whether to perform affine model-based motion compensation is determined based on the following flags in the bitstream.
[0255] `sps_affine_enabled_flag` indicates whether affine-based motion compensation can be used in inter-frame prediction. If `sps_affine_enabled_flag` is 0, sequence-by-sequence suppression is applied, preventing affine-based motion compensation. Furthermore, `inter_affine_flag` and `cu_affine_type_flag` are not transmitted in the CU (coded block) syntax of the encoded video sequence. If `sps_affine_enabled_flag` is 1, affine-based motion compensation can be used in the encoded video sequence.
[0256] `sps_affine_type_flag` indicates whether motion compensation based on a six-parameter affine model can be used in inter-frame prediction. If `sps_affine_type_flag` is 0, motion compensation based on a six-parameter affine model is suppressed. Additionally, `cu_affine_type_flag` is not transmitted in the CU syntax of the encoded video sequence. If `sps_affine_type_flag` is 1, motion compensation based on a six-parameter affine model can be used in the encoded video sequence. If `sps_affine_type_flag` is not present, it is set to 0.
[0257] When decoding P-strips or B-strips, in the CU that is currently being processed, if `inter_affine_flag` is 1, motion compensation based on an affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed. If `inter_affine_flag` is 0, the affine model is not used for the CU that is currently being processed. If `inter_affine_flag` does not exist, it is set to 0.
[0258] When decoding P-strips or B-strips, in the CU that is currently being processed, if cu_affine_type_flag is 1, motion compensation based on a six-parameter affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed. If cu_affine_type_flag is 0, motion compensation based on a four-parameter affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed.
[0259] In motion compensation based on affine models, since reference indices or motion vectors are derived on a sub-block basis, motion compensation prediction signals are generated using the reference indices or motion vectors that are the objects of processing on a sub-block basis.
[0260] The four-parameter affine model is as follows: the motion vector of the sub-block is derived from the four parameters of the horizontal and vertical components of the motion vectors of the two control points, and motion compensation is performed on a sub-block basis.
[0261] (Second Implementation)
[0262] Next, the image encoding apparatus and image decoding apparatus according to the second embodiment will be described. The structure is the same as that of the image encoding apparatus and image decoding apparatus according to the first embodiment, but the processing steps of the history merging candidate derivation units 345 and 445 are different. The processing steps of the history merging candidate derivation units 345 and 445 are as follows: Figure 39 As shown in the flowchart, to replace Figure 30 The flowchart illustrates these differences.
[0263] <Historical Merging Candidate Export Processing in the Second Implementation>
[0264] use Figure 39 The flowchart illustrates the historical merging candidate export process of the image encoding device and the image decoding device involved in the second embodiment.
[0265] The second embodiment differs from the first embodiment in that it only compares spatial merging candidate A1 derived from the block A1 adjacent to the left side of the encoded block being processed, spatial merging candidate B1 derived from the block B1 adjacent to the right side, and each element specified by the offset value hMvpIdxOffset from the element following the historical predicted motion vector candidate list (the most recently added element). This is in contrast to the historical merging candidate derivation processing of the image encoding apparatus and image decoding apparatus according to the first embodiment. Figure 30 The flowchart is different from that of the flowchart. Figure 39 China will respectively Figure 30 Steps S2306, S2307, and S2309 are changed to steps S2326, S2327, and S2329, and the rest of the process is the same. In the second embodiment, firstly, an initialization process is performed ( Figure 39 Step S2301). Set the value of FALSE for each element of isPruned[i] starting from 0 and the value of numCurrMergeCand for the variable numOrigMergeCand to the number of elements registered in the current merge candidate list.
[0266] Next, the initial value of index hMvpIdx is set to 1, and the process is repeated from this initial value to NumHmvpCand. Figure 39 The addition process from step S2303 to step S2311 ( Figure 39 Steps S2302 to S2312). If the number of features registered in the current merge candidate list, numCurrMergeCand, is not below (MaxNumMergeCand-1), then the historical merge candidate export process ends because merge candidates have been added to all features in the merge candidate list. Figure 39 Step S2303: No). If the number of features registered in the current merge candidate list, numCurrMergeCand, is less than (MaxNumMergeCand-1) ( Figure 39 If step S2303 is true, then the processing after step S2304 will be executed.
[0267] First, set the value of sameMotion to FALSE. Figure 39 Step S2304). Next, when the index hMvpIdx is below the predetermined offset value hMvpIdxOffset, that is, when confirming whether it is a feature not included in the merge candidate list ( Figure 39 Step S2305: Yes), set the initial value of index i to 0, and execute the operation from this initial value up to the smaller of 1 and numOrigMergeCand-1, which is a predetermined upper limit value. Figure 39 Processing steps S2327 and S2308 ( Figure 39 (S2326-S2329). Here, in the second embodiment, only the spatial merging candidate A1 derived from the block A1 adjacent to the left of the coded block of the processing object, the spatial merging candidate B1 adjacent to the right, and the number of elements specified by the offset value hMvpIdxOffset from the elements following the historical predicted motion vector candidate list (the most recently added elements) are compared. The predetermined upper limit value of 1 is because the spatial merging candidate A1 derived from the block A1 adjacent to the left of the coded block of the processing object or the spatial merging candidate B1 adjacent to the right can only be stored as the 0th and 1st elements starting from 0 in the merging candidate list. Spatial merging candidates A1 and B1 are compared with the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] starting from 0 in the historical predicted motion vector candidate list. Figure 39 Step S2327). Compare whether the values of all the constituent elements (inter-frame prediction mode, L0 and L1 reference indices, L0 and L1 motion vectors) of the merge candidates are the same. Here, the same value for merge candidates means that the values of all the constituent elements (inter-frame prediction mode, L0 and L1 reference indices, L0 and L1 motion vectors) of the merge candidates are the same. Therefore, in the comparison process of step S2327, when the i-th element mergeCandList[i] of the merge candidate list, starting from 0, is spatial merge candidate A1 derived from the block A1 adjacent to the left or spatial merge candidate B1 adjacent to the right, and isPruned[i] is FALSE, compare whether the values of all the constituent elements (inter-frame prediction mode, L0 and L1 reference indices 0, L0 and L1 motion vectors) of mergeCandList[i] and HmvpCandList[NumHMvpCand-hMvpIdx] are the same. In the case of the same value ( Figure 39 Step S2327: Yes), set both sameMotion and isPruned[i] to TRUE. Figure 39 Step S2308). Additionally, the flag isPruned[i] indicates that the i-th element in the merged candidate list (starting from 0) has the same value as an element in the historical predicted motion vector candidate list. In the case where they are not the same value ( Figure 39 Step S2327: No), skip step S2308. From Figure 39 After the repeated processing of steps S2326 to S2329 is completed, compare whether sameMotion is FALSE. Figure 39 Step S2310), when sameMotion is FALSE (false) Figure 39 Step S2310: Yes), add the (NumHmvpCand-hMvpIdx)th feature HmvpCandList[NumHmvpCand-hMvpIdx] from the historical predicted motion vector candidate list to mergeCandList[numCurrMergeCand] of the merge candidate list starting from 0, and increment numCurrMergeCand by 1. Figure 39 Step S2311). Increment the index hMvpIdx by 1 ( Figure 39 Step S2302) is performed. Figure 39 Repeat steps S2302 to S2312.
[0268] After confirming all elements of the historical predicted motion vector candidate list, or after merging candidates have been added to all elements of the merged candidate list, the export process of this historical merged candidate is completed.
[0269] In the second embodiment, it is described that spatial merging candidates A1 and B1 stored in the merging candidate list are compared with elements of the historical predicted motion vector candidate list. However, it is also possible to store spatial merging candidates A1 and B1 in memory outside the merging candidate list, and compare the spatial merging candidates A1 and B1 stored in memory outside the merging candidate list with elements of the historical predicted motion vector candidate list.
[0270] (Third Implementation)
[0271] Next, the image encoding apparatus and image decoding apparatus according to the third embodiment will be described. The structure is the same as that of the image encoding apparatus and image decoding apparatus according to the first embodiment, but the same element confirmation processing step in the historical predicted motion vector candidate list initialization / update processing step of the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side is different. The same element confirmation processing step in the historical predicted motion vector candidate list initialization / update processing step of the first embodiment is replaced by... Figure 27 The flowchart of the third embodiment shows the same element confirmation process in the historical prediction motion vector candidate list initialization / update processing step, as follows: Figure 40 The flowchart illustrates these differences.
[0272] <Same Element Confirmation Processing Step in the Initialization / Update Processing Step of the Historical Prediction Motion Vector Candidate List in the Third Implementation>
[0273] use Figure 40 The flowchart illustrates the same element confirmation process step in the historical prediction motion vector candidate list initialization / update processing step of the image encoding device and image decoding device involved in the third embodiment.
[0274] In the third embodiment, the following aspects differ from the first embodiment: In the historical predicted motion vector candidate list update process, when the maximum number of elements has been added to the historical predicted motion vector candidate list HmvpCandList, that is, when the current number of historical predicted motion vector candidates NumHmvpCand reaches the maximum number of historical predicted motion vector candidates MaxNumHmvpCand, the first element contained in the historical predicted motion vector candidate list (i.e., the 0th element counting from 0) is not compared; only the elements after the 1st element are compared. The elements contained in the historical predicted motion vector candidate list consist of inter-frame prediction modes, reference indices, and motion vectors. By not comparing the first element contained in the historical predicted motion vector candidate list, the maximum number of element comparisons is limited to (MaxNumHmvpCand-1), reducing the maximum processing load associated with element comparisons. The same element confirmation process as in the historical predicted motion vector candidate list initialization / update process of the first embodiment is also included. Figure 27 The difference in the flowchart is that, Figure 27 Steps S2122 and S2125 are in Figure 40 The steps are changed to S2132 and S2135 respectively, and the rest of the process is the same.
[0275] In the third embodiment, when the value of the historical predicted motion vector candidate number NumHmvpCand is 0 ( Figure 40 Step S2121: No), the historical predicted motion vector candidate list HmvpCandList is empty, there are no duplicate candidates, therefore, skip. Figure 40 Steps S2132 to S2135 conclude the same feature confirmation process. This applies when the value of the historical predicted motion vector candidate number NumHmvpCand is greater than 0 ( Figure 40 Step S2121: Yes), the historical predicted motion vector index hMvpIdx ranges from 0 or 1 to NumHmvpCand-1, and the processing of step S2123 is repeated. Figure 40 Steps S2132 to S2135). First, if the current number of historical predicted motion vector candidates, NumHmvpCand, is less than the maximum number of historical predicted motion vector candidates, MaxNumHmvpCand, the first element in the historical predicted motion vector candidate list, i.e., the 0th element counting from 0 (historical predicted motion vector candidate), is compared, and therefore hMvpIdx is set to 0. On the other hand, if the current number of historical predicted motion vector candidates, NumHmvpCand, reaches the maximum number of historical predicted motion vector candidates, MaxNumHmvpCand, the first element in the historical predicted motion vector candidate list, i.e., the 0th element counting from 0 (historical predicted motion vector candidate), is not compared, and therefore hMvpIdx is set to 1 (…). Figure 40 Step S2132). Next, compare whether the hMvpIdx-th element HmvpCandList[hMvpIdx] in the historical predicted motion vector candidate list, starting from 0, is the same as the inter-frame prediction information candidate hMvpCand of the registered object. Figure 40 Step S2123). Under the same circumstances ( Figure 40 Step S2123: Yes), set the value of TRUE to the flag identicalCandExist indicating whether the same candidate exists, set the value of hMvpIdx to the index removeIdx of the deleted object, and end the same feature confirmation process. In the case of different... Figure 40 Step S2123: No), increment hMvpIdx by 1. If the historical predicted motion vector index hMvpIdx is below NumHmvpCand-1, then proceed with the processing after step S2123. Figure 40 Steps S2132 to S2135).
[0276] (Fourth Implementation)
[0277] Next, the image encoding apparatus and image decoding apparatus according to the fourth embodiment will be described. The structure is the same as that of the image encoding apparatus and image decoding apparatus according to the first embodiment, but the same element confirmation processing step in the historical predicted motion vector candidate list initialization / update processing step of the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side is different. The same element confirmation processing step in the historical predicted motion vector candidate list initialization / update processing step of the first embodiment is replaced by... Figure 27 The flowchart of the fourth embodiment shows the same element confirmation process in the historical prediction motion vector candidate list initialization / update processing step, as follows: Figure 41 The flowchart illustrates these differences.
[0278] <The same element confirmation process in the historical predicted motion vector candidate list initialization / update processing steps involved in the fourth embodiment>
[0279] use Figure 41 The flowchart illustrates the same element confirmation process step in the historical prediction motion vector candidate list initialization / update processing step of the image encoding device and image decoding device involved in the fourth embodiment.
[0280] In the fourth embodiment, unlike the first and third embodiments, the historical predicted motion vector candidate list update process compares elements in descending order, starting from the last element in the historical predicted motion vector candidate list. Furthermore, unlike the first embodiment, in the historical predicted motion vector candidate list update process, when the maximum number of elements has been added to the historical predicted motion vector candidate list HmvpCandList, the first elements in the historical predicted motion vector candidate list are not compared; that is, the 0th element (counting from 0) is not compared, but only elements from the 1st element onwards are compared. The elements included in the historical predicted motion vector candidate list consist of inter-frame prediction modes, reference indices, and motion vectors. Because the first elements in the historical predicted motion vector candidate list are not compared, the maximum number of element comparisons is limited to (MaxNumHmvpCand-1), reducing the maximum processing load associated with element comparisons.
[0281] Similarly, in the fourth embodiment, when the value of the historical predicted motion vector candidate number NumHmvpCand is 0 ( Figure 41 Step S2151: No), the historical predicted motion vector candidate list HmvpCandList is empty, there are no duplicate candidates, therefore, skip. Figure 41 Steps S2152 to S2155 conclude the same feature confirmation process. When the value of the historical predicted motion vector candidate number NumHmvpCand is greater than 0 ( Figure 41 In step S2152 (if yes), the index i is increased from 1 to the smaller of the maximum number of historical predicted motion vector candidates MaxNumHmvpCand-1 and the number of historical predicted motion vector candidates NumHmvpCand, and the processing of step S2153 is repeated. Figure 41 Steps S2152 to S2155). First, compare whether the NumHmvpCand-i element HmvpCandList[NumHmvpCand-i] in the historical predicted motion vector candidate list is the same as the inter-frame prediction information candidate hMvpCand of the registered object (starting from 0). Figure 41 Step S2153). Under the same circumstances ( Figure 41 Step S2153: Yes), set the value of TRUE to the flag identicalCandExist indicating whether the same candidate exists, set the value of NumHmvpCand-i to the index removeIdx of the deleted object, and end the same feature confirmation process. Figure 41 Step S2154). In different cases ( Figure 41 Step S2153: No), increment i by 1. If i is less than the smaller of the maximum number of historical predicted motion vector candidates MaxNumHmvpCand-1 and the number of historical predicted motion vector candidates NumHmvpCand, then proceed with the processing after step S2153. Figure 41 Steps S2152 to S2155). Additionally, when index i is 1 as the initial value, the NumHmvpCand-i element HmvpCandList[NumHmvpCand-i] in the historical predicted motion vector candidate list, starting from 0, represents the last element registered in the historical predicted motion vector candidate list, and the elements in the historical predicted motion vector candidate list are displayed in descending order as index i increments by 1. The maximum index i is (MaxNumHmvpCand-1), therefore the first element HMVPCandList[0] in the historical predicted motion vector candidate list is not compared.
[0282] All of the above-described embodiments can also be combined in multiple ways.
[0283] In all the embodiments described above, the bitstream output by the image encoding device has a specific data format so that it can be decoded according to the encoding method used in the embodiment. Furthermore, the image decoding device corresponding to the image encoding device is capable of decoding the bitstream of this specific data format.
[0284] When using wired or wireless networks to exchange bitstreams between an image encoding device and an image decoding device, the bitstream can be converted into a data format suitable for transmission over the communication line for transmission. In this case, a transmitting device is provided that converts the bitstream output from the image encoding device into encoded data in a data format suitable for transmission over the communication line and transmits the encoded data to the network; and a receiving device receives the encoded data from the network and restores the encoded data to a bitstream for supply to the image decoding device. The transmitting device includes: a memory for buffering the bitstream output from the image encoding device; a packet processing unit for packetizing the bitstream; and a transmitting unit for transmitting the packetized encoded data via the network. The receiving device includes: a receiving unit for receiving the packetized encoded data via the network; a memory for buffering the received encoded data; and a packet processing unit for packetizing the encoded data to generate a bitstream and supplying the bitstream to the image decoding device.
[0285] Alternatively, the display device can be configured by adding a display unit that displays the image decoded by the image decoding device. In this case, the display unit reads the decoded image signal generated by the decoded image signal overlap unit 207 and stored in the decoded image memory 208, and displays it on the screen.
[0286] Alternatively, an imaging unit can be added to the structure to input the captured image into an image encoding device, thus functioning as an imaging device. In this case, the imaging unit inputs the captured image signal into the block segmentation unit 101.
[0287] Figure 37 An example of the hardware structure of the encoding / decoding apparatus according to this embodiment is shown. The encoding / decoding apparatus includes the structure of the image encoding apparatus and the image decoding apparatus according to embodiments of the present invention. The encoding / decoding apparatus 9000 has a CPU 9001, an encoder / decoder IC 9002, an I / O interface 9003, a memory 9004, an optical disc drive 9005, a network interface 9006, and a video interface 9009, and the various parts are connected via a bus 9010.
[0288] The image encoding unit 9007 and the image decoding unit 9008 are typically installed as a codec IC 9002. In the image encoding apparatus according to embodiments of the present invention, image encoding processing is performed by the image encoding unit 9007, and in the image decoding apparatus according to embodiments of the present invention, image decoding processing is performed by the image decoding unit 9008. The I / O interface 9003 is implemented, for example, via a USB interface, and is connected to an external keyboard 9104, mouse 9105, etc. The CPU 9001 controls the encoding / decoding apparatus 9000 to execute the user-desired action based on user operations input through the I / O interface 9003. These operations, performed by the user via the keyboard 9104, mouse 9105, etc., include selecting which function to perform (encoding or decoding), setting the encoding quality, setting the input / output destination of the bitstream, and setting the input / output destination of the image.
[0289] When a user wishes to reproduce an image recorded on the disc recording medium 9100, the optical disc drive 9005 reads a bitstream from the inserted disc recording medium 9100 and sends the read bitstream to the image decoding unit 9008 of the codec IC 9002 via the bus 9010. The image decoding unit 9008 performs image decoding processing in the image decoding apparatus according to embodiments of the present invention on the input bitstream and sends the decoded image to an external monitor 9103 via the video interface 9009. Additionally, the codec apparatus 9000 has a network interface 9006 and can connect to an external distribution server 9106 and a portable terminal 9107 via the network 9101. When a user wishes to reproduce an image recorded on the distribution server 9106 or the mobile terminal 9107 instead of an image recorded on the disc recording medium 9100, the network interface 9006 obtains the bitstream from the network 9101 instead of reading the bitstream from the input disc recording medium 9100. Furthermore, if a user wishes to reproduce an image recorded in memory 9004, the image decoding process in the image decoding apparatus according to an embodiment of the present invention is performed on the bitstream recorded in memory 9004.
[0290] When a user wishes to encode and record an image captured by an external camera 9102 in memory 9004, the video interface 9009 inputs the image from the camera 9102 and sends it to the image encoding unit 9007 of the codec IC 9002 via bus 9010. The image encoding unit 9007 performs image encoding processing according to the image encoding apparatus of this invention on the image input via the video interface 9009 and generates a bitstream. Then, the bitstream is sent to memory 9004 via bus 9010. When the user wishes to record the bitstream on the disc recording medium 9100 instead of memory 9004, the optical disc drive 9005 writes the bitstream to the inserted disc recording medium 9100.
[0291] It is also possible to implement a hardware structure that has an image encoding device but no image decoding device, or a hardware structure that has an image decoding device but no image encoding device. Such a hardware structure can be implemented, for example, by replacing the codec IC9002 with an image encoding unit 9007 or an image decoding unit 9008, respectively.
[0292] The processing related to the above encoding and decoding can, of course, be implemented using hardware transmission, storage, and reception devices, and can be implemented through firmware stored in ROM (Read-Only Memory), flash memory, or software such as computers. This firmware or software program can be provided by recording it on a readable recording medium such as a computer, or by providing it from a server via wired or wireless networks, or by providing it as data broadcast via terrestrial wave or satellite digital broadcasting.
[0293] The present invention has been described above based on embodiments. The embodiments are illustrative; various modifications can be made to the combination of these constituent elements and processing steps, and such modifications are also within the scope of the present invention, as will be understood by those skilled in the art.
[0294] Symbol Explanation
[0295] 100 Image encoding device, 101 Block segmentation unit, 102 Inter-frame prediction unit, 103 Intra-frame prediction unit, 104 Decoded image memory, 105 Prediction method determination unit, 106 Residual generation unit, 107 Orthogonal transform / quantization unit, 108 Bit string encoding unit, 109 Inverse quantization / inverse orthogonal transform unit, 110 Decoded image signal overlay unit, 111 Encoded information storage memory, 200 Image decoding device, 201 Bit string decoding unit, 202 Block segmentation unit, 203 Inter-frame prediction unit, 204 Intra-frame prediction unit, 205 Encoded information storage memory, 206 Inverse quantization / inverse orthogonal transform unit, 207 Decoded image signal overlay unit, 208 Decoded image memory.< / poc>
Claims
1. An image coding apparatus that encodes a moving image in blocks and generates a bitstream using inter-frame prediction based on inter-frame prediction information, the image coding apparatus being characterized by comprising: The encoding information storage unit saves the inter-frame prediction information used in the inter-frame prediction of the coded blocks into the historical prediction motion vector candidate list; The spatial merging candidate derivation unit derives spatial merging candidates from the inter-frame prediction information of blocks spatially adjacent to the coded object block, and registers the spatial merging candidates into the merging candidate list. as well as The historical merging candidate derivation unit derives historical merging candidates from the inter-frame prediction information stored in the historical predicted motion vector candidate list, and registers the historical merging candidates into the merging candidate list. The historical merging candidate derivation unit compares the inter-frame prediction information stored in the historical prediction motion vector candidate list with the inter-frame prediction information of the spatial merging candidate. When at least one of the inter-frame prediction mode, the reference index of L0 and L1, and the value of the motion vector of L0 and L1, which are components of the inter-frame prediction information, is different, it is used as the historical merging candidate. For inter-frame prediction information that is earlier than the predetermined number of frames, it does not compare it with the inter-frame prediction information of the spatial merging candidate and is used as the historical merging candidate.
2. An image coding method, which uses inter-frame prediction based on inter-frame prediction information to encode a moving image in blocks and generate a bitstream, the image coding method being characterized by comprising: The encoding information saving step saves the inter-frame prediction information used in the inter-frame prediction of the coded blocks to the historical prediction motion vector candidate list; The spatial merging candidate derivation step involves deriving spatial merging candidates from the inter-frame prediction information of blocks spatially adjacent to the coded object block, and registering the spatial merging candidates in the merging candidate list. as well as The historical merging candidate export step involves exporting historical merging candidates from the inter-frame prediction information stored in the historical predicted motion vector candidate list, and registering the historical merging candidates into the merging candidate list. In the historical merging candidate derivation step, for a predetermined number of inter-frame prediction information stored in the historical predicted motion vector candidate list, starting from the end, a comparison is made with the inter-frame prediction information of the spatial merging candidate. If at least one of the inter-frame prediction mode, the reference index of L0 and L1, and the value of the motion vector of L0 and L1, which are components of the inter-frame prediction information, is different, it is used as the historical merging candidate. For inter-frame prediction information that is earlier than the predetermined number, starting from the end, no comparison is made with the inter-frame prediction information of the spatial merging candidate, and it is used as the historical merging candidate.
3. An image decoding apparatus for decoding an encoded bit string, wherein the encoded bit string is a bit string encoded using inter-frame prediction on a block-by-block basis for a moving image, the image decoding apparatus being characterized in that it comprises: The encoding information storage unit saves the inter-frame prediction information used in the inter-frame prediction of the decoded blocks into the historical prediction motion vector candidate list; The spatial merging candidate derivation unit derives spatial merging candidates from the inter-frame prediction information of blocks spatially adjacent to the decoded object block, and registers the spatial merging candidates in the merging candidate list. as well as The historical merging candidate derivation unit derives historical merging candidates from the inter-frame prediction information stored in the historical predicted motion vector candidate list, and registers the historical merging candidates into the merging candidate list. The historical merging candidate derivation unit compares the inter-frame prediction information stored in the historical prediction motion vector candidate list with the inter-frame prediction information of the spatial merging candidate. When at least one of the inter-frame prediction mode, the reference index of L0 and L1, and the value of the motion vector of L0 and L1, which are components of the inter-frame prediction information, is different, it is used as the historical merging candidate. For inter-frame prediction information that is earlier than the predetermined number of frames, it does not compare it with the inter-frame prediction information of the spatial merging candidate and is used as the historical merging candidate.
4. An image decoding method for decoding an encoded bit string, wherein the encoded bit string is a bit string encoded using inter-frame prediction on a block-by-block basis for a moving image, the image decoding method being characterized by comprising: The encoding information saving step saves the inter-frame prediction information used in the inter-frame prediction of the decoded blocks into the historical prediction motion vector candidate list; The spatial merging candidate derivation step involves deriving spatial merging candidates from the inter-frame prediction information of blocks spatially adjacent to the decoded object block, and registering the spatial merging candidates in the merging candidate list. as well as The historical merging candidate export step involves exporting historical merging candidates from the inter-frame prediction information stored in the historical predicted motion vector candidate list, and registering the historical merging candidates into the merging candidate list. In the historical merging candidate derivation step, for a predetermined number of inter-frame prediction information stored in the historical predicted motion vector candidate list, starting from the end, a comparison is made with the inter-frame prediction information of the spatial merging candidate. If at least one of the inter-frame prediction mode, the reference index of L0 and L1, and the value of the motion vector of L0 and L1, which are components of the inter-frame prediction information, is different, it is used as the historical merging candidate. For inter-frame prediction information that is earlier than the predetermined number, starting from the end, no comparison is made with the inter-frame prediction information of the spatial merging candidate, and it is used as the historical merging candidate.
5. A method for storing a bit stream, comprising generating a bit stream by performing the image encoding method of claim 2, and storing the bit stream.
6. A method for transmitting a bit stream, comprising generating a bit stream by performing the image encoding method of claim 2, and transmitting the bit stream.
Citation Information
Patent Citations
Image encoding apparatus and method, image decoding apparatus and method, and storage medium
CN113055690A
Moving image coding / decoding device using moving compensation inter-frame prediction system employing affine transformation
JP1997172644A
Video image decoding device, video image decoding method, and video image decoding program
JP2013090033A
Image decoding device, image decoding method, and image decoding program
JP2014200023A