Moving picture decoding apparatus and method, and moving picture encoding apparatus and method
By deriving motion information candidates that are spatially and temporally close to the decoded object block, the problem of high processing load in image encoding and decoding is solved, and efficient image encoding and decoding processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JVC KENWOOD CORP
- Filing Date
- 2019-12-20
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies suffer from excessive processing load and reduced coding efficiency due to image transformations during image encoding and decoding.
By deriving motion information candidates that are spatially and temporally close to the decoded object block, and by deriving historical motion information candidates from the motion information of the already decoded block, direct comparison between motion information is avoided, thus reducing the processing load.
It achieves efficient image encoding and decoding processing under low load.
Smart Images

Figure CN116389751B_ABST
Abstract
Description
[0001] This application is a divisional application based on the invention patent with application number 201980050793.9, application date December 20, 2019, applicant JVC Kenwood Corporation, entitled "Motion Image Decoding Device, Motion Image Decoding Method, Motion Image Decoding Program, Motion Image Encoding Device, Motion Image Encoding Method and Motion Image Encoding Program". Technical Field
[0002] This invention relates to image encoding and decoding techniques for segmenting images into blocks and making predictions. Background Technology
[0003] In image encoding and decoding, the image to be processed is divided into a set of pixels of a predetermined number, i.e., blocks, and processed in blocks. By dividing the image into appropriate blocks and setting intra-frame prediction and inter-frame prediction appropriately, encoding efficiency is improved.
[0004] In the encoding / decoding of moving images, encoding efficiency is improved by performing inter-frame prediction based on already encoded / decoded images. Patent Document 1 discloses a technique that applies affine transformation during inter-frame prediction. In moving images, objects often undergo deformations such as magnification, reduction, and rotation; by applying the technique in Patent Document 1, efficient encoding can be achieved.
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent document 1: Japanese Patent Application Publication No. 9-172644. Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] However, the technology in Patent Document 1 involves image transformations, resulting in a high processing load. In view of the above problems, this invention provides a low-load and efficient encoding technology.
[0010] means for solving problems
[0011] To address the aforementioned problems, one aspect of the present invention provides a motion image decoding apparatus, comprising: a spatial motion information candidate deriving unit, which derives spatial motion information candidates from motion information of blocks spatially close to the target block being decoded; a temporal motion information candidate deriving unit, which derives temporal motion information candidates from motion information of blocks temporally close to the target block being decoded; and a historical motion information candidate deriving unit, which derives historical motion information candidates from a memory storing motion information of decoded blocks, wherein the temporal motion information candidates are not compared with either the spatial motion information candidates or the historical motion information candidates.
[0012] In addition, another aspect of the motion image decoding method of the present invention includes the following steps: deriving spatial motion information candidates from motion information of blocks spatially close to the decoding target block; deriving temporal motion information candidates from motion information of blocks temporally close to the decoding target block; and deriving historical motion information candidates from a memory that holds motion information of decoded blocks, wherein the temporal motion information candidates are not compared with any of the spatial motion information candidates and the historical motion information candidates.
[0013] In addition, another aspect of the motion picture decoding program of the present invention enables a computer to function as the following: a spatial motion information candidate deriving unit, which derives spatial motion information candidates from motion information of blocks spatially close to the decoded object block; a temporal motion information candidate deriving unit, which derives temporal motion information candidates from motion information of blocks temporally close to the decoded object block; and a historical motion information candidate deriving unit, which derives historical motion information candidates from a memory storing motion information of decoded blocks, wherein the temporal motion information candidates are not compared with any of the spatial motion information candidates and the historical motion information candidates.
[0014] The effects of the invention
[0015] According to the present invention, high-efficiency image encoding / decoding processing can be achieved with low overhead. Attached Figure Description
[0016] Figure 1 This is a block diagram of an image encoding apparatus according to an embodiment of the present invention;
[0017] Figure 2 This is a block diagram of an image decoding apparatus according to an embodiment of the present invention;
[0018] Figure 3 This is a flowchart illustrating the actions of splitting tree blocks;
[0019] Figure 4 This is a diagram illustrating the process of segmenting an input image into tree blocks;
[0020] Figure 5 This is a diagram illustrating the z-scan.
[0021] Figure 6A This is a diagram showing the segmented shape of the block;
[0022] Figure 6B This is a diagram showing the segmented shape of the block;
[0023] Figure 6C This is a diagram showing the segmented shape of the block;
[0024] Figure 6D This is a diagram showing the segmented shape of the block;
[0025] Figure 6E This is a diagram showing the segmented shape of the block;
[0026] Figure 7 This is a flowchart illustrating the action of dividing a block into four parts;
[0027] Figure 8 This is a flowchart used to illustrate the action of dividing a block into 2 or 3 parts;
[0028] Figure 9 It is a syntax used to describe the shape of block segmentation;
[0029] Figure 10A This is a diagram used to illustrate intra-frame prediction;
[0030] Figure 10B This is a diagram used to illustrate intra-frame prediction;
[0031] Figure 11 This is a diagram used to illustrate the reference block for inter-frame prediction;
[0032] Figure 12 It is the syntax used to describe the prediction pattern of coded blocks;
[0033] Figure 13 It is a graph showing the correspondence between syntactic elements and patterns related to inter-frame prediction;
[0034] Figure 14 This is a diagram used to illustrate affine transformation motion compensation with two control points;
[0035] Figure 15 This is a diagram used to illustrate affine transformation motion compensation with three control points;
[0036] Figure 16 yes Figure 1 A block diagram showing the detailed structure of the inter-frame prediction unit 102;
[0037] Figure 17 yes Figure 16A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 301;
[0038] Figure 18 yes Figure 16 A block diagram showing the detailed structure of the typical merging mode derivation section 302;
[0039] Figure 19 It is used for explanation Figure 16 A flowchart of the normal prediction motion vector pattern export process of the normal prediction motion vector pattern export unit 301.
[0040] Figure 20 This is a flowchart illustrating the typical processing steps for deriving motion vector patterns from predicted data.
[0041] Figure 21 This is a flowchart illustrating the typical steps involved in exporting data using a merge pattern.
[0042] Figure 22 yes Figure 2 A block diagram showing the detailed structure of the inter-frame prediction unit 203;
[0043] Figure 23 yes Figure 22 A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 401;
[0044] Figure 24 yes Figure 22 A block diagram showing the detailed structure of the typical merging mode derivation section 402;
[0045] Figure 25 It is used for explanation Figure 22 A flowchart of the normal prediction motion vector pattern export process of the normal prediction motion vector pattern export unit 401.
[0046] Figure 26 This is a diagram illustrating the initialization / update process of the historical predicted motion vector candidate list;
[0047] Figure 27 This is a flowchart of the same feature confirmation process in the initialization / update process of the historical predicted motion vector candidate list.
[0048] Figure 28 This is a flowchart of the feature shifting process in the initialization / update process of the historical predicted motion vector candidate list.
[0049] Figure 29 This is a flowchart illustrating the steps involved in deriving candidate motion vectors from historical predictions.
[0050] Figure 30 This is a flowchart illustrating the steps involved in exporting historical merge candidates;
[0051] Figure 31A This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0052] Figure 31B This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0053] Figure 31C This is a diagram illustrating an example of the historical predicted motion vector candidate list update process;
[0054] Figure 32 This is a graph used to illustrate motion compensation prediction when the reference image (RefL0Pic) of L0 is at a time before the processing object image (CurPic) in L0 prediction;
[0055] Figure 33 This is a diagram used to illustrate motion compensation prediction when the reference image for L0 prediction is at a time after the image of the object being processed in L0 prediction.
[0056] Figure 34 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction is at a time before the object image being processed, and the reference image for L1 prediction is at a time after the object image being processed.
[0057] Figure 35 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction and the reference image for L1 prediction are at a time before the object image being processed in dual prediction.
[0058] Figure 36 This is a diagram used to illustrate the prediction direction of motion compensation prediction when the reference image for L0 prediction and the reference image for L1 prediction are at a time after the object image being processed in dual prediction.
[0059] Figure 37 This is a diagram illustrating an example of the hardware structure of the encoding / decoding apparatus according to an embodiment of the present invention;
[0060] Figure 38 This is the second embodiment of the present invention. Figure 16 A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 301;
[0061] Figure 39 This is the second embodiment of the present invention. Figure 22 A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 401;
[0062] Figure 40 This is the third embodiment of the present invention. Figure 16A block diagram showing the detailed structure of the typical predicted motion vector mode derivation unit 301;
[0063] Figure 41 This is the third embodiment of the present invention. Figure 22 A block diagram showing the detailed structure of the typical predictive motion vector pattern derivation unit 401. Detailed Implementation
[0064] Define the technologies and technical terms used in this embodiment.
[0065] <Tree Block>
[0066] In this implementation, the image to be encoded / decoded is divided equally into segments of a predetermined size. This unit is defined as a tree block. Figure 4 In this design, the size of the tree block is set to 128×128 pixels, but the size is not limited to this and can be set to any size. The tree blocks, which are processed (corresponding to encoding objects in encoding processing and decoding objects in decoding processing), are switched according to the raster scan order, i.e., from left to right and from top to bottom. Each tree block can be further recursively divided internally. The block that becomes the object of encoding and decoding after recursively dividing the tree block is defined as an encoding block. Furthermore, tree blocks and encoding blocks are collectively referred to as blocks. By performing appropriate block division, efficient encoding can be achieved. The size of the tree block can be a fixed value predetermined in the encoding and decoding devices, or a structure can be adopted where the size of the tree block determined by the encoding device is transmitted to the decoding device. Here, the maximum size of the tree block is set to 128×128 pixels, and the minimum size of the tree block is set to 16×16 pixels. Similarly, the maximum size of the encoding block is set to 64×64 pixels, and the minimum size of the encoding block is set to 4×4 pixels.
[0067] <Prediction Pattern>
[0068] The intra-frame prediction (MODE_INTRA) and inter-frame prediction (MODE_INTER) are switched on a per-processing-object-coded-block basis. The intra-frame prediction (MODE_INTRA) predicts based on the processed image signal of the object image, while the inter-frame prediction predicts based on the image signal of the processed image.
[0069] In the encoding process, the processed image is used to decode the encoded signal to obtain the image, image signal, tree block, block, coded block, etc., and in the decoding process, it is used to decode the image, image signal, tree block, block, coded block, etc.
[0070] The mode that identifies the intra-frame prediction (MODE_INTRA) and inter-frame prediction (MODE_INTER) is defined as the prediction mode (PredMode). The prediction mode (PredMode) is represented by a value for either intra-frame prediction (MODE_INTRA) or inter-frame prediction (MODE_INTER).
[0071] Inter-frame prediction
[0072] In inter-frame prediction, which predicts based on image signals from processed images, multiple processed images can be used as reference pictures. To manage these multiple reference pictures, two reference lists, L0 (reference list 0) and L1 (reference list 1), are defined, each using a reference index to determine the reference picture. In a P-slice, L0 prediction (Pred_L0) can be used. In a B-slice, L0 prediction (Pred_L0), L1 prediction (Pred_L1), and double prediction (Pred_BI) can be used. L0 prediction (Pred_L0) is an inter-frame prediction that references the reference pictures managed by L0, and L1 prediction (Pred_L1) is an inter-frame prediction that references the reference pictures managed by L1. Double prediction (Pred_BI) is an inter-frame prediction that simultaneously performs L0 and L1 predictions and references individual reference pictures managed by each of L0 and L1. The information determining L0 prediction, L1 prediction, and double prediction is defined as the inter-frame prediction mode. Regarding the constants and variables with the subscript LX appended to the output in subsequent processing, the premise is that they will be processed as L0 and L1.
[0073] <Predicting motion vector patterns>
[0074] The predicted motion vector mode is a mode that transmits the index, differential motion vector, inter-frame prediction mode, reference index, and inter-frame prediction information for determining the predicted motion vector and deciding on the target block. The predicted motion vector is derived from the predicted motion vector candidate and the index used to determine the predicted motion vector. The predicted motion vector candidate is derived from the processed blocks adjacent to the target block or from blocks in the processed image that are located at the same position as or near the target block.
[0075] Merge Mode
[0076] The merging mode is as follows: without transmitting differential motion vectors or reference indices, the inter-frame prediction information of the processing object block is derived based on the inter-frame prediction information of processed blocks adjacent to the processing object block or blocks in the processed image that are located at the same position or nearby (neighboring) to the processing object block.
[0077] Processed blocks adjacent to the target block, along with their inter-frame prediction information, are defined as spatial merge candidates. Blocks belonging to the processed image that are located at the same position as or near the target block, along with their inter-frame prediction information derived from that block, are defined as temporal merge candidates. Each merge candidate is registered in a merge candidate list, and the merge candidate used in the prediction of the target block is determined by a merge index.
[0078] <Adjacent blocks>
[0079] Figure 11 This diagram illustrates the reference blocks used to derive inter-frame prediction information in both the predicted motion vector mode and the merge mode. A0, A1, A2, B0, B1, B2, and B3 are processed blocks adjacent to the object block. T0 is a block within the processed image blocks that is located at the same position or in the vicinity (nearby) of the object block in the object image.
[0080] A1 and A2 are blocks located to the left of the processing object's encoding block and adjacent to it. B1 and B3 are blocks located above the processing object's encoding block and adjacent to it. A0, B0, and B2 are blocks located to the lower left, upper right, and upper left of the processing object's encoding block, respectively.
[0081] The details of how adjacent blocks are handled in the prediction motion vector mode and the merging mode are described later.
[0082] <Affine Transformation Motion Compensation>
[0083] Affine transformation motion compensation involves dividing the coded block into sub-blocks of predetermined units and determining motion vectors for each sub-block individually. The motion vectors of each sub-block are derived based on one or more control points. These control points are derived from inter-frame prediction information of processed blocks adjacent to the target block or blocks within the processed image that are located at the same position or nearby (neighboring) to the target block. In this embodiment, the size of the sub-block is set to 4×4 pixels, but the size of the sub-block is not limited to this; motion vectors can also be derived in pixels.
[0084] Figure 14 An example of affine transformation motion compensation with two control points is shown. In this case, the two control points have two parameters: a horizontal component and a vertical component. Therefore, the affine transformation with two control points is called a four-parameter affine transformation. Figure 14 CP1 and CP2 are control points.
[0085] Figure 15An example of affine transformation motion compensation with three control points is shown. In this case, the three control points have two parameters: a horizontal component and a vertical component. Therefore, the affine transformation with three control points is called a six-parameter affine transformation. Figure 15 CP1, CP2, and CP3 are control points.
[0086] Affine transformation motion compensation can be used in either the predictive motion vector mode or the merging mode. A mode in which affine transformation motion compensation is applied in the predictive motion vector mode is defined as the sub-block predictive motion vector mode, and a mode in which affine transformation motion compensation is applied in the merging mode is defined as the sub-block merging mode.
[0087] <Syntax of Inter-Frame Prediction>
[0088] use Figure 12 and Figure 13 The syntax related to inter-frame prediction is explained.
[0089] Figure 12 `merge_flag` indicates whether the processed object coding block is set to merge mode or predictive motion vector mode. `merge_affine_flag` indicates whether to apply sub-block merging mode in the processed object coding block in merge mode. `inter_affine_flag` indicates whether to apply sub-block predictive motion vector mode in the processed object coding block in predictive motion vector mode. `cu_affine_type_flag` is a flag used to determine the number of control points in sub-block predictive motion vector mode.
[0090] Figure 13 The values of each syntactic element and their corresponding prediction methods are shown. `merge_flag = 1` and `merge_affine_flag = 0` correspond to the normal merge mode. The normal merge mode is a merge mode that is not a sub-block merge. `merge_flag = 1` and `merge_affine_flag = 1` correspond to the sub-block merge mode. `merge_flag = 0` and `inter_affine_flag = 0` correspond to the normal predicted motion vector mode. The normal predicted motion vector mode is a predicted motion vector merge that is not a sub-block predicted motion vector mode. `merge_flag = 0` and `inter_affine_flag = 1` correspond to the sub-block predicted motion vector mode. When `merge_flag = 0` and `inter_affine_flag = 1`, `cu_affine_type_flag` is further passed to determine the number of control points.
[0091] <poc>
[0092] The Picture Order Count (POC) is a variable associated with the picture to be encoded, and it is set to an incrementing value of 1 corresponding to the output order of the pictures. Based on the POC value, it is possible to determine whether the pictures are the same, the order of the pictures in the output sequence, and the distance between the exported pictures. For example, if two pictures have the same POC value, they can be considered the same picture. If two pictures have different POC values, the picture with the smaller POC value is determined to be the first image output, and the difference between the POC values of the two pictures represents the distance between them along the timeline.
[0093] (First Implementation)
[0094] The image encoding apparatus 100 and the image decoding apparatus 200 according to the first embodiment of the present invention will be described.
[0095] Figure 1 This is a block diagram of the image encoding apparatus 100 according to the first embodiment. The image encoding apparatus 100 of the embodiment includes a block segmentation unit 101, an inter-frame prediction unit 102, an intra-frame prediction unit 103, a decoded image memory 104, a prediction method determination unit 105, a residual generation unit 106, an orthogonal transform / quantization unit 107, a bit string encoding unit 108, an inverse quantization / inverse orthogonal transform unit 109, a decoded image signal overlap unit 110, and an encoding information storage memory 111.
[0096] The block segmentation unit 101 recursively segments the input image to generate coded blocks. The block segmentation unit 101 includes a 4-segmentation unit and a 2-3 segmentation unit. The 4-segmentation unit segments the blocks to be segmented in both the horizontal and vertical directions, while the 2-3 segmentation unit segments the blocks to be segmented in either the horizontal or vertical direction. The block segmentation unit 101 sets the generated coded blocks as processing target coded blocks and provides the image signal of the processing target coded blocks to the inter-frame prediction unit 102, the intra-frame prediction unit 103, and the residual generation unit 106. Furthermore, the block segmentation unit 101 provides information representing the determined recursive segmentation structure to the bit string encoding unit 108. Detailed operation of the block segmentation unit 101 will be described later.
[0097] The inter-frame prediction unit 102 performs inter-frame prediction for the target coded block. Based on the inter-frame prediction information stored in the encoded information storage memory 111 and the decoded image signal stored in the decoded image memory 104, the inter-frame prediction unit 102 derives multiple candidate inter-frame prediction information, selects a suitable inter-frame prediction mode from the derived candidates, and provides the selected inter-frame prediction mode and the corresponding prediction image signal to the prediction method determination unit 105. The detailed structure and operation of the inter-frame prediction unit 102 will be described later.
[0098] The intra-prediction unit 103 performs intra-prediction of the target coded block. The intra-prediction unit 103 uses the decoded image signal stored in the decoded image memory 104 as a reference pixel, and generates a predicted image signal based on intra-prediction information such as the intra-prediction mode stored in the encoding information storage memory 111. In intra-prediction, the intra-prediction unit 103 selects a suitable intra-prediction mode from multiple intra-prediction modes and provides the selected intra-prediction mode and the predicted image signal corresponding to the selected intra-prediction mode to the prediction method determination unit 105.
[0099] Figure 10A and Figure 10B An example of intra-frame prediction is shown. Figure 10A This diagram illustrates the correspondence between the prediction direction of intra-prediction and the intra-prediction mode number. For example, intra-prediction mode 50 generates an intra-predicted image by copying reference pixels in the vertical direction. Intra-prediction mode 1 is a DC mode, which sets the pixel values of all pixels in the processed object block to the average of the reference pixels. Intra-prediction mode 0 is a Planar mode (two-dimensional mode), which generates a two-dimensional intra-prediction image based on reference pixels in both the vertical and horizontal directions. Figure 10B This is an example of an intra-prediction image generated in intra-prediction mode 40. The intra-prediction unit 103 copies the value of a reference pixel in the direction indicated by the intra-prediction mode to each pixel of the processing object block. If the reference pixel in the intra-prediction mode is not at an integer position, the intra-prediction unit 103 determines the reference pixel value by interpolation based on the reference pixel values at surrounding integer positions.
[0100] The decoded image memory 104 stores the decoded image generated by the decoded image signal overlap portion 110. The decoded image memory 104 provides the stored decoded image to the inter-frame prediction unit 102 and the intra-frame prediction unit 103.
[0101] The prediction method determination unit 105 evaluates each of intra-frame prediction and inter-frame prediction by using coding information and the amount of coding of the residual, the amount of distortion between the predicted image signal and the image signal to be processed, etc., to determine the optimal prediction mode. In the case of intra-frame prediction, the prediction method determination unit 105 provides intra-frame prediction information such as the intra-frame prediction mode as coding information to the bit string coding unit 108. In the case of inter-frame prediction merging mode, the prediction method determination unit 105 provides inter-frame prediction information such as merging index and information indicating whether it is a sub-block merging mode (sub-block merging flag) as coding information to the bit string coding unit 108. In the case of inter-frame prediction motion vector prediction mode, the prediction method determination unit 105 provides inter-frame prediction information such as inter-frame prediction mode, prediction motion vector index, L0 and L1 reference indices, differential motion vector, and information indicating whether it is a sub-block prediction motion vector mode (sub-block prediction motion vector flag) as coding information to the bit string coding unit 108. In addition, the prediction method determination unit 105 provides the determined coding information to the coding information storage memory 111. The prediction method determination unit 105 provides the predicted image signal to the residual generation unit 106 and the decoded image signal overlap unit 110.
[0102] The residual generation unit 106 generates a residual by subtracting the predicted image signal from the image signal of the object being processed, and provides it to the orthogonal transformation / quantization unit 107.
[0103] The orthogonal transformation / quantization unit 107 performs orthogonal transformation and quantization on the residual according to the quantization parameters to generate the orthogonal transformation / quantization residual, and provides the generated residual to the bit string encoding unit 108 and the inverse quantization / inverse orthogonal transformation unit 109.
[0104] In addition to information about sequences, images, stripes, and coding block units, the bit string encoding unit 108 encodes coding information corresponding to the prediction method determined by the prediction method determination unit 105 for each coding block. Specifically, the bit string encoding unit 108 encodes the prediction mode PredMode for each coding block. When the prediction mode is inter-frame prediction (MODE_INTER), the bit string encoding unit 108 encodes coding information (inter-frame prediction information) such as the flag indicating whether it is a merging mode, the sub-block merging flag, the merging index when it is a merging mode, the inter-frame prediction mode when it is not a merging mode, the predicted motion vector index, information related to the differential motion vector, and the sub-block predicted motion vector flag, according to the prescribed syntax (bit string syntax rules), to generate a first bit string. When the prediction mode is intra-frame prediction (MODE_INTRA), the bit string encoding unit encodes coding information (intra-frame prediction information) such as the intra-frame prediction mode, according to the prescribed syntax (bit string syntax rules), to generate a first bit string. Furthermore, the bit string encoding unit 108 performs entropy encoding on the orthogonal transform and quantized residuals according to a prescribed syntax to generate a second bit string. The bit string encoding unit 108 then multiplexes the first and second bit strings according to a prescribed syntax to output a bit stream.
[0105] The inverse quantization / inverse quadrature transformation unit 109 performs inverse quantization and inverse quadrature transformation on the residual provided by the quadrature transformation / quantization unit 107 to calculate the residual, and provides the calculated residual to the decoded image signal overlap unit 110.
[0106] The decoded image signal overlay unit 110 overlays the predicted image signal corresponding to the decision of the prediction method determination unit 105 with the residual obtained by inverse quantization and inverse quadrature transformation performed by the inverse quantization / inverse quadrature transformation unit 109 to generate a decoded image, which is then stored in the decoded image memory 104. Alternatively, the decoded image signal overlay unit 110 may also perform filtering processing on the decoded image to reduce block distortion and other distortions caused by encoding before storing it in the decoded image memory 104.
[0107] The encoding information storage memory 111 stores encoding information such as the prediction mode (inter-frame prediction or intra-frame prediction) determined by the prediction method determination unit 105. In the case of inter-frame prediction, the encoding information stored in the encoding information storage memory 111 includes the determined motion vector, the reference index of reference lists L0 and L1, and the historical predicted motion vector candidate list. Furthermore, in the case of inter-frame prediction merging mode, the encoding information stored in the encoding information storage memory 111 includes, in addition to the above information, inter-frame prediction information such as a merging index and information indicating whether it is a sub-block merging mode (sub-block merging flag). Furthermore, in the case of inter-frame prediction predicting motion vector mode, the encoding information stored in the encoding information storage memory 111 includes, in addition to the above information, inter-frame prediction information such as the inter-frame prediction mode, predicted motion vector index, differential motion vector, and information indicating whether it is a sub-block predicting motion vector mode (sub-block predicting motion vector flag). In the case of intra-frame prediction, the encoding information stored in the encoding information storage memory 111 includes intra-frame prediction information such as the determined intra-frame prediction mode.
[0108] Figure 2 It means and Figure 1 The image encoding apparatus is shown in the block diagram of the image decoding apparatus according to the embodiments of the present invention. The image decoding apparatus of the embodiments includes a bit string decoding unit 201, a block segmentation unit 202, an inter-frame prediction unit 203, an intra-frame prediction unit 204, an encoded information storage memory 205, an inverse quantization / inverse quadrature transformation unit 206, a decoded image signal overlap unit 207, and a decoded image memory 208.
[0109] Figure 2 Image decoding device decoding processing and Figure 1 The decoding processing is internally configured within the image encoding device, therefore Figure 2 The encoded information storage memory 205, the inverse quantization / inverse quadrature transform unit 206, the decoded image signal overlap unit 207, and the decoded image memory 208 each have structures that are consistent with... Figure 1 The functions of each structure of the image encoding device, including the encoding information storage memory 111, the inverse quantization / inverse quadrature transformation unit 109, the decoded image signal overlap unit 110, and the decoded image memory 104, are respectively defined.
[0110] The bitstream provided to the bitstream decoding unit 201 is separated according to prescribed syntax rules. The bitstream decoding unit 201 decodes the separated first bitstream to obtain information on the sequence, image, stripe, coded block unit, and coded block unit encoding information. Specifically, the bitstream decoding unit 201 decodes the prediction mode PredMode on a coded block basis. The prediction mode PredMode is determined to be either inter-frame prediction (MODE_INTER) or intra-frame prediction (MODE_INTRA). When the prediction mode is inter-frame prediction (MODE_INTER), the bitstream decoding unit 201 decodes the encoding information (inter-frame prediction information) related to the flag indicating whether it is a merge mode, the merge index in the case of merge mode, the sub-block merge flag, the inter-frame prediction mode in the case of prediction motion vector mode, the prediction motion vector index, the differential motion vector, and the sub-block prediction motion vector flag, according to the prescribed syntax. The encoding information (inter-frame prediction information) is then provided to the encoding information storage memory 205 via the inter-frame prediction unit 203 and the block segmentation unit 202. When the prediction mode is intra-prediction (MODE_INTRA), the coded information (intra-prediction information) such as the intra-prediction mode is decoded according to the prescribed syntax, and the coded information (intra-prediction information) is provided to the coded information storage memory 205 via the inter-prediction unit 203 or the intra-prediction unit 204 and the block segmentation unit 202. The bit string decoding unit 201 decodes the separated second bit string, calculates the residual after orthogonal transformation / quantization, and provides the residual after orthogonal transformation / quantization to the inverse quantization / inverse orthogonal transformation unit 206.
[0111] When the prediction mode PredMode of the coded block being processed is the prediction motion vector mode in inter-frame prediction (MODE_INTER), the inter-frame prediction unit 203 uses the coded information of the decoded image signal stored in the coded information storage memory 205 to derive multiple candidate prediction motion vectors, and registers the derived candidate prediction motion vectors in the candidate prediction motion vector list described later. The inter-frame prediction unit 203 selects the prediction motion vector corresponding to the prediction motion vector index provided by the bit string decoding unit 201 from the multiple candidate prediction motion vectors registered in the candidate prediction motion vector list, calculates the motion vector based on the differential motion vector decoded by the bit string decoding unit 201 and the selected prediction motion vector, and stores the calculated motion vector along with other coded information in the coded information storage memory 205. Here, the encoding information of the coded blocks to be provided / saved includes the prediction mode PredMode, flags indicating whether L0 and L1 prediction are used (predFlagL0[xP][yP], predFlagL1[xP][yP]), reference indices for L0 and L1 (refIdxL0[xP][yP], refIdxL1[xP][yP]), motion vectors for L0 and L1 (mvL0[xP][yP], mvL1[xP][yP]), etc. Here, xP and yP represent the indices of the top-left pixel position of the coded block within the image. When the prediction mode PredMode is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode is L0 prediction (Pred_L0), the flag predFlagL0 indicating whether L0 prediction is used is 1, and the flag predFlagL1 indicating whether L1 prediction is used is 0. When the inter-frame prediction mode is L1 prediction (Pred_L1), the flag predFlagL0 indicating whether to use L0 prediction is 0, and the flag predFlagL1 indicating whether to use L1 prediction is 1. When the inter-frame prediction mode is double prediction (Pred_BI), both the flag predFlagL0 indicating whether to use L0 prediction and the flag predFlagL1 indicating whether to use L1 prediction are 1. Furthermore, when the prediction mode PredMode of the coded block being processed is the merge mode in inter-frame prediction (MODE_INTER), merge candidates are derived.Using the encoding information of the decoded coded blocks stored in the encoding information storage memory 205, multiple merging candidates are derived and registered in the merging candidate list (described later). From the multiple merging candidates registered in the merging candidate list, a merging candidate corresponding to the merging index provided by the bit string decoding unit 201 is selected. Inter-frame prediction information, including the flags predFlagL0[xP][yP], predFlagL1[xP][yP] indicating whether to utilize the selected merging candidate for L0 and L1 prediction, the reference indices refIdxL0[xP][yP], refIdxL1[xP][yP] for L0 and L1, and the motion vectors mvL0[xP][yP], mvL1[xP][yP] for L0 and L1, are stored in the encoding information storage memory 205. Here, xP and yP are indices representing the position of the top-left pixel of the coded block within the image. The detailed structure and operation of the inter-frame prediction unit 203 will be described later.
[0112] When the prediction mode PredMode of the encoded block being processed is intra-prediction (MODE_INTRA), the intra-prediction unit 204 performs intra-prediction. The intra-prediction mode is included in the encoded information decoded by the bit string decoding unit 201. Based on the intra-prediction mode included in the decoded information decoded by the bit string decoding unit 201, the intra-prediction unit 204 generates a predicted image signal based on the decoded image signal stored in the decoded image memory 208, and provides the generated predicted image signal to the decoded image signal overlay unit 207. Since the intra-prediction unit 204 corresponds to the intra-prediction unit 103 of the image encoding apparatus 100, it performs the same processing as the intra-prediction unit 103.
[0113] The inverse quantization / inverse quadrature transformation unit 206 performs inverse quadrature transformation and inverse quantization on the residual after quadrature transformation / quantization decoded by the bit string decoding unit 201 to obtain the residual after inverse quadrature transformation / inverse quantization.
[0114] The decoded image signal overlay unit 207 decodes the image signal by overlaying the predicted image signal obtained by inter-frame prediction by inter-frame prediction unit 203 or intra-frame prediction by intra-frame prediction unit 204, and the residual after inverse quadrature transformation / inverse quantization by inverse quantization / inverse quadrature transformation unit 206, and stores the decoded image signal in the decoded image memory 208. When storing the image in the decoded image memory 208, the decoded image signal overlay unit 207 can also perform filtering processing, such as reducing block distortion caused by encoding, before storing the image in the decoded image memory 208.
[0115] Next, the operation of the block segmentation unit 101 in the image encoding device 100 will be explained. Figure 3 This is a flowchart illustrating the process of segmenting an image into tree blocks and further segmenting each tree block. First, the input image is segmented into tree blocks of a predetermined size (step S1001). For each tree block, it is scanned in a predetermined order, i.e., raster scan order (step S1002), and the interior of the tree block to be processed is segmented (step S1003).
[0116] Figure 7 This is a flowchart illustrating the detailed actions of the segmentation process in step S1003. First, it is determined whether to divide the block of the object to be processed into 4 parts (step S1101).
[0117] If it is determined that the processing object block 4 should be divided, the processing object block 4 is divided (step S1102). For each block obtained by dividing the processing object block, it is scanned in the Z-scan order, that is, the order of top left, top right, bottom left, and bottom right (step S1103). Figure 5 This is an example of Z-scan order. Figure 6A 601 is an example of processing object block 4 after splitting. Figure 6A The numbers 0 to 3 in step 601 indicate the processing order. Then, for each block divided in step S1101, the process is executed recursively. Figure 7 The segmentation process (step S1104).
[0118] If it is determined that the object block to be processed will not be divided into 4 parts, then 2-3 parts will be performed (step S1105).
[0119] Figure 8 This is a flowchart showing the detailed actions of the 2-3 segmentation process in step S1105. First, it is determined whether to perform 2-3 segmentation on the block to be processed, that is, whether to perform either 2-segmentation or 3-segmentation (step S1201).
[0120] If it is determined that the block to be processed will not be split into 2-3 segments, that is, if it is determined that no splitting will be performed, the splitting process ends (step S1211). That is, for blocks obtained by recursive splitting, no further recursive splitting is performed.
[0121] If it is determined that the block of the object to be processed should be divided into 2-3, then it is determined whether to further divide the block of the object to be processed into 2 (step S1202).
[0122] If it is determined that the processing object block should be divided into two parts, it is determined whether to divide the processing object block vertically (step S1203). Based on the result, the processing object block is divided into two parts vertically (step S1204), or the processing object block is divided into two parts horizontally (step S1205). As a result of step S1204, the processing object block is as follows: Figure 6B As shown in 602, it is divided into two parts, upper and lower (vertical direction). As a result of step S1205, the processed object block is as follows: Figure 6D As shown in 604, it is divided into two parts, left and right (horizontal direction).
[0123] In step S1202, if it is not determined that the processing object block is to be divided into two parts, that is, if it is determined that it is to be divided into three parts, then it is determined whether to divide the processing object block into top, middle, and bottom (vertical direction) (step S1206). Based on this result, the processing object block is divided into three parts in the top, middle, and bottom (vertical direction) (step S1207), or the processing object block is divided into three parts in the left, middle, and right (horizontal direction) (step S1208). In the result of step S1207, the processing object block is as follows: Figure 6C As shown in 603, it is divided into three parts: upper, middle, and lower (vertical direction). In the result of step S1208, the processed object block is as follows: Figure 6E As shown in 605, it is divided into three parts: left, middle, and right (horizontal direction).
[0124] After executing any one of steps S1204, S1205, S1207, or S1208, the blocks into which the processing object block is divided are scanned in order from left to right and from top to bottom (step S1209). Figures 6B to 6E The numbers 0 to 2, from 602 to 605, indicate the processing order. For each segmented block, the process is executed recursively. Figure 8 The 2-3 segmentation process (step S1210).
[0125] The recursive block partitioning described here can also limit whether partitioning is necessary based on the number of partitions or the size of the block being processed. The information limiting whether partitioning is necessary can be implemented by pre-agreeing between the encoding and decoding devices without transmitting the information, or by the encoding device determining whether partitioning is necessary and recording the information in a bit string before transmitting it to the decoding device.
[0126] When a block is divided, the block before the division is called the parent block, and the blocks after the division are called child blocks.
[0127] Next, the operation of the block segmentation unit 202 in the image decoding apparatus 200 will be described. The block segmentation unit 202 segments tree blocks according to the same processing steps as the block segmentation unit 101 in the image encoding apparatus 100. However, the difference is that in the block segmentation unit 101 of the image encoding apparatus 100, the optimal block segmentation shape is determined by applying optimization methods such as optimal shape estimation or distortion rate optimization based on image recognition. In contrast, the block segmentation unit 202 in the image decoding apparatus 200 determines the block segmentation shape by decoding the block segmentation information recorded in the bit string.
[0128] Figure 9 The syntax (bit string syntax rules) related to block partitioning in the first embodiment is shown. `coding_quadtree()` represents the syntax involved in the 4-partitioning of the block. `multi_type_tree()` represents the syntax involved in the 2-partitioning or 3-partitioning of the block. `qt_split` is a flag indicating whether the block is 4-partitioned. When the block is 4-partitioned, `qt_split` = 1; otherwise, `qt_split` = 0. In the case of 4-partitioning (`qt_split` = 1), the 4-partitioned blocks are recursively partitioned into 4 parts (`coding_quadtree(0), coding_quadtree(1), coding_quadtree(2), coding_quadtree(3)`, where 0 to 3 correspond to... Figure 6A (601). Without a 4-splitting operation (qt_split = 0), subsequent splits are determined according to multi_type_tree(). mtt_split is a flag indicating whether further splitting is required. Furthermore, if splitting is required (mtt_split = 1), the flag indicating whether the split is vertical or horizontal is transmitted, namely mtt_split_vertical, and the flag indicating whether to perform a 2-splitting or 3-splitting operation, namely mtt_split_binary, is transmitted. mtt_split_vertical = 1 indicates a vertical split, and mtt_split_vertical = 0 indicates a horizontal split. mtt_split_binary = 1 indicates a 2-splitting operation, and mtt_split_binary = 0 indicates a 3-splitting operation. In the case of a 2-splitting operation (mtt_split_binary = 1), the blocks after the 2-splitting are recursively split (multi_type_tree(0), multi_type_tree(1), where the 0 to 1 of the independent variable corresponds to Figures 6B to 6D (Numbers 602 or 604). In the case of 3-partition (mtt_split_binary = 0), the 3-partitioned blocks are recursively split (multi_type_tree(0), multi_type_tree(1), multi_type_tree(2), 0 to 2 correspond to...). Figure 6B 603 or Figure 6E (Number 605). Hierarchical block splitting is performed by recursively calling multi_type_tree until mtt_split = 0.
[0129] Inter-frame prediction
[0130] The inter-frame prediction method in the implementation method Figure 1 The inter-frame prediction unit 102 of the image coding apparatus and Figure 2 It is implemented in the inter-frame prediction unit 203 of the image decoding device.
[0131] The inter-frame prediction method according to the implementation method is described with reference to the accompanying drawings. The inter-frame prediction method is implemented in either the encoding or decoding process on a block-by-block basis.
[0132] <Explanation of the inter-frame prediction unit 102 on the coding side>
[0133] Figure 16 It is shown Figure 1 A diagram showing the detailed structure of the inter-frame prediction unit 102 of the image coding apparatus. The normal prediction motion vector pattern derivation unit 301 derives multiple normal prediction motion vector candidates to select a prediction motion vector and calculates the difference motion vector between the selected prediction motion vector and the detected motion vector. The detected inter-frame prediction pattern, reference index, motion vector, and calculated difference motion vector constitute the inter-frame prediction information for the normal prediction motion vector pattern. This inter-frame prediction information is provided to the inter-frame prediction pattern determination unit 305. The detailed structure and processing of the normal prediction motion vector pattern derivation unit 301 will be described later.
[0134] In the normal merging mode derivation unit 302, multiple normal merging candidates are derived, and a normal merging candidate is selected to obtain inter-frame prediction information for the normal merging mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305. The detailed structure and processing of the normal merging mode derivation unit 302 will be described later.
[0135] In the sub-block prediction motion vector mode derivation unit 303, multiple sub-block prediction motion vector candidates are derived to select a sub-block prediction motion vector, and the difference motion vector between the selected sub-block prediction motion vector and the detected motion vector is calculated. The detected inter-frame prediction mode, reference index, motion vector, and calculated difference motion vector constitute the inter-frame prediction information of the sub-block prediction motion vector mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305.
[0136] In the sub-block merging mode derivation unit 304, multiple sub-block merging candidates are derived, and a sub-block merging candidate is selected to obtain inter-frame prediction information for the sub-block merging mode. This inter-frame prediction information is provided to the inter-frame prediction mode determination unit 305.
[0137] The inter-frame prediction mode determination unit 305 determines inter-frame prediction information based on the inter-frame prediction information provided by the normal prediction motion vector mode derivation unit 301, the normal merging mode derivation unit 302, the sub-block prediction motion vector mode derivation unit 303, and the sub-block merging mode derivation unit 304. The inter-frame prediction information corresponding to the determination result is provided from the inter-frame prediction mode determination unit 305 to the motion compensation prediction unit 306.
[0138] The motion compensation prediction unit 306 performs inter-frame prediction on the reference image signal stored in the decoded image memory 104 based on the determined inter-frame prediction information. The detailed structure and processing of the motion compensation prediction unit 306 will be described later.
[0139] <Explanation of the inter-frame prediction unit 203 on the decoding side>
[0140] Figure 22 It is shown Figure 2 A diagram showing the detailed structure of the inter-frame prediction unit 203 of the image decoding device.
[0141] Normally, the predictive motion vector mode derivation unit 401 derives multiple normally predicted motion vector candidates to select a predicted motion vector, and calculates the sum of the selected predicted motion vector and the decoded differential motion vector as the motion vector. The decoded inter-frame prediction mode, reference index, and motion vector constitute the inter-frame prediction information of the normally predicted motion vector mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408. The detailed structure and processing of the normally predicted motion vector mode derivation unit 401 will be described later.
[0142] In the normal merging mode derivation unit 402, multiple normal merging candidates are derived to select a normal merging candidate, thereby obtaining inter-frame prediction information for the normal merging mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408. The detailed structure and processing of the normal merging mode derivation unit 402 will be described later.
[0143] In the sub-block prediction motion vector mode derivation unit 403, multiple sub-block prediction motion vector candidates are derived to select a sub-block prediction motion vector. The selected sub-block prediction motion vector is calculated as the sum of the selected sub-block prediction motion vector and the decoded differential motion vector, which is then used as the motion vector. The decoded inter-frame prediction mode, reference index, and motion vector constitute the inter-frame prediction information of the sub-block prediction motion vector mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408.
[0144] In the sub-block merging mode derivation unit 404, multiple sub-block merging candidates are derived to select a sub-block merging candidate, thereby obtaining inter-frame prediction information for the sub-block merging mode. This inter-frame prediction information is provided to the motion compensation prediction unit 406 via switch 408.
[0145] In the motion compensation prediction unit 406, inter-frame prediction is performed on the reference image signal stored in the decoded image memory 208 based on the determined inter-frame prediction information. The detailed structure and processing of the motion compensation prediction unit 406 are the same as those of the motion compensation prediction unit 306 on the encoding side.
[0146] <Typically predicted motion vector pattern derivation part (usually AMVP)>
[0147] Figure 17 The normal prediction motion vector pattern derivation unit 301 includes a spatial prediction motion vector candidate derivation unit 321, a time prediction motion vector candidate derivation unit 322, a historical prediction motion vector candidate derivation unit 323, a prediction motion vector candidate supplementation unit 325, a normal motion vector detection unit 326, a prediction motion vector candidate selection unit 327, and a motion vector subtraction unit 328.
[0148] Figure 23 The general prediction motion vector pattern derivation unit 401 includes a spatial prediction motion vector candidate derivation unit 421, a time prediction motion vector candidate derivation unit 422, a historical prediction motion vector candidate derivation unit 423, a prediction motion vector candidate supplementation unit 425, a prediction motion vector candidate selection unit 426, and a motion vector addition unit 427.
[0149] Use respectively Figure 19 , Figure 25 The flowchart describes the processing steps of the normal prediction motion vector pattern derivation unit 301 on the encoding side and the normal prediction motion vector pattern derivation unit 401 on the decoding side. Figure 19 This is a flowchart illustrating the normal predicted motion vector pattern derivation processing steps of the normal motion vector pattern derivation unit 301 based on the encoding side. Figure 25 This is a flowchart illustrating the normal predicted motion vector pattern export processing steps of the normal motion vector pattern export unit 401 based on the decoding side.
[0150] <Commonly Predictive Motion Vector Pattern Derivation Unit (typically AMVP): Explanation of the Encoding Side>
[0151] refer to Figure 19 The typical steps for deriving predicted motion vector patterns on the encoding side are explained. Figure 19 The instructions for the processing steps sometimes omit... Figure 19 The word "usually" is shown.
[0152] First, the typical motion vector detection unit 326 detects the typical motion vector for each inter-frame prediction mode and reference index. Figure 19 Step S100).
[0153] Next, the spatial prediction motion vector candidate derivation unit 321, the temporal prediction motion vector candidate derivation unit 322, the historical prediction motion vector candidate derivation unit 323, the prediction motion vector candidate supplementation unit 325, the prediction motion vector candidate selection unit 327, and the motion vector subtraction unit 328 calculate, for each L0 and L1, the differential motion vector of the motion vector used in the inter-frame prediction of the normal prediction motion vector mode. Figure 19 Steps S101 to S106). Specifically, when the prediction mode PredMode of the processed object block is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode is L0 prediction (Pred_L0), the candidate list of predicted motion vectors for L0, mvpListL0, is calculated, the predicted motion vector mvpL0 is selected, and the differential motion vector mvdL0 of the motion vector mvL0 of L0 is calculated. When the inter-frame prediction mode of the processed object block is L1 prediction (Pred_L1), the candidate list of predicted motion vectors for L1, mvpListL1, is calculated, the predicted motion vector mvpL1 is selected, and the differential motion vector mvdL1 of the motion vector mvL1 of L1 is calculated. When the inter-frame prediction mode for processing object blocks is dual prediction (Pred_BI), L0 prediction and L1 prediction are performed simultaneously. The candidate list of predicted motion vectors for L0, mvpListL0, is calculated. The predicted motion vector of L0, mvpL0, is selected. The differential motion vector of L0, mvL0, is calculated. The candidate list of predicted motion vectors for L1, mvpListL1, is calculated. The predicted motion vector of L1, mvpL1, is calculated. The differential motion vector of L1, mvL1, is calculated.
[0154] Differential motion vector calculations are performed separately for L0 and L1, but the process is common to both. Therefore, in the following explanation, L0 and L1 will be represented as a common LX. In the calculation of the differential motion vector for L0, X of LX is 0, and in the calculation of the differential motion vector for L1, X of LX is 1. Furthermore, in the calculation of the differential motion vector for LX, if information from another list is referenced instead of LX, this other list will be represented as LY.
[0155] When using the motion vector mvLX of LX ( Figure 19 Step S102: Yes), calculate the candidate predicted motion vectors of LX, and construct the candidate predicted motion vector list mvpListLX( Figure 19 Step S103). Multiple candidate predicted motion vectors are derived from the spatial predicted motion vector candidate derivation unit 321, the temporal predicted motion vector candidate derivation unit 322, the historical predicted motion vector candidate derivation unit 323, and the predicted motion vector candidate supplementation unit 325 in the normal predicted motion vector pattern derivation unit 301, constructing a predicted motion vector candidate list mvpListLX. (About...) Figure 19 The detailed processing steps of step S103 are as follows: Figure 20 The flowchart is described later.
[0156] Next, the predicted motion vector candidate selection unit 327 selects the predicted motion vector mvpLX of LX from the predicted motion vector candidate list mvpListLX of LX. Figure 19 Step S104). Here, in the candidate list of predicted motion vectors mvpListLX, a certain element (the i-th element counting from 0) is represented as mvpListLX[i]. Calculate each differential motion vector, which is the difference between the motion vector mvLX and the candidate mvpListLX[i] of each predicted motion vector stored in the candidate list of predicted motion vectors mvpListLX. For each element (predicted motion vector candidate) in the candidate list of predicted motion vectors mvpListLX, calculate the coding amount when encoding these differential motion vectors. Then, among the elements registered in the candidate list of predicted motion vectors mvpListLX, select the candidate mvpListLX[i] of the predicted motion vector with the smallest coding amount as the predicted motion vector mvpLX, and obtain the index i. If there are multiple candidates for the predicted motion vector that will become the smallest generated code amount in the candidate list of predicted motion vectors mvpListLX, the candidate mvpListLX[i] represented by the smallest index i in the candidate list of predicted motion vectors mvpListLX is selected as the best predicted motion vector mvpLX, and that index i is obtained.
[0157] Next, the motion vector subtraction unit 328 subtracts the selected predicted motion vector mvpLX of LX from the motion vector mvLX of LX, denoted as mvdLX = mvLX - mvpLX, to calculate the differential motion vector mvdLX of LX. Figure 19 Step S105).
[0158] <Typical Predictive Motion Vector Pattern Derivation Unit (Typical AMVP): Decoding Side Explanation>
[0159] Next, refer to Figure 25 The typical prediction motion vector mode processing steps on the decoding side are explained. On the decoding side, the spatial prediction motion vector candidate derivation unit 421, the temporal prediction motion vector candidate derivation unit 422, the historical prediction motion vector candidate derivation unit 423, and the prediction motion vector candidate supplementation unit 425 calculate, for each L0 and L1, the motion vectors used in inter-frame prediction in the typical prediction motion vector mode. Figure 25 Steps S201 to S206). Specifically, when the prediction mode PredMode of the processing object block is inter-frame prediction (MODE_INTER) and the inter-frame prediction mode of the processing object block is L0 prediction (Pred_L0), the candidate list of predicted motion vectors for L0, mvpListL0, is calculated, the predicted motion vector mvpL0 is selected, and the motion vector mvL0 of L0 is calculated. When the inter-frame prediction mode of the processing object block is L1 prediction (Pred_L1), the candidate list of predicted motion vectors for L1, mvpListL1, is calculated, the predicted motion vector mvpL1 is selected, and the motion vector mvL1 of L1 is calculated. When the inter-frame prediction mode for processing object blocks is dual prediction (Pred_BI), L0 prediction and L1 prediction are performed simultaneously. The candidate list of predicted motion vectors for L0, mvpListL0, is calculated. The predicted motion vector of L0, mvpL0, is selected, and the motion vector of L0, mvL0, is calculated. The candidate list of predicted motion vectors for L1, mvpListL1, is calculated, and the predicted motion vector of L1, mvpL1, is calculated. The motion vector of L1, mvL1, is calculated separately.
[0160] Similar to the encoding side, motion vector calculations are also performed on L0 and L1 separately on the decoding side, but L0 and L1 are processed in a common manner. Therefore, in the following description, L0 and L1 are denoted as a common LX. LX represents the inter-prediction mode used for inter-frame prediction of the coded block being processed. In the process of calculating the motion vector of L0, X is 0, and in the process of calculating the motion vector of L1, X is 1. In addition, in the process of calculating the motion vector of LX, if information from another reference list is referenced instead of the same reference list as the LX being calculated, this other reference list is denoted as LY.
[0161] When using the motion vector mvLX of LX ( Figure 25 Step S202: Yes), calculate the candidates for predicted motion vectors of LX, and construct the candidate list of predicted motion vectors of LX, mvpListLX( Figure 25 Step S203). Multiple candidates for predicted motion vectors are calculated from the spatial predicted motion vector candidate derivation unit 421, the temporal predicted motion vector candidate derivation unit 422, the historical predicted motion vector candidate derivation unit 423, and the predicted motion vector candidate supplementation unit 425 in the normal predicted motion vector pattern derivation unit 401, and a predicted motion vector candidate list mvpListLX is constructed. (About...) Figure 25 The detailed processing steps of step S203 are as follows: Figure 20 The flowchart is described later.
[0162] Next, the prediction motion vector candidate selection unit 426 retrieves the candidate mvpListLX[mvpIdxLX] of the prediction motion vector that corresponds to the index mvpIdxLX of the prediction motion vector provided by the bit string decoding unit 201 from the prediction motion vector candidate list mvpListLX, and selects it as the selected prediction motion vector mvpLX. Figure 25 Step S204).
[0163] Next, the motion vector addition unit 427 performs an addition operation on the differential motion vector mvdLX and the predicted motion vector mvpLX of LX provided by the bit string decoding unit 201, denoted as mvLX = mvpLX + mvdLX, to calculate the motion vector mvLX of LX. Figure 25 Step S205).
[0164] <Commonly Predictive Motion Vector Pattern Derivative (AMVP): Motion Vector Prediction Methods>
[0165] Figure 20 This is a flowchart illustrating the processing steps of the general prediction motion vector pattern derivation process, which has a common function in the general prediction motion vector pattern derivation unit 301 of the image encoding apparatus and the general prediction motion vector pattern derivation unit 401 of the image decoding apparatus according to embodiments of the present invention.
[0166] Both the normal prediction motion vector pattern derivation unit 301 and the normal prediction motion vector pattern derivation unit 401 have a prediction motion vector candidate list mvpListLX. The prediction motion vector candidate list mvpListLX forms a list structure and is provided with a storage area that stores prediction motion vector indices representing positions within the prediction motion vector candidate list and the prediction motion vector candidates corresponding to those indices as elements. The prediction motion vector indices start from 0, and the prediction motion vector candidates are stored in the storage area of the prediction motion vector candidate list mvpListLX. In this embodiment, it is assumed that the prediction motion vector candidate list mvpListLX can register at least two prediction motion vector candidates (inter-frame prediction information). Furthermore, the variable numCurrMvpCand, representing the number of prediction motion vector candidates registered in the prediction motion vector candidate list mvpListLX, is set to 0.
[0167] Spatial prediction motion vector candidate derivation units 321 and 421 derive candidates for predicted motion vectors from the block adjacent to the left. In this process, reference is made to the block adjacent to the left ( Figure 11 The inter-frame prediction information (A0 or A1), including flags indicating whether the predicted motion vector candidate can be used, motion vectors, reference indices, etc., is used to derive the predicted motion vector mvLXA, and then the exported mvLXA is added to the predicted motion vector candidate list mvpListLX. Figure 20 Step S301). Additionally, X is 0 during L0 prediction and 1 during L1 prediction (the same applies below). Next, the spatial prediction motion vector candidate derivation units 321 and 421 derive candidates for predicted motion vectors from the block adjacent to the upper side. In this process, reference is made to the block adjacent to the upper side ( Figure 11 The inter-frame prediction information (B0, B1, or B2) indicates whether the predicted motion vector candidate can be used, along with the motion vector, reference index, etc., to derive the predicted motion vector mvLXB. If the derived mvLXA and mvLXB are not equal, then mvLXB is added to the predicted motion vector candidate list mvpListLX. Figure 20 Step S302). Figure 20 The processing steps S301 and S302 are the same except that the positions and number of the referenced adjacent blocks are different. They derive the flag availableFlagLXN, which indicates whether the predicted motion vector candidate of the coded block can be used, as well as the motion vector mvLXN and the reference index refIdxN (N represents A or B, the same below).
[0168] Next, the historical predicted motion vector candidate derivation units 323 and 423 add the historical predicted motion vector candidates registered in the historical predicted motion vector candidate list HmvpCandList to the predicted motion vector candidate list mvpListLX. Figure 20 Step S303). This will be used later. Figure 29 The flowchart describes the details of the registration processing step in step S303.
[0169] Next, the temporal prediction motion vector candidate derivation units 322 and 422 derive candidates for predicted motion vectors from blocks in images at times different from the current processing target image. In this process, the availability label LXCol, motion vector mvLXCol, reference index refIdxCol, and reference list listCol of the predicted motion vector candidates indicating whether the coded blocks of the images at different times can be utilized are derived. mvLXCol is then added to the predicted motion vector candidate list mvpListLX. Figure 20 Step S304).
[0170] Furthermore, it is assumed that the processing of time-predicted motion vector candidate deriving units 322 and 422, which are in units of sequence (SPS), image (PPS), or strip, can be omitted.
[0171] Next, before satisfying the predicted motion vector candidate supplementation unit 325 and 425, the predicted motion vector candidate with predetermined values such as (0, 0) is added. Figure 20 (S305).
[0172] <Normal Merge Pattern Export Section (Normal Merge)>
[0173] Figure 18 The normal merging mode derivation section 302 includes a spatial merging candidate derivation section 341, a temporal merging candidate derivation section 342, an average merging candidate derivation section 344, a historical merging candidate derivation section 345, a merging candidate supplementation section 346, and a merging candidate selection section 347.
[0174] Figure 24 The normal merging mode derivation section 402 includes a spatial merging candidate derivation section 441, a temporal merging candidate derivation section 442, an average merging candidate derivation section 444, a historical merging candidate derivation section 445, a merging candidate supplementation section 446, and a merging candidate selection section 447.
[0175] Figure 21 This is a flowchart illustrating the steps of a normal merging mode export process that have a common function in the normal merging mode export unit 302 of the image encoding apparatus and the normal merging mode export unit 402 of the image decoding apparatus according to embodiments of the present invention.
[0176] The following sections will explain each process in turn. Furthermore, unless otherwise specified, the following explanations will focus on the case where the slice_type is B-slice, but this also applies to the case of P-slice. However, when the slice_type is P-slice, since only L0 prediction (Pred_L0) exists as the inter-frame prediction mode, and L1 prediction (Pred_L1) and double prediction (Pred_BI) do not exist, the processing surrounding L1 can be omitted.
[0177] Both the normal merge mode derivation unit 302 and the normal merge mode derivation unit 402 have a merge candidate list, mergeCandList. The merge candidate list mergeCandList forms a list structure and includes a merge index indicating the position within the merge candidate list, and a storage area for storing the merge candidates corresponding to the index as elements. The merge index number starts from 0, and merge candidates are stored in the storage area of the merge candidate list mergeCandList. In subsequent processing, it is assumed that the merge candidate registered at merge index i in the merge candidate list mergeCandList is represented by mergeCandList[i]. In this embodiment, it is assumed that the merge candidate list mergeCandList can register at least six merge candidates (inter-frame prediction information). Furthermore, the variable numCurrMergeCand, representing the number of merge candidates registered in the merge candidate list mergeCandList, is set to 0.
[0178] In the spatial merging candidate derivation unit 341 and the spatial merging candidate derivation unit 441, based on the encoding information stored in the encoding information storage memory 111 of the image encoding device or the encoding information storage memory 205 of the image decoding device, the candidates are processed in the order of B1, A1, B0, A0, B2 from the blocks adjacent to the left and top of the processing target block. Figure 11 Export space merge candidates from B1, A1, B0, A0, B2, and register the exported space merge candidates in the MergeCandList. Figure 21 (Step S401 in the previous section). Here, N is defined to represent any one of the spatial merge candidate B1, A1, B0, A0, B2, or the temporal merge candidate Col. The following are derived: a flag availableFlagN indicating whether the inter-frame prediction information of block N can be used as a spatial merge candidate; a reference index refIdxL0N for L0 and a reference index refIdxL1N for L1 of spatial merge candidate N; an L0 prediction flag predFlagL0N indicating whether L0 prediction is performed; an L1 prediction flag predFlagL1N indicating whether L1 prediction is performed; the motion vector mvL0N for L0; and the motion vector mvL1N for L1. However, in this embodiment, since the merge candidates are derived without referring to the inter-frame prediction information of the blocks contained in the coded block to be processed, spatial merge candidates using the inter-frame prediction information of the blocks contained in the coded block to be processed are not derived. Figure 11 (B1, A1, B0, A0, B2)
[0179] Next, in the time merging candidate export unit 342 and the time merging candidate export unit 442, time merging candidates from images at different times are exported, and the exported time merging candidates are registered in the merge candidate list mergeCandList. Figure 21 Step S402). Derive the availableFlagCol, which indicates whether time merging candidates can be used; the predFlagL0Col, which indicates whether time merging candidates should be predicted for L0; the predFlagL1Col, which indicates whether L1 prediction should be made; and the motion vectors mvL0Col for L0 and mvL1Col for L1.
[0180] Furthermore, the processing of the time merging candidate export unit 342 and the time merging candidate export unit 442, which are based on sequences (SPS), images (PPS), or stripes, can be omitted.
[0181] Next, in the historical merge candidate derivation unit 345 and the historical merge candidate derivation unit 445, the historical predicted motion vector candidates registered in the historical predicted motion vector candidate list HmvpCandList are registered in the merge candidate list mergeCandList. Figure 21 Step S403).
[0182] Furthermore, if the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is less than the maximum number of merge candidates, MaxNumMergeCand, then the number of merge candidates registered in the merge candidate list mergeCandList numCurrMergeCand is used as the upper limit of the maximum number of merge candidates, MaxNumMergeCand, to derive historical merge candidates, and these candidates are then registered in the merge candidate list mergeCandList.
[0183] Next, in the average merge candidate derivation section 344 and the average merge candidate derivation section 444, average merge candidates are derived from the merge candidate list mergeCandList, and the exported average merge candidates are added to the merge candidate list mergeCandList. Figure 21 Step S404).
[0184] Furthermore, if the number of merge candidates registered in the merge candidate list numCurrMergeCand is less than the maximum number of merge candidates MaxNumMergeCand, the number of merge candidates registered in the merge candidate list numCurrMergeCand is capped at the maximum number of merge candidates MaxNumMergeCand, and an average number of merge candidates is derived and registered in the merge candidate list mergeCandList.
[0185] Here, the average merge candidate is a new merge candidate that has a motion vector obtained by averaging the motion vectors of the first and second merge candidates registered in the merge candidate list mergeCandList according to each L0 prediction and L1 prediction.
[0186] Next, in merge candidate supplementation section 346 and merge candidate supplementation section 446, when the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is less than the maximum number of merge candidates, MaxNumMergeCand, the number of merge candidates registered in the merge candidate list mergeCandList, numCurrMergeCand, is used as the upper limit of the maximum number of merge candidates, MaxNumMergeCand, to derive additional merge candidates, and these candidates are registered in the merge candidate list mergeCandList. Figure 21 Step S405). Using the maximum number of merge candidates (MaxNumMergeCand) as the upper limit, in the P-strip, add merge candidates with a prediction mode of L0 prediction (Pred_L0) where the motion vector has a value of (0,0). In the B-strip, add merge candidates with a prediction mode of double prediction (Pred_BI) where the motion vector has a value of (0,0). The reference index when adding merge candidates is different from the reference index already added.
[0187] Next, in the merge candidate selection unit 347 and merge candidate selection unit 447, merge candidates are selected from those registered in the merge candidate list mergeCandList. On the encoding side, the merge candidate selection unit 347 selects merge candidates by calculating the amount of code and the amount of distortion, and provides the merge index representing the selected merge candidate and the inter-frame prediction information of the merge candidate to the motion compensation prediction unit 306 via the inter-frame prediction mode determination unit 305. On the other hand, on the decoding side, the merge candidate selection unit 447 selects merge candidates based on the decoded merge index and provides the selected merge candidates to the motion compensation prediction unit 406.
[0188] <Update historical prediction motion vector candidate list>
[0189] Next, the initialization and update methods of the historical predicted motion vector candidate list HmvpCandList possessed by the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side are described in detail. Figure 26 This is a flowchart illustrating the steps involved in initializing / updating the candidate list of historical predicted motion vectors.
[0190] In this embodiment, it is assumed that the update of the historical predicted motion vector candidate list HmvpCandList is performed in the encoding information storage memory 111 and the encoding information storage memory 205. Alternatively, a historical predicted motion vector candidate list update unit may be provided in the inter-frame prediction unit 102 and the inter-frame prediction unit 203 to perform the update of the historical predicted motion vector candidate list HmvpCandList.
[0191] At the beginning of the strip, the historical predicted motion vector candidate list HmvpCandList is initially set. On the encoding side, if the prediction method determination unit 105 selects the normal predicted motion vector mode or the normal merging mode, the historical predicted motion vector candidate list HmvpCandList is updated. On the decoding side, if the prediction information decoded by the bit string decoding unit 201 is the normal predicted motion vector mode or the normal merging mode, the historical predicted motion vector candidate list HmvpCandList is updated.
[0192] Inter-frame prediction information used during inter-frame prediction in either the normal motion vector prediction mode or the normal merging mode is registered in the historical prediction motion vector candidate list hmvpCandList as inter-frame prediction information candidate hMvpCand. Inter-frame prediction information candidate hMvpCand includes the reference index refIdxL0 for L0 and the reference index refIdxL1 for L1, the L0 prediction flag predFlagL0 indicating whether L0 prediction is performed, the L1 prediction flag predFlagL1 indicating whether L1 prediction is performed, the motion vector mvL0 for L0, and the motion vector mvL1 for L1.
[0193] If, among the elements (i.e., inter-frame prediction information) registered in the historical prediction motion vector candidate list HmvpCandList stored in the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side, there exists inter-frame prediction information with the same value as the inter-frame prediction information candidate hMvpCand, then that element is deleted from the historical prediction motion vector candidate list HmvpCandList. On the other hand, if there is no inter-frame prediction information with the same value as the inter-frame prediction information candidate hMvpCand, then the element at the beginning of the historical prediction motion vector candidate list HmvpCandList is deleted, and the inter-frame prediction information candidate hMvpCand is added to the end of the historical prediction motion vector candidate list HmvpCandList.
[0194] The number of elements in the historical predicted motion vector candidate list HmvpCandList possessed by the encoding information storage memory 111 on the encoding side and the encoding information storage memory 205 on the decoding side of the present invention is set to 6.
[0195] First, initialize the historical predicted motion vector candidate list HmvpCandList( on a strip basis). Figure 26 Step S2101). At the beginning of the strip, all elements in the historical predicted motion vector candidate list HmvpCandList are set to empty, and the value of the number of historical predicted motion vector candidates (current candidate number) NumHmvpCand registered in the historical predicted motion vector candidate list HmvpCandList is set to 0.
[0196] Furthermore, although it is assumed that the initialization of the historical predicted motion vector candidate list HmvpCandList is implemented in strip units (the initial coded blocks of the strip), it can also be implemented in image units, tile units, or tree block rows units.
[0197] Next, the following update process for the historical predicted motion vector candidate list HmvpCandList is repeated for each coded block within the strip. Figure 26 Steps S2102 to S2111).
[0198] First, initial settings are performed on a block-by-block basis. The flag `identicalCandExist`, indicating the existence of identical candidates, is set to `FALSE`, and the index of the candidate to be deleted, `removeIdx`, is set to "0". Figure 26 Step S2103).
[0199] Determine whether there are inter-frame prediction information candidates hMvpCand for the registered object. Figure 26 Step S2104). If the prediction method determination unit 105 on the encoding side determines it to be a normal prediction motion vector mode or a normal merging mode, or if the bit string decoding unit 201 on the decoding side decodes it to be a normal prediction motion vector mode or a normal merging mode, the inter-frame prediction information is set as the inter-frame prediction information candidate hMvpCand for the registered object. If the prediction method determination unit 105 on the encoding side determines it to be an intra-frame prediction mode, a sub-block prediction motion vector mode, or a sub-block merging mode, or if the bit string decoding unit 201 on the decoding side decodes it to be an intra-frame prediction mode, a sub-block prediction motion vector mode, or a sub-block merging mode, the historical prediction motion vector candidate list HmvpCandList is not updated, and there is no inter-frame prediction information candidate hMvpCand for the registered object. If there is no inter-frame prediction information candidate hMvpCand for the registered object, steps S2105 to S2106 are skipped. Figure 26 Step S2104: No). If there is a candidate hMvpCand for inter-frame prediction information of the registered object, proceed with the processing after step S2105. Figure 26 Step S2104: Yes).
[0200] Next, it is determined whether there exists an element in each element of the historical predicted motion vector candidate list HmvpCandList that has the same value as the inter-frame prediction information candidate hMvpCand of the registered object (inter-frame prediction information), that is, whether there is a matching element. Figure 26 Step S2105). Figure 27 This is a flowchart of the same element confirmation process. When the value of the historical predicted motion vector candidate number NumHmvpCand is 0 ( Figure 27 Step S2121: No), the historical predicted motion vector candidate list HmvpCandList is empty. Since there are no identical candidates, this step is skipped. Figure 27 Steps S2122 to S2125 conclude the same feature confirmation process. This applies when the value of the historical predicted motion vector candidate number NumHmvpCand is greater than 0 ( Figure 27 Step S2121: Yes), the historical predicted motion vector index hMvpIdx is from 0 to NumHmvpCand-1, and the processing of step S2123 is repeated. Figure 27 Steps S2122 to S2125). First, compare whether the hMvpIdx-th element HmvpCandList[hMvpIdx] in the historical predicted motion vector candidate list (starting from 0) is the same as the inter-frame prediction information candidate hMvpCand. Figure 27 Step S2123). Under the same circumstances ( Figure 27 Step S2123: If yes, set the flag `identicalCandExist`, indicating whether there are identical candidates, to TRUE; set the current historical predicted motion vector index `hMvpIdx`, representing the location of the deleted object's index `removeIdx`, to the value of the index `hMvpIdx`; and end the identical feature confirmation process. In the case of different features ( Figure 27 Step S2123: No), increase hMvpIdx by 1. If the historical predicted motion vector index hMvpIdx is below NumHmvpCand-1, then proceed with the processing after step S2123.
[0201] Return again Figure 26 The flowchart describes the process of shifting and adding elements to the historical predicted motion vector candidate list HmvpCandList. Figure 26 Step S2106). Figure 28 yes Figure 26 The flowchart for the feature shifting / addition process in step S2106 of the historical predicted motion vector candidate list HmvpCandList is as follows: First, it is determined whether to add new features after removing features stored in the historical predicted motion vector candidate list HmvpCandList, or to add new features without removing existing features. Specifically, it compares whether the flag identicalCandExist, indicating the existence of identical candidates, is TRUE or whether NumHmvvpCand is 6. Figure 28 Step S2141). Under the condition that either the flag `identicalCandExist` indicating the existence of identical candidates is TRUE or the current number of candidates `NumHmvpCand` is 6 (…), Figure 28 Step S2141: Yes), after removing the features stored in the historical predicted motion vector candidate list HmvpCandList, add new features. Set the initial value of index i to the value of removeIdx+1. Repeat the feature shifting process of step S2143 from this initial value to NumHmvpCand. Figure 28 Steps S2142 to S2144). By copying the elements of HmvpCandList[i] to HmvpCandList[i-1], the elements are shifted forward ( Figure 28 Step S2143), increment i by 1 ( Figure 28 Steps S2142 to S2144). Next, add the inter-frame prediction information candidate hMvpCand to the (NumHmvpCand-1)th HmvpCandList[NumHmvpCand-1], which is the last corresponding to the historical prediction motion vector candidate list, starting from 0. Figure 28 Step S2145) ends the feature shifting / addition process of the historical predicted motion vector candidate list HmvpCandList. On the other hand, if neither of the conditions that the flag identicalCandExist indicating the existence of identical candidates is TRUE nor NumHmvpCand is 6 is met ( Figure 28 Step S2141: No), instead of removing the features stored in the historical predicted motion vector candidate list HmvpCandList, add the inter-frame prediction information candidate hMvpCand to the end of the historical predicted motion vector candidate list. Figure 28 Step S2146). Here, the last element of the historical predicted motion vector candidate list is the NumHmvpCand-th HmvpCandList[NumHmvpCand], counting from 0. Additionally, NumHmvpCand is incremented by 1, ending the shifting and addition process for elements in the historical predicted motion vector candidate list HmvpCandList.
[0202] Figure 31 is a diagram illustrating an example of the update process for the historical predicted motion vector candidate list. When adding a new feature to the historical predicted motion vector candidate list HmvpCandList, which already has six registered features (inter-frame prediction information), the new inter-frame prediction information is compared sequentially with the features preceding them in the historical predicted motion vector candidate list HmvpCandList. Figure 31A If the new feature has the same value as the third feature (HMVP2) from the beginning of the historical predicted motion vector candidate list (HmvpCandList), then remove feature HMVP2 from the historical predicted motion vector candidate list (HmvpCandList), shift (copy) the subsequent features HMVP3 to HMVP5 one by one forward, and add the new feature to the end of the historical predicted motion vector candidate list (HmvpCandList). Figure 31B Complete the update of the historical predicted motion vector candidate list HmvpCandList. Figure 31C ).
[0203] <Historical Predicted Motion Vector Candidate Export Processing>
[0204] Next, a detailed explanation will be provided as... Figure 20 Step S304 is a method for deriving historical predicted motion vector candidates from the historical predicted motion vector candidate list HmvpCandList. Figure 20 The processing step S304 is a common process in the historical prediction motion vector candidate derivation unit 323 of the normal prediction motion vector pattern derivation unit 301 on the encoding side and the historical prediction motion vector candidate derivation unit 423 of the normal prediction motion vector pattern derivation unit 401 on the decoding side. Figure 29 This is a flowchart illustrating the steps involved in deriving candidate motion vectors from historical predictions.
[0205] If the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (which is 2 here), or if the historical number of predicted motion vector candidates, NumHmvpCand, is 0 ( Figure 29 Step S2201: No), omitted Figure 29 The processing from steps S2202 to S2209 ends the historical predicted motion vector candidate export processing step. When the current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX (2), and the value of the historical predicted motion vector candidate number, NumHmvpCand, is greater than 0 (…), the process continues. Figure 29 Step S2201: Yes), execute Figure 29 The processing steps S2202 to S2209.
[0206] Next, repeat. Figure 29 The processing in steps S2203 to S2208 continues until either index i is between 1 and 4 and either the historical predicted motion vector candidate number numCheckedHMVPCand is a smaller value. Figure 29 Steps S2202 to S2209). When the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (more than 2). Figure 29 Step S2203: No), omitted Figure 29 The processing steps S2204 to S2209 conclude the historical predicted motion vector candidate export processing step. The current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX, which is 2 ( Figure 29 If step S2203 is true, execute the following steps: Figure 29 The processing after step S2204.
[0207] Next, for Y values of 0 and 1 (L0 and L1), the processes from steps S2205 to S2207 are performed respectively. Figure 29 Steps S2204 to S2208). When the current number of predicted motion vector candidates, numCurrMvpCand, is greater than or equal to the maximum number of features in the predicted motion vector candidate list mvpListLX (more than 2). Figure 29 Step S2205: No), omitted Figure 29 The processing steps S2206 to S2209 conclude the historical predicted motion vector candidate export processing step. This applies when the current number of predicted motion vector candidates, numCurrMvpCand, is less than the maximum number of features in the predicted motion vector candidate list mvpListLX (2). Figure 29 Step S2205 in the process: Yes, execute Figure 29 The processing after step S2206 in the process.
[0208] Next, in the historical predicted motion vector candidate list HmvpCandList, the element that has the same reference index as the reference index refIdxLX of the encoded / decoded object motion vector, and is different from any element in the predicted motion vector list mvpListLX ( Figure 29 Step S2206 is to add the motion vector of the historical predicted motion vector candidate HmvpCandList[NumHmvpCand-i] to the element mvpListLX[numCurrMvpCand] starting from 0 in the predicted motion vector candidate list. Figure 29 Step S2207), and increment the current number of predicted motion vector candidates numCurrMvpCand by 1. This occurs when there are no features in the historical predicted motion vector candidate list HmvpCandList with the same reference index refIdxLX as the motion vector of the encoded / decoded object, and no features that are different from any feature in the predicted motion vector list mvpListLX. Figure 29 If step S2206 is not selected, skip the addition process in step S2207.
[0209] Both L0 and L1 are performed as above Figure 29 Processing steps S2205 to S2207 ( Figure 29 Steps S2204 to S2208). Increment index i by 1. If index i is less than or equal to 4 and either the historical predicted motion vector candidate number NumHmvpCand is smaller, repeat the processing after step S2203. Figure 29 Steps S2202 to S2209).
[0210] <Historical Merge Candidate Export Processing>
[0211] Next, a detailed explanation will be provided as... Figure 21 Step S404 is a method for deriving historical merge candidates from the historical merge candidate list HmvpCandList. Figure 21 The processing step S404 is a common process in the historical merge candidate derivation section 345 of the normal merge mode derivation section 302 on the encoding side and the historical merge candidate derivation section 445 of the normal merge mode derivation section 402 on the decoding side. Figure 30 This is a flowchart illustrating the steps involved in the historical merge candidate export process.
[0212] First, perform initialization processing ( Figure 30 Step S2301). Set the value of FALSE for each element from 0 to (numCurrMergeCand-1)th element in isPruned[i], and set the variable numOrigMergeCand to the number of elements registered in the current merge candidate list, numCurrMergeCand.
[0213] Next, the initial value of index hMvpIdx is set to 1, and the process is repeated from this initial value to NumHmvpCand. Figure 30 The addition process from step S2303 to step S2310 ( Figure 30 Steps S2302 to S2311). If the number of features registered in the current merge candidate list, numCurrMergeCand, is not below (MaxNumMergeCand-1), then the historical merge candidate export process ends because all features have been added to the merge candidate list as merge candidates. Figure 30 Step S2303: No). If the number of features registered in the current merge candidate list, numCurrMergeCand, is less than (MaxNumMergeCand-1), then proceed with the processing after step S2304. Set the value of sameMotion to FALSE. Figure 30 (Step S2304). Next, the initial value of index i is set to 0, and the process continues from this initial value to numOrigMergeCand-1. Figure 30 Processing of steps S2306 and S2307 ( Figure 30 (S2305~S2308). Compare whether the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] in the historical motion vector prediction candidate list (starting from 0) is the same as the i-th element mergeCandList[i] in the merged candidate list (starting from 0). Figure 30 Step S2306).
[0214] The term "same value" for merge candidates refers to the fact that all the constituent elements (inter-frame prediction mode, reference index, and motion vector) of a merge candidate have the same value. This applies when the merge candidates have the same value and isPruned[i] is FALSE. Figure 30 Step S2306: Yes), sameMotion and isPruned[i] are both set to TRUE. Figure 30 Step S2307). In the case that the values are not the same ( Figure 30 If step S2306 is not specified, skip step S2307. Figure 30 After the repeated processing of steps S2305 to S2308 is completed, compare whether sameMotion is FALSE. Figure 30 Step S2309), in the case that sameMotion is FALSE (false) Figure 30 Step S2309: Yes), that is, since the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHvpCand-hMvpIdx] from the historical predicted motion vector candidate list does not exist in mergeCandList, therefore, the (NumHmvpCand-hMvpIdx)th element HmvpCandList[NumHmvpCand-hMvpIdx] from the historical predicted motion vector candidate list is added to the mergeCandList[numCurrMergeCand] of the merge candidate list, and numCurrMergeCand is incremented by 1. Figure 30 Step S2310). Increment the index hMvpIdx by 1 ( Figure 30 Step S2302) is performed. Figure 30 Repeat steps S2302 to S2311.
[0215] After confirming all elements in the historical predicted motion vector candidate list, or adding all elements in the merge candidate list to the merge candidate list, the export process of the historical merge candidate is completed.
[0216] <Motion Compensation Predictive Processing>
[0217] The motion compensation prediction unit 306 acquires the position and size of the block of the object currently being predicted during encoding. Additionally, the motion compensation prediction unit 306 acquires inter-frame prediction information from the inter-frame prediction mode determination unit 305. Based on the acquired inter-frame prediction information, a reference index and motion vector are derived. After acquiring an image signal that shifts the reference image determined by the reference index in the decoded image memory 104 from the same position as the image signal of the predicted block by the amount of motion vector, a prediction signal is generated.
[0218] In inter-frame prediction, when the inter-frame prediction mode is L0 or L1 prediction, which involves prediction from a single reference image, the prediction signal obtained from one reference image is set as the motion-compensated prediction signal. When the inter-frame prediction mode is BI prediction, which involves prediction from two reference images, the signal obtained by weighted averaging of the prediction signals obtained from the two reference images is set as the motion-compensated prediction signal, and this motion-compensated prediction signal is provided to the prediction method determination unit 105. Here, the weighted average ratio of the two predictions is set to 1:1, but other ratios can also be used for weighted averaging. For example, the closer the image to be predicted is to the reference image, the larger the weighting ratio. Alternatively, a table corresponding to combinations of image intervals and weighting ratios can be used to calculate the weighting ratio.
[0219] The motion compensation prediction unit 406 has the same function as the motion compensation prediction unit 306 on the encoding side. The motion compensation prediction unit 406 acquires inter-frame prediction information from the normal prediction motion vector pattern derivation unit 401, the normal merging pattern derivation unit 402, the sub-block prediction motion vector pattern derivation unit 403, and the sub-block merging pattern derivation unit 404 via switch 408. The motion compensation prediction unit 406 provides the acquired motion compensation prediction signal to the decoded image signal overlay unit 207.
[0220] <About Inter-Frame Prediction Mode>
[0221] The process of making a prediction based on a single reference image is defined as single prediction. In the case of single prediction, a prediction is made using either an L0 prediction or an L1 prediction, which uses either of the two reference images registered in the reference list L0 or L1.
[0222] Figure 32 This shows the case where the reference image (RefL0Pic) of L0 in a single prediction is at a moment before the processing object image (CurPic). Figure 33 This illustrates the case where the reference image for the L0 prediction in a single prediction is at a time after the processed object image. Similarly, it is also possible to... Figure 32 and Figure 33 The reference image for L0 prediction is replaced with the reference image for L1 prediction (RefL1Pic) for single prediction.
[0223] The process of making predictions based on two reference images is defined as dual prediction. In the case of dual prediction, L0 prediction and L1 prediction are used to describe BI prediction. Figure 34 This illustrates the case where the reference image for L0 prediction is at a time before the object image is processed, and the reference image for L1 prediction is at a time after the object image is processed. Figure 35 This shows the situation where the reference images for L0 prediction and L1 prediction in a dual prediction are at a time before the object image is processed. Figure 36 This shows the case where the reference images for L0 prediction and L1 prediction in a dual prediction scenario are at a time after the processed object image.
[0224] Thus, the relationship between the prediction category and time of L0 / L1 can be used when L0 is not limited to the past direction and L1 is not limited to the future direction. Furthermore, in the case of dual prediction, the same reference image can be used to perform both L0 and L1 predictions. Additionally, it is determined whether motion compensation prediction is performed using single or dual prediction based on information indicating whether L0 or L1 prediction is used (e.g., flags).
[0225] <About the Reference Index>
[0226] In embodiments of the present invention, to improve the accuracy of motion compensation prediction, the optimal reference image can be selected from multiple reference images during motion compensation prediction. Therefore, the reference image used in motion compensation prediction is used as a reference index, and the reference index is encoded into the bitstream along with the differential motion vector.
[0227] <Motion compensation processing based on typical predicted motion vector patterns>
[0228] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when inter-frame prediction information based on the normal prediction motion vector mode derivation unit 301 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0229] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to the normal prediction motion vector mode derivation unit 401 during decoding, motion compensation prediction unit 406 acquires the inter-frame prediction information based on the normal prediction motion vector mode derivation unit 401, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0230] <Motion compensation processing based on the usual merging pattern>
[0231] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when inter-frame prediction information based on the normal merging mode derivation unit 302 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0232] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to the normal merging mode derivation unit 402 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on the normal merging mode derivation unit 402, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0233] <Motion Compensation Processing Based on Sub-Block Predicted Motion Vector Patterns>
[0234] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when the inter-frame prediction information based on the sub-block prediction motion vector mode derivation unit 303 is selected in the inter-frame prediction mode determination unit 305, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0235] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to sub-block prediction motion vector mode derivation unit 403 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on sub-block prediction motion vector mode derivation unit 403, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to decoded image signal overlay unit 207.
[0236] <Motion Compensation Processing Based on Sub-Block Merging Pattern>
[0237] As in Figure 16 As also shown in the inter-frame prediction unit 102 on the encoding side, when the inter-frame prediction mode determination unit 305 selects the inter-frame prediction information based on the sub-block merging mode derivation unit 304, the motion compensation prediction unit 306 obtains the inter-frame prediction information from the inter-frame prediction mode determination unit 305, derives the inter-frame prediction mode, reference index, and motion vector of the block to be processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the prediction method determination unit 105.
[0238] Similarly, as in Figure 22 As also shown in the inter-frame prediction unit 203 on the decoding side, when switch 408 is connected to sub-block merging mode derivation unit 404 during decoding, motion compensation prediction unit 406 acquires inter-frame prediction information based on sub-block merging mode derivation unit 404, derives the inter-frame prediction mode, reference index, and motion vector of the block currently being processed, and generates a motion compensation prediction signal. The generated motion compensation prediction signal is provided to the decoded image signal overlay unit 207.
[0239] Motion compensation processing based on affine transformation prediction
[0240] In both the normal prediction motion vector mode and the normal merging mode, affine model-based motion compensation can be used based on the following flags. These flags are reflected in the following flags based on the inter-frame prediction conditions determined by the inter-frame prediction mode determination unit 305 during the encoding process, and are encoded into the bitstream. During the decoding process, whether to perform affine model-based motion compensation is determined based on the following flags in the bitstream.
[0241] `sps_affine_enabled_flag` indicates whether affine-based motion compensation can be used in inter-frame prediction. If `sps_affine_enabled_flag` is 0, sequence-by-sequence suppression is applied, preventing affine-based motion compensation. Furthermore, `inter_affine_flag` and `cu_affine_type_flag` are not transmitted in the CU (coded block) syntax of the encoded video sequence. If `sps_affine_enabled_flag` is 1, affine-based motion compensation can be used in the encoded video sequence.
[0242] `sps_affine_type_flag` indicates whether motion compensation based on a six-parameter affine model can be used in inter-frame prediction. If `sps_affine_type_flag` is 0, motion compensation based on a six-parameter affine model is suppressed. Additionally, `cu_affine_type_flag` is not transmitted in the CU syntax of the encoded video sequence. If `sps_affine_type_flag` is 1, motion compensation based on a six-parameter affine model can be used in the encoded video sequence. If `sps_affine_type_flag` is not present, it is set to 0.
[0243] When decoding P-strips or B-strips, in the CU that is currently being processed, if `inter_affine_flag` is 1, motion compensation based on an affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed. If `inter_affine_flag` is 0, the affine model is not used for the CU that is currently being processed. If `inter_affine_flag` does not exist, it is set to 0.
[0244] When decoding P-strips or B-strips, in the CU that is currently being processed, if cu_affine_type_flag is 1, motion compensation based on a six-parameter affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed. If cu_affine_type_flag is 0, motion compensation based on a four-parameter affine model is used to generate the motion compensation prediction signal for the CU that is currently being processed.
[0245] In motion compensation based on affine models, since reference indices or motion vectors are derived on a sub-block basis, motion compensation prediction signals are generated using the reference indices or motion vectors that are the objects of processing on a sub-block basis.
[0246] The four-parameter affine model is as follows: the motion vector of the sub-block is derived from the four parameters of the horizontal and vertical components of the motion vectors of the two control points, and motion compensation is performed on a sub-block basis.
[0247] In this embodiment, during the derivation of the candidate list for predicted motion vectors in a typical motion vector prediction mode, candidates are added in the order of spatial predicted motion vector candidates, historical predicted motion vector candidates, and temporal predicted motion vector candidates. By adopting this structure, the following effects can be achieved.
[0248] 1. In the historical predicted motion vector candidate derivation process, a feature verification step is performed between the features already added to the predicted motion vector candidate list and the features in the historical predicted motion vector candidate list. Only if they are different are the features added to the historical predicted motion vector candidate list, thus ensuring that the predicted motion vector candidate lists each contain different features. Furthermore, spatial predicted motion vector candidates using spatial correlation and historical predicted motion vector candidates using processing history have different characteristics. Therefore, the possibility of providing multiple predicted motion vector candidates with different characteristics increases, improving coding efficiency.
[0249] 2. The historical predicted motion vector candidate derivation process performs the same feature verification as the spatial predicted motion vector candidate, but not the same feature verification as the temporal predicted motion vector candidate. Therefore, by limiting the number of times the same feature verification is performed, the processing load associated with the derivation of the predicted motion vector candidate list can be reduced.
[0250] 3. The temporal prediction motion vector candidate derivation process does not perform the same feature verification as the spatial prediction motion vector candidate and the historical prediction motion vector candidate. Therefore, historical prediction motion vector candidates and temporal prediction motion vector candidates can be derived independently. This enables improved throughput based on parallel processing.
[0251] (Second Implementation)
[0252] In the second embodiment, in the generation of the candidate list of predicted motion vectors for the usual predicted motion vector pattern, instead of exporting time-predicted motion vector candidates, candidates are added in the order of spatial-predicted motion vector candidates and historical-predicted motion vector candidates.
[0253] Figure 38 It is the second embodiment. Figure 16 A block diagram showing the detailed structure of the typical predictive motion vector pattern derivation unit 301.
[0254] Figure 39 It is the second embodiment. Figure 22 A block diagram showing the detailed structure of the typical predictive motion vector pattern derivation unit 401.
[0255] In the second embodiment, since a list of predicted motion vector candidates is generated without deriving temporally predicted motion vector candidates, the processing load can be reduced. Furthermore, in the normal predicted motion vector mode, the coding efficiency is not reduced because the list of predicted motion vector candidates is sufficiently filled with historical predicted motion vector candidates.
[0256] (Third Implementation)
[0257] In the third embodiment, during the generation of the candidate list for predicted motion vectors in a typical predicted motion vector pattern, candidates are added in the order of spatial predicted motion vector candidates, temporal predicted motion vector candidates, and historical predicted motion vector candidates. Here, in the historical predicted motion vector candidate derivation process, the same elements as those in the spatial and temporal predicted motion vector candidates are not verified.
[0258] Figure 40 It is the third embodiment. Figure 16 A block diagram showing the detailed structure of the typical predictive motion vector pattern derivation unit 301.
[0259] Figure 41 It is the third embodiment. Figure 22 A block diagram showing the detailed structure of the typical predictive motion vector pattern derivation unit 401.
[0260] In the third embodiment, similar to the first embodiment, the number of times the same element is checked can be limited, thus reducing the processing load associated with deriving the predicted motion vector candidate list. Furthermore, by adding temporal predicted motion vector candidates to the predicted motion vector candidate list in a higher order than historical predicted motion vector candidates, a predicted motion vector candidate list with high coding efficiency can be generated by prioritizing temporal predicted motion vector candidates with high prediction efficiency over historical predicted motion vector candidates without suppressing the processing load by checking for the same element of the predicted motion vector among different types of candidates (spatial predicted motion vector candidates, temporal predicted motion vector candidates, and historical predicted motion vector candidates).
[0261] All of the above-described embodiments can also be combined in multiple ways.
[0262] In all the embodiments described above, the bitstream output by the image encoding device has a specific data format so that it can be decoded according to the encoding method used in the embodiment. Furthermore, the image decoding device corresponding to the image encoding device is capable of decoding the bitstream of this specific data format.
[0263] When using wired or wireless networks to exchange bitstreams between an image encoding device and an image decoding device, the bitstream can be converted into a data format suitable for transmission over the communication line for transmission. In this case, a transmitting device is provided that converts the bitstream output from the image encoding device into encoded data in a data format suitable for transmission over the communication line and transmits the encoded data to the network; and a receiving device receives the encoded data from the network and restores the encoded data to a bitstream for supply to the image decoding device. The transmitting device includes: a memory for buffering the bitstream output from the image encoding device; a packet processing unit for packetizing the bitstream; and a transmitting unit for transmitting the packetized encoded data via the network. The receiving device includes: a receiving unit for receiving the packetized encoded data via the network; a memory for buffering the received encoded data; and a packet processing unit for packetizing the encoded data to generate a bitstream and supplying the bitstream to the image decoding device.
[0264] Alternatively, the display device can be configured by adding a display unit that displays the image decoded by the image decoding device. In this case, the display unit reads the decoded image signal generated by the decoded image signal overlap unit 207 and stored in the decoded image memory 208, and displays it on the screen.
[0265] Alternatively, an imaging unit can be added to the structure to input the captured image into an image encoding device, thus functioning as an imaging device. In this case, the imaging unit inputs the captured image signal into the block segmentation unit 101.
[0266] Figure 37 An example of the hardware structure of the encoding / decoding apparatus according to this embodiment is shown. The encoding / decoding apparatus includes the structure of the image encoding apparatus and the image decoding apparatus according to embodiments of the present invention. The encoding / decoding apparatus 9000 has a CPU 9001, an encoder / decoder IC 9002, an I / O interface 9003, a memory 9004, an optical disc drive 9005, a network interface 9006, and a video interface 9009, and the various parts are connected via a bus 9010.
[0267] The image encoding unit 9007 and the image decoding unit 9008 are typically installed as a codec IC 9002. In the image encoding apparatus according to embodiments of the present invention, image encoding processing is performed by the image encoding unit 9007, and in the image decoding apparatus according to embodiments of the present invention, image decoding processing is performed by the image decoding unit 9008. The I / O interface 9003 is implemented, for example, via a USB interface, and is connected to an external keyboard 9104, mouse 9105, etc. The CPU 9001 controls the encoding / decoding apparatus 9000 to execute the user-desired action based on user operations input through the I / O interface 9003. These operations, performed by the user via the keyboard 9104, mouse 9105, etc., include selecting which function to perform (encoding or decoding), setting the encoding quality, setting the input / output destination of the bitstream, and setting the input / output destination of the image.
[0268] When a user wishes to reproduce an image recorded on the disc recording medium 9100, the optical disc drive 9005 reads a bitstream from the inserted disc recording medium 9100 and sends the read bitstream to the image decoding unit 9008 of the codec IC 9002 via the bus 9010. The image decoding unit 9008 performs image decoding processing in the image decoding apparatus according to embodiments of the present invention on the input bitstream and sends the decoded image to an external monitor 9103 via the video interface 9009. Additionally, the codec apparatus 9000 has a network interface 9006 and can connect to an external distribution server 9106 and a portable terminal 9107 via the network 9101. When a user wishes to reproduce an image recorded on the distribution server 9106 or the mobile terminal 9107 instead of an image recorded on the disc recording medium 9100, the network interface 9006 obtains the bitstream from the network 9101 instead of reading the bitstream from the input disc recording medium 9100. Furthermore, if a user wishes to reproduce an image recorded in memory 9004, the image decoding process in the image decoding apparatus according to an embodiment of the present invention is performed on the bitstream recorded in memory 9004.
[0269] When a user wishes to encode and record an image captured by an external camera 9102 in memory 9004, the video interface 9009 inputs the image from the camera 9102 and sends it to the image encoding unit 9007 of the codec IC 9002 via bus 9010. The image encoding unit 9007 performs image encoding processing according to the image encoding apparatus of this invention on the image input via the video interface 9009 and generates a bitstream. Then, the bitstream is sent to memory 9004 via bus 9010. When the user wishes to record the bitstream on the disc recording medium 9100 instead of memory 9004, the optical disc drive 9005 writes the bitstream to the inserted disc recording medium 9100.
[0270] It is also possible to implement a hardware structure that has an image encoding device but no image decoding device, or a hardware structure that has an image decoding device but no image encoding device. Such a hardware structure can be implemented, for example, by replacing the codec IC9002 with an image encoding unit 9007 or an image decoding unit 9008, respectively.
[0271] The processing related to the above encoding and decoding can, of course, be implemented using hardware transmission, storage, and reception devices, and can be implemented through firmware stored in ROM (Read-Only Memory), flash memory, or software such as computers. This firmware or software program can be provided by recording it on a readable recording medium such as a computer, or by providing it from a server via wired or wireless networks, or by providing it as data broadcast via terrestrial wave or satellite digital broadcasting.
[0272] The present invention has been described above based on embodiments. The embodiments are illustrative; various modifications can be made to the combination of these constituent elements and processing steps, and such modifications are also within the scope of the present invention, as will be understood by those skilled in the art.
[0273] Symbol Explanation
[0274] 100 Image encoding device, 101 Block segmentation unit, 102 Inter-frame prediction unit, 103 Intra-frame prediction unit, 104 Decoded image memory, 105 Prediction method determination unit, 106 Residual generation unit, 107 Orthogonal transform / quantization unit, 108 Bit string encoding unit, 109 Inverse quantization / inverse orthogonal transform unit, 110 Decoded image signal overlay unit, 111 Encoded information storage memory, 200 Image decoding device, 201 Bit string decoding unit, 202 Block segmentation unit, 203 Inter-frame prediction unit, 204 Intra-frame prediction unit, 205 Encoded information storage memory, 206 Inverse quantization / inverse orthogonal transform unit, 207 Decoded image signal overlay unit, 208 Decoded image memory.< / poc>
Claims
1. A moving image decoding device, characterized in that, include: The spatial motion information candidate derivation unit derives spatial motion information candidates from the motion information of blocks that are spatially close to the decoded object block; The temporal motion information candidate derivation section derives temporal motion information candidates from the motion information of blocks that are temporally close to the decoded object block. as well as The historical motion information candidate derivation unit derives historical motion information candidates from the memory that holds the motion information of the decoded blocks. The historical motion information candidate and the spatial motion information candidate are compared in terms of motion information, but the historical motion information candidate is not compared with the temporal motion information candidate, nor is the spatial motion information candidate compared with the temporal motion information candidate.
2. A moving image decoding method, which is a method in a moving image decoding device, characterized in that it includes the following steps: Derive spatial motion information candidates from the motion information of blocks that are spatially close to the decoded object block; Derive temporal motion information candidates from the motion information of blocks that are temporally close to the decoded object block; as well as Export historical motion information candidates from the memory that holds the motion information of the decoded blocks. The historical motion information candidate and the spatial motion information candidate are compared in terms of motion information, but the historical motion information candidate is not compared with the temporal motion information candidate, nor is the spatial motion information candidate compared with the temporal motion information candidate.
3. A motion image encoding device, characterized in that, include: The spatial motion information candidate derivation unit derives spatial motion information candidates from the motion information of blocks that are spatially close to the encoded object block; The temporal motion information candidate derivation section derives temporal motion information candidates from the motion information of blocks that are temporally close to the encoded object block. as well as The historical motion information candidate derivation unit derives historical motion information candidates from the memory that holds the motion information of the encoded blocks. The historical motion information candidate and the spatial motion information candidate are compared in terms of motion information, but the historical motion information candidate is not compared with the temporal motion information candidate, nor is the spatial motion information candidate compared with the temporal motion information candidate.
4. A moving image encoding method, which is a method in a moving image encoding apparatus, characterized in that it includes the following steps: Derive spatial motion information candidates from the motion information of blocks that are spatially close to the encoded object block; The step of deriving temporal motion information candidates from the motion information of blocks that are temporally close to the encoded object block; as well as The step of deriving historical motion information candidates from the memory that holds motion information of encoded blocks. The historical motion information candidate and the spatial motion information candidate are compared in terms of motion information, but the historical motion information candidate is not compared with the temporal motion information candidate, nor is the spatial motion information candidate compared with the temporal motion information candidate.
5. A method for storing a bit stream, characterized in that, Includes the following steps: Generate a bitstream by performing the motion image encoding method according to claim 4; And to store the bit stream in a recording medium.
6. A method for transmitting a bit stream, characterized in that, Includes the following steps: Generate a bitstream by performing the motion image encoding method according to claim 4; And to transmit the bit stream.