Predictive image generation device, motion picture decoding device, and motion picture encoding device
By performing motion compensation and filtering on moving images, and utilizing the weighted average of the filter coefficients and the derivation of motion vectors, the problems of motion compensation accuracy and coding efficiency are solved. This reduces the storage of filter coefficients and the amount of differential vector code, thereby improving coding efficiency.
Patent Information
- Application Number
- CN202210849182.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-02-01
- Filing Date
- 2017-01-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2037-01-26
AI Technical Summary
In recent years, the improved precision of motion compensation processing in motion image coding and decoding technologies has led to an increase in filtering coefficients and storage requirements. Furthermore, the code size of differential vectors increases when using high-precision motion vectors, and coding efficiency may not necessarily improve.
A predicted image is generated by performing motion compensation on the reference image. The image is then filtered by a filtering unit, and the motion vector is derived by a motion vector derivation unit based on the size of the prediction block, quantization parameters, and the accuracy of the motion vector switching flag, using the weighted average of the filtering coefficients.
It solves the problems of motion compensation accuracy and coding efficiency, reduces the storage of filter coefficients and the amount of differential vector code, and improves coding efficiency.
Smart Images

Figure CN115209158B_ABST
Abstract
Description
[0001] This application is a divisional application of the PCT national phase application filed on January 26, 2017, with national application number 201780008957.2. Technical Field
[0002] This invention relates to a predictive image generation apparatus, a motion picture decoding apparatus, and a motion picture encoding apparatus. Background Technology
[0003] In order to efficiently transmit or record moving images, a moving image encoding device is used to generate encoded data by encoding the moving images, and a moving image decoding device is used to generate decoded images by decoding the encoded data.
[0004] Specific moving image coding methods include those proposed in H.264 / MPEG-4.AVC and HEVC (High-Efficiency Video Coding).
[0005] In this motion picture coding method, the image (picture) that constitutes the motion picture is managed by a hierarchical structure formed by slices obtained by segmenting the image, coding units (sometimes also called coding units) obtained by segmenting the slices, and prediction units (PUs) and transform units (TUs) that are blocks obtained by segmenting the coding units, thereby encoding and decoding by blocks.
[0006] Furthermore, in this motion picture coding method, a prediction image is typically generated based on a locally decoded image obtained by encoding / decoding the input image, while the prediction residual (sometimes called a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Methods for generating the prediction image include inter-frame prediction and intra-frame prediction.
[0007] In addition, non-patent literature 1 can be cited as a technology for motion image encoding and decoding in recent years.
[0008] Existing technical documents
[0009] Non-patent literature
[0010] Non-patent document 1: Video / JVET, "Algorithm Description of Joint ExplorationTestModel 1 (JEM 1)", INTERNATIONAL ORGANIZATION FOR STANDARDIZATIONORGANISATION INTERNATIONALE DE NORMALISATION ISO / IEC JTC1 / SC29 / WG11 CODING OFMOVING PICTURES AND AUDIO, ISO / IEC JTC1 / SC29 / WG11 / N15790, October 2015, Geneva, CH. Summary of the Invention
[0011] The problem the invention aims to solve
[0012] In recent years, motion image coding and decoding technologies have incorporated motion compensation filtering during the motion compensation process for generating predicted images. However, this has led to a significant challenge: the higher the accuracy of motion compensation, the more filter coefficients are required, and consequently, the greater the storage space needed to pre-store these coefficients.
[0013] Furthermore, recent motion picture coding and decoding techniques have employed methods that use high-precision motion vectors to generate predicted images. However, using high-precision motion vectors increases the amount of code required for the difference vectors, thus creating a second challenge: coding efficiency may not necessarily improve.
[0014] The present invention provides an image decoding apparatus, an image encoding apparatus, and a predictive image generation apparatus capable of solving at least one of the first and second problems described above.
[0015] Technical solution
[0016] To address the aforementioned issues, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a filtering unit that operates on an image with motion vector applied to it at a resolution of 1 / Mac pixels. The filtering unit operates on the image with motion vector applied by applying filtering coefficients mcFilter[i][k] specified by a phase i (i is an integer between 0 and Mac-1) and a filtering coefficient position k (k is an integer between 0 and Ntaps-1, where Ntaps is the number of taps). The filtering coefficients mcFilter[i][k] are related to a weighted average of filtering coefficients mcFilter[p][k] (P≠i) and mcFilter[q][k] (Q≠i).
[0017] Furthermore, in order to solve the above-mentioned problems, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a motion vector derivation unit that derives motion vectors by adding or subtracting difference vectors from the predictive vectors according to the predictive blocks. The motion vector derivation unit switches the accuracy of the motion vectors derived relative to the predictive blocks according to the size of the predictive blocks.
[0018] Furthermore, in order to solve the above-mentioned problems, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a motion vector derivation unit that derives motion vectors by adding or subtracting difference vectors from the predictive vectors according to the predictive blocks. The motion vector derivation unit switches the accuracy of the motion vectors derived relative to the predictive blocks according to the magnitude of quantization parameters related to the predictive blocks.
[0019] Furthermore, in order to solve the above-mentioned problems, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a motion vector derivation unit that derives a motion vector by adding or subtracting an inversely quantized differential vector from a predictive vector. The motion vector derivation unit switches the precision of the inverse quantization processing for the differential vector based on the quantization value of the quantized differential vector.
[0020] Furthermore, in order to solve the above-mentioned problems, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a motion vector derivation unit that derives a motion vector by adding or subtracting a difference vector from a predictive vector. When a flag indicating the precision of the motion vector is displayed at a first value, the motion vector derivation unit switches the precision of the inverse quantization processing for the difference vector based on the quantization value of the quantized difference vector. When a flag indicating the precision of the motion vector is displayed at a second value, the inverse quantization processing for the difference vector is performed with a fixed precision regardless of the quantization value of the quantized difference vector.
[0021] Furthermore, to address the aforementioned issues, a predictive image generation apparatus of the present invention generates a predictive image by performing motion compensation on a reference image. The predictive image generation apparatus includes a motion vector derivation unit that derives a motion vector by adding or subtracting a difference vector from a predictive vector. When a flag indicating the precision of the motion vector is displayed as a first value, the motion vector derivation unit switches between setting the precision of the inverse quantization processing of the difference vector to a first precision or a second precision based on the quantization value of the quantized difference vector. When the flag indicating the precision of the motion vector is displayed as a second value, the motion vector derivation unit switches between setting the precision of the inverse quantization processing of the difference vector to a third precision or a fourth precision based on the quantization value of the quantized difference vector. At least one of the first and second precisions is higher than the third and fourth precisions.
[0022] Beneficial effects
[0023] Based on the above structure, at least one of the first and second issues mentioned above can be solved. Attached Figure Description
[0024] Figure 1 This is a diagram showing the hierarchical structure of the encoded stream data in this embodiment.
[0025] Figure 2 It is a diagram representing the pattern of PU segmentation. Figure 2 (a)~(h) represent the partition shapes for PU partitioning modes of 2N×2N, 2N×N, 2N×nU, 2N×nD, N×2N, nL×2N, nR×2N, and N×N, respectively.
[0026] Figure 3 This is a concept map representing an example of a list of reference images.
[0027] Figure 4 This is a concept map representing an example of a reference image.
[0028] Figure 5This is a schematic diagram showing the configuration of the image decoding apparatus in this embodiment.
[0029] Figure 6 This is a schematic diagram showing the configuration of the inter-frame prediction parameter decoding unit in this embodiment.
[0030] Figure 7 This is a schematic diagram showing the configuration of the merging prediction parameter derivation unit in this embodiment.
[0031] Figure 8 This is a schematic diagram showing the configuration of the AMVP prediction parameter derivation unit in this embodiment.
[0032] Figure 9 This is a conceptual diagram representing an example of a vector candidate.
[0033] Figure 10 This is a schematic diagram showing the configuration of the inter-frame prediction parameter decoding control unit in this embodiment.
[0034] Figure 11 This is a schematic diagram showing the configuration of the inter-frame prediction image generation unit in this embodiment.
[0035] Figure 12 This is a block diagram illustrating the configuration of the image encoding apparatus in this embodiment.
[0036] Figure 13 This is a schematic diagram showing the configuration of the inter-frame prediction parameter coding unit in this embodiment.
[0037] Figure 14 This is a schematic diagram illustrating the configuration of an image transmission system according to an embodiment of the present invention.
[0038] Figure 15 This is a flowchart illustrating the process of inter-frame prediction syntax decoding performed by the inter-frame prediction parameter decoding control unit of this embodiment.
[0039] Figure 16 This is a flowchart illustrating an example of differential vector decoding processing in this embodiment.
[0040] Figure 17 This is a flowchart illustrating other examples of differential vector decoding processing in this embodiment.
[0041] Figure 18 This is a flowchart illustrating the motion vector derivation process performed by the inter-frame prediction parameter decoding unit in this embodiment.
[0042] Figure 19 This is a flowchart illustrating an example of the difference vector derivation process in this embodiment.
[0043] Figure 20This is a flowchart illustrating an example of the predictive vector loop processing in this embodiment.
[0044] Figure 21 This is a flowchart specifically illustrating an example of the motion vector scale derivation process in this embodiment.
[0045] Figure 22 This is a flowchart that more specifically illustrates other examples of the motion vector scale derivation process of this embodiment.
[0046] Figure 23 This is a flowchart that more specifically illustrates other examples of the motion vector scale derivation process of this embodiment.
[0047] Figure 24 (a)~ Figure 24 (c) is a table showing the relationship between the basic vector accuracy of this embodiment and the parameter (shiftS) indicating the motion vector accuracy set (switched) according to the block size of the object block.
[0048] Figure 25 (a)~ Figure 25 (c) is a table showing the relationship between the parameter (shiftS) that indicates the motion vector accuracy set (switched) according to the block size and motion vector accuracy flag of the object block in this embodiment.
[0049] Figure 26 (a)~ Figure 26 (c) is a table showing the parameter (shiftS) indicating the motion vector accuracy set (switched) according to the QP of this embodiment.
[0050] Figure 27 (a) and Figure 27 (b) is a table showing the motion vector precision (shiftS) set (switched) according to the QP and motion vector precision flag in this embodiment.
[0051] Figure 28 This is a graph showing the relationship between the quantized difference vector and the inverse quantized difference vector in this embodiment.
[0052] Figure 29 This is a flowchart that more specifically illustrates other examples of the motion vector scale derivation process of this embodiment.
[0053] Figure 30 This is a block diagram showing the specific configuration of the motion compensation unit in this embodiment.
[0054] Figure 31 This is a diagram showing an example of the filtering coefficients in this embodiment.
[0055] Figure 32(a) is a diagram showing an example of how the motion compensation filter unit of this embodiment calculates the filter coefficients for odd-numbered phases from the filter coefficients for even-numbered phases. Figure 32 (b) is a diagram showing an example of how the motion compensation filter unit of this embodiment calculates the filter coefficients for even phases from the filter coefficients for odd phases.
[0056] Figure 33 This diagram illustrates the configuration of a transmitting device equipped with the aforementioned image encoding apparatus and a receiving device equipped with the aforementioned image decoding apparatus. Figure 33 (a) indicates a transmitting device equipped with an image encoding device. Figure 33 (b) indicates a receiving device equipped with an image decoding device.
[0057] Figure 34 This diagram illustrates the configuration of a recording device equipped with the aforementioned image encoding device and a playback device equipped with the aforementioned image decoding device. Figure 34 (a) indicates a recording device equipped with an image encoding device. Figure 34 (b) indicates a reproduction device equipped with an image decoding device. Detailed Implementation
[0058] (First Implementation)
[0059] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0060] Figure 14 This is a schematic diagram showing the configuration of the image transmission system 1 of this embodiment.
[0061] Image transmission system 1 is a system that transmits code obtained by encoding an image of an object and displays an image obtained by decoding the transmitted code. Image transmission system 1 is composed of an image encoding device (moving picture encoding device) 11, a network 21, an image decoding device (moving picture decoding device) 31, and an image display device 41.
[0062] The image coding apparatus 11 receives a signal T indicating whether the image is a single layer or multiple layers. A layer is a concept used to distinguish multiple images when there are more than one image constituting a given moment. For example, encoding the same image across multiple layers with different image quality and resolution is called scalable coding, and encoding images from different viewpoints across multiple layers is called view scalable coding. Coding efficiency is significantly improved when prediction is performed between images in multiple layers (inter-layer prediction, inter-viewpoint prediction). Furthermore, even without prediction (simulcast), encoded data can be aggregated.
[0063] Network 21 transmits the encoded stream Te generated by image encoding device 11 to image decoding device 31. Network 21 is the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or a combination thereof. Network 21 is not necessarily limited to a two-way communication network; it can also be a one-way or two-way communication network that transmits broadcast waves such as terrestrial digital broadcasts and satellite broadcasts. Furthermore, network 21 can also be replaced by storage media such as DVDs (Digital Versatile Discs) and Blu-ray Discs (registered trademarks) that record the encoded stream Te.
[0064] The image decoding device 31 decodes the encoded stream Te transmitted by the network 21 to generate one or more decoded layer images Td (decoded viewpoint images Td).
[0065] The image display device 41 displays all or a portion of one or more decoded layer images Td generated by the image decoding device 31. For example, in viewpoint-scalable encoding, a three-dimensional image (stereoscopic image) or a free viewpoint image is displayed when the entire image is displayed, and a two-dimensional image is displayed when only a portion is displayed. The image display device 41 may include, for example, a liquid crystal display or an organic EL (electroluminescence) display. Furthermore, in spatially scalable encoding and signal-to-noise ratio (SNR) scalable encoding, when the image decoding device 31 and the image display device 41 have high processing power, a high-quality enhancement layer image is displayed. Conversely, when the image decoding device 31 and the image display device 41 have only low processing power, a base layer image with high processing and display capabilities that does not require enhancement layer levels is displayed.
[0066] <Structure of encoded stream Te>
[0067] Before providing a detailed description of the image encoding device 11 and the image decoding device 31 of this embodiment, the data structure of the encoded stream Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be explained.
[0068] Figure 1 This is a diagram representing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te exemplarily contains a sequence and multiple images that constitute the sequence. Figure 1 (a)~ Figure 1(f) is a diagram representing the sequence layer of the given sequence SEQ, the picture layer of the specified picture PICT, the slice layer of the specified slice S, the slice data layer of the specified slice data, the coding tree layer of the coding tree unit contained in the slice data, and the coding unit layer of the coding tree (CU: Coding Unit) contained in the coding tree.
[0069] (Sequence layer)
[0070] In the sequence layer, an image decoding device 31 is defined as a set of data referenced for decoding the sequence SEQ (hereinafter also referred to as the object sequence) of the object being processed. The sequence SEQ is as follows: Figure 1 As shown in (a), it includes the Video Parameter Set, Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture PICT, and Supplemental Enhancement Information (SEI). Here, the value following the # indicates the layer ID. Figure 1 The example shown contains coded data with #0 and #1, i.e., layer 0 and layer 1, but the type and number of layers do not depend on this.
[0071] The Video Parameter Set (VPS) defines a set of common coding parameters for multiple motion pictures in a motion picture composed of multiple layers, as well as a set of coding parameters for the multiple layers contained in the motion picture and the layers associated with each layer.
[0072] The Sequence Parameter Set (SPS) specifies a set of encoding parameters that the image decoding device 31 refers to when decoding the object sequence. For example, it specifies the width and height of the image. It should be noted that multiple SPSs can exist. In this case, any one of the multiple SPSs is selected from the PPS.
[0073] The image parameter set (PPS) specifies a set of encoding parameters that the image decoding device 31 refers to when decoding each image in the object sequence. For example, it includes a reference value for the quantization width used for image decoding (pic_init_qp_minus26) and a flag indicating the application of weighted prediction (weighted_pred_flag). It should be noted that multiple PPSs can exist. In this case, any one of the multiple PPSs is selected from the images in the object sequence.
[0074] (Image layer)
[0075] In the image layer, an image decoding device 31 is specified to refer to a set of data for decoding the image PICT (hereinafter also referred to as the object image) of the object being processed. The image PICT is as follows: Figure 1 As shown in (b), it contains slices S0 to SNS-1 (NS is the total number of slices contained in the image PICT).
[0076] It should be noted that, in the following descriptions, where there is no need to distinguish between slices S0 to SNS-1, the code suffix may sometimes be omitted. Furthermore, the same applies to other data contained in the encoded stream Te that are marked with a suffix.
[0077] (Slice layer)
[0078] Within the slice layer, an image decoding device 31 is defined to reference a set of data for decoding the slice S (also called the object slice) of the object being processed. Slice S is as follows: Figure 1 As shown in (c), it includes the slice header SH and the slice data SDATA.
[0079] The slice header SH contains a group of encoded parameters referenced by the image decoding device 31 to determine the decoding method for the object slice. The slice type specification information (slice_type) is an example of the encoded parameters contained in the slice header SH.
[0080] Slice types that can be specified by slice type specification information include: (1) I slices that use only intra-frame prediction during encoding, (2) P slices that use unidirectional prediction or intra-frame prediction during encoding, and (3) B slices that use unidirectional prediction, bidirectional prediction, or intra-frame prediction during encoding.
[0081] It should be noted that the slice header SH may also contain a reference (pic_parameter_set_id) to the image parameter set PPS contained in the above sequence layer.
[0082] (Sliced data layer)
[0083] In the slice data layer, an image decoding device 31 is defined to refer to a set of data for decoding the slice data SDATA of the object being processed. The slice data SDATA is as follows: Figure 1 As shown in (d), it contains a Coded Tree Block (CTB). A CTB is a fixed-size (e.g., 64×64) block that makes up a slice, and is sometimes also called the Largest Cording Unit (LCU) or Coded Tree Unit (CTU).
[0084] (Encoding Tree Layer)
[0085] Encoding tree layers such as Figure 1 As shown in (e), an image decoding device 31 is provided with a set of data referenced for decoding the encoded tree blocks of the processed object. The encoded tree blocks are segmented through recursive quadtree partitioning. The nodes of the tree structure obtained through recursive quadtree partitioning are called the coding tree. The intermediate nodes of the quadtree are coded quadtrees (CQTs), and the encoded tree block itself is defined as the top-level CQT. Each CQT contains a split flag (split_flag). When split_flag is 1, the tree is divided into four encoded tree units (CQTs). When split_flag is 0, the encoded tree unit (CQT) is not partitioned, but instead has a coded unit (CU) as a node. The coded unit (CU) is the terminal node of the encoded tree layer, and no further partitioning occurs at that layer. The coded unit (CU) becomes the basic unit of the encoding process.
[0086] Furthermore, with the size of the coding tree block (CTB) being 64×64 pixels, the size of the coding unit can be any one of 64×64 pixels, 32×32 pixels, 16×16 pixels, and 8×8 pixels.
[0087] (Encoding unit layer)
[0088] Encoding unit layer such as Figure 1 As shown in (f), an image decoding device 31 is provided with a set of data referenced for decoding the encoding units of the object being processed. Specifically, the encoding unit consists of an encoding tree, a prediction tree, a transform tree, and a CU header CUF. The encoding tree specifies segmentation flags, segmentation patterns, prediction modes, etc.
[0089] The prediction tree defines the prediction information (refer to image index, motion vector, etc.) for each prediction block obtained by dividing a coding unit into one or more prediction blocks. In other words, a prediction block is one or more non-repeating regions that constitute a coding unit. Furthermore, the prediction tree contains one or more prediction blocks obtained through the above-described division. It should be noted that the prediction units obtained by further dividing the prediction blocks are referred to as "sub-blocks" below. A sub-block (prediction block) consists of one or more pixels. When the size of the prediction block and the sub-block are equal, there is one sub-block in the prediction block. When the size of the prediction block is larger than the size of the sub-block, the prediction block is divided into sub-blocks. For example, if the prediction block is 8×8 and the sub-block is 4×4, the prediction block is divided horizontally into two parts and vertically into two parts, thus forming four sub-blocks.
[0090] Prediction processing is performed on this prediction block (sub-block). Hereinafter, the prediction block, which is used as the unit of prediction, will also be referred to as the prediction unit (PU).
[0091] There are two main types of segmentation in prediction trees: intra-frame prediction and inter-frame prediction. Intra-frame prediction is prediction within the same image, while inter-frame prediction refers to prediction processing performed between different images (e.g., between display times or between layer images).
[0092] In the case of intra-frame prediction, the segmentation methods are 2N×2N (same size as the coding unit) and N×N.
[0093] Furthermore, in the case of inter-frame prediction, the segmentation method is encoded according to the PU segmentation mode (part_mode) of the coded data. Segmentation options include 2N×2N (same size as the coding unit), 2N×N, 2N×nU, 2N×nD, N×2N, nL×2N, nR×2N, and N×N. It should be noted that 2N×nU indicates that a 2N×2N coding unit is divided into two regions, 2N×0.5N and 2N×1.5N, from top to bottom. 2N×nD indicates that a 2N×2N coding unit is divided into two regions, 2N×1.5N and 2N×0.5N, from top to bottom. nL×2N indicates that a 2N×2N coding unit is divided into two regions, 0.5N×2N and 1.5N×2N, from left to right. nR×2N represents dividing a 2N×2N coding unit into two regions, 1.5N×2N and 0.5N×1.5N, from left to right. The number of divisions can be any one of one, two, or four, therefore the number of PUs contained in the CU can be one to four. These PUs are denoted as PU0, PU1, PU2, and PU3, respectively.
[0094] Figure 2 (a)~ Figure 2 In (h), the location of the boundary of the PU segment in the CU is shown in the specific map for each segmentation type.
[0095] It should be noted that, Figure 2 (a) indicates a 2N×2N PU partitioning mode without CU partitioning.
[0096] also, Figure 2 (b) Figure 2 (c) and Figure 2 (d) represents the shape of the partitions in the cases where the PU partitioning modes are 2N×N, 2N×nU, and 2N×nD, respectively. Hereinafter, the partitions in the cases where the PU partitioning modes are 2N×N, 2N×nU, and 2N×nD are collectively referred to as horizontal partitions.
[0097] also, Figure 2(e) Figure 2 (f), and Figure 2 (g) represents the shape of the partition when the PU partitioning mode is N×2N, nL×2N, and nR×2N, respectively. Hereinafter, partitions with PU partitioning types of N×2N, nL×2N, and nR×2N will be collectively referred to as vertical partitions.
[0098] In addition, both horizontal and vertical partitions are collectively referred to as rectangular partitions.
[0099] also, Figure 2 (h) represents the shape of the partition when the PU partitioning mode is N×N. Based on the shape of its partitions, it will also... Figure 2 (a) and Figure 2 The PU segmentation pattern of (h) is called square segmentation. Furthermore, it will also... Figure 2 (b)~ Figure 2 (g) The PU segmentation pattern is called non-square segmentation.
[0100] In addition, Figure 2 (a)~ Figure 2 In (h), each region is assigned a number to represent its identifier, and the regions are processed sequentially according to the order of these identifiers. That is, the identifier indicates the scanning order of the partitions.
[0101] In addition, Figure 2 (a)~ Figure 2 In (h), the upper left corner is set as the reference point (origin) of CU.
[0102] Furthermore, in the transform tree, the coding unit is divided into one or more transform blocks, and the position and size of each transform block are defined. In other words, a transform block is one or more non-repeating regions that constitute a coding unit. Moreover, the transform tree contains one or more transform blocks obtained through the above-described division.
[0103] In the segmentation of the transform tree, there are segments that allocate regions of the same size as the coding unit as transform blocks, and segments that are performed recursively by quadtree segmentation, similar to the segmentation of tree blocks described above.
[0104] The transformation process is performed according to this transformation block. Hereinafter, the transformation block, which is used as the unit of transformation, will also be referred to as the transformation unit (TU).
[0105] (Prediction parameters)
[0106] The predicted image of the prediction unit is derived from the prediction parameters attached to the prediction unit. These prediction parameters include intra-frame prediction parameters and inter-frame prediction parameters. The following explains the inter-frame prediction parameters. The inter-frame prediction parameters consist of prediction list utilization flags predFlagL0 and predFlagL1, reference image indices refIdxL0 and refIdxL1, and vectors mvL0 and mvL1. The prediction list utilization flags predFlagL0 and predFlagL1 indicate whether to use the respective reference image lists, referred to as L0 lists and L1 lists. When the value is 1, the corresponding reference image list is used. It should be noted that when referred to as "flags indicating whether it is ××" in this specification, 1 represents the case of ××, and 0 represents the case of not ××. In logical NOT, logical multiplication, etc., 1 is treated as true, and 0 is treated as false (the same applies below). However, in actual devices and methods, other values can also be used as true and false values. When using two reference image lists, i.e., when predFlagL0 = 1 and predFlagL1 = 1, it corresponds to dual prediction. Conversely, when using one reference image list, i.e., when (predFlagL0, predFlagL1) = (1, 0) or (predFlagL0, predFlagL1) = (0, 1), it corresponds to single prediction. It should be noted that the information regarding the prediction list utilization flag can also be represented by the inter-frame prediction flag inter_pred_idc, which will be described later. Typically, the prediction list utilization flag is used in the prediction image generation unit (prediction image generation apparatus) 308 and the prediction parameter memory 307, which will be described later. Furthermore, when decoding information about which reference image list is used based on the encoded data, the inter-frame prediction flag inter_pred_idc is used.
[0107] Syntax elements used to derive the inter-prediction parameters contained in the encoded data include, for example, the segmentation mode (part_mode), merge flag (merge_flag), merge index (merge_idx), inter-prediction flag (inter_pred_idc), reference image index (refIdxLX), prediction vector index (mvp_LX_idx), and difference vector (mvdLX).
[0108] (Refer to an example of a list of images)
[0109] Next, an example of a list of reference images will be explained. The list of reference images is stored in the reference image memory 306 ( Figure 5 (A column formed by referring to the images.) Figure 3This is a conceptual diagram representing an example of a list of reference images. In the list of reference images 601, five rectangles arranged in a left-right column represent reference images. The codes P1, P2, Q0, P3, and P4, shown from left to right, represent the codes for each reference image. The P in P1, etc., represents viewpoint P, and the Q in Q0 represents a viewpoint Q different from viewpoint P. The suffixes P and Q indicate the image sequence number POC. The downward arrow directly below refIdxLX indicates the reference image index refIdxLX, which is the index in the reference image storage 306 that references reference image Q0.
[0110] (Refer to the example in the image)
[0111] Next, an example of the reference image used in deriving vectors will be explained. Figure 4 This is a concept diagram representing an example of a reference image. In Figure 4 In the diagram, the horizontal axis represents the time of display, and the vertical axis represents the viewpoint. Figure 4 The six rectangles arranged in two rows and three columns represent images. In the bottom row, the second rectangle from the left represents the image of the object to be decoded (the object image), and the remaining five rectangles represent reference images. Reference image Q0, indicated by an upward arrow from the object image, is an image displayed at the same time as the object image but from a different viewpoint. Reference image Q0 is used in displacement prediction based on the object image. Reference image P1, indicated by an arrow pointing left from the object image, is a past image from the same viewpoint as the object image. Reference image P2, indicated by an arrow pointing right from the object image, is a future image from the same viewpoint as the object image. Reference image P1 or P2 is used in motion prediction based on the object image. (Inter-frame prediction flags and prediction list utilization flags)
[0112] The relationship between the inter-frame prediction flag and the prediction list utilization flags predFlagL0 and predFlagL1 is as described below, and they can be interchanged. Therefore, either the prediction list utilization flag or the inter-frame prediction flag can be used as an inter-frame prediction parameter. Furthermore, in the following, decisions using the prediction list utilization flag can be replaced with the inter-frame prediction flag. Conversely, decisions using the inter-frame prediction flag can also be replaced with the prediction list utilization flag.
[0113] Inter-frame prediction flag = (predFlagL1<<1) + predFlagL0
[0114] predFlagL0 = Inter-frame prediction flag &1
[0115] predFlagL1 = Inter-frame prediction flag >> 1
[0116] Here, >> means right shift and << means left shift.
[0117] (Merger Prediction and AMVP Prediction)
[0118] The decoding (encoding) methods for prediction parameters include merge prediction mode and AMVP (Adaptive Motion Vector Prediction) mode. The merge flag is used to identify these. Both merge prediction mode and AMVP mode use the prediction parameters of already processed blocks to derive the prediction parameters of the target PU. Merge prediction mode directly uses the derived prediction parameters of nearby PUs without including the prediction list with the predFlagLX (or inter-prediction flag inter_pred_idc), the reference image index refIdxLX, and the motion vector mvLX in the encoded data. Conversely, AMVP mode includes the inter-prediction flag inter_pred_idc, the reference image index refIdxLX, and the motion vector mvLX in the encoded data. It should be noted that the motion vector mvLX is encoded as the prediction vector index mvp_LX_idx and the difference vector mvdLX that identify the prediction vector mvpLX.
[0119] The inter-frame prediction flag `inter_pred_idc` represents the type and number of reference images, taking any value from `Pred_L0`, `Pred_L1`, and `Pred_Bi`. `Pred_L0` and `Pred_L1` indicate the use of reference images stored in lists called the L0 list and L1 list, respectively, and both use a single reference image (single prediction). Predictions using the L0 and L1 lists are respectively called L0 predictions and L1 predictions. `Pred_Bi` indicates the use of two reference images (double prediction), and both reference images stored in the L0 and L1 lists are used. The prediction vector index `mvp_LX_idx` represents the index of the prediction vector, and the reference image index `refIdxLX` represents the index of the reference image stored in the reference image list. It should be noted that `LX` is a notation used when there is no distinction between L0 and L1 predictions. By replacing `LX` with `L0` and `L1`, a distinction is made between parameters for the L0 list and parameters for the L1 list. For example, refIdxL0 is the reference image index used for L0 prediction, refIdxL1 is the reference image index used for L1 prediction, and refIdx(refIdxLX) is a tag used when there is no distinction between refIdxL0 and refIdxL1.
[0120] The merge index merge_idx is an index indicating whether to use any of the prediction parameters from the prediction parameter candidates (merge candidates) derived from the processed block as the prediction parameters for the decoded object block.
[0121] It should be noted that an "object block" can be a prediction block that is one level higher than multiple prediction blocks, or it can be a coding unit that contains the aforementioned multiple prediction blocks.
[0122] (Motion vector and displacement vector)
[0123] In motion vector mvLX, there are two main components: a narrow motion vector (representing the offset between blocks in two images at different times) and a disparity vector (representing the offset between two blocks at the same time). In the following explanation, we will not distinguish between motion vector and disparity vector, and will simply refer to them as motion vector mvLX. The prediction vector and difference vector associated with motion vector mvLX are referred to as prediction vector mvpLX and difference vector mvdLX, respectively. Whether motion vector mvLX and difference vector mvdLX are motion vectors or disparity vectors is identified using the reference image index refIdxLX that accompanies the vector.
[0124] (Composition of an image decoding device)
[0125] Next, the configuration of the image decoding device 31 in this embodiment will be described. Figure 5 This is a schematic diagram showing the configuration of the image decoding apparatus 31 of this embodiment. The image decoding apparatus 31 is configured to include an entropy decoding unit 301, a prediction parameter decoding unit (predictive image generation device) 302, a reference image memory (reference image storage unit, frame memory) 306, a prediction parameter memory (predictive parameter storage unit, frame memory) 307, a prediction image generation unit 308, an inverse quantization / inverse DCT unit 311, an addition unit 312, and a residual storage unit 313 (residual recording unit).
[0126] Furthermore, the prediction parameter decoding unit 302 is configured to include an inter-frame prediction parameter decoding unit (motion vector derivation unit) 303 and an intra-frame prediction parameter decoding unit 304. The prediction image generation unit 308 is configured to include an inter-frame prediction image generation unit 309 and an intra-frame prediction image generation unit 310.
[0127] The entropy decoding unit 301 performs entropy decoding on the externally input encoded stream Te, and decodes each code (syntax element) separately. The separated code contains prediction information for generating the predicted image and residual information for generating the difference image, etc.
[0128] The entropy decoding unit 301 outputs a portion of the separated code to the prediction parameter decoding unit 302. This portion of the separated code includes, for example, the prediction mode PredMode, the segmentation mode part_mode, the merge flag merge_flag, the merge index merge_idx, the inter-frame prediction flag inter_pred_idc, the reference image index refIdxLX, the prediction vector index mvp_LX_idx, and the difference vector mvdLX. Control over which code to decode is based on the instruction from the prediction parameter decoding unit 302. The entropy decoding unit 301 outputs quantization coefficients to the inverse quantization / inverse DCT unit 311. These quantization coefficients are obtained by performing a Discrete Cosine Transform (DCT) on the residual signal and then quantizing it during the encoding process.
[0129] The inter-frame prediction parameter decoding unit 303 decodes the inter-frame prediction parameters based on the code input from the entropy decoding unit 301 and with reference to the prediction parameters stored in the prediction parameter memory 307.
[0130] The inter-frame prediction parameter decoding unit 303 outputs the decoded inter-frame prediction parameters to the prediction image generation unit 308, and stores them in the prediction parameter memory 307. Details of the inter-frame prediction parameter decoding unit 303 will be described below.
[0131] The intra-prediction parameter decoding unit 304 decodes the intra-prediction parameters based on the code input from the entropy decoding unit 301 and with reference to the prediction parameters stored in the prediction parameter memory 307. Intra-prediction parameters are parameters used in the process of predicting image patches within an image, such as the intra-prediction mode (IntraPredMode). The intra-prediction parameter decoding unit 304 outputs the decoded intra-prediction parameters to the prediction image generation unit 308 and stores them in the prediction parameter memory 307.
[0132] The intra-prediction parameter decoding unit 304 can also derive intra-prediction modes that differ in luminance and chrominance. In this case, the intra-prediction parameter decoding unit 304 decodes the luminance prediction mode IntraPredModeY as the luminance prediction parameter and the chrominance prediction mode IntraPredModeC as the chrominance prediction parameter. The luminance prediction mode IntraPredModeY is mode 35, corresponding to planar prediction (0), DC prediction (1), and directional prediction (2-34). The chrominance prediction mode IntraPredModeC uses any one of planar prediction (0), DC prediction (1), directional prediction (2-34), and LM mode (35). The intra-prediction parameter decoding unit 304 decodes a flag indicating whether IntraPredModeC is the same mode as the luminance mode. If the flag indicates that it is the same mode as the luminance mode, IntraPredModeY is assigned to IntraPredModeC. Furthermore, if the flag indicates a mode different from the luminance mode, the intra-prediction parameter decoding unit 304 can also decode the plane prediction (0), DC prediction (1), direction prediction (2-34), and LM mode (35) as IntraPredModeC.
[0133] The reference image memory 306 stores the reference image blocks (reference image blocks) generated by the addition unit 312 in a predetermined location according to the image and blocks of the decoding object.
[0134] The prediction parameter memory 307 stores prediction parameters in predetermined locations according to the images and blocks of the decoding object. Specifically, the prediction parameter memory 307 stores the inter-frame prediction parameters decoded by the inter-frame prediction parameter decoding unit 303, the intra-frame prediction parameters decoded by the intra-frame prediction parameter decoding unit 304, and the prediction mode predMode separated by the entropy decoding unit 301. Among the stored inter-frame prediction parameters, for example, there are prediction list utilization flag predFlagLX (inter-frame prediction flag inter_pred_idc), reference image index refIdxLX, and motion vector mvLX.
[0135] The prediction image generation unit 308 receives a prediction mode, predMode, input from the entropy decoding unit 301, and prediction parameters are input from the prediction parameter decoding unit 302. Furthermore, the prediction image generation unit 308 reads a reference image from the reference image memory 306. Under the prediction mode represented by predMode, the prediction image generation unit 308 uses the input prediction parameters and the read reference image to generate a prediction image block P (predicted image).
[0136] Here, when the prediction mode `predMode` represents the inter-frame prediction mode, the inter-frame prediction image generation unit 309 uses the inter-frame prediction parameters input from the inter-frame prediction parameter decoding unit 303 and the read reference image to generate a prediction image block P through inter-frame prediction. The prediction image block P corresponds to the prediction unit `PU`. `PU` is equivalent to a part of the image formed by multiple pixels that have undergone prediction processing as described above; that is, it is equivalent to the decoding object block that has undergone one prediction processing.
[0137] The inter-frame prediction image generation unit 309, relative to the prediction list, uses a list of reference images (L0 list or L1 list) with the flag predFlagLX set to 1. Based on the reference image indicated by the reference image index refIdxLX, it reads the reference image block located at the position shown by the motion vector mvLX based on the decoded object block from the reference image memory 306. The inter-frame prediction image generation unit 309 predicts the read reference image block to generate a prediction image block P. The inter-frame prediction image generation unit 309 outputs the generated prediction image block P to the adder 312.
[0138] When the prediction mode `predMode` indicates intra-prediction mode, the intra-prediction image generation unit 310 performs intra-prediction using intra-prediction parameters input from the intra-prediction parameter decoding unit 304 and read reference images. Specifically, the intra-prediction image generation unit 310 reads reference image blocks that are to be decoded from the reference image memory 306 and are located within a predetermined range from the decoded target block in the already decoded blocks. The predetermined range is any one of the adjacent blocks (left, upper left, upper, upper right) when the decoded target blocks are moved sequentially in a so-called raster scan order, depending on the intra-prediction mode. The raster scan order is the order in which each row in each image is moved sequentially from left to right from top to bottom.
[0139] The intra-prediction image generation unit 310 predicts the read-out reference image block in the prediction mode shown in IntraPredMode, and generates a predicted image block. The intra-prediction image generation unit 310 outputs the generated predicted image block P to the addition unit 312.
[0140] In the intra-prediction parameter decoding unit 304, when deriving intra-prediction modes that differ in luminance and chrominance, the intra-prediction image generation unit 310 generates a luminance prediction image block using any one of planar prediction (0), DC prediction (1), or directional prediction (2-34) based on the luminance prediction mode IntraPredModeY. Furthermore, the intra-prediction image generation unit 310 generates a chrominance prediction image block using any one of planar prediction (0), DC prediction (1), directional prediction (2-344), or LM mode (35) based on the chrominance prediction mode IntraPredModeC.
[0141] The inverse quantization / inverse DCT unit 311 dequantizes the quantization coefficients input from the entropy decoding unit 301 to obtain the DCT coefficients. The inverse quantization / inverse DCT unit 311 performs an inverse DCT (Inverse Discrete Cosine Transform) on the obtained DCT coefficients to calculate the decoded residual signal. The inverse quantization / inverse DCT unit 311 outputs the calculated decoded residual signal to the adder unit 312 and the residual storage unit 313.
[0142] The addition unit 312 adds the signal values of the predicted image block P input from the inter-frame prediction image generation unit 309 and the intra-frame prediction image generation unit 310, and the decoded residual signal input from the inverse quantization / inverse DCT unit 311, pixel by pixel, to generate a reference image block. The addition unit 312 stores the generated reference image block in the reference image memory 306, and outputs the decoded layer image Td obtained by integrating the reference image block to the outside.
[0143] (The structure of the inter-frame prediction parameter decoding unit)
[0144] Next, the configuration of the inter-frame prediction parameter decoding unit 303 will be explained.
[0145] Figure 6 This is a schematic diagram showing the configuration of the inter-frame prediction parameter decoding unit 303 in this embodiment. The inter-frame prediction parameter decoding unit 303 is configured to include an inter-frame prediction parameter decoding control unit (motion vector derivation unit) 3031, an AMVP prediction parameter derivation unit 3032, an addition unit 3035, and a merging prediction parameter derivation unit 3036.
[0146] The inter-frame prediction parameter decoding control unit 3031 instructs the entropy decoding unit 301 to decode the code (syntax elements) associated with inter-frame prediction, and extracts the code (syntax elements) contained in the encoded data, such as the segmentation mode part_mode, merge flag merge_flag, merge index merge_idx, inter-frame prediction flag inter_pred_idc, reference image index refIdxLX, prediction vector index mvp_LX_idx, and difference vector mvdLX.
[0147] The inter-frame prediction parameter decoding control unit 3031 first extracts the merging flag. When the inter-frame prediction parameter decoding control unit 3031 extracts a certain syntax element, it indicates to the entropy decoding unit 301 that the syntax element is decoded, and the matching syntax element is read from the encoded data. Here, when the value indicated by the merging flag is 1, i.e., indicating the merge prediction mode, the inter-frame prediction parameter decoding control unit 3031 extracts the merge index merge_idx as the prediction parameter for merge prediction. The inter-frame prediction parameter decoding control unit 3031 outputs the extracted merge index merge_idx to the merge prediction parameter derivation unit 3036.
[0148] When the merge flag (merge_flag) is 0, indicating AMVP prediction mode, the inter-frame prediction parameter decoding control unit 3031 extracts AMVP prediction parameters from the encoded data using the entropy decoding unit 301. Examples of AMVP prediction parameters include the inter-frame prediction flag (inter_pred_idc), the reference image index (refIdxLX), the prediction vector index (mvp_LX_idx), and the difference vector (mvdLX). The inter-frame prediction parameter decoding control unit 3031 outputs the prediction list derived from the extracted inter-frame prediction flag (inter_pred_idc) to the AMVP prediction parameter derivation unit 3032 and the prediction image generation unit 308 using the flag (predFlagLX) and the reference image index (refIdxLX). Figure 5 In addition, the prediction parameter memory 307 is stored in the prediction parameter memory. Figure 5 The inter-frame prediction parameter decoding control unit 3031 outputs the extracted prediction vector index mvp_LX_idx to the AMVP prediction parameter derivation unit 3032. The inter-frame prediction parameter decoding control unit 3031 outputs the extracted difference vector mvdLX to the addition unit 3035.
[0149] Figure 7This is a schematic diagram showing the configuration of the merging prediction parameter derivation unit 3036 in this embodiment. The merging prediction parameter derivation unit 3036 includes a merging candidate derivation unit 30361 (prediction vector calculation unit) and a merging candidate selection unit 30362. The merging candidate storage unit 303611 stores the merging candidates input from the merging candidate derivation unit 30361. It should be noted that the merging candidates include a prediction list constructed using the flag predFlagLX, the motion vector mvLX, and the reference image index refIdxLX. In the merging candidate storage unit 303611, the index is assigned to the stored merging candidates according to a predetermined rule.
[0150] The candidate merging derivation unit 30361 directly derives merging candidates using the motion vectors of neighboring blocks that have already undergone decoding and the reference image index refIdxLX. Alternatively, affine prediction can also be used to derive merging candidates. This method will be described in detail below. The candidate merging derivation unit 30361 can use affine prediction in the spatial merging candidate derivation process, temporal merging (inter-frame merging) candidate derivation process, combined merging candidate derivation process, and zero-merging candidate derivation process, as described later. It should be noted that affine prediction is performed on a sub-block basis, and the prediction parameters are stored in the prediction parameter memory 307 on a sub-block basis. Alternatively, affine prediction can also be performed on a pixel basis.
[0151] (Spatial Merging Candidate Derivation Processing)
[0152] As part of the spatial merging candidate derivation process, the merging candidate derivation unit 30361 reads the prediction parameters (predFlagLX, motion vector mvLX, and reference image index refIdxLX) stored in the prediction parameter memory 307 according to prescribed rules, and derives the read prediction parameters as merging candidates. The read prediction parameters are the prediction parameters of each block (e.g., all or part of the blocks connected to the lower left, upper left, and upper right ends of the decoded object block) within a predetermined range from the decoded object block. The merging candidates derived by the merging candidate derivation unit 30361 are stored in the merging candidate storage unit 303611.
[0153] (Time merging candidate derivation processing)
[0154] As part of the temporal merging derivation process, the merging candidate derivation unit 30361 reads the prediction parameters of a block in a reference image, including the lower right coordinates of the decoded object block, from the prediction parameter memory 307, and uses them as merging candidates. The reference image can be specified using, for example, the reference image index refIdxLX specified in the slice header, or by using the smallest reference image index refIdxLX among the reference image indices refIdxLX of the blocks adjacent to the decoded object block. The merging candidates derived by the merging candidate derivation unit 30361 are stored in the merging candidate storage unit 303611.
[0155] (Combined candidate derivation processing)
[0156] As part of the combination and merging derivation process, the merging candidate derivation unit 30361 derives the combined merging candidate by combining the vectors of two different completed merging candidates that have already been derived and stored in the merging candidate storage unit 303611 and the reference image index as vectors L0 and L1, respectively. The merging candidates derived by the merging candidate derivation unit 30361 are stored in the merging candidate storage unit 303611.
[0157] (Zero-merging candidate derivation processing)
[0158] As part of the zero-merge candidate derivation process, the merge candidate derivation unit 30361 derives merge candidates whose reference image index refIdxLX is 0 and whose X and Y components of the motion vector mvLX are both 0. The merge candidates derived by the merge candidate derivation unit 30361 are stored in the merge candidate storage unit 303611.
[0159] The merge candidate selection unit 30362 selects the merge candidate stored in the merge candidate storage unit 303611 that has an index corresponding to the merge index merge_idx input from the inter-frame prediction parameter decoding control unit 3031 as the inter-frame prediction parameter of the target PU. The merge candidate selection unit 30362 stores the selected merge candidate in the prediction parameter memory 307 and outputs it to the prediction image generation unit 308. Figure 5 ).
[0160] Figure 8 This is a schematic diagram showing the configuration of the AMVP prediction parameter derivation unit 3032 in this embodiment. The AMVP prediction parameter derivation unit 3032 includes a vector candidate derivation unit 3033 (vector calculation unit) and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 reads the vectors (motion vectors or displacement vectors) stored in the prediction parameter memory 307 based on the reference image index refIdx, and uses them as the prediction vector mvpLX. The read vector is the vector of each block (for example, all or part of the blocks that are connected to the lower left, upper left, and upper right ends of the decoding object block, respectively) within a predetermined range from the decoding object block.
[0161] The vector candidate selection unit 3034 selects the vector candidate indicated by the prediction vector index mvp_LX_idx input from the inter-frame prediction parameter decoding control unit 3031 from the vector candidates read by the vector candidate derivation unit 3033, and uses it as the prediction vector mvpLX. The vector candidate selection unit 3034 outputs the selected prediction vector mvpLX to the addition unit 3035.
[0162] Furthermore, the vector candidate selection unit 3034 may also employ a configuration that performs cyclic processing on the selected prediction vector mvpLX, as described later.
[0163] The vector candidate storage unit 30331 stores the vector candidates input from the vector candidate derivation unit 3033. It should be noted that the vector candidates are composed of the prediction vector mvpLX. In the vector candidate storage unit 30331, indices are assigned to the stored vector candidates according to a predetermined rule.
[0164] The vector candidate derivation unit 3033 uses affine prediction to derive vector candidates. The vector candidate derivation unit 3033 can also use affine prediction for spatial vector candidate derivation processing, temporal vector (inter-frame vector) candidate derivation processing, combined vector candidate derivation processing, and zero vector candidate derivation processing, as described later. It should be noted that affine prediction is performed on a sub-block basis, and prediction parameters are stored in the prediction parameter memory 307 on a sub-block basis. Alternatively, affine prediction can also be performed on a pixel basis.
[0165] Figure 9 This is a conceptual diagram representing an example of a vector candidate. Figure 9 The predicted vector list 602 shown is a list formed by multiple vector candidates derived in the vector candidate derivation unit 3033. In the predicted vector list 602, the five rectangles arranged in a left-right column represent the regions indicating the predicted vectors. The downward arrow directly below the second mvp_LX_idx from the left and the mvpLX below it indicate that the predicted vector index mvp_LX_idx is the index of the reference vector mvpLX in the prediction parameter memory 307.
[0166] Vector candidates are generated based on the vectors of the blocks referenced by the vector candidate selection unit 3034. The blocks referenced by the vector candidate selection unit 3034 are blocks that have completed the decoding process, or blocks within a predetermined range from the target block (e.g., adjacent blocks). It should be noted that adjacent blocks include not only blocks that are spatially adjacent to the target block, such as the left block and the top block, but also blocks that are temporally adjacent to the target block, such as blocks obtained from blocks that are at the same position as the target block but displayed at a different time.
[0167] The addition unit 3035 adds the prediction vector mvpLX input from the AMVP prediction parameter derivation unit 3032 and the difference vector mvdLX input from the inter-frame prediction parameter decoding control unit 3031 to calculate the motion vector mvLX. The addition unit 3035 outputs the calculated motion vector mvLX to the prediction image generation unit 308. Figure 5 ).
[0168] Figure 10 This is a schematic diagram showing the configuration of the inter-frame prediction parameter decoding control unit 3031 in this embodiment. The inter-frame prediction parameter decoding control unit 3031 includes an additional prediction flag decoding unit 30311, a merge index decoding unit 30312, a vector candidate index decoding unit 30313, and (not shown) a segmentation mode decoding unit, a merge flag decoding unit, an inter-frame prediction flag decoding unit, a reference image index decoding unit, and a vector difference decoding unit. The segmentation mode decoding unit, merge flag decoding unit, merge index decoding unit, inter-frame prediction flag decoding unit, reference image index decoding unit, vector candidate index decoding unit 30313, and vector difference decoding unit decode the segmentation mode part_mode, merge flag merge_flag, merge index merge_idx, inter-frame prediction flag inter_pred_idc, reference image index refIdxLX, prediction vector index mvp_LX_idx, and difference vector mvdLX, respectively. (Inter-frame prediction image generation unit 309)
[0169] Figure 11 This is a schematic diagram showing the configuration of the inter-frame prediction image generation unit 309 in this embodiment. The inter-frame prediction image generation unit 309 is configured to include a motion compensation unit 3091 and a weight prediction unit 3094.
[0170] (Motion compensation)
[0171] The motion compensation unit 3091, based on the prediction list input from the inter-frame prediction parameter decoding unit 303, uses the flag predFlagLX, the reference image index refIdxLX, and the motion vector mvLX to read from the reference image memory 306 the block located at a position offset by the motion vector mvLX from the position of the decoded target block of the reference image specified by the reference image index refIdxLX, thereby generating a motion-compensated image. Here, if the precision of the motion vector mvLX is not integer precision, a filter called motion compensation filtering is performed to generate pixels at fractional positions, generating the motion-compensated image. Hereinafter, the motion-compensated image predicted by L0 is referred to as predSamplesL0, and the motion-compensated image predicted by L1 is referred to as predSamplesL1. Without distinguishing between the two, it is referred to as predSamplesLX.
[0172] (Weight Prediction)
[0173] The weight prediction unit 3094 generates a predicted image block P (predicted image) by multiplying the input motion displacement image predSamplesLX by weight coefficients. The input motion displacement image predSamplesLX is an image that has undergone residual prediction when residual prediction is performed. When one of the reference list utilization flags (predFlagL0 or predFlagL1) is 1 (in the case of single prediction), without using weight prediction, the following processing is performed to make the input motion displacement image predSamplesLX (LX is L0 or L1) consistent with the number of pixels.
[0174] predSamples[x][y]=Clip3(0, (1<<bitDepth)-1,(predSamplesLX[x][y]+offset1)> >shift1)
[0175] Here, shift1 = 14-bitDepth, offset1 = 1 << (shift1 - 1).
[0176] Furthermore, when both flags (predFlagL0 or predFlagL1) in the reference list are 1 (in the case of dual prediction), without using weighted prediction, the following process is performed to average the input motion displacement images predSamplesL0 and predSamplesL1 and make their average consistent with the number of pixels.
[0177] predSamples[x][y]=Clip3(0, (1<<bitDepth)-1,(predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)> >shift2)
[0178] Here, shift2 = 15-bitDepth, offset2 = 1 << (shift2 - 1).
[0179] Furthermore, in the case of single prediction, when performing weighted prediction, the weight prediction unit 3094 derives the weight prediction coefficient w0 and the offset value o0 from the encoded data and performs the following processing.
[0180] predSamples[x][y]=Clip3(0, (1<<bitDepth)-1,((predSamplesLX[x][y]*w0+2log2WD-1)> >log2WD)+o0)
[0181] Here, log2WD is a variable representing the specified shift amount.
[0182] Furthermore, in the case of dual prediction, when performing weighted prediction, the weight prediction unit 3094 derives the weight prediction coefficients w0, w1, o0, o1 from the encoded data and performs the following processing.
[0183] predSamples[x][y]=Clip3(0, (1< <bitDepth)-1,(predSamplesL0[x][y]*w0+predSamplesL1[x][y]*w1+((o0+o1+1)<<log2WD))> >(log2WD+1))
[0184] <Motion Vector Decoding Processing>
[0185] The following is for reference Figures 15-27 The motion vector decoding process of this embodiment will be described in detail.
[0186] As can be clearly seen from the above description, the motion vector decoding process of this embodiment includes the process of decoding the syntax elements related to inter-frame prediction (also known as motion syntax decoding process) and the process of deriving motion vectors (motion vector derivation process).
[0187] (Motion grammar decoding processing)
[0188] Figure 15 This is a flowchart illustrating the process of inter-frame prediction syntax decoding performed by the inter-frame prediction parameter decoding control unit 3031. Figure 15 In the following description, unless otherwise specified, each process is performed by the inter-frame prediction parameter decoding control unit 3031.
[0189] First, in step S101, the merge flag is decoded. In step S102, it is determined whether merge_flag ! = 0.
[0190] If merge_flag! = 0 (which is Y in S102), the merge index merge_idx is decoded in S103, and the process proceeds to the motion vector derivation in merge mode (S201). Figure 18 (a)).
[0191] If merge_flag! = 0 is false (N in S102), the inter-frame prediction flag inter_pred_idc is decoded in S104, the reference image index refIdxL0 is decoded in S105, the differential vector syntax mvdL0 is decoded in S106, and the prediction vector index mvp_L0_idx is decoded in S107.
[0192] In S108, the reference image index refIdxL1 is decoded; in S109, the difference vector syntax mvdL1 is decoded; in S110, the prediction vector index mvp_L1_idx is decoded, and the process proceeds to motion vector derivation in AMVP mode (S301). Figure 18 (b)).
[0193] It should be noted that when the inter-frame prediction flag `inter_pred_idc` is 0, i.e., when L0 prediction (PRED_L0) is indicated, the processing steps S108 to S110 are not required. On the other hand, when the inter-frame prediction flag `inter_pred_idc` is 1, i.e., when L1 prediction (PRED_L1) is indicated, the processing steps S105 to S107 are not required. Furthermore, when the inter-frame prediction flag `inter_pred_idc` is 2, i.e., when double prediction (PRED_B) is indicated, steps S105 to S110 are executed.
[0194] (Differential vector decoding processing)
[0195] Figure 16 This is a flowchart that more specifically illustrates the differential vector decoding process in steps S106 and S109 above. Up to this point, the horizontal and vertical components of the motion vector and the differential vector mvdLX are not distinguished and are simply labeled as mvLX and mvdLX. Here, in order to clarify the cases where the syntax of horizontal and vertical components is required and the cases where the processing of horizontal and vertical components is required, [0] and [1] are used to label each component.
[0196] like Figure 16 As shown, firstly, in step S10611, the syntax mvdAbsVal[0] representing the absolute value of the difference between the horizontal motion vectors is decoded from the encoded data. Secondly, in step S10612, it is determined whether the absolute value of the difference between the (horizontal) motion vectors is 0.
[0197] mvdAbsVal[0]! = 0.
[0198] If the absolute value of the horizontal motion vector difference, mvdAbsVal[0]! = 0, is true (Y in S10612), in S10614, the syntax mv_sign_flag[0] representing the code (positive or negative) of the horizontal motion vector difference is decoded from the encoded data, and proceeds to S10615. On the other hand, if mvdAbsVal[0]! = 0, is false (N in S10612), in S10613, mv_sign_flag[0] is inferred to be 0, and proceeds to S10615.
[0199] Next, in step S10615, the syntax mvdAbsVal[1] representing the absolute value of the vertical motion vector difference is decoded, and in step S10612, it is determined whether the absolute value of the (vertical) motion vector difference is 0.
[0200] mvdAbsVal[1]! = 0.
[0201] If mvdAbsVal[1]! = 0 is true (Y in S10616), the syntax mv_sign_flag[1] representing the code (positive or negative) of the vertical motion vector difference is decoded from the encoded data in S10618. On the other hand, if mvdAbsVal[1]! = 0? is false (N in S10616), the syntax mv_sign_flag[1] representing the code (positive or negative) of the vertical motion vector difference is set to 0 in S10617.
[0202] In the above, the absolute value of the motion vector difference, mvdAbsVal, and the code of the motion vector difference, mvd_sign_flag, are represented by a vector formed by {horizontal component, vertical component}, with the horizontal component selected as [0] and the vertical component selected as [1]. Alternatively, the vertical component could be selected as [0] and the horizontal component as [1]. Furthermore, the vertical component is processed after the horizontal component, but the order of processing is not limited to this. For example, the vertical component can be processed first, followed by the horizontal component (the same applies below).
[0203] Figure 17 It means to utilize and in Figure 16 The flowcharts illustrate examples of different processing methods, specifically the decoding of the difference vector in steps S106 and S109. Figure 16 The steps already explained in [the document], in Figure 17 The same symbols are also assigned to them, thus omitting the explanation.
[0204] exist Figure 17In the example shown, the further decoding of the motion vector precision flag mvd_dequant_flag is similar to... Figure 16 different.
[0205] That is, in Figure 17 In the example shown, following the already explained S10617 and S10618, S10629 derives the variable nonZeroMV, which represents whether the difference vector is 0, to determine whether the difference vector is 0.
[0206] nonZeroMV! = 0?
[0207] Here, the variable nonZeroMV can be derived as follows.
[0208] nonZeroMV=mvdAbsVal[0]+mvdAbsVal[1]
[0209] When nonZeroMV! = 0 is true (Y in S10629), that is, when the difference vector is not 0, the motion vector precision flag mvd_dequant_flag is decoded from the encoded data in S10630. Furthermore, when nonZeroMV! = 0 is false (N in S10629), mvd_dequant_flag is set to 0 without decoding it from the encoded data. In other words, mvd_dequant_flag is decoded only when nonZeroMV! = 0, except when the difference vector is 0.
[0210] It should be noted that the motion vector precision flag `mvd_dequant_flag` is used to toggle the precision of the motion vectors. Furthermore, when this flag is set to select whether to set the motion vector precision to full pixel, it can be written as (or marked as) `integer_mv_flag`.
[0211] (Derivation and processing of motion vectors)
[0212] Next, use Figures 18-27 The derivation of motion vectors is explained.
[0213] Figure 18 This is a flowchart illustrating the motion vector derivation process performed by the inter-frame prediction parameter decoding unit 303 in this embodiment.
[0214] (Derivation of motion vectors under the merged prediction mode)
[0215] Figure 18(a) is a flowchart illustrating the motion vector derivation process under the merging prediction mode. For example... Figure 18 As shown in (a), in S201, the merge candidate derivation unit 30361 derives the merge candidate list mergeCandList. In S202, the merge candidate selection unit 30362 selects the merge candidate mvLX specified by the merge index merge_idx based on mergeCandList[merge_idx]. For example, it is derived by mvLX = mergeCandList[merge_idx]. (Motion vector derivation processing in AMVP mode)
[0216] In AMVP mode, the differential motion vector mvdLX is derived from the decoded syntax mvdAbsVal and mv_sign_flag, and the differential motion vector mvdLX is added to the prediction vector mvpLX to derive the motion vector mvLX. In the syntax description, mvdAbsVal[0], mvdAbsVal[1], etc. and [0], [1] are used to distinguish between horizontal and vertical components. However, for simplicity, the components are not distinguished in the following text and are simply referred to as mvdAbsVal, etc. In fact, the motion vector has horizontal and vertical components. Therefore, the processing of each component is performed sequentially without distinguishing between components.
[0217] on the other hand, Figure 18 (b) is a flowchart illustrating the motion vector derivation process in AMVP mode. For example... Figure 18 (b) As shown, in S301, the vector candidate derivation unit 3033 derives the motion vector prediction sublist mvpListLX. In S302, the vector candidate selection unit 3034 selects the motion vector candidate (predicted vector, predicted motion vector) mvpLX = mvpListLX[mvp_LX_idx] specified by the prediction vector index mvp_LX_idx.
[0218] Next, in S303, the inter-frame prediction parameter decoding control unit 3031 derives the differential vector mvdLX. For example... Figure 18 As shown in S304 of (b), the vector candidate selection unit 3034 can also perform cyclic processing on the selected prediction vector. Next, in S305, the prediction vector mvpLX and the difference vector mvdLX are added in the addition unit 3035 to calculate the motion vector mvLX. That is, through...
[0219] mvLX=mvpLX+mvdLX
[0220] Calculate mvLX.
[0221] (Difference vector derivation)
[0222] Next, use Figure 19 to illustrate the differential vector derivation process. Figure 19 is a flowchart that more specifically represents the differential vector derivation process in step S303 above. The differential vector derivation process consists of the following two processes. Inverse quantization process (PS_DQMV): Inverse quantization is performed on the value decoded from the encoded data and quantized, that is, the absolute value of the motion vector difference mvdAbsVal (quantized value), and the absolute value of the motion vector difference mvdAbsVal with a specific precision (for example, the basic vector precision described later) is derived. Code assignment process (PS_SIGN): Determine the code of the derived absolute value of the motion vector difference mvdAbsVal, and derive the motion vector difference mvdLX.
[0223] In Figure 19 the following description of the description, unless otherwise specified, each process is performed by the inter-frame prediction parameter decoding control unit 3031.
[0224] As Figure 19 shown, in S3031, the motion vector scale shiftS, which is a parameter specifying the motion vector precision, is derived. In S3032, it is judged whether the motion vector scale is >0. When the motion vector scale >0 is true, that is, shiftS>0 (Y in S3032), in S3033, for example, the differential vector is inverse quantized by using a shift operation of shiftS. Here, the shift operation is more specifically, for example, a process of shifting the quantized absolute value of the motion vector difference mvdAbsVal to the left by shiftS. By
[0225] mvdAbsVal = mvdAbsVal << shiftS formula (Scale)
[0226] it is carried out (process PS_DQMV0).
[0227] Then, in S3034, the code assignment process of the differential vector is performed, and it advances to S3041. It should be noted that this code assignment process (process PS_SIGN) is through
[0228] mvdLX = mvdAbsVal * (1 - 2 * mv_sign_flag) formula (sign)
[0229] This is done as follows. That is, the motion vector difference mvdLX is derived from the absolute value of the motion vector difference mvdAbsVal according to the value of mv_sign_flag. It should be noted that when the motion vector scale > 0 is false, i.e., shiftS = 0 (N in S3032), the process proceeds to S3034 without going through S3033. It should be noted that the inverse quantization of the difference vector with the shift applied when the value 0 (shiftS = 0) is used does not affect the value of the difference vector. Therefore, even when the motion vector scale > 0 is false, it is also possible to adopt a configuration where S3033 is not skipped, but rather S3033 is performed after deriving the motion vector scale as 0 (shiftS = 0).
[0230] In addition, since the motion vector scale generally uses a value of 0 or more, when the motion vector scale is other than 0, the motion vector scale is always positive (> 0). Thus, the determination of "motion vector scale!= 0" can also be used instead of the determination of "motion vector scale > 0". It should be noted that in this specification, the same applies to other processes where the determination of "motion vector scale > 0" is made.
[0231] (Prediction Vector Loop Processing)
[0232] Next, use Figure 20 to describe the prediction vector loop processing (predicted motion vector loop processing). Figure 20 is a flowchart that more specifically represents the prediction vector loop processing in step S304 above. In the following description of the description in Figure 20 , unless otherwise specified, each process is performed by the vector candidate selection unit 3034. As shown in Figure 20 , the motion vector scale is derived in S3041, and it is judged in S3042 whether the motion vector scale > 0. When the motion vector scale > 0 is true (Y in S3042), that is, when the inverse quantization of the difference vector according to the motion vector scale is performed, in S3043, the predicted motion vector mvpLX can also be looped based on the motion vector scale, that is, through
[0233] mvpLX = round(mvpLX, shiftS)
[0234] to perform a loop processing (processing PS_PMVROUND). Here, round(mvpLX, shiftS) represents a function that performs a loop processing on the predicted motion vector mvpLX using shiftS. For example, the loop processing can use the following formulas (SHIFT - 1) to (SHIFT - 4), etc., to set the predicted motion vector mvpLX to a value (discrete value) of 1 << shiftS units.
[0235] After S304, proceed to S305. In S305, derive the motion vector mvLX from the prediction vector mvpLX and the difference vector mvdLX. It should be noted that if the motion vector scale > 0 is false (N in S3042), the prediction motion vector mvpLX proceeds to S305 without looping, and the motion vector mvLX is derived.
[0236] (Derivation of motion vector scale using motion vector accuracy flags)
[0237] Next, use Figure 21 The derivation of motion vector scale using the motion vector accuracy flag (PS_P0) is explained. Figure 21 More specifically, S3031 mentioned above (refer to...) Figure 19 ) and S3041 (refer to Figure 20 The flowchart for the derivation of motion vector scales in (). Figure 21 For ease of explanation, the processing of S3041 is specifically illustrated, but it can also be applied to S3031. Figure 21 The processing shown.
[0238] exist Figure 21 In the following description, unless otherwise specified, each process is performed by the inter-frame prediction parameter decoding control unit 3031.
[0239] like Figure 21 As shown, in S304111, it is determined whether the motion vector precision flag mvd_dequant_flag satisfies the condition.
[0240] `mvd_dequant_flag!` = 0. When `mvd_dequant_flag!` = 0 is true (Y in S304111), for example, when `mvd_dequant_flag` = 1, in S304112, for example, `shiftS` is set to be equal to the motion vector base precision `mvBaseAccu` (>0), a parameter representing the reference for motion vector precision, and the motion vector precision is set to full pixels. Here, the value of `mvBaseAccu` is, for example, 2. After S304112, proceed to S3042. When the motion vector precision flag `mvd_dequant_flag!` = 0 is false (N in S304111), for example, when `mvd_dequant_flag` = 0, in S304113, `shiftS` is set to 0, and proceed to S3042. In this case, the motion vector precision is set to 1 / 4 pixel. It should be noted that in S304112, for example, shiftS can also be set to the motion vector base precision mvBaseAccu-1, which represents the baseline parameter for motion vector precision. In this case, when mvd_dequant_flag is 1 (other than 0), the precision of the motion vector is set to half a pixel.
[0241] It should be noted that, in the above description, the motion vector precision is reduced when `mvd_dequant_flag` is 1, and maintained when `mvd_dequant_flag` is 0. Alternatively, the value of `mvd_dequant_flag` could be set to another value, for example, 0 to reduce motion vector precision and 1 to maintain it. That is, for the flags shown in this specification, their values and the content they represent can be set to any combination.
[0242] Thus, based on the above structure, the motion vector precision is switched by referring to the motion vector precision flag, thereby enabling the use of motion vectors with more appropriate precision. On the other hand, since the motion vector precision flag needs to be included in the encoded data, the amount of code increases, and there may be cases where the encoding efficiency is not as improved as expected.
[0243] The following is an example of how motion vectors can be constructed to improve coding efficiency and use appropriate precision.
[0244] (The motion vector scale of the object block was derived using the block size)
[0245] As one example of constructing motion vectors that improve coding efficiency and use appropriate precision (deriving PS_P1A), using... Figure 22A description will be given of a motion vector scale derivation process that utilizes the block size of an object block. Figure 22 This is a flowchart that more specifically represents the motion vector scale derivation process in S3031 (refer to Figure 19 ) and S3041 (refer to Figure 20 ). In Figure 22 , for the sake of convenience in explanation, the process of S3041 is specifically exemplified, but the process shown in Figure 22 can also be applied to S3031.
[0246] In the following description in the description of Figure 22 , unless otherwise specified, each process is performed by the inter-frame prediction parameter decoding control unit 3031.
[0247] As shown in Figure 22 , in S304121, it is judged whether the block size blkW satisfies
[0248] blkW < TH (TH is a prescribed threshold).
[0249] When blkW < TH is true (Y in S304121), that is, when the block size blkW is small, in S304122, it is set to
[0250] shiftS = shiftM,
[0251] and it proceeds to S3042. In addition, when blkW < TH is false (N in S304121), that is, when the block size blkW is large, in S304123, it is set to
[0252] shiftS = shiftN,
[0253] and it proceeds to S3042. Here, shiftM and shiftN are scale parameters that satisfy shiftM > shiftN, and shiftN can also be 0.
[0254] It should be noted that Figure 22 the above-described process can also be expressed by the following formula.
[0255] shiftS = (blkW < TH)? shiftM : shiftN (formula P1A)
[0256] It should be noted that in a configuration where there is a case where the width blkW and the height blkH of the block size are different, as the threshold determination of the block size, blkW + blkH < TH can also be used instead of blkW < TH. It should be noted that the above-described changes can also be appropriately applied to other processes in this specification.
[0257] Furthermore, the branch determination is not limited to < (larger); ≤ (lower) can also be used as an equivalent construct. Alternatively, the branches of Y and N can be set to > and ≥ respectively. It should be noted that the above modifications can also be appropriately applied to other processing in this specification.
[0258] The above describes how the motion vector scale is derived based on the size of the block, with larger block sizes resulting in smaller values for the motion vector scale (thus improving motion vector accuracy). As in the example above, in a configuration that categorizes block sizes based on their size and switches the motion vector scale accordingly, the number of block size categories is not limited to two; it can also be three or more.
[0259] Thus, based on the above configuration, the precision of the differential vector can be switched according to the block size. For example, when the block size is larger than a specified value, a high-precision vector can be used, and when the block size is smaller than the specified value, a low-precision motion vector can be used. In this way, by switching the precision of the motion vector according to the block size, a motion vector with more appropriate precision can be used.
[0260] Furthermore, in the above configuration, the precision of the differential vector can be switched without utilizing the motion vector precision flag. Therefore, there is no need to encode and decode the motion vector precision flag, thereby reducing the amount of encoded data. Furthermore, this improves encoding efficiency. (Motion vector scale derivation processing utilizing block size and motion vector precision flag)
[0261] Furthermore, as another example of configuration (deriving PS_P1B), it is then used... Figure 23 The derivation of motion vector scale using the block size of the object block and the motion vector precision flag is explained. Figure 23 This is a flowchart that more specifically illustrates the derivation process of the motion vector scale in S3031 and S3041 described above. Figure 23 For ease of explanation, the processing of S3041 is specifically illustrated, but it can also be applied to S3031. Figure 23 The processing shown.
[0262] like Figure 23 As shown, in S304131, it is determined whether mvd_dequant_flag! = 0. If mvd_dequant_flag! = 0 is false (N in S304131), that is, if mvd_dequant_flag is 0, then in S304132, it is determined whether the block size blkW satisfies the condition.
[0263] blkW < TH (where TH is a specified threshold). When the block size blkW < TH is true (Y in S304132), that is, when the block size blkW is small,
[0264] In S304133, it is set to
[0265] shiftS = shiftM,
[0266] Proceed to S3042. Additionally, when the block size blkW < TH is false (N in S304132), that is, when the block size blkW is large, in S304133, a value different from the value (shiftM) in the case of a smaller block value is set as the motion vector scale, and it is set to
[0267] shiftS = shiftN,
[0268] Proceed to S3042.
[0269] Furthermore, when mvd_dequant_flag!= 0 is true (Y in S304131), that is, when mvd_dequant_flag is 1,
[0270] It is set to shiftS = shiftL, and proceed to S3042. Here, shiftL, shiftM, and shiftN are scale parameters that satisfy shiftL ≥ shiftM > shiftN, and shiftN can also be 0. shiftN = 0 is equivalent to not performing inverse quantization on the motion vector (differential motion vector) (encoding the motion vector that has become less rough through the quantization scale). It should be noted that when shiftL = mvBaseAccu and mvd_dequant_flag is 1, it is also appropriate to set it to full pixels.
[0271] It should be noted that Figure 23 Part of the processing of S304132 to S304134 (equivalent to Figure 22 ) can also be expressed by the above (formula P1A).
[0272] Figure 23 Overall, the entire processing of
[0273] shiftS = mvd_dequant_flag!= 0? shiftL : (blkW < TH)? shiftM : shiftN (formula P1B)
[0274] As described above, the following configuration is formed: a mode in which multiple precisions of motion vectors are used according to the block size, and the switching is performed such that the value of the motion vector scale becomes smaller (the motion vector precision is increased) as the block size becomes larger. Note that the classification of the block size is not limited to two classifications, and a configuration classified into three or more may also be adopted.
[0275] According to the above configuration, the precision of the motion vector is determined by referring to both the block size and the motion vector precision flag, and thus, a motion vector with a more appropriate precision can be used. For example, when the precision of the motion vector is represented by an integer precision (low precision) by the motion vector precision flag, the precision of the differential vector is set to the integer precision (low precision) regardless of the block size. In addition, when the precision of the motion vector is represented by a fractional precision (high precision) by the motion vector precision flag, the precision of the differential vector can be further switched according to the block size.
[0276] Therefore, according to the above configuration, an improvement in coding efficiency can be achieved by using a motion vector with a more appropriate precision.
[0277] Note that the derivation of the above motion vector can be described in another way as follows. That is, the inter-frame prediction parameter decoding unit 303 (motion vector derivation unit) derives the motion vector by adding or subtracting the differential vector to / from the prediction vector for each prediction block. The inter-frame prediction parameter decoding unit 303 switches the precision of the motion vector derived for the prediction block (particularly, the shift value used to derive the absolute value of the motion vector difference) according to the size of the prediction block.
[0278] In addition, the motion vector derived by the process of the above motion vector derivation can be expressed by the following formula. That is, when the motion vector to be derived is labeled as mvLX, the prediction vector is labeled as mvpLX, the differential vector is labeled as mvdLX, and the rounding process is labeled as round(), the shift amount shiftS can be determined according to the size of the prediction block, and
[0279] mvLX = round(mvpLX) + (mvdLX << shiftS)
[0280] is used to determine mvLX.
[0281] Note that the inverse quantization of the differential vector mvdLX represented by the above (mvdLX << shiftS) term can also be performed in the absolute value of the motion vector difference. That is, as described below, a configuration in which the inverse quantization process of the absolute value of the motion vector difference mvdAbsVal is performed and the code process is also performed may be adopted.
[0282] mvdAbsVal = mvdAbsVal(= |qmvd|) << shiftS
[0283] mvdLX = mvdAbsVal * (1 - 2 * mv_sign_flag)
[0284] mvLX = round(mvpLX) + mvdLX
[0285] In the above, it is expressed in terms of the composition of updating the variables mvdAbsVal and mvdLX. However, when it is expressed using "'" for the sake of clarity of the process, it can be expressed in the following manner.
[0286] mvdAbsVal' = mvdAbsVal(= |qmvd|) << shiftS
[0287] mvdLX' = mvdAbsVal' * (1 - 2 * mv_sign_flag)
[0288] mvLX = round(mvpLX) + mvdLX'
[0289] In addition, as described in the parentheses () above, in addition to mvdAbsVal, the absolute value of the differential motion vector before inverse quantization (after quantization) can also be represented by qmvd.
[0290] (Various specific examples of loop processing)
[0291] Loop processing was mentioned in the above description. The specific examples of loop processing do not limit this embodiment. For example, using
[0292] round(mvpLX) = (mvpLX >> shiftS << shiftS) ··· (SHIFT - 1)
[0293] is fine. In addition, the variable inside round() is not limited to mvpLX. In addition, for loop processing, in addition to the above example, examples using the following formula can also be cited. For example, an offset value
[0294] offsetS = 1 << (shiftS - 1),
[0295] is set as
[0296] round(mvpLX) = ((mvpLX + offsetS) >> shiftS) << shiftS ··· (SHIFT - 2). In addition,
[0297] round(mvpLX)=mvpLX>0? (mvpLX>>shiftS)<<shiftS:-(((-mvpLX)> >shiftS)< <shiftS)···(SHIFT-3)
[0298] That is, when the object of the loop is negative, the following structure can also be used: temporarily transform it to a positive value by multiplying by -1, and then perform the same process as (formula: SHIFT-1), and then transform it to a negative value by multiplying by -1.
[0299] also,
[0300] round(mvpLX)=mvpLX>0? ((mvpLX+offsetS)>>shiftS)<<shiftS:-((((-mvpLX+offsetS))> >shiftS)< <shiftS)···(SHIFT-4)
[0301] That is, it is also possible to use a combination of the processing of formula (SHIFT-2) and formula (SHIFT-3).
[0302] (An example of motion vector precision using block size switching)
[0303] The following uses Figure 24 A specific example of motion vector precision (derivation process P1A) using the block size switching of object blocks is explained. Figure 24 (a)~ Figure 24 (c) is a table showing the relationship between the basic vector precision and the parameter (shiftS) indicating the motion vector precision set (switched) according to the block size of the object block. For example, the inter-frame prediction parameter decoding control unit 3031 can use... Figure 24 (a)~ Figure 24 The above is performed in the manner shown in example (c). Figure 22 The processing shown in S304121 to S304123.
[0304] Note that in this specification, the concept of "basic vector" is introduced, and mvBaseAccu is used to represent the parameter specifying the accuracy of this basic vector. In addition, the "basic vector" is assumed to be decoded with an accuracy of 1<<mvBaseAccu. Here, the term "basic vector" is only for convenience, and the "basic vector" is just a concept introduced as a reference for specifying the accuracy of the motion vector. In this patent, the accuracy of the vector when input to the motion compensation filtering unit 30912 is given in terms of the basic vector accuracy. For example, when mvBaseAccu = 2, the basic vector is set to be processed with an accuracy of 1 / 4 (= 1 / (1<<mvBaseAccu)) pixel. In the motion compensation filtering unit 30912, filtering is performed using a set of filtering coefficients of phases 0 to M-1 (M = (1<<mvBaseAccu)) (filtering coefficients of 0 to M-1). In addition, a configuration may be adopted in which the motion vector derived (utilized) is stored in the prediction parameter memory 108 using the accuracy of this basic vector.
[0305] Figure 24 (a) is a table showing the relationship between the block size of the target block, the basic vector accuracy, and the parameter shiftS indicating the motion vector accuracy when the motion vector accuracy is switched to two values. In Figure 24 In the example shown by "I" in (a), mvBaseAccu = 3, and the accuracy of the basic vector is 1 / 8 pel. In the example shown by "I", when the block size blkW of the target block satisfies blkW >= 64, shiftS = 0 is set, and the motion vector accuracy becomes 1 / 8 pel. On the other hand, when the block size blkW satisfies blkW < 64, shiftS = 1 is set, and the motion vector accuracy becomes 1 / 4 pel.
[0306] In Figure 24 (a) the example shown by "II", mvBaseAccu = 4, and the accuracy of the basic vector is 1 / 16 pel. In the example shown by "II", when the block size blkW of the target block satisfies blkW >= 64, shiftS = 0 is set, and the motion vector accuracy becomes 1 / 16 pel. On the other hand, when the block size blkW satisfies blkW < 64, shiftS = 2 is set, and the motion vector accuracy becomes 1 / 4 pel.
[0307] In Figure 24In the example shown in "III" of (a), mvBaseAccu = 6, and the precision of the base vector is 1 / 64pel. In the example shown in "III", when the block size blkW of the object block satisfies blkW >= 64, shiftS is set to 0, and the motion vector precision becomes 1 / 64pel. On the other hand, when the block size blkW satisfies blkW < 64, shiftS is set to 4, and the motion vector precision becomes 1 / 4pel.
[0308] Figure 24 (b) is a table showing the relationship between the block size, basic vector precision, and parameter (shiftS) indicating the motion vector precision of the object block when the motion vector precision is switched to three values.
[0309] exist Figure 24 In the example shown in (b), the basic vector precision is 1 / 64 pel. When the block size blkW of the object block satisfies blkW>=64, shiftS is set to 0, and the motion vector precision becomes 1 / 64 pel. Furthermore, when the block size blkW satisfies blkW>=32 && blkW<64, shiftS is set to 2, and the motion vector precision becomes 1 / 16 pel. Moreover, when the block size blkW satisfies blkW<32, shiftS is set to 4, and the motion vector precision becomes 1 / 4 pel.
[0310] Figure 24 (c) is a table showing the relationship between the block size, basic vector precision, and parameter (shiftS) indicating the motion vector precision of the object block when the motion vector precision is switched to five values.
[0311] exist Figure 24 In the example shown in (c), the basic vector precision is 1 / 64 pel. When the block size blkW of the object block satisfies blkW>=128, shiftS is set to 0, and the motion vector precision becomes 1 / 64 pel. Furthermore, when the block size blkW satisfies blkW>=64 && blkW<128, shiftS is set to 1, and the motion vector precision becomes 1 / 32 pel. Furthermore, when the block size blkW satisfies blkW>=32 && blkW<64, shiftS is set to 2, and the motion vector precision becomes 1 / 16 pel. Furthermore, when the block size blkW satisfies blkW>=16 && blkW<32, shiftS is set to 3, and the motion vector precision becomes 1 / 8 pel. Furthermore, when the block size blkW satisfies blkW<16, shiftS is set to 4, and the motion vector precision becomes 1 / 4 pel.
[0312] (Utilization of Motion Vector Precision MVQStep)
[0313] The inter-frame prediction parameter decoding control unit 3031 may also be configured to perform inverse quantization by multiplying the quantization step size (Quantization step size) MVQStep of the motion vector, instead of using the left shift performed by the motion vector scale shiftS. That is, inverse quantization can be performed by the following formula instead of formula (Scale).
[0314] mvdAbsVal = mvdAbsVal * MVQStep formula (QStep)
[0315] Here, MVQStep and shiftS satisfy
[0316] shiftS = log2(MVQStep)
[0317] the relationship. This is equivalent to the following.
[0318] MVQStep = 1 << shiftS = 2 shifts
[0319] Furthermore, when the accuracy of the basic motion vector is 1 / 8, when the quantization step MVQStep is 1, the accuracy (MVStep) of the encoded motion vector becomes 1 / 8, and when the quantization step MVQStep is 2, the accuracy (MVStep) of the encoded motion vector becomes 1 / 4. Therefore, when the accuracy of the basic motion vector is set to 1 / mvBaseAccu, in the quantization step MVQStep, the accuracy MVStep of the encoded motion vector becomes 1 / mvBaseAccu * MVQStep. For example, when using the product of MVQStep instead of the shift using the quantization scale shiftS to perform Figure 24 (a) the switching of the motion vector accuracy shown by "I", when the block size blkW satisfies blkW >= 64, set MVQStep = 1 (= 1 << ShiftS = 1 << 0 = 1). In other words, set MVStep = 1 / 8 = (1 / mvBaseAccu * MVQStep = 1 / 8 * 1). On the other hand, when the block size blkW satisfies blkW < 64, set MVQStep = 2 (= 1 << ShiftS = 1 << 1 = 2). In other words, set MVStep = 1 / 4 = (1 / 8 * 2).
[0320] In addition, when using MVQStep to perform Figure 24In the case of switching the motion vector precision as shown in (b), when the block size blkW satisfies blkW >= 64, MVQStep is set to 1 (= 1 << shiftS = 1 << 0). In other words, MVStep is set to 1 / 64 = (1 / 64 * 1). In addition, when the block size blkW satisfies blkW >= 32 && blkW < 64, MVQStep is set to 4 (= 1 << ShiftS = 1 << 2 = 4). In other words, MVStep is set to 1 / 16 = (1 / 64 * 4). In addition, when the block size blkW satisfies blkW < 32, MVQStep is set to 16 (= 1 << ShiftS = 1 << 4). In other words, MVStep is set to 1 / 4 = (1 / 64 * 16).
[0321] In addition, when using MVQStep for Figure 24 In the case of switching the motion vector precision as shown in (c), when the block size blkW satisfies blkW >= 128, MVQStep is set to 1. In other words, MVStep is set to 1 / 64. In addition, when the block size blkW satisfies blkW >= 64 && blkW < 128, MVQStep is set to 2. In other words, MVStep is set to 1 / 32. In addition, when the block size blkW satisfies blkW >= 32 && blkW < 64, MVQStep is set to 3. In other words, MVStep is set to 1 / 16. In addition, when the block size blkW satisfies blkW >= 16 && blkW < 32, MVQStep is set to 4. In other words, MVStep is set to 1 / 8. In addition, when the block size blkW < 16 is satisfied, MVQStep is set to 5. In other words, MVStep is set to 1 / 4. (Example of motion vector precision using block size and motion vector precision flag switching)
[0322] Next, use Figure 25 A specific example of the motion vector precision (derivation process PS_P1B) switched using the block size and the mvd_dequant_flag as the motion vector precision flag is described. Figure 25 (a) to Figure 25 (c) is a table showing the relationship with the parameter (shiftS) indicating the motion vector precision set (switched) by the block size of the object block and the motion vector precision flag. It should be noted that in Figure 25 (a) to Figure 25 (c), an example of setting the basic vector precision to 1 / 16 is shown, but any value can be applied to the value of the basic vector precision. The inter-frame prediction parameter decoding control unit 3031 can also use Figure 25 (a) to Figure 25(c) Perform the above in the manner of the example shown Figure 23 The processing of S304131 to S304135 shown.
[0323] In Figure 25 In the example shown in (a), when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is greater than a specified value (large block size), shiftS is set to 0 and the motion vector precision becomes 1 / 16 pel. In addition, when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is less than the specified value (small block size), shiftS is set to 2 and the motion vector precision becomes 1 / 4 pel. On the other hand, when the motion vector precision flag is mvd_dequant_flag = 1, shiftS is set to 4 and the motion vector precision becomes 1 pel (full pixel). When expressed as an equation, it is as follows.
[0324] shiftS = mvd_dequant_flag!= 0? 4 : (blkW < TH)? 2 : 0 (equivalent to equation P1B)
[0325] In Figure 25 In the example shown in (b), when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is greater than a specified value (large block size), shiftS is set to 0 and the motion vector precision MVStep becomes 1 / 16 pel. In addition, when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is less than the specified value (small block size), shiftS is set to 2 and the motion vector precision becomes 1 / 4 pel. On the other hand, when the motion vector precision flag is mvd_dequant_flag = 1 and the block size is greater than the specified value (large block size), shiftS is set to 3 and the motion vector precision becomes 1 / 2 pel (half pixel). In addition, when the motion vector precision flag is mvd_dequant_flag = 1 and the block size is less than the specified value (small block size), shiftS is set to 4 and the motion vector precision becomes 1 pel (full pixel).
[0326] In Figure 25In the example shown in (c), when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is larger than the specified value (large block size), shiftS is set to 0, and the motion vector precision becomes 1 / 16 pel. Furthermore, when the motion vector precision flag is mvd_dequant_flag = 0 and the block size is smaller than the specified value (small block size), shiftS is set to 1, and the motion vector precision becomes 1 / 8 pel. On the other hand, when the motion vector precision flag is mvd_dequant_flag = 1 and the block size is larger than the specified value (large block size), shiftS is set to 2, and the motion vector precision becomes 1 / 4 pel. Furthermore, when the motion vector precision flag is mvd_dequant_flag = 1 and the block size is smaller than the specified value (small block size), shiftS is set to 3, and the motion vector precision becomes 1 / 2 pel (half a pixel).
[0327] (Example of motion vector precision switching using QP)
[0328] It should be noted that in the above example, the configuration for deriving shiftS based on the block size of the object block was described. Alternatively, the inter-frame prediction parameter decoding control unit 3031 (motion vector derivation unit) can derive shiftS based on the QP (Quantization Parameter) as a quantization parameter instead of the block size of the object block (derivation processing PS_P2A). In particular, by deriving shiftS based on the magnitude of QP (or the predicted value of QP), a high-precision motion vector (smaller shiftS) is used when QP is small, and a low-precision motion vector (larger shiftS) is used when QP is large. For example, QP is determined based on a predetermined value; thus, a high-precision motion vector is used when QP is less than the predetermined value, and a low-precision motion vector is used otherwise.
[0329] Here, use Figure 26 An example of motion vector precision using QP switching is illustrated. Figure 26 (a)~ Figure 26 (c) is a table representing the parameter (shiftS) indicating the accuracy of the motion vector set (switched) according to the QP. For example, the inter-frame prediction parameter decoding control unit 3031 can also use... Figure 26 (a)~ Figure 26 The difference vector derivation is performed in the manner shown in example (c). Figure 26 (a)~ Figure 26 The description in (c) does not specifically mention the value of the basic vector precision, but any value of the basic vector precision can be used. Figure 26Examples (a), (b), and (c) show examples of switching two, three, and five values based on QP, respectively, but the number of switches (the number of categories in QP) is not limited to these. Furthermore, the domain values used for QP classification are not limited to the examples shown in the figures.
[0330] Figure 26 (a) is a table showing the relationship between QP and motion vector accuracy (shiftS) when the motion vector accuracy is switched to two values. Figure 26 In the example shown in (a), when QP is small (QP<24), shiftS is set to 0. On the other hand, when QP is large (QP>=24), shiftS is set to a value greater than that when QP is small, which is shiftS=1.
[0331] exist Figure 26 The example shown in (b) is a table illustrating the relationship between QP and the parameter (shiftS) indicating the motion vector precision when the motion vector precision is switched to three values. Figure 26 As shown in (b), shiftS is set to 0 when QP is small (QP<12). Furthermore, shiftS is set to 1 when QP is moderate (QP>=12 && QP<24). Furthermore, shiftS is set to 2 when QP is large (QP>=36).
[0332] Figure 26 (c) is a table showing the correspondence between QP and the parameter (shiftS) indicating the motion vector precision when the motion vector precision is switched to five values. For example... Figure 26 As shown in (c), when QP < 12, shiftS is set to 0. Furthermore, when QP >= 12 && QP < 18, shiftS is set to 1. Furthermore, when QP >= 18 && QP < 24, shiftS is set to 2. Furthermore, when QP >= 24 && QP < 36, shiftS is set to 3. Furthermore, when QP >= 36, shiftS is set to 4.
[0333] Based on the above structure, the precision of the motion vector derived for the prediction block is switched according to the magnitude of the quantization parameter. Therefore, a predicted image can be generated using motion vectors with appropriate precision. It should be noted that this structure can also be used in conjunction with a structure that switches the precision of the motion vector based on a flag.
[0334] (Utilization of Motion Vector Precision MVQStep)
[0335] In addition, the inter-frame prediction parameter decoding control unit 3031 may also adopt the following configuration: instead of the above-mentioned shiftS, MVQStep is derived based on the block size of the object block as the motion vector accuracy.
[0336] For example, using MVQStep Figure 26 (a) shows the switching of motion vector precision. When QP < 24, MVQStep is set to 16. In other words, MVStep is set to 1 / 16. On the other hand, when QP >= 24, MVQStep is set to 4. In other words, MVStep is set to 1 / 4.
[0337] In addition, using MVQStep Figure 26 (b) shows the switching of motion vector precision. When QP < 12, MVQStep is set to 64. In other words, MVStep is set to 1 / 64. Furthermore, when QP >= 12 and QP < 36, MVQStep is set to 16. In other words, MVStep is set to 1 / 16. Additionally, when QP < 36, MVQStep is set to 4. In other words, MVStep is set to 1 / 4.
[0338] In addition, using MVQStep Figure 26 (c) When switching motion vector precision, if QP < 12, set MVQStep = 64.
[0339] In other words, MVStep is set to 1 / 64. Furthermore, when QP >= 12 && QP < 18, MVQStep is set to 32. In other words, MVStep is set to 1 / 32. Furthermore, when QP >= 18 && QP < 24, MVQStep is set to 16. In other words, MVStep is set to 1 / 16. Furthermore, when QP >= 24 && QP < 36, MVQStep is set to 8. In other words, MVStep is set to 1 / 8. Furthermore, when QP >= 36, MVQStep is set to 4. In other words, MVStep is set to 1 / 4.
[0340] (An example of switching motion vector precision using QP and motion vector precision flags)
[0341] Next, use Figure 27 An example is given of using QP and mvd_dequant_flag as a marker of motion vector precision to switch motion vector precision (derived processing PS_P2B). Figure 27 (a) and Figure 27(b) is a table showing the motion vector precision (shiftS) set (switched) according to QP and motion vector precision flag. The inter-frame prediction parameter decoding control unit 3031 can also... Figure 27 The differential vector derivation is performed in the manner shown in examples (a) and (b).
[0342] exist Figure 27 (a) and Figure 27 (b) shows an example of setting the base vector precision to 1 / 16pel (mvBaseAccu = 4), but the value of the base vector precision (mvBaseAccu) can be any value.
[0343] exist Figure 27 In the example shown in (a), when the motion vector precision flag mvd_dequant_flag is 1 (other than 0), the value of shiftS is fixed regardless of QP. Furthermore, when the motion vector precision flag mvd_dequant_flag is 0, the value of the motion vector scale shiftS is determined based on QP. For example, when the motion vector precision flag mvd_dequant_flag = 0 and QP is less than a specified value (small QP), shiftS is set to 0, and the motion vector precision becomes 1 / 16 pel. Furthermore, when the motion vector precision flag mvd_dequant_flag = 0 and QP is greater than a specified value (large QP), shiftS is set to 2, and the motion vector precision becomes 1 / 4 pel. On the other hand, when the motion vector precision flag mvd_dequant_flag = 1, shiftS is fixedly set to mvBaseAccu (= 4), and the motion vector precision becomes 1 pel (full pixel). Thus, even when it changes according to QP, it is appropriate to set shiftS for the case where the motion vector precision flag mvd_dequant_flag is 1 to be greater than shiftS for the other cases (where mvd_dequant_flag is 0) (setting the motion vector precision to low precision).
[0344] exist Figure 27(b) illustrates an example of deriving the motion vector scale shiftS from QP even when the motion vector precision flag is 1 (other than 0). Specifically, when the motion vector precision flag is mvd_dequant_flag = 0 and QP is less than a specified value (small QP), shiftS is set to 0, and the motion vector precision becomes 1 / 16 pel. Furthermore, when the motion vector precision flag is mvd_dequant_flag = 0 and QP is greater than a specified value (large QP), shiftS is set to 4, and the motion vector precision becomes 1 pel. On the other hand, when the motion vector precision flag is mvd_dequant_flag = 1 and QP is less than a specified value (small QP), shiftS is set to 3, and the motion vector precision becomes 1 / 2 pel (half a pixel). Furthermore, when the motion vector precision flag is mvd_dequant_flag = 1 and QP is greater than a specified value (large QP), shiftS is set to mvBaseAccu (= 4), and the motion vector precision becomes 1 pel (full pixel). Therefore, with the motion vector precision flag mvd_dequant_flag set to 1, it is more appropriate to switch between half-pixels and full pixels.
[0345] Furthermore, the motion vector derived through the above-described derivation of motion vectors can be expressed as follows. That is, when the motion vector of the derived object is labeled as mvLX, the prediction vector as mvpLX, the difference vector as mvdLX, and the loop processing as round(), the shift amount shiftS is determined according to the size of the prediction block, and then...
[0346] mvLX = round(mvpLX) + (mvdLX < <shiftS)
[0347] Determine mvLX.
[0348] <Differential Vector Inverse Quantization Processing>
[0349] The following is for reference Figure 28 as well as Figure 29 The inverse quantization processing of the difference vector in this embodiment will be explained.
[0350] Unless otherwise specified, the processing described below is performed by the inter-frame prediction parameter decoding control unit 3031.
[0351] (Example 1 of inverse quantization: Inverse quantization of the difference vector to the precision of the motion vector corresponding to the completed quantization)
[0352] The following describes the nonlinear inverse quantization processing performed by the inter-frame prediction parameter decoding control unit 3031 on the quantized differential vector qmvd (quantized value, quantized differential vector).
[0353] It should be noted that the difference vector after quantization is equivalent to the absolute value of the difference vector, mvdAbsVal, at the time point obtained by decoding the encoded data (before inverse quantization), and the absolute value of qmvd is mvdAbsVal. It should also be noted that in... Figure 28 In the example shown, regardless of whether the difference vector is negative or positive, the quantization value qmvd of the difference vector is set to a value that is valid for both positive and negative values in order to clarify the image. On the other hand, in actual processing, the absolute value of the difference vector can be used as qmvd, i.e., it can be set as qmvd = mvdAbsVal. In the following explanation, qmvd is treated as an absolute value.
[0354] Figure 28 This is a graph showing the relationship between the quantized differential vector and the inverse quantized differential vector in this processing example. Figure 28 The horizontal axis of the graph shown represents the quantized differential vector, i.e., qmvd (the value obtained by decoding the overquantized and encoded differential vectors without inverse quantization, i.e., the quantized value of the differential vector). Figure 28 The vertical axis of the graph shown represents the inverse-quantized differential vector (also simply called the inverse-quantized differential vector) mvd (= inverse-quantized mvdAbsVal). The inter-frame prediction parameter decoding control unit 3031 is as follows... Figure 28 As shown in the graph, the quantized differential vector qmvd is inversely quantized.
[0355] mvdAbsVal=mvdAbsVal(=qmvd)< <shiftS
[0356] Subsequently, the addition unit 3035 derives the motion vector by adding or subtracting the inversely quantized difference vector from the predicted vector. For example, by...
[0357] mvdLX=mvdAbsVal*(1-2*mv_sign_flag)
[0358] mvLX = round(mvpLX) + mvdLX
[0359] And the derivation.
[0360] right Figure 28 The graph shown is explained in detail. For example... Figure 28As shown, the inter-frame prediction parameter decoding control unit 3031 switches the precision of the inverse quantization process of the quantized motion vector according to the magnitude relationship between the quantized motion vector difference decoded from the encoded data and a specified value (dTH).
[0361] For example, when the absolute value of the differential vector mvd is small, the precision of the motion vector is set to be high, and when the absolute value of the differential vector mvd is large, the precision of the motion vector is set to be low.
[0362] In other words, when the differential vector mvd is near zero (when the absolute value of the differential vector mvd is small), compared with the case where the absolute value of the differential vector mvd is far from near zero, the change in the inverse quantization differential vector mvd is smaller with respect to the change in the quantized differential vector qmvd.
[0363] When the differential vector mvd is far from near zero (when the absolute value of the motion vector difference is large), compared with the case where the differential vector mvd is near zero, the change in the inverse quantization differential vector is larger with respect to the change in the quantized differential vector qmvd.
[0364] This can be achieved by the following configuration. That is, when the absolute value of the quantized differential vector, i.e., qmvd, is less than a specified value (threshold) dTH (or less), the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization of the quantized differential vector qmvd specified by a specified slope (scale factor). And when the quantized differential vector qmvd is greater than or equal to the specified value dTH, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization specified by a slope obtained by shifting the specified slope to the left by the motion vector scale shiftS. Here, the above-mentioned specified slope can be, for example, 1.
[0365] In summary, it is as follows (derivation process Q2). That is, the inter-frame prediction parameter decoding control unit 3031 does not perform inverse quantization when the absolute value of the quantized differential vector is small (qmvd < dTH).
[0366] Or only performs basic inverse quantization implemented by multiplication by K (or left shift according to log2(K)), and
[0367] mvdAbsVal = K * qmvd ··· Equation Q1
[0368] Derives mvdAbsVal.
[0369] In addition, when the absolute value of the quantized differential vector is large (satisfying qmvd >= dTH), in addition to basic inverse quantization, additional inverse quantization is further performed by a prescribed inverse quantization scale shiftS, and
[0370] mvdAbsVal = K * (dTH + (qmvd - dTH) << shiftS) ··· Equation Q2
[0371] derive mvdAbsVal.
[0372] It should be noted that dTH appears in the equation because the values in Equation Q1 and Equation Q2 are connected in such a way that they become equal when qmvd = dTH. When paying attention to the coefficient (slope) of qmvd, note that it becomes K * 1 << shiftS, that is, note that it is better to make the inverse quantization scale large by shiftS.
[0373] Here, mvdAbsVal is the absolute value of the inverse quantized differential vector, and K represents a prescribed proportionality coefficient. As described above, K can be set to 1, or it can be not set as such. The product based on K can be achieved by left shift according to log2(K). It should be noted that when K = 1, only when the quantized differential vector qmvd >= dTH is satisfied, inverse quantization based on shiftS is performed, and when qmvd is small, inverse quantization based on shiftS is not performed.
[0374] In addition, specifically, the following configuration can be adopted: in the above-mentioned Equation Q1 and Q2, the basic vector accuracy is set to 1 / 8 pel (mvBaseAccu = 3), shiftS = 1, and dTH = 16. In this case, when the quantized differential vector qmvd is 16 or more (equivalent to the motion vector accuracy being 2 pel or more), qmvd is left-shifted by shiftS = 1 to perform inverse quantization, and the motion vector accuracy is set to 1 / 4 pel. That is, when qmvd is large, a configuration of setting the motion vector accuracy to a lower value can be adopted.
[0375] In addition, in other examples, the following configuration can also be adopted: in the above arithmetic expressions Q1 and Q2, the basic vector accuracy is set to 1 / 16 pel (mvBaseAccu = 4), shiftS = 1, and dTH = 16. In this case, when the quantized differential vector qmvd is 16 or more (corresponding to a motion vector accuracy of 1 pel or more), qmvd is left-shifted by shiftS = 1 for inverse quantization, and the motion vector accuracy is set to 1 / 8 pel. That is, when qmvd is large, a configuration in which the motion vector accuracy is set to a lower value can be adopted.
[0376] It should be noted that if the arithmetic expressions Q1 and Q2 are represented by one arithmetic expression, the inter-frame prediction parameter decoding control unit 3031 can also be said to be through
[0377] mvdAbsVal = min(qmvd, dTH) + max(0, (qmvd - dTH) << shiftS) ··· Arithmetic expression Q3
[0378] to derive the composition of mvdAbsVal as the absolute value of the differential vector.
[0379] In addition, as another configuration, when the quantized differential vector qmvd is less than (or equal to) the threshold value dTH, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization specific to the slope obtained through
[0380] 1 << shiftS1
[0381] In addition, when the quantized differential vector qmvd is greater than or equal to (or greater than) the threshold value dTH, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization specific to the slope obtained through
[0382] 1 << shiftS2
[0383] Here, shiftS1 and shiftS2 may be equal to each other or not.
[0384] According to the above configuration, the accuracy of the inverse quantization process for the differential vector is switched according to the value of the quantized differential vector. Therefore, a prediction image can be generated using a motion vector with a more appropriate accuracy. In addition, a reduction in the code amount of the differential vector can be achieved, thereby improving the coding efficiency. (Inverse quantization process example 2A: An example of inverse quantizing the differential vector with a motion vector accuracy corresponding to the motion vector accuracy flag and the quantized differential vector)
[0385] Next, an example of inverse quantization of the differential vector with the motion vector precision corresponding to the motion vector precision flag, i.e., mvd_dequant_flag, and the differential vector after quantization is described (derivation process PS_P2A).
[0386] In this processing example, when the motion vector precision flag satisfies mvd_dequant_flag = 1, by
[0387] mvdAbsVal = qmvd << shiftA ··· Equation Q4
[0388] mvdAbsVal is derived.
[0389] On the other hand, when the motion vector precision flag satisfies mvd_dequant_flag = 0 and the quantized differential vector qmvd < the specified value dTHS, by
[0390] mvdAbsVal = qmvd ··· Equation Q5
[0391] mvdAbsVal is derived.
[0392] In addition, when the motion vector precision flag satisfies mvd_dequant_flag = 0 and the quantized differential vector qmvd >= the specified value dTHS, by
[0393] mvdAbsVal = dTHS + (qmvd - dTHS) << shiftS ··· Equation Q6
[0394] mvdAbsVal is derived.
[0395] That is, non-linear inverse quantization is performed when the motion vector precision flag mvd_dequant_flag == 0, and linear quantization is performed when the motion vector precision flag mvd_dequant_flag == 1.
[0396] In summary, the following equation is obtained.
[0397] mvdAbsVal = mvd_quant_flag == 1?
[0398] qmvd << shiftA :
[0400] qmvd < dTHS? qmvd : dTHS + (qmvd - dTHS) << shiftS
[0401] In other words, when the flag indicating the precision of the motion vector is displayed at a first value (mvd_dequant_flag == 0), the inter-frame prediction parameter decoding control unit 3031 switches the precision of the inverse quantization process for the differential vector according to the value (quantization value) of the quantized differential vector. When the flag indicating the precision of the motion vector is displayed at a second value (mvd_dequant_flag == 1), the inverse quantization process for the differential vector is performed with a fixed precision regardless of the quantization value of the quantized differential vector.
[0402] For example, in the above formulas Q4 to Q6, when the basic vector precision is set to 1 / 8 pel (mvBaseAccu = 3) and shiftA = 3, shiftS = 1, and dTHS = 16, when the motion vector precision flag mvd_dequant_flag = 1 is satisfied, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization by shifting the quantized motion vector qmvd (absolute value of the differential motion vector) left by shiftA (= 3 bits), regardless of qmvd. That is, by setting the motion vector precision to full pixels, it is fixed to be lower than when the motion vector precision is set to mvd_dequant_flag = 0. On the other hand, when the motion vector precision flag mvd_dequant_flag = 0 is satisfied, and qmvd is a specified threshold value of 16 (equivalent to 2 pel) or higher, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization by shifting the quantized motion vector qmvd left by shiftS (= 1 bit). That is, the motion vector precision is set to 1 / 4pel, which is lower than the case where qmvd is less than the specified threshold of 16.
[0403] Furthermore, in other examples, when the basic vector precision is set to 1 / 16 pel (mvBaseAccu = 4) and shiftA = 4, shiftS = 2, and dTHS = 16 in the above formulas Q4 to Q6, when the motion vector precision flag mvd_dequant_flag = 1, the inter-frame prediction parameter decoding control unit 3031 shifts the quantized motion vector qmvd left by shiftA (= 4 bits) to perform inverse quantization on the motion vector, regardless of qmvd. That is, the motion vector precision is set to full pixels, and the motion vector precision is set to a lower value. On the other hand, when the motion vector precision flag mvd_dequant_flag = 0, and the specified threshold qmvd is 16 (equivalent to 1 pel) or higher, the quantized motion vector qmvd is shifted left by shiftS (= 2 bits) to perform inverse quantization on the motion vector. That is, the motion vector precision is set to 1 / 4 pel, and the motion vector precision is set to a lower value.
[0404] It should be noted that in the above-described configurations of Q4 to Q6, the specified threshold value and the values of the inverse quantization scales (shiftS, shiftA) may not be limited to the above examples and other values may be used.
[0405] According to the above configuration, a predicted image can be generated using a more appropriate accuracy of the motion vector. Therefore, the prediction accuracy is improved, and thus the coding efficiency is improved.
[0406] (Inverse quantization processing example 2B: Another example of inverse quantizing the differential vector with the motion vector accuracy corresponding to the motion vector accuracy flag and the quantized differential vector)
[0407] Next, another example of inverse quantizing the differential vector with the motion vector accuracy corresponding to the motion vector accuracy flag and the quantized differential vector (derivation process Q2B) will be described.
[0408] In this processing example, when the motion vector accuracy flag is mvd_dequant_flag = 1 and the quantized differential vector qmvd < the specified value dTHA, by
[0409] mvdAbsVal = qmvd << shiftA1 ··· Equation Q7
[0410] mvdAbsVal is derived.
[0411] In addition, when the motion vector accuracy flag is mvd_dequant_flag = 1 and the quantized differential vector qmvd >= the specified value dTHA, by
[0412] mvdAbsVal = dTHA << shiftA1 + (qmvd - dTHA) << shiftA2 ··· Equation Q8
[0413] mvdAbsVal is derived.
[0414] On the other hand, when the motion vector accuracy flag is mvd_dequant_flag = 0 and the quantized differential vector qmvd < the specified value dTHS, by
[0415] mvdAbsVal = qmvd ··· Equation Q9
[0416] mvdAbsVal is derived.
[0417] In addition, when the motion vector accuracy flag is mvd_dequant_flag = 0 and the quantized differential vector qmvd >= dTHS, by
[0418] mvdAbsVal = dTHS + (qmvd - dTHS) << shiftS ··· Equation Q10
[0419] Derive mvdAbsVal.
[0420] That is, regardless of whether the motion vector precision flag mvd_dequant_flag == 0 or the motion vector precision flag mvd_dequant_flag == 1, non - linear inverse quantization of the quantized differential vector is performed.
[0421] In summary, it becomes the following formula.
[0422] mvdAbsVal = mvd_quant_flag == 1?
[0423] qmvd < dTHA? qmvd << shiftA1 : dTHA << shiftA1 + (qmvd - dTHA) << shiftA2 :
[0425] qmvd < dTHS? qmvd : dTHS + (qmvd - dTHS) << shiftS
[0426] In other words, when the flag indicating the precision of the motion vector indicates the first value (the case of mvd_dequant_flag == 0), the inter - frame prediction parameter decoding control unit 3031 switches whether to set the precision of the inverse quantization process for the differential vector to the first precision or the second precision according to the quantization value of the quantized differential vector (the value qmvd before inverse quantization). When the flag indicating the precision of the motion vector indicates the second value (the case of mvd_dequant_flag == 1), it switches whether to set the precision of the inverse quantization process for the differential vector to the third precision or the fourth precision according to the quantization value of the quantized differential vector. At least one of the above - mentioned first precision and second precision is higher than the above - mentioned third precision and fourth precision.
[0427] For example, in the above formulas Q7 to Q10, if the basic vector precision is set to 1 / 8 pel and shiftA1 = 2, shiftA2 = 3, dTHA = 4, shiftS = 2, and dTHS = 16, and the motion vector precision flag mvd_dequant_flag = 1 and the quantized differential vector qmvd is less than 4, the inter-frame prediction parameter decoding control unit 3031 shifts the quantized differential vector qmvd (absolute value of the differential motion vector) to the left by shiftA1 = 2 to perform inverse quantization on the motion vector. That is, the motion vector precision is set to 1 / 2 pel, which is a lower setting.
[0428] Furthermore, when the motion vector precision flag mvd_dequant_flag = 1 and the quantized differential vector qmvd is dTHA = 4 or higher, the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization on the motion vector by shifting the quantized differential vector qmvd (absolute value of the differential motion vector) to the left by shiftA2 = 3. That is, the motion vector precision is set to 1pel, and the motion vector precision is set to a lower value.
[0429] On the other hand, when the motion vector precision flag mvd_dequant_flag = 0 and the quantized differential vector qmvd is less than 16, the inter-frame prediction parameter decoding control unit 3031 sets the motion vector precision to 1 / 8 of the basic vector precision.
[0430] Furthermore, when the motion vector precision flag mvd_dequant_flag = 0 and the quantized differential vector qmvd is dTHS = 16 or higher, the inter-frame prediction parameter decoding control unit 3031 shifts the quantized differential vector qmvd left by shiftS = 2 to perform inverse quantization on the motion vector. That is, the motion vector precision is set to 1 / 2pel, and the motion vector precision is set to a lower value.
[0431] Based on the above configuration, motion vectors with more appropriate precision can be used to generate predicted images. Therefore, prediction accuracy is improved, thereby increasing coding efficiency.
[0432] (Example 3 of inverse quantization: inverse quantization of the difference vector corresponding to the completed quantization and iterative processing of the prediction vector)
[0433] Next, an example of cyclic processing of the prediction vector in the case of inverse quantization of the difference vector based on the completed quantization difference vector will be explained.
[0434] In this processing example, when the inter-frame prediction parameter decoding control unit 3031 performs inverse quantization processing on the differential vector with lower precision, it derives the motion vector by adding or subtracting the inverse-quantized differential vector from the prediction vector that has undergone a loop process by the vector candidate selection unit 3034.
[0435] For example, it can be set according to Equation Q3 described in the above (inverse quantizing the differential vector with the motion vector precision corresponding to the quantized differential vector) as
[0436] mvdAbsVal = qmvd + (qmvd - dTH) << shiftS ··· Equation Q20.
[0437] When the quantized differential vector, i.e., qmvd, is greater than or equal to the specified value dTH, the motion vector mvLX is derived as the sum of the prediction vector mvpLX that has undergone a loop process and the differential vector mvdLX. That is, it is derived as
[0438] mvLX = round(mvpLX, shiftS) + mvdLX. Here, through round(mvpLX, shiftS), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftS units. Here, the process of deriving mvdLX from mvdAbsVal is as described in the processing of the code assignment PS_SIGN.
[0439] On the other hand, when the quantized differential vector, i.e., qmvd, is less than the specified value dTH, the motion vector mvLX is derived as the sum of the prediction vector mvpLX and the differential vector mvdLX. That is, it is derived as
[0440] mvLX = mvpLX + mvdLX.
[0441] (Inverse Quantization Processing Example 4: Inverse Quantization of the Differential Vector Corresponding to the Motion Vector Precision Flag and the Value of the Quantized Differential Vector and Loop Processing of the Prediction Vector)
[0442] In this processing example, the inter-frame prediction parameter decoding control unit 3031 derives the motion vector by adding or subtracting the inverse-quantized differential vector from the prediction vector that has undergone a loop process by the vector candidate selection unit 3034 according to the motion vector precision flag and the quantized differential vector.
[0443] For example, when the motion vector precision flag is mvd_dequant_flag = 1, by
[0444] mvdAbsVal = qmvd << shiftA
[0445] Derive mvdAbs. Then, the motion vector mvLX is obtained by
[0446] mvLX = round(mvpLX, shiftA) + mvdLX
[0447] where the derivation is the sum of the predicted vector mvpLX and the differential vector mvdLX that have undergone loop processing. Here, through round(mvpLX, shiftA), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftA units. Here, the process of deriving mvdLX from mvdAbsVal is as described for the code assignment process PS_SIGN.
[0448] On the other hand, the case where the motion vector precision flag is mvd_dequant_flag = 0 is as follows. When the quantized differential vector qmvd is less than the specified value dTH, the absolute value of the differential vector mvdAbsVal is derived to be equal to the quantized differential vector qmvd. That is, it is derived by mvdAbsVal = qmvd. Then, the motion vector mvLX becomes the sum of the predicted vector mvpLX and the differential vector mvdLX. That is, it is obtained by
[0449] mvLX == mvpLX + mvdLX
[0450] Moreover, when the quantized differential vector qmvd is greater than or equal to the specified value dTH, the absolute value of the differential vector mvdAbsVal is obtained by
[0451] mvdAbsVal = dTHS + (qmvd - dTHS) << shiftS
[0452] and then derived. Then, the motion vector mvLX becomes the sum of the predicted vector mvpLX that has undergone loop processing and the differential vector mvdLX. That is, it is obtained by
[0453] mvLX = round(mvpLX, shiftS) + mvdLX
[0454] Here, through round(mvpLX, shiftS), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftS units. Here, the process of deriving mvdLX from mvdAbsVal is as described for the code assignment process PS_SIGN.
[0455] (Inverse quantization process example 5: Another example of inverse quantization of the differential vector corresponding to the motion vector precision flag and the quantized differential vector and loop processing of the predicted vector)
[0456] Next, other examples of the inverse quantization of the differential vector corresponding to the motion vector precision flag and the completed quantization of the differential vector and the loop processing of the prediction vector will be described.
[0457] In this processing example, when the inter-frame prediction parameter decoding control unit 3031 performs the inverse quantization processing for the differential vector with a precision other than the highest precision among the first precision, the second precision, the third precision, and the fourth precision, the addition unit 3035 derives the motion vector by adding or subtracting the inverse-quantized differential vector to or from the prediction vector that has undergone the loop processing by the vector candidate selection unit 3034.
[0458] A detailed description will be given of an example of this embodiment.
[0459] When the motion vector precision flag is mvd_dequant_flag = 1 and the completed quantization differential vector satisfies less than the specified value dTHA, by
[0460] mvdAbsVal = qmvd << shiftA1
[0461] Derive mvdAbsVal.
[0462] Then, the motion vector mvLX is derived as the sum of the prediction vector mvpLX that has undergone the loop processing and the differential vector mvdLX. That is, it is derived as
[0463] mvLX = round(mvpLX, shiftA1) + mvdLX. Here, by round(mvpLX, shiftA1), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftA1 units.
[0464] In addition, when the motion vector precision flag is mvd_dequant_flag = 1 and the completed quantization differential vector satisfies the specified value dTHA or more, by
[0465] mvdAbsVal = specified value dTHA << shiftA1 + (qmvd - dTHA) << shiftA2
[0466] Derive mvdAbsVal.
[0467] Then, the motion vector mvLX is derived as the sum of the prediction vector mvpLX that has undergone the loop processing and the differential vector mvdLX. That is, it is derived as
[0468] mvLX = round(mvpLX, shiftA2) + mvdLX. Here, by round(mvpLX, shiftA2), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftA2 units. Here, the process of deriving mvdLX from mvdAbsVal is as described for the code assignment process PS_SIGN.
[0469] On the other hand, the case where the motion vector precision flag is mvd_dequant_flag = 0 is as follows. When the quantized differential vector qmvd is less than the specified value dTH, the absolute value mvdAbsVal of the differential vector is derived to be equal to the quantized differential vector qmvd. That is, it is derived by mvdAbsVal = qmvd. Then, the motion vector mvLX becomes the sum of the prediction vector mvpLX and the differential vector mvdLX. That is, by
[0470] mvLX = mvpLX + mvdLX
[0471] And it is derived. In addition, when the quantized differential vector qmvd is greater than or equal to the specified value dTH, the absolute value mvdAbsVal of the differential vector is by
[0472] mvdAbsVal = dTHS + (qmvd - dTHS) << shiftS
[0473] And it is derived. Then, the motion vector mvLX becomes the sum of the prediction vector mvpLX that has undergone the rounding process and the differential vector mvdLX. That is, by
[0474] mvLX = round(mvpLX, shiftS) + mvdLX
[0475] And it is derived. Here, by round(mvpLX, shiftS), the motion vector precision of mvpLX is reduced to the precision of 1 << shiftS units. Here, the process of deriving mvdLX from mvdAbsVal is as described for the code assignment process PS_SIGN.
[0476] (Flowchart of the motion vector scale derivation process using the quantized differential vector)
[0477] Figure 29 More specifically represents the motion vector scale derivation process in S3031 (refer to Figure 19 ) and S3041 (refer to Figure 20 ). In Figure 29 , for the convenience of explanation, the process of S3041 is specifically exemplified, but it can also be applied to S3031 Figure 22The processing shown.
[0478] like Figure 29 As shown, in S304131, it is determined whether the quantized differential vector qmvd < the specified value dTH is satisfied. If the quantized differential vector qmvd < the specified value dTH is false (N in S304131), shiftS is set to M in S304132. If the quantized differential vector qmvd < the specified value dTH is true (Y in S304131), shiftS is set to 0 in S304133. Then, proceed to S3042.
[0479] Motion Compensation Filtering
[0480] The following is for reference Figures 30-32 The motion compensation filter of the motion compensation unit 3091 will be explained.
[0481] Figure 30 This is a block diagram showing the specific structure of the motion compensation unit 3091. For example... Figure 30 As shown, the motion compensation unit 3091 includes a motion vector application unit 30911, a motion compensation filtering unit (filtering unit) 30912, and a filtering coefficient memory 30913.
[0482] The motion vector application unit 30911, based on the prediction list input from the inter-frame prediction parameter decoding unit 303, uses the flag predFlagLX, the reference image index refIdxLX, and the motion vector mvLX to read from the reference image memory 306 the block located at a position offset by the motion vector mvLX from the position of the decoded object block of the reference image specified by the reference image index refIdxLX as the starting point, thereby generating an image after applying the motion vector.
[0483] When the motion vector mvLX is not integer precision but 1 / M pixel precision (M is a natural number greater than 2), the image after applying the motion vector will also have 1 / M pixel precision.
[0484] When the motion vector mvLX is not of integer precision, the motion compensation filter unit 30912 generates the aforementioned motion compensation image (predSamplesL0 in the case of L0 predicted motion compensation image; predSamplesL1 in the case of L1 predicted motion compensation image; and predSamplesLX when the two are not distinguished).
[0485] When the motion vector mvLX has integer precision, the motion compensation filter unit 30912 does not affect the image after the motion vector has been applied, and the image after the motion vector has been applied directly becomes the motion compensation image.
[0486] The filter coefficient memory 30913 stores motion compensation filter coefficients decoded from the encoded data. More specifically, the filter coefficient memory 30913 stores at least a portion of the filter coefficients associated with i from the filter coefficients mcFilter[i][k] (where i is an integer greater than or equal to M-1 and k is an integer greater than or equal to Ntaps-1) used by the motion compensation filter 30912.
[0487] (Filtering coefficients)
[0488] Here, use Figure 31 The details of the filter coefficients mcFilter[i][k] are explained. Figure 31 This is a diagram showing an example of the filtering coefficients in this embodiment.
[0489] Figure 31 The filter coefficients mcFilter[i][k] shown illustrate the case where the total number of phases (i = 0-15) of the image after applying the motion vectors is 16, and the number of filter taps is 8 (8 taps (k = 0-7)). In this example, the total number of phases (i = 0-15) is 16. With a total of 16 phases, the precision of the motion vectors is 1 / 16 pixel. That is, with a total of M phases, the precision of the motion vectors is 1 / M pixel.
[0490] For example, shown in Figure 31 The topmost filter coefficients {0, 0, 0, 64, 0, 0, 0, 0} represent the filter coefficients at each position when phase i = 0. Here, coefficient position refers to the relative position of the pixel that makes the filter coefficient effective. Similarly, shown in Figure 31 The filtering coefficients of the other layers are the filtering coefficients for each coefficient position of other phases (i = 1-15).
[0491] The total number of filter coefficients is the value obtained by multiplying the number of filter taps by the number of phases (i.e., the reciprocal of the precision of the motion vector).
[0492] (Calculation of filter coefficients)
[0493] The filter coefficients mcFilter[i][k] used by the motion compensation filter unit (filter unit) 30912 may also include filter coefficients calculated using filter coefficients mcFilter[p][k] (p≠i) and filter coefficients mcFilter[q][k] (q≠i). Details of the calculation example for filter coefficients mcFilter[i][k] are described below.
[0494] (Calculation Example 1: Calculate the filter coefficient of phase i from the average of the filter coefficients of phase i-1 and i+1)
[0495] use Figure 32 (a) and Figure 32 (b) An example of calculating the filter coefficients in this embodiment is explained.
[0496] Figure 32 (a) shows an example of the motion compensation filter unit 30912 calculating the filter coefficients for other phases (odd phases in this case) from the filter coefficients of a portion (even-numbered) of the phases, and using the calculated filter coefficients. Figure 32 In (a), the filter coefficients for even-numbered phases are indicated by underlining.
[0497] exist Figure 32 In the example shown in (a), the filter coefficient memory 30913 stores the filter coefficients for even-numbered phases. Then, the filter coefficients for odd-numbered phase i are calculated by averaging the filter coefficients for even-numbered phases i-1 and i+1. That is, when i%2 = 1 (i divided by 2 leaves a remainder of 1), it becomes mcFilter[i][k] = (mcFilter[i-1][k] + mcFilter[i+1][k]) / 2. Furthermore, in even-numbered phase i, mcFilter[i][k] stored in the filter coefficient memory 30913 is used as the filter coefficient. That is, when i%2 = 0 (i divided by 2 leaves a remainder of 0), it becomes mcFilter[i][k] = mcFilter[i][k].
[0498] In addition, when i%2=1, it can also be set as mcFilter[i][k]=(mcFilter[i-1][k]+mcFilter[i+1][k])>>1.
[0499] Alternatively, the filter coefficients of a portion of the phases of odd numbers can be stored in the filter coefficient memory 30913, and the motion compensation filter unit 30912 can use the stored filter coefficients as the composition of the filter coefficients.
[0500] It should be noted that the above is equivalent to the basic filter coefficient mcFilterC being stored in the filter coefficient memory 30913, and the composition of the filter coefficient mcFitler actually used being derived by the following formula.
[0501] mcFilter[i][k]=mcFilterC[i>>1][k](i=0, 2, 4,..., 2n, 2n+1, n=7)
[0502] mcFilter[i][k]=(mcFilterC[i>>1][k]+mcFilterC[(i>>1)+1][k]) / 2(i=1, 3, 5,..., 2n+1, n=6)
[0503] Here, the division represented by / 2 can be set as >>1.
[0504] For example, in Figure 32 In the example shown in (a), the following table can be used as mcFilterC.
[0505] mcFilterC[][] =
[0506] {
[0507] {0, 0, 0, 64, 0, 0, 0, 0},
[0508] {-1, 2, -5, 62, 8, -3, 1, 0}
[0509] {-1, 4, -10, 58, 17, -5, 1, 0}
[0510] {-1, 3, -9, 47, 31, -10, 4, -1}
[0511] {-1, 4, -11, 40, 40, -11, 4, -1},
[0512] {-1, 4, -10, 31, 47, -9, 3, -1}
[0513] {0, 1, -5, 17, 58, -10, 4, -1}
[0514] {0, 1, -3, 8, 62, -5, 2, -1}
[0515] {0, 1, -2, 4, 63, -3, 1, 0}
[0516] }
[0517] on the other hand, Figure 32(b) shows an example of how the motion compensation filter unit 30912 calculates the filter coefficients for other phases (even phases) from the filter coefficients of a portion (in this case, odd-numbered phases). Figure 32 In (b), the filter coefficients for odd-phase numbers are indicated by underlining.
[0518] exist Figure 32 In example (b), the filter coefficient memory 30913 stores the filter coefficients for odd-numbered phases. Then, the filter coefficients for even-numbered phase i are calculated by averaging the filter coefficients for odd-numbered phases i-1 and i+1. That is, when i%2 = 0 (i divided by 2 leaves a remainder of 0), it becomes mcFilter[i][k] = (mcFilter[i-1][k] + mcFilter[i+1][k]) / 2. Furthermore, in odd-numbered phase i, mcFilter[i][k] stored in the filter coefficient memory 30913 is used. That is, when i%2 = 1 (i divided by 2 leaves a remainder of 1), it becomes mcFilter[i][k] = mcFilter[i][k].
[0519] In addition, when i%2=0, it can also be set as mcFilter[i][k]=(mcFilter[i-1][k]+mcFilter[i+1][k])>>1.
[0520] The above is equivalent to storing the basic filter coefficient mcFilterC in the filter coefficient memory 30913, and deriving the composition of the filter coefficient mcFitler actually used through the following formula.
[0521] mcFilter[i][k]=mcFilterC[i>>1][k](i=0, 1, 3, 5,..., 2n+1, n=7)
[0522] mcFilter[i][k]=(mcFilterC[i>>1][k]+mcFilterC[(i>>1)+1][k]) / 2(i=0, 2, 4, 6,..., 2n+1, n=7)
[0523] Here, it can also be set to / 2>>1.
[0524] For example, in Figure 32 In the example shown in (b), the following table can be used as mcFilterC.
[0525] mcFilterC[][] =
[0526] {
[0527] {0, 0, 0, 64, 0, 0, 0, 0},
[0528] {0, 1, -3, 63, 4, -2, 1, 0}
[0529] {-1, 3, -8, 60, 13, -4, 1, 0}
[0530] {-1, 4, -11, 52, 26, -8, 3, -1},
[0531] {-1, 4, -11, 45, 34, -10, 4, -1},
[0532] {-1, 4, -10, 34, 45, -11, 4, -1},
[0533] {-1, 3, -8, 26, 52, -11, 4, -1}
[0534] {0, 1, -4, 13, 60, -8, 3, -1}
[0535] {0, 1, -2, 4, 63, -3, 1, 0}
[0536] }
[0537] Alternatively, the filter coefficients of a portion of the even-numbered phases can be stored in the filter coefficient memory 30913, and the motion compensation filter unit 30912 can use the stored filter coefficients as the composition of the filter coefficients.
[0538] (Calculation Example 2: Calculate the filter coefficient of phase i by linear interpolation of the filter coefficients of other phases before and after the current phase)
[0539] Next, an example will be explained of how the motion compensation filter unit 30912 calculates the filter coefficient of phase i by linear interpolation of the filter coefficients of other phases before and after it. The motion compensation filter unit 30912 calculates the filter coefficient of phase i using the following formula.
[0540] mcFilter[i][k]=((Nw)*mcFilter[i0][k]+w*mcFilter[i1][k])>>log(N)
[0541] Here, i0 = (i / N)*N, i1 = i0 + N, w = (i%N), and N is an integer greater than or equal to 2.
[0542] That is, among the above filter coefficients Filter[i][k], there are filter coefficients that satisfy mcFilter[i][k]=((Nw)*mcFilter[i0][k]+w*mcFilter[i1][k])>>log2(N), i0=(i / N)*N, i1=i0+N, w=(i%N) and N is an integer greater than 2.
[0543] The above is equivalent to storing the basic filter coefficient mcFilterC in the filter coefficient memory 30913, and deriving the composition of the filter coefficient mcFitler actually used through the following formula.
[0544] The above is equivalent to storing the basic filter coefficient mcFilterC in the filter coefficient memory 30913, and deriving the composition of the filter coefficient mcFitler actually used through the following formula.
[0545] mcFilter[i][k]=mcFilterC[i>>log2(N)][k](i=N*n)
[0546] mcFilter[i][k]=((Nw)*mcFilterC[i>>log2(N)][k]+w*mcFilter[(i>>log2(N))+1][k])>>log2(N)(i!=N*n)
[0547] Alternatively, it can also be composed of the following components.
[0548] mcFilter[i][k]=mcFilterC[i>>log2(N)][k](i=0,N*n+1)
[0549] mcFilter[i][k]=((Nw)*mcFilterC[i>>log2(N)][k]+w*mcFilter[(i>>log2(N))+1][k])>>log2(N)(i!=N*n+1)
[0550] Based on the configuration shown in the calculation example above, it is not necessary to store all motion compensation filter coefficients in the filter coefficient memory 30913. Therefore, the amount of memory used to store filter coefficients can be reduced. Furthermore, since only a portion of the motion compensation filter coefficients can be included in the encoded data, the amount of code in the encoded data is reduced, thereby expecting an improvement in encoding efficiency.
[0551] (Composition of an image encoding device)
[0552] Next, the configuration of the image encoding device 11 in this embodiment will be described. Figure 12This is a block diagram illustrating the configuration of the image encoding apparatus 11 according to this embodiment. The image encoding apparatus 11 is configured to include a prediction image generation unit 101, a subtraction unit 102, a DCT / quantization unit 103, an entropy encoding unit 104, an inverse quantization / inverse DCT unit 105, an addition unit 106, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a prediction parameter encoding unit 111, and a residual storage unit 313 (residual recording unit). The prediction parameter encoding unit 111 is configured to include an inter-frame prediction parameter encoding unit 112 and an intra-frame prediction parameter encoding unit 113.
[0553] The prediction image generation unit 101 generates prediction image blocks P for each viewpoint image of the layer image T input from the outside, by blocks (regions formed by dividing the image). Here, the prediction image generation unit 101 reads reference image blocks from the reference image memory 109 based on prediction parameters input from the prediction parameter encoding unit 111. The prediction parameters input from the prediction parameter encoding unit 111 are, for example, motion vectors or displacement vectors. The prediction image generation unit 101 reads the reference image blocks located at the positions indicated by the motion vectors or displacement vectors predicted with the encoded object block as the starting point. For the read reference image blocks, the prediction image generation unit 101 generates prediction image blocks P using one of a plurality of prediction methods. The prediction image generation unit 101 outputs the generated prediction image blocks P to the subtraction unit 102. It should be noted that the prediction image generation unit 101 operates in the same manner as the prediction image generation unit 308 already described, therefore, details of the generation of prediction image blocks P are omitted.
[0554] In order to select a prediction method, for example, the prediction method that minimizes the error value, the prediction image generation unit 101 selects a prediction method that minimizes the error value. The error value is based on the difference between the signal value of each pixel in the block contained in the image and the signal value of each pixel corresponding to the predicted image block P. The method for selecting the prediction method is not limited to this.
[0555] Several prediction methods exist: intra-frame prediction, motion prediction, and merge prediction. Motion prediction is the prediction between display times, as described above in the inter-frame prediction method. Merge prediction uses a reference picture block that has already been encoded and is identical to a block located within a predetermined range from the encoded object block, along with prediction parameters.
[0556] When intra-frame prediction is selected, the prediction image generation unit 101 outputs the prediction mode IntrapredMode, which represents the intra-frame prediction mode used when generating the prediction image block P, to the prediction parameter encoding unit 111.
[0557] When motion prediction is selected, the prediction image generation unit 101 stores the motion vector mvLX used when generating the prediction image block P in the prediction parameter memory 108 and outputs it to the inter-frame prediction parameter coding unit 112. The motion vector mvLX represents the vector from the position of the target block to the position of the reference image block when generating the prediction image block P. The information representing the motion vector mvLX includes information representing the reference image (e.g., reference image index refIdxLX, image sequence number POC), and may also include information representing the prediction parameters. Furthermore, the prediction image generation unit 101 outputs the prediction mode predMode, representing the inter-frame prediction mode, to the prediction parameter coding unit 111.
[0558] When merge prediction is selected, the prediction image generation unit 101 outputs the merge index merge_idx, representing the selected reference image block, to the inter-frame prediction parameter coding unit 112. Furthermore, the prediction image generation unit 101 outputs the prediction mode predMode, representing the merge prediction mode, to the prediction parameter coding unit 111.
[0559] Furthermore, the predictive image generation unit 101 may also have a configuration that generates motion compensation filter coefficients referenced by the motion compensation unit 3091 of the image decoding device 31.
[0560] Furthermore, the predictive image generation unit 101 may also have a configuration corresponding to the motion vector precision switching described in the image decoding device 31. That is, the predictive image generation unit 101 may also switch the motion vector precision according to the block size and QP, etc. In addition, it may be possible to use a configuration that encodes the motion vector precision flag mvd_dequant_flag referenced when switching the motion vector precision in the image decoding device 31.
[0561] The subtraction unit 102 subtracts the signal value of the predicted image block P input from the predicted image generation unit 101 from the signal value of the block corresponding to the externally input layer image T, pixel by pixel, to generate a residual signal. The subtraction unit 102 outputs the generated residual signal to the DCT / quantization unit 103 and the encoding parameter determination unit 110.
[0562] The DCT / quantization unit 103 performs DCT on the residual signal input from the subtraction unit 102 to calculate the DCT coefficients. The DCT / quantization unit 103 quantizes the calculated DCT coefficients to obtain the quantization coefficients. The DCT / quantization unit 103 outputs the obtained quantization coefficients to the entropy encoding unit 104 and the inverse quantization / inverse DCT unit 105.
[0563] In the entropy coding unit 104, quantization coefficients are input from the DCT / quantization unit 103, and coding parameters are input from the coding parameter determination unit 110. The input coding parameters include, for example, codes such as the reference image index refIdxLX, the prediction vector index mvp_LX_idx, the difference vector mvdLX, the prediction mode predMode, and the merge index merge_idx.
[0564] It should be noted that the entropy encoding unit 104 may also be configured to perform a nonlinear inverse quantization process corresponding to the nonlinear inverse quantization process described in the image decoding device 31 before encoding the difference vector mvdLX, i.e., a nonlinear quantization process for the difference vector.
[0565] The entropy coding unit 104 performs entropy coding on the input quantization coefficients and coding parameters to generate a coding stream Te, and outputs the generated coding stream Te to the outside.
[0566] The inverse quantization / inverse DCT unit 105 performs inverse quantization on the quantization coefficients input from the DCT / quantization unit 103 to obtain the DCT coefficients. The inverse quantization / inverse DCT unit 105 performs inverse DCT on the obtained DCT coefficients to calculate the decoding residual signal. The inverse quantization / inverse DCT unit 105 outputs the calculated decoding residual signal to the adder unit 106.
[0567] The addition unit 106 adds the signal value of the predicted image block P input from the predicted image generation unit 101 to the signal value of the decoded residual signal input from the inverse quantization / inverse DCT unit 105, pixel by pixel, to generate a reference image block. The addition unit 106 stores the generated reference image block in the reference image memory 109.
[0568] The prediction parameter memory 108 stores the prediction parameters generated by the prediction parameter encoding unit 111 in a predetermined location according to the image and block of the encoded object.
[0569] The reference image memory 109 stores the reference image blocks generated by the addition unit 106 in a predetermined location according to the image and block of the encoded object.
[0570] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters are the aforementioned prediction parameters and parameters generated in relation to these prediction parameters that become objects to be encoded. The prediction image generation unit 101 uses these sets of encoding parameters respectively to generate prediction image blocks P.
[0571] The encoding parameter determination unit 110 calculates the amount of information and the cost value indicating the encoding error for each of the multiple sets. The cost value is, for example, the sum of the code amount and the squared error multiplied by a coefficient λ. The code amount is the amount of information in the encoded stream Te obtained by entropy encoding the quantization error and the encoding parameters. The squared error is the sum of the squares of the residual values of the residual signal calculated in the subtraction unit 102 among the pixels. The coefficient λ is a real number greater than a predetermined zero. The encoding parameter determination unit 110 selects the set of encoding parameters whose calculated cost value is the smallest. Therefore, the entropy encoding unit 104 outputs the selected set of encoding parameters as the encoded stream Te to the outside, instead of outputting the set of unselected encoding parameters.
[0572] The prediction parameter encoding unit 111 derives the prediction parameters used when generating the prediction image based on the parameters input from the prediction image generation unit 101, and encodes the derived prediction parameters to generate a set of encoded parameters. The prediction parameter encoding unit 111 outputs the generated set of encoded parameters to the entropy encoding unit 104.
[0573] The prediction parameter encoding unit 111 stores the prediction parameters corresponding to the parameters selected by the encoding parameter determination unit 110 from the set of generated encoding parameters in the prediction parameter memory 108.
[0574] When the prediction mode predMode input from the prediction image generation unit 101 indicates an inter-frame prediction mode, the prediction parameter coding unit 111 activates the inter-frame prediction parameter coding unit 112. When the prediction mode predMode indicates an intra-frame prediction mode, the prediction parameter coding unit 111 activates the intra-frame prediction parameter coding unit 113.
[0575] The inter-frame prediction parameter coding unit 112 derives inter-frame prediction parameters based on the prediction parameters input from the coding parameter determination unit 110. The inter-frame prediction parameter coding unit 112, as a component for deriving the inter-frame prediction parameters, includes the inter-frame prediction parameter decoding unit 303 (see reference 110). Figure 5 The structure of the inter-frame prediction parameter coding unit 112 is the same as that used to derive the inter-frame prediction parameters. The structure of the inter-frame prediction parameter coding unit 112 will be described below.
[0576] The intra-frame prediction parameter coding unit 113 determines the intra-frame prediction mode IntraPredMode shown by the prediction mode predMode input from the coding parameter determination unit 110 as a set of inter-frame prediction parameters. (Structure of the inter-frame prediction parameter coding unit)
[0577] Next, the configuration of the inter-frame prediction parameter coding unit 112 will be described. The inter-frame prediction parameter coding unit 112 is a unit corresponding to the inter-frame prediction parameter decoding unit 303.
[0578] Figure 13 This is a schematic diagram showing the configuration of the inter-frame prediction parameter coding unit 112 in this embodiment.
[0579] The inter-frame prediction parameter coding unit 112 is composed of a merging prediction parameter derivation unit 1121, an AMVP prediction parameter derivation unit 1122, a subtraction unit 1123, and a prediction parameter integration unit 1126.
[0580] The merged prediction parameter derivation unit 1121 has the same features as the merged prediction parameter derivation unit 3036 described above (see reference). Figure 7 The AMVP prediction parameter derivation unit 1122 has the same structure as the AMVP prediction parameter derivation unit 3032 described above (see reference 3032). Figure 8 The same composition.
[0581] In the merging prediction parameter derivation unit 1121, when the prediction mode predMode input from the prediction image generation unit 101 indicates a merging prediction mode, a merging index merge_idx is input from the encoding parameter determination unit 110. The merging index merge_idx is output to the prediction parameter integration unit 1126. The merging prediction parameter derivation unit 1121 reads from the prediction parameter memory 108 the reference image index refIdxLX and motion vector mvLX of the reference block indicated by the merging index merge_idx in the merging candidate. The merging candidate is a reference block located within a predetermined range from the encoding object block that becomes the encoding target (e.g., a reference block connected to the lower left, upper left, and upper right ends of the encoding object block), and is a reference block that has completed the encoding process.
[0582] The AMVP prediction parameter derivation unit 1122 has the same characteristics as the AMVP prediction parameter derivation unit 3032 described above (refer to...). Figure 8 The same composition.
[0583] That is, in the AMVP prediction parameter derivation unit 1122, when the prediction mode predMode input from the prediction image generation unit 101 represents the inter-frame prediction mode, the motion vector mvLX is input from the coding parameter determination unit 110. The AMVP prediction parameter derivation unit 1122 derives the prediction vector mvpLX based on the input motion vector mvLX. The AMVP prediction parameter derivation unit 1122 outputs the derived prediction vector mvpLX to the subtraction unit 1123. It should be noted that the reference image index refIdx and the prediction vector index mvp_LX_idx are output to the prediction parameter integration unit 1126.
[0584] The subtraction unit 1123 subtracts the prediction vector mvpLX input from the AMVP prediction parameter derivation unit 1122 from the motion vector mvLX input from the autoencoder parameter determination unit 110, generating a difference vector mvdLX. The difference vector mvdLX is output to the prediction parameter integration unit 1126.
[0585] When the prediction mode predMode input from the prediction image generation unit 101 indicates a merged prediction mode, the prediction parameter integration unit 1126 outputs the merge index merge_idx input from the coding parameter determination unit 110 to the entropy coding unit 104.
[0586] When the prediction mode predMode input from the prediction image generation unit 101 indicates the inter-frame prediction mode, the prediction parameter integration unit 1126 performs the next processing.
[0587] The prediction parameter integration unit 1126 integrates the reference image index refIdxLX and the prediction vector index mvp_LX_idx input from the encoding parameter determination unit 110, and the difference vector mvdLX input from the subtraction unit 1123. The prediction parameter integration unit 1126 outputs the integrated code to the entropy coding unit 104.
[0588] It should be noted that the inter-frame prediction parameter coding unit 112 may also include an inter-frame prediction parameter coding control unit (not shown). The inter-frame prediction parameter coding control unit decodes the code (syntax elements) related to inter-frame prediction indicated by the entropy coding unit 104, and encodes the code (syntax elements) contained in the encoded data, such as the segmentation mode part_mode, merge flag merge_flag, merge index merge_idx, inter-frame prediction flag inter_pred_idc, reference image index refIdxLX, prediction vector index mvp_LX_idx, and difference vector mvdLX.
[0589] In this case, the inter-frame prediction parameter coding control unit 1031 includes a merge index coding unit (and... Figure 10 (corresponding to the merged index decoding unit 30312), vector candidate index encoding unit (and) Figure 10It is composed of a vector candidate index decoding unit (corresponding to 30313), a segmentation mode coding unit, a merge flag coding unit, an inter-frame prediction flag coding unit, a reference image index coding unit, and a vector difference coding unit. The segmentation mode coding unit, merge flag coding unit, merge index coding unit, inter-frame prediction flag coding unit, reference image index coding unit, vector candidate index coding unit, and vector difference coding unit encode the segmentation mode part_mode, merge flag merge_flag, merge index merge_idx, inter-frame prediction flag inter_pred_idc, reference image index refIdxLX, prediction vector index mvp_LX_idx, and difference vector mvdLX, respectively.
[0590] It should be noted that a portion of the image encoding device 11 and image decoding device 31 described above, such as the entropy decoding unit 301, prediction parameter decoding unit 302, prediction image generation unit 101, DCT / quantization unit 103, entropy encoding unit 104, inverse quantization / inverse DCT unit 105, encoding parameter determination unit 110, prediction parameter encoding unit 111, entropy decoding unit 301, prediction parameter decoding unit 302, prediction image generation unit 308, and inverse quantization / inverse DCT unit 311, can be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the computer system can read and execute the program recorded on the recording medium. It should be noted that the "computer system" mentioned here refers to a computer system built into any one of the image encoding devices 11-11h and image decoding devices 31-31h, employing hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into the computer system. Furthermore, a "computer-readable recording medium" can include: a medium that dynamically stores a program for a short period of time, such as a communication line in the case of transmitting a program via a network such as the Internet or a communication line such as a telephone line; or a medium that stores a program for a fixed period of time, such as volatile memory within a computer system that serves as a server or client in such a case. In addition, the aforementioned program can be a program used to implement the aforementioned functions, or it can be a program that can further combine the aforementioned functions with programs already recorded in the computer system to implement them.
[0591] Furthermore, part or all of the image encoding device 11 and image decoding device 31 in the above embodiments can be implemented as integrated circuits such as LSI (Large Scale Integration). Each functional block of the image encoding device 11 and image decoding device 31 can be individually processorized, or part or all can be integrated for processorization. Moreover, the method of integrated circuit implementation is not limited to LSI; it can also be implemented using dedicated circuits or general-purpose processors. Furthermore, if advancements in semiconductor technology lead to integrated circuit technologies that replace LSI, integrated circuits based on such technologies can also be used.
[0592] The above description, with reference to the accompanying drawings, details one embodiment of the invention. However, the specific configuration is not limited to the above description, and various design changes can be made without departing from the spirit of the invention.
[0593] [Application Example]
[0594] The image encoding device 11 and image decoding device 31 described above can be mounted on various devices for transmitting, receiving, recording, and reproducing moving images. It should be noted that the moving images can be natural moving images captured by a camera or the like, or artificial moving images (including CG and GUI) generated by a computer or the like.
[0595] First, refer to Figure 33 The following describes the situation in which the image encoding device 11 and the image decoding device 31 described above can be used for the transmission and reception of moving images.
[0596] Figure 33 (a) is a block diagram showing the configuration of the transmitting device PROD_A equipped with the image encoding device 11. For example... Figure 33 As shown in (a), the transmitting device PROD_A includes: an encoding unit PROD_A1 that obtains encoded data by encoding a moving image; a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave based on the encoded data obtained by the encoding unit PROD_A1; and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The image encoding device 11 described above is used as the encoding unit PROD_A1.
[0597] The transmitting device PROD_A may further include: a camera PROD_A4 for capturing moving images, a recording medium PROD_A5 for recording moving images, an input terminal PROD_A6 for inputting moving images from an external source, and an image processing unit A7 for generating or processing images, serving as a source of moving images input to the encoding unit PROD_A1. Figure 33(a) illustrates that the transmitting device PROD_A has all of these components, but some can be omitted.
[0598] It should be noted that the recording medium PROD_A5 can be a medium for recording unencoded motion images, or a medium for recording motion images encoded using a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the decoding unit (not shown) that decodes the encoded data read from the recording medium PROD_A5 according to the recording encoding method is located between the recording medium PROD_A5 and the encoding unit PROD_A1.
[0599] Figure 33 (b) is a block diagram showing the configuration of the receiving device PROD_B equipped with the image decoding device 31. For example... Figure 33 As shown in (b), the receiving device PROD_B includes: a receiving unit PROD_B1 for receiving a modulated signal, a demodulation unit PROD_B2 for obtaining coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 for obtaining a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The image decoding device 31 described above is used as the decoding unit PROD_B3.
[0600] The receiving device PROD_B may further include a display PROD_B4 for displaying moving images, a recording medium PROD_B5 for recording moving images, and an output terminal PROD_B6 for outputting moving images to the outside, serving as the supply destination for the moving images output by the decoding unit PROD_B3. Figure 33 (b) illustrates that the receiving device PROD_B has all these components, but some can be omitted.
[0601] It should be noted that the recording medium PROD_B5 can be a medium for recording unencoded motion images, or it can be a medium encoded using a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the encoding unit (not shown) that encodes the motion images obtained from the decoding unit PROD_B3 according to the recording encoding method is located between the decoding unit PROD_B3 and the recording medium PROD_B5.
[0602] It should be noted that the transmission medium for modulated signals can be wireless or wired. Furthermore, the transmission scheme for modulated signals can be broadcast (where the destination does not pre-determine a specific transmission scheme) or communication (where the destination has a pre-determined transmission scheme). That is, the transmission of modulated signals can be achieved through any of the following: wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
[0603] For example, a terrestrial digital broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wireless broadcasting. Similarly, a cable television broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via cable broadcasting.
[0604] Furthermore, servers (workstations, etc.) and clients (TV receivers, personal computers, smartphones, etc.) using internet-based VOD (Video On Demand) services, moving image sharing services, etc., are examples of transmitting devices PROD_A and receiving devices PROD_B that transmit and receive modulated signals via communication (typically, either wireless or wired is used as the transmission medium in a LAN, and wired is used in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Additionally, smartphones also include multi-functional portable telephone terminals.
[0605] It should be noted that, in addition to decoding and displaying the encoded data downloaded from the server, the client of the motion picture sharing service also has the function of encoding motion pictures captured by a camera and uploading them to the server. That is, the client of the motion picture sharing service performs the functions of both the sending device PROD_A and the receiving device PROD_B.
[0606] Next, refer to Figure 34 The following describes the situation where the image encoding device 11 and the image decoding device 31 described above can be used for recording and reproducing moving images.
[0607] Figure 34 (a) is a block diagram showing the configuration of the recording device PROD_C equipped with the image encoding device 11 described above. Figure 34 As shown in (a), the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 into the recording medium PROD_M. The image encoding device 11 described above is used as the encoding unit PROD_C1.
[0608] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the recording device PROD_C, such as HDD (Hard Disk Drive) or SSD (Solid State Drive); it can also be (2) a type of recording medium connected to the recording device PROD_C, such as SD memory card or USB (Universal Serial Bus) flash memory; or it can be (3) a type of recording medium loaded into a drive device (not shown) built into the recording device PROD_C, such as DVD (Digital Versatile Disc) or BD (Blu-ray Disc).
[0609] Furthermore, the recording device PROD_C may also include: a camera PROD_C3 for capturing moving images, an input terminal PROD_C4 for inputting moving images from the outside, a receiving unit PROD_C5 for receiving moving images, and an image processing unit C6 for generating or processing images, serving as a source of moving images input to the encoding unit PROD_C1. Figure 34 (a) illustrates that the recording device PROD_C has all of these components, but some can be omitted.
[0610] It should be noted that the receiving unit PROD_C5 can receive unencoded motion images, as well as encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the transmission decoding unit (not shown) that decodes the encoded data encoded using the transmission encoding method between the receiving unit PROD_C5 and the encoding unit PROD_C1.
[0611] Examples of such recording devices PROD_C include DVD recorders, BD recorders, and HDD (Hard Disk Drive) recorders (in which case the input terminal PROD_C4 or the receiving unit PROD_C5 becomes the main source of moving images). Furthermore, portable camcorders (in which case the camera PROD_C3 becomes the main source of moving images), personal computers (in which case the receiving unit PROD_C5 or the image processing unit C6 becomes the main source of moving images), and smartphones (in which case the camera PROD_C3 or the receiving unit PROD_C5 becomes the main source of moving images) are also examples of such recording devices PROD_C.
[0612] Figure 34(b) is a block diagram showing the configuration of the playback device PROD_D equipped with the aforementioned image decoding device 31. Figure 34 As shown in (b), the playback device PROD_D includes: a readout unit PROD_D1 that reads coded data written to the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the coded data read out by the readout unit PROD_D1. The image decoding device 31 described above is used as the decoding unit PROD_D2.
[0613] It should be noted that the recording medium PROD_M can be (1) a recording medium built into the playback device PROD_D, such as HDD or SSD; (2) a recording medium connected to the playback device PROD_D, such as SD memory card or USB flash drive; or (3) a recording medium loaded into a drive device (not shown) built into the playback device PROD_D, such as DVD or BD.
[0614] Furthermore, the playback device PROD_D may also include a display PROD_D3 for displaying moving images, an output terminal PROD_D4 for outputting moving images to the outside, and a transmission unit PROD_D5 for transmitting moving images, serving as the supply destination for the moving images output by the decoding unit PROD_D2. Figure 34 (b) illustrates that the reproduction device PROD_D has all of these components, but some can be omitted.
[0615] It should be noted that the transmitting unit PROD_D5 can transmit unencoded motion images, or it can transmit encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the encoding unit (not shown) that encodes the motion images using the transmission encoding method between the decoding unit PROD_D2 and the transmitting unit PROD_D5.
[0616] Examples of such playback devices PROD_D include DVD players, BD players, HDD players, etc. (in this case, the output terminal PROD_D4 connected to a TV receiver, etc., becomes the main destination for the moving images). Furthermore, TV receivers (in this case, the display PROD_D3 becomes the main destination for the moving images), digital signage (also called electronic billboards, electronic bulletin boards, etc., where the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images), desktop PCs (in this case, the output terminal PROD_D4 or the transmitter PROD_D5 becomes the main destination for the moving images), laptop or tablet PCs (in this case, the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images), and smartphones (in this case, the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images) are also examples of such playback devices PROD_D.
[0617] (Hardware implementation and software implementation)
[0618] Furthermore, each of the aforementioned image decoding device 31 and image encoding device 11 can be implemented in hardware using logic circuits formed on an integrated circuit (IC chip), or in software using a CPU (Central Processing Unit).
[0619] In the latter case, each of the aforementioned devices includes: a CPU that executes commands for programs that perform various functions; a ROM (Read Only Memory) that stores the programs; a RAM (Random Access Memory) that expands the programs; and a memory that stores the programs and various data, etc., such as a storage device (recording medium). Then, the object of the present invention can also be achieved by supplying a recording medium containing the software that performs the aforementioned functions, i.e., the program code (executable form program, intermediate code program, source program) of the control program of each of the aforementioned devices that is readable by a computer, to each of the aforementioned devices, and the computer (or CPU, MPU) reads the program code recorded on the recording medium and executes it.
[0620] As recording media, examples include magnetic tapes, cassette tapes, and other tape types; disks such as floppy disks (registered trademark) and hard disks; optical discs such as CD-ROMs (Compact Disc Read-Only Memory), MO discs (Magneto-Optical Disc), MD discs (Mini Disc), DVDs (Digital Versatile Disc), CD-R discs (CD Recordable), and Blu-ray discs (registered trademark); cards such as IC cards (including memory cards) and optical cards; semiconductor memory types such as mask ROMs, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable and Programmable Read-Only Memory, registered trademark) and flash ROMs; or logic circuits such as PLDs (Programmable Logic Devices) and FPGAs (Field Programmable Gate Arrays).
[0621] Furthermore, the aforementioned devices can be configured to connect to a communication network, via which the program code can be supplied. The communication network need only be capable of transmitting program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, Virtual Private Network, telephone line network, mobile communication network, satellite communication network, etc., can be used. Moreover, the transmission medium constituting the communication network only needs to be a medium capable of transmitting program code, and is not limited to a specific configuration or type. For example, it can be used in wired networks such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV lines, telephone lines, and ADSL (Asymmetric Digital Subscriber Line) lines, as well as in wireless networks such as IrDA (Infrared Data Association), infrared (like remote controls), Bluetooth (registered trademark), IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance, registered trademark), mobile phone networks, satellite lines, and terrestrial digital networks. It should be noted that this invention can also be implemented as a computer data signal embedded in a carrier wave, which electronically transmits the aforementioned program code.
[0622] This invention is not limited to the embodiments described above, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical solutions with appropriate modifications within the scope of the claims are also included within the technical scope of this invention.
[0623] (Cross-reference to related applications)
[0624] This application claims priority to Japanese Patent Application No. 2016-017444, filed on February 1, 2016, the entire contents of which are incorporated herein by reference.
[0625] Industrial availability
[0626] This invention is preferably applicable to an image decoding apparatus that decodes encoded data obtained by encoding image data, and to an image encoding apparatus that generates encoded data obtained by encoding image data. Furthermore, it is preferably applicable to a data structure of encoded data generated by the image encoding apparatus and referenced by the image decoding apparatus.
[0627] Symbol Explanation
[0628] 11: Image coding device (moving picture coding device)
[0629] 31: Image decoding device (moving image decoding device)
[0630] 302: Prediction parameter decoding unit (prediction image generation device)
[0631] 303: Inter-frame prediction parameter decoding unit (motion vector derivation unit)
[0632] 308: Predictive image generation unit (predictive image generation device)
[0633] 3031: Inter-frame prediction parameter decoding control unit (motion vector derivation unit)
[0634] 30912: Compensation Filtering Unit (Filtering Unit)
Claims
1. A predictive image generation apparatus for generating a predictive image, the predictive image generation apparatus comprising: The inter-frame prediction parameter decoding control circuit derives the modified differential motion vector by using the motion vector difference value, and decodes a flag indicating the accuracy of the motion vector from the encoded data when the motion vector difference value is not equal to zero. as well as The predictive image generation circuit generates a predictive image based on a modified differential motion vector using motion vectors. in, The inter-frame prediction parameter decoding control circuit uses the flag to determine the shift value for the modification process of the motion vector difference value.
2. A video decoding device, comprising: The predictive image generation apparatus according to claim 1, in, The video decoding device decodes the encoded object image by adding the residual image to the predicted image or subtracting the residual image from the predicted image.
3. A video encoding apparatus, comprising: The predictive image generation apparatus according to claim 1, in, The video encoding device encodes the residual between the predicted image and the image to be encoded.
Citation Information
Patent Citations
Compressor
JP2016017444A
Moving picture signal coding method, decoding method, coding apparatus, and decoding apparatus
CN1882099A