Encoding device, decoding device, and non-transitory computer-readable recording medium

The encoding device uses inter-frame prediction and motion compensation to enhance compression efficiency and reduce processing load in video coding, addressing the limitations of existing standards like HEVC.

TWI931798BActive Publication Date: 2026-07-11PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW113129233
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-04-27
Filing Date
2018-04-25
Publication Date
2026-07-11
Estimated Expiration
2038-04-24

AI Technical Summary

Technical Problem

Existing video coding standards like HEVC face challenges in achieving further improvements in compression efficiency and reducing processing load.

Method used

An encoding device utilizing inter-frame prediction with motion compensation and gradient images to derive local motion estimation values for sub-block units, enhancing compression efficiency and reducing processing load.

Benefits of technology

The solution achieves improved compression efficiency and reduced processing load in video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMG-2_DRAW_113129233-A0304-14-0001-1
    Figure IMG-2_DRAW_113129233-A0304-14-0001-1
  • Figure IMG-2_DRAW_113129233-A0304-14-0002-2
    Figure IMG-2_DRAW_113129233-A0304-14-0002-2
  • Figure IMG-2_DRAW_113129233-A0304-14-0003-3
    Figure IMG-2_DRAW_113129233-A0304-14-0003-3
Patent Text Reader

Abstract

An encoding apparatus is provided that uses inter-frame prediction to encode a block of an image. The apparatus includes a processor and memory. The processor uses the memory to perform motion compensation using motion vectors corresponding to each of two reference images, thereby obtaining two prediction images from the two reference images and two gradient images corresponding to the two prediction images. Using sub-block units obtained by segmenting the block of the image to be encoded, local motion estimates are derived using the two prediction images and the two gradient images. Finally, using the two prediction images, the two gradient images, and the local motion estimates of the sub-block units, a final prediction image of the block of the image to be encoded is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Invention Field This disclosure relates to an encoding and decoding method for images that utilizes inter-frame prediction. Prior Technology

[0002] Background of the Invention A video coding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding). Previous technical documents

[0003] Non-patent literature Non-patent document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention

[0004] Invention Summary The problem the invention aims to solve This encoding and decoding technology requires further improvements in compression efficiency and a reduction in processing load.

[0005] Therefore, this disclosure provides an encoding device, decoding device, encoding method, or decoding method that can achieve further improvement in compression efficiency and reduction in processing load. The means to solve the problem

[0006] The encoding apparatus of the present invention is an encoding apparatus that uses inter-frame prediction to encode an encoding object block of an image. It includes a processor and a memory. The processor uses the memory to perform motion compensation using motion vectors corresponding to each of two reference images, thereby obtaining two prediction images from the two reference images and two gradient images corresponding to the two prediction images from the two reference images. Using sub-block units obtained by segmenting the encoding object block, local motion estimation values ​​are derived using the two prediction images and the two gradient images. The final prediction image of the encoding object block is generated using the two prediction images, the two gradient images, and the local motion estimation values ​​of the sub-block units.

[0007] Furthermore, these comprehensive or specific forms can be realized using systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or any combination of systems, methods, integrated circuits, computer programs, and recording media. Invention Effects

[0008] This disclosure may provide an encoding device, decoding device, encoding method, or decoding method that can achieve further improvements in compression efficiency and reduction in processing load. Simple Explanation of the Diagram

[0009] Figure 1 is a block diagram showing the functional configuration of the encoding device in Embodiment 1.

[0010] Figure 2 is an example of block partitioning in implementation form 1.

[0011] Figure 3 is a table showing the transformation basis functions corresponding to each transformation type.

[0012] Figure 4A is an example diagram showing the shape of a filter used in ALF.

[0013] Figure 4B is another example of the shape of a filter used in ALF.

[0014] Figure 4C is another example of the shape of a filter used in ALF.

[0015] Figure 5A shows the 67 in-frame prediction patterns in the display frame prediction.

[0016] Figure 5B is a flowchart illustrating the overview of the predictive image correction process performed by OBMC processing.

[0017] Figure 5C is a conceptual diagram illustrating the general outline of predictive image correction processing performed by OBMC processing.

[0018] Figure 5D is a diagram showing one example of FRUC.

[0019] Figure 6 is a diagram illustrating pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0020] Figure 7 is a diagram illustrating the pattern matching (template matching) between a template in the current image and a block in a reference image.

[0021] Figure 8 is a diagram used to illustrate a model that assumes uniform linear motion.

[0022] Figure 9A is a diagram illustrating the derivation of the motion vector of a sub-block unit based on the motion vectors of a plurality of adjacent blocks.

[0023] Figure 9B is a diagram illustrating the general process of motion vector derivation through the merging mode.

[0024] Figure 9C is a conceptual diagram illustrating the outline of DMVR processing.

[0025] Figure 9D is a schematic diagram illustrating a predictive image generation method that utilizes luminance correction processing performed by LIC processing.

[0026] Figure 10 is a block diagram showing the functional configuration of the decoding device in Embodiment 1.

[0027] Figure 11 is a flowchart showing the inter-frame prediction in implementation configuration 2.

[0028] Figure 12 is a conceptual diagram illustrating inter-frame prediction in implementation form 2.

[0029] Figure 13 is a conceptual diagram illustrating one example of the reference range for the motion compensation filter and gradient filter in embodiment 2.

[0030] Figure 14 is a conceptual diagram illustrating one example of the reference range of the motion compensation filter in Modification 1 of Embodiment 2.

[0031] Figure 15 is a conceptual diagram illustrating one example of the reference range of the gradient filter in Modification 1 of Embodiment 2.

[0032] Figure 16 is an example diagram showing the shape of the pixel referenced in the derivation of the local motion estimation value in Modification 2 of Embodiment 2.

[0033] Figure 17 is an overall diagram of the content delivery system that implements the content delivery service.

[0034] Figure 18 is an example of the encoding construction in adjustable encoding.

[0035] Figure 19 is an example of the encoding construction in adjustable encoding.

[0036] Figure 20 is an example of a webpage display screen.

[0037] Figure 21 is an example of a webpage display screen.

[0038] Figure 22 shows an example of a smartphone.

[0039] Figure 23 is a block diagram showing an example of the structure of a smartphone. Implementation

[0040] Forms used to implement inventions The following describes the specific implementation details with reference to the diagrams.

[0041] Furthermore, the embodiments described below are all examples showing comprehensive or specific implementations. The values, shapes, materials, constituent elements, the arrangement and connection of constituent elements, steps, and the order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claim. Also, among the constituent elements of the following embodiments, those constituent elements not recorded in the independent claim items representing the highest-level concept are described as arbitrary constituent elements. (Implementation Form 1)

[0042] First, an overview of Embodiment 1 is provided as an example of an encoding and decoding apparatus applicable to the processing and / or configuration described in the various embodiments of this disclosure. However, Embodiment 1 is merely one example of an encoding and decoding apparatus applicable to the processing and / or configuration described in the various embodiments of this disclosure, and the processing and / or configuration described in the various embodiments of this disclosure may also be implemented in encoding and decoding apparatuses different from Embodiment 1.

[0043] When applying the processing and / or configuration described in the various embodiments of this disclosure to Embodiment 1, any of the following may also be performed, for example. (1) For the encoding or decoding device of embodiment 1, the constituent elements that correspond to the constituent elements described in each of the various embodiments of the present disclosure among the plurality of constituent elements constituting the encoding or decoding device are replaced with the constituent elements described in each of the various embodiments of the present disclosure. (2) For the encoding or decoding device of embodiment 1, for a portion of the plurality of constituent elements constituting the encoding or decoding device, after any change such as addition, replacement, or deletion of the processing applied or implemented, the constituent element corresponding to the constituent element described in each embodiment of the present disclosure may be replaced with the constituent element described in each embodiment of the present disclosure. (3) For the method implemented by the encoding or decoding apparatus of embodiment 1, after any changes such as adding processing and / or replacing or deleting a portion of the processing included in the method, the processing corresponding to the processing described in each embodiment of the present disclosure can be replaced with the processing described in each embodiment of the present disclosure. (4) A portion of the constituent elements constituting the encoding or decoding device of embodiment 1 may be combined with the following constituent elements to implement the device: constituent elements described in the various embodiments of this disclosure, constituent elements having one part of the functions of the constituent elements described in the various embodiments of this disclosure, or constituent elements that perform one part of the processing performed by the constituent elements described in the various embodiments of this disclosure. (5) A component having a portion of the functions of a component having a portion of the components constituting the encoding or decoding device of embodiment 1, or a component having a portion of the processing performed by a component having a portion of the components constituting the encoding or decoding device of embodiment 1, may be combined with the following components: the components described in the various embodiments of this disclosure, the components having a portion of the functions of the components described in the various embodiments of this disclosure, or the components having a portion of the processing performed by the components described in the various embodiments of this disclosure. (6) For the method implemented by the encoding or decoding apparatus of embodiment 1, the processing corresponding to the processing described in each of the various embodiments of the present disclosure can be replaced with the processing described in each of the various embodiments of the present disclosure among the plurality of processing included in the method. (7) A portion of the processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1 may be combined with the processing described in the various embodiments of this disclosure.

[0044] Furthermore, the implementation methods of the processing and / or configurations described in the various embodiments disclosed herein are not limited to the examples described above. For example, they may also be implemented in devices used for different purposes than the motion picture / image encoding device or motion picture / image decoding device disclosed in Embodiment 1, or the processing and / or configurations described in each embodiment may be implemented individually. Additionally, the processing and / or configurations described in different embodiments may be combined for implementation. [Overview of Encoding Device]

[0045] First, an overview of the encoding device in Embodiment 1 will be described. Figure 1 is a block diagram showing the functional configuration of the encoding device 100 in Embodiment 1. The encoding device 100 is a motion image / image encoding device that encodes motion images / images in block units.

[0046] As shown in Figure 1, the encoding device 100 is a device for encoding images in block units, and includes a segmentation unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an in-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128.

[0047] The encoding device 100 can be implemented using, for example, a general-purpose processor and memory. In this case, when the processor executes the software program stored in memory, the processor functions as a segmentation unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an in-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128. Furthermore, the encoding device 100 can also be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the in-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.

[0048] The following describes each component included in the encoding device 100. [Divider]

[0049] The segmentation unit 102 divides each image contained in the input dynamic image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 can first segment the image into fixed-size blocks (e.g., 128×128). These fixed-size blocks are called coding tree units (CTUs). Furthermore, the segmentation unit 102 can segment each fixed-size block into variable-size blocks (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block segmentation. These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs). Moreover, in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks within the image can be processed as CUs, PUs, or TUs.

[0050] Figure 2 is a diagram showing an example of block partitioning in Implementation 1. In Figure 2, solid lines represent block boundaries formed by quadtree block partitioning, and dashed lines represent block boundaries formed by binary tree block partitioning.

[0051] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into four square blocks of 64×64 (quadtree block partitioning).

[0052] The top-left 64×64 block is further vertically divided into two 32×64 rectangular blocks, and the left 32×64 block is further vertically divided into two 16×64 rectangular blocks (binary tree block partitioning). The result is that the top-left 64×64 block is divided into two 16×64 blocks (11 and 12) and a 32×64 block (13).

[0053] The 64×64 block in the upper right corner is horizontally divided into two rectangular 64×32 blocks, 14 and 15 (binary tree block division).

[0054] The bottom left 64×64 block is divided into four 32×32 square blocks (quadtree block partitioning). Among these four 32×32 blocks, the top left and bottom right blocks are further divided. The top left 32×32 block is vertically divided into two 16×32 rectangular blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block partitioning). The bottom right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block partitioning). As a result, the bottom left 64×64 block can be divided into: 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.

[0055] The 64x64 block 23 in the lower right corner is not divided.

[0056] As shown above, in Figure 2, block 10 is divided into 13 variable-size blocks 11-23 based on the recursive quad-tree and binary tree block partitioning. This partitioning is called QTBT (quad-tree plus binary tree) partitioning.

[0057] Furthermore, although Figure 2 shows the division of one block into four or two blocks (quadtree or binary tree block partitioning), the partitioning is not limited to these. For example, one block can also be divided into three blocks (ternary tree block partitioning). This type of partitioning, which includes ternary tree block partitioning, is called MBT (multi-type tree) partitioning. [Subtraction Section]

[0058] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) using block units divided by the segmentation unit 102. In other words, the subtraction unit 104 calculates the prediction error (also called the residual) of the encoded target block (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0059] The original signal is the input signal of the encoding device 100, and it is a signal representing the images of each picture that constitutes the moving image (e.g., luminance signal and two chroma signals). Hereinafter, the signal representing the image will sometimes be referred to as a sample. [Transition Section]

[0060] The conversion unit 106 converts the prediction error in the spatial region into conversion coefficients in the frequency region and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs a predefined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial region, for example.

[0061] Furthermore, the transformation unit 106 can also appropriately select a transformation type from a plurality of transformation types, and use a transformation basis function corresponding to the selected transformation type to convert the prediction error into transformation coefficients. This transformation is sometimes called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).

[0062] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. In Figure 3, N represents the number of input pixels. The selection of a transformation type from these multiple transformation types can be determined, for example, by the type of prediction (intra-prediction and inter-prediction) or by the intra-prediction mode.

[0063] This information, which indicates whether EMT or AMT is applicable (e.g., an AMT flag) and the selected conversion type, can be signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0064] Furthermore, the transformation unit 106 can also re-transform the transformation coefficients (transformation results). This re-transformation is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 performs re-transformation on each sub-block (e.g., a 4×4 sub-block), where such sub-blocks are sub-blocks contained within blocks of transformation coefficients corresponding to in-box prediction errors. Information indicating whether NSST is applicable, and information related to the transformation matrix used in NSST, is signaled at the CU level. Moreover, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0065] Here, a separable transformation refers to a method of separating the input in each direction and performing multiple transformations in a manner equivalent to the number of dimensions of the input. A non-separable transformation refers to a method of summing up two or more dimensions and treating them as a single dimension when the input is multidimensional, and then performing the transformation on that sum.

[0066] Examples of non-separable transformations include treating a 4×4 block as an array of 16 elements and performing a transformation on that array using a 16×16 transformation matrix.

[0067] Similarly, after treating a 4×4 input block as an array with 16 elements, performing a transformation such as a complex number of Givens rotations on that array (the Hypercube Givens Transform) is also an example of a non-separable transformation. [Quantitative Department]

[0068] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scan order and quantizes the conversion coefficients according to the quantization parameters (QP) corresponding to the scanned conversion coefficients. Furthermore, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0069] The predetermined order is the order in which the conversion coefficients are quantized / dequantized. For example, the predetermined scan order is defined by an ascending frequency order (from low to high frequency) or a descending frequency order (from high to low frequency).

[0070] The quantization parameter defines the quantization step size (quantization width). For example, increasing the value of the quantization parameter will increase the quantization step size. In other words, increasing the value of the quantization parameter will increase the quantization error. [Entropy Coding Department]

[0071] The entropy coding unit 110 generates a coded signal (coded bit stream) by performing variable-length encoding on the quantization coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110 performs arithmetic encoding on the binary signal, for example, by binarizing the quantization coefficients. [De-quantization Department]

[0072] The inverse quantization unit 112 performs inverse quantization on the input, i.e., the quantization coefficients, from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantization coefficients of the current block in a predetermined scan order. Furthermore, the inverse quantization unit 112 outputs the inverse quantization conversion coefficients of the current block to the inverse conversion unit 114. [Reverse Conversion Section]

[0073] The inverse conversion unit 114 restores the prediction error by performing an inverse conversion on the input, i.e., the conversion coefficients, from the inverse quantization unit 112. Specifically, the inverse conversion unit 114 restores the prediction error of the current block by performing an inverse conversion on the conversion coefficients corresponding to the conversion performed by the conversion unit 106. Then, the inverse conversion unit 114 outputs the restored prediction error to the addition unit 116.

[0074] Furthermore, since the prediction error of the restoration loses information due to quantization, it is not consistent with the prediction error calculated by the subtraction unit 104. That is, the prediction error of the restoration includes quantization error. [Addition Department]

[0075] The addition unit 116 performs an addition operation on the input (prediction error) from the inverse conversion unit 114 and the input (prediction sample) from the prediction control unit 128, thereby reconstructing the current block. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. Sometimes the reconstructed block is also referred to as a local decoding block. [Block Memory]

[0076] Block memory 118 is a storage unit for storing blocks referenced in in-frame prediction and also blocks within the encoded object image (hereinafter referred to as the current image). Specifically, block memory 118 stores the reconstructed blocks output from addition unit 116. [Loop Filtering Section]

[0077] The loop filtering unit 120 performs loop filtering on the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. A loop filter is a filter used within the encoding loop (in-loop filter), and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0078] In ALF, a least squares error filter for removing coding distortion can be applied. For example, a filter can be selected from a plurality of filters based on the direction and activity of the local gradient for each of the 2×2 sub-blocks in the current block.

[0079] Specifically, firstly, sub-blocks (e.g., 2×2 sub-blocks) can be classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is based on the direction and activity of the gradient. For example, the classification value C (e.g., C = 5D + A) can be calculated using the gradient direction value D (e.g., 0~2 or 0~4) and the gradient activity value A (e.g., 0~4). Then, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).

[0080] The direction value D of the gradient can be derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantizing the result.

[0081] Based on the results of this classification, the filter to be used for a sub-block can be determined from a plurality of filters.

[0082] For the shape of filters used in ALF, circular symmetrical shapes can be used, for example. Figures 4A-4C show several examples of filter shapes used in ALF. Figure 4A shows a 5×5 diamond-shaped filter, Figure 4B shows a 7×7 diamond-shaped filter, and Figure 4C shows a 9×9 diamond-shaped filter. The information showing the filter shape is signaled at the image layer. Furthermore, the signaling of the filter shape information is not limited to the image layer and can also be at other layers (e.g., sequence layer, fragment layer, tile layer, CTU layer, or CU layer).

[0083] The on / off state of ALF (Alternating Current Image) is determined at, for example, the picture layer or the CU (Chip Array) layer. For instance, the decision to apply ALF for luminance is made at the CU layer, while the decision for chroma is made at the picture layer. Information indicating whether ALF is on or off can be signaled at either the picture layer or the CU layer. Furthermore, the signaling of ALF on / off information is not limited to the picture layer or the CU layer; it can also be at other layers (e.g., sequence layer, fragment layer, tile layer, or CTU layer).

[0084] A set of coefficients for a selectable plurality of filters (up to, for example, 15 or 25 filters) is signaled at the image level. Furthermore, the signaling of the coefficient set is not limited to the image level, but can also be at other levels (e.g., sequence level, segment level, tile level, CTU level, CU level, or sub-block level). [Frame Memory]

[0085] Frame memory 122 is a storage unit used to store reference images used for inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, frame memory 122 stores reconstructed blocks that have been filtered by the loop filter unit 120. [In-frame prediction section]

[0086] The in-frame prediction unit 124 performs in-frame prediction (also known as in-image prediction) of the current block by referring to the blocks within the current image stored in the block memory 118, thereby generating a prediction signal (in-frame prediction signal). Specifically, the in-frame prediction unit 124 performs in-frame prediction by referring to samples (e.g., brightness values, color difference values) of blocks adjacent to the current block, thereby generating an in-frame prediction signal, and outputs the in-frame prediction signal to the prediction control unit 128.

[0087] For example, the in-frame prediction unit 124 uses one of a plurality of predefined in-frame prediction patterns to perform in-frame prediction. The plurality of in-frame prediction patterns includes one or more non-directional prediction patterns and a plurality of directional prediction patterns.

[0088] One or more non-directional prediction modes include, for example, the planar prediction mode and DC prediction mode specified in the H.265 / HEVC (High-Efficiency Video Coding) specification (Non-Patent Document 1).

[0089] Multiple directional prediction modes include, for example, the 33 directional prediction modes specified in the H.265 / HEVC specification. Furthermore, multiple directional prediction modes may further include 32 additional directional prediction modes (a total of 65 directional prediction modes). Figure 5A shows 67 in-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in in-frame prediction. Solid arrows represent the 33 directions specified in the H.265 / HEVC specification, and dashed arrows represent the additional 32 directions.

[0090] Furthermore, in the in-frame prediction of chromatic difference blocks, luminance blocks can also be referenced. That is, the chromatic difference component of the current block can be predicted based on its luminance component. This in-frame prediction is sometimes called CCLM (cross-component linear model) prediction. This in-frame prediction mode for chromatic difference blocks referencing luminance blocks (e.g., called the CCLM mode) can also be added as an in-frame prediction mode for a single chromatic difference block.

[0091] The in-box prediction unit 124 can also correct the predicted pixel values ​​based on the gradient of reference pixels in the horizontal / vertical directions. This in-box prediction with accompanying correction is sometimes referred to as PDPC (position dependent intra prediction combination). Information indicating the applicability of PDPC (such as a PDPC flag) is signaled at, for example, the CU level. Furthermore, the signaling of this information is not limited to the CU level and can also be at other levels (such as sequence level, image level, segment level, tile level, or CTU level). [Inter-frame Prediction Department]

[0092] The inter-frame prediction unit 126 refers to a reference image stored in the frame memory 122, which is different from the current image, to perform inter-frame prediction (also called inter-picture prediction) for the current block, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed on a unit of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 126 performs motion search (motion estimation) within the reference image for the current block or sub-block. Then, the inter-frame prediction unit 126 uses the motion information (e.g., motion vectors) obtained from the motion search to perform motion compensation, thereby generating the inter-frame prediction signal for the current block or sub-block. Then, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.

[0093] Motion information used for motion compensation is signaled. In the signaling of motion vectors, motion vector predictors can also be used. That is, the difference between motion vectors and motion vector predictors can also be signaled.

[0094] Furthermore, motion information from not only the current block obtained through motion search but also that from adjacent blocks can be used to generate inter-frame prediction signals. Specifically, this can be achieved by weighted summing of the prediction signals based on motion information obtained through motion search and the prediction signals based on motion information from adjacent blocks, using sub-block units within the current block as the basis for generating inter-frame prediction signals. This type of inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0095] In this OBMC mode, information displaying the size of sub-blocks used in OBMC (e.g., referred to as OBMC block size) is signaled at the sequence level. Furthermore, information displaying whether OBMC mode is applicable (e.g., referred to as OBMC flag) is signaled at the CU level. Moreover, the signaling level for this information is not limited to the sequence level and CU level; it can also be other levels (e.g., image level, fragment level, tile level, CTU level, or sub-block level).

[0096] To illustrate this more specifically with the OBMC model, Figures 5B and 5C are flowcharts and conceptual diagrams illustrating the general outline of predictive image correction processing performed via OBMC.

[0097] First, motion vectors (MVs) allocated to the encoded object blocks are used to obtain a predicted image (Pred) formed by general motion compensation.

[0098] Next, the motion vector (MV_L) of the encoded left adjacent block is applied to the encoded target block to obtain the prediction image (Pred_L), and the aforementioned prediction image and Pred_L are weighted and overlapped to perform the first correction of the prediction image.

[0099] Similarly, the motion vector (MV_U) of the upper adjacent block after encoding is applied to the encoding target block to obtain the prediction image (Pred_U), and the prediction image that has undergone the first correction is weighted and overlapped with Pred_U to perform the second correction of the prediction image, and then set as the final prediction image.

[0100] Furthermore, although the two-stage correction method using the left and upper adjacent blocks has been explained here, it can also be configured to use the right or lower adjacent blocks to perform more corrections than the two-stage method.

[0101] Furthermore, the overlapping area is not the entire pixel area of ​​the block, but only a part of the area near the block boundary.

[0102] Furthermore, although the correction process for the predicted image from a single reference image has been described here, the same applies to the correction of the predicted image from multiple reference images. After obtaining the predicted images that have been corrected from each reference image, the resulting predicted images are further superimposed to form the final predicted image.

[0103] Furthermore, the aforementioned processing target block can be a predicted block unit, or a sub-block unit formed by further dividing the predicted block.

[0104] One method for determining whether OBMC processing is applicable is to use an OBMC flag (obmc_flag) that indicates whether OBMC processing is applicable. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a motion-complex region. If it does, a value of 1 is set as the OBMC flag (obmc_flag), and OBMC processing is applied for encoding. If it does not belong to a motion-complex region, a value of 0 is set as the OBMC flag (obmc_flag), and OBMC processing is not applied for encoding. On the other hand, in a decoding device, the OBMC flag (obmc_flag) described in the stream is decoded, thereby switching whether OBMC processing is applicable and decoding is performed according to the value of the flag.

[0105] Furthermore, motion information can also be exported from the decoding device without being signaled. For example, the merge mode specified in the H.265 / HEVC standard can be used. Alternatively, motion information can be exported by performing motion search on the decoding device side. In this case, motion search can be performed without using the pixel values ​​of the current block.

[0106] Here, we will explain the motion search mode performed on the decoding device side. Sometimes, the motion search mode performed on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0107] Figure 5D shows an example of FRUC processing. First, by referring to the motion vectors of coded blocks spatially or temporally adjacent to the current block, a list of multiple candidate blocks, each with motion vector predictors, is generated (this can also be common to the merge list). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, the evaluation value of each candidate in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0108] Then, the motion vector for the current block can be derived based on the selected candidate motion vectors. Specifically, for example, the selected candidate motion vector (best candidate MV) can be directly derived as the motion vector for the current block. Alternatively, the motion vector for the current block can be derived by pattern matching in the surrounding area of ​​the position in the reference image corresponding to the selected candidate motion vector. That is, the surrounding area of ​​the best candidate MV can be searched in the same way, and if an MV with a better evaluation value is found, the best candidate MV is updated to the aforementioned MV and used as the final MV for the current block. Furthermore, it is also possible to configure the block to not perform this process.

[0109] When processing at the sub-block level, the processing can also be set to be exactly the same.

[0110] Furthermore, the evaluation value is calculated by matching the patterns between the regions within the reference image corresponding to the motion vector and the specified regions to obtain the difference value of the reconstructed image. Moreover, in addition to the difference value, other information can also be used to calculate the evaluation value.

[0111] For pattern matching, either first-order pattern matching or second-order pattern matching can be used. Sometimes, first-order pattern matching and second-order pattern matching are referred to as bilateral matching and template matching, respectively.

[0112] In the first pattern matching, pattern matching is performed between two blocks within two different reference images, which are also along the motion trajectory of the current block. Therefore, in the first pattern matching, the predetermined region used to calculate the aforementioned candidate evaluation value is a region within another reference image along the motion trajectory of the current block.

[0113] Figure 6 illustrates an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory. As shown in Figure 6, in the first pattern matching, the best matching pair is searched among two blocks within two different reference images (Ref0, Ref1) along the motion trajectory of the current block, thereby deriving two motion vectors (MV0, MV1). Specifically, for the current block, the difference between two reconstructed images is derived, and an evaluation value is calculated using the obtained difference values. The two reconstructed images are the reconstructed image at a specified position within the first encoded complete reference image (Ref0) specified by the candidate MV, and the reconstructed image at a specified position within the second encoded complete reference image (Ref1) specified by the symmetrical MV obtained by scaling the candidate MV at time intervals. Among the multiple candidate MVs, the candidate MV with the best evaluation value is selected as the final MV.

[0114] Under the assumption of continuous motion trajectories, it is assumed that the motion vectors (MV0, MV1) of two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, if the current image is located between the two reference images in time, and the temporal distances from the current image to the two reference images are equal, then in the first-pattern matching, a mirror-symmetric bidirectional motion vector will be derived.

[0115] In the second pattern matching, pattern matching is performed between a template in the current image (a block adjacent to the current block in the current image (e.g., the adjacent block above and / or to the left)) and a block in the reference image. Therefore, in the second pattern matching, the predetermined region used to calculate the aforementioned candidate evaluation value is a block adjacent to the current block in the current image.

[0116] Figure 7 illustrates an example of pattern matching (template matching) between a template in the current image and a block in the reference image. As shown in Figure 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference image (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current image (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded completed regions of the left and top adjacent regions or either of them, and the reconstructed image at the same position in the encoded completed reference image (Ref0) specified by the candidate MV, is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value among the multiple candidate MVs is selected as the best candidate MV.

[0117] Information indicating whether FRUC mode applies (e.g., an FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode applies (e.g., when the FRUC flag is true), information regarding the pattern matching method (first pattern matching or second pattern matching) (e.g., an FRUC mode flag) is signaled at the CU level. Moreover, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, fragment level, tile level, CTU level, or sub-block level).

[0118] Here, we will explain the mode for deriving motion vectors based on a model that assumes uniform linear motion. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.

[0119] Figure 8 illustrates a model assuming constant linear motion. In Figure 8, (vx, vy) represents the velocity vector, while τ0 and τ1 represent the time distances between the current image (Cur Pic) and two reference images (Ref 0, Ref 1), respectively. (MVx 0, MVy 0) represents the motion vector corresponding to reference image Ref 0, and (MVx 1, MVy 1) represents the motion vector corresponding to reference image Ref 1.

[0120] At this point, under the assumption of constant linear motion of velocity vector (vx,vy), (MVx 0,MVy 0) and (MVx 1,MVy 1) are expressed as (v xτ 0,vy τ 0) and (-v xτ 1,-vy τ 1) respectively, and the following optical flow equation (1) holds. [Mathematical Expression 1]

[0121] Here, I(k) represents the luminance value of the reference image k (k=0,1) after motion compensation. This optical flow equation states that the sum of (i), (ii), and (iii) is zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of this optical flow equation and the Hermite interpolation formula, the motion vector of the block unit obtained from the merge list, etc., can be corrected in pixels.

[0122] Furthermore, motion vectors can also be derived on the decoding device side using methods different from those derived from motion vectors derived from a model assuming constant linear motion. For example, motion vectors can be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0123] Here, we will explain the mode of deriving motion vectors in sub-block units based on the motion vectors of a plurality of adjacent blocks. This mode is called the affine motion compensation prediction mode.

[0124] Figure 9A is a diagram illustrating the derivation of motion vectors for sub-block units based on motion vectors of multiple adjacent blocks. In Figure 9A, the current block is a 4×4 sub-block comprising 16 blocks. Here, the motion vector v0 of the upper left control point of the current block is derived from the motion vectors of adjacent blocks, and the motion vector v1 of the upper right control point of the current block is derived from the motion vectors of adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vectors (vx, vy) of each sub-block within the current block are derived using the following equation (2). [Mathematical Expression 2]

[0125] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents a pre-defined weighting coefficient.

[0126] In this affine motion compensation prediction mode, the method for deriving the motion vectors of the top-left and top-right control points can also include several different modes. The information displaying this affine motion compensation prediction mode (which may be referred to as, for example, affine flags) is signaled at the CU level. Furthermore, the signaling of the information displaying this affine motion compensation prediction mode is not limited to the CU level; it can also be at other levels (such as sequence level, image level, segment level, tile level, CTU level, or sub-block level). [Forecasting and Control Department]

[0127] The prediction control unit 128 selects either the in-frame prediction signal or the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0128] Here, an example of deriving motion vectors from an encoded object image using a merge mode is illustrated. Figure 9B is a diagram illustrating the overview of the motion vector deriving process performed using a merge mode.

[0129] First, a list of candidate predicted MVs is generated. Candidate predicted MVs include spatially adjacent predicted MVs, temporally adjacent predicted MVs, combined predicted MVs, and zero predicted MVs. The aforementioned spatially adjacent predicted MVs are the MVs of multiple encoded completed blocks spatially adjacent to the encoded object block. The aforementioned temporally adjacent predicted MVs are the MVs of blocks near the position of the encoded object block projected in the encoded completed reference image. The aforementioned combined predicted MVs are MVs generated by combining the MV values ​​of spatially adjacent predicted MVs and temporally adjacent predicted MVs. The aforementioned zero predicted MVs are MVs with a value of zero.

[0130] Next, the MV to be encoded is determined by selecting one of the multiple predicted MVs that have been registered in the predicted MV list.

[0131] Furthermore, in the variable-length encoding section, the signal that will show which predicted MV has been selected, namely merge_idx, is described and encoded in the stream.

[0132] Furthermore, the predicted MVs listed in the predicted MV list illustrated in Figure 9B are just one example. They can also be configured with a different number of MVs than those shown in the figure, or with a variety of MVs that do not include any of the predicted MVs shown in the figure, or with predicted MVs that are not included in the predicted MVs shown in the figure.

[0133] Furthermore, the MV of the encoded object block derived through the merging mode can also be used for the DMVR processing described later, thereby determining the final MV.

[0134] Here, we will illustrate an example of using DMVR processing to determine the MV.

[0135] Figure 9C is a conceptual diagram illustrating the outline of DMVR processing.

[0136] First, the most suitable MVP in the processing object block is set as the candidate MV. According to the aforementioned candidate MV, reference pixels are obtained from the processed image in the L0 direction (i.e., the first reference image) and the processed image in the L1 direction (i.e., the second reference image). The template is generated by averaging the values ​​of each reference pixel.

[0137] Next, using the aforementioned template, the surrounding areas of the candidate MVs for the first and second reference images are searched, and the MV with the lowest cost is selected as the final MV. Furthermore, the cost value is calculated using the differences between the pixel values ​​of the template and the pixel values ​​of the search area, as well as the MV value.

[0138] Furthermore, the general outline of the processing described herein is essentially the same in both the encoding and decoding devices.

[0139] Furthermore, even if it is not the processing described here, as long as it is a process that can search for the surrounding elements of candidate MVs to derive the final MV, other processing methods can also be used.

[0140] Here, we will explain the mode of generating predicted images using LIC processing.

[0141] Figure 9D is a schematic diagram illustrating a predictive image generation method that utilizes luminance correction processing via LIC processing.

[0142] First, export the MV of the reference image used to obtain the corresponding block of the encoded object from the encoded image, i.e., the reference image.

[0143] Next, for the encoded object block, the brightness pixel values ​​of the surrounding reference area are extracted using the left and top adjacent encoded data, as well as the brightness pixel values ​​at the same position in the reference image specified by MV. Information on how the brightness values ​​change in the reference image and the encoded object image is then extracted, and the brightness correction parameters are calculated.

[0144] For the reference image within the reference image specified by MV, brightness correction is performed using the aforementioned brightness correction parameters, thereby generating a predicted image relative to the encoded object block.

[0145] Furthermore, the shape of the aforementioned surrounding reference area in Figure 9D is just an example; other shapes may also be used.

[0146] Furthermore, although the process of generating a prediction image from a single reference image has been described here, the same applies to generating a prediction image from multiple reference images. The prediction image is generated after performing brightness correction processing on the reference images obtained from each reference image in the same way.

[0147] One method for determining whether LIC processing is applicable is by using a LIC flag (lic_flag) to indicate whether LIC processing is applicable. Specifically, in an encoding device, it is determined whether the area to be encoded belongs to a region where brightness has changed. If it does, a value of 1 is set as the LIC flag (lic_flag), and LIC processing is applied for encoding. If it does not belong to a region where brightness has changed, a value of 0 is set as the LIC flag (lic_flag), and encoding is performed without applying LIC processing. On the other hand, in a decoding device, the LIC flag (lic_flag) described in the stream is decoded, thereby switching whether LIC processing is applied based on the value of the LIC flag during decoding.

[0148] Other methods for determining whether LIC processing is applicable include, for example, determining whether LIC processing is applicable to surrounding blocks. As a specific example, when the encoded target block is in merge mode, it is determined whether the surrounding encoded blocks selected during the MV export in merge mode processing have been processed and encoded using LIC, and the application of LIC processing is switched accordingly. Furthermore, in this case, the decoding process is exactly the same. [Overview of the Decoding Device]

[0149] Next, an overview of the decoding apparatus capable of decoding the encoded signal (encoded bit stream) output from the aforementioned encoding device 100 will be described. Figure 10 is a block diagram showing the functional configuration of the decoding apparatus 200 in Embodiment 1. The decoding apparatus 200 is a motion picture / image decoding apparatus that decodes motion pictures / images in block units.

[0150] As shown in Figure 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse conversion unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an in-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.

[0151] The decoding device 200 can be implemented using, for example, a general-purpose processor and memory. In this case, when the processor executes the software program stored in memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse conversion unit 206, an adder 208, a loop filter 212, an in-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. Alternatively, the decoding device 200 can be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse conversion unit 206, the adder 208, the loop filter 212, the in-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.

[0152] The following describes the constituent elements included in the decoding device 200. [Entropy Decoding Department]

[0153] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202 performs arithmetic decoding on the binary signal from the encoded bitstream, for example. Then, the entropy decoding unit 202 debinarizes the binary signal. In doing so, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. [De-quantization Department]

[0154] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the input, i.e., the decoded target block (hereinafter referred to as the current block), from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient of the current block according to the quantization parameter corresponding to that quantization coefficient. Furthermore, the inverse quantization unit 204 outputs the inverse quantization coefficients (i.e., conversion coefficients) of the current block to the inverse conversion unit 206. [Reverse Conversion Section]

[0155] The inverse conversion unit 206 restores the prediction error by performing an inverse conversion on the input, i.e., the conversion coefficient, from the inverse quantization unit 204.

[0156] If the information that has been interpreted from the encoded bit stream indicates that EMT or AMT is applicable (e.g., the AMT flag is true), the inverse conversion unit 206 will perform inverse conversion on the conversion coefficients of the current block based on the information indicating the interpreted conversion type.

[0157] For example, if the information that has been interpreted from the encoded bit stream is applicable to NSST, the inverse conversion unit 206 will apply an inverse reconversion to the conversion coefficients. [Addition Department]

[0158] The addition unit 208 performs an addition operation on the input (prediction error) from the inverse conversion unit 206 and the input (prediction sample) from the prediction control unit 220, thereby reconstructing the current block. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212. [Block Memory]

[0159] Block memory 210 is a storage unit used to store blocks referenced in in-frame prediction and also blocks within the decoded object image (hereinafter referred to as the current image). Specifically, block memory 210 stores the reconstructed blocks output from addition unit 208. [Loop Filtering Section]

[0160] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.

[0161] When the information on ALF being turned on / off, as interpreted from the encoded bitstream, shows that ALF is on, one filter can be selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter can be applied to the reconstructed block. [Frame Memory]

[0162] Frame memory 214 is a storage unit used to store reference images used for inter-frame prediction, and is sometimes called a frame buffer. Specifically, frame memory 214 stores reconstructed blocks that have been filtered by loop filter 212. [In-frame prediction section]

[0163] The intra-frame prediction unit 216 generates a prediction signal (intra-frame prediction signal) by performing intra-frame prediction based on the intra-frame prediction pattern decoded from the encoded bit stream and referring to the blocks within the current image stored in the block memory 210. Specifically, the intra-frame prediction unit 216 performs intra-frame prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, thereby generating an intra-frame prediction signal, and outputs the intra-frame prediction signal to the prediction control unit 220.

[0164] Furthermore, when the in-frame prediction mode of the reference luminance block is selected in the in-frame prediction of the chromatic difference block, the in-frame prediction unit 216 can also predict the chromatic difference component of the current block based on the luminance component of the current block.

[0165] Furthermore, when the information interpreted from the encoded bit stream indicates the applicability of PDPC, the in-frame prediction unit 216 will correct the predicted pixel values ​​in the in-frame based on the gradient of the reference pixels in the horizontal / vertical direction. [Inter-frame Prediction Department]

[0166] The inter-frame prediction unit 218 predicts the current block by referring to a reference image stored in the frame memory 214. The prediction is performed on a unit of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 218 uses motion information (e.g., motion vectors) interpreted from the encoded bit stream to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.

[0167] Furthermore, when the information interpreted from the encoded bit stream is applicable to OBMC mode, the inter-frame prediction unit 218 will use not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks to generate inter-frame prediction signals.

[0168] Furthermore, when the information interpreted from the encoded bit stream is applicable to FRUC mode, the inter-frame prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) interpreted from the encoded stream, thereby deriving motion information. Then, the inter-frame prediction unit 218 uses the derived motion information to perform motion compensation.

[0169] Furthermore, when using BIO mode, the inter-frame prediction unit 218 derives motion vectors based on a model assuming constant velocity linear motion. Also, when displaying information interpreted from the encoded bitstream using affine motion compensation prediction mode, the inter-frame prediction unit 218 derives motion vectors in sub-block units based on the motion vectors of multiple adjacent blocks. [Forecasting and Control Department]

[0170] The prediction control unit 220 selects either the in-frame prediction signal or the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208. (Implementation Form 2)

[0171] Next, we will explain Embodiment 2. This embodiment concerns inter-frame prediction in the so-called BIO mode. In this embodiment, the motion vector of the block unit is corrected not at the pixel unit, but at the sub-block unit, which differs from Embodiment 1 described above. The following explanation will focus on the differences between this embodiment and Embodiment 1.

[0172] Furthermore, since the structure of the encoding and decoding devices in this embodiment is essentially the same as that in Embodiment 1, illustrations and descriptions are omitted. [Inter-frame prediction]

[0173] Figure 11 is a flowchart showing the inter-frame prediction in Embodiment 2. Figure 12 is a conceptual diagram illustrating the inter-frame prediction in Embodiment 2. The following processing is performed by the inter-frame prediction unit 126 of the encoding device 100 or the inter-frame prediction unit 218 of the decoding device 200.

[0174] As shown in Figure 11, firstly, a block-unit loop processing (S101~S111) is performed on multiple blocks within the encoded / decoded object image (current image 1000). In Figure 12, the encoded / decoded object block is selected from the multiple blocks as the current block 1001.

[0175] In the loop processing of block units, the processed images, namely the first reference image 1100 (L0) and the second reference image 1200 (L1), are processed in loop units (S102~S106).

[0176] In the loop processing of the reference image unit, firstly, motion vectors for obtaining block units of the predicted image from the reference image are derived (S103). In Figure 12, a first motion vector 1110 (MV_L0) is derived or obtained for the first reference image 1100, and a second motion vector 1210 (MV_L1) is derived or obtained for the second reference image 1200. Methods for deriving motion vectors include the general inter-frame prediction mode, the merging mode, and the FRUC mode. For example, in the case of the general inter-frame prediction mode, the motion vector is derived by motion search in the encoding device 100, and the motion vector is obtained from the bit stream in the decoding device 200.

[0177] Next, motion compensation is performed using the exported or acquired motion vectors, thereby obtaining a predicted image from the reference image (S104). In Figure 12, motion compensation is performed using the first motion vector 1110, thereby obtaining a first predicted image 1140 from the first reference image 1100. Furthermore, motion compensation is performed using the second motion vector 1210, thereby obtaining a second predicted image 1240 from the second reference image 1200.

[0178] In motion compensation, a motion compensation filter can be applied to a reference image. The motion compensation filter is an interpolation filter used to obtain a predicted image with sub-pixel precision. In the first reference image 1100 of FIG12, pixels within a first interpolation reference range 1130, which includes the pixels of the first prediction block 1120 and its surrounding pixels, are referenced using a motion compensation filter relative to the first prediction block 1120 specified by the first motion vector 1110. Similarly, in the second reference image 1200, pixels within a second interpolation reference range 1230, which includes the pixels of the second prediction block 1220 and its surrounding pixels, are referenced using a motion compensation filter relative to the second prediction block 1220 specified by the second motion vector 1210.

[0179] Furthermore, the first interpolation reference range 1130 and the second interpolation reference range 1230, in normal frame prediction without processing using local motion estimation value, are included in the first normal reference range and the second normal reference range referenced for motion compensation of the current block 1001. The first normal reference range is included in the first reference image 1100, and the second normal reference range is included in the second reference image 1200. In normal frame prediction, motion vectors are derived in block units, for example, by motion search, and motion compensation is performed in block units using the derived motion vectors, and the motion-compensated image is directly used as the final predicted image. That is, local motion estimation value is not used in normal frame prediction. Moreover, the first interpolation reference range 1130 and the second interpolation reference range 1230 can also be consistent with the first normal reference range and the second normal reference range.

[0180] Next, a gradient image corresponding to the predicted image is obtained from the reference image (S105). Each pixel in the gradient image has a gradient value that represents the slope of the brightness or color difference in space. The gradient value can be obtained by applying a gradient filter to the reference image. In the first reference image 1100 of FIG12, the pixels of the first gradient reference range 1135, which includes the pixels of the first prediction block 1120 and its surrounding pixels, are referenced by the gradient filter used for the first prediction block 1120. This first gradient reference range 1135 is included in the first interpolation reference range 1130. Similarly, in the second reference image 1200, the gradient filter refers to the pixels of the second gradient reference range 1235, which includes the pixels of the second prediction block 1220 and its surrounding pixels. This second gradient reference range 1235 is included in the second interpolation reference range 1230.

[0181] If the acquisition of the predicted image and gradient image of each of the first and second reference images ends, the loop processing of the reference image unit ends (S106). Afterwards, loop processing of sub-block units that further divide the block is performed (S107~S110). Each of the plurality of sub-blocks has a size smaller than the current block (e.g., 4×4 pixels).

[0182] In the loop processing of the sub-block unit, firstly, a local motion estimation value 1300 for the sub-block is derived using the first prediction image 1140 and the second prediction image 1240, and the first gradient image 1150 and the second gradient image 1250 obtained from the first reference image 1100 and the second reference image 1200 (S108). For example, in each of the first prediction image 1140 and the second prediction image 1240, and the first gradient image 1150 and the second gradient image 1250, a local motion estimation value 1300 is derived for the sub-block by referring to the pixels contained in the predicted sub-block. The predicted sub-block refers to the region within the first prediction block 1120 and the second prediction block 1220 corresponding to the sub-block within the current block 1001. The local motion estimation value is sometimes also called the corrected motion vector.

[0183] Next, using the pixel values ​​of the first predicted image 1140 and the second predicted image 1240, the gradient values ​​of the first gradient image 1150 and the second gradient image 1250, and the local motion estimation value 1300, the final predicted image 1400 of the sub-block is generated (S109). If the generation of the final predicted image for each of the sub-blocks contained in the current block is completed, the final predicted image of the current block is generated, and the loop processing of the sub-block unit is completed (S110).

[0184] Furthermore, if the loop processing of the block unit ends (S111), then the processing in Figure 11 ends.

[0185] Furthermore, the motion vector of the current block unit can be directly assigned to each sub-block, thereby obtaining the prediction image and gradient image using the sub-block unit. [Reference range for motion compensation filters and gradient filters]

[0186] Here, the reference range for motion compensation filters and gradient filters is explained.

[0187] Figure 13 is a conceptual diagram illustrating one example of the reference range for the motion compensation filter and gradient filter in embodiment 2.

[0188] In Figure 13, each of the plurality of circular symbols represents a pixel. Figure 13 also uses the example of setting the current block size to 8×8 pixels and the sub-block size to 4×4 pixels.

[0189] Reference range 1131 is the reference range (for example, an 8×8 pixel rectangular range) for displaying the motion compensation filter applicable to the top left pixel 1122 of the first prediction block 1120. Reference range 1231 is the reference range (for example, an 8×8 pixel rectangular range) for displaying the motion compensation filter applicable to the top left pixel 1222 of the second prediction block 1220.

[0190] Furthermore, reference range 1132 is a reference range (for example, a rectangular range of 6×6 pixels) for displaying the gradient filter applied to the top left pixel 1122 of the first prediction block 1120. Reference range 1232 is a reference range (for example, a rectangular range of 6×6 pixels) for displaying the gradient filter applied to the top left pixel 1222 of the second prediction block 1220.

[0191] For other pixels within the first prediction block 1120 and the second prediction block 1220, motion compensation filters and gradient filters are applied while referring to pixels within a reference range of the same size corresponding to the position of each pixel. As a result, to obtain the first prediction image 1140 and the second prediction image 1240, pixels within the first interpolation reference range 1130 and the second interpolation reference range 1230 can be referenced. Similarly, to obtain the first gradient image 1150 and the second gradient image 1250, pixels within the first gradient reference range 1135 and the second gradient reference range 1235 can be referenced. [Effects, etc.]

[0192] As described above, the encoding and decoding apparatus of this embodiment can derive local motion estimation values ​​in sub-block units. Therefore, prediction errors can be reduced by utilizing local motion estimation values ​​in sub-block units, and the processing load or processing time is reduced compared to deriving local motion estimation values ​​in pixel units.

[0193] Furthermore, according to the encoding and decoding apparatus of this embodiment, the interpolation reference range can be included in the normal reference range. Therefore, in the generation of the final predicted image using the local motion estimation value of sub-block units, it is not necessary to read new pixel data from the frame memory for motion compensation, thus suppressing the increase in memory capacity and memory bandwidth.

[0194] Furthermore, according to the encoding and decoding apparatus of this embodiment, the gradient reference range can be included in the interpolation reference range. Therefore, it is not necessary to read new pixel data from the frame memory in order to obtain the gradient image, and the increase in memory capacity and memory bandwidth can be suppressed.

[0195] This sample may also be implemented in combination with at least a portion of other samples disclosed herein. Furthermore, a portion of the processing described in the flowchart of this sample, a portion of the configuration of the apparatus, and a portion of the syntax may also be implemented in combination with other samples. (Variation 1 of Implementation 2)

[0196] Next, variations of the motion compensation filter and gradient filter will be explained in detail while referring to the diagram. Furthermore, in the following variation 1, since the processing of the second predicted image is similar to the processing of the first predicted image, the explanation will be omitted or simplified as appropriate. [Motion Compensation Filter]

[0197] First, the motion compensation filter will be explained. Figure 14 is a conceptual diagram of one example of the reference range used to illustrate the motion compensation filter in Modification 1 of Embodiment 2.

[0198] Here, we will take the case of applying a 1 / 4 pixel motion compensation filter in the horizontal direction and a 1 / 2 pixel motion compensation filter in the vertical direction to the first prediction block 1120 as an example. The motion compensation filter is a so-called 8-tap filter, which can be represented by the following mathematical formula (3). [Mathematical Expression 3]

[0199] Here, Ik[x,y] represents the pixel values ​​of the first predicted image with fractional pixel precision when k is 0, and the pixel values ​​of the second predicted image with fractional pixel precision when k is 1. Pixel values ​​refer to the values ​​possessed by a pixel, such as brightness or color difference values ​​in the predicted image. w0.25 and w0.5 are weighting coefficients for displaying 1 / 4 pixel precision and 1 / 2 pixel precision, respectively. I0k[x,y] represents the pixel values ​​of the first predicted image with integer pixel precision when k is 0, and the pixel values ​​of the second predicted image with integer pixel precision when k is 1.

[0200] For example, when applying the motion compensation filter of mathematical formula (3) to the top left pixel 1122 of Figure 14, the values ​​of the pixels arranged in the horizontal direction within the reference range 1131A can be weighted and summed in each column, and the summed result from the complex column can be further weighted and summed.

[0201] Thus, in this variant, the motion compensation filter used for the top-left pixel 1122 is the pixel referenced to the reference range 1131A. The reference range 1131A is a rectangular area extending 3 pixels to the left, 4 pixels to the right, 3 pixels up, and 4 pixels down from the top-left pixel 1122.

[0202] This motion compensation filter can be applied to all pixels within the first prediction block 1120. Thus, in the motion compensation filter used for the first prediction block 1120, pixels within the first interpolation reference range 1130A can be referenced.

[0203] The motion compensation filter is applied to the second prediction block 1220 in the same way as to the first prediction block 1120. That is, the pixel in reference range 1231A can be referenced for the top left pixel 1222, and the pixel in the second prediction block 1220 as a whole is referenced to the second interpolation reference range 1230A. Gradient Filter

[0204] Next, the gradient filter will be explained. Figure 15 is a conceptual diagram illustrating one example of the reference range of the gradient filter in Modification 1 of Embodiment 2.

[0205] The gradient filter in this variation is a so-called 5-tap filter, which can be represented by the following mathematical expressions (4) and (5). [Mathematical Expression 4] [Mathematical Expression 5]

[0206] Here, Ixk[x,y] displays the horizontal gradient values ​​of each pixel in the first gradient image when k is 0, and the horizontal gradient values ​​of each pixel in the second gradient image when k is 1. Iyk[x,y] displays the vertical gradient values ​​of each pixel in the first gradient image when k is 0, and the vertical gradient values ​​of each pixel in the second gradient image when k is 1. w is the weighting coefficient.

[0207] For example, when applying the gradient filter of mathematical formulas (4) and (5) to the top-left pixel 1122 in Figure 15, the horizontal gradient value is calculated by weighted summation of the following pixel values: the pixel values ​​of the five pixels arranged horizontally including the top-left pixel 1122, and the pixel values ​​of the predicted image with integer pixel precision. Similarly, the vertical gradient value is calculated by weighted summation of the following pixel values: the pixel values ​​of the five pixels arranged vertically including the top-left pixel 1122, and the pixel values ​​of the predicted image with integer pixel precision. In this case, the weighting coefficients have values ​​that are opposite in sign to the pixels above or below, or left or right, with the top-left pixel 1122 as the symmetrical point.

[0208] Thus, in this variant, the gradient filter used for the top-left pixel 1122 is a pixel that references the reference range 1132A. The reference range 1132A has a cross shape extending two pixels in all directions from the top-left pixel 1122.

[0209] This gradient filter can be applied to all pixels within the first prediction block 1120. Thus, in the motion compensation filter used for the first prediction block 1120, pixels within the first gradient reference range 1135A can be referenced.

[0210] The gradient filter is applied to the second prediction block 1220 in the same way as to the first prediction block 1120. That is, the pixel in reference range 1232A can be referenced for the top left pixel 1222, and in the second prediction block 1220 as a whole, the pixel in reference range 1235A is referenced.

[0211] Furthermore, when the motion vector within a specified reference range displays fractional pixel positions, the pixel values ​​of the reference ranges 1132A and 1232A of the gradient filter can be converted to fractional pixel precision pixel values, and the gradient filter can be applied to the converted pixel values. Alternatively, the gradient filter can be applied to pixel values ​​with integer pixel precision, where the gradient filter has the following coefficient values: the value obtained by multiplying the coefficient value used for converting to fractional pixel precision with the coefficient value used to derive the gradient value. In this case, the gradient filter differs for each fractional pixel position. [Derivation of Local Motion Estimation Values ​​for Sub-Block Units]

[0212] Next, the derivation of the local motion estimation value of the sub-block unit will be explained. Specifically, the derivation of the local motion estimation value of the top-left sub-block among the multiple sub-blocks contained in the current block will be used as an example.

[0213] In this variation, the horizontal local motion estimate value u and the vertical local motion estimate value v of the sub-block are derived based on the following mathematical formula (6). [Mathematical Expression 6]

[0214] Here, sG xG y, sG x 2, sG y 2, sG xdI, and sG ydI are values ​​calculated in sub-block units and are respectively calculated according to the following mathematical formula (7). [Mathematical Expression 7]

[0215] Here, Ω is the set of coordinates of all pixels within a sub-block of the prediction block. Gx[i,j] represents the sum of the horizontal gradient values ​​of the first and second gradient images, and Gy[i,j] represents the sum of the vertical gradient values ​​of the first and second gradient images. ΔI[i,j] represents the difference between the first and second prediction images. w[i,j] represents the weighting coefficient dependent on the pixel position within the prediction sub-block. For example, the same weighting coefficient can be used for all pixels within the prediction sub-block.

[0216] Specifically, Gx[i,j], Gy[i,j], and ΔI[i,j] can be represented by the following mathematical expression (8). [Mathematical Expression 8]

[0217] As shown above, the local motion estimation value can be calculated in sub-block units. [Generation of the final predicted image]

[0218] Next, the generation of the final predicted image will be explained. The pixel values ​​p[x,y] of the final predicted image are calculated using the pixel values ​​I0[x,y] of the first predicted image and I1[x,y] of the second predicted image, and according to the following mathematical formula (9). [Mathematical Expression 9]

[0219] Here, b[x,y] represents the correction value for each pixel. In mathematical formula (9), the pixel values ​​p[x,y] of the final predicted image are calculated by right-shifting the sum of the pixel values ​​I0[x,y] of the first predicted image, the pixel values ​​I1[x,y] of the second predicted image, and the correction value b[x,y]. Furthermore, the correction value b[x,y] is represented by the following mathematical formula (10). [Mathematical Expression 10]

[0220] In mathematical formula (10), the correction value b[x,y] is calculated by adding the following results: the difference between the horizontal gradient values ​​(Ix0[x,y]-Ix1[x,y]) between the first gradient image and the second gradient image multiplied by the horizontal local motion estimation value (u), and the difference between the vertical gradient values ​​(Iy0[x,y]-Iy1[x,y]) between the first gradient image and the second gradient image multiplied by the vertical local motion estimation value (v).

[0221] Furthermore, the operational expressions described using mathematical expressions (6) to (10) are just one example. Any operational expression with the same effect is acceptable, even if it is a mathematical expression different from that operational expression. [Effects, etc.]

[0222] As described above, using the motion compensation filter and gradient filter of this variant, local motion estimation values ​​can also be derived in sub-block units. If the local motion estimation values ​​of sub-block units derived in this way are used to generate the final predicted image of the current block, the same effect as in embodiment 2 described above can be obtained.

[0223] This sample may also be implemented in combination with at least a portion of other samples disclosed herein. Furthermore, a portion of the processing described in the flowchart of this sample, a portion of the configuration of the apparatus, and a portion of the syntax may also be implemented in combination with other samples. (Variation 2 of Implementation 2)

[0224] In Embodiment 2 and its variation 1 described above, although the local motion estimation value is derived by referring to all pixels contained in the predicted sub-block within the predicted sub-block corresponding to the current block, it is not limited to this. For example, it is also possible to refer to only a portion of the pixels contained in the predicted sub-block.

[0225] Therefore, in this variation, the explanation focuses on the case where, in deriving the local motion estimation value of a sub-block unit, only a portion of the pixels contained in the predicted sub-block are considered. For example, in the mathematical formula (7) of the above variation 1, the set of coordinates of a portion of the pixels within the predicted sub-block is used instead of the set of coordinates of all pixels contained in the predicted sub-block, i.e., Ω. Various types of sets of coordinates of pixels within the predicted sub-block can be used as the set of coordinates of pixels within the predicted sub-block.

[0226] Figure 16 is an example diagram showing the pattern of the pixels referenced in the derivation of the local motion estimation value in Variation 2 of Embodiment 2. In Figure 16, the shaded circular marks within the predicted sub-blocks 1121 or 1221 indicate the referenced pixels, while the unshaded circular marks indicate the unreferenced pixels.

[0227] Each of the seven pixel patterns in Figures 16(a) to (g) is a pixel representing a portion of the plurality of pixels contained in the predicted sub-block 1121 or 1221. Furthermore, the seven pixel patterns are distinct from one another.

[0228] In Figures 16(a) to (c), only 8 pixels out of the 16 pixels contained in the predicted sub-block 1121 or 1221 are referenced. Also, in Figures 16(d) to (g), only 4 pixels out of the 16 pixels contained in the predicted sub-block 1121 or 1221 are referenced. That is, in Figures 16(a) to (c), 8 pixels out of the 16 pixels are removed, and in Figures 16(d) to (g), 12 pixels out of the 16 pixels are removed.

[0229] More specifically, in Figure 16(a), eight pixels are referenced that are offset from each other by one pixel in the horizontal / vertical direction. In Figure 16(b), the left and right pixel pairs of two pixels arranged in the horizontal direction are interactively referenced in the vertical direction. In Figure 16(c), the four central pixels and four corner pixels within the predicted sub-block 1121 or 1221 are referenced.

[0230] Furthermore, in Figures 16(d) and (e), the pixels in the first and third rows from the left are referenced by 2 pixels per row. In Figure 16(f), the four corner pixels are referenced. In Figure 16(g), the four central pixels are referenced.

[0231] Alternatively, a pixel pattern can be appropriately selected from these predefined plurality of pixel patterns based on the two predicted images. For example, a pixel pattern containing a number of pixels corresponding to the representative gradient values ​​of the two predicted images can be selected. Specifically, if the representative gradient value is smaller than the threshold, a pixel pattern containing 4 pixels (e.g., any one of (d) to (g)) can be selected; otherwise, a pixel pattern containing 8 pixels (e.g., any one of (a) to (c)) can be selected.

[0232] When selecting a pixel pattern from a plurality of pixel patterns, the local motion estimate of the sub-block can be derived by referring to the pixels in the predicted sub-block shown by the selected pixel pattern.

[0233] Furthermore, information about the selected pixel pattern can also be written into the bitstream. In this case, the decoding device only needs to obtain information from the bitstream and select the pixel pattern based on the obtained information. Information about the selected pixel pattern can be written into the header of, for example, a block, segment, image, or stream unit.

[0234] As described above, the encoding and decoding apparatus of this embodiment can derive the local motion estimation value by referring only to a portion of the pixels contained in the predicted sub-block, using sub-block units. Therefore, compared to referring to all of the pixels, the processing load or processing time can be further reduced.

[0235] Furthermore, according to the encoding and decoding apparatus of this embodiment, the local motion estimation value can be derived by referring only to the pixels contained in a pixel pattern selected from a plurality of pixel patterns, in sub-block units. Thus, by switching pixel patterns to refer to the derived pixels suitable for the local motion estimation value of the sub-block, the prediction error can be reduced.

[0236] This sample may also be implemented in combination with at least a portion of other samples disclosed herein. Furthermore, a portion of the processing described in the flowchart of this sample, a portion of the configuration of the apparatus, and a portion of the syntax may also be implemented in combination with other samples. (Other variations of implementation form 2)

[0237] The above description, while describing one or more types of encoding and decoding devices according to embodiments and variations thereof, is not limited to these embodiments and variations. Any modifications conceivable to those skilled in the art that pertain to this invention, applied to these embodiments or variations thereof, without departing from the spirit of this disclosure, may also be included within the scope of one or more types of this disclosure.

[0238] For example, although the motion compensation filter in Embodiment 2 and its variation 1 described above has 8 taps, it is not limited to this. The number of taps in the motion compensation filter can be any number of taps, as long as the interpolation reference range is included in the normal reference range.

[0239] Furthermore, although the number of taps in Embodiment 2 and its variation 1 is 6 pixels or 5 pixels, it is not limited to this. Any other number of taps is possible as long as the gradient reference range is included in the interpolation reference range.

[0240] Furthermore, although in the above-described embodiment 2 and its variation 1, the first gradient reference range and the second gradient reference range are included in the first interpolation reference range and the second interpolation reference range, it is not limited to this. For example, the first gradient reference range may also be consistent with the first interpolation reference range, and the second gradient reference range may also be consistent with the second interpolation reference range.

[0241] Furthermore, when deriving the local motion estimation value in sub-block units, pixel values ​​can be weighted to prioritize the value of the central pixel of the predicted sub-block. That is, in deriving the local motion estimation value, the values ​​of multiple pixels contained in the predicted sub-block can be weighted in each of the first and second prediction blocks, and in this case, the pixel located closer to the center of the predicted sub-block can have a larger weight. More specifically, for example, in Variation 1 of Embodiment 2, the weighting coefficient w[i,j] in mathematical formula (7) has a larger value the closer the coordinate value is to the center of the predicted sub-block.

[0242] Furthermore, when deriving the local motion estimation value in sub-block units, pixels within adjacent prediction sub-blocks belonging to the same prediction block can also be referenced. That is, in each of the first and second prediction blocks, in addition to the plurality of pixels contained in the prediction sub-block, pixels contained in other prediction sub-blocks adjacent to the prediction sub-block within that prediction block can also be referenced to derive the local motion estimation value of the sub-block unit.

[0243] Furthermore, the reference range of the motion compensation filter and gradient filter in the above-described embodiment 2 and its variant 1 is merely illustrative and need not be limited thereto.

[0244] Furthermore, although seven pixel patterns are illustrated in variation 2 of embodiment 2 described above, it is not a limitation. Pixel patterns obtained by rotating each of the seven pixel patterns can also be used.

[0245] Furthermore, the weighting coefficient value in Variation 1 of Embodiment 2 is just one example and is not limited to it. Also, the block size and sub-block size in Embodiment 2 and its variations are just examples and are not limited to 8×8 pixel size and 4×4 pixel size. Even with other sizes, inter-frame prediction can still be performed in the same way as in Embodiment 2 and its variations.

[0246] This sample may also be implemented in combination with at least a portion of other samples disclosed herein. Furthermore, a portion of the processing described in the flowchart of this sample, a portion of the configuration of the apparatus, and a portion of the syntax may also be implemented in combination with other samples. (Implementation Form 3)

[0247] In the above embodiments, each functional block can typically be implemented using an MPU and memory. Furthermore, the processing performed by each functional block is usually achieved by having a program execution unit such as a processor read and execute software (programs) recorded on a recording medium such as ROM. This software can be distributed via download or by recording it on a recording medium such as semiconductor memory. Alternatively, each functional block can, of course, be implemented using hardware (dedicated circuitry).

[0248] Furthermore, the processing described in each embodiment can be implemented centrally using a single device (system) or distributedly using multiple devices. Also, the processor executing the above program can be single or multiple. That is, it can be either centralized or distributed processing.

[0249] The present invention is not limited to the above embodiments and various modifications can be made, and such modifications are also included within the scope of the present invention.

[0250] Furthermore, examples of applications of the motion image encoding method (image encoding method) or motion image decoding method (image decoding method) shown in the above embodiments and systems utilizing them will be described herein. The system is characterized by having an image encoding device utilizing the image encoding method, an image decoding device utilizing the image decoding method, and an image encoding / decoding device possessing both. Other components of the system may be appropriately modified depending on the circumstances. [Usage Example]

[0251] Figure 17 is a diagram showing the overall structure of the content delivery system ex100 that implements the content delivery service. The communication service provision area is divided into desired sizes, and fixed radio stations, i.e., base stations ex106, ex107, ex108, ex109, and ex110, are respectively set in each cell.

[0252] In this content delivery system ex100, various devices such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, and smartphones ex115 can be connected to the Internet ex101 via Internet service provider ex102 or communication network ex104, and base stations ex106~ex110. The content delivery system ex100 can also be configured to combine and connect any of the above-mentioned elements. Alternatively, the devices can be directly or indirectly connected to each other via telephone networks or short-range wireless communication without using base stations ex106~ex110, which function as fixed radio stations. Furthermore, the streaming server ex103 connects to the various devices such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, and smartphones ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 connects to terminals in the hotspot within the aircraft ex117 via satellite ex116.

[0253] Furthermore, wireless access points or hotspots can be used to replace base stations ex106~ex110. Also, the streaming server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or Internet service provider ex102, and can also connect directly to the aircraft ex117 without going through the satellite ex116.

[0254] The EX113 camera is a digital camera or similar device capable of still and video photography. The EX115 smartphone can be a smartphone, mobile phone, or PHS (Personal Handyphone System) compatible with mobile communication systems commonly referred to as 2G, 3G, 3.9G, 4G, and the future 5G.

[0255] The home appliance ex118 can be a refrigerator, or a machine included in a home fuel cell cogeneration system, etc.

[0256] In the content delivery system ex100, live streaming is made possible by connecting a terminal with photography capabilities to the streaming server ex103 via a base station ex106, etc. During live streaming, the terminal (computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and terminal inside an airplane ex117, etc.) performs the encoding processing described in the aforementioned embodiments on the still or moving images captured by the user using the terminal, and multiplexes the encoded image data and the corresponding audio data (encoded from the image data) before transmitting the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of this disclosure.

[0257] On the other hand, the streaming server ex103 streams the content data sent to requesting clients. Clients refer to terminals such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or airplanes ex117 that can decode the data that has undergone the aforementioned encoding process. Each machine that has received the sent data will decode and play it. That is, each machine functions as an image decoding device, as described in this disclosure. [Distributed processing]

[0258] Furthermore, the streaming server ex103 can also be a plurality of servers or a plurality of computers that distribute, process, or record data for transmission. For example, the streaming server ex103 can be implemented using a CDN (Content Delivery Network), or it can be implemented using a network connecting numerous edge servers distributed around the world. On a CDN, physically proximate edge servers are dynamically allocated according to the client. Furthermore, latency can be reduced by caching the content and sending it to the edge server. Moreover, since processing can be distributed among multiple edge servers or the transmission subject can be switched to other edge servers to bypass the obstructed parts of the network and continue transmission when an error occurs or the communication state changes due to increased traffic, high-speed and stable transmission can be achieved.

[0259] Furthermore, beyond the distributed processing of the transmission itself, the encoding of captured data can be performed on each terminal, on the server side, or distributed among different terminals. For example, encoding typically involves two processing loops. In the first loop, the complexity or encoding amount of the image within a frame or scene unit can be detected. In the second loop, processing can be performed to maintain image quality and improve encoding efficiency. For instance, by having the terminal perform the first encoding process and the server receiving the content perform the second encoding process, the processing load on each terminal can be reduced, and the quality and efficiency of the content can be improved. In this case, if there is a requirement for near-real-time reception and decoding, the data encoded in the first step by one terminal can be received and played by other terminals, thus enabling more flexible real-time transmission.

[0260] As another example, cameras like the EX113 extract features from images and compress the data related to those features as metadata before sending it to the server. The server performs compression based on the meaning of the image, such as determining the importance of the object from the features and adjusting the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed on the terminal, while more computationally demanding encoding such as CABAC (Context-Referenced Adaptive Binary Arithmetic Coding) can be performed on the server.

[0261] Furthermore, as another example, in sports fields, shopping malls, or factories, there may be multiple images of almost identical scenes captured by multiple terminals. In such cases, the multiple terminals that have taken photos, along with other terminals and servers that haven't taken photos as needed, can be used to distribute the encoding and processing by assigning them to different units, such as GOP (Group of Pictures), image units, or tile units that have already divided the images. This reduces latency and achieves greater real-time performance.

[0262] Furthermore, since multiple image data sets depict almost identical scenes, they can be managed and / or instructed on the server to coordinate and reference the image data captured by each terminal. Alternatively, the server can receive encoded data from each terminal and change the reference relationships between the multiple data sets, or correct or replace the images themselves and re-encode them. This allows the generation of a stream with improved quality and efficiency for each piece of data.

[0263] Furthermore, the server can also transcode the image data after changing its encoding method before sending the image data. For example, the server can convert MPEG encoding to VP encoding, or H.264 to H.265.

[0264] Thus, encoding processing can be performed via a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the subject of processing, some or all of the processing performed on the server can be performed on the terminal, and some or all of the processing performed on the terminal can be performed on the server. Furthermore, the same applies to decoding processing. [3D, Multi-angle]

[0265] In recent years, there has been a growing trend of integrating and utilizing images or videos captured by multiple cameras such as the EX113 and / or smartphones such as the EX115, which are used almost simultaneously, to capture different scenes or to capture the same scene from different angles. The images captured by each device are integrated based on the relative positional relationship between the devices or the fact that the feature points contained in the images are in the same area.

[0266] The server not only encodes two-dimensional moving images, but can also automatically encode still images or at user-specified times based on scene analysis of the moving images and transmit them to the receiving terminal. Furthermore, given the relative positional relationships between camera terminals, the server can generate not only two-dimensional moving images, but also three-dimensional shapes of a scene from images taken from different angles of the same scene. Moreover, the server can further encode three-dimensional data generated using point clouds, and can select from or reconstruct images from multiple terminals based on the results of using 3D data to identify or track people or targets, and then generate images to be transmitted to the receiving terminal.

[0267] In this way, users can freely select images corresponding to each camera terminal to enjoy the scene, or enjoy content created by cutting out images from any viewpoint from three-dimensional data reconstructed from multiple images or videos. In addition, similar to the images, sound can also be picked up from multiple different angles, and the server can also coordinate with the images to multiplex and transmit sound and images from specific angles or spaces.

[0268] Furthermore, in recent years, content that establishes a correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has gradually become more widespread. In the case of VR images, the server can create separate viewpoint images for the right and left eyes, and use Multi-View Coding (MVC) to encode the images between the different viewpoints with permissible references, or encode them as different streams without mutual reference. When decoding the different streams, they can be synchronized for playback to recreate the virtual three-dimensional space according to the user's viewpoint.

[0269] In the case of AR images, the server overlays virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device can also acquire or retain virtual object information and 3D data, and generate a 2D image in response to the user's viewpoint movement, seamlessly connecting them to create overlay data. Alternatively, the decoding device can transmit the user's viewpoint movement to the server in addition to virtual object information, and the server can use this viewpoint movement received from the 3D data held on the server to create overlay data, encode it, and send it to the decoding device. Furthermore, the overlay data can have an α value for transparency in addition to RGB. The server sets the α value of the portion outside the target generated from the 3D data to 0, and encodes it while that portion is transparent. Alternatively, the server can set a predetermined RGB value as the background, such as using chroma keying, and generate data where the portion outside the target is the background color.

[0270] Similarly, the decoding of transmitted data can be performed on client terminals, on the server side, or distributed among them. For example, one terminal can temporarily send a receiving request to the server, while other terminals receive the requested content, decode it, and then transmit the decoded signal to a device with a display. By distributing the processing and selecting appropriate content without relying on the performance of the communicating terminals themselves, high-quality data can be played. Furthermore, as another example, large-format image data can be received using a TV, and segments of the image, such as individual tiles, can be decoded and displayed on the viewer's personal terminal. This allows for the sharing of the entire image and enables users to conveniently identify their assigned area or areas requiring more detailed examination.

[0271] Furthermore, the following scenarios can be envisioned in the future: regardless of whether indoors or outdoors, and in situations where multiple wireless communications of short, medium, or long ranges are available, content can be seamlessly received while switching appropriate data for the connected communication using transmission system specifications such as MPEG-DASH. This allows users to freely select and switch between their own terminals and decoding or display devices such as monitors installed indoors or outdoors. Furthermore, decoding can be performed simultaneously by switching between the terminal to be decoded and the terminal to be displayed, based on the user's location information. This also allows for the display of map information on the walls or ground of nearby buildings with embedded display devices while moving towards a destination. Moreover, it is possible to switch the bit rate of received data based on the ease of access to encoded data on the network by caching encoded data to servers that can be accessed from the receiving terminal in a short time, or copying it to edge servers within a content delivery server. [Adjustable Encoding]

[0272] Regarding content switching, the adjustable stream used for compression encoding, as shown in Figure 18 and applied in the aforementioned embodiments, will be explained. While it is acceptable for a server to have multiple streams with the same content but different qualities as individual streams, the feature of a time / space adjustable stream, achieved through layered encoding as illustrated, can also be utilized to switch the content composition. That is, by allowing the decoding side to determine which layer to decode based on intrinsic factors such as performance and extrinsic factors such as the state of the communication band, the decoding side can freely switch between decoding low-resolution and high-resolution content. For example, when watching a continuation of a video viewed on a mobile smartphone (ex115) at home using an internet TV device, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.

[0273] Furthermore, as mentioned above, in addition to encoding the image layer by layer and implementing scalability with an enhancement layer above the base layer, the enhancement layer can also include meta-information such as statistical information about the image. The decoding side then uses this meta-information to perform super-resolution on the base layer image, thereby generating high-quality content. Super-resolution can be either an increase in the signal-to-noise ratio (SN) at the same resolution or an increase in resolution. Meta-information includes information about linear or nonlinear filter coefficients used in a specific super-resolution process, or information about parameter values ​​in filtering, machine learning, or least squares operations used in a specific super-resolution process.

[0274] Alternatively, the image can be segmented into tiles based on the meaning of the target within the image, and the decoding side selects the tiles to be decoded, thereby decoding only a portion of the area. Furthermore, by storing the target's attributes (people, vehicles, balls, etc.) and its position within the image (coordinates within the same image, etc.) as metadata, the decoding side can specify the desired target's location based on the metadata and determine the tile containing that target. For example, as shown in Figure 19, the metadata can be stored using a different data storage structure than pixel data, such as SEI information in HEVC. This metadata represents, for example, the position, size, or color of the main target.

[0275] Alternatively, metadata can be stored in units consisting of multiple images, such as streams, sequences, or random access units. This allows the decoder to obtain information such as the time a specific person appears in the image and compare it with the information in the image unit, thereby identifying the image containing the specific target and its position within the image. [Webpage Optimization]

[0276] Figure 20 shows an example of a webpage display on a computer such as ex111. Figure 21 shows an example of a webpage display on a smartphone such as ex115. As shown in Figures 20 and 21, when a webpage contains multiple links to image content (i.e., linked images), its appearance will vary depending on the viewing device. When multiple linked images are visible on the screen, until the user explicitly selects a linked image, or the linked image is near the center of the screen, or the entire linked image enters the screen, the display device (decoding device) displays static images or I-pictures (Intra Picture / I-Picture) with each content as linked images, or displays images in the form of GIF animations as multiple static images or I-pictures, or only receives the basic layer to decode and display the image.

[0277] When a user has already selected a linked image, the display device will prioritize decoding the base layer. Furthermore, as long as there is information in the HTML that constitutes the webpage that is adjustable content, the display device can also decode up to the enhancement layer. Also, to ensure real-time performance, before selection or when communication bandwidth is very limited, the display device can reduce the delay between the decoding and display of the initial image (the delay from the start of content decoding to the start of display) by decoding and displaying only forward reference images (I-images, P-images (predictive pictures), and only forward reference B-images (bidirectionally predictive pictures)). Additionally, the display device can also intentionally ignore the image reference relationships and set all B-images and P-images as forward references for coarse decoding, and then perform normal decoding by adding more received images over time. [Autonomous Driving]

[0278] Furthermore, when transmitting or receiving static map or image data such as two-dimensional or three-dimensional map information for the purpose of autonomous driving or driving support of automobiles, the receiving terminal can also receive information such as weather or construction as metadata in addition to image data belonging to more than one layer, and decode accordingly. Moreover, metadata can also belong to layers, or it can simply be multiplexed with image data.

[0279] At this point, since the vehicle, drone, or airplane containing the receiving terminal is mobile, seamless reception and decoding can be achieved while switching between base stations ex106~ex110 by transmitting the receiving terminal's location information when a reception request is made. Furthermore, the receiving terminal can dynamically switch the level of metadata reception and map information updates based on user selection, user status, and communication frequency band conditions.

[0280] As mentioned above, in the content delivery system ex100, the client can receive the encoded information transmitted by the user in real time and decode and play it. [Sending Personal Content]

[0281] Furthermore, in the content delivery system ex100, not only can high-quality, long-duration content from video providers be sent via unicast or multicast, but also low-quality, short-duration content from individuals. Moreover, this type of personal content is expected to continue to increase in the future. To further enhance the quality of personal content, the server can perform encoding processing after editing. This can be achieved through, for example, the following configuration.

[0282] The server performs real-time or cumulative processing during and after shooting, identifying and processing photographic errors, scene searching, meaning analysis, and target detection from the original image or encoded data. Then, based on the identification results, the server performs manual or automatic editing as follows: correcting out-of-focus or camera shake, deleting less important scenes such as those with lower brightness or out-of-focus areas, emphasizing target edges, and adjusting tones. The server encodes the edited data based on the editing results. Furthermore, it is well known that longer shooting times can lead to lower viewership; therefore, the server automatically edits not only less important scenes but also scenes with less dynamic content, based on image processing results, to ensure the content is within a specific time frame appropriate for the shooting time. Alternatively, the server can generate and encode a digest based on the meaning analysis results of the scene.

[0283] Furthermore, in personal content, there are cases where including content as is would infringe on copyright, moral rights, or portrait rights. There are also situations where sharing exceeds the intended scope, causing inconvenience for the individual. Therefore, it's possible to, for example, have the server intentionally change the faces of people at the periphery of the image, or the interior of a house, into out-of-focus images and encode them. Additionally, the server can identify whether a face different from a pre-registered person is captured in the encoded image, and if so, blur the face. Alternatively, as pre- or post-processing before encoding, from a copyright perspective, it's feasible to allow users to specify the people or background areas they want to process, and then have the server replace the specified areas with other images or blur the focus. If it's a person, the face can be replaced while tracking the person in a moving image.

[0284] Furthermore, since personal content with smaller data volumes requires greater real-time performance, the decoding device prioritizes receiving and decoding the base layer, even though bandwidth also plays a role. The decoding device can also receive enhancement layers during this period, and in cases of loop playback or playback more than twice, the enhancement layers are included to play high-definition video. Any stream with adjustable encoding like this can provide an experience where, although the initial view is a coarse animation, the stream gradually becomes more refined (smart), improving the image quality. Besides adjustable encoding, the same experience can be provided even by combining a coarse stream from the first playback with a second stream encoded based on the first animation. [Other usage examples]

[0285] Furthermore, these encoding or decoding processes are generally handled within the LSIex500 chip present in each terminal. The LSIex500 can be a single chip or a combination of multiple chips. Additionally, software for encoding or decoding motion graphics can be installed on a recording medium (CD-ROM, flexible disk, or hard disk, etc.) that can be read by a computer such as the ex111, and the software can then be used for encoding or decoding. Moreover, if the ex115 smartphone has a camera, motion graphics data captured by that camera can also be transmitted. In this case, the motion graphics data is encoded using the LSIex500 chip present in the ex115 smartphone.

[0286] Furthermore, the LSIex500 can also be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or whether it has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal will download the codec or application software, and then obtain and play the content.

[0287] Furthermore, not limited to the content delivery system ex100 via the Internet ex101, at least one of the aforementioned embodiments of a motion graphics encoding device (image encoding device) or motion graphics decoding device (image decoding device) can also be installed in a digital playback system. Since multiplexed data with multiplexed image and sound is carried on the radio waves of the playback system using satellites or the like for transmission and reception, there is a difference that the configuration of the content delivery system ex100, which is more prone to unicast, is more suitable for multicast. However, the encoding and decoding processes can still be applied in the same way. [Hardware Configuration]

[0288] Figure 22 shows a diagram of the smartphone ex115. Figure 23 shows a configuration example of the smartphone ex115. The smartphone ex115 includes: an antenna ex450 for transmitting and receiving radio waves between itself and a base station ex110; a camera unit ex465 for capturing images and still images; and a display unit ex458 for displaying decoded data, including images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 further includes: a touch panel (operation unit ex466), a speaker (sound output unit ex457), a microphone (sound input unit ex456), a memory unit ex467 for storing captured images or still images, recorded audio, received images or still images, emails, and other encoded or decoded data, and a slot unit ex464 serving as an interface with a SIM ex468 for authentication of various data access, primarily via a network. Furthermore, external memory can be used instead of the memory unit ex467.

[0289] Furthermore, the main control unit ex460, which integrates the control of the display unit ex458 and the operation unit ex466, is connected to the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / splitting unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.

[0290] When the power button is turned on by the user, the power circuit section ex461 will start the smartphone ex115 into an operational state by supplying power to various parts from the battery pack.

[0291] The smartphone ex115 processes calls and data communications under the control of the main control unit ex460, which includes a CPU, ROM, and RAM. During a call, the voice signal processing unit ex454 converts the audio signal received by the voice input unit ex456 into a digital audio signal, which is then spread-spectrum processed by the modulation / demodulation unit ex452. Next, the transmission / reception unit ex451 performs digital-to-analog conversion and frequency conversion before transmitting the signal through the antenna ex450. Furthermore, the received data is amplified and undergoes frequency conversion and analog-to-digital conversion. The modulation / demodulation unit ex452 performs despread-spectrum processing, which is then converted back into an analog audio signal by the voice signal processing unit ex454 and output by the voice output unit ex457. In data communication mode, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 through the operation unit ex466 of the main unit, and the same transmission and reception processing is performed. When transmitting images, still images, or images and sound in data communication mode, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the motion image encoding method shown in the above embodiments, and sends the encoded image data to the multiplexing / splitting unit ex453. Furthermore, the sound signal processing unit ex454 encodes the sound signal received by the sound input unit ex456 in the images or still images captured by the camera unit ex465, and sends the encoded sound data to the multiplexing / splitting unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes encoded image data and encoded audio data in a predetermined manner, and performs modulation and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data through the antenna ex450.

[0292] When images attached to emails or online chats, or images linked to web pages, are received, in order to decode the multiplexed data received through antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into bit streams of image data and bit streams of audio data. Then, through synchronization bus ex470, the encoded image data is supplied to image signal processing unit ex455, and the encoded audio data is supplied to audio signal processing unit ex454. Image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method shown in the above embodiments, and displays the image or still image contained in the linked motion picture file from display unit ex458 through display control unit ex459. Audio signal processing unit ex454 decodes the audio signal and outputs audio from audio output unit ex457. Furthermore, given the widespread availability of real-time streaming, sound may be played in places where it is not appropriate to emit sound, depending on the user's circumstances. Therefore, ideally, the initial configuration should only play video data without any audio signal. Audio can also be played synchronously only when the user selects video data.

[0293] Furthermore, although the EX115 smartphone was used as an example here, the following three assembly forms can be considered as terminals: in addition to a transmit / receive terminal with both an encoder and a decoder, there is also a transmit terminal with only an encoder, and a receive terminal with only a decoder. Moreover, in digital broadcasting systems, although the description focuses on receiving or transmitting multiplexed data that has already multiplexed audio data within video data, multiplexed data can also multiplex text data related to the video in addition to audio data, and can also receive or transmit video data itself rather than multiplexed data.

[0294] Furthermore, although the ex460 main control unit, which includes the CPU, is designed to control encoding or decoding processing, most terminals also have a GPU. Therefore, it is possible to utilize the GPU's performance by using memory shared by the CPU and GPU, or by managing addresses in memory that can be shared, to process a wider area at once. This can shorten encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient when processing motion search, deblocking filter, SAO (Sample Adaptive Offset), conversion, and quantization, which is performed on a per-image basis instead of using the CPU. Industrial availability

[0295] This disclosure can be used in, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, or digital video cameras.

[0296] From the above discussion, it will be understood that the present invention can be embodied in various embodiments, including but not limited to the following:

[0297] Item 1: An encoding apparatus that uses inter-frame prediction to encode object blocks of an image, the encoding apparatus comprising: Processor; and Memory The aforementioned processor utilizes the aforementioned memory. Motion compensation is performed using the motion vector corresponding to each of the two reference images, thereby obtaining two predicted images from the aforementioned two reference images. Two gradient images corresponding to the two predicted images are obtained from the two reference images mentioned above. Using the sub-block units obtained by segmenting the aforementioned encoded object block, local motion estimates are derived from the aforementioned two prediction images and two gradient images. The final predicted image of the aforementioned coded object block is generated by using the aforementioned two predicted images, the aforementioned two gradient images, and the local motion estimation value of the aforementioned sub-block unit.

[0298] Item 2: The encoding apparatus as described in Item 1, wherein, in acquiring the aforementioned two prediction images, when acquiring prediction images with fractional pixel precision, in each of the aforementioned two reference images, pixels within the interpolation reference range surrounding the prediction block specified by the aforementioned motion vector are referenced, and the fractional pixel precision pixels are interpolated. The aforementioned interpolation reference range, in ordinary frame prediction without processing using the aforementioned local motion estimation value, is included in the ordinary reference range referenced for motion compensation of the aforementioned coded object block.

[0299] Item 3: The encoding device as described in Item 2, wherein the aforementioned interpolation reference range is consistent with the aforementioned general reference range.

[0300] Item 4: The encoding apparatus as described in Item 2, wherein in acquiring the aforementioned two gradient images, pixels referencing the gradient reference range surrounding the aforementioned prediction block are used in each of the aforementioned two reference images. The aforementioned gradient reference range is included in the aforementioned interpolation reference range.

[0301] Item 5: The encoding device as described in Item 4, wherein the aforementioned gradient reference range is consistent with the aforementioned interpolation reference range.

[0302] Item 6: The encoding apparatus as described in Item 1, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, the values ​​of a plurality of pixels contained in the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, are used in a weighted manner. Among the aforementioned plurality of pixels, the pixel located closer to the center of the aforementioned region has a greater weight.

[0303] Item 7: The encoding apparatus of Item 1, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, in addition to the pixels contained in the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, the pixels contained in other predicted sub-blocks adjacent to the aforementioned predicted sub-block are also referenced, and are contained in other predicted sub-blocks contained in the predicted block specified by the aforementioned motion vector.

[0304] Item 8: The encoding device of any of Items 1 to 7, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, only a portion of the pixels of the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, are referenced.

[0305] Item 9: The encoding device as described in Item 8, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, (i) Select a pixel pattern from a plurality of pixel patterns that represent a subset of the plurality of pixels contained in the aforementioned predicted sub-block, and that are distinct from each other. (ii) Referring to the pixels within the aforementioned predicted sub-block shown in the selected pixel pattern, the aforementioned local motion estimation value of the aforementioned sub-block is derived. The aforementioned processor further writes information that will display the selected pixel pattern into a bit stream.

[0306] Item 10: The encoding device as described in Item 8, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, (i) From a plurality of pixel patterns that represent a subset of the plurality of pixels contained in the aforementioned prediction sub-block, and which are distinct from each other, a pixel pattern is appropriately selected based on the aforementioned two prediction images. (ii) Refer to the pixels in the aforementioned predicted sub-block shown by the selected aforementioned pixel pattern to derive the aforementioned local motion estimation value of the aforementioned sub-block.

[0307] Item 11: An encoding method that uses inter-frame prediction to encode object blocks in an image, wherein the encoding method performs the following processing: Motion compensation is performed using the motion vector corresponding to each of the two reference images, thereby obtaining two predicted images from the aforementioned two reference images; Obtain two gradient images corresponding to the two predicted images from the two reference images mentioned above; Using the sub-block units obtained by segmenting the aforementioned encoded object block, local motion estimates are derived using the aforementioned two prediction images and two gradient images. The final predicted image of the aforementioned coded object block is generated by using the aforementioned two predicted images, the aforementioned two gradient images, and the local motion estimation value of the aforementioned sub-block unit.

[0308] Item 12: A decoding apparatus that uses inter-frame prediction to decode a decoding object block of an image, the decoding apparatus comprising: Processor; and Memory The aforementioned processor utilizes the aforementioned memory. Motion compensation is performed using the motion vector corresponding to each of the two reference images, thereby obtaining two predicted images from the aforementioned two reference images. Two gradient images corresponding to the two predicted images are obtained from the two reference images mentioned above. Using the sub-block units obtained by segmenting the aforementioned decoded object block, local motion estimates are derived using the aforementioned two prediction images and two gradient images. The final predicted image of the aforementioned decoded object block is generated by using the aforementioned two predicted images, the aforementioned two gradient images, and the aforementioned local motion estimation value of the sub-block unit.

[0309] Item 13: The decoding apparatus as described in Item 12, wherein, in acquiring the aforementioned two prediction images, when acquiring prediction images with fractional pixel precision, in each of the aforementioned two reference images, pixels within the interpolation reference range surrounding the prediction block specified by the aforementioned motion vector are referenced, and the fractional pixel precision pixels are interpolated. The aforementioned interpolation reference range, in normal frame prediction without processing using the aforementioned local motion estimation value, is included in the normal reference range referenced for motion compensation of the aforementioned decoded object block.

[0310] Item 14: The decoding device as described in Item 13, wherein the aforementioned interpolation reference range is consistent with the aforementioned general reference range.

[0311] Item 15: The decoding apparatus as described in Item 13, wherein in acquiring the aforementioned two gradient images, pixels referencing the gradient reference range surrounding the aforementioned prediction block are used in each of the aforementioned two reference images. The aforementioned gradient reference range is included in the aforementioned interpolation reference range.

[0312] Item 16: The decoding device as described in Item 15, wherein the aforementioned gradient reference range is consistent with the aforementioned interpolation reference range.

[0313] Item 17: The decoding apparatus as described in Item 12, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, the values ​​of a plurality of pixels contained in the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, are used in a weighted manner. Among the aforementioned plurality of pixels, the pixel located closer to the center of the aforementioned region has a greater weight.

[0314] Item 18: The decoding apparatus of Item 12, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, in addition to the pixels contained in the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, the pixels contained in other predicted sub-blocks adjacent to the aforementioned predicted sub-block are also referenced, and are contained in other predicted sub-blocks contained in the predicted sub-block specified by the aforementioned motion vector.

[0315] Item 19: The decoding apparatus of any of Items 12 to 18, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, only a portion of the pixels of the region corresponding to the aforementioned sub-block, i.e., the predicted sub-block, are referenced.

[0316] Item 20: The decoding apparatus as described in Item 19, wherein the aforementioned processor further obtains information on the display pixel pattern from the bit stream. In the derivation of the aforementioned local motion estimation value, it is in each of the aforementioned two reference images, (i) Based on the aforementioned information, select a pixel pattern from a plurality of pixel patterns that display a portion of the plurality of pixels contained in the aforementioned predicted sub-block, and which are all different from each other. (ii) Refer to the pixels in the aforementioned predicted sub-block shown by the selected aforementioned pixel pattern to derive the aforementioned local motion estimation value of the aforementioned sub-block.

[0317] Item 21: The decoding apparatus as described in Item 19, wherein in the derivation of the aforementioned local motion estimation value, in each of the aforementioned two reference images, (i) From a plurality of distinct pixel patterns that represent a subset of the plurality of pixels contained in the aforementioned prediction sub-block, a pixel pattern is appropriately selected based on the aforementioned two prediction images. (ii) Refer to the pixels in the aforementioned predicted sub-block shown by the selected aforementioned pixel pattern to derive the aforementioned local motion estimation value of the aforementioned sub-block.

[0318] Item 22: A decoding method that uses inter-frame prediction to decode the decoding object blocks of an image, wherein the decoding method performs the following processing: Motion compensation is performed using the motion vector corresponding to each of the two reference images, thereby obtaining two predicted images from the aforementioned two reference images; Obtain two gradient images corresponding to the two predicted images from the two reference images mentioned above; Using the sub-block units obtained by segmenting the aforementioned decoded object block, local motion estimates are derived using the aforementioned two prediction images and the aforementioned two gradient images; The final predicted image of the aforementioned decoded object block is generated by using the aforementioned two predicted images, the aforementioned two gradient images, and the aforementioned local motion estimation value of the sub-block unit.

[0319] Blocks 10-23 100: Encoding device 102: Segmentation 104: Subtraction Section 106: Conversion Section 108: Quantitative Department 110: Entropy Coding Department 112,204: Inverse Quantization Section 114,206: Inverter Unit 116,208: Addition Department 118,210: Block memory 120, 212: Loop Filter Section 122,214: Box memory 124,216: In-frame prediction section 126,218: Inter-frame prediction unit 128,220: Predictive Control Department 200: Decoding device 202: Entropy Decoding Department 1000: Current image 1001: Current block 1100: First reference image 1110: First motion vector 1120: Predicted Block 1 1121,1221: Predicted sub-blocks 1122, 1222: Top left pixel 1130, 1130A: First interpolation reference range 1131,1131A,1132,1132A,1231,1231A,1232,1232A: Reference range 1135, 1135A: First gradient reference range 1140: First Predicted Image 1150: First gradient image 1200: Second reference image 1210: Second motion vector 1220: 2nd Prediction Block 1230, 1230A: Second interpolation reference range 1235, 1235A: Second gradient reference range 1240: Second Predicted Image 1250: Second gradient image 1300: Local motion estimation value 1400: Final Predicted Image ex100: Content Supply System ex101: Internet ex102: Internet Service Provider ex103: Streaming Server ex104: Communication Network ex106, ex107, ex108, ex109, ex110: Base Station ex111: Computer ex112: Game console ex113: Camera ex114: Home Appliances ex115: Smartphone ex116: Satellite ex117: Airplane ex450: Antenna ex451: Transmitter / Receiver ex452: Modulation / Demodulation Unit ex453: Multiplexing / Separation Section ex454: Audio Signal Processing Unit EX455: Image Signal Processing Unit ex456: Audio Input Section ex457: Audio Output Section ex458: Display Section ex459: Display Control Unit EX460: Main Control Unit ex461: Power Supply Circuit Section ex462: Operation Input Control Unit ex463: Camera Interface ex464: Slot section ex465: Camera Department ex466: Operations Department ex467: Memory Section ex468:SIM ex470: Bus ex500:LSI MV0, MV1, MVx0, MVy0, MVx1, MVy1, v0, v1: Motion vectors Ref0, Ref1: Refer to the images S101~S111: Steps TD0,TD1,τ0,τ1: Distance in time

Claims

1. An encoding apparatus for encoding a current block of a current image, the encoding apparatus comprising: a processor; and a memory, wherein the processor utilizes the memory to: in a first frame prediction mode in which a first final predicted image is derived without using local motion estimation: obtain a first predicted image of the current block with fractional pixel precision by referring to pixels within a first general reference range in a first reference image to which the current block is referenced; obtain a second predicted image of the current block with fractional pixel precision by referring to pixels within a second general reference range in a second reference image to which the current block is referenced; and generate the first final predicted image of the current block using the first predicted image and the second predicted image; in a second frame prediction mode in which a second final predicted image is derived using local motion estimation: obtain a third predicted image of the current block with fractional pixel precision by referring to pixels within a first interpolation reference range in the first reference image; A fourth predicted image with fractional-pixel precision for the current block is obtained by referencing pixels within a second interpolation reference range in the aforementioned second reference image; a first gradient image is obtained, comprising first gradient values ​​of pixels contained in the aforementioned first reference image, the first gradient values ​​being generated based on the aforementioned pixels within the aforementioned first interpolation reference range; a second gradient image is obtained, comprising second gradient values ​​of pixels contained in the aforementioned second reference image, the second gradient values ​​being generated based on the aforementioned pixels within the aforementioned second interpolation reference range; local motion estimates for each sub-block are derived based on the aforementioned third and fourth predicted images and the aforementioned first and second gradient images, the sub-blocks being obtained by segmenting the aforementioned current block; and the aforementioned second final predicted image for the current block is generated using the aforementioned third and fourth predicted images and the aforementioned local motion estimates derived for each of the aforementioned sub-blocks. The aforementioned first and second interpolation reference ranges are respectively included in the aforementioned first and second normal reference ranges. The aforementioned second inter-frame prediction mode is performed using the pixel data of the aforementioned first and second normal reference ranges, without reading additional pixel data of the aforementioned first and second interpolation reference ranges that are newer than the aforementioned pixel data of the aforementioned first and second normal reference ranges. The aforementioned first gradient value and the aforementioned second gradient value represent horizontal gradients.

2. A decoding apparatus for decoding a current block of a current image, the decoding apparatus comprising: a processor; and a memory, wherein the processor utilizes the memory to: in a first frame prediction mode in which a first final predicted image is derived without using local motion estimation; obtain a first predicted image of the current block with fractional pixel precision by referring to pixels within a first general reference range in a first reference image to which the current block is referenced; obtain a second predicted image of the current block with fractional pixel precision by referring to pixels within a second general reference range in a second reference image to which the current block is referenced; and generate the first final predicted image of the current block using the first predicted image and the second predicted image; in a second frame prediction mode in which a second final predicted image is derived using local motion estimation; obtain a third predicted image of the current block with fractional pixel precision by referring to pixels within a first interpolation reference range in the first reference image; A fourth predicted image with fractional-pixel precision for the current block is obtained by referencing pixels within a second interpolation reference range in the aforementioned second reference image; a first gradient image is obtained, comprising first gradient values ​​of pixels contained in the aforementioned first reference image, the first gradient values ​​being generated based on the aforementioned pixels within the aforementioned first interpolation reference range; a second gradient image is obtained, comprising second gradient values ​​of pixels contained in the aforementioned second reference image, the second gradient values ​​being generated based on the aforementioned pixels within the aforementioned second interpolation reference range; local motion estimates for each sub-block are derived based on the aforementioned third and fourth predicted images and the aforementioned first and second gradient images, the sub-blocks being obtained by segmenting the aforementioned current block; and the aforementioned second final predicted image for the current block is generated using the aforementioned third and fourth predicted images and the aforementioned local motion estimates derived for each of the aforementioned sub-blocks. The aforementioned first and second interpolation reference ranges are respectively included in the aforementioned first and second normal reference ranges. The aforementioned second inter-frame prediction mode is performed using the pixel data of the aforementioned first and second normal reference ranges, without reading additional pixel data of the aforementioned first and second interpolation reference ranges that are newer than the aforementioned pixel data of the aforementioned first and second normal reference ranges. The aforementioned first gradient value and the aforementioned second gradient value represent horizontal gradients.

3. A non-transitory computer-readable recording medium for use with a computer, the recording medium recording a computer program that causes the computer to execute: In a first frame prediction mode that derives a first final predicted image without using local motion estimation: A first predicted image of the current block with fractional pixel precision is obtained by referring to pixels within a first general reference range in a first reference image referenced by the current block; A second predicted image of the current block with fractional pixel precision is obtained by referring to pixels within a second general reference range in a second reference image referenced by the current block; and the first predicted image and the second predicted image are used to generate the first final predicted image of the current block; In a second frame prediction mode that derives a second final predicted image using local motion estimation: A third predicted image of the current block with fractional pixel precision is obtained by referring to pixels within a first interpolation reference range in the first reference image; A fourth predicted image with fractional-pixel precision for the current block is obtained by referencing pixels within a second interpolation reference range in the aforementioned second reference image; a first gradient image is obtained, comprising first gradient values ​​of pixels contained in the aforementioned first reference image, the first gradient values ​​being generated based on the aforementioned pixels within the aforementioned first interpolation reference range; a second gradient image is obtained, comprising second gradient values ​​of pixels contained in the aforementioned second reference image, the second gradient values ​​being generated based on the aforementioned pixels within the aforementioned second interpolation reference range; local motion estimates for each sub-block are derived based on the aforementioned third and fourth predicted images and the aforementioned first and second gradient images, the sub-blocks being obtained by segmenting the aforementioned current block; and the aforementioned second final predicted image for the current block is generated using the aforementioned third and fourth predicted images and the aforementioned local motion estimates derived for each of the aforementioned sub-blocks. The aforementioned first and second interpolation reference ranges are respectively included in the aforementioned first and second normal reference ranges. The aforementioned second inter-frame prediction mode is performed using the pixel data of the aforementioned first and second normal reference ranges, without reading additional pixel data of the aforementioned first and second interpolation reference ranges that are newer than the aforementioned pixel data of the aforementioned first and second normal reference ranges. The aforementioned first gradient value and the aforementioned second gradient value represent horizontal gradients.