Encoding device, decoding device, bit stream generating device, and method for transmitting bit stream
Patent Information
- Application Number
- TW114143839
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-25
- Filing Date
- 2019-04-23
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2039-04-22
AI Technical Summary
Existing encoding methods, such as H.265/HEVC, face challenges in improving encoding efficiency, particularly in inter-frame prediction processing.
The proposed solution involves using a circuit and memory to import a reference motion vector and write a difference motion vector, which includes diagonal directions and distances, to enhance encoding and decoding processes.
This approach leads to more accurate motion vector derivation in inter-frame prediction, thereby improving coding efficiency.
Smart Images

Figure TWG2TB001909099_001 
Figure TWG2TB001909099_002 
Figure TWG2TB001909099_003
Abstract
Description
[Technical Field]
[0001] Field of the Invention This disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method. [Previous Technology]
[0002] Background of the Invention Previously, H.265 existed as a specification for encoding moving images. H.265 is also known as HEVC (High Efficiency Video Coding).
[0003] Preliminary Technology Documents Non-Patent Documents Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) [Summary of the Invention]
[0004] Summary of the Invention Problem to be Solved by the Invention In such an encoding and decoding method, it is hoped that the encoding efficiency can be improved.
[0005] The purpose of this disclosure is to provide an encoding device, decoding device, encoding method or decoding method that can improve encoding efficiency.
[0006] Means for solving the problem The present invention discloses a type of encoding device having a circuit and a memory. The circuit uses the memory to import a reference motion vector in the inter-frame prediction processing of the object block, and writes a difference (delta) motion vector, which displays one of a plurality of directions including a diagonal direction and the distance calculated from the reference motion vector, into the encoding signal. The object block is encoded using the difference motion vector and the reference motion vector.
[0007] A decoding device of this disclosure includes a circuit and a memory. The circuit uses the memory to import a reference motion vector in the inter-frame prediction processing of the object block, and analyzes a difference motion vector that displays one of a plurality of directions including a diagonal direction and a distance calculated from the reference motion vector. The circuit uses the difference motion vector and the reference motion vector to decode the object block.
[0008] Furthermore, such general or specific forms can be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs and recording media.
[0009] Further benefits and advantages provided by the disclosed embodiments are evident from the specification and drawings. Such benefits and advantages are sometimes brought about individually by the various embodiments or features of the specification and drawings, and it is not necessary to provide all of them in order to obtain more than one benefit and advantage.
[0010] Effects of the Invention This disclosure can provide an encoding device, decoding device, encoding method or decoding method that can improve encoding efficiency.
Implementation Method
[0053] For example, in one embodiment of the invention, an encoding device includes a circuit and a memory. The circuit uses the memory to import a reference motion vector in the inter-frame prediction processing of the object block, and writes a difference motion vector, which displays one of a plurality of directions including a diagonal direction and a distance calculated from the reference motion vector, into an encoding signal. The object block is encoded using the difference motion vector and the reference motion vector.
[0054] Thus, the coding apparatus derives a more accurate motion vector in the inter-frame prediction processing. This improves the coding efficiency of the inter-frame prediction processing. Therefore, the coding apparatus can improve coding efficiency.
[0055] Here, for example, the inter-frame prediction processing of the aforementioned object block is an inter-frame prediction processing using a merged mode, and the aforementioned reference motion vector is imported by selecting one of the candidates from a list of candidates for multiple motion vectors for the aforementioned object block.
[0056] Furthermore, for example, in one type of decoding device disclosed herein, the aforementioned circuit, in the aforementioned inter-frame prediction processing, predetermines the aforementioned plurality of directions, and according to the obtained prediction parameters, selects a group of mutually perpendicular first and second directions among the plurality of directions, including the aforementioned oblique direction, and selects either the aforementioned first direction or the aforementioned second direction in the previously selected group as the aforementioned one direction.
[0057] Furthermore, for example, in the aforementioned inter-frame prediction process, the aforementioned circuit selects a first motion vector from a list of candidates for a plurality of motion vectors for the aforementioned object block as the aforementioned reference motion vector, and uses the direction of the aforementioned first motion vector to derive a first direction and a second direction that are perpendicular to each other as the aforementioned plurality of directions, and takes one of the aforementioned first direction and the second direction as the aforementioned one direction to parse the aforementioned differential motion vector.
[0058] Furthermore, for example, in the aforementioned inter-frame prediction processing, the aforementioned circuit uses the direction of the aforementioned first movement vector to derive a direction that is approximately the same as the direction of the aforementioned first movement vector, namely the aforementioned first direction, and derives the aforementioned mutually perpendicular first direction and second direction.
[0059] For example, in the aforementioned inter-frame prediction process, the aforementioned circuit predetermines the aforementioned plurality of directions and derives the direction that is closest to the aforementioned first movement vector among the aforementioned plurality of directions as the aforementioned first direction.
[0060] Furthermore, for example, the circuit includes a memory, the circuit uses the memory to import a reference motion vector in the inter-frame prediction processing of the object block, and analyzes a difference motion vector that displays one of a plurality of directions including the diagonal direction and the distance calculated from the reference motion vector, and uses the difference motion vector and the reference motion vector to decode the object block.
[0061] Furthermore, one type of decoding apparatus disclosed herein is: the inter-frame prediction processing of the aforementioned object block is an inter-frame prediction processing using a merged mode, and the aforementioned reference motion vector is imported by selecting one of the candidates from a list of candidates for a plurality of motion vectors for the aforementioned object block.
[0062] Thus, the decoding device derives a more accurate motion vector in the inter-frame prediction processing. This improves the coding efficiency of the inter-frame prediction processing. Therefore, the decoding device can improve coding efficiency.
[0063] For example, in the aforementioned inter-frame prediction process, the aforementioned circuit predetermines the aforementioned plurality of directions, and selects a group of mutually perpendicular first and second directions among the plurality of directions including the aforementioned oblique direction, based on the obtained prediction parameters, and selects either the aforementioned first direction or the aforementioned second direction in the previously selected group as the aforementioned one direction.
[0064] Furthermore, for example, in the aforementioned inter-frame prediction process, the aforementioned circuit selects a first motion vector from a list of candidates for a plurality of motion vectors for the aforementioned object block as the aforementioned reference motion vector, and uses the direction of the aforementioned first motion vector to derive a first direction and a second direction that are perpendicular to each other as the aforementioned plurality of directions, and takes one of the aforementioned first direction and the second direction as the aforementioned one direction to resolve the aforementioned differential motion vector.
[0065] Furthermore, for example, in the aforementioned inter-frame prediction processing, the aforementioned circuit uses the direction of the aforementioned first movement vector to derive a direction that is approximately the same as the direction of the aforementioned first movement vector, namely the aforementioned first direction, and derives the aforementioned mutually perpendicular first direction and second direction.
[0066] For example, in the aforementioned inter-frame prediction process, the aforementioned circuit predetermines the aforementioned plurality of directions and derives the direction that is closest to the aforementioned first movement vector among the aforementioned plurality of directions as the aforementioned first direction.
[0067] Furthermore, for example, in the aforementioned inter-frame prediction process, the aforementioned circuit predetermines the aforementioned plurality of directions and selects one of the aforementioned plurality of directions by means of the flags included in the obtained prediction parameters.
[0068] Furthermore, for example, in one type of encoding method disclosed herein, a reference motion vector is imported in the inter-frame prediction processing of the object block, and a differential motion vector showing one of a plurality of directions including the diagonal direction and the distance calculated from the aforementioned reference motion vector is written into the encoding signal, and the aforementioned differential motion vector and the aforementioned reference motion vector are used to encode the aforementioned object block.
[0069] Thus, this coding method derives a more accurate motion vector in inter-frame prediction processing. This improves the coding efficiency of inter-frame prediction processing. Therefore, this coding method can improve coding efficiency.
[0070] For example, in one type of decoding method disclosed herein, a reference motion vector is imported into the inter-frame prediction processing of the object block, and a difference motion vector is parsed, which displays one of a plurality of directions including the diagonal direction and the distance calculated from the aforementioned reference motion vector. The aforementioned difference motion vector and the aforementioned reference motion vector are used to decode the aforementioned object block.
[0071] Thus, this decoding method derives a more accurate motion vector in inter-frame prediction processing. This improves the coding efficiency of inter-frame prediction processing. Therefore, this decoding method can improve coding efficiency.
[0072] Furthermore, such general or specific forms can be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs and recording media.
[0073] The following will explain the specific implementation details while referring to the diagram.
[0074] Furthermore, the embodiments described below are all general or specific examples. The values, shapes, materials, constituent elements, the arrangement and connection of constituent elements, steps, and the order of steps shown in the following embodiments are examples, and their purpose is not to limit the scope of the patent application. Also, any constituent element in the following embodiments that is not described in the independent claim representing the highest-level concept is treated as an arbitrary constituent element.
[0075] (Embodiment 1) First, an overview of Embodiment 1 will be provided as an example of an encoding and decoding apparatus for which the processing and / or configuration described in the various embodiments of this disclosure can be applied. However, Embodiment 1 is only one example of an encoding and decoding apparatus for which the processing and / or configuration described in the various embodiments of this disclosure can be applied, and the processing and / or configuration described in the various embodiments of this disclosure can also be implemented in encoding and decoding apparatuses different from Embodiment 1.
[0076] When applying the processing and / or configuration described in the various embodiments of this disclosure to Embodiment 1, any of the following may be performed: (1) For the encoding or decoding device of Embodiment 1, the constituent elements of the plurality of constituent elements constituting the encoding or decoding device that correspond to the constituent elements described in the various embodiments of this disclosure are replaced with the constituent elements described in the various embodiments of this disclosure; (2) For the encoding or decoding device of Embodiment 1, after arbitrarily changing the constituent elements of the plurality of constituent elements constituting the encoding or decoding device by adding, replacing, deleting, or otherwise altering the functions or processing to be implemented, the constituent elements corresponding to the constituent elements described in the various embodiments of this disclosure are replaced with the constituent elements described in the various embodiments of this disclosure; (3) (3) The method implemented by the encoding or decoding apparatus of Embodiment 1, after any changes such as replacement or deletion are made to the addition of processing and / or a part of the processing included in the plurality of processing, replaces the processing described in the various embodiments of this disclosure with the processing corresponding to the processing described in the various embodiments of this disclosure; (4) A part of the constituent elements constituting the encoding or decoding apparatus of Embodiment 1 is combined and implemented with the constituent elements described in the various embodiments of this disclosure, the constituent elements having a part of the functions of the constituent elements described in the various embodiments of this disclosure, or the constituent elements implementing a part of the processing of the constituent elements described in the various embodiments of this disclosure; (5) (5) Combining and implementing a component having a portion of the functions of a portion of the components constituting the encoding or decoding apparatus of Embodiment 1, or a component implementing a portion of the processing performed by a portion of the components constituting the encoding or decoding apparatus of Embodiment 1, with the components described in the various embodiments of this disclosure, a component having a portion of the functions of the components described in the various embodiments of this disclosure, or a component implementing a portion of the processing performed by the components described in the various embodiments of this disclosure; (6) Replacing the processing described in the various embodiments of this disclosure with the processing corresponding to the processing described in the various embodiments of this disclosure among the plurality of processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1; (7) Combining and implementing a portion of the processing among the plurality of processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1 with the processing described in the various embodiments of this disclosure. Furthermore, the implementation methods of the processing and / or configuration described in the various embodiments of this disclosure are not limited to the above examples. For example, it can be implemented in a device used for a different purpose than the motion picture / image encoding device or motion picture / image decoding device disclosed in Embodiment 1, or the processing and / or configuration described in each embodiment can be implemented alone. Furthermore, the processing and / or configuration described in different embodiments can also be combined and implemented.
[0077] [Overview of Encoding Device] First, an overview of the encoding device in Embodiment 1 will be described. Figure 1 is a block diagram showing the functional configuration of the encoding device 100 in Embodiment 1. The encoding device 100 is a motion picture / image encoding device that encodes motion pictures / images in block units.
[0078] As shown in FIG1, the encoding device 100 is an apparatus for encoding images in block units, comprising a segmentation unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128.
[0079] The encoding device 100 can be implemented using, for example, a general-purpose processor and memory. In this case, when the processor executes the software program stored in the memory, the processor functions as a segmentation unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128. Alternatively, the encoding device 100 can be implemented using one or more dedicated electronic circuits, and the aforementioned electronic circuits correspond to the segmentation unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.
[0080] The following describes the constituent elements included in the encoding device 100.
[0081] [Segmentation Unit] The segmentation unit 102 segments each image contained in the input dynamic image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the image into blocks of a fixed size (e.g., 128×128). These fixed-size blocks are sometimes called coding tree units (CTUs). Then, the segmentation unit 102 segments each fixed-size block into blocks of variable size (e.g., 64×64 or less) according to the recursive quadtree and / or binary tree block segmentation. These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs). Furthermore, in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the image can be processed as CUs, PUs, or TUs.
[0082] Figure 2 is a diagram showing an example of block partitioning in Embodiment 1. In Figure 2, solid lines represent block boundaries of quaternion tree block partitioning, and dashed lines represent block boundaries of binary tree block partitioning.
[0083] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into 4 square blocks of 64×64 (quadtree block partitioning).
[0084] The top left 64x64 block is further vertically divided into two rectangular 32x64 blocks, and the left 32x64 block is further vertically divided into two rectangular 16x64 blocks (binary tree block partitioning). As a result, the top left 64x64 block is divided into two 16x64 blocks 11 and 12, and a 32x64 block 13.
[0085] The upper right 64×64 block is horizontally divided into two rectangular 64×32 blocks 14 and 15 (binary tree block division).
[0086] The lower left 64×64 block is divided into four 32×32 square blocks (quaternion tree block division). Of the four 32×32 blocks, the upper left and lower right blocks are further divided. The upper left 32×32 block is vertically divided into two 16×32 rectangular blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the lower left 64×64 block is divided into 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.
[0087] The bottom right 64×64 block 23 is not divided.
[0088] As above, in Figure 2, block 10 is divided into 13 variable-size blocks 11 to 23 according to the recursive quadtree and binary tree block partitioning. Such partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.
[0089] Furthermore, as shown in Figure 2, although one block is divided into four or two blocks (quadtree or binary tree block division), the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree block division). This kind of division, which includes ternary tree blocks, is sometimes called MBT (multi-type tree) division.
[0090] [Subtraction Unit] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in units of blocks divided by the segmentation unit 102. That is, the subtraction unit 104 calculates the prediction error (also called residual) of the encoded target block (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.
[0091] The original signal is the input signal of the encoding device 100, and is a signal representing the images of each picture constituting the dynamic image (e.g., luminance signal and two chroma signals). In the following description, the signal representing the image is sometimes also referred to as a sample.
[0092] [Conversion Unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs a discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example, in advance.
[0093] Furthermore, the transformation unit 106 can also adaptively select a transformation type from a plurality of transformation types, and then use the transformation basis function corresponding to the selected transformation type to convert the prediction error into transformation coefficients. Such a transformation is sometimes called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0094] The multiple conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 shows a table of basis functions corresponding to each conversion type. In Figure 3, N represents the number of input pixels. When selecting a conversion type from these multiple conversion types, it may depend on, for example, the type of prediction (intra-frame prediction and inter-frame prediction), or on the intra-frame prediction mode.
[0095] This information indicating whether EMT or AMT is applied (e.g., referred to as the AMT flag) and information indicating the selected conversion type are signaled at the CU level. Furthermore, the signaling of such information is not limited to the CU level, but can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0096] Furthermore, the transformation unit 106 can also perform a re-transformation on the transformation coefficients (transformation results). Such a re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 re-transforms each sub-block (e.g., a 4×4 sub-block) contained in the block of transformation coefficients corresponding to the intra-frame prediction error. Information indicating whether to apply NSST and information about the transformation matrix used for NSST are signaled at the CU level. Moreover, the signaling of such information is not limited to the CU level, but can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0097] Here, Separable transformation refers to a method of performing multiple transformations by separating the input in each direction according to the number of dimensions, while Non-Separable transformation refers to a method of treating two or more dimensions as one dimension and performing a transformation at once when the input is multidimensional.
[0098] For example, the following example can be given as an example of a non-separable transformation: when the input is a 4×4 block, the aforementioned block is regarded as an array with 16 elements, and the aforementioned array is transformed with a 16×16 transformation matrix.
[0099] Similarly, the case of treating the 4×4 input block as an array with 16 elements and then performing multiple Givens rotations (Hypercube Givens Transform) on the aforementioned array is also a non-separable transformation example.
[0100] [Quantization Unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scan order and quantizes the conversion coefficients according to the quantization parameters (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the dequantization unit 112.
[0101] The predetermined order is the order used for quantization / dequantization of the conversion coefficients. For example, the predetermined scanning order is defined by ascending frequency (from low frequency to high frequency) or descending frequency (from high frequency to low frequency).
[0102] The quantization parameter is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter is increased, the quantization step size will also increase. In other words, if the value of the quantization parameter is increased, the quantization error will increase.
[0103] [Entropy Encoding Unit] The entropy encoding unit 110 performs variable-length encoding on the quantization coefficients input from the quantization unit 108, thereby generating an encoded signal (encoded bit stream). Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.
[0104] [Dequantization Unit] The dequantization unit 112 dequantizes the quantization coefficients input from the quantization unit 108. Specifically, the dequantization unit 112 dequantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the dequantization unit 112 outputs the dequantized conversion coefficients of the current block to the deconversion unit 114.
[0105] [Inverse Conversion Unit] The inverse conversion unit 114 inversely converts the conversion coefficients input from the inverse quantization unit 112, thereby restoring the prediction error. Specifically, the inverse conversion unit 114 restores the prediction error of the current block by performing an inverse conversion on the conversion coefficients corresponding to the conversion of the conversion unit 106. Then, the inverse conversion unit 114 outputs the restored prediction error to the addition unit 116.
[0106] Furthermore, since the restored prediction error loses information due to quantization, it will not be consistent with the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes quantization error.
[0107] [Addition Unit] The addition unit 116 reconstructs the current block by adding the prediction error input from the inversion conversion unit 114 to the prediction sample input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes also called a local decoding block.
[0108] [Block Memory] Block memory 118 is a memory unit used to store blocks that are referenced in intra-frame prediction and are blocks within the encoded object image (hereinafter referred to as the current image). Specifically, block memory 118 stores reconstructed blocks output from the arithmetic unit 116.
[0109] [Loop Filtering Unit] The loop filtering unit 120 applies loop filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering is a filter used within the coding loop (in-loop filter), including filters such as deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF).
[0110] In ALF, a least square error filter is applied to remove coding distortion. For example, for each 2×2 sub-block in the current block, one filter is applied, which is selected from a plurality of filters based on the direction and activity of the local gradient.
[0111] Specifically, firstly, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple groups (e.g., 15 or 25 groups). The classification of sub-blocks is based on the direction and activity of the gradient. For example, the classification value C (e.g., C = 5D + A) is calculated using the gradient direction value D (e.g., 0~2 or 0~4) and the gradient activity value A (e.g., 0~4). Then, the sub-blocks are classified into multiple groups (e.g., 15 or 25 groups) based on the classification value C.
[0112] The gradient direction value D is derived, for example, by comparing the gradients of a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by adding the gradients of a plurality of directions and then quantizing the result.
[0113] Based on such classification results, the filter to be used for the sub-block is determined from a plurality of filters.
[0114] As the filter shape used in ALF, a circularly symmetrical shape can be used, for example. Figures 4A to 4C are diagrams showing several examples of filter shapes used in ALF. Figure 4A shows a 5×5 diamond-shaped filter, Figure 4B shows a 7×7 diamond-shaped filter, and Figure 4C shows a 9×9 diamond-shaped filter. The information displaying the filter shape is signaled at the image level. Furthermore, the signaling of the information displaying the filter shape is not limited to the image level, but can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0115] The on / off state of ALF is determined at, for example, the picture level or the CU level. For example, for brightness, the decision to apply ALF is made at the CU level, while for chromatic aberration, the decision is made at the picture level. The information indicating whether ALF is on or off is signaled at the picture level or the CU level. Furthermore, the signaling of the information indicating whether ALF is on or off is not limited to the picture level or the CU level, but can also be at other levels (such as the sequence level, slice level, tile level, or CTU level).
[0116] The coefficient set of the selectable multiple filters (e.g., up to 15 or 25 filters) is signaled at the picture level. Furthermore, the signaling of the coefficient set is not limited to the picture level, but can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0117] [Frame Memory] Frame memory 122 is a memory unit used to store reference images used for inter-frame prediction, and is sometimes also called a frame buffer. Specifically, frame memory 122 stores reconstructed blocks filtered by loop filter unit 120.
[0118] [Intra-frame prediction unit] The intra-frame prediction unit 124 refers to the blocks in the current image stored in the block memory 118 to perform intra-frame prediction (also known as intra-frame prediction) of the current block, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 refers to the samples (e.g., luminance values, chrominance values) of the blocks adjacent to the current block to perform intra-frame prediction, thereby generating an intra-frame prediction signal, and outputting the intra-frame prediction signal to the prediction control unit 128.
[0119] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of predefined intra-prediction modes. The plurality of intra-prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0120] One or more non-directional prediction modes include, for example, Planar prediction mode and DC prediction mode as specified in H.265 / HEVC (High-Efficiency Video Coding) specification (Non-Patent Document 1).
[0121] A plurality of directional prediction modes includes, for example, the 33 directional prediction modes specified in the H.265 / HEVC specification. Furthermore, in addition to the 33 directions, the plurality of directional prediction modes may further include 32 directional prediction modes (a total of 65 directional prediction modes). Figure 5A is a diagram showing 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. Solid arrows represent the 33 directions specified in the H.265 / HEVC specification, and dashed arrows represent the additional 32 directions.
[0122] Furthermore, in the intra-frame prediction of chromatic aberration blocks, luminance blocks can also be referenced. That is, the chromatic aberration component of the current block can also be predicted based on the luminance component of the current block. Such intra-frame prediction is sometimes called CCLM (cross-component linear model) prediction. This intra-frame prediction mode of chromatic aberration blocks that references luminance blocks (e.g., called CCLM mode) can also be added as one of the intra-frame prediction modes of chromatic aberration blocks.
[0123] The intra-prediction unit 124 can also correct the intra-predicted pixel values based on the gradients of reference pixels in the horizontal / vertical directions. Intra-prediction accompanied by this correction is sometimes called PDPC (position-dependent intra-prediction combination). Information indicating whether PDPC has been applied (e.g., a PDPC flag) is signaled at, for example, the CU level. Furthermore, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).
[0124] [Inter-frame prediction unit] The inter-frame prediction unit 126 refers to a reference image stored in the frame memory 122, which is different from the current image, to perform inter-frame prediction (also called inter-frame prediction) for the current block, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed on a unit basis, either the current block or a sub-block within the current block (e.g., a 4×4 block). For example, the inter-frame prediction unit 126 performs motion estimation within the reference image for the current block or sub-block. Then, the inter-frame prediction unit 126 uses the motion information (e.g., motion vector) obtained by motion estimation to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block. Then, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0125] Motion information used for motion compensation is signaled. The signaling of the motion vector can also be done using a motion vector predictor. That is, the difference between the motion vector and the motion vector predictor can also be signaled.
[0126] Furthermore, not only can the motion information of the current block obtained through motion estimation be used, but also the motion information of adjacent blocks can be used to generate inter-frame prediction signals. Specifically, the prediction signal based on the motion information obtained through motion estimation and the prediction signal based on the motion information of adjacent blocks can be weighted and added together to generate inter-frame prediction signals on a per-sub-block basis within the current block. Such inter-frame prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0127] In this OBMC mode, information displaying the size of sub-blocks used for OBMC (e.g., referred to as OBMC block size) is signaled at the sequence level. Furthermore, information displaying whether OBMC mode is applied (e.g., referred to as OBMC flag) is signaled at the CU level. Moreover, the signaling level of this information is not limited to the sequence level and CU level; it can also be other levels (e.g., image level, slice level, tile level, CTU level, or sub-block level).
[0128] A more detailed explanation of the OBMC mode. Figures 5B and 5C are flowcharts and conceptual diagrams illustrating the overview of the predicted image correction process using OBMC.
[0129] First, the motion vector (MV) assigned to the encoded object block is used to obtain the prediction image (Pred) with normal motion compensation.
[0130] Next, the movement vector (MV_L) of the left adjacent block that has been encoded is applied to the encoded target block to obtain the prediction image (Pred_L). The prediction image and Pred_L are then weighted and overlapped to perform the first correction of the prediction image.
[0131] Similarly, the movement vector (MV_U) of the upper adjacent block that has been encoded is applied to the encoded target block to obtain the prediction image (Pred_U). The prediction image that has undergone the first correction mentioned above is weighted and overlapped with Pred_U to perform the second correction of the prediction image, and it is used as the final prediction image.
[0132] Furthermore, although the two-stage correction method using the left adjacent block and the top adjacent block has been explained here, it can also be configured to use the right adjacent block or the bottom adjacent block to perform more than two stages of correction.
[0133] Furthermore, the overlapping area may not be the entire pixel area of the block, but only a part of the area near the block boundary.
[0134] Furthermore, although the prediction image correction process from one reference image has been explained here, the same applies to the correction of prediction images from multiple reference images. After obtaining the prediction images corrected from each reference image, the obtained prediction images are further superimposed to serve as the final prediction image.
[0135] Furthermore, the aforementioned processing target block can be a predicted block unit or a sub-block unit formed by further dividing the predicted block.
[0136] As a method for determining whether to apply OBMC processing, for example, there is a method using obmc_flag, where obmc_flag is a signal indicating whether to apply OBMC processing. Specifically, in the encoding device, it can be determined whether the area to be encoded belongs to a region with complex motion. If it does, the value of obmc_flag is set to 1, OBMC processing is applied, and encoding is performed. If it does not belong to a region with complex motion, the value of obmc_flag is set to 0, and OBMC processing is not applied, and encoding is performed. On the other hand, in the decoding device, the application of OBMC processing and decoding is switched according to the value of obmc_flag recorded in the stream.
[0137] Furthermore, motion information can also be exported without signaling at the decoding device side. For example, the merging mode specified in the H.265 / HEVC standard can also be used. Alternatively, motion estimation can be performed at the decoding device side to export motion information. In this case, motion estimation is performed without using the pixel values of the current block.
[0138] Here, the mode of motion estimation performed on the decoding device side will be explained. This mode of motion estimation performed on the decoding device side is sometimes called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0139] Figure 5D shows an example of FRUC processing. First, referencing the movement vectors of coded blocks spatially or temporally adjacent to the current block, a plurality of candidate lists (which can also be common to the merge list) are generated, each with a movement vector predictor. Second, the best candidate MV is selected from the plurality of candidate MVs registered in the candidate lists. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0140] Then, based on the selected candidate movement vector, the movement vector for the current block is derived. Specifically, for example, the selected candidate movement vector (best candidate MV) is directly exported as the movement vector for the current block. Alternatively, for example, pattern matching can be performed on the surrounding area of the position in the reference image corresponding to the selected candidate movement vector to derive the movement vector for the current block. That is, the surrounding area of the best candidate MV can also be estimated in the same way, and if there is an MV with a better evaluation value, the best candidate MV is updated to the aforementioned MV, and the aforementioned MV is used as the final MV of the current block. Furthermore, this process can also be omitted.
[0141] When processing on a sub-block basis, the exact same processing can also be performed.
[0142] Furthermore, the evaluation value is calculated by using style matching between the region within the reference image corresponding to the motion vector and the predetermined region to obtain the difference value of the reconstructed image. Furthermore, in addition to the difference value, information other than the difference value can also be used to calculate the evaluation value.
[0143] As a style matching method, either the first style matching or the second style matching can be used. The first style matching and the second style matching are sometimes referred to as bilateral matching and template matching, respectively.
[0144] The first style matching is performed between two blocks: two blocks in two different reference images and two blocks along the motion trajectory of the current block. Therefore, in the first style matching, the region in other reference images along the motion trajectory of the current block is used as the predetermined region for calculating the evaluation value of the above candidates.
[0145] Figure 6 illustrates an example of pattern matching (bidirectional matching) between two blocks along a movement track. As shown in Figure 6, in the first pattern matching, two movement vectors (MV0, MV1) are derived by estimating the most matching pair among two blocks in two different reference images (Ref0, Ref1) along the movement track of the current block. Specifically, the difference between the reconstructed image at a specified position in the first encoded reference image (Ref0) and the reconstructed image at a specified position in the second encoded reference image (Ref1) is derived for the current block, and the evaluation value is calculated using the obtained difference value. The first encoded reference image is the image specified by the candidate MV, and the second encoded reference image is the image specified by the symmetrical MV after scaling the candidate MV using the display time interval. The candidate MV with the best evaluation value is selected as the final MV from among the multiple candidate MVs.
[0146] Under the assumption of a continuous movement trajectory, the movement vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, if the current image is located between the two reference images in time, and if the temporal distances from the current image to the two reference images are equal, then in the first style matching, a mirror-symmetric bidirectional movement vector will be derived.
[0147] In the second style matching, style matching is performed between the template in the current image (the block in the current image that is adjacent to the current block (e.g., the top and / or left adjacent block)) and the block in the reference image. Therefore, in the second style matching, the block adjacent to the current block in the current image is used as the predetermined area for calculating the evaluation value of the above candidates.
[0148] Figure 7 is a diagram illustrating an example of style matching (template matching) between a template in the current image and a block in the reference image. As shown in Figure 7, in the second style matching, the movement vector of the current block is derived by estimating the block in the reference image (Ref0) that best matches the block adjacent to the current block in the current image (Cur block). Specifically, for the current block, the difference between the reconstructed image of the encoded regions of the left and top adjacent regions or one of them and the reconstructed image at the same position in the encoded reference image (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value among the multiple candidate MVs is selected as the best candidate MV.
[0149] Information indicating whether to apply this FRUC pattern (e.g., the FRUC flag) is signaled at the CU level. Furthermore, when an FRUC pattern is applied (e.g., when the FRUC flag is true), information indicating the method of pattern matching (first pattern matching or second pattern matching) (e.g., the FRUC pattern flag) is signaled at the CU level. Moreover, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, CTU level, or sub-block level).
[0150] Here, the mode for deriving the translation vector is explained based on the model assuming uniform linear motion. This mode is sometimes called the BIO (bi-directional optical flow) mode.
[0151] Figure 8 is a diagram used to illustrate the model assuming uniform linear motion. In Figure 8, (vx, vy) represents the velocity vector, and τ0 and τ1 represent the time distances between the current image (Cur Pic) and the two reference images (Ref0, Ref1), respectively. (MVx0, MVy0) represents the movement vector corresponding to the reference image Ref0, and (MVx1, MVy1) represents the movement vector corresponding to the reference image Ref1.
[0152] At this point, under the assumption that the velocity vector (vx,vy) is in constant linear motion, (MVx0,MVy0) and (MVx1,MVy1) are expressed as (vxτ0,vyτ0) and (-vxτ1,-vyτ1) respectively, and the following optical flow equation (1) holds. [Number 1]
[0153] Here, I(k) represents the brightness value of the reference image k (k=0,1) after motion compensation. The aforementioned optical flow equation indicates that the sum of (i), (ii), and (iii) equals zero, where (i) is the temporal derivative of the brightness value, (ii) is the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) is the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the aforementioned optical flow equation and Hermite interpolation, the movement vector of the block unit obtained from the merge list, etc., is corrected in pixels.
[0154] Furthermore, a different method can be used to derive the movement vector on the decoding device side than the method of deriving the movement vector based on a model that assumes uniform linear motion. For example, the movement vector can also be derived on a sub-block basis based on the movement vectors of a plurality of adjacent blocks.
[0155] Here, a mode for deriving movement vectors on a sub-block basis based on the movement vectors of a plurality of adjacent blocks is explained. This mode is sometimes called the affine motion compensation prediction mode.
[0156] Figure 9A is used to illustrate the derivation of the movement vector of a sub-block unit based on the movement vectors of a plurality of adjacent blocks. In Figure 9A, the current block contains 16 4×4 sub-blocks. Here, the movement vector v0 of the upper left corner control point of the current block is derived based on the movement vectors of the adjacent blocks, and the movement vector v1 of the upper right corner control point of the current block is derived based on the movement vectors of the adjacent sub-blocks. Then, using the two movement vectors v0 and v1, the movement vectors (vx, vy) of each sub-block within the current block are derived by the following equation (2). [Equation 2]
[0157] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents the weighting coefficient determined in advance.
[0158] Such an affine motion compensation prediction pattern may also include several patterns with different methods for deriving the movement vectors of the upper left and upper right control points. This indicates that the information of such an affine motion compensation prediction pattern (e.g., affine flags) is signaled at the CU level. Furthermore, this indicates that the signaling of the information of this affine motion compensation prediction pattern is not limited to the CU level, but can also be at other levels (e.g., sequence level, image level, slice level, tile level, CTU level, or sub-block level).
[0159] [Prediction Control Unit] The prediction control unit 128 selects either the intra-frame prediction signal or the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0160] Here, an example of deriving the motion vector of an encoded object image using a merging mode is explained. Figure 9B is a diagram illustrating the outline of the motion vector deriving process using a merging mode.
[0161] First, a list of candidate predicted MVs with a predicted MV is generated. The candidate predicted MVs are: spatial adjacency predicted MVs, which are the MVs of multiple encoded blocks that are spatially located around the encoded object block; temporal adjacency predicted MVs, which are the MVs of blocks near the location of the encoded object block projected in the encoded reference image; combined predicted MVs, which are the MVs generated by combining the MV values of spatial adjacency predicted MVs and temporal adjacency predicted MVs; and zero predicted MVs, which are MVs with a value of zero, etc.
[0162] Next, select one prediction MV from the plurality of prediction MVs registered in the prediction MV list and determine it as the MV of the encoded object block.
[0163] Further, in the variable length encoding section, the signal indicating which prediction MV has been selected, namely merge_idx, is recorded in the stream and encoded.
[0164] Furthermore, the number of predicted MVs registered in the predicted MV list as illustrated in Figure 9B may be different from the number shown in the figure, or it may be a composition that does not include a portion of the predicted MVs shown in the figure, or it may be a composition that includes predicted MVs other than the types shown in the figure.
[0165] Furthermore, the MV of the encoded object block derived by the merging mode can also be used to perform the DMVR processing described later, thereby determining the final MV.
[0166] Here, an example of using DMVR processing to determine MV will be explained.
[0167] Figure 9C is a conceptual diagram used to illustrate the outline of DMVR processing.
[0168] First, the best MVP set in the processing object block is taken as the candidate MV. Based on the aforementioned candidate MV, reference pixels are obtained from the processed image in the L0 direction, i.e. the first reference image, and the processed image in the L1 direction, i.e. the second reference image, respectively. The average of each reference pixel is used to generate a template.
[0169] Next, using the aforementioned template, the surrounding areas of the candidate MVs for the first and second reference images are estimated respectively, and the MV with the lowest cost is determined as the final MV. Furthermore, the cost value is calculated using the difference between each pixel value of the template and each pixel value of the estimated area, as well as the MV value, etc.
[0170] Furthermore, the processing generalities described herein are essentially the same in both the encoding and decoding devices.
[0171] Furthermore, as long as the processing can estimate the peripherals of the candidate MV and derive the final MV, other processing can be used instead of the processing already described here.
[0172] Here, the mode of generating a prediction image using LIC processing will be explained.
[0173] Figure 9D is a diagram illustrating the general outline of a predictive image generation method that uses brightness correction processing with LIC processing.
[0174] First, export the MV used to obtain the reference image corresponding to the encoded object block from the encoded image, i.e., the reference image.
[0175] Secondly, for the encoded object block, the brightness pixel values of the left and top adjacent encoded surrounding reference areas, and the brightness pixel values of the same position in the reference image specified by MV, are used to extract information indicating how the brightness value changes in the reference image and the encoded object image, and the brightness correction parameters are calculated.
[0176] The aforementioned brightness correction parameters are used to perform brightness correction processing on the reference image within the reference image specified by MV, thereby generating a prediction image for the encoded object block.
[0177] Furthermore, the shape of the surrounding reference area mentioned above in Figure 9D is one example, and other shapes may also be used.
[0178] Furthermore, although the process of generating a prediction image from a single reference image has been described here, when generating a prediction image from multiple reference images, the same method is used to perform brightness correction processing on the reference images obtained from each reference image to generate the prediction image.
[0179] As a method for determining whether to apply LIC processing, for example, there is a method using lic_flag, where lic_flag is a signal indicating whether to apply LIC processing. Specifically, in the encoding device, it can be determined whether the area to be encoded belongs to a region where brightness changes occur. If it does, the value of lic_flag is set to 1, LIC processing is applied, and encoding is performed. If it does not belong to a region where brightness changes occur, the value of lic_flag is set to 0, LIC processing is not applied, and encoding is performed. On the other hand, in the decoding device, the application of LIC processing and decoding is switched based on the value of lic_flag recorded in the stream.
[0180] As another method for determining whether to apply LIC processing, there is a method based on whether surrounding blocks have applied LIC processing. For a specific example, when the encoded object block is in merge mode, it is determined whether the surrounding encoded blocks selected during the MV export process have applied LIC processing and been encoded, and then the application of LIC processing and encoding is switched accordingly. Furthermore, in this example, the decoding process is exactly the same.
[0181] [Summary of Decoding Device] Next, an overview of a decoding device capable of decoding the encoded signal (encoded bit stream) output from the above-described encoding device 100 will be described. Figure 10 is a block diagram showing the functional configuration of the decoding device 200 in Embodiment 1. The decoding device 200 is a motion picture / image decoding device that decodes motion pictures / images in blocks.
[0182] As shown in FIG10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse conversion unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.
[0183] The decoding device 200 can be implemented using, for example, a general-purpose processor and memory. In this case, when the processor executes the software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse conversion unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. Alternatively, the decoding device 200 can be implemented using one or more dedicated electronic circuits, and the aforementioned electronic circuits correspond to the entropy decoding unit 202, the inverse quantization unit 204, the inverse conversion unit 206, the addition unit 208, the loop filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0184] The following describes the constituent elements included in the decoding device 200.
[0185] [Entropy Decoding Unit] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202 performs arithmetic decoding on the binary signal from the encoded bitstream, for example. Then, the entropy decoding unit 202 debinarizes the binary signal. Herein, the entropy decoding unit 202 outputs the quantization coefficients to the dequantization unit 204 in blocks.
[0186] [Dequantization Unit] The dequantization unit 204 dequantizes the quantization coefficients of the decoded target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the dequantization unit 204 dequantizes each quantization coefficient of the current block according to the quantization parameter corresponding to that quantization coefficient. Then, the dequantization unit 204 outputs the dequantized quantization coefficients (i.e., conversion coefficients) of the current block to the inverse conversion unit 206.
[0187] [Inverse Conversion Unit] The inverse conversion unit 206 inversely converts the conversion coefficients input from the inverse quantization unit 204, thereby restoring the prediction error.
[0188] For example, if the information interpreted from the encoded bit stream indicates the application of EMT or AMT (e.g., the AMT flag is true), the inverse conversion unit 206 inversely converts the conversion coefficient of the current block according to the interpreted information indicating the conversion type.
[0189] Furthermore, for example, if the information interpreted from the encoded bit stream indicates that NSST is applied, the inverse conversion unit 206 applies inverse reconversion to the conversion coefficients.
[0190] [Addition Unit] The addition unit 208 reconstructs the current block by adding the prediction error input from the inversion unit 206 to the prediction sample input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0191] [Block Memory] Block memory 210 is a memory unit used to store blocks that are referenced in intra-frame prediction and are blocks within the decoded target image (hereinafter referred to as the current image). Specifically, block memory 210 stores reconstructed blocks output from the arithmetic unit 208.
[0192] [Loop Filtering Unit] The loop filtering unit 212 applies loop filtering to the block reconstructed by the arithmetic unit 208 and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.
[0193] If the information indicating ALF on / off is interpreted from the encoded bit stream, and it indicates that ALF is on, then one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0194] [Frame Memory] Frame memory 214 is a memory unit used to store reference images used for inter-frame prediction, and is sometimes called a frame buffer. Specifically, frame memory 214 stores reconstructed blocks filtered by loop filter unit 212.
[0195] [Intra-frame prediction unit] The intra-frame prediction unit 216 performs intra-frame prediction based on the intra-frame prediction pattern interpreted from the coded bit stream, and references blocks within the current image stored in the block memory 210, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 216 performs intra-frame prediction by referencing samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, thereby generating an intra-frame prediction signal, and outputs the intra-frame prediction signal to the prediction control unit 220.
[0196] Furthermore, when the intra-prediction mode of the reference luminance block is selected in the intra-prediction of the chromatic difference block, the intra-prediction unit 216 can also predict the chromatic difference component of the current block based on the luminance component of the current block.
[0197] Furthermore, if the information interpreted from the encoded bit stream indicates the application of PDPC, the intra-prediction unit 216 corrects the pixel value after intra-prediction based on the gradient of the reference pixel in the horizontal / vertical direction.
[0198] [Inter-frame prediction unit] The inter-frame prediction unit 218 predicts the current block by referring to a reference image stored in the frame memory 214. The prediction is performed on a unit of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 218 uses motion information (e.g., motion vector) interpreted from the coded bit stream to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0199] Furthermore, if the information interpreted from the coded bit stream indicates that the OBMC mode is applied, the inter-frame prediction unit 218 can not only use the motion information of the current block obtained by motion estimation, but also use the motion information of adjacent blocks to generate an inter-frame prediction signal.
[0200] Furthermore, if the information interpreted from the coded bitstream indicates the application of the FRUC mode, the inter-frame prediction unit 218 performs motion estimation based on the pattern matching method (bidirectional matching or template matching) interpreted from the coded bitstream, thereby deriving motion information. Then, the inter-frame prediction unit 218 uses the derived motion information to perform motion compensation.
[0201] Furthermore, if the BIO mode is applied, the inter-frame prediction unit 218 derives the motion vector based on a model assuming constant-speed linear motion. Also, if the information interpreted from the encoded bitstream indicates that an affine motion compensation prediction mode is applied, the inter-frame prediction unit 218 derives the motion vector on a sub-block basis based on the motion vectors of a plurality of adjacent blocks.
[0202] [Prediction Control Unit] The prediction control unit 220 selects either the intra-frame prediction signal or the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208.
[0203] [First State Sample of Inter-Frame Prediction Processing] The first to fifth state samples of this embodiment describe the case where inter-frame prediction processing is performed by introducing motion vector derivation processing into the encoding device 100 and the decoding device 200. Furthermore, this technique is called UMVE (Ultimate Motion Vector Expression), and is sometimes also called MMVD (Merge mode Motion Vector Difference).
[0204] Although the following description uses the operation of the decoding device 200 as an example, the operation of the encoding device 100 is the same.
[0205] FIG11 is a flowchart showing an internal processing example of the decoding apparatus 200 of the first state sample of Embodiment 1. FIG12 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the first state sample of Embodiment 1.
[0206] First, as shown in FIG11, the inter-frame prediction unit 218 of the decoding device 200 analyzes one of the mutually perpendicular directions, a first direction and a second direction, and the difference motion vector calculated from the reference motion vector. (S1001) More specifically, the inter-frame prediction unit 218 analyzes the information interpreted from the coded bit stream to import the reference motion vector and analyzes the difference motion vector.
[0207] Here, the differential movement vector is information added to the reference movement vector, and is an extension of the reference movement vector. More specifically, the differential movement vector displays the difference between the movement vectors in one of a plurality of directions, including the diagonal direction, relative to the reference movement vector, and is displayed in terms of direction and magnitude (step). In this example, the differential movement vector is information presented by the distance calculated from the reference vector in one of two mutually perpendicular directions among a plurality of predetermined directions.
[0208] Furthermore, the reference motion vector may be, for example, one predicted MV selected from the list of predicted MVs using the merging mode. In other words, the inter-frame prediction processing in this state can also be inter-frame prediction processing using the merging mode. Moreover, the inter-frame prediction unit 218 may also select one motion vector candidate from the list (predicted MV list) displaying a plurality of motion vector candidates for the target block, thereby importing the motion vector as the reference.
[0209] In the example shown in Figure 12, it is assumed that the reference movement vector is located at the origin of the X and Y axes, and the first or second diagonal direction corresponds to a direction tilted 45 degrees towards the X and Y axes. In the example shown in Figure 12, the first direction includes not only the right diagonal direction (+45°) in the first quadrant, but also the direction extending towards the third quadrant (+225°). Similarly, the second direction includes not only the left diagonal direction (-45°) in the fourth quadrant, but also the direction extending towards the second quadrant (-225°). Furthermore, as shown in Figure 12, the first and second directions are perpendicular, and the first direction is along the diagonal of the X and Y axes. Moreover, the differential movement vector is a vector located in either the positive or negative direction of the first or second direction, and has a size (step) at the position indicated by the circle.
[0210] Furthermore, in this example, the inter-frame prediction unit 218 can obtain the direction and magnitude of the display differential motion vector from a predetermined list displaying a plurality of directions and a plurality of magnitudes by analyzing the direction and magnitude parameters. Moreover, the predetermined list displays a plurality of directions and a plurality of magnitudes when the reference motion vector is located at the origin of the X-axis and Y-axis. Regarding the direction parameters, different values can be assigned to the positive direction of the first direction, the negative direction of the first direction, the positive direction of the second direction, and the negative direction of the second direction.
[0211] Next, as shown in FIG11, the inter-frame prediction unit 218 uses at least the differential motion vector parsed in step S1001 to decode the current block (S1002). More specifically, the inter-frame prediction unit 218 uses the differential motion vector parsed in step S1001 and the reference motion vector to perform motion compensation, thereby decoding the current block (object block), that is, generating the prediction signal of the current block.
[0212] Furthermore, although the above description addresses the case where, for example, the MV of the decoded block in the reference image ahead is registered as a candidate for prediction MV in the prediction MV list, and one prediction MV in the prediction MV list is set as the reference vector for the current block in the decoding one-way prediction mode, it is not limited to this.
[0213] It can also be set as the current block for decoding bidirectional prediction mode. In this case, simply set the above differential movement vector as the first differential movement vector, and use the second differential movement vector, which is symmetrical to the first differential movement vector about the origin, to decode the current block of bidirectional prediction mode. For example, if the first differential movement vector for the first reference image in the L0 direction is [i, j], then the second differential movement vector for the second reference image in the L1 direction is [-i, -j].
[0214] [Second State of Inter-Frame Prediction Processing] In this state, the inter-frame prediction unit 218 of the decoding device 200 processes the resolution differential shift vector in two stages. Hereinafter, the differences from the first state will be explained.
[0215] FIG13 is a flowchart showing an internal processing example of the decoding apparatus 200 of the second state sample of Embodiment 1. FIG14 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the second state sample of Embodiment 1.
[0216] First, as shown in FIG13, the inter-frame prediction unit 218 of the decoding device 200 analyzes the prediction parameters and other parameters, and selects a group of mutually perpendicular directions, including the oblique direction, or a group of mutually perpendicular directions, including the horizontal direction, from a plurality of predetermined directions (S10011).
[0217] Here, FIG14 will be used for a detailed explanation. The inter-frame prediction unit 218 parses the prediction parameters and other parameters obtained from the information interpreted from the coded bit stream, and as a first stage, selects two directions from the four directions shown in FIG14. FIG14 shows the first direction, the second direction, the third direction, and the fourth direction as four directions. Furthermore, in the example shown in FIG14, the first direction includes not only the right diagonal direction (+45°) in the first quadrant, but also the direction extending towards the third quadrant (+225°). Similarly, the second direction includes not only the left diagonal direction (-45°) in the fourth quadrant, but also the direction extending towards the second quadrant (-225°). The third direction is shown as a horizontal direction including both positive and negative sides, and the fourth direction is shown as a vertical direction including both positive and negative sides.
[0218] In this case, the inter-frame prediction unit 218 selects a group of mutually perpendicular first and second directions from a plurality of predetermined directions based on the obtained prediction parameters.
[0219] Next, as shown in FIG13, the inter-frame prediction unit 218 selects one direction from the group selected in step S10011, and parses the selected direction and the difference in distance calculated from the reference movement vector (S10012). For example, the inter-frame prediction unit 218 may select one direction from the group selected in step S10011 by parsing the direction index and the magnitude index. Thus, the inter-frame prediction unit 218 selects either the first direction or the second direction in the group selected in step S10011 as one direction.
[0220] Here, FIG. 14 will be used for a detailed explanation. As the second stage, the inter-frame prediction unit 218 parses the values represented by the direction index and magnitude index of the parameters obtained from the information interpreted from the coded bit stream, and selects one direction from the two selected directions shown in FIG. 14, either positive or negative, and determines its magnitude. In this way, the inter-frame prediction unit 218 can obtain a difference motion vector showing the selected direction and the distance calculated from the motion vector used as a reference.
[0221] [Third-State Sample of Inter-Frame Prediction Processing] Figure 15 is a flowchart showing an internal processing example of the decoding apparatus 200 of Embodiment 1, which is the third-state sample. Figure 16 is a diagram showing an example of the prediction MV list of the third-state sample of Embodiment 1. Figure 17 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the third-state sample of Embodiment 1. Hereinafter, the explanation will focus on the differences from the first and second-state samples.
[0222] First, as shown in FIG15, the inter-frame prediction unit 218 of the decoding device 200 selects the first motion vector from the list of motion vectors (S10013). More specifically, the inter-frame prediction unit 218 selects the first motion vector as the reference motion vector from the list displaying a plurality of candidate motion vectors for the current block.
[0223] Here, the list of movement vectors, that is, a list showing multiple candidate movement vectors for the current block, is, for example, the list of predicted MVs used in inter-frame prediction processing using merge mode. In this case, the inter-frame prediction unit 218 can select, for example, the first movement vector shown in Figure 17(a) by parsing, for example, the parameter showing the candidate index (candidate IDX) as shown in Figure 16.
[0224] Next, the inter-frame prediction unit 218 uses at least the direction of the first movement vector selected in step S10013 to derive a first direction and a second direction that are perpendicular to each other (S10014). For example, as shown in FIG17(b), the derived first direction is a direction that is approximately the same as the direction of the first movement vector, and the second direction is a direction perpendicular to the first direction. In other words, the inter-frame prediction unit 218 uses the direction of the first movement vector selected in step S10013 to derive a direction that is approximately the same as the direction of the first movement vector, i.e., the first direction, and derives a first direction and a second direction that are perpendicular to each other. A detailed example of this derivation process will be described later.
[0225] Next, the inter-frame prediction unit 218 analyzes a difference motion vector that displays one of the first and second directions and the distance calculated from the first motion vector (S10015). In this example, the inter-frame prediction unit 218 designates one of the first and second directions as the "one direction" and analyzes the difference motion vector, wherein the difference motion vector displays the derived direction (either positive or negative) and the distance calculated from the first motion vector. The difference motion vector is presented using a direction index and a magnitude index.
[0226] Hereinafter, using Figures 17 to 19, a detailed example of the derivation process for a direction that is approximately the same as the direction of the first movement vector will be explained. Figure 18 is a diagram showing an example of a lookup table for the third state of Embodiment 1. Figure 19 is a table showing an example of the correspondence between the size index and the shift value for the third state of Embodiment 1 according to pixel precision.
[0227] First, in the first step of the export processing, the inter-frame prediction unit 218 uses the component values (x value, y value) of the selected first motion vector to determine the minimum value displayed as "min" and the maximum value displayed as "max". Here, in the example shown in FIG17(a), the component values of the selected first motion vector are [m, n]. For example, when the component values of the first motion vector are [100, 64], "min" = 64 and "max" = 100.
[0228] Next, in the second step of the export process, the inter-frame prediction unit 218 uses the ratio of the maximum and minimum values determined in the first step as a parameter to export the first integer value. For example, the inter-frame prediction unit 218 performs a division operation of 64*min / max, and exports the first value by calculating an integer close to the result of the operation. Furthermore, the inter-frame prediction unit 218 may also use a binary search method as a method to replace the division operation and export the first value. For example, if "max" = 100 and "min" = 64, the first value is exported as 40 after searching 6 times in the range [0, 64].
[0229] Next, in the third step of the export processing, the inter-frame prediction unit 218 can use the lookup table shown in FIG18 to determine the second and third values from the first value derived in the second step. For example, the inter-frame prediction unit 218 can use the lookup table shown in FIG18 to determine the second value = 434 and the third value = 271 from the first value = 40 derived in the second step. Furthermore, the values in the lookup table shown in FIG18 are generated based on the assumption that the precision of the motion vector is 16, that is, based on the assumption that the movement distance of 1 pixel is 16. Also, it is generated based on the range of the first value being 0 to 64. Therefore, the values in the lookup table shown in FIG18 may be different when the precision of the motion vector is different, or when the range of the first value is different.
[0230] Next, in the fourth step of the export processing, the inter-frame prediction unit 218 parses the size index of the differential shift vector and uses the table shown in FIG19 to determine the shift value. Then, the inter-frame prediction unit 218 uses the determined shift value to replace the coordinates shown by the second and third values obtained in the third step with the xy coordinates of FIG17(b). For example, the inter-frame prediction unit 218 can use the determined shift value to perform a shift operation as defined by Equation (3), update the second and third values obtained in the third step, thereby replacing the coordinates shown by the second and third values with the xy coordinates of FIG17(b). [Number 3]
[0231] For example, the size index resolved by the inter-frame prediction unit 218 is 1 (magnitude index=1), and as shown in Figure 19, the shift value is 6. Also, in the third step, the second value = 434 and the third value = 271 are obtained. In this case, the inter-frame prediction unit 218 can perform the shift operation defined by equation (3), updating the second value to (434+32)>>6=7, and updating the third value to (271+32)>>6=4.
[0232] Next, in the fifth step of the export processing, when max = abs(m), the inter-frame prediction unit 218 only needs to update the ratio of the second value to the third value (second value / third value) to have the same sign as the ratio of the component values of the first motion vector (x value / y value). In this case, the inter-frame prediction unit 218 can decide to display the coordinates of the first direction as [second value, third value]. On the other hand, when max = abs(m), the inter-frame prediction unit 218 only needs to update the ratio of the second value to the third value (second value / third value) to have the same sign as the inverse ratio of the component values of the first motion vector (y value / x value). In this case, the inter-frame prediction unit 218 can decide to display the coordinates of the first direction as [third value, second value].
[0233] In the example used in steps 1 to 4, the component value of the first movement vector is [100, 64], and max = abs(100), so the inter-frame prediction unit 218 can determine the coordinates of the first direction as [7, 4]. Therefore, the coordinates of the second direction are displayed as [-4, 7].
[0234] Furthermore, in step S10015, the inter-frame prediction unit 218 can also parse the size index and direction index obtained from the prediction parameters, thereby setting one of the first direction and the second direction, and either positive or negative, as the direction of the display differential movement vector, and obtaining its size. For example, when the inter-frame prediction unit 218 parses the size index = 1 and the direction index = 0 from the prediction parameters, the direction of the display differential movement vector is the first direction, and it is presented as [7, 4].
[0235] [Fourth State Sample of Inter-Frame Prediction Processing] In the third state sample, although the method of using the direction of the first movement vector to derive the direction that is approximately the same as the direction of the first movement vector as the first direction has been described, it is not limited to this. It is also possible to derive the direction that is closest to the direction of the first movement vector from among a plurality of predetermined directions as the first direction. Hereinafter, this case as the fourth state sample will be described.
[0236] FIG20 is a flowchart showing an internal processing example of the decoding apparatus 200 of the fourth state sample of Embodiment 1. FIG21 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the fourth state sample of Embodiment 1.
[0237] As shown in FIG20, the inter-frame prediction unit 218 of the decoding device 200 derives the direction closest to the direction of the first movement vector from a plurality of predetermined directions as the first direction among the first and second directions that are perpendicular to each other (S100141).
[0238] Here, Figure 21 is used for a detailed explanation.
[0239] Figure 21 shows a plurality of predetermined directions: a first direction, a second direction, a third direction, and a fourth direction. Furthermore, the first direction includes not only the right-hand diagonal direction (+45°) in the first quadrant but also the direction extending into the third quadrant (+225°). Similarly, the second direction includes not only the left-hand diagonal direction (-45°) in the fourth quadrant but also the direction extending into the second quadrant (-225°). The third direction is shown as a horizontal direction encompassing both positive and negative sides, and the fourth direction is shown as a vertical direction encompassing both positive and negative sides.
[0240] As shown in FIG21, the inter-frame prediction unit 218 derives the first direction and the second direction from a plurality of predetermined directions.
[0241] More specifically, firstly, the inter-frame prediction unit 218 selects the first motion vector as shown in FIG21(a). Next, the inter-frame prediction unit 218 determines which of a plurality of predetermined directions the direction of the first motion vector shown in FIG21(a) is closest to. For example, the inter-frame prediction unit 218 may also determine whether the ratio (m / n) of the component values of the first motion vector shown in FIG21(a) falls within a predetermined range (+45°±22.5°) including the first direction as shown in FIG21(b). That is, the inter-frame prediction unit 218 can determine whether it is closest to the first direction by determining whether (32 / 13)>(m / n)>(32 / 77).
[0242] Next, the inter-frame prediction unit 218 determines that the direction of the first motion vector is closest to the first direction when the above relationship is satisfied. That is, the inter-frame prediction unit 218 derives the direction closest to the direction of the first motion vector from a plurality of predetermined directions as the first direction.
[0243] Thus, the inter-frame prediction unit 218 can use the direction of the first movement vector to derive a first direction and a second direction that are perpendicular to each other. In this way, the inter-frame prediction unit 218 can derive a first direction and a second direction that are perpendicular to each other from a plurality of predetermined directions.
[0244] Furthermore, step S100141 conforms to other examples of the derivation processes in the first and second directions in step S10014 above. The processing from step S10015 onwards will be omitted from the description.
[0245] [Fifth State of Inter-Frame Prediction Processing] In this state, the inter-frame prediction unit 218 of the decoding device 200 performs processing to directly select one direction from a plurality of predetermined directions and resolves the differential shift vector. Hereinafter, the differences from the first to fourth states will be explained.
[0246] FIG22 is a flowchart showing an internal processing example of the decoding apparatus 200 of the fifth state sample of Embodiment 1. FIG23 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the fifth state sample of Embodiment 1.
[0247] First, as shown in FIG22, the inter-frame prediction unit 218 of the decoding device 200 will analyze the difference in distance between the selected direction from a plurality of predetermined directions including the oblique direction and the distance calculated from the reference movement vector (S2001).
[0248] In this example, firstly, the inter-frame prediction unit 218 parses the information interpreted from the coded bit stream and imports a motion vector as a reference. Next, the inter-frame prediction unit 218 evaluates all predetermined multiple directions, including their magnitudes (steps), and selects a motion vector in one direction.
[0249] Here, Figure 23 will be used for a specific explanation. Figure 23 shows a plurality of predetermined directions, including the first to eighth directions, which include the diagonal direction. In the example shown in Figure 23, the first direction can be represented as [0, 1], the second direction as [1, 1], the third direction as [1, 0], and the fourth direction as [1, -1]. Similarly, the fifth direction can be represented as [0, -1], the sixth direction as [-1, -1], the seventh direction as [-1, 0], and the eighth direction as [-1, 1].
[0250] The inter-frame prediction unit 218 can also select the direction at least by means of the flags included in the acquired prediction parameters. That is, the inter-frame prediction unit 218 can also parse the movement direction syntax that displays the selected direction in the acquired prediction parameters, thereby selecting the direction. Furthermore, the inter-frame prediction unit 218 can also parse the size index included in the acquired prediction parameters, thereby selecting the size of the movement vector in one direction as the differential movement vector from a predetermined list. Here, the predetermined list is a list that displays the relationship between the size index and the size of the movement vector in that direction, and is provided in the form of a table, for example, as shown in FIG19.
[0251] Next, as shown in FIG22, the inter-frame prediction unit 218 uses at least the differential motion vector parsed in step S2001 to decode the current block (S2002). More specifically, the inter-frame prediction unit 218 uses the differential motion vector parsed in step S2001 and the reference motion vector to perform motion compensation, thereby decoding the current block (object block), that is, generating the prediction signal of the current block.
[0252] [Effects of State 1 to State 5] Based on State 1 to State 5, by introducing motion vector derivation processing into the inter-frame prediction processing disclosed herein, motion compensation can be performed using motion vectors with higher precision than the reference motion vectors. This improves the coding efficiency of inter-frame prediction processing.
[0253] More specifically, according to the first to fifth state samples, the decoding device, etc., imports (selects) a reference motion vector in the inter-frame prediction processing and parses the differential motion vector, wherein the aforementioned differential motion vector represents the difference between the reference motion vector and the direction and step size. For example, in the inter-frame prediction processing using the merging mode, the decoding device, etc., selects one prediction MV from the prediction MV list, thereby importing a reference motion vector. Then, using encoding parameters, etc., the differential motion vector, which represents the difference between the prediction MV and the one prediction MV in terms of direction and step size, is parsed.
[0254] Thus, the decoding device can obtain the differential motion vector, which is an extension of the motion vector based on a prediction MV, and is the difference information between the motion vector with higher accuracy than the motion vector based on the reference and the motion vector based on the reference.
[0255] Then, the decoding device and the like use the differential movement vector obtained after parsing and the movement vector as a reference to perform movement compensation, thereby generating a prediction signal of the object block (decoding the object block).
[0256] Thus, the decoding device derives a more accurate motion vector in inter-frame prediction processing. This improves the coding efficiency of inter-frame prediction processing.
[0257] As described above, based on the first to fifth state samples, since a more accurate motion vector can be derived in the inter-frame prediction processing, it is possible to realize an encoding device, decoding device, encoding method or decoding method that can improve the coding efficiency of the inter-frame prediction processing.
[0258] Furthermore, all the constituent elements recorded in the first state sample are not always required, and only a portion of the constituent elements of the first state sample may be present.
[0259] [Installation Example of Encoding Device] FIG24 is a block diagram showing an installation example of the encoding device 100 of Embodiment 1. The encoding device 100 includes a circuit 160 and a memory 162. For example, a plurality of components of the encoding device 100 shown in FIG1 are mounted by the circuit 160 and the memory 162 shown in FIG24.
[0260] Circuit 160 is a circuit for performing information processing and is also a circuit capable of accessing memory 162. For example, circuit 160 is a dedicated or general-purpose electronic circuit for encoding moving images. Circuit 160 can also be a processor like a CPU. Furthermore, circuit 160 can also be an assembly of multiple electronic circuits. Furthermore, for example, circuit 160 can also perform the functions of multiple components of the encoding device 100 shown in FIG1, other than the component for storing information.
[0261] Memory 162 is a dedicated or general-purpose memory that stores information used by circuit 160 to encode dynamic images. Memory 162 can be an electronic circuit or connected to circuit 160. Memory 162 can also be contained within circuit 160. Memory 162 can also be a collection of multiple electronic circuits. Memory 162 can be a magnetic disk or optical disk, or it can be presented as a memory device or recording medium. Memory 162 can be either non-volatile memory or volatile memory.
[0262] For example, memory 162 may store both the motion image to be encoded and the bit string corresponding to the motion image to be encoded. Also, memory 162 may store the program used by circuit 160 to encode the motion image.
[0263] Furthermore, for example, memory 162 can also function as a component for storing information among the plurality of components of the encoding device 100 shown in FIG1, etc. Specifically, memory 162 can also function as block memory 118 and frame memory 122 shown in FIG1. More specifically, memory 162 can also store reconstructed blocks and reconstructed images, etc.
[0264] Furthermore, the encoding device 100 may not necessarily include all of the plurality of constituent elements shown in FIG1, etc., nor may it necessarily perform all of the aforementioned plurality of processes. A portion of the plurality of constituent elements shown in FIG1, etc., may be included in other devices, and a portion of the aforementioned plurality of processes may also be performed by other devices. Moreover, in the encoding device 100, by including a portion of the plurality of constituent elements shown in FIG1, etc., and performing a portion of the aforementioned plurality of processes, encoding efficiency is improved.
[0265] Hereinafter, an example of the operation of the encoding device 100 shown in FIG24 will be shown. In the following example of operation, FIG25 is a flowchart showing an example of the operation of the encoding device 100 shown in FIG24. For example, when the encoding device 100 shown in FIG24 encodes a moving image, it performs the operation shown in FIG25.
[0266] Specifically, the circuit 160 of the encoding device 100 uses the memory 162 to perform the following processing. That is, firstly, in the inter-frame prediction processing of the object block, the circuit 160 imports a reference motion vector (S311) and writes a difference motion vector, which displays one of a predetermined plurality of directions including the diagonal direction and the distance calculated from the reference motion vector, into the encoding signal (S312). Secondly, the circuit 160 uses the difference motion vector obtained in step S312 and the reference motion vector to encode the object block (S313).
[0267] In this way, the coding apparatus 100 derives a more accurate motion vector in the inter-frame prediction processing, thereby improving the coding efficiency of the inter-frame prediction processing. Therefore, the coding apparatus 100 can improve coding efficiency.
[0268] [Installation Example of Decoding Device] FIG26 is a block diagram showing an installation example of the decoding device 200 of Embodiment 1. The decoding device 200 includes a circuit 260 and a memory 262. For example, a plurality of components of the decoding device 200 shown in FIG10 are mounted by the circuit 260 and the memory 262 shown in FIG26.
[0269] Circuit 260 is a circuit for performing information processing and is also a circuit capable of accessing memory 262. For example, circuit 260 is a dedicated or general-purpose electronic circuit for decoding moving images. Circuit 260 can also be a processor like a CPU. Furthermore, circuit 260 can also be an assembly of multiple electronic circuits. Furthermore, for example, circuit 260 can also perform the functions of multiple components of the decoding device 200 shown in FIG10, other than the component used for storing information.
[0270] Memory 262 is a dedicated or general-purpose memory that stores information used by circuit 260 to decode moving images. Memory 262 can be an electronic circuit or connected to circuit 260. Memory 262 can also be contained within circuit 260. Memory 262 can also be a collection of multiple electronic circuits. Memory 262 can be a magnetic disk or optical disk, or it can be presented as a memory device or recording medium. Memory 262 can be either non-volatile memory or volatile memory.
[0271] For example, memory 262 may store both a bit string corresponding to an encoded moving image and a moving image corresponding to a decoded bit string. Furthermore, memory 262 may also store the program used by circuit 260 to decode the moving image.
[0272] Furthermore, for example, memory 262 can also function as one of the constituent elements of the decoding device 200 shown in FIG10, for storing information. Specifically, memory 262 can also function as block memory 210 and frame memory 214 shown in FIG10. More specifically, memory 262 can also store reconstructed blocks and reconstructed images, etc.
[0273] Furthermore, the decoding device 200 may not necessarily include all of the plurality of components shown in FIG10, etc., nor may it necessarily perform all of the aforementioned plurality of processes. A portion of the plurality of components shown in FIG10, etc., may be included in other devices, and a portion of the aforementioned plurality of processes may also be performed by other devices. Moreover, in the decoding device 200, by including a portion of the plurality of components shown in FIG10, etc., and performing a portion of the aforementioned plurality of processes, motion compensation can be performed efficiently.
[0274] Hereinafter, an example of the operation of the decoding device 200 shown in FIG26 will be shown. FIG27 is a flowchart showing an example of the operation of the decoding device 200 shown in FIG26. For example, when decoding a moving image, the decoding device 200 shown in FIG26 performs the operation shown in FIG27.
[0275] Specifically, the circuit 260 of the decoding device 200 uses the memory 262 to perform the following processing. That is, firstly, in the inter-frame prediction processing of the object block, the circuit 260 imports a reference motion vector (S411) and parses a difference motion vector that displays one of a plurality of directions, including the diagonal direction, and the distance calculated from the reference motion vector (S412). Secondly, the circuit 260 uses the memory 262 and the difference motion vector parsed in step S412 and the reference motion vector to decode the object block (S413).
[0276] In this way, the decoding device 200 derives a more accurate motion vector in the inter-frame prediction processing. This improves the coding efficiency of the inter-frame prediction processing. Therefore, the decoding device 200 can improve coding efficiency.
[0277] [Supplement] Furthermore, the encoding device 100 and the decoding device 200 of this embodiment can be used as an image encoding device and an image decoding device, respectively, or as a motion picture encoding device and a motion picture decoding device. Alternatively, the encoding device 100 and the decoding device 200 can be used as inter-frame prediction devices (inter-picture prediction devices), respectively.
[0278] That is, the encoding device 100 and the decoding device 200 may also correspond only to the intra-frame prediction unit (intra-frame prediction unit) 126 and the inter-frame prediction unit (inter-frame prediction unit) 218, respectively. Furthermore, other components such as the conversion unit 106 and the inverse conversion unit 206 may also be included in other devices.
[0279] Furthermore, in this embodiment, each component can be constructed by dedicated hardware or implemented by executing software programs suitable for each component. Each component can also be implemented by the following method: the program execution unit of the CPU or processor reads the software program recorded on the recording medium such as the hard disk or semiconductor memory and executes it.
[0280] Specifically, each of the encoding device 100 and the decoding device 200 may also include a processing circuitry and a memory device electrically connected to the processing circuitry and accessible from the processing circuitry. For example, the processing circuitry corresponds to circuitry 160 or 260, and the memory device corresponds to memory 162 or 262.
[0281] The processing circuit includes at least one of dedicated hardware and a program execution unit, and uses a memory device to perform processing. Furthermore, when the processing circuit includes a program execution unit, the memory device stores the software program executed by the program execution unit.
[0282] Here, the software that implements the encoding device 100 or decoding device 200 of this embodiment is the following program.
[0283] That is, the program can also enable the computer to perform the following encoding method: First, in the inter-frame prediction processing of the object block, a reference motion vector is imported, and a differential motion vector, which displays one of the multiple directions including the diagonal direction and the distance calculated from the aforementioned reference motion vector, is written into the encoding signal. Second, the aforementioned object block is encoded using the differential motion vector and the aforementioned reference motion vector.
[0284] Alternatively, the program may also cause the computer to perform the following decoding method: First, in the inter-frame prediction processing of the object block, import the reference motion vector, and parse the difference motion vector, which displays one of the multiple directions including the diagonal direction and the distance calculated from the aforementioned reference motion vector. Second, use the aforementioned difference motion vector and the aforementioned reference motion vector to decode the aforementioned object block.
[0285] Furthermore, each component can also be a circuit as described above. These circuits can form a single integrated circuit or be separate individual circuits. Moreover, each component can be implemented using a general-purpose processor or a dedicated processor.
[0286] Furthermore, the processing performed by a specific component can also be performed by other components. Furthermore, the order of processing can be changed, and multiple processes can be performed in parallel. Furthermore, the encoding and decoding device can also include an encoding device 100 and a decoding device 200.
[0287] The ordinal numbers 1 and 2 used in the description can also be replaced appropriately. Furthermore, new ordinal numbers can be assigned to constituent elements, etc., or ordinal numbers can be removed.
[0288] Although the forms of the encoding device 100 and the decoding device 200 have been described above according to the embodiments, the forms of the encoding device 100 and the decoding device 200 are not limited to this embodiment. As long as they do not depart from the spirit of this disclosure, various modifications that can be conceived by those skilled in the art of the present invention are applied to this embodiment, or forms constructed by combining the constituent elements of different embodiments, can be included within the scope of the forms of the encoding device 100 and the decoding device 200.
[0289] This sample may also be combined with at least a portion of other samples disclosed herein and implemented. Furthermore, a portion of the processing described in the flowchart of this sample, a portion of the device configuration, a portion of the syntax, etc., may also be combined with other samples and implemented.
[0290] (Embodiment 2) In the above embodiments, each functional block can typically be implemented using an MPU and memory. Furthermore, the processing performed by each functional block can typically be implemented by a program execution unit such as a processor reading and executing software (programs) recorded on a recording medium such as ROM. This software can be distributed by downloading, or recorded on a recording medium such as semiconductor memory and distributed. Alternatively, each functional block can, of course, be implemented using hardware (dedicated circuitry).
[0291] Furthermore, the processing described in each embodiment can be implemented by using a single device (system) for centralized processing, or by using multiple devices for distributed processing. Also, the processors executing the above program can be either singular or multiple. That is, centralized processing or distributed processing is possible.
[0292] The present disclosure is not limited to the above embodiments and various modifications are possible, which are also included within the scope of the present disclosure.
[0293] Further, examples of applications of the dynamic image encoding method (image encoding method) or dynamic image decoding method (image decoding method) shown in the above embodiments and systems using them will be described. The system is characterized by having: an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Other components of the system can be appropriately modified depending on the circumstances.
[0294] [Usage Example] Figure 28 is a diagram showing the overall structure of the content delivery system ex100 that implements the content publishing service. The area for providing communication services is divided into desired sizes, and fixed wireless stations, i.e., base stations ex106, ex107, ex108, ex109, and ex110, are set up in each cell.
[0295] In this content delivery system ex100, various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 through Internet service provider ex102 or communication network ex104 and base stations ex106~ex110. The content delivery system ex100 can also be connected by combining any of the above elements. The devices can also be directly or indirectly connected to each other through telephone networks or short-range wireless networks without using fixed wireless base stations ex106~ex110. Furthermore, streaming server ex103 is connected to various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 through Internet ex101, etc. Furthermore, the streaming server ex103 connects to terminals and other devices in the hotspot within the aircraft ex117 via satellite ex116.
[0296] Furthermore, wireless access points or hotspots can also be used to replace base stations ex106~ex110. Also, the streaming server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, and can also connect directly to the aircraft ex117 without going through the satellite ex116.
[0297] The camera ex113 is a digital camera or similar device capable of capturing still images and moving images. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that generally supports 2G, 3G, 3.9G, 4G, and the future 5G mobile communication systems.
[0298] Home Appliances EX118 refers to refrigerators, or machines included in a home fuel cell cogeneration system, etc.
[0299] In the content supply system ex100, a terminal with photography capabilities can be connected to the streaming server ex103 via a base station ex106, thereby enabling live streaming. During live streaming, the terminal (computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and terminal within an airplane ex117, etc.) performs the encoding processing described in the above embodiments on still images or moving images captured by the user using the terminal, and multiplexes the encoded image data with the corresponding encoded audio data, sending the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device as described in this disclosure.
[0300] On the other hand, the streaming server ex103 streams content data to requesting clients. Clients can be terminals such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or airplanes ex117 that can decode the encoded data. Each machine that receives the published data decodes and plays it. That is, each machine functions as an image decoding device as described in this disclosure.
[0301] [Distributed Processing] Furthermore, the streaming server ex103 can also be multiple servers or multiple computers, and data is processed, recorded, and published in a distributed manner. For example, the streaming server ex103 can also be implemented using a CDN (Contents Delivery Network), which connects numerous edge servers distributed around the world to each other via the Internet. In a CDN, physically close edge servers are dynamically allocated according to the client. Furthermore, by caching and publishing content on these edge servers, latency can be reduced. Moreover, in the event of an error or when the communication status changes due to increased traffic, high-speed and stable publishing can be achieved because multiple edge servers can be used for distributed processing, the publishing entity can be switched to other edge servers, or the network portion where the obstacle occurs can be bypassed to continue publishing.
[0302] Furthermore, the encoding processing of captured data can be performed not only on each terminal but also on the server side, or shared among different terminals. For example, encoding typically involves two processing loops. In the first loop, the complexity or amount of encoding of the image, measured in frames or scenes, is detected. In the second loop, processing is performed to maintain image quality and improve encoding efficiency. For instance, the terminal performs the first encoding process, while the server receiving the content performs the second encoding process. This reduces the processing load on each terminal and improves the quality and efficiency of the content. In this case, if there is a requirement for near-instantaneous reception and decoding, other terminals can also receive and play the data that has already been encoded by the terminal in the first iteration, thus enabling more flexible real-time publishing.
[0303] For other examples, cameras such as the ex113 extract features from images, compress the relevant feature data, and send it as metadata to a server. The server uses the features to determine the importance of objects and adjust the quantization precision accordingly, compressing the image based on its meaning. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server compresses it again. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed on the terminal, while more demanding encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding) can be performed on the server.
[0304] To further illustrate this point, in stadiums, shopping malls, or factories, there are sometimes multiple images of almost identical scenes captured by multiple terminals. In such cases, the encoding processing is distributed among the multiple terminals performing the photography, as well as other terminals not yet capturing images, and servers, using units such as GOP (Group of Picture), image units, or tile units derived from image segmentation. This reduces latency and achieves greater real-time performance.
[0305] Furthermore, since the multiple image data sets depict nearly identical scenes, the server can manage and / or instruct the references between the image data captured by each terminal. Alternatively, the server can receive encoded data from each terminal, change the reference relationships between the multiple data sets, or correct or replace the images themselves before re-encoding them. This allows for the generation of streams with improved quality and efficiency for each data set.
[0306] Furthermore, the server can also transcode the image data after changing its encoding method before publishing the image data. For example, the server can convert the encoding method of the MPEG system to the VP system, or convert H.264 to H.265.
[0307] As such, encoding processing can be performed via a terminal or one or more servers. Therefore, although the terms "server" or "terminal" are used below to refer to the subject of processing, some or all of the processing performed by the server can also be performed by the terminal, and some or all of the processing performed by the terminal can also be performed by the server. Furthermore, the same applies to decoding processing.
[0308] [3D, Multi-angle] In recent years, there has been a growing trend of integrating and utilizing images or videos of different scenes captured by multiple cameras such as the EX113 and / or smartphones such as the EX115, which are almost synchronized with each other, or images of the same scene captured from different angles. The images captured by each terminal are integrated based on the relative positional relationship between the terminals or the fact that the feature points contained in the images are consistent regions.
[0309] The server can not only encode 2D moving images, but also automatically or at user-specified times encode still images based on scene analysis of the moving images, and send them to the receiving terminal. Furthermore, when the server can obtain the relative positional relationship between the camera terminals, it can generate the 3D shape of the scene not only based on 2D moving images, but also based on images of the same scene taken from different angles. Moreover, the server can also encode 3D data generated by point clouds, etc., and can also use 3D data to identify or track people or objects, and based on the identification or tracking results, select or reconstruct images from images taken by multiple terminals to generate images to be sent to the receiving terminal.
[0310] Thus, users can arbitrarily select images corresponding to each camera terminal to appreciate the scene, and can also appreciate the content of images cut from any viewpoint from 3D data reconstructed using multiple images or images. Furthermore, similar to images, sound can also be picked up from multiple different angles, and the server, in conjunction with the images, multiplexes and sends sound and images from specific angles or spaces.
[0311] Furthermore, in recent years, content that aligns the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular. In the case of VR images, the server can create separate viewpoint images for the right and left eyes. Multi-View Coding (MVC) and other techniques are used to encode images that allow reference between different viewpoints, or they can be encoded into different streams without mutual reference. When decoding different streams, the virtual 3D space is recreated according to the user's viewpoint, allowing the streams to be synchronized and played.
[0312] In the case of AR images, the server overlays virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device can also acquire or retain virtual object information and 3D data, and generate a 2D image in response to the user's viewpoint movement, creating overlay data by smoothly combining these. Alternatively, in addition to requesting virtual object information, the decoding device can also send the user's viewpoint movement to the server. The server uses the 3D data held on the server to create overlay data in conjunction with the received viewpoint movement, encodes the overlay data, and publishes it to the decoding device. Furthermore, the overlay data can also have an α value representing penetration in addition to RGB. The server can set the α value of the parts other than the object created from the 3D data to 0, and encode the parts in a state of penetration. Alternatively, the server can set the background to a predetermined RGB value, similar to a chroma key, and generate data where the parts other than the object are set to the background color.
[0313] Similarly, the decoding of published data can be performed on the client side (each terminal), on the server side, or distributed among terminals. For example, one terminal can temporarily send a receiving request to the server, and other terminals can receive the content requested and decode it, sending a signal indicating completion to a device with a display. By distributing the processing and selecting appropriate content without relying on the performance of the communicating terminals themselves, high-quality data can be played. Furthermore, as another example, large-format image data can be received by a television or similar device, and a portion of the image, such as a segmented tile, can be decoded and displayed on the viewer's personal terminal. This allows for sharing of the overall image and identification of one's own area of responsibility or areas requiring more detailed examination.
[0314] Furthermore, it is anticipated that in the future, in situations where multiple short-range, medium-range, or long-range wireless communications can be used regardless of whether the communication is indoors or outdoors, content can be received seamlessly by utilizing publishing system standards such as MPEG-DASH, while switching appropriate data for the connected communications. This allows users to switch between decoding devices or display devices, such as monitors installed indoors or outdoors, without being limited to their own terminals. Furthermore, decoding can be performed while switching between the terminal to be decoded and the terminal to be displayed, based on location information, etc. This allows users to move towards their destination while simultaneously displaying map information on a part of the wall or floor of a building next to a displayable device. Additionally, the bit rate of the received data can be switched based on the ease of access to the encoded data on the network. The aforementioned ease of access to the encoded data refers to whether the encoded data is cached to a server that can be accessed briefly from the receiving terminal, or copied to an edge server in a content delivery service, etc.
[0315] [Adaptive Coding] Regarding content switching, an adaptive stream using the dynamic image coding method described in the above embodiments, as shown in Figure 29, will be used for compression coding. While it is acceptable for a server to have multiple streams with the same content but different qualities as individual streams, they can also be configured as layered encodings as shown in the figure to achieve temporal / spatial adaptability and utilize their characteristics to switch content. In other words, the decoding side determines which layer to decode based on intrinsic performance factors and extrinsic factors such as communication band status. This allows the decoding side to freely switch between low-resolution and high-resolution content during decoding. For example, after watching a video on a smartphone (ex115) while on the go, if you want to watch it on an internet TV device at home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0316] Furthermore, in addition to the above-described structure where images are encoded layer by layer and an enhancement layer exists above the base layer, the enhancement layer may contain basic metadata such as statistical information of the image, and the decoder may use this metadata to perform super-resolution on the base layer image, thereby generating high-quality content. Super-resolution can be an improvement in the SN ratio within the same resolution, or an increase in resolution. The metadata includes information on linear or nonlinear filter coefficients used in specific super-resolution processing, or information on parameter values in filtering, machine learning, or least squares operations used in specific super-resolution processing.
[0317] Alternatively, the structure can be as follows: the image is segmented into tiles according to the meaning of objects within the image, and the decoding side selects the tiles to be decoded, thereby decoding only a portion of the area. Furthermore, object attributes (people, vehicles, balls, etc.) and their positions within the image (coordinate positions within the same image, etc.) are stored as metadata. This allows the decoding side to specify the desired object's position based on the metadata and determine the tiles containing that object. For example, as shown in Figure 30, metadata is stored using a data storage structure different from pixel data, such as the SEI message in HEVC. The aforementioned metadata represents, for example, the position, size, or color of the main object.
[0318] Alternatively, metadata can be stored in units consisting of multiple images, such as streaming units, sequences, or random access units. In this way, the decoding side can obtain information such as the time when a specific person appears in the image, and can use the information of the image unit to identify the image in which the object exists and the position of the object within the image.
[0319] [Web Page Optimization] Figure 31 shows an example of a web page display on a computer such as ex111. Figure 32 shows an example of a web page display on a smartphone such as ex115. As shown in Figures 31 and 32, web pages sometimes contain multiple links to image content. Depending on the browsing device, the viewing result will be different. When multiple links are visible on the screen, until the user explicitly selects a link, or until the link is near the center of the screen or the entire link is within the screen, the display device (decoding device) displays still images or I-images of each content as link images, or displays images like GIF animations using multiple still images or I-images, or simply receives the base layer and decodes and displays the image.
[0320] When a user selects a linked image, the display device prioritizes decoding the base layer. Furthermore, when the HTML of a webpage contains information indicating it is adaptable content, the display device can also decode up to the enhancement layer. Also, to ensure real-time performance, before selection or when communication bandwidth is very limited, the display device only decodes and displays images that are referenced to the preceding image (I images, P images, and only B images that are referenced to the preceding image). This reduces the delay between the decoding and display of the initial image (the delay from the start of content decoding to the start of display). Alternatively, the display device can deliberately ignore the reference relationship of images, causing all B and P images to be referenced to the preceding image and decoded coarsely, then performing normal decoding as time passes and more images are received.
[0321] [Autonomous Driving] Furthermore, when transmitting or receiving still images or data such as 2D or 3D map information for the purpose of autonomous driving or supporting driving, the receiving terminal can receive not only image data belonging to one or more layers, but also metadata such as weather or construction information, and decode them accordingly. Moreover, metadata may belong to layers, or it may simply be multiplexed with image data.
[0322] In this case, since the vehicle, drone, or aircraft containing the receiving terminal is mobile, the receiving terminal can seamlessly receive and decode while switching base stations ex106~ex110 by sending its location information when a receiving request is made. Furthermore, the receiving terminal can dynamically switch the level of metadata reception or map information updates based on user selection, user status, or the status of the communication band.
[0323] As described above, in the content supply system ex100, the client can receive the encoded information sent by the user in real time, decode it and play it.
[0324] [Personal Content Publishing] Furthermore, in the content supply system ex100, not only high-quality, long-duration content from video publishers, but also low-quality, short-duration content from individuals can be unicast or multicast. It is foreseeable that such personal content will gradually increase in the future. To further enhance the quality of personal content, the server can also perform editing and encoding processes. This can be achieved through, for example, the following configuration.
[0325] The server performs real-time or cumulative processing during and after photography, identifying and processing photographic errors, scene estimation, meaning analysis, and object detection from the original image or encoded data. Then, based on the identification results, the server manually or automatically performs the following edits: correcting blur or camera shake, deleting scenes of low importance such as those with lower brightness or out of focus compared to other images, emphasizing object edges, and changing color tones. The server encodes the edited data based on the editing results. Furthermore, given that viewership decreases if the photography time is too long, the server can also automatically edit scenes with less dynamic content, in addition to those of low importance, based on the image processing results, to create content within a specific time frame to accommodate the photography time. Additionally, the server can generate and encode a digest based on the meaning analysis results of the scene.
[0326] Furthermore, there are cases where direct broadcasting of personal content may infringe on copyrights, moral rights, or portrait rights, and there are also situations where the scope of sharing exceeds the intended scope, which is inconvenient for the individual. Therefore, the server can, for example, intentionally change the focus of images such as faces or homes around the perimeter of the screen to out-of-focus images before re-encoding. Also, the server can identify whether the face displayed in the image being encoded is different from a pre-registered person, and if so, perform processing such as adding mosaic to the face. Furthermore, as pre-processing or post-processing for encoding, users can also specify the people or background areas in the image they want to process based on copyright and other considerations, and the server will perform processing such as replacing the specified area with another image or blurring the focus. If it is a person, the image of the face can be replaced while tracking the person in the dynamic image.
[0327] Furthermore, viewing personal content with smaller data volumes requires real-time processing. Therefore, although video bandwidth is also a factor, the decoding device prioritizes receiving the base layer for decoding and playback. The decoding device can also receive the enhancement layer during this period, and in cases of loop playback or playback more than twice, high-definition video including the enhancement layer is played. With adaptive encoding in this way, the following experience can be provided: while there is a rough animation at the beginning of viewing or before selection, the stream gradually becomes smarter and the image improves. In addition to adaptive encoding, combining the rough stream from the first playback with a second stream encoded based on the first animation into a single stream can also provide the same experience.
[0328] [Other Usage Examples] Furthermore, such encoding or decoding processing is generally performed in the LSIex500 of each terminal. The LSIex500 can be a single chip or composed of multiple chips. Alternatively, software for encoding or decoding motion images can be installed on a recording medium (CD-ROM, floppy disk, or hard disk, etc.) that can be read by a computer such as ex111, and the software can be used for encoding or decoding. Furthermore, when the smartphone ex115 is equipped with a camera, motion image data captured by the camera can also be sent. This motion image data is data that has been encoded in the LSIex500 of the smartphone ex115.
[0329] Furthermore, the LSIex500 can also be configured to download and activate application software. In this case, the terminal first determines whether it corresponds to the content's encoding method or whether it has the capability to execute a specific service. When the terminal does not correspond to the content's encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software, and then obtains and plays the content.
[0330] Furthermore, not only in the content delivery system ex100 via the Internet ex101, but also in the digital broadcasting system, at least one of the aforementioned embodiments of motion graphics encoding devices (image encoding devices) or motion graphics decoding devices (image decoding devices) can be installed. Since it uses satellites or other means to transmit and receive multiplexed data that has already been multiplexed for images and sound via radio waves, it is easier to perform unicasting than the content delivery system ex100. The difference is that it is suitable for multicasting, but the same applications can be performed for encoding and decoding processing.
[0331] [Hardware Configuration] Figure 33 shows a diagram of the smartphone ex115. Figure 34 shows a configuration example of the smartphone ex115. The smartphone ex115 includes: an antenna ex450 for transmitting and receiving radio waves between itself and the base station ex110; a camera unit ex465 for capturing images and still images; and a display unit ex458 for displaying decoded data such as images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 further includes: an operation unit ex466, which is a touch panel, etc.; a sound output unit ex457, which is a speaker for outputting sound or audio; a sound input unit ex456, which is a microphone for inputting sound, etc.; a memory unit ex467, which can store encoded or decoded data such as captured images or still images, recorded audio, received images or still images, and emails; and a slot unit ex464, which is the interface with SIMex468, which can be used to authenticate specific users for access to various data, including network data. Furthermore, an external memory unit ex467 can also be used to replace the memory unit ex467.
[0332] Furthermore, the main control unit ex460, power circuit unit ex461, operation input control unit ex462, image signal processing unit ex455, camera interface unit ex463, display control unit ex459, modulation / demodulation unit ex452, multiplexing / splitting unit ex453, audio signal processing unit ex454, slot unit ex464, and memory unit ex467 that coordinately control the display unit ex458 and operation unit ex466 are connected via bus ex470.
[0333] When the power button is turned on by the user, the power circuit section ex461 supplies power to each part from the battery pack, thereby starting the smartphone ex115 into an operational state.
[0334] The smartphone ex115 performs call and data communication processing under the control of the main control unit ex460, which includes a CPU, ROM, and RAM. During a call, the audio signal received by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. This signal is then spread-spectrum processed by the modulation / demodulation unit ex452, and digital-to-analog conversion and frequency conversion processing are applied by the transmitting / receiving unit ex451 before being transmitted through the antenna ex450. Furthermore, the received data is amplified, subjected to frequency conversion and analog-to-digital conversion processing, despread-spectrum processed by the modulation / demodulation unit ex452, converted into an analog audio signal by the audio signal processing unit ex454, and then output from the audio output unit ex457. In data communication mode, text, still images, or video data are sent to the main control unit ex460 via the operation unit ex466 of the main unit and the operation input control unit ex462, and the sending and receiving processes are performed in the same way. When sending images, still images, or images and sound in data communication mode, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the motion image encoding method shown in the above embodiments, and sends the encoded image data to the multiplexing / splitting unit ex453. In addition, the sound signal processing unit ex454 encodes the sound signal received by the sound input unit ex456 when the camera unit ex465 captures images or still images, and sends the encoded sound data to the multiplexing / splitting unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a predetermined manner, and then applies modulation and conversion processing to the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmitting / receiving unit ex451, and transmits them through the antenna ex450.
[0335] When receiving images attached to emails or chat rooms, or images linked to web pages, in order to decode the multiplexing data received through the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexing data into a bitstream of image data and a bitstream of audio data. The encoded image data is then supplied to the image signal processing unit ex455 via the synchronization bus ex470, and the encoded audio data is supplied to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method described in the above embodiments, and displays the image or still image contained in the linked motion picture file on the display unit ex458 via the display control unit ex459. Furthermore, the audio signal processing unit ex454 decodes the audio signal and outputs audio from the audio output unit ex457. Furthermore, given the widespread availability of live streaming, depending on the user's situation, the playback of audio may be deemed inappropriate from a social perspective. Therefore, as an initial configuration, it is advisable to configure the system to play only video signals and no audio signals. Alternatively, audio can be played only when the user performs actions such as clicking on video data.
[0336] Furthermore, although the smartphone ex115 is used as an example here, as a terminal, in addition to a transmitting and receiving terminal with both an encoder and a decoder, three installation forms can also be considered: a transmitting terminal with only an encoder and a receiving terminal with only a decoder. Furthermore, although the case of receiving or transmitting multiplexed data such as audio data within video data in a digital broadcasting system has been explained, in addition to audio data, multiplexed data can also include text data associated with the video, and the video data itself can also be received or transmitted, rather than being multiplexed data.
[0337] Furthermore, although the case of the CPU's main control unit (ex460) controlling encoding or decoding processing has been described, many terminals also have GPUs. Therefore, it is also possible to utilize the GPU's performance by using shared memory between the CPU and GPU, or memory with managed addresses for shared use, to process large areas at once. This can shorten encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to use the GPU instead of the CPU to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and conversion / quantization processing at the same time, such as on a per-image basis.
[0338] Industrial Applicability This disclosure can be used in, for example, television sets, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, video conferencing systems, or electronic mirrors. [Simplified Explanation of the Diagram]
[0011] Figure 1 is a block diagram showing the functional configuration of the encoding device in Embodiment 1.
[0012] Figure 2 is a diagram showing an example of block division in implementation form 1.
[0013] Figure 3 is a table showing the transformation basis functions corresponding to each transformation type.
[0014] Figure 4A is a diagram showing one example of the shape of the filter used in ALF.
[0015] Figure 4B is a diagram showing another example of the shape of the filter used in ALF.
[0016] Figure 4C is a diagram showing another example of the shape of the filter used in ALF.
[0017] Figure 5A is a diagram showing 67 intra-prediction modes in intra-frame prediction.
[0018] Figure 5B is a flowchart illustrating the outline of the prediction image correction process using OBMC.
[0019] Figure 5C is a conceptual diagram illustrating the outline of predictive image correction processing using OBMC.
[0020] Figure 5D is a diagram showing one example of FRUC.
[0021] Figure 6 is a diagram used to illustrate style matching (bidirectional matching) between two blocks along the movement track.
[0022] Figure 7 is a diagram used to illustrate style matching (template matching) between the template in the current image and the block in the reference image.
[0023] Figure 8 is a diagram used to illustrate the model assuming uniform linear motion.
[0024] Figure 9A is a diagram used to illustrate how to derive the movement vector of a sub-block unit based on the movement vectors of a plurality of adjacent blocks.
[0025] Figure 9B is a diagram illustrating the outline of the process derived using the move vector of the merge mode.
[0026] Figure 9C is a conceptual diagram used to illustrate the outline of DMVR processing.
[0027] Figure 9D is a diagram illustrating the general outline of a predictive image generation method that uses brightness correction processing with LIC processing.
[0028] Figure 10 is a block diagram showing the functional configuration of the decoding device in Embodiment 1.
[0029] Figure 11 is a flowchart showing an internal processing example of the decoding device in the first state of embodiment 1.
[0030] Figure 12 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the first state of implementation mode 1.
[0031] Figure 13 is a flowchart showing an internal processing example of the decoding device in the second state of embodiment 1.
[0032] Figure 14 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the second state of implementation mode 1.
[0033] Figure 15 is a flowchart showing an internal processing example of the decoding device in the third state of embodiment 1.
[0034] Figure 16 is a diagram showing one example of the predicted MV list for the third state of implementation mode 1.
[0035] Figure 17 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the third state of implementation mode 1.
[0036] Figure 18 is a diagram showing an example of a look-up table in the third state of embodiment 1.
[0037] Figure 19 is a table showing an example of the correspondence between the size index and the shift value of the corresponding pixel precision in the third state of embodiment 1.
[0038] Figure 20 is a flowchart showing an internal processing example of the decoding device in the fourth state of Embodiment 1.
[0039] Figure 21 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the fourth state of implementation mode 1.
[0040] Figure 22 is a flowchart showing an internal processing example of the decoding device in the fifth state of embodiment 1.
[0041] Figure 23 is an explanatory diagram of the differential shift vector used in the inter-frame prediction processing of the fifth state of implementation mode 1.
[0042] Figure 24 is a block diagram showing an installation example of the encoding device in Embodiment 1.
[0043] Figure 25 is a flowchart showing an example of the operation of the encoding device in Embodiment 1.
[0044] Figure 26 is a block diagram showing an installation example of the decoding device in Embodiment 1.
[0045] Figure 27 is a flowchart showing an example of the operation of the decoding device in Embodiment 1.
[0046] Figure 28 is a diagram of the overall structure of the content supply system that implements content publishing services.
[0047] Figure 29 is a diagram showing one example of the encoding construction in scalable encoding.
[0048] Figure 30 is a diagram showing one example of the coding structure in adaptive coding.
[0049] Figure 31 is an example of a webpage display screen.
[0050] Figure 32 is an example of a webpage display screen.
[0051] Figure 33 is a diagram showing an example of a smartphone.
[0052] Figure 34 is a block diagram showing an example of the structure of a smartphone.
Claims
1. An encoding device, comprising: Circuit; The aforementioned circuit uses the aforementioned memory to perform the following operations: selecting a reference movement vector from the candidates of movement vectors; Choose the direction of the differential movement vector from the different candidates, the candidates including (i) a first direction, (ii) a second direction perpendicular to the first direction, and (iii) a third direction perpendicular to the second direction; determine the size of the differential movement vector, the size of which depends on the precision of the differential movement vector; write an index into a bit stream, the index indicating the size; and use the differential movement vector to modify the movement vector used as a reference to encode the current block.
2. A decoding device, comprising: Circuit; The aforementioned circuit uses the aforementioned memory to perform the following operations: selecting a reference movement vector from the candidates of movement vectors; selecting the direction of the differential movement vector from the different candidates, the candidates including (i) a first direction, (ii) a second direction perpendicular to the first direction, and (iii) a third direction perpendicular to the second direction; determining the size of the aforementioned differential movement vector based on (i) the index stored in the bit stream and (ii) the precision of the aforementioned differential movement vector; and using the aforementioned differential movement vector to change the aforementioned reference movement vector to decode the current block.
3. A bitstream generation device, comprising: Circuit; The aforementioned circuitry uses the aforementioned memory to generate a bitstream containing an index indicating size. This index is used for decoding processing, which includes: selecting a reference movement vector from candidates of movement vectors; selecting the direction of a differential movement vector from different candidates, including (i) a first direction, (ii) a second direction perpendicular to the first direction, and (iii) a third direction perpendicular to the second direction; determining the size of the differential movement vector based on (i) the aforementioned index stored in the aforementioned bitstream and (ii) the precision of the aforementioned differential movement vector; and using the aforementioned differential movement vector to modify the reference movement vector to decode the current block.
4. A method for sending a bitstream, comprising: Choose the reference movement vector from the candidates for movement vectors; Choose the direction of the differential movement vector from the different candidates, the candidates including (i) a first direction, (ii) a second direction perpendicular to the first direction, and (iii) a third direction perpendicular to the second direction; determine the size of the differential movement vector, the size of which depends on the precision of the differential movement vector; write an index into the bit stream, the index indicating the size; use the differential movement vector to modify the movement vector used as a reference to encode the current block; and send the bit stream.
Citation Information
Patent Citations
Efficient rounding for deblocking
TW201709744A
Encoding device, decoding device, encoding method and decoding method
TW201811027A
Moving image decoding apparatus, moving image decoding method
TW201813394A
Piperazine derivatives having multimodal activity against pain
US20170001967A1
Method, application processor, and mobile terminal for processing reference image
US20170201751A1