Encoding device and decoding device
By employing predictive block interpolation and gradient block generation with motion compensation and padding techniques, the encoding and decoding devices address the need for improved compression efficiency and reduced processing load in video coding, specifically in inter-prediction functions.
Patent Information
- Application Number
- JP2024166362
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-13
- Filing Date
- 2024-09-25
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2038-11-30
AI Technical Summary
There is a demand for further improvements in compression efficiency and reduction in processing load in video encoding and decoding techniques, particularly in inter-prediction functions that build predictions for a current frame based on reference frames.
The encoding and decoding devices utilize a predictive block interpolation process, gradient block generation, motion compensation, and padding techniques to enhance inter-prediction, including processes like OBMC, DMVR, and interpolation to improve compression efficiency and reduce processing load.
These methods reduce processing load and enhance compression efficiency by improving encoding and decoding processes, accelerating speeds, and optimizing component selection in video coding systems.
Smart Images

Figure 0007753488000011 
Figure 0007753488000012 
Figure 0007753488000013
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to video coding, and more particularly to video encoding and decoding systems, components, and methods that perform inter-prediction functions that predict a current frame based on a reference frame. [Background technology]
[0002] A video coding standard known as High-Efficiency Video Coding (HEVC) has been standardized by the Joint Collaborative Team on Video Coding (JCT-VC).
[0003] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there is a constant need to provide improvements and optimizations in video coding technology to handle ever-increasing amounts of digital video data in various applications. This disclosure relates to further advances, improvements, and optimizations in video coding, particularly in inter-prediction functions that build predictions for a current frame based on reference frames. Summary of the Invention [Problem to be solved by the invention]
[0004] There is a demand for further improvements in compression efficiency and reduction in processing load in encoding / decoding techniques.
[0005] Therefore, the present disclosure provides an encoding device, a decoding device, an encoding method, and a decoding method that can reduce the processing load and further improve compression efficiency. [Means for solving the problem]
[0006] An encoding device according to one aspect of the present disclosure includes a memory and a circuit connected to the memory. The circuit, in operation, generates a predictive block by performing an interpolation process using values of samples included in a reference picture, generates a gradient block of the same size as the predictive block using the predictive block, generates a predicted image using the gradient block, and encodes a current block based on the predicted image. The generation of the gradient block includes a process of calculating a gradient value indicating a difference between a value of a right sample adjacent to the right of a position of a target sample included in the predictive block and a value of a left sample adjacent to the left of the target sample. A first gradient value at the left end of the gradient block is calculated using the value of a first sample at the same position as the first gradient value in the predictive block as the left sample, and a second gradient value adjacent to the right of the position of the first gradient value is calculated using the value of the first sample as the left sample.
[0007] One aspect of the present disclosure is a coding device that codes a current block in a picture using inter prediction, the coding device including a processor and a memory, and the processor uses the memory to perform the following steps: obtaining two predicted images from two reference pictures by performing motion compensation using motion vectors corresponding to the two reference pictures, obtaining two gradient images corresponding to the two predicted images from the two reference pictures; deriving a local motion estimation value using the two gradient images and the two predicted images in sub-blocks obtained by dividing the current block; and generating a final predicted image of the current block using the local motion estimation value of the sub-blocks, the two gradient images, and the two predicted images.
[0008] According to another aspect of the present disclosure, there is provided an image coding apparatus comprising: a circuit and a memory coupled to the circuit, the circuit performing, in operation, steps including at least a prediction process using motion vectors from another picture, predicting a first block of prediction samples for a current block of a picture, padding the first block of prediction samples to form a second block of prediction samples larger than the first block, calculating at least gradients using the second block of prediction samples, and encoding the current block using at least the calculated gradients.
[0009] According to another aspect of the present disclosure, there is provided an image coding device comprising a circuit and a memory coupled to the circuit, the circuit performing, in operation, steps including at least a prediction process using a motion vector from another picture, predicting a first block of prediction samples for a current block of a picture, padding the first block of prediction samples to form a second block of prediction samples larger than the first block, performing an interpolation process using the second block of prediction samples, and encoding the current block using at least a block resulting from the interpolation process.
[0010] According to another aspect of the present disclosure, there is provided an image coding device including a circuit and a memory coupled to the circuit, the circuit performing, in operation, steps including at least a prediction process using a motion vector from another picture, predicting a first block of prediction samples for a current block of a picture, padding a second block of prediction samples adjacent to the current block to form a third block of prediction samples, performing an On-the-Band Multiplication (OBMC) process using at least the first and third block of prediction samples, and encoding the current block using at least a block resulting from the OBMC process.
[0011] According to another aspect of the present disclosure, there is provided an image coding device including a circuit and a memory connected to the circuit, wherein the circuit performs, in operation, steps including at least a prediction process using a first motion vector from another picture, predicting a first block of prediction samples for a current block of the picture, deriving a second motion vector for the current block by a dynamic motion vector refreshing (DMVR) process using at least the first motion vector, performing an interpolation process including a padding process on the current block using the second motion vector, and encoding the current block using at least the block resulting from the interpolation process.
[0012] According to another aspect of the present disclosure, there is provided an image decoding device comprising a circuit and a memory coupled to the circuit, the circuit performing, in operation, steps including at least a prediction process using a motion vector from another picture, predicting a first block of prediction samples for a current block of a picture, padding the first block of prediction samples to form a second block of prediction samples larger than the first block, calculating at least gradients using the second block of prediction samples, and decoding the current block using at least the calculated gradients.
[0013] According to another aspect of the present disclosure, there is provided an image decoding device comprising a circuit and a memory coupled to the circuit, the circuit performing, in operation, steps including at least a prediction process using a motion vector from another picture, predicting a first block of prediction samples for a current block of a picture, padding the first block of prediction samples to form a second block of prediction samples larger than the first block, performing an interpolation process using the second block of prediction samples, and decoding the current block using at least a block resulting from the interpolation process.
[0014] According to another aspect of the present disclosure, there is provided an image decoding device including a circuit and a memory connected to the circuit, wherein the circuit performs, in operation, steps including at least a prediction process using a motion vector from another picture, predicting a first block of prediction samples for a current block of a picture, padding a second block of prediction samples adjacent to the current block to form a third block of prediction samples, performing an On-the-Band Multiplication (OBMC) process using at least the first and third block of prediction samples, and decoding the current block using at least a block resulting from the OBMC process.
[0015] According to another aspect of the present disclosure, there is provided an image decoding device including a circuit and a memory connected to the circuit, wherein the circuit performs, in operation, steps including at least a prediction process using a first motion vector from another picture, predicting a first block of prediction samples for a current block of the picture, deriving a second motion vector for the current block by a dynamic motion vector refreshing (DMVR) process using at least the first motion vector, performing an interpolation process including a padding process on the current block using the second motion vector, and decoding the current block using at least a block resulting from the interpolation process.
[0016] According to another aspect of the present disclosure, there is provided an image encoding method enabling an image encoding device to perform steps according to various aspects of the present disclosure as described herein.
[0017] According to another aspect of the present disclosure, there is provided an image decoding method, which enables an image decoding device to perform steps according to various aspects of the present disclosure as described herein.
[0018] These comprehensive and specific embodiments may be realized using a system, a method, an integrated circuit, a computer program, or a computer-readable medium such as a CD-ROM, or may be realized by a combination of the system, the method, the integrated circuit, the computer program, and the medium. [Effects of the Invention]
[0019] The present disclosure can provide an encoding device, a decoding device, an encoding method, and a decoding method that can reduce processing load and further improve compression efficiency.
[0020] Some implementations of the present disclosure can improve encoding efficiency, simplify encoding / decoding processes, accelerate encoding / decoding process speed, and efficiently select appropriate components / operations to be used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.
[0021] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. Benefits and / or advantages may be obtained by the various embodiments and features of the specification and drawings individually, and it is not necessary for all of the various embodiments and features of the specification and drawings to be present in order to obtain one or more such benefits and / or advantages.
[0022] It should be noted that the generic or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any combination thereof. [Brief explanation of the drawings]
[0023] In the drawings, the same reference numbers refer to like elements, and the sizes and relative positions of elements in the drawings are not necessarily drawn to scale. [Figure 1] FIG. 1 is a block diagram showing a functional configuration of an encoding device according to an embodiment. [Figure 2]FIG. 2 is a diagram showing an example of block division. [Figure 3] FIG. 3 is a table showing the transformation basis functions corresponding to each transformation type. [Figure 4A] FIG. 4A is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter). [Figure 4B] FIG. 4B is a diagram showing another example of the shape of the filter used in ALF. [Figure 4C] FIG. 4C is a diagram showing another example of the shape of the filter used in ALF. [Figure 5A] FIG. 5A is a diagram showing 67 intra prediction modes in intra prediction. [Figure 5B] FIG. 5B is a flowchart for explaining an outline of the predictive image correction process using OBMC (overlapped block motion compensation) processing. [Figure 5C] FIG. 5C is a conceptual diagram for explaining an outline of the predicted image correction process using the OBMC process. [Figure 5D] FIG. 5D is a diagram showing an example of FRUC (frame rate up-conversion). [Figure 6] FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 8] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. [Figure 9A] FIG. 9A is a diagram for explaining derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 9B] FIG. 9B is a diagram for explaining an outline of the motion vector derivation process in the merge mode. [Figure 9C]FIG. 9C is a conceptual diagram for explaining an outline of DMVR (dynamic motion vector refreshing) processing. [Figure 9D] FIG. 9D is a diagram for explaining an outline of a predicted image generation method using luminance correction processing by LIC (local illumination compensation) processing. [Figure 10] FIG. 10 is a block diagram illustrating a functional configuration of a decoding device according to an embodiment. [Figure 11] FIG. 11 is a flowchart showing an inter prediction process according to another embodiment. [Figure 12] FIG. 12 is a conceptual diagram used to explain inter prediction according to the embodiment shown in FIG. [Figure 13] FIG. 13 is a conceptual diagram used to explain an example of the reference ranges of the gradient filter and the motion compensation filter according to the embodiment shown in FIG. [Figure 14] FIG. 14 is a conceptual diagram used to explain an example of the reference range of the motion compensation filter according to the first modification of the embodiment shown in FIG. [Figure 15] FIG. 15 is a conceptual diagram used to explain an example of the reference range of the gradient filter according to the first modification of the embodiment shown in FIG. [Figure 16] FIG. 16 is a diagram showing an example of a pixel pattern to be referred to by deriving a local motion detection value according to the second modification of the embodiment shown in FIG. [Figure 17] FIG. 17 is a flowchart illustrating an example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture. [Figure 18A] 18A is a conceptual diagram illustrating an example of picture order counts of a picture and another picture, in which the picture may be designated as the current picture, and the other picture may be designated as the first picture or the second picture. [Figure 18B]FIG. 18B is a conceptual diagram illustrating another example of a picture order count of a picture and another picture. [Figure 18C] FIG. 18C is a conceptual diagram illustrating yet another example of a picture order count of a picture and another picture. [Figure 19A] FIG. 19A illustrates an example of padding directions for padding blocks of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 19B] FIG. 19B illustrates another example of padding directions for padding blocks of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 20A] FIG. 20A illustrates an example process for padding a block of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 20B] FIG. 20B illustrates another example process for padding a block of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 20C] FIG. 20C illustrates another example process for padding a block of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 20D] FIG. 20D illustrates yet another example of a process for padding a block of prediction samples according to the example image encoding / decoding method as illustrated in FIG. [Figure 21A] FIG. 21A shows an example of a gradient filter for a block. [Figure 21B] FIG. 21B shows an example of multiple gradient filters for a block. [Figure 21C] FIG. 21C shows yet another example of a gradient filter for a block. [Figure 21D] FIG. 21D shows yet another example of a gradient filter for a block. [Figure 22] FIG. 22 is a flowchart showing an embodiment of an image encoding / decoding method according to the example shown in FIG. [Figure 23]FIG. 23 is a conceptual diagram showing an embodiment of the image encoding / decoding method shown in FIG. [Figure 24] FIG. 24 is a flowchart showing another embodiment of the image encoding / decoding method according to the example shown in FIG. [Figure 25] FIG. 25 is a conceptual diagram showing an embodiment of the image encoding / decoding method shown in FIG. [Figure 26] FIG. 26 shows an example of a block generated by the padding process in the embodiment of the image encoding / decoding method shown in FIGS. [Figure 27A] FIG. 27A is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture. [Figure 27B] FIG. 27B is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture. [Figure 27C] FIG. 27C is a flowchart illustrating yet another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture. [Figure 28] FIG. 28 is a conceptual diagram showing an embodiment of the image encoding / decoding method as shown in FIG. 27A. [Figure 29] FIG. 29 is a conceptual diagram showing an embodiment of the image encoding / decoding method as shown in FIG. 27B. [Figure 30] FIG. 30 is a conceptual diagram showing an embodiment of the image encoding / decoding method as shown in FIG. 27C. [Figure 31] FIG. 31 shows an example of adjacent blocks of the current block. [Figure 32A] FIG. 32A illustrates an example process for padding a block of prediction samples according to the example image encoding / decoding method as illustrated in FIGS. 27A, 27B, and 27C. [Figure 32B]FIG. 32B illustrates another example process for padding blocks of prediction samples according to the example image encoding / decoding method as illustrated in FIGS. 27A, 27B, and 27C. [Figure 32C] FIG. 32C illustrates another example process for padding blocks of prediction samples according to the example image encoding / decoding method illustrated in FIGS. 27A, 27B, and 27C. [Figure 33A] FIG. 33A shows examples of padding directions for padding blocks according to the example image encoding / decoding methods shown in FIGS. 17, 22, 24, 27A, 27B, and 27C. [Figure 33B] FIG. 33B shows another example of a padding direction for padding blocks according to the example image encoding / decoding methods shown in FIGS. 17, 22, 24, 27A, 27B, and 27C. [Figure 33C] FIG. 33C shows another example of a padding direction for padding blocks according to the example image encoding / decoding methods shown in FIGS. 17, 22, 24, 27A, 27B, and 27C. [Figure 33D] FIG. 33D shows yet another example of a padding direction for padding blocks according to the example image encoding / decoding methods shown in FIGS. 17, 22, 24, 27A, 27B, and 27C. [Figure 34] Figure 34 shows alternative examples of blocks of prediction samples for the example image encoding / decoding methods as shown in Figures 17, 22, 24, 27A, 27B, and 27C, in which the blocks of prediction samples have non-rectangular shapes. [Figure 35] FIG. 35 is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture. [Figure 36] FIG. 36 is a conceptual diagram showing an embodiment of the image encoding / decoding method shown in FIG. [Figure 37A]FIG. 37A illustrates an example process for padding a block of prediction samples according to a second motion vector for the example image encoding / decoding method as illustrated in FIG. [Figure 37B] FIG. 37B illustrates another example process for padding a block of prediction samples according to a second motion vector for the example image encoding / decoding method as illustrated in FIG. [Figure 38] 38 is a flowchart showing yet another alternative example of an image encoding / decoding method in which an inter prediction function is used to generate a prediction of a current block of a picture based on a reference block of another picture. The flowchart in Figure 38 is similar to the flowchart in Figure 35, except that the DMVR (dynamic motion vector refreshing) process in step 3804 further includes a padding process. [Figure 39] FIG. 39 is a conceptual diagram showing an embodiment of the image encoding / decoding method shown in FIG. [Figure 40] FIG. 40 is a block diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 41] FIG. 41 is a conceptual diagram showing an example of a coding structure for scalable coding. [Figure 42] FIG. 42 is a conceptual diagram showing an example of a coding structure for scalable coding. [Figure 43] FIG. 43 is a conceptual diagram showing an example of a display screen of a web page. [Figure 44] FIG. 44 is a conceptual diagram showing an example of a display screen of a web page. [Figure 45] FIG. 45 is a block diagram showing an example of a smartphone. [Figure 46] FIG. 46 is a block diagram showing an example of the configuration of a smartphone. DETAILED DESCRIPTION OF THE INVENTION
[0024] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection forms, steps, step relationships and order, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Therefore, components disclosed in the following embodiments but not recited in the independent claims defining the broadest inventive concept may be understood as optional components.
[0025] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with regard to the processes and / or configurations applied to the embodiments, any of the following may be implemented.
[0026] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.
[0027] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, or deleting a function or process. For example, any function or process may be replaced with or combined with another function or process described in any of the aspects of the present disclosure.
[0028] (3) In the method implemented by the encoding device or decoding device of the embodiment, some of the processes included in the method may be arbitrarily modified, such as by addition, replacement, deletion, etc. For example, any process in the method may be replaced with or combined with another process described in any of the aspects of the present disclosure.
[0029] (4) Some of the components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in any of the aspects of the present disclosure.
[0030] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.
[0031] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the multiple processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.
[0032] (7) Some of the processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.
[0033] (8) The implementation of the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for a purpose other than the video encoding or video decoding disclosed in the embodiments.
[0034] [Encoding device] First, an overview of a coding device according to an embodiment will be described. Fig. 1 is a block diagram showing the functional configuration of a coding device 100 according to an embodiment. The coding device 100 is a video coding device that codes a video on a block-by-block basis.
[0035] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0036] The encoding device 100 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Alternatively, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0037] Each component included in the encoding device 100 will be described below.
[0038] [Divided part] The division unit 102 divides each picture included in the input video into multiple blocks and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides the picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks may be called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). In various processing examples, CUs, PUs, and TUs do not need to be distinguished from one another, and some or all of the blocks in a picture may serve as the processing units of CUs, PUs, and TUs.
[0039] Fig. 2 is a diagram showing an example of block division in an embodiment, in which solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.
[0040] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).
[0041] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.
[0042] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14 and 15 (binary tree block division).
[0043] The lower-left 64x64 block is divided into four square 32x32 blocks (quadtree block decomposition). Of the four 32x32 blocks, the upper-left and lower-right blocks are further divided. The upper-left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block decomposition). The lower-right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block decomposition). As a result, the lower-left 64x64 block is divided into 16x32 block 16, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.
[0044] The bottom right 64x64 block 23 is not split.
[0045] 2, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. This type of division is sometimes called QTBT (quad-tree plus binary tree) division.
[0046] In Fig. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.
[0047] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (a prediction sample input from a prediction control unit 128 described below) from the original signal (original sample) input from the division unit 102, for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also referred to as a residual) of a block to be coded (hereinafter referred to as a current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.
[0048] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing an image may also be referred to as a sample.
[0049] [Conversion section] The transform unit 106 transforms the spatial domain prediction errors into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction errors.
[0050] The transform unit 106 may adaptively select a transform type from among a plurality of transform types and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).
[0051] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Fig. 3 is a table showing transform basis functions corresponding to each transform type. In Fig. 3, N represents the number of input pixels. Selection of a transform type from among these multiple transform types may depend, for example, on the type of prediction (intra prediction or inter prediction) or the intra prediction mode.
[0052] Such information indicating whether EMT or AMT is applied (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transformation type are usually signaled at the CU level, but the signaling of this information does not need to be limited to the CU level and may be at other levels (e.g., the bit sequence level, picture level, slice level, tile level, or CTU level).
[0053] Furthermore, the transform unit 106 may retransform the transform coefficients (transform results). Such retransformation may be referred to as an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transform unit 106 performs retransformation for each sub-block (e.g., 4x4 sub-block) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether to apply NSST and information regarding the transform matrix used for NSST are typically signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may also be at other levels (e.g., the sequence level, picture level, slice level, tile level, or CTU level).
[0054] Separable transformation and non-separable transformation may be applied to the transformation unit 106. Separable transformation is a method of separating the input into directions for the number of dimensions and performing transformation multiple times, and non-separable transformation is a method of treating two or more dimensions of a multi-dimensional input as one dimension and performing transformation all at once.
[0055] For example, an example of a non-separable transformation is when the input is a 4x4 block, it is treated as a single array with 16 elements, and a 16x16 transformation matrix is used to perform transformation processing on that array.
[0056] Another example of a non-separable transformation is to treat a 4x4 input block as a single array with 16 elements, and then perform a transformation (e.g., a Hypercube Givens Transform) on the array by performing multiple Givens rotations.
[0057] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on quantization parameters (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.
[0058] The predetermined scanning order is an order for quantizing / dequantizing transform coefficients, for example, the predetermined scanning order is defined as an ascending order (low frequency to high frequency) or a descending order (high frequency to low frequency).
[0059] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. In other words, as the value of the quantization parameter increases, the quantization error also increases.
[0060] [Entropy coding section] The entropy coding unit 110 generates a coded signal (coded bitstream) based on the quantized coefficients input from the quantization unit 108. Specifically, for example, the entropy coding unit 110 binarizes the quantized coefficients, arithmetically codes the binary signal, and outputs a compressed bitstream or sequence.
[0061] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. The inverse quantization unit 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0062] [Inverse conversion section] The inverse transform unit 114 restores prediction errors (residuals) by inverse transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores prediction errors of the current block by performing an inverse transform on the transform coefficients corresponding to the transform performed by the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction errors to the adder unit 116.
[0063] Note that the restored prediction error usually loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error usually contains a quantization error.
[0064] [Addition section] The adder 116 reconstructs a current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.
[0065] [Block Memory] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are in the picture to be coded (referred to as the "current picture"). Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.
[0066] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0067] ALF applies a least squares error filter to remove coding artifacts, for example, for each 2x2 sub-block in the current block, one filter selected from multiple filters based on local gradient direction and activity.
[0068] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The sub-blocks are classified based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, the sub-blocks are classified into a plurality of classes based on the classification value C.
[0069] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived, for example, by adding gradients in multiple directions and quantizing the sum.
[0070] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.
[0071] The filter shape used in ALF is, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing several examples of filter shapes used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the filter shape is usually signaled at the picture level. Note that signaling of the information indicating the filter shape does not need to be limited to the picture level, and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0072] Whether to turn on or off ALF may be determined at the picture level or the CU level. For example, whether to apply ALF for luma may be determined at the CU level, and whether to apply ALF for chroma may be determined at the picture level. Information indicating whether ALF is on or off is usually signaled at the picture level or the CU level. Note that signaling of information indicating whether ALF is on or off does not need to be limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).
[0073] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level, although the signaling of coefficient sets need not be limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0074] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may also be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0075] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also referred to as intra-picture prediction) of the current block with reference to blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0076] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes typically includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0077] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.
[0078] The plurality of directional prediction modes includes, for example, 33 prediction modes defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may include 32 prediction modes in addition to the 33 directions (65 directional prediction modes in total).
[0079] 5A is a conceptual diagram showing all 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra-prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the two "non-directional" prediction modes are not shown in FIG. 5A).
[0080] In various processing examples, a luminance block may be referenced in intra prediction of a chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. This intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of the chrominance block that references such a luminance block (e.g., called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0081] The intra prediction unit 124 may correct pixel values after intra prediction based on gradients of reference pixels in the horizontal / vertical directions. Intra prediction involving such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is usually signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0082] [Inter prediction section] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also referred to as inter prediction) on the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or current sub-block to find a reference block or sub-block within the reference picture that most closely matches the current block or sub-block. Then, the inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion estimation to obtain motion information (e.g., a motion vector) that compensates for (or predicts) the movement or change from the reference block or sub-block to the current block or sub-block, and generates an inter prediction signal for the current block or sub-block based on the motion information. The inter prediction unit 126 then outputs the generated inter prediction signal to the prediction control unit 128.
[0083] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms, such as a motion vector, or as a difference between a motion vector and a predicted motion vector.
[0084] Note that an inter-prediction signal may be generated using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks. Specifically, an inter-prediction signal may be generated for each sub-block in the current block by weighting and adding a prediction signal based on the motion information obtained by motion estimation (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture). Such inter-prediction (motion compensation) may be called OBMC (overlapped block motion compensation).
[0085] In the OBMC mode, information indicating the size of a sub-block for OBMC (e.g., referred to as an OBMC block size) may be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as an OBMC flag) may be signaled at the CU level. Note that the signaling level of this information is not limited to the sequence level and the CU level, and may be at another level (e.g., the picture level, slice level, tile level, CTU level, or sub-block level).
[0086] The OBMC mode will now be described in more detail. Figures 5B and 5C are a flowchart and a conceptual diagram illustrating the predicted image correction process using the OBMC process.
[0087] Referring to Figure 5C, first, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to the current block to be coded. In Figure 5C, the arrow "MV" points to a reference picture, indicating what the current block in the current picture references to obtain the predicted image.
[0088] Next, the motion vector (MV_L) already derived for the coded left neighboring block is applied (reused) to the current block to obtain a predicted image (Pred_L). The motion vector (MV_L) is indicated by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by overlapping the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between the neighboring blocks.
[0089] Similarly, a motion vector (MV_U) already derived for the coded upper adjacent block is applied (reused) to the current block to be coded to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. Then, a second correction of the predicted image is performed by superimposing the predicted image Pred_U on the predicted images (i.e., Pred and Pred_L) that have been corrected the first time. This has the effect of blending the boundaries between adjacent blocks, in one aspect. The predicted image obtained by the second correction is the final predicted image of the current block, in which the boundaries with the adjacent blocks have been blended (smoothed).
[0090] Although a two-stage correction method using the left adjacent block and the above adjacent block has been described here, it is also possible to configure a method in which correction is performed more than two times using the right adjacent block or the below adjacent block.
[0091] The area to be superimposed does not have to be the pixel area of the entire block, but may be only a part of the area near the block boundary.
[0092] Here, the OBMC predicted image correction process has been described, in which additional predicted images Pred_L and Pred_U are superimposed based on one reference picture to obtain one predicted image Pred. However, when a predicted image is corrected based on multiple reference pictures, the same process may be applied to each of the multiple reference pictures. In such a case, the OBMC image correction based on multiple reference pictures is performed to obtain a corrected predicted image from each reference picture, and then the obtained multiple corrected predicted images are further superimposed to obtain a final predicted image.
[0093] In OBMC, the unit of the current block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.
[0094] As a method for determining whether to apply OBMC processing, for example, there is a method using obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, an encoding device determines whether a block to be encoded belongs to an area with complex motion, and if it belongs to an area with complex motion, sets the value "1" as obmc_flag and performs encoding by applying OBMC processing, and if it does not belong to an area with complex motion, sets the value "0" as obmc_flag and performs encoding without applying OBMC processing. On the other hand, a decoding device decodes obmc_flag described in a stream (i.e., a compressed sequence), and switches whether to apply OBMC processing depending on the value, and performs decoding.
[0095] The motion information may be derived on the decoding device side without being signaled from the encoding device side. For example, a merge mode defined in the H.265 / HEVC standard may be used. Alternatively, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation may be performed on the decoding device side without using pixel values of the current block.
[0096] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.
[0097] An example of the FRUC process is shown in Figure 5D. First, a list of multiple candidates (which may be the same as the merge list) each having a predicted motion vector (MV) is generated by referring to the motion vectors of coded blocks spatially or temporally adjacent to the current block. Next, a best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value of each candidate MV included in the candidate list is calculated, and one candidate MV is selected based on the evaluation value.
[0098] Then, a motion vector for the current block is derived based on the motion vector of the selected candidate. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as the motion vector for the current block. Also, for example, the motion vector for the current block may be derived by performing pattern matching in a peripheral area of a position in the reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed using pattern matching in the reference picture and an evaluation value for the peripheral area of the best candidate MV, and if an MV with a better evaluation value is found, the best candidate MV may be updated to that MV, and this may be used as the final MV for the current block. It is also possible to configure the system without performing the process of updating to an MV with a better evaluation value.
[0099] The same processing may be performed when processing is performed in sub-block units.
[0100] The evaluation value may be calculated by various methods. For example, a reconstructed image of an area in a reference picture corresponding to the motion vector may be compared with a reconstructed image of a predetermined area (for example, as described later, this may be an area of another reference picture or an area of an adjacent block in the current picture), and the difference in pixel values between the two reconstructed images may be calculated and used as the evaluation value of the motion vector. The evaluation value may also be calculated using other information in addition to the difference value.
[0101] Next, an example of pattern matching will be described in detail. First, one candidate MV included in a candidate MV list (e.g., a merge list) is selected as a starting point for search by pattern matching. For example, first pattern matching or second pattern matching can be used as pattern matching. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.
[0102] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture that is along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the candidate for an area in a reference picture.
[0103] 6 is a diagram illustrating an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of blocks that best match among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of a current block (Cur block). Specifically, for the current block, a difference is derived between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by a candidate MV and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV by the display time interval, and an evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, which can lead to good results.
[0104] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (CurPic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures temporally and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives two mirror-symmetric bidirectional motion vectors.
[0105] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.
[0106] 7 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of a current block is derived by searching a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in the current picture (Cur Pic). Specifically, a difference is derived between a reconstructed image of both or either of the coded areas adjacent to the left and / or above the current block and a reconstructed image at the same position in the coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the multiple candidate MVs can be selected as the best candidate MV.
[0107] Information indicating whether such a FRUC mode is applied (e.g., referred to as an FRUC flag) may be signaled at the CU level. Furthermore, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating a pattern matching method (e.g., first pattern matching or second pattern matching) (e.g., referred to as an FRUC mode flag) may be signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, the picture level, the slice level, the tile level, the CTU level, or the sub-block level).
[0108] Next, we will explain how to derive a motion vector. First, we will explain a mode in which a motion vector is derived based on a model that assumes uniform linear motion. This mode is sometimes called BIO (bi-directional optical flow) mode.
[0109] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v y) denotes a velocity vector, and τ0 and τ1 denote the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0 and Ref1), respectively. (MVx0,MVy0) denotes a motion vector corresponding to reference picture Ref0, and (MVx1,MVy1) denotes a motion vector corresponding to reference picture Ref1.
[0110] At this time, the velocity vector (v x ,v y ), (MVx0,MVy0) and (MVx1,MVy1) are respectively (v x τ0,v y τ0) and (-v x τ1,-v y τ1), and the following optical flow equation (1) holds:
[0111]
number
[0112] where I (k) denotes the luminance value of reference image k (k=0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-based motion vectors obtained from a merge list or the like may be corrected pixel-by-pixel.
[0113] Note that the decoding device may derive motion vectors using a method other than that based on a model assuming constant-velocity linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.
[0114] Next, a mode in which a motion vector is derived for each sub-block based on the motion vectors of multiple neighboring blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.
[0115] FIG. 9A is a diagram for explaining the derivation of a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v x ,v y ) is derived.
[0116]
number
[0117] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting coefficient.
[0118] The affine motion compensation prediction mode may include several modes in which the methods of deriving the motion vectors of the upper-left and upper-right corner control points are different. Information indicating the affine motion compensation prediction mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating the affine motion compensation prediction mode does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0119] [Predictive control unit] The prediction control unit 128 selects either an intra-prediction signal (a signal output from the intra-prediction unit 124) or an inter-prediction signal (a signal output from the inter-prediction unit 126), and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.
[0120] As shown in FIG. 1 , in various processing examples, the prediction control unit 128 may output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bitstream and perform the same prediction process as that performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or any index, flag, or value based on or indicating the prediction process performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128.
[0121] In some embodiments, the prediction control unit 128 operates in merge mode and optimizes the motion vector calculated for the current picture using the intra prediction signal from the intra predictor 124 and the inter prediction signal from the inter predictor 126. Figure 9B shows an example of a process for deriving a motion vector for the current picture in merge mode.
[0122] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include spatially adjacent prediction MVs, which are MVs held by multiple coded blocks spatially located around the target block, temporally adjacent prediction MVs, which are MVs held by blocks in the vicinity of the target block projected onto the coded reference picture, joint prediction MVs, which are MVs generated by combining the MV values of the spatially adjacent prediction MVs and the temporally adjacent prediction MVs, and zero prediction MVs, which are MVs with a value of zero.
[0123] Next, one predicted MV is selected from the plurality of predicted MVs registered in the predicted MV list, and is determined as the MV for the target block.
[0124] Furthermore, the variable length coding unit encodes merge_idx, which is a signal indicating which predicted MV has been selected, into the stream.
[0125] Note that the predicted MVs registered in the predicted MV list described in Figure 9B are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include predicted MVs other than the types of predicted MVs shown in the figure.
[0126] The final MV may be determined by performing a DMVR (decoder motion vector refinement) process, which will be described later, using the MV of the current block derived in the merge mode.
[0127] FIG. 9C is a conceptual diagram illustrating an example of DMVR processing for determining an MV.
[0128] First, the optimal MVP set for the current block (for example, in merge mode) is set as the candidate MV. Then, reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 direction, according to the candidate MV (L0). Similarly, reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 direction, according to the candidate MV (L1). A template is generated by averaging these reference pixels.
[0129] Next, the template is used to search the surrounding areas of the candidate MVs in the first reference picture (L0) and the second reference picture (L1), and the MV with the smallest cost is determined as the final MV. Note that the cost value may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search area, the candidate MV value, etc.
[0130] Typically, the encoding device and the decoding device described below basically have the same configuration and operation for the processing described here.
[0131] Any processing may be used, not limited to the processing example described here, as long as it is a processing that can search around the candidate MVs and derive the final MV.
[0132] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.
[0133] FIG. 9D is a conceptual diagram for explaining an example of a predicted image generation method using luminance correction processing by LIC processing.
[0134] First, the MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.
[0135] Next, information indicating how the luminance values of the current block have changed between the reference picture and the current picture is extracted. This extraction is performed based on the luminance pixel values of the coded left-adjacent reference area (peripheral reference area) and the coded upper-adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV. Then, the information indicating how the luminance values have changed is used to calculate luminance correction parameters.
[0136] A predicted image for the current block is generated by performing luminance correction processing that applies the luminance correction parameters to a reference image in a reference picture specified by the MV.
[0137] The shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.
[0138] Although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures, and a luminance correction process may be performed on the reference images obtained from each reference picture in the same manner as described above before generating a predicted image.
[0139] One method for determining whether to apply LIC processing is to use lic_flag, which is a signal indicating whether to apply LIC processing. As a specific example, an encoding device determines whether the current block belongs to an area where a luminance change occurs, and if the current block belongs to an area where a luminance change occurs, sets the value of lic_flag to "1" and performs encoding by applying LIC processing, and if the current block does not belong to an area where a luminance change occurs, sets the value of lic_flag to "0" and performs encoding without applying LIC processing. On the other hand, a decoding device may decode lic_flag described in the stream, and switch whether to apply LIC processing depending on the value and perform decoding.
[0140] Another method for determining whether to apply LIC processing is to determine whether LIC processing has been applied to neighboring blocks.As a specific example, when the current block is in merge mode, it is determined whether the neighboring coded blocks selected when deriving MV in merge mode processing have been coded using LIC processing.Depending on the result, whether to apply LIC processing is switched and coding is performed.In this example, the same processing is also applied to the processing on the decoding device side.
[0141] [Overview of the decoding device] Next, an overview will be given of a decoding device capable of decoding the coded signal (coded bitstream) output from the above coding device 100. Fig. 10 is a block diagram showing the functional configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a video decoding device that decodes video on a block-by-block basis.
[0142] As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0143] The decoding device 200 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0144] Each component included in the decoding device 200 will be described below.
[0145] [Entropy Decoding] The entropy decoding unit 202 entropy-decodes the coded bitstream. Specifically, for example, the entropy decoding unit 202 arithmetically decodes the coded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. The entropy decoding unit 202 outputs quantized coefficients to the inverse quantization unit 204 on a block-by-block basis. The entropy decoding unit 202 may output prediction parameters included in the coded bitstream (see FIG. 1 ) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the coding device side.
[0146] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of the block to be decoded (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. The inverse quantization unit 204 then outputs the inverse quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0147] [Inverse conversion section] The inverse transform unit 206 restores prediction errors (residuals) by inverse transforming the transform coefficients input from the inverse quantization unit 204.
[0148] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.
[0149] Also, for example, if the information decoded from the coded bitstream indicates that NSST is to be applied, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.
[0150] [Addition section] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transformer 206 and the prediction sample input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0151] [Block Memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter referred to as a current picture). Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208.
[0152] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.
[0153] If the information indicating ALF on / off read from the encoded bitstream indicates that ALF is on, one filter is selected from multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.
[0154] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.
[0155] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction based on the intra prediction mode interpreted from the encoded bitstream, by referring to blocks in the current picture stored in the block memory 210. Specifically, the intra prediction unit 216 generates the intra prediction signal by performing intra prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0156] Note that when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0157] Furthermore, when information interpreted from the coded bitstream (for example, prediction parameters output from the entropy decoding unit 202) indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical directions.
[0158] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the coded bitstream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.
[0159] If the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation, but also the motion information of neighboring blocks.
[0160] Furthermore, if the information interpreted from the coded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the coded bitstream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.
[0161] Furthermore, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming constant-velocity linear motion. Furthermore, when information interpreted from the coded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks.
[0162] [Predictive control unit] The prediction control unit 220 selects either an intra-prediction signal or an inter-prediction signal, and outputs the selected signal as a prediction signal to the adder 208. The prediction control unit 220 may also perform various functions and processes, such as the merge mode (see FIG. 9B), DMVR processing (see FIG. 9C), and LIC processing (see FIG. 9D), as described above, with reference to the prediction control unit 128 on the encoding device side. Overall, the configurations, functions, and processing of the prediction control unit 220, intra prediction unit 216, and inter prediction unit 218 on the decoding device side may correspond to the configurations, functions, and processing of the prediction control unit 128, intra prediction unit 124, and inter prediction unit 126 on the encoding device side.
[0163] [Interpretation Details] An example of another embodiment of inter prediction according to the present application will be described below. This embodiment relates to so-called BIO mode inter prediction. In this embodiment, the motion vector of a block is corrected not in units of pixels but in units of sub-blocks, similar to the embodiment shown in FIGS. 1 to 10. In describing this embodiment, attention will be focused on the differences from the embodiment shown in FIGS. 1 to 10.
[0164] The configurations of the encoding device and decoding device in this embodiment are substantially the same as those in the embodiment shown in FIGS. 1 to 10, and therefore depiction and description of these configurations will be omitted.
[0165] [Inter Prediction] Fig. 11 is a flowchart showing inter prediction in this embodiment. Fig. 12 is a conceptual diagram used to explain inter prediction in this embodiment. The following processing is performed by the inter prediction unit 126 of the encoding device 100 and the inter prediction unit 218 of the decoding device 200.
[0166] As shown in Fig. 11, first, block loop processing is performed on a plurality of blocks in a picture to be coded / decoded (current picture 1000) (S101 to S111). In Fig. 12, a block to be coded / decoded is selected as a current block 1001 from the plurality of blocks.
[0167] In the block loop process, loop processing is performed on the reference pictures, ie, the first reference picture 1100 (L0) and the second reference picture 1200 (L1), which are processed pictures (S102 to S106).
[0168] In the reference picture loop process, first, a block motion vector is derived or obtained to obtain a predicted image from a reference picture (S103). In FIG. 12, a first motion vector 1110 (MV_L0) for a first reference picture 1100 is derived or obtained, and a second motion vector 1210 (MV_L1) for a second reference picture 1200 is derived or obtained. Examples of available motion vector derivation methods include a normal inter prediction mode, a merge mode, and a FRUC mode. In the normal inter prediction mode, the encoding device 100 derives a motion vector by motion search, and the decoding device 200 obtains the motion vector from a bitstream.
[0169] Next, a predicted image is obtained from the reference picture by performing motion compensation using the derived or acquired motion vector (S104). In Fig. 12, a first predicted image 1140 is obtained from the first reference picture 1100 by performing motion compensation using a first motion vector 1110. Also, a second predicted image 1240 is obtained from the second reference picture 1200 by performing motion compensation using a second motion vector 1210.
[0170] In motion compensation, a motion compensation filter is applied to a reference picture. The motion compensation filter is an interpolation filter used to obtain a predicted image with sub-pixel accuracy. Using a first reference picture 1100 in FIG. 12, pixels in a first interpolation range 1130 including pixels in a first prediction block 1120 and surrounding pixels are referenced by the motion compensation filter in the first prediction block 1120 indicated by a first motion vector 1110. Also, using a second reference picture 1200, pixels in a second interpolation range 1230 including pixels in a second prediction block 1220 and surrounding pixels are referenced by the motion compensation filter in the second prediction block 1220 indicated by a second motion vector 1210.
[0171] The first interpolation reference range 1130 and the second interpolation range 1230 include a first normal reference range and a second normal reference range used to perform motion compensation on the current block 1001 in normal inter prediction, which uses local motion estimation values. The first normal reference range is included in the first reference picture 1100, and the second normal reference range is included in the second reference picture 1200. In normal inter prediction, a motion vector is derived for a block using motion estimation, motion compensation is performed for the derived motion vector for the block, and the motion-compensated image is used as the final predicted image. In other words, local motion estimation values are not used in normal inter prediction. The first interpolation reference range 1130 and the second interpolation reference range 1230 may be the same as the first normal reference range and the second normal reference range.
[0172] Next, a gradient image corresponding to the predicted image is obtained from a reference picture (S105). Each pixel in the gradient image has a gradient value indicating the spatial gradient of luma or chroma. The gradient value is obtained by applying a gradient filter to the reference picture. In the first reference picture 1100 in FIG. 12 , pixels in a first gradient reference range 1135, which includes pixels in a first prediction block 1120 and surrounding pixels, are referenced using the gradient filter for the first prediction block 1120. The first gradient reference range 1135 is included in the first interpolation reference range 1130. In the second reference picture 1200, pixels in a second gradient reference range 1235, which includes pixels in a second prediction block 1220 and surrounding pixels, are referenced using the gradient filter for the second prediction block 1220. The second gradient reference range 1235 is included in the second interpolation reference range 1230.
[0173] When the predicted image and the gradient image are obtained from both the first reference picture and the second reference picture, the reference picture loop process ends (S106). After that, the loop process is performed on a plurality of sub-blocks obtained by further dividing the block (S107 to S110). The size of each sub-block is smaller than the size of the current block (for example, 4×4 pixels).
[0174] In the sub-block loop process, first, a local motion estimation value 1300 is derived using the first predicted image 1140, the second predicted image 1240, the first gradient image 1150, and the second gradient image 1250 obtained from the first reference picture 1100 and the second reference picture 1200 (S108). For example, the local motion estimation value 1300 is derived for a sub-block by referring to pixels in a predicted sub-block in each of the first predicted image 1140, the second predicted image 1240, the first gradient image 1150, and the second gradient image 1250. The predicted sub-block is an area in the first predicted block 1140 or the second predicted block 1240 that corresponds to a sub-block of the current block 1001. The local motion estimation value is also called a correction motion vector.
[0175] Next, a final predicted image 1400 for the sub-block is generated using the pixel values of the first predicted image 1140 and the second predicted image 1240, the gradient values of the first gradient image 1150 and the second gradient image 1250, and the local motion estimation value 1300. When a final predicted image has been generated for each sub-block of the current block, a final predicted image for the current block is generated and the sub-block loop process ends (S110).
[0176] When the block loop process ends (S111), the process in FIG. 11 ends.
[0177] It should be noted that the predicted image and gradient image can be obtained at the sub-block level by assigning to each sub-block a motion vector for the sub-block of the current block.
[0178] [Reference Range of Motion Compensation Filter and Gradient Filter in the Example of the Present Embodiment] An example of the reference ranges of the motion compensation filter and the gradient filter according to this embodiment will be described below.
[0179] FIG. 13 is a conceptual diagram used to explain an example of the reference ranges of the gradient filter and the motion compensation filter according to the example of this embodiment.
[0180] In Fig. 13, each circle represents a sample. In the example shown in Fig. 13, the size of the current block is 8x8 samples, and the size of the sub-block is 4x4 samples.
[0181] The reference range 1131 is the reference range (e.g., a square range of 8x8 samples) of the motion compensation filter applied to the top-left sample 1122 of the first prediction block 1120. The reference range 1231 is the reference range (e.g., a square range of 8x8 samples) of the motion compensation filter applied to the top-left sample 1222 of the second prediction block 1220.
[0182] The reference range 1132 is the reference range (e.g., a square range of 6x6 samples) of the gradient filter applied to the top-left sample 1122 of the first prediction block 1120. The reference range 1232 is the reference range (e.g., a square range of 6x6 samples) of the gradient filter applied to the top-left sample 1222 of the second prediction block 1220.
[0183] The motion compensation filter and gradient filter are applied to the other samples of the first prediction block 1120 and the second prediction block 1220, while referencing samples within a reference range of the same size at a position corresponding to each sample. As a result, samples within the first interpolation range 1130 and the second interpolation range 1230 are referenced to obtain the first predicted image 1140 and the second predicted image 1240. Also, samples within the first gradient range 1135 and the second gradient range 1235 are referenced to obtain the first gradient image 1150 and the second gradient image 1250.
[0184] [Effects, etc.] In this way, the encoding device and decoding device according to this embodiment can derive local motion estimation values for sub-blocks, thereby reducing prediction errors using local motion estimation values in sub-sample units, while reducing processing load or processing time compared to deriving local motion estimation values in sample units.
[0185] Furthermore, the encoding device and decoding device according to this embodiment can use an interpolation reference range within the normal reference range. Therefore, when generating a final predicted image using local motion estimation values in sub-blocks, there is no need to load new sample data from the frame memory during motion compensation, which can prevent increases in memory capacity and memory bandwidth.
[0186] Furthermore, the encoding device and decoding device of this embodiment can use the gradient reference range within the normal reference range, which eliminates the need to load new sample data from the frame memory to obtain a gradient image, thereby suppressing increases in memory capacity and memory bandwidth.
[0187] This embodiment may be combined with at least some aspects of other embodiments, and some processes, some elements of the device, and some syntax described in the flowcharts of this embodiment may be combined with aspects of other embodiments.
[0188] [Modification 1 of this embodiment] In the following, variations of the gradient filter and the motion compensation filter according to the present embodiment will be described in detail. In Variation 1, the processing for the second predicted image is similar to the processing for the first predicted image, and further description will be omitted or simplified.
[0189] [Motion compensation filter] First, the motion compensation filter will be described. Fig. 14 is a conceptual diagram used to explain an example of the reference range of the motion compensation filter in the first modification of this embodiment.
[0190] In the following description, a motion compensation filter with 1 / 4 samples in the horizontal direction and 1 / 2 samples in the vertical direction is applied to the first prediction block 1120. This motion compensation filter is a so-called 8-tap filter, and is expressed by the following equation (3).
[0191]
number
[0192] where I k [x,y] represents a sample value in the first predicted image with sub-sample accuracy when k is 0, and represents a sample value in the second predicted image with sub-sample accuracy when k is 1. The sample value is a value for a sample, for example, a luminance value or a chrominance value in the predicted image. Here, w 0.25 and w 0.5 is the weighting factor for 1 / 4 sample accuracy and 1 / 2 sample accuracy. I0 kWhen k is 0, [x, y] represents a sample value in the first predicted image at full sample precision, and when k is 1, [x, y] represents a sample value in the second predicted image at full sample precision.
[0193] For example, when the motion compensation filter in equation (3) is applied to the top left sample 1122 in Figure 14, the values of the horizontally arranged samples within the reference range 1131A are weighted and added for each row, and the addition results for those rows are also weighted and added.
[0194] In this modification, the motion compensation filter for the upper left sample 1122 references samples within a reference range 1131A. The reference range 1131A is a rectangular range that continues from the upper left sample 1122 by three samples to the left, four samples to the right, three samples above, and four samples below.
[0195] This motion compensation filter is applied to all samples in the first prediction block 1120. Thus, the samples in the first interpolation reference range 1130A are referenced by the motion compensation filter for the first prediction block 1120.
[0196] This motion compensation filter is applied to the second predicted block 1220 in the same way as the first predicted block 1120. That is, samples in the reference range 1231A are referenced to the top-left sample 1222, and samples in the second interpolation range 1230A are referenced throughout the second predicted block 1220.
[0197] Gradient Filter The gradient filter will be described below: Fig. 15 is a conceptual diagram used to explain an example of the reference range of the gradient filter in the first modification of this embodiment.
[0198] The gradient filter in this modification is a so-called 5-tap filter, and is expressed by the following equations (4) and (5).
[0199]
number
number
[0200] where I x k [x,y] represents the horizontal gradient value for each sample in the first gradient image when k is 0, and represents the horizontal gradient value for each sample in the second gradient image when k is 1. y k [x,y] represents the vertical gradient value for each sample in the first gradient image when k is 0, and represents the vertical gradient value for each sample in the second gradient image when k is 1.
[0201] For example, when the gradient filters in Equations (4) and (5) are applied to the top-left sample 1122 in Fig. 15, the horizontal gradient sample value is the sample value of five samples arranged in the horizontal direction including the top-left sample 1122, and is calculated by weighting and adding the sample values in the predicted image at whole sample precision. In this case, the weighting coefficients have positive or negative values for the samples above or below and to the left or right of the top-left sample 1122.
[0202] In this modification, the gradient filter for the top left sample 1122 references samples within a reference range 1132A. The reference range 1132A is cross-shaped, extending two samples above and below the top left sample 1122 and to the left and right.
[0203] This gradient filter is applied to all samples in the first prediction block 1120. Thus, samples within the first gradient reference range 1135A are referenced by the motion compensation filter for the first prediction block 1120.
[0204] This gradient filter is applied to the second predicted block 1220 in the same way as the first predicted block 1120. That is, samples in the reference range 1232A are referenced to the top-left sample 1222, and samples in the second gradient range 1235A are referenced throughout the second predicted block 1220.
[0205] If the motion vector indicating the reference range indicates a sub-sample position, the sample values in the reference ranges 1132A and 1232A of the gradient filter may be converted to sample values with sub-sample precision, and the gradient filter may be applied to the converted sample values. Alternatively, a gradient filter having coefficient values obtained by convolving coefficient values for conversion to sub-sample precision with coefficient values for deriving gradient values may be applied to the sample values with full sample precision. In this case, a different gradient filter is used for each sub-sample position.
[0206] [Deriving local motion estimation values for sub-blocks] The following is a description of the derivation of the local motion estimation value for a sub-block: In this example, the local motion estimation value is derived for the top left sub-block of the sub-blocks in the current block.
[0207] In this modification, a horizontal local motion estimation value u and a vertical local motion estimation value v are derived for each sub-block based on the following equation (6).
[0208]
number
[0209] where sG x , sG y , sG x 2 ,,sG y 2 , sG x dI and sG y dI is a value calculated in the sub-block and is derived based on the following equation (7).
[0210]
number
[0211] where Ω is the set of coordinates for all samples in the prediction sub-block in the region / area of the prediction block corresponding to the sub-block. x [i,j] denotes the sum of the horizontal gradient value of the first gradient image and the horizontal gradient value of the second gradient image, and G y [i,j] denotes the sum of the vertical gradient value of the first gradient image and the vertical gradient value of the second gradient image. ΔI[i,j] denotes the difference value between the first predicted image and the second predicted image. w[i,j] denotes a weighting factor that depends on the sample position in the predicted sub-block. However, the same weighting factor may be used for all samples in the predicted sub-block.
[0212] More specifically, G x [i,j], G y [i,j] and ΔI[i,j] are expressed by the following equation (8).
[0213]
number
[0214] In this way, local motion estimation values are derived at the sub-block level.
[0215] [Generating the final predicted image] The following is a description of how the final predicted image is generated. Each sample value p[x,y] in the final predicted image is calculated by multiplying the sample value I in the first predicted image by 0 [x,y] and sample value I in the second predicted image 1 It is derived based on the following equation (9) using [x, y].
[0216]
number
[0217] Here, b[x,y] represents the correction value for each sample. In equation (9), the sample value I 0 [x,y] and sample value I in the second predicted image 1 Each sample value p[x,y] in the final predicted image is calculated by shifting the sum of [x,y] and the correction value b[x,y] one bit to the right. The correction value b[x,y] is expressed by the following equation (10).
[0218]
number
[0219] In equation (10), the difference in horizontal gradient values between the first gradient image and the second gradient image (I x 0 [x,y]-I x 1 [x,y]) is multiplied by the horizontal local motion detection value (u), and the difference in vertical gradient values between the first gradient image and the second gradient image (I y 0 [x,y]-I y 1 The correction value b[x,y] is calculated by multiplying the vertical local motion detection value (v) by the vertical local motion detection value (v) and adding up the products.
[0220] The calculations described using equations (6) through (10) are merely examples. Any other equation may be used with the same effect.
[0221] [Effects, etc.] In this way, by using the motion compensation filter and gradient filter of this modification, local motion estimation values can be derived for each sub-block. The local motion estimation values of the sub-blocks derived in this way are used to generate a final predicted image for the current block, thereby achieving the same results as in this embodiment.
[0222] This embodiment may be combined with at least some aspects of other embodiments, and some processes, some elements of the device, and some syntax described in the flowcharts of this embodiment may be combined with aspects of other embodiments.
[0223] [Modification 2 of this embodiment] In this embodiment and Variation 1 of this embodiment, when deriving a local motion estimation value, all samples in a prediction sub-block in a prediction block corresponding to a sub-block in a current block are referenced. However, this disclosure is not limited to these examples. For example, only some samples in the prediction sub-block may be referenced. This scenario is described in the following paragraph as Variation 2 of this embodiment.
[0224] In the description of this modification, when deriving a local motion estimation value in a sub-block, only some samples in the prediction sub-block are referenced. For example, in equation (7) in modification 1, instead of the coordinate set Ω for all samples in the prediction sub-block, a coordinate set for some samples in the prediction sub-block is used. Various patterns can be used as the coordinate set for some samples in the prediction sub-block.
[0225] Fig. 16 is a diagram showing an example of a sample pattern referenced by deriving a local motion estimation value in Modification 2 of the embodiment. In Fig. 16, cross-hatched circles in the prediction block 1121 or 1221 indicate referenced samples, and non-cross-hatched circles indicate unreferenced samples.
[0226] 16(a) to 16(g), each of the seven sample patterns indicates several samples in the prediction sub-block 1121 or 1221. These seven sample patterns are all different.
[0227] In (a) to (c) of Figure 16, only 8 samples are referenced out of the 16 samples in the prediction sub-block 1121 or 1221. In (d) to (g) of Figure 16, only 4 samples are referenced out of the 16 samples in the prediction sub-block 1121 or 1221. That is, in (a) to (c) of Figure 16, 8 samples are thinned out out of the 16 samples, and in (d) to (g) of Figure 16, 12 samples are thinned out out of the 16 samples.
[0228] More specifically, in Fig. 16(a), eight samples, i.e., every other sample in the horizontal and vertical directions, are referenced. In Fig. 16(b), two pairs of samples, which are located on the left and right sides horizontally but alternately arranged vertically, are referenced. In Fig. 16(c), in prediction sub-block 1121 or 1221, four samples in the center and four samples at the corners are referenced.
[0229] In Figures 16(d) and (e), two samples are referenced in the first and third horizontal rows and the first and third vertical rows, respectively. In Figure 16(f), four corner samples are referenced. In Figure 16(g), four center samples are referenced.
[0230] A reference pattern may be adaptively selected from a plurality of predetermined patterns based on the two predicted images. For example, a sample pattern including samples corresponding to representative gradient values in the two predicted images may be selected. More specifically, if the representative gradient value is smaller than a threshold, a sample pattern including four samples (e.g., one of (d) to (g)) may be selected. If the representative gradient value is smaller than a threshold, a sample pattern including eight samples (e.g., one of (a) to (c)) may be selected.
[0231] When one sample pattern is selected from a plurality of sample patterns, a local motion estimation value for the subblock is derived by referring to the samples in the prediction subblock that indicate the selected sample pattern.
[0232] Alternatively, information indicating the selected sample pattern may be written to the bitstream. In this case, the decoding device may obtain information from the bitstream and select a sample pattern based on the obtained information. The information indicating the selected sample pattern may be written to a header for each block, slice, picture, or stream.
[0233] In this way, the encoding device and decoding device according to this embodiment can derive local motion estimation values for a sub-block by referring to only some samples in the prediction sub-block, thereby reducing the processing load or processing time compared to when all samples are referenced.
[0234] Furthermore, the encoding device and decoding device according to this embodiment can derive a local motion estimation value for a sub-block by referring to only samples in one sample pattern selected from multiple sample patterns. By switching the sample pattern, it is possible to refer to samples suitable for deriving a local motion estimation value for the sub-block, thereby reducing prediction errors.
[0235] This embodiment may be combined with at least some aspects of other embodiments, and some processes, some elements of the device, and some syntax described in the flowcharts of this embodiment may be combined with aspects of other embodiments.
[0236] [Another Modification of the Present Embodiment] The encoding device and the decoding device according to one or more aspects of the present disclosure have been described with reference to embodiments and variants of the embodiments. However, the present disclosure is not limited to these embodiments and variants of the embodiments. Those skilled in the art will be able to easily conceive of various modifications that can be made to these embodiments and variants of the embodiments without departing from the spirit and scope of the present disclosure and that fall within the scope of one or more aspects of the present disclosure.
[0237] For example, the number of taps in the motion compensation filter used in the example of this embodiment and the first modification of this embodiment is 8 samples. However, the present disclosure is not limited to this example. The number of taps in the motion compensation filter may be any number of taps as long as the interpolation reference range is included in the normal reference range.
[0238] The number of taps in the gradient filter used in the example of this embodiment and the first modified example of this embodiment is 5 or 6 samples. However, the present disclosure is not limited to this example. The number of taps in the gradient filter may be any number of taps as long as the gradient reference range is included in the interpolation reference range.
[0239] In the example of this embodiment and the first modification of this embodiment, the first gradient reference range and the second gradient reference range are included in the first interpolation reference range and the second interpolation reference range. However, the present disclosure is not limited to this example. For example, the first gradient reference range may coincide with the first interpolation reference range, and the second gradient reference range may coincide with the second interpolation reference range.
[0240] When deriving a local motion estimation value in a sub-block, sample values may be weighted so that values for samples at the center of the prediction sub-block are favored. Specifically, when deriving a local motion estimation value, values for multiple samples in the prediction sub-block may be weighted in both the first prediction block and the second prediction block. In this case, a larger weight may be given to a sample at the center of the prediction sub-block. That is, a sample at the center of the prediction sub-block may be weighted with a larger value than a value used for a sample outside the center of the prediction sub-block. More specifically, in Variation 1 of this embodiment, the weighting coefficient w[i,j] in Equation (7) may be used to weight coordinate values closer to the center of the prediction sub-block.
[0241] When deriving a local motion estimation value in a sub-block, samples belonging to adjacent prediction sub-blocks may be referenced. Specifically, in both the first prediction sub-block and the second prediction sub-block, samples in adjacent prediction sub-blocks may be additionally referenced in addition to samples in the prediction sub-block to derive a local motion estimation value in the sub-block.
[0242] The reference ranges of the motion compensation filter and the gradient filter in the examples of this embodiment and the first variant of this embodiment are for illustrative purposes only, and the present disclosure is not necessarily limited to these examples.
[0243] In the second modification of the present embodiment, seven sample patterns are given as an example. However, the present disclosure is not limited to these sample patterns. For example, sample patterns obtained by rotating each of the seven sample patterns may be used.
[0244] The values of the weighting coefficients in the first modification of this embodiment are merely examples, and the present disclosure is not limited to these examples. Furthermore, the block sizes and sub-block sizes in this embodiment and both modifications of this embodiment are also merely examples. This disclosure is not limited to the 8x8 sample size and the 4x4 sample size. Inter prediction can be performed using other sizes, similar to the example of this embodiment and both modifications of this embodiment.
[0245] This embodiment may be combined with at least some aspects of other embodiments, and some processes, some elements of the device, and some syntax described in the flowcharts of this embodiment may be combined with aspects of other embodiments.
[0246] As described with reference to Figures 4 and 5, for the reference range of surrounding samples to which the gradient filter is applied to obtain gradient values, the sample values in the reference block indicated by the motion vector assigned to the prediction block used in the prediction process may be compensated to a different value than the sample values described with reference to Figures 4 and 5.
[0247] For example, if the motion vector assigned to the prediction block indicates a fractional sample position, the sample value corresponding to the fractional sample position may be compensated and used.
[0248] However, if the filtering process references surrounding samples to obtain sample values corresponding to fractional sample positions, it may not be necessary to reference surrounding samples in a wider range than that described with respect to FIGS.
[0249] To solve the above problems, the present disclosure provides, for example, the following method.
[0250] FIG. 17 is a flow chart illustrating an example image encoding / decoding method 1700 that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture.
[0251] In an embodiment of the present application, the encoding method is implemented by an encoding device 100 as shown in Figure 1 and described in the corresponding description. In an embodiment, the encoding method is implemented by an inter prediction unit 126 of the encoding device 100 in cooperation with other components of the encoding device 100.
[0252] Similarly, in an embodiment of the present application, the decoding method is implemented by a decoding device 200 as shown in Figure 10 and described in the corresponding description. In an embodiment, the decoding method is implemented by an inter prediction unit 128 of the decoding device 200 in cooperation with other components of the decoding device 200.
[0253] As shown in FIG. 17, the encoding method according to this example includes the following steps.
[0254] Step 1702: Predict a first block of prediction samples for a current block of a picture, where predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture.
[0255] Step 1704: Pad the first block of prediction samples to form a second block of prediction samples, where the second block is larger than the first block.
[0256] Step 1706: The second block of prediction samples is used to calculate at least the gradient.
[0257] Step 1708: Encode the current block using at least the calculated gradient.
[0258] In an embodiment, another picture may have a picture order count (POC) that is different from the POC of the picture in the time domain of the stream. As shown in Figures 18A, 18B, and 18C, a picture may be set as the current picture, and the other picture may be set as the first picture or the second picture. For example, in Figure 18A, if the other picture is the first picture, the POC of the other picture is smaller than the POC of the picture in the left time domain but larger than the POC of the picture in the right time domain. In Figure 18B, if the other picture is the first picture, the POC of the other picture is smaller than the POC of the picture in both the left and right time domains. In Figure 18C, if the other picture is the first picture, the POC of the other picture is larger than the POC of the picture in both the left and right time domains.
[0259] In other embodiments, the other picture may be a coding reference picture that is temporally and / or spatially adjacent to the picture, and that may have a POC that is smaller or larger than the POC of the picture, temporally and / or spatially.
[0260] The current block may be arbitrarily selected within a picture. For example, the current block may be a square 4x4 sample block. The size of the current block may be changed to suit actual prediction accuracy needs.
[0261] Similar to the motion vector shown in Figure 5C, the motion vector in this example of the coding method points from the current block to another picture, which may be the first or second picture shown in Figures 18A to 18C. In embodiments, the motion vector may be received by the inter predictor 126 from another picture, especially if the other picture is a coded reference picture coded with the same motion vector.
[0262] In step 1702, the inter prediction unit 126 of the encoding device 100 predicts a first block of prediction samples for the current block of the picture using a motion vector from another picture, using at least the prediction process described above.
[0263] In an embodiment, the first block of prediction samples may be the same as the prediction block used in the prediction process as described above for a prediction mode such as merge mode or inter prediction mode.
[0264] In other embodiments, the first block of prediction samples may be the same as the prediction block used in the motion compensation process as described above for a prediction mode such as merge mode or inter prediction mode.
[0265] In step 1704, the inter prediction unit 126 of the encoding device 100 pads the first block of prediction samples to form a second block of prediction samples. The second block of prediction samples is larger than the first block of prediction samples. In this manner, more information is retained in the second block for subsequent processing, which may advantageously provide improved processing for the inter prediction method and thereby reduce memory bandwidth accesses for the inter prediction method, which is hardware friendly.
[0266] For example, step 1704 may be performed by inter prediction unit 126 of encoding device 100 by padding samples on at least two sides of the first block of prediction samples, where at least two sides of the first block of prediction samples are not orthogonal.
[0267] In some embodiments, the inter predictor 126 of the encoding device 100 may pad at least two sides of the first block of prediction samples in a parallel padding direction. As shown in Figure 19A, the inter predictor 126 of the encoding device 100 may pad two vertical sides of the first 4x4 block of prediction samples 1902 in a manner indicated by the dotted lines in Figure 19A. Alternatively, as shown in Figure 19B, the inter predictor 126 of the encoding device 100 may pad two horizontal sides of the first 4x4 block of prediction samples 1904 in a manner indicated by the dotted lines in Figure 19B.
[0268] In some other embodiments, the inter prediction unit 126 of the encoding device 100 may pad four sides of the first block of prediction samples, as illustrated in Figures 20A, 20B, 20C, and 20D. Figures 20A, 20B, 20C, and 20D also illustrate four padding methods that can be used to pad the first block of prediction samples. Those skilled in the art will appreciate that these padding methods may also be used to pad any sample block.
[0269] In an embodiment, as shown in Figure 20A, padding may be performed by mirroring a first block of predictive samples 2002 in a manner indicated by the dotted lines in Figure 20A. For example, inter predictor 126 of encoding device 100 may pad the first block of predictive samples 2002 by mirroring the first block of predictive samples 2002 symmetrically about each edge of the first block of predictive samples 2002. The mirroring may include symmetrically copying sample values (e.g., "A," "B," etc.) of the first block of predictive samples 2002 to samples surrounding the first block of predictive samples 2002 to form a second block of predictive samples 2000.
[0270] In another embodiment, as shown in Figure 20B, padding may be performed by copying the first block of prediction samples 2004 in a manner indicated by the dotted lines in Figure 20B. For example, the inter prediction unit 126 of the encoding device 100 may pad the first block of prediction samples 2004 by copying sample values (e.g., "A," "B," etc.) of samples located at each edge of the first block of prediction samples 2004 to corresponding samples surrounding and adjacent to each of the samples located at each edge of the first block of prediction samples 2004 to form the second block of prediction samples 2001.
[0271] The corresponding samples may be samples of one row adjacent to the first block. The fewer the number of rows of padded samples, the higher the accuracy of the second block 2001. The corresponding samples may be samples of multiple rows, including one row adjacent to the first block. Specifically, the corresponding samples are adjacent to each of the samples located on each edge of the first block 2004. The greater the number of rows of padded samples, the more memory bandwidth can be reduced. The number of rows of padded samples may be set based on, for example, the size of the first block, or may be determined in advance in accordance with a standard.
[0272] In another embodiment, as shown in FIG. 20C , padding may be performed by padding samples surrounding the first block of prediction samples with a fixed value (e.g., “Q” as shown in FIG. 20C ) in a manner indicated by the dotted line in FIG. 20C . For example, the fixed value may be selected from at least one of 0, 128, 512, a positive integer, the mean value of the first block of prediction samples, and the median value of the first block of prediction samples. The mean and median values may be associated with the sample values of the first block of prediction samples. In the example of FIG. 20C , the fixed value is a positive integer Q, and the inter prediction unit 126 of the encoding device 100 may pad the first block of prediction samples by padding samples surrounding the first block of prediction samples with the fixed value Q to form the second block of prediction samples 2050.
[0273] In other embodiments, padding may be performed by applying a function to the first block of prediction samples. Examples of functions may be filters, polynomial functions, exponential functions, clipping functions, etc. For simplicity, this example is not shown.
[0274] In yet another embodiment, as shown in Figure 20D, padding may be performed by any combination of mirroring, copying, padding a first value, and performing a function on the prediction samples, as described above. In the example of Figure 20D, the inter prediction unit 126 of the encoding device 100 may pad the first block of prediction samples 2008 by a combination of copying the first block of prediction samples 2008 to samples that are a first constant distance away from the first block 2008 and padding samples that are a second constant distance away from the first block 2008 with a fixed value Q to form a second block of prediction samples 2051.
[0275] In the example shown in Figures 20A to 20D, inter prediction unit 126 of encoding device 100 pads a first 4x4 block of prediction samples 2002, 2004, 2006, 2008 to form a second 8x8 block of prediction samples 2000, 2001, 2050, 2051.
[0276] Referring again to the examples of Figures 19A and 19B, it can be seen that inter prediction unit 126 of encoding device 100 pads 4x4 first blocks of prediction samples 1902, 1904 by mirroring first blocks of prediction samples 1902, 1904 to form second blocks of prediction samples 1900, 1901. In Figure 19A, second block of prediction samples 1900 is an 8x4 sample block. In Figure 19B, second block of prediction samples 1901 is a 4x8 sample block.
[0277] Therefore, those skilled in the art will appreciate that the size of the second block of prediction samples desired from the padding process can be varied depending on the actual prediction accuracy needs. Advantageously, the padding process reduces data for subsequent inter-prediction processing.
[0278] Therefore, the padding step 1704 of the present encoding method can be carefully designed to pad a first block of prediction samples of size M×N to form a second block of prediction samples of any desired size (M+d1)×(N+d2) based on the actual prediction accuracy needs based on the processes / techniques described above with respect to Figures 19A, 19B, and 20A-20D and / or different padding directions. Those skilled in the art will understand that M may be the same as or different from N, while d1 and d2 are greater than or equal to 0.
[0279] Thanks to the padding, small-sized data can be padded to a larger size so as to contain more information and generate an accurate prediction for the current block. In this way, the encoding method is advantageous in that it requires only small-sized data as input, thereby reducing memory bandwidth access for inter-prediction and providing a more hardware-friendly process.
[0280] In step 1706, the inter predictor 126 of the encoding device 100 calculates at least gradients using the second block of prediction samples. In an embodiment, when calculating at least gradients, the inter predictor 126 of the encoding device 100 may apply a gradient filter to the second block of prediction samples to generate at least derivative values to function as at least one gradient. An example of applying a gradient filter to generate at least one gradient is described above in the gradient filter section. In this manner, data in the current block of the current picture is referenced by data from a reference block of the reference picture being coded using respective gradients (i.e., derivative values) between the reference block and the current block. Such steps advantageously facilitate reducing memory bandwidth accesses for inter prediction and provide a more hardware-friendly process.
[0281] Two examples of gradient filters are shown in Figures 21A and 21B, respectively. In Figure 21A, one gradient filter {2, -9, 0, 9, -2} is applied to all sample values in the second block of predicted samples, regardless of the fractional part of the motion vector. In Figure 21A, the process of applying the gradient filter {2, -9, 0, 9, -2} is similar to that described above in the gradient filter section. In the example of Figure 21A, equations (4) and (5) are used for all motion vectors.
[0282] In Figure 21B, there are nine gradient filters applied to the sample values in the second block of prediction samples according to the fractional part of the motion vector: {8, -39, -3, 46, -17, 5}, {8, -32, -13, 50, -18, 5}, {7, -27, -20, 54, -19, 5}, {6, -21, -29, 57, -18, 5}, {4, -17, -36, 60, -15, 4}, {3, -9, -44, 61, -15, 4}, {1, -4, -48, 61, -13, 3}, {0, 1, -54, 60, -9, 2}, and {-1, 4, -57, 57, -4, 1}. In the example of Figure 21B, the fractional part of the motion vector is a subsample of the motion vector, e.g., 1 / 4 sample. Therefore, different gradient filters are selected based on the sub-sample portion of the motion vector. For example, if the horizontal motion vector is 1 + (1 / 4) samples, the fractional portion of the motion vector is 1 / 4. Then, gradient filters {4, -17, -36, 60, -15, 4} are selected to calculate the horizontal gradient. w[i] in equation (4) is replaced with {4, -17, -36, 60, -15, 4}. In another example, if the vertical motion vector is 4 + (7 / 16) samples, the fractional portion of the motion vector is 7 / 16. In this example, gradient filters {0, 1, -54, 60, -9, 2} are selected to calculate the vertical gradient. w[i] in equation (5) is replaced with {0, 1, -54, 60, -9, 2}.
[0283] Figures 21C and 21D show two other examples of gradient filters: In Figure 21C, one gradient filter {-1, 0, 1} is applied to all sample values in the second block of predicted samples, regardless of the fractional part of the motion vector.
[0284] In Figure 21D, one gradient filter {-1,1} is applied to all sample values in the second block of predicted samples, regardless of the fractional part of the motion vector. In Figures 21C and 21D, the process of applying each gradient filter {-1,0,1} or {-1,1} is similar to that described above in the gradient filter section.
[0285] In step 1708, the inter predictor 126 of the encoding device 100 encodes the current block of the picture using at least the calculated gradients.
[0286] As shown in example 1700 of Figure 17, the steps in the decoding method are similar to those in the encoding method, except for the final step 1708. In this final step, the encoding device 100 encodes the current block using at least the gradients, while the decoding device 200 decodes the current block using at least the gradients.
[0287] The above examples and embodiments use a predictive block (alternately referred to as a block of predictive samples) as input to predict the current block. Below, we further describe another embodiment of the encoding method, in which two predictive blocks are used as input to predict the current block.
[0288] When two prediction blocks (i.e., two blocks of prediction samples) are used as input to predict the current block, step 1702 as described above further includes predicting another block of prediction samples for the current block using another motion vector from another picture. The process of predicting the other block of prediction samples is the same as the process of predicting the first block of prediction samples. Furthermore, padding step 1704 as described above includes padding the other block of prediction samples to form a further block of prediction samples.
[0289] An embodiment 2200 of an image encoding / decoding method for predicting a current block using two reference pictures is shown in Fig. 22. In this embodiment, in addition to using an additional prediction block as input, an interpolation process is also developed for sub-sample accuracy. The embodiment includes the following steps:
[0290] Step 2202: Predict a first block and a second block of prediction samples for a current block of a picture using a prediction process that uses at least a first motion vector and a second motion vector from a different picture.
[0291] Step 2204: Pad the first and second blocks of prediction samples to form third and fourth blocks of prediction samples.
[0292] Step 2206: Perform interpolation processing on the third and fourth blocks of predicted samples.
[0293] Step 2208: The gradient is calculated using the result of the interpolation process.
[0294] Step 2210: The current block is coded / decoded using at least the result of the interpolation process and the calculated gradient.
[0295] Steps 2202 to 2210 are illustrated in a conceptual diagram 2300 in Figure 23. The encoding method embodiment 2200 as shown in Figures 22 and 23 can be performed by the image encoding device 100. Note that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 22 and 23.
[0296] In step 2202, a first block of prediction samples for a current block of a picture is predicted using at least a prediction process that predicts the samples using a first motion vector from another picture. In an embodiment, the first motion vector may point to a first picture, which is different from the current picture.
[0297] Similarly, a second block of prediction samples for the current block of the picture is predicted using at least a prediction process that predicts the samples using a second motion vector from another picture, which may point to a second picture that is different from the current picture.
[0298] In one example, the first picture may be different from the second picture, while in another example, the first picture may be the same as the second picture.
[0299] In one example, as shown in FIG. 18A, at least one of the picture order counts (POC) of the first picture and the second picture is smaller than the POC of the current picture, and the other of the picture order counts (POC) of the first picture and the second picture is larger than the POC of the current picture.
[0300] In another example, as shown in FIG. 18B, the POC of the first picture and the second picture may be smaller than the POC of the current picture.
[0301] In yet another example, the POC of the first and second pictures may be greater than the POC of the current picture, as shown in FIG. 18C.
[0302] The POCs of the first picture, the second picture, and the current picture shown in Figures 18A to 18C are in the time domain and indicate the coding order of these pictures in a data stream. The coding order may be the same as or different from the playback order of these pictures in a data stream (e.g., a video clip). Those skilled in the art will understand that the POCs may also refer to the order of these pictures in the spatial domain.
[0303] In one example, the first and second blocks of the prediction samples predicted in step 2202 may be the same as the reference blocks used in the motion compensation process performed for a prediction mode such as merge mode or inter prediction mode.
[0304] In one example, the first and second blocks of prediction samples may be equal to or greater than the current block.
[0305] As shown in Figure 23, an embodiment of the prediction step 2202 of Figure 22 is shown in step 2302. In this embodiment, the current block has a size of M x N, where M may be the same as or different from N. The inter prediction unit 126 of the encoding device 100 may predict a first block and a second block of prediction samples, each having a size of (M + d1) x (N + d2), where d1 and d2 may be equal to or greater than 0. An example of the first block of prediction samples is shown by a dotted line. The second block of prediction samples has the same size as the first block. It will be understood that the second block of prediction samples may have a size different from that of the first block of prediction samples.
[0306] Predicting the first and second blocks with a larger size prediction sample than the current block seems more advantageous since more information can be included in the prediction sample, contributing to more accurate prediction results.
[0307] In step 2204, the first block of prediction samples is padded to form a third block of prediction samples, and similarly, the second block of prediction samples is padded to form a fourth block of prediction samples.
[0308] In one example, the third block of prediction samples is larger than the first block of prediction samples, and the fourth block of prediction samples is larger than the second block of prediction samples. As shown in padding step 2304 of Figure 23, the third and fourth blocks each have a size of (M + d1 + d3) x (N + d2 + d4), where d3 and d4 are greater than 0. An example of the third block of prediction samples is shown by a dotted line. The fourth block of prediction samples has the same size as the third block. As mentioned above, it will be understood that the fourth block of prediction samples may have a different size than the third block of prediction samples.
[0309] In one example, as shown in FIG. 20A, the padding step may include mirroring the first and second blocks of prediction samples. In another example, as shown in FIG. 20B, the padding step may include duplicating the first and second blocks of prediction samples. In another example, the padding step may include padding a fixed value as shown in FIG. 20C. The fixed value may be at least one of 0, 128, 512, a positive integer, the mean value of the prediction samples, and the median value of the prediction samples. In some examples, the padding step may include performing a function on the prediction samples. Examples of functions may be a filter, a polynomial function, an exponential function, and a clipping function. In some examples, as shown in FIG. 20D, the padding step may include any combination of mirroring, duplication, padding a fixed value, and performing a function on the prediction samples.
[0310] In step 2206, an interpolation process is performed on the third and fourth blocks of prediction samples. The interpolation process may include applying an interpolation filter to the third and fourth blocks of prediction samples according to the first and second motion vectors, respectively.
[0311] In one example, the interpolation filter may be the same as the filter used in the motion compensation process performed for the prediction mode, such as merge mode, inter prediction mode, etc. In another example, the interpolation filter may be different from the filter used in the motion compensation process performed for the prediction mode, such as those described above.
[0312] As shown in step 2306 of FIG. 23, the interpolation process is performed on the third and fourth blocks of the prediction samples to generate resulting blocks. Each resulting block has a size of (M+d5)×(N+d6), where d5 and d6 are greater than 0. An example of a resulting block of the interpolation process on the third block of the prediction samples is shown by a dotted line. A similar example of a resulting block of the interpolation process on the fourth block of the prediction samples is shown. A portion of each resulting block of the interpolation process may be used to process the current block. Here, the size of this portion is M×N. This portion may be the same as the block resulting from the motion compensation process performed for a prediction mode such as merge mode or inter prediction mode.
[0313] In step 2208, the resulting interpolated blocks are used to calculate gradients. The gradient calculation may include applying a gradient filter to the sample values in the resulting interpolated blocks to generate derivative values. As shown in step 2308 of Figure 23, the calculated gradients may form blocks of size MxN.
[0314] In one example, as shown in FIG. 21A, there may be only one gradient filter applied to all sample values within the block resulting from the interpolation process, regardless of the fractional parts of the first and second motion vectors.
[0315] In one example, as shown in FIG. 21B, there may be nine gradient filters applied to the sample values in the block resulting from the interpolation process depending on the fractional parts of the first and second motion vectors.
[0316] In step 2210, the current block is coded using at least the blocks resulting from the interpolation process and the calculated gradients. An example of a coding block is shown in step 2310 of FIG.
[0317] Note that the terms "encoding" and "processing" used in the description of step 2210 for the encoding method performed by the image encoding device 100 can be replaced with the term "decoding" in step 2210 for the decoding method performed by the image decoding device 200.
[0318] In this embodiment, by introducing a padding process into inter prediction, the encoding method and decoding method advantageously reduce the memory bandwidth access of the inter prediction process. In addition, by using the padding process, the present embodiment can generate a sufficient prediction result by performing only one interpolation process. In this way, the present application advantageously removes additional interpolation processes from the inter prediction process, which is hardware-friendly.
[0319] Note that interpolation filters with different tap counts may be used in the interpolation process performed on the third block and the fourth block. For example, in the above example, only an 8-tap interpolation filter is used to obtain a block resulting from the interpolation process having a size of (M+d5)×(N+d6). However, the number of taps of the interpolation filter used to obtain a first region of size M×N within the block resulting from the interpolation process may be different from the number of taps of the interpolation filter used to obtain a second region, which is a region within the block resulting from the interpolation process other than the first region and is used to generate a gradient block. For example, the number of taps of the interpolation filter used to obtain the second region may be fewer than the number of taps of the interpolation filter used to obtain the first region. Specifically, if the number of taps of the interpolation filter used to obtain the first region is 8, the number of taps of the interpolation filter used to obtain the second region may be fewer than 8. This allows for reduced processing load while preventing degradation of image quality.
[0320] Another embodiment of an image encoding / decoding method for predicting a current block using two reference pictures is shown in Figure 24. In this embodiment, the interpolation process is performed before the padding process. The embodiment includes the following steps:
[0321] Step 2402: Predict a first block and a second block of prediction samples for a current block of a picture using a prediction process that uses at least a first motion vector and a second motion vector from a different picture.
[0322] Step 2404: Perform interpolation on the first and second blocks of predicted samples.
[0323] Step 2406: Pad the results of the interpolation process to form a third block and a fourth block of predicted samples.
[0324] Step 2408: Calculate the gradient using the third and fourth blocks of prediction samples.
[0325] Step 2410: The current block is coded / decoded using at least the result of the interpolation process and the calculated gradient.
[0326] Steps 2402 to 2410 are conceptually illustrated in Figure 25. The encoding method embodiment 2400 as shown in Figures 24 and 25 can be performed by the image encoding device 100. Note that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 24 and 25.
[0327] In step 2402, a first block of prediction samples for a current block of a picture is predicted using at least a prediction process that predicts samples for the current block using a first motion vector from another picture. In an embodiment, the first motion vector may point to a first picture, where the first picture is different from the current picture.
[0328] Similarly, a second block of prediction samples for the current block of the picture is predicted using at least a prediction process that predicts the samples using a second motion vector from another picture, which may point to a second picture that is different from the current picture.
[0329] In one example, the first picture may be different from the second picture, while in another example, the first picture may be the same as the second picture.
[0330] In one example, as shown in FIG. 18A, at least one of the picture order counts (POC) of the first picture and the second picture is smaller than the POC of the current picture, and the other of the picture order counts (POC) of the first picture and the second picture is larger than the POC of the current picture.
[0331] In another example, as shown in FIG. 18B, the POC of the first picture and the second picture may be smaller than the POC of the current picture.
[0332] In yet another example, the POC of the first and second pictures may be greater than the POC of the current picture, as shown in FIG. 18C.
[0333] The POCs of the first picture, the second picture, and the current picture shown in Figures 18A to 18C are in the time domain and indicate the coding order of these pictures in a data stream. The coding order may be the same as or different from the playback order of these pictures in a data stream (e.g., a video clip). Those skilled in the art will understand that the POCs may also refer to the order of these pictures in the spatial domain.
[0334] In one example, the first and second blocks of the prediction samples predicted in step 2402 may be the same as the reference blocks used in the motion compensation process performed for a prediction mode such as merge mode or inter prediction mode.
[0335] In one example, the first and second blocks of prediction samples may be equal to or greater than the current block.
[0336] As shown in Figure 25, an embodiment of the prediction step 2402 of Figure 24 is shown in step 2502. In this embodiment, the current block has a size of M x N, where M may be the same as or different from N. The inter prediction unit 126 of the encoding device 100 may predict a first block and a second block of prediction samples, each having a size of (M + d1) x (N + d2), where d1 and d2 may be equal to or greater than 0. An example of the first block of prediction samples is shown by a dotted line. The second block of prediction samples has the same size as the first block. As mentioned above, it will be understood that the second block of prediction samples may have a size different from that of the first block of prediction samples.
[0337] Predicting the first and second blocks with a larger size prediction sample than the current block seems more advantageous since more information can be included in the prediction sample, contributing to more accurate prediction results.
[0338] In step 2404, an interpolation process is performed on the first block and the second block of prediction samples. The interpolation process may include applying an interpolation filter to the first block and the second block of prediction samples according to the first motion vector and the second motion vector, respectively.
[0339] In one example, the interpolation filter may be the same as the filter used in the motion compensation process performed for the prediction mode, such as merge mode, inter prediction mode, etc. In another example, the interpolation filter may be different from the filter used in the motion compensation process performed for the prediction mode, such as those described above.
[0340] As shown in step 2504 of Figure 25, an interpolation process is performed on the first and second blocks of prediction samples to generate a resulting interpolation block. Each resulting interpolation block has a size of M x N. The resulting interpolation block may be used to process the current block, or may be the same as the resulting motion compensation block for a prediction mode such as merge mode or inter prediction mode.
[0341] In step 2406, the block resulting from the interpolation performed on the first block is padded to form a third block of predicted samples, which is larger than the current block. Similarly, the block resulting from the interpolation performed on the second block is padded to form a fourth block of predicted samples, which is larger than the current block.
[0342] As shown in step 2506 of Figure 25, an example third block of prediction samples may have a size of (M + d3) x (N + d4), where d3 and d4 are greater than 0. The fourth block of prediction samples has the same size. As mentioned above, it will be appreciated that the fourth block of prediction samples may have a different size than the third block of prediction samples.
[0343] In one example, as shown in FIG. 20A, the padding step may include mirroring the blocks resulting from the interpolation process. In another example, as shown in FIG. 20B, the padding step may include duplicating the blocks resulting from the interpolation process. In another example, the padding step may include padding with a fixed value as shown in FIG. 20C. The fixed value may be at least one of 0, 128, 512, a positive integer, the mean value of the predicted samples, and the median value of the predicted samples. In some examples, the padding step may include applying a function to the blocks resulting from the interpolation process. Examples of the function may include a filter, a polynomial function, an exponential function, and a clipping function. In another example, as shown in FIG. 26, the padding step may include using predicted samples of the first block and the second block.
[0344] In another example, the padding step may include applying a second interpolation filter to the first block and the second block, the second interpolation filter being different from the interpolation filter performed in step 2404, and the second interpolation filter having fewer taps than the interpolation filter performed in step 2404. In another example, as shown in Figure 20D, the padding step may include any combination of mirroring, copying, padding with a fixed value, performing a function on the blocks resulting from the interpolation process, padding samples with the first block and the second block, and applying a second interpolation filter to the first block and the second block.
[0345] In step 2408, a gradient is calculated using the third and fourth blocks of prediction samples. The gradient calculation may include applying a gradient filter to the sample values in the third and fourth blocks of prediction samples to generate derivative values. The gradient filter is as described above with respect to Figures 21A and 21B. As shown in step 2508 of Figure 25, the calculated gradients may form blocks of size MxN.
[0346] In step 2410, the current block is coded using at least the blocks resulting from the interpolation process and the calculated gradients. An example of a coding block is shown in step 2510 of FIG.
[0347] Note that the terms "encoding" and "processing" used in the description of step 2410 for the encoding method performed by the image encoding device 100 can be replaced with the term "decoding" in step 2410 for the decoding method performed by the image decoding device 200.
[0348] In this embodiment, by introducing a padding process into inter prediction, the encoding and decoding methods advantageously reduce memory bandwidth accesses in the inter prediction process. In addition, by using the padding process, the present embodiment can generate sufficient prediction results with only one interpolation process. In this way, the present application advantageously removes additional interpolation processes from the inter prediction process, which is hardware-friendly. This may maintain the same number of interpolation filter operations as those used in the motion compensation process performed for prediction modes such as merge mode and inter prediction mode.
[0349] FIG. 27A is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction for a current block of a picture.
[0350] Steps 2702A to 2708A of Figure 27A are shown in the conceptual diagram of Figure 28. An embodiment 2700A of the encoding method as shown in Figures 27A and 28 can be performed by the image encoding device 100. Note that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 27A and 28.
[0351] As shown in example 2700A, steps 2702A and 2704A are the same as steps 1702 and 1704 as described with respect to Figure 17. An embodiment of the first block of prediction samples predicted in step 2702A of Figure 27A is shown in step 2802 of Figure 28. In the embodiment shown in step 2802, a first block of prediction samples of size (M+d1) x (N+d2) is predicted for a current block of size M x N. Those skilled in the art will understand that M may be the same as or different from N, while d1 and d2 may be greater than, equal to, or less than 0.
[0352] An embodiment of a second block of prediction samples formed by padding the first block of prediction samples as per step 2704A of Figure 27A is shown in step 2804 of Figure 28. As shown in step 2804 of Figure 28, padding step 2704A of Figure 27A can be carefully designed to pad the first block of prediction samples of size (M+d1)×(N+d2) to form a second block of prediction samples having any desired size (M+d1+d3)×(N+d2+d4) based on actual prediction accuracy needs based on different padding directions and / or processes / techniques described above with respect to Figures 19A, 19B, and 20A-20D. Those skilled in the art will understand that d3 and d4 are greater than 0.
[0353] In step 2706A, an interpolation process is performed on the second block of prediction samples. The interpolation process may include applying an interpolation filter to the second block of prediction samples, respectively, in response to the first motion vector. In one example, the interpolation filter may be the same as the filter used in the motion compensation process performed for a prediction mode such as merge mode or inter prediction mode. In another example, the interpolation filter may be different from the filter used in the motion compensation process performed for a prediction mode such as those described above. An embodiment of the result of the interpolation process performed in step 2706A of FIG. 27A is shown in step 2806 of FIG. 28.
[0354] In step 2708A, the current block is coded using at least the blocks resulting from the interpolation process performed on the second block of predicted samples. An embodiment of the current block coded using the blocks resulting from the interpolation process as shown in step 2708A of Figure 27A is shown in step 2808 of Figure 28.
[0355] Note that the terms “encoding” and “processing” used in the description of step 2708A for the encoding method performed by image encoding device 100 can be replaced with the term “decoding” in step 2708A for the decoding method performed by image decoding device 200.
[0356] FIG. 27B is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture and neighboring blocks of the current block.
[0357] Steps 2702B to 2708B of Figure 27B are illustrated in the conceptual diagram of Figure 29. An embodiment 2700B of the encoding method as shown in Figures 27B and 29 can be performed by the image encoding device 100. Note that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 27B and 29.
[0358] In step 2701B, a first block of prediction samples is predicted for a current block of a picture, and the prediction step includes at least a prediction process using a first motion vector from another picture. As shown in Figure 29, an embodiment of the first block of prediction samples predicted in step 2702B of Figure 27B is shown in step 2902.
[0359] In step 2704B, the second block of predicted samples is padded to form a third block of predicted samples. The second block may be adjacent to the current block. An embodiment of the second block is shown in step 2904 of Figure 29. In this embodiment, the second block has a size of M x N, where M may be the same as or different from N.
[0360] As shown in step 2906 of Figure 29, the inter prediction unit 126 of the encoding device 100 may be configured to perform step 2704B and pad the second block of prediction samples to form a third block of prediction samples. The third block of prediction samples may have a size of (M + d3) x N, where d3 may be equal to or greater than 0. An example of the third block of prediction samples is shown by a dotted line.
[0361] In the example shown in Figure 29, the second block of predictive samples has the same size as the first block. It will be appreciated that the second block of predictive samples may have a different size than the first block of predictive samples.
[0362] 29, the second block is the left block adjacent to the first block. It should be understood that the second block may be at least one of the upper block, left block, right block, lower block, upper left block, upper right block, lower left block, and lower right block adjacent to the first block. An explanation of the position candidate for the second block is shown in FIG. 31, where the second block is adjacent to the current block.
[0363] In one example, as shown in FIG. 32A, padding in step 2704B may include mirroring the samples of the second block. In another example, as shown in FIG. 32B, padding in step 2704B may include duplicating the samples of the second block. In another example, as shown in FIG. 32C, padding in step 2704B may include padding with a fixed value. Here, the fixed value may be at least one of 0, 128, 512, a positive integer, the mean value of the second block, and the median value of the second block. In another example, padding in step 2704B may include performing a function on the samples of the second block. Examples of functions may include a filter, a polynomial function, an exponential function, and a clipping function. In another example, padding in step 2704B may include any combination of mirroring, duplicating, padding with a first value, and performing a function on the samples of the second block.
[0364] In the example of Figure 29, the padding in step 2704B may include padding samples on only one side of the second block, as shown in Figure 33A. In another example, the padding in step 2704B may include padding samples on two sides of the second block, where the two sides of the second block are parallel, as shown in Figure 33C, or orthogonal, as shown in Figure 33B. In another example, the padding in step 2704B may include padding samples on three or more sides of the second block, as shown in Figure 33D.
[0365] In step 2706B, an overlapped block motion compensation (OBMC) process is performed using at least the first and third blocks of prediction samples. The OBMC process is as described in the previous paragraphs of this application. As shown in step 2908 of Figure 29, the OBMC process generates an OBMC block OBMC.
[0366] In step 2708B, the current block is coded using at least the results of the OBMC process. An example of a coded block is shown in step 2910 of FIG.
[0367] Note that the terms "encoding" and "processing" used in the description of step 2708B for the encoding method performed by image encoding device 100 can be replaced with the term "decoding" in step 2708B for the decoding method performed by image decoding device 200.
[0368] In this embodiment, thanks to the introduction of padding into the OBMC process, the encoding and decoding methods advantageously reduce memory bandwidth accesses for inter prediction processes.
[0369] FIG. 27C is a flowchart illustrating yet another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture.
[0370] Steps 2702C to 2710C of Figure 27C are shown in the conceptual diagram of Figure 30. The encoding method embodiment 2700C as shown in Figures 27C and 30 can be performed by the image encoding device 100. Note that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 24 and 25.
[0371] As shown in example 2700C, steps 2702C and 2704C are the same as steps 1702 and 1704 as described with respect to FIG.
[0372] As shown in Figure 30, an embodiment of prediction step 2702C of Figure 27C is shown in step 3002. In this embodiment, the current block has a size of M x N, where M may be the same as or different from N. The inter prediction unit 126 of the encoding device 100 may predict a first block of prediction samples having a size of (M + d1) x (N + d2), where d1 and d2 may be equal to or greater than 0. As shown in step 3002 of Figure 30, the current block may have one or more neighboring blocks available for use later in this example of the method.
[0373] As shown in Figure 30, predicting the first block of prediction samples with a size larger than the current block seems more advantageous because more information can be included in the prediction sample, contributing to more accurate prediction results. An example of the first block of prediction samples is shown by the dotted line in step 3002 of Figure 30.
[0374] As shown in step 3004 of Figure 30, the inter prediction unit 126 of the encoding device 100 may be configured to perform step 2704C and pad the first block of prediction samples to form a second block of prediction samples. The second block of prediction samples may have a size of (M+d1+d3)×(N+d2+d4), where d3 and d4 may be equal to or greater than 0. An example of the second block of prediction samples is indicated by a thin dotted line in step 3004 of Figure 30, while an example of the first block of prediction samples is indicated by a thick dotted line.
[0375] As shown in Figure 30, forming a second block of prediction samples with a larger size than the first block seems more advantageous, since more information can be included in the padding block of prediction samples, contributing to more accurate prediction results.
[0376] In step 2706C, an interpolation process is performed on the second block of prediction samples. The interpolation process may include applying an interpolation filter to the second block of prediction samples in response to the first motion vector. In one example, the interpolation filter may be the same filter as that used in the motion compensation process performed for a prediction mode such as merge mode or inter prediction mode. In another example, the interpolation filter may be different from the filter used in the motion compensation process performed for the prediction mode described above.
[0377] As shown in step 3006 of Figure 30, an interpolation process performed on the second block of prediction samples may generate a resulting block of the interpolation process. The resulting block of the interpolation process may have a size of (M+d5) x (N+d6), where d5 and d6 are greater than 0. An example of a resulting block of the interpolation process performed on the third block of prediction samples is indicated by a dotted line in step 3006 of Figure 30. A portion of the resulting block of the interpolation process may be used in processing the current block. Here, the size of this portion may be M x N, as shown in step 3008 of Figure 30. This portion may be the same as a block resulting from a motion compensation process performed for a prediction mode such as merge mode or inter prediction mode.
[0378] In step 2708C, the current block is coded using at least the block resulting from the interpolation performed on the second block of predicted samples. An example of a coding block is shown in step 3008 of Figure 30.
[0379] Concurrently with, subsequent to, or prior to step 2708C, step 2710C performs OBMC processing on one or more neighboring blocks of the current block, which may use at least the blocks resulting from the interpolation processing.
[0380] The OBMC processing in step 2708C is as described in the preceding paragraphs of this application. As shown in step 2906 of Figure 29, the OBMC processing generates one or more OBMC blocks between one or more adjacent blocks and the current block. By using one or more OBMC blocks between one or more adjacent blocks and the current block generated at a time, the method of the present disclosure advantageously reduces memory bandwidth accesses (i.e., data fetched from off-chip memory, DRAM) required for the OBMC processing.
[0381] Note that the terms “encoding” and “processing” used in the description of step 2708C for the encoding method performed by image encoding device 100 can be replaced with the term “decoding” in step 2708C for the decoding method performed by image decoding device 200.
[0382] In the present disclosure, the blocks of prediction samples described in the above examples and embodiments may be replaced with non-rectangular shaped portions of prediction samples, examples of which may include at least one of triangular shaped portions, L-shaped portions, pentagonal shaped portions, hexagonal shaped portions, and polygonal shaped portions, as shown in Figure 34.
[0383] Those skilled in the art will understand that the non-rectangular shaped portions are not limited to the shaped portions shown in Figure 34. Furthermore, the shaped portions shown in Figure 34 may be freely combined.
[0384] FIG. 35 is a flowchart illustrating another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture.
[0385] Steps 3502 to 3508 of Figure 35 are illustrated conceptually in Figure 36. An embodiment 3500 of the encoding method as shown in Figures 35 and 36 can be performed by the image encoding device 100. It will be appreciated that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 35 and 36.
[0386] In step 3502, a first block of prediction samples for a current block of a picture is predicted, where the prediction step includes at least a prediction process using a first motion vector from another picture. As shown in Figure 36, an embodiment of the first block of prediction samples predicted in step 3502 of Figure 35 is shown in step 3602. In this embodiment, the current block has a size of M x N, where M may be the same as or different from N. The inter prediction unit 126 of the encoding device 100 may predict the first block of prediction samples having a size of (M + d1) x (N + d2), where d1 and d2 may be equal to or greater than 0.
[0387] In step 3504, a second motion vector for the current block is derived using at least the first motion vector through DMVR processing as described in the previous paragraph. An embodiment of the second motion vector is shown in step 3604 of Figure 36. In this embodiment, the second motion vector is represented by a dotted line pointing from the current block to the first block.
[0388] In step 3506, an interpolation process may be performed on the current block using the second motion vector. The interpolation process may include a padding process. In an embodiment of step 3506, the inter prediction unit 126 of the encoding device 100 may be configured to pad the first block of prediction samples according to the second motion vector to generate a second block of prediction samples, and to perform an interpolation process on the second block of prediction samples using at least the second motion vector. As shown in step 3606 of Figure 36, the second block of prediction samples may have a size of (M + d1 + d3) × (N + d2 + d4), where d3 and d4 may be equal to or greater than 0.
[0389] In the embodiment shown in Figure 36, the second block is an L-shaped portion adjacent to the first block, as shown in Figure 34. It will be understood that the second block may also be at least one of the top block, left block, right block, bottom block, top-left block, top-right block, bottom-left block, and bottom-right block adjacent to the first block. Alternatively, the second block may be a triangular, L-shaped, pentagonal, hexagonal, or polygonal portion adjacent to the first block.
[0390] 37A and 37B show two examples of the padding process in step 3506.
[0391] In the example shown in Figure 37A, the padding process in step 3506 may include padding samples on only one side of the first block using the second motion vector derived in step 3504 to form a second block of predicted samples.
[0392] In the example shown in Figure 37B, the padding process in step 3506 may include padding samples along two sides of the first block to form a second block of predicted samples using the second motion vector derived in step 3504. In the example of Figure 37B, the two sides of the first block are orthogonal, as also shown in Figure 33B. Alternatively, the two sides of the first block are parallel, as shown in Figure 33C. In another example, as shown in Figure 33D, the padding process in step 3506 may include padding samples along three or more sides of the second block.
[0393] In some examples, as shown in FIG. 32A, the padding process in step 3506 may include mirroring the samples of the first block. In other examples, as shown in FIG. 32B, the padding process in step 3506 may include duplicating the samples of the second block. In other examples, as shown in FIG. 32C, the padding process in step 3506 may include padding a fixed value. Here, the fixed value may be at least one of 0, 128, 512, a positive integer, the mean value of the second block, and the median value of the second block. In other examples, the padding process in step 3506 may include performing a function on the samples of the second block. Examples of the function may be a filter, a polynomial function, an exponential function, and a clipping function. In other examples, the padding process in step 3506 may include any combination of mirroring, duplicating, padding a first value, and performing a function on the samples of the second block.
[0394] As shown in step 3606, an interpolation process is performed on the second block of prediction samples. The interpolation process may include applying an interpolation filter to the second block of prediction samples in response to the second motion vector. In one example, the interpolation filter may be the same as the filter used in the motion compensation process performed for the prediction mode, such as merge mode or inter prediction mode. In another example, the interpolation filter may be different from the filter used in the motion compensation process performed for the prediction mode, such as the above.
[0395] An interpolation resultant block may be generated by performing an interpolation process on the second block of prediction samples, as shown in step 3606 of Figure 36. The size of the interpolation resultant block is MxN.
[0396] In step 3508, the current block is coded using at least the block resulting from the interpolation process performed on the second block of predicted samples in step 3506. An example of a coding block is shown in step 3608 of Figure 36.
[0397] Note that the terms "encoding" and "processing" used in the description of step 3508 for the encoding method performed by the image encoding device 100 can be replaced with the term "decoding" in step 3508 for the decoding method performed by the image decoding device 200.
[0398] In this embodiment, thanks to the introduction of padding into the OBMC process, the encoding and decoding methods advantageously reduce memory bandwidth access for the DMVR process.
[0399] FIG. 38 is a flowchart illustrating yet another alternative example of an image encoding / decoding method that uses inter prediction functionality to generate a prediction of a current block of a picture based on a reference block of another picture.
[0400] Steps 3802, 3804, 3806 and 3808 of Figure 38 are illustrated in the conceptual diagram of Figure 39. An embodiment 3800 of the encoding method as shown in Figures 38 and 39 can be performed by the image encoding device 100. It will be appreciated that the decoding method performed by the image decoding device 200 is the same as the encoding method performed by the image encoding device 100 as shown in Figures 38 and 39.
[0401] As mentioned above, steps 3802, 3804, 3806, and 3808 of Figure 38 are similar to those of Figure 35, except that the dynamic motion vector refreshing (DMVR) process in step 3804 further includes a padding process. In the embodiment shown in Figure 39, the padding process in step 3804 includes padding a first block of prediction samples according to a first motion vector. In this regard, step 3804 derives a second motion vector for the current block using at least the first motion vector based on the padded first block via the DMVR process as described in the previous paragraph. An embodiment of the second motion vector is illustrated in step 3904 of Figure 39 by a dotted line pointing from the current block to the first block.
[0402] Note that the terms "encoding" and "processing" used in the description of step 3808 for the encoding method performed by the image encoding device 100 can be replaced with the term "decoding" in step 3808 for the decoding method performed by the image decoding device 200.
[0403] In this embodiment, thanks to the introduction of padding into the OBMC process, the encoding and decoding methods advantageously reduce memory bandwidth access for the DMVR process.
[0404] In this application, the term "block" described in the above examples and embodiments may be replaced with the term "prediction unit." Also, the term "block" described in each aspect may be replaced with the term "sub-prediction unit." Also, the term "block" described in each aspect may be replaced with the term "coding unit."
[0405] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can usually be realized by an MPU (micro processing unit), memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memory. It is also possible to realize each functional block by hardware (dedicated circuitry). Various combinations of hardware and software may be employed.
[0406] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Furthermore, the processor that executes the program may be a single processor or multiple processors. In other words, centralized processing or distributed processing may be performed.
[0407] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.
[0408] Furthermore, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems implementing the application examples will be described. Such systems may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device including both. Other configurations of such systems can be appropriately changed depending on the situation.
[0409] [Usage example] 40 is a diagram showing the overall configuration of an appropriate content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.
[0410] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may connect any combination of the above devices. In various implementations, the devices may be connected to each other directly or indirectly via a telephone network or short-range wireless communication, without the intervention of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 may be connected to a terminal in a hotspot on an airplane ex117, etc., via a satellite ex116.
[0411] Note that wireless access points, hotspots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.
[0412] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handy-phone System) compatible with mobile communication systems such as 2G, 3G, 3.9G, 4G, and the upcoming 5G.
[0413] The home appliance ex114 is a refrigerator or an appliance included in a home fuel cell cogeneration system.
[0414] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content captured by a user using the terminal, may multiplex the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and may transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.
[0415] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client may be a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117, which is capable of decoding the encoded data. Each device that receives the distributed data may decode and play the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.
[0416] [Distributed processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented as a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server can be dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing latency. Furthermore, when certain types of errors occur or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the failed portion of the network, thereby achieving high-speed and stable distribution.
[0417] In addition to the distributed processing of the distribution itself, the encoding of captured data can be performed on each device, on the server side, or shared among devices. For example, encoding generally involves two processing loops. The first loop detects the image complexity or code size for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a device can perform the first encoding process, and the server that receives the content can perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each device. In this case, if there is a request for near-real-time reception and decoding, the data encoded by a device can be received and played back on another device, enabling more flexible real-time distribution.
[0418] As another example, the camera ex113 or the like extracts features (quantities of features or characteristics) from an image, compresses the data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning of the image (or the importance of the content), for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server recompresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a heavy processing load such as CABAC (context-adaptive binary arithmetic coding).
[0419] As another example, in a stadium, shopping mall, factory, etc., there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.
[0420] Since multiple video data are of nearly the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referential. The server may also receive encoded data from each terminal and change the reference relationships between multiple data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.
[0421] Furthermore, the server may perform transcoding to change the encoding format of the video data before distributing it. For example, the server may convert an MPEG-based encoding format into a VP-based encoding format (e.g., VP9), or convert H.264 to H.265.
[0422] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.
[0423] [3D, multi-angle] Images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other are increasingly being integrated for use. The videos taken by each device can be integrated based on the relative positions of the devices acquired separately or on areas where feature points in the videos match.
[0424] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a 3D shape of the scene based on not only the 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point clouds, etc., or may select or reconstruct images to be transmitted to the receiving terminal from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.
[0425] In this way, a user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting device, or can enjoy content in which a video from a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.
[0426] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, and may perform encoding that allows reference between the viewpoint images using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.
[0427] In the case of AR images, the server may superimpose virtual object information in the virtual space onto camera information in the real space based on the 3D position or the movement of the user's viewpoint. The decoding device may acquire or store virtual object information and 3D data, generate a 2D image according to the movement of the user's viewpoint, and smoothly connect the images to create superimposed data. Alternatively, the decoding device may send the user's viewpoint movement to the server in addition to a request for virtual object information. The server may create superimposed data based on the viewpoint movement received from the 3D data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data typically has an α value indicating transparency in addition to RGB. The server may set the α value of parts other than the object created from the 3D data to 0, for example, to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object. The predetermined RGB value may be predetermined.
[0428] Similarly, the decoding of distributed data may be performed by the client (e.g., a terminal), by the server, or by both parties. As an example, a terminal may first send a reception request to a server, and then another terminal may receive and decode content according to the request, and then transmit the decoded signal to a device with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-capable terminals themselves, it is possible to reproduce data with high image quality. As another example, large-sized image data may be received on a TV or the like, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of responsibility or an area they wish to view in more detail.
[0429] In situations where multiple short-, medium-, or long-range wireless communications are available, both indoors and outdoors, it may be possible to seamlessly receive content using distribution system standards such as MPEG-DASH. Users may freely select and switch between decoding and display devices, such as their own devices and indoor / outdoor displays, in real time. Decoding can also be performed by switching between decoding and display devices using location information. This allows information to be mapped and displayed on a part of the wall or ground of a neighboring building with an embedded display device while the user is moving toward their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as if the encoded data is cached on a server that can be quickly accessed from the receiving device or copied to an edge server in a content delivery service.
[0430] [Scalable Coding] Regarding content switching, we will explain it using a scalable stream, as shown in Figure 41, compressed and encoded using the video encoding method described in each of the above embodiments. The server may have multiple streams with the same content but different qualities as individual streams, or it may be configured to switch content by taking advantage of the characteristics of a temporally / spatially scalable stream, which is achieved by encoding the content separately into layers, as shown in the figure. In other words, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as communication bandwidth status, allowing the decoding side to freely switch between low-resolution content and high-resolution content. For example, if a user wants to continue watching a video they were watching on their smartphone ex115 while on the go on a device such as an Internet TV after returning home, the device can simply decode the same stream up to a different layer, thereby reducing the burden on the server.
[0431] Furthermore, as described above, pictures are coded for each layer, and in addition to the configuration in which scalability is achieved by an enhancement layer above the base layer, the enhancement layer may include meta-information based on image statistics, etc. The decoding side may generate high-quality content by super-resolving pictures in the base layer based on the meta-information. Super-resolution may improve the signal-to-noise ratio while maintaining and / or expanding the resolution. The meta-information may include information for specifying linear or non-linear filter coefficients used in super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least-squares calculations used in super-resolution processing.
[0432] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of objects in the image. The decoding side selects tiles to decode, thereby decoding only a portion of the area. Furthermore, by storing the object's attributes (such as a person, a car, or a ball) and its position in the video (such as a coordinate position in the same image) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in Figure 42, the meta information may be stored using a data storage structure different from pixel data, such as an SEI (supplemental enhancement information) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.
[0433] Meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in the video, and by combining the picture-by-picture information with the time information, it can identify the picture in which the object exists and determine the position of the object within the picture.
[0434] [Webpage optimization] FIG. 43 is a diagram showing an example of a web page display screen on a computer ex111 or the like. FIG. 44 is a diagram showing an example of a web page display screen on a smartphone ex115 or the like. As shown in FIGS. 43 and 44, a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device on which the page is viewed. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture contained in each content as a link image until the user explicitly selects the link image, or until the link image approaches the center of the screen or until the entire link image is within the screen, or may display a video such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the video.
[0435] When a link image is selected by a user, the display device performs decoding, for example, giving top priority to the base layer. Note that if the HTML constituting the web page contains information indicating that the content is scalable, the display device may also decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device decodes and displays only forward-referenced pictures (I pictures, P pictures, and forward-reference-only B pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Furthermore, the display device may intentionally ignore the picture reference relationships and roughly decode all B and P pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.
[0436] [Autonomous driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.
[0437] In this case, since a vehicle, drone, or airplane including a receiving terminal is moving, the receiving terminal can transmit location information of the receiving terminal, thereby realizing seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.
[0438] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.
[0439] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distribution companies, but also low-quality, short-duration content from individuals via unicast or multicast. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, using the following configuration.
[0440] During shooting, either in real time or after accumulating and shooting, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image data or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.
[0441] Personal content may contain content that, if left as is, violates copyright, moral rights, or portrait rights, and may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. Furthermore, the server may recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons. The server may replace the specified area with another image or blur the focus. If the image contains a person, the server may track the person in the video and replace the image of the person's face.
[0442] Because viewing personal content with small data volumes requires real-time performance, the decoding device may first receive the base layer as a top priority and decode and play it back, depending on the bandwidth. The decoding device may also receive the enhancement layer during this time, and if the content is played back more than once, such as when playback is looped, play back high-quality video including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smoother and the image quality improves. In addition to scalable encoding, a similar experience can also be provided by configuring a single stream consisting of a rough stream played the first time and a second stream that is encoded with reference to the first video.
[0443] [Other application examples] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 possessed by each terminal. The LSIex500 (see FIG. 40) may be a single chip or may be configured with multiple chips. It is also possible to incorporate video encoding or decoding software into some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by the computer ex111, and perform the encoding or decoding process using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data may be data encoded by the LSIex500 possessed by the smartphone ex115.
[0444] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal may download a codec or application software and then acquire and play the content.
[0445] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, the content supply system ex100 is more suited to multicast than the content supply system ex100, which is more suited to unicast, but similar applications are possible with regard to encoding and decoding processes.
[0446] [Hardware configuration] FIG. 45 is a diagram illustrating further details of the smartphone ex115 illustrated in FIG. 40. FIG. 46 is a diagram illustrating an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and other images received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing encoded data or decoded data such as captured video or still images, recorded voice, received video or still images, and email, and a slot unit ex464 that serves as an interface with a SIM ex468 for identifying users and authenticating access to various data, including networks. In addition, an external memory may be used instead of the memory unit ex467.
[0447] A main control unit ex460 that can comprehensively control the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.
[0448] When the power key is turned on by a user, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state and supplies power to each unit from the battery pack.
[0449] The smartphone ex115 processes calls, data communications, and other communications under the control of a main control unit ex460, which includes a CPU, ROM, RAM, and other components. During a call, the audio signal collected by the audio input unit ex456 is converted to a digital audio signal by the audio signal processing unit ex454, which then performs spectrum spread processing on the modulation / demodulation unit ex452. The resulting signal is then transmitted via the antenna ex450. The received data is then amplified, subjected to frequency conversion and analog-to-digital conversion, subjected to spectrum despreading processing on the modulation / demodulation unit ex452, and converted to an analog audio signal by the audio signal processing unit ex454, which then outputs the resulting signal from the audio output unit ex457. During data communications, text, still images, or video data can be transmitted via the operation input control unit ex462 under the control of the main control unit ex460, based on the operation of the main unit's operation unit ex466, etc. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the video encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the video or still image is being captured by the camera unit ex465, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the multiplexed video data and audio data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, before transmitting the multiplexed video data and audio data via the antenna ex450. The predetermined method may be determined in advance.
[0450] In the case of receiving video attached to an e-mail or chat, or video linked to a web page, for example, the multiplexed data received via the antenna ex450 is decoded by the multiplexing / demultiplexing unit ex453, which separates the multiplexed data into a video data bitstream and an audio data bitstream. The multiplexing / demultiplexing unit ex453 then demultiplexes the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and displays the video or still images contained in the linked video file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. As real-time streaming becomes increasingly common, audio playback may be socially inappropriate depending on the user's circumstances. Therefore, it is preferable that the initial setting be a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.
[0451] Although the smartphone ex115 has been used as an example, other implementations of the terminal are possible, such as a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed onto video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Furthermore, the video data itself may be received or transmitted instead of the multiplexed data.
[0452] While the main control unit ex460, which includes a CPU, has been described as controlling the encoding and decoding processes, various terminals often include a GPU. Therefore, a configuration in which a memory shared by the CPU and GPU, or a memory with addresses managed for common use, is also possible, leveraging the GPU's performance to process a large area at once. This shortens encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation and quantization processes at a picture level or other unit in the GPU rather than the CPU.
[0453] Those skilled in the art will appreciate that various modifications and / or variations may be made to the present disclosure as set forth in the specific embodiments without departing from the spirit or scope of the disclosure as broadly described. Accordingly, the present embodiments are to be considered illustrative and not restrictive.
[0454] The present disclosure provides various features, including:
[0455] 1. A coding device for coding a current block in a picture using inter prediction, comprising: a processor; a memory; The processor uses memory to By performing motion compensation using motion vectors corresponding to the two reference pictures, two predicted images are obtained from the two reference pictures; Obtain two gradient images corresponding to the two predicted images from the two reference pictures; deriving a local motion estimation value using two gradient images and two predicted images in a sub-block obtained by dividing the target block; An encoding device that generates a final predicted image for a current block using a local motion estimation value of a sub-block, two gradient images, and two predicted images.
[0456] 2. When two reference images are obtained with sub-pixel accuracy in the step of obtaining two predicted images, pixels are interpolated with sub-pixel accuracy by referring to pixels within an interpolation reference range surrounding the predicted block indicated by the motion vector in each of the two reference pictures; An encoding device according to statement 1, wherein the interpolation reference range is included in the normal reference range in order to perform motion compensation on the block to be processed in normal inter prediction, in which processing is performed using local motion detection values.
[0457] 3. The encoding device according to statement 2, wherein the interpolation reference range coincides with the normal reference range.
[0458] 4. An encoding device described in statement 2 or 3, wherein in the step of obtaining two gradient images, pixels are referenced within a gradient reference range surrounding the prediction block in each of the two reference pictures, and the gradient reference range is included in the interpolation reference range.
[0459] 5. The encoding device according to statement 4, wherein the gradient reference range corresponds to the interpolation reference range.
[0460] 6. An encoding device described in any one of statements 1 to 5, wherein in the step of deriving a local motion estimation value, the values of pixels included in a predicted sub-block within an area corresponding to a sub-block in each of two reference pictures are weighted and used, and among these pixels, pixels located in the center of the area are weighted with a value greater than the value used for pixels other than the center of the area.
[0461] 7. An encoding device as described in statement 6, in which, in the step of deriving a local motion detection value, in addition to pixels within a predictive sub-block of an area corresponding to a sub-block in each of two reference pictures, pixels within another predictive sub-block adjacent to the predictive sub-block included in the predictive block indicated by the motion vector are referenced.
[0462] 8. An encoding device described in any one of statements 6 to 7, in which, in the step of deriving a local motion estimation value, in each of two reference pictures, only some of the pixels in the predicted sub-block of the area corresponding to the sub-block are referenced.
[0463] 9. In the step of deriving a local motion estimate, (i) selecting, in each of two reference pictures, a pixel pattern from a plurality of different pixel patterns, the pixel pattern indicating some pixels in the prediction sub-block; (ii) referencing pixels in the prediction sub-block that exhibit the selected pixel pattern to derive a local motion estimate for the sub-block; 9. The encoding device of claim 8, wherein the processor is configured to write information indicating the selected pixel pattern to the bitstream.
[0464] 10. In the step of deriving a local motion estimate, (i) adaptively selecting a pixel pattern from a plurality of different pixel patterns representing some pixels in the prediction sub-block based on two predicted images in each of two reference pictures; (ii) The encoding device described in statement 8, wherein pixels in a predictive sub-block that exhibit a selected pixel pattern are referenced to derive a local motion estimation value for the sub-block.
[0465] 11. A coding method for coding a current block in a picture using inter prediction, comprising: By performing motion compensation using motion vectors corresponding to the two reference pictures, two predicted images are obtained from the two reference pictures; Obtain two gradient images corresponding to the two predicted images from the two reference pictures; deriving a local motion estimation value using two gradient images and two predicted images in a sub-block obtained by dividing the target block; An encoding method that generates a final predicted image for a current block using local motion estimation values of sub-blocks, two gradient images, and two predicted images.
[0466] 12. A decoding device for decoding a current block in a picture using inter prediction, comprising: a processor; a memory; The processor, using the memory, performs motion compensation using motion vectors corresponding to the two reference pictures, thereby obtaining two predicted images from the two reference pictures; Obtain two gradient images corresponding to the two predicted images from the two reference pictures; deriving a local motion estimation value using two gradient images and two predicted images in a sub-block obtained by dividing the target block; A decoding device that generates a final predicted image of a current block using a local motion estimation value of a sub-block, two gradient images, and two predicted images.
[0467] 13. A decoding device as described in statement 12, wherein when two reference images are obtained with sub-pixel accuracy in the step of obtaining two predicted images, pixels are interpolated with sub-pixel accuracy by referencing pixels within an interpolation reference range surrounding the predicted block indicated by the motion vector in each of the two reference pictures, and the interpolation reference range is included in the normal reference range in order to perform motion compensation on the block to be processed using normal inter prediction in which processing is performed using local motion detection values.
[0468] 14. The decoding device according to statement 13, wherein the interpolation reference range coincides with the normal reference range.
[0469] 15. A decoding device described in statement 13 or 14, wherein in the step of obtaining two gradient images, pixels are referenced within a gradient reference range surrounding the prediction block in each of two reference pictures, and the gradient reference range is included in the interpolation reference range.
[0470] 16. The decoding device of statement 15, wherein the gradient reference range coincides with the interpolation reference range.
[0471] 17. A decoding device described in any one of statements 12 to 16, wherein in the step of deriving a local motion estimation value, the values of pixels included in a predicted sub-block within an area corresponding to a sub-block in each of two reference pictures are weighted and used, and among these pixels, pixels located in the center of the area are weighted with a value greater than the value used for pixels other than the center of the area.
[0472] 18. A decoding device as described in statement 17, in which, in the step of deriving a local motion detection value, in addition to pixels within a predictive sub-block of an area corresponding to a sub-block in each of two reference pictures, pixels within another predictive sub-block adjacent to the predictive sub-block included in the predictive block indicated by the motion vector are referenced.
[0473] 19. A decoding device described in any one of statements 17 to 18, in which, in the step of deriving a local motion estimation value, in each of two reference pictures, only some of the pixels in the predicted sub-block of the area corresponding to the sub-block are referenced.
[0474] 20. The processor obtains information indicating the selected pixel pattern from the bitstream; In the step of deriving a local motion estimate, A decoding device as described in statement 19, which (i) selects a pixel pattern from a plurality of different pixel patterns indicating some pixels in a predictive sub-block in each of two reference pictures based on the acquired information, and (ii) references pixels in the predictive sub-block indicating the selected pixel pattern to derive a local motion estimation value for the sub-block.
[0475] 21. In the step of deriving a local motion estimate, (i) adaptively selecting a pixel pattern from a plurality of different pixel patterns representing some pixels in the prediction sub-block based on two predicted images in each of two reference pictures; (ii) A decoding device as described in statement 19, which references pixels in a predicted sub-block that exhibits a selected pixel pattern to derive a local motion estimation value for the sub-block.
[0476] 22. A decoding method for decoding a current block in a picture using inter prediction, comprising: By performing motion compensation using motion vectors corresponding to the two reference pictures, two predicted images are obtained from the two reference pictures; Obtain two gradient images corresponding to the two predicted images from the two reference pictures; deriving a local motion estimation value using two gradient images and two predicted images in a sub-block obtained by dividing the target block; A decoding method that generates a final predicted image for a current block using local motion estimation values of the sub-block, two gradient images, and two predicted images.
[0477] 23. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; Calculating at least the gradient using the second block of prediction samples; The image coding apparatus codes the current block using at least the calculated gradient.
[0478] 24. The image encoding device according to claim 23, wherein the first block of prediction samples is a prediction block used in a prediction process performed for a prediction mode that is a merge mode or an inter prediction mode.
[0479] 25. The image encoding device according to claim 23, wherein the first block of prediction samples is a reference block used in motion compensation processing performed for a prediction mode that is merge mode or inter prediction mode.
[0480] 26. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 24. The image encoding apparatus of claim 23, wherein at least two sides of the first block of predictive samples are padded to form a second block of predictive samples, and at least two sides of the first block are non-orthogonal.
[0481] 27. The circuit should, at least when calculating the gradient: 24. The image coding apparatus of claim 23, further comprising applying a gradient filter to the second block of prediction samples to generate at least a derivative value.
[0482] 28. The circuit includes a step of padding a first block of predicted samples to form a second block of predicted samples, the step comprising: 24. The image coding apparatus of claim 23, further comprising mirroring the predicted samples of the first block to form the second block of predicted samples.
[0483] 29. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 24. The image coding apparatus of claim 23, further comprising: copying predictive samples of the first block to form the second block of predictive samples.
[0484] 30. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 24. The image encoding device of claim 23, wherein the first block of predicted samples is padded with a fixed value to form a second block of predicted samples, and the fixed value may be 0, 128, a positive integer, the average value of the first block of predicted samples, or the median value of the first block of predicted samples.
[0485] 31. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 24. The image encoding device of claim 23, further comprising: performing a function on a first block of prediction samples to form a second block of prediction samples.
[0486] 32. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 24. The image encoding device of claim 23, further comprising: forming the second block of predicted samples by combining at least two of mirroring, duplicating, padding with a fixed value, and performing a function on the first block of predicted samples.
[0487] 33. The circuit, when predicting a first block of prediction samples for a current block of a picture, further comprises: 24. The image encoding device of claim 23, wherein the step of predicting another block of prediction samples for a current block of the picture and predicting another block of prediction samples includes at least a prediction process using another motion vector from another picture.
[0488] 34. The image coding apparatus of claim 33, wherein the other picture has a picture order count that is different from the picture order count of the other picture and / or the picture order count of the picture.
[0489] 35. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 35. The image encoding apparatus of claim 34, further comprising padding another block of prediction samples to form a further block of prediction samples.
[0490] 36. The image encoding device of claim 23, wherein after the circuit pads the first block of prediction samples to form the second block of prediction samples, the circuit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0491] 37. The image encoding device of claim 23, wherein before the circuit pads the first block of prediction samples to form the second block of prediction samples, the circuit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0492] 38. The circuit performs the following when encoding the current block using at least the calculated gradient: 24. The image coding device of claim 23, wherein the current block is coded using the block of predicted samples resulting from the interpolation process and at least the calculated gradients.
[0493] 39. The circuit should, at least when calculating the gradient: 39. The image coding apparatus of claim 38, further comprising applying one or more gradient filters to the block of prediction samples resulting from the interpolation process to generate one or more derivative values.
[0494] 40. Circuit and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; performing interpolation using the second block of predicted samples; An image encoding device that encodes a current block using at least a block resulting from an interpolation process.
[0495] 41. In operation, a circuit 41. The image encoding device according to claim 40, wherein an OBMC process is performed to predict one or more neighboring blocks of a current block, and the OBMC process uses at least a block resulting from an interpolation process.
[0496] 42. In operation, the circuit includes: when padding a first block of predicted samples to form a second block of predicted samples: 41. The image encoding device of claim 40, wherein two sides of a first block of predicted samples are padded with samples, the two sides of the first block being parallel to each other.
[0497] 43. In operation, the circuit includes: when padding a first block of predicted samples to form a second block of predicted samples: 41. The image encoding device of claim 40, wherein three or more sides of the first block of predicted samples are padded with samples.
[0498] 44. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the second block of predicted samples to form a third block of predicted samples, the second block being adjacent to the current block; performing OBMC processing using at least the first and third blocks of prediction samples; An image coding device that codes a current block using at least a block resulting from an OBMC process.
[0499] 45. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a first motion vector from another picture; deriving a second motion vector for the current block by DMVR processing using at least the first motion vector; performing an interpolation process on the current block using the second motion vector, the interpolation process including a padding process; An image encoding device that encodes a current block using at least a block resulting from an interpolation process.
[0500] 46. When the circuit performs interpolation on the current block, padding the first block of predicted samples according to the second motion vector to generate a second block of predicted samples; 46. The image coding device of claim 45, wherein the interpolation is performed using at least the second block of prediction samples.
[0501] 47. The image encoding device of claim 46, wherein the circuit, when padding the first block of prediction samples to generate the second block of prediction samples, pads one or more sides of the first block of prediction samples according to the second motion vector.
[0502] 48. The image encoding device of claim 45, wherein the circuit performs padding on the first block of prediction samples in response to the first motion vector when DMVR processing is used to derive the second motion vector.
[0503] 49. A division unit that receives an original picture and divides it into blocks in operation; a first addition unit that receives the blocks from the division unit, predicts them from a prediction control unit, subtracts each predicted value from a corresponding block, and outputs a residual; a transform unit that transforms the residual output from the adder unit and outputs a transform coefficient; a quantization unit that, in operation, quantizes the transform coefficients to generate quantized transform coefficients; an entropy coding unit that, in operation, codes the quantized transform coefficients to generate a bitstream; an inverse quantization transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; a second adder that, in operation, adds the residual output from the inverse quantization transform unit and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: an inter prediction unit configured to generate a prediction of a current block based on a reference block in a previously coded reference picture; and a prediction control unit coupled to a memory; In operation, when generating a prediction of a current block based on a reference block in a previously coded reference picture, the inter predictor: predicting a first block of prediction samples for a current block, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from a coded reference picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; Calculating at least the gradient using the second block of prediction samples; The image coding apparatus codes the current block using at least the calculated gradient.
[0504] 50. The image encoding device of claim 49, wherein the first block of prediction samples is a prediction block used in a prediction process performed for a prediction mode that is merge mode or inter prediction mode.
[0505] 51. The image encoding device of claim 49, wherein the first block of prediction samples is a reference block used in motion compensation processing performed for a prediction mode that is merge mode or inter prediction mode.
[0506] 52. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 50. The image encoding apparatus of claim 49, wherein at least two sides of the first block of predictive samples are padded to form a second block of predictive samples, and wherein at least two sides of the first block are not orthogonal.
[0507] 53. The inter prediction unit, at least when calculating gradients, 50. The image encoding apparatus of claim 49, further comprising applying a gradient filter to the second block of prediction samples to generate at least a derivative value.
[0508] 54. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 50. The image encoding apparatus of claim 49, further comprising mirroring the predicted samples of the first block to form the second block of predicted samples.
[0509] 55. When the inter prediction unit pads the first block of prediction samples to form the second block of prediction samples, 50. The image encoding apparatus of claim 49, further comprising: copying the predictive samples of the first block to form the second block of predictive samples.
[0510] 56. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 50. The image encoding apparatus of claim 49, wherein the first block of predicted samples is padded with a fixed value to form a second block of predicted samples, and the fixed value may be 0, 128, a positive integer, the average value of the first block of predicted samples, or the median value of the first block of predicted samples.
[0511] 57. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 50. The image encoding apparatus of claim 49, further comprising: performing a function on a first block of prediction samples to form a second block of prediction samples.
[0512] 58. When the inter predictor pads the first block of predicted samples to form the second block of predicted samples, 50. The image encoding apparatus of claim 49, further comprising: combining at least two of mirroring, duplicating, padding a first value, and performing a function on the first block of predicted samples to form the second block of predicted samples.
[0513] 59. When predicting a first block of prediction samples for a current block of a picture, the inter prediction unit further 50. The image encoding device of claim 49, wherein the step of predicting another block of prediction samples for the current block and predicting another block of prediction samples includes at least a prediction process using another motion vector from another coded reference picture.
[0514] 60. The image encoding device of claim 59, wherein another encoded reference picture has a picture order count that is different from the picture order count of the encoded reference picture and / or the picture order count of the original picture.
[0515] 61. When the inter predictor pads the first block of predicted samples to form the second block of predicted samples, it further comprises: 61. The image encoding apparatus of claim 60, further comprising padding another block of prediction samples to form a further block of prediction samples.
[0516] 62. The image encoding device of claim 49, wherein after the inter prediction unit pads the first block of prediction samples to form the second block of prediction samples, the inter prediction unit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0517] 63. The image encoding device of claim 49, wherein before the inter prediction unit pads the first block of predicted samples to form the second block of predicted samples, the inter prediction unit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0518] 64. When encoding a current block using at least the calculated gradients, the inter predictor: 50. The image coding device of claim 49, wherein the current block is coded using the block of prediction samples resulting from the interpolation process and at least the calculated gradients.
[0519] 65. The inter prediction unit, at least when calculating gradients, 65. An image encoding apparatus as claimed in claim 64, further comprising applying one or more gradient filters to the block of prediction samples resulting from the interpolation process to generate one or more derivative values.
[0520] 66. A division unit, in operation, receives an original picture and divides it into blocks; a first addition unit, in operation, receives the blocks from the division unit, predicts them from a prediction control unit, subtracts each prediction from a corresponding block, and outputs a residue; a transform unit that transforms the residual output from the adder unit and outputs a transform coefficient; a quantization unit that, in operation, quantizes the transform coefficients to generate quantized transform coefficients; an entropy coding unit that, in operation, codes the quantized transform coefficients to generate a bitstream; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; a second adder that, in operation, adds the residual output from the inverse quantization transform unit and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: an inter prediction unit configured to generate a prediction of a current block based on a reference block in a previously coded reference picture; and a prediction control unit coupled to a memory; In operation, when generating a prediction of a current block based on a reference block in a previously coded reference picture, the inter predictor: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; performing interpolation using the second block of predicted samples; An image encoding device that encodes a current block using at least a block resulting from an interpolation process.
[0521] 67. The inter prediction unit operates as follows: 67. The image encoding device according to claim 66, wherein an OBMC process is performed to predict one or more neighboring blocks of a current block, and the OBMC process uses at least a block resulting from an interpolation process.
[0522] 68. In operation, the inter predictor may pad a first block of predicted samples to form a second block of predicted samples by: 67. The image encoding apparatus of claim 66, wherein two sides of a first block of predicted samples are padded with samples, the two sides of the first block being parallel to each other.
[0523] 69. In operation, the inter predictor may pad a first block of predicted samples to form a second block of predicted samples by: 67. The image encoding device of claim 66, wherein samples are padded on three or more sides of the first block of predicted samples.
[0524] 70. A division unit, in operation, receives an original picture and divides it into blocks; a first addition unit, in operation, receives the blocks from the division unit, predicts them from a prediction control unit, subtracts each prediction from a corresponding block, and outputs a residual; a transform unit that transforms the residual output from the adder unit and outputs a transform coefficient; a quantization unit that, in operation, quantizes the transform coefficients to generate quantized transform coefficients; an entropy coding unit that, in operation, codes the quantized transform coefficients to generate a bitstream; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; a second adder that, in operation, adds the residual output from the inverse quantization transform unit and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: an inter prediction unit configured to generate a prediction of a current block based on a reference block in a previously coded reference picture; and a prediction control unit coupled to a memory; In operation, when generating a prediction of a current block based on a reference block in a previously coded reference picture, the inter predictor: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the second block of predicted samples to form a third block of predicted samples, the second block being adjacent to the current block; performing OBMC processing using at least the first and third blocks of prediction samples; An image coding device that codes a current block using at least a block resulting from an OBMC process.
[0525] 71. A division unit, in operation, receives an original picture and divides it into blocks; a first addition unit, in operation, receives the blocks from the division unit, predicts them from a prediction control unit, subtracts each prediction from a corresponding block, and outputs a residue; a transform unit that transforms the residual output from the adder unit and outputs a transform coefficient; a quantization unit that, in operation, quantizes the transform coefficients to generate quantized transform coefficients; an entropy coding unit that, in operation, codes the quantized transform coefficients to generate a bitstream; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; a second adder that, in operation, adds the residual output from the inverse quantization transform unit and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: an inter prediction unit configured to generate a prediction of a current block based on a reference block in a previously coded reference picture; and a prediction control unit coupled to a memory; In operation, when generating a prediction of a current block based on a reference block in a previously coded reference picture, the inter predictor: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a first motion vector from another picture, and deriving a second motion vector for the current block by a DMVR process using at least the first motion vector; performing an interpolation process on the current block using the second motion vector, the interpolation process including a padding process; An image encoding device that encodes a current block using at least a block resulting from an interpolation process.
[0526] 72. When performing interpolation processing on a current block, the inter prediction unit: padding the first block of predicted samples according to the second motion vector to generate a second block of predicted samples; 72. The image coding device of claim 71, wherein the interpolation is performed using at least the second block of prediction samples.
[0527] 73. An image encoding device as described in claim 72, wherein the inter prediction unit, when padding a first block of prediction samples to generate a second block of prediction samples, pads one or more sides of the first block of prediction samples according to the second motion vector.
[0528] 74. The image encoding device according to claim 71, wherein the inter prediction unit, when deriving the second motion vector using DMVR processing, performs padding processing on the first block of predicted samples in accordance with the first motion vector.
[0529] 75. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; Calculating at least the gradient using the second block of prediction samples; An image decoding device that decodes a current block using at least the calculated gradient.
[0530] 76. The image decoding device according to claim 75, wherein the first block of prediction samples is a prediction block used in a prediction process performed for a prediction mode that is a merge mode or an inter prediction mode.
[0531] 77. The image decoding device according to claim 75, wherein the first block of prediction samples is a reference block used in motion compensation processing performed for a prediction mode that is merge mode or inter prediction mode.
[0532] 78. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding apparatus of claim 75, wherein at least two sides of a first block of predictive samples are padded to form a second block of predictive samples, and at least two sides of the first block are not orthogonal.
[0533] 79. The circuit should, at least when calculating the gradient: 76. The image decoding apparatus of claim 75, further comprising: applying a gradient filter to the second block of prediction samples to generate at least a derivative value.
[0534] 80. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding apparatus of claim 75, further comprising: mirroring the predicted samples of the first block to form the second block of predicted samples.
[0535] 81. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding apparatus of claim 75, further comprising: copying predictive samples of the first block to form the second block of predictive samples.
[0536] 82. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding device of claim 75, wherein a first block of predicted samples is padded with a fixed value to form a second block of predicted samples, and the fixed value may be 0, 128, a positive integer, the average value of the first block of predicted samples, or the median value of the first block of predicted samples.
[0537] 83. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding apparatus of claim 75, further comprising: performing a function on a first block of prediction samples to form a second block of prediction samples.
[0538] 84. The circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 76. The image decoding apparatus of claim 75, wherein the second block of predictive samples is formed by combining at least two of mirroring, duplicating, padding the first value, and performing a function on the first block of predictive samples.
[0539] 85. The circuit, when predicting a first block of prediction samples for a current block of a picture, further comprises: 76. The image decoding device of claim 75, wherein the step of predicting another block of prediction samples for a current block of the picture and predicting another block of prediction samples includes at least a prediction process using another motion vector from another picture.
[0540] 86. The image decoding apparatus of claim 85, wherein the other picture has a picture order count that is different from the picture order count of the other picture and / or the picture order count of the picture.
[0541] 87. The circuit, when padding the first block of predicted samples to form a second block of predicted samples, further comprises: 87. The image decoding apparatus of claim 86, further comprising padding another block of predictive samples to form a further block of predictive samples.
[0542] 88. The image decoding device of claim 75, wherein after the circuit pads the first block of prediction samples to form the second block of prediction samples, the circuit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0543] 89. The image decoding device of claim 75, wherein before the circuit pads the first block of prediction samples to form the second block of prediction samples, the circuit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0544] 90. The circuit performs the following when decoding the current block using at least the calculated gradient: 76. The image decoding device of claim 75, wherein the current block is decoded using the block of prediction samples resulting from the interpolation process and at least the calculated gradients.
[0545] 91. The circuit should, at least when calculating the gradient: 91. An image decoding apparatus as claimed in claim 90, further comprising: applying one or more gradient filters to a block of prediction samples resulting from the interpolation process to generate one or more derivative values.
[0546] 92. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; performing interpolation using the second block of predicted samples; An image decoding device that decodes a current block using at least a block resulting from an interpolation process.
[0547] 93. A circuit, in operation, 93. The image decoding device according to claim 92, wherein an OBMC process is performed to predict one or more neighboring blocks of a current block, and the OBMC process uses at least a block resulting from an interpolation process.
[0548] 94. In operation, the circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 93. The image decoding apparatus of claim 92, wherein samples are padded on two sides of a first block of predicted samples, the two sides of the first block being parallel to each other.
[0549] 95. In operation, the circuit includes, when padding a first block of predicted samples to form a second block of predicted samples: 93. The image decoding device of claim 92, wherein samples are padded on three or more sides of the first block of predicted samples.
[0550] 96. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the second block of predicted samples to form a third block of predicted samples, the second block being adjacent to the current block; performing OBMC processing using at least the first and third blocks of prediction samples; An image decoding device that decodes a current block using at least a block resulting from an OBMC process.
[0551] 97. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a first motion vector from another picture; deriving a second motion vector for the current block by DMVR processing using at least the first motion vector; performing an interpolation process on the current block using the second motion vector, the interpolation process including a padding process; An image decoding device that decodes a current block using at least a block resulting from an interpolation process.
[0552] 98. When the circuit performs interpolation on the current block, padding the first block of predicted samples according to the second motion vector to generate a second block of predicted samples; 98. An image decoding device as claimed in claim 97, wherein interpolation is performed using at least the second block of prediction samples.
[0553] 99. The image decoding device of claim 98, wherein when padding the first block of predictive samples to generate the second block of predictive samples, the circuit pads one or more sides of the first block of predictive samples according to the second motion vector.
[0554] 100. The image decoding device of claim 97, wherein the circuit performs padding on the first block of predicted samples according to the first motion vector when deriving the second motion vector using DMVR processing.
[0555] 101. An entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain quantized transform coefficients; an inverse quantization transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; an adder that, in operation, adds the residual output from the inverse quantization transformer and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: a prediction control unit coupled to an inter prediction unit that generates a prediction of a current block of a picture based on a reference block in a decoded reference picture and a memory; In operation, when generating a prediction of a current block based on a reference block in a decoded reference picture, the inter predictor: predicting a first block of prediction samples for a current block, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from a decoded reference picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; Calculating at least the gradient using the second block of prediction samples; An image decoding device that decodes a current block using at least the calculated gradient.
[0556] 102. The image decoding device according to claim 101, wherein the first block of prediction samples is a prediction block used in a prediction process performed for a prediction mode that is a merge mode or an inter prediction mode.
[0557] 103. The image decoding device according to claim 101, wherein the first block of prediction samples is a reference block used in motion compensation processing performed for a prediction mode that is merge mode or inter prediction mode.
[0558] 104. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. The image decoding apparatus of claim 101, wherein at least two sides of the first block of predictive samples are padded to form a second block of predictive samples, and at least two sides of the first block are not orthogonal.
[0559] 105. The inter prediction unit performs the following at least when calculating gradients: 102. The image decoding apparatus of claim 101, further comprising: applying a gradient filter to the second block of prediction samples to generate at least a derivative value.
[0560] 106. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. The image decoding apparatus of claim 101, further comprising mirroring the predictive samples of the first block to form the second block of predictive samples.
[0561] 107. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. The image decoding apparatus of claim 101, further comprising: copying predictive samples of a first block to form a second block of predictive samples.
[0562] 108. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. The image decoding device of claim 101, wherein a first block of predicted samples is padded with a fixed value to form a second block of predicted samples, and the fixed value may be 0, 128, a positive integer, the average value of the first block of predicted samples, or the median value of the first block of predicted samples.
[0563] 109. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. An image decoding apparatus as claimed in claim 101, further comprising: performing a function on a first block of prediction samples to form a second block of prediction samples.
[0564] 110. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 102. The image decoding apparatus of claim 101, wherein forming the second block of predictive samples combines at least two of mirroring, duplicating, padding the first value, and performing a function on the first block of predictive samples.
[0565] 111. When predicting a first block of prediction samples for a current block of a picture, the inter prediction unit further 102. The image decoding device of claim 101, wherein the step of predicting another block of prediction samples for the current block and predicting another block of prediction samples includes at least a prediction process using another motion vector from another decoded reference picture.
[0566] 112. The image decoding apparatus of claim 111, wherein another decoded reference picture has a picture order count that is different from the picture order count of the decoded reference picture and / or the picture order count of the original picture.
[0567] 113. When padding a first block of predicted samples to form a second block of predicted samples, the inter predictor: 113. The image decoding apparatus of claim 112, further comprising padding another block of predictive samples to form a further block of predictive samples.
[0568] 114. The image decoding device of claim 101, wherein after the inter prediction unit pads the first block of prediction samples to form the second block of prediction samples, the inter prediction unit performs interpolation processing for a prediction mode that is a merge mode or an inter prediction mode.
[0569] 115. The image decoding device of claim 101, wherein before the inter prediction unit pads the first block of prediction samples to form the second block of prediction samples, the inter prediction unit performs an interpolation process for a prediction mode that is a merge mode or an inter prediction mode.
[0570] 116. When decoding a current block using at least the calculated gradients, the inter prediction unit: 102. The image decoding device of claim 101, wherein the current block is decoded using the block of prediction samples resulting from the interpolation process and at least the calculated gradients.
[0571] 117. The inter prediction unit performs at least the following when calculating gradients: 117. An image decoding device as claimed in claim 116, further comprising applying one or more gradient filters to the block of prediction samples resulting from the interpolation process to generate one or more derivative values.
[0572] 118. An entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain quantized transform coefficients; an inverse quantization transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; an adder that, in operation, adds the residual output from the inverse quantization transformer and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: a prediction control unit coupled to an inter prediction unit that generates a prediction of a current block of a picture based on a reference block in a decoded reference picture and a memory; In operation, when generating a prediction of a current block based on a reference block in a decoded reference picture, the inter predictor: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the first block of predicted samples to form a second block of predicted samples, the second block being larger than the first block; performing interpolation using the second block of predicted samples; An image decoding device that decodes a current block using at least a block resulting from an interpolation process.
[0573] 119. A circuit, in operation, 119. The image decoding device according to claim 118, wherein an OBMC process is performed to predict one or more neighboring blocks of a current block, and the OBMC process uses at least a block resulting from an interpolation process.
[0574] 120. In operation, the circuit includes: when padding a first block of predicted samples to form a second block of predicted samples: 119. The image decoding apparatus of claim 118, wherein samples are padded on two sides of a first block of predicted samples, the two sides of the first block being parallel to each other.
[0575] 121. In operation, the circuit includes: when padding a first block of predicted samples to form a second block of predicted samples: 119. The image decoding device of claim 118, wherein samples are padded on three or more sides of the first block of predicted samples.
[0576] 122. An entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain quantized transform coefficients; an inverse quantization transform unit that, in operation, inverse quantizes the quantized transform coefficients to obtain transform coefficients and inverse transforms the transform coefficients to obtain residuals; an adder that, in operation, adds the residual output from the inverse quantization transformer and the predicted value output from the prediction control unit to reconstruct a block; In operation, the method includes: a prediction control unit coupled to an inter prediction unit that generates a prediction of a current block of a picture based on a reference block in a decoded reference picture and a memory; In operation, when generating a prediction of a current block based on a reference block in a decoded reference picture, the inter predictor: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a motion vector from another picture; padding the second block of predicted samples to form a third block of predicted samples, the second block being adjacent to the current block; performing OBMC processing using at least the first and third blocks of prediction samples; An image decoding device that decodes a current block using at least a block resulting from an OBMC process.
[0577] 123. Circuits and a memory connected to the circuit; In operation, the circuit: predicting a first block of prediction samples for a current block of a picture, wherein the step of predicting the first block of prediction samples includes at least a prediction process using a first motion vector from another picture; deriving a second motion vector for the current block by DMVR processing using at least the first motion vector; performing an interpolation process on the current block using the second motion vector, the interpolation process including a padding process; An image decoding device that decodes a current block using at least a block resulting from an interpolation process.
[0578] 124. When the circuit performs interpolation on the current block, padding the first block of predicted samples according to the second motion vector to generate a second block of predicted samples; 124. An image decoding device as claimed in claim 123, wherein interpolation is performed using at least the second block of prediction samples.
[0579] 125. The image decoding device of claim 124, wherein when padding the first block of prediction samples to generate the second block of prediction samples, the circuit pads one or more sides of the first block of prediction samples according to the second motion vector.
[0580] 126. The image decoding device of claim 123, wherein the circuit performs padding on the first block of prediction samples according to the first motion vector when deriving the second motion vector using DMVR processing.
[0581] 127. An image coding method enabling an image coding device to carry out the steps according to any one of claims 1 to 48.
[0582] 128. An image decoding method enabling an image decoding device to execute the steps according to any one of claims 75 to 100. [Industrial Applicability]
[0583] The present disclosure is applicable to, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, and the like. [Explanation of symbols]
[0584] 100 Encoding device 102 Division 104 Subtraction section 106 Conversion unit 108 Quantization section 110 Entropy coding unit 112, 204 Inverse quantization section 114, 206 Inverse conversion unit 116, 208 Addition section 118, 210 block memory 120, 212 Loop filter section 122, 214 frame memory 124, 216 Intra prediction section 126, 218 Inter prediction section 128, 220 Predictive control section 200 Decryption Device 202 Entropy Decoding Unit 1000 Current Picture 1001 Current Block 1100 First Reference Picture 1110 First motion vector 1120 First prediction block 1121, 1221 prediction subblocks 1122, 1222 top left pixel 1130, 1130A First interpolation reference range 1131, 1131A, 1132, 1132A, 1231, 1231A, 1232, 1232A Reference Range 1135, 1135A First gradient reference range 1140 First predicted image 1150 First gradient image 1200 Second Reference Picture 1210 Second motion vector 1220 Second prediction block 1230, 1230A Second interpolation reference range 1235, 1235A Secondary slope reference range 1240 Second predicted image 1250 Second Gradient Image 1300 local motion detection value 1400 First predicted image
Claims
1. Memory and a circuit connected to the memory; The circuit, in operation, generating a prediction block by performing an interpolation process using values of samples included in a reference picture; using the predicted block to generate a gradient block of the same size as the predicted block; generating a predicted image using said gradient block; encoding the current block based on the predicted image; The generation of the gradient block comprises: a process of calculating a gradient value indicating a difference between a value of a right sample adjacent to the right of a target sample included in the prediction block and a value of a left sample adjacent to the target sample to the left, a first gradient value at the left end of the gradient block is calculated using, as the left sample, the value of a first sample at the same position as the first gradient value in the prediction block; a second gradient value adjacent to the right of the location of the first gradient value is calculated using the value of the first sample as the left sample; Encoding device.
2. Memory and a circuit connected to the memory; The circuit, in operation, generating a prediction block by performing an interpolation process using values of samples included in a reference picture; using the predicted block to generate a gradient block of the same size as the predicted block; generating a predicted image using said gradient block; Decoding a current block based on the predicted image; The generation of the gradient block comprises: a process of calculating a gradient value indicating a difference between a value of a right sample adjacent to the right of a target sample included in the prediction block and a value of a left sample adjacent to the target sample to the left, a first gradient value at the left end of the gradient block is calculated using, as the left sample, the value of a first sample at the same position as the first gradient value in the prediction block; a second gradient value adjacent to the right of the location of the first gradient value is calculated using the value of the first sample as the left sample; Decryption device.
Citation Information
Patent Citations
Video decoding device, video encoding device, prediction image generation device and motion vector derivation device
WO2018230493A1