Encoding device, decoding device, encoding method and decoding method
Patent Information
- Application Number
- KR1020247034651
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-05-19
- Filing Date
- 2018-05-14
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2038-05-14
Smart Images

Figure 112024112879772-PAT00033_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to an encoding device, etc., that encodes a moving image by performing picture-to-picture prediction. Background Technology
[0002] Conventionally, H.265 exists as a standard for encoding video images. H.265 is also known as HEVC (High Efficiency Video Coding). Prior art literature
[0003] H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) The problem to be solved
[0004] However, in the encoding and decoding of video images, if predictive processing is not performed efficiently, the throughput increases, and there is a possibility that processing delays may occur.
[0005] Therefore, the present disclosure provides an encoding device, etc., capable of efficiently performing prediction processing. means of solving the problem
[0006] An encoding device according to one aspect of the present disclosure is an encoding device that encodes a moving image using picture-to-picture prediction, comprising a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) process is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, depending on whether a bi-directional optical flow (BIO) process is applied to the generation of a predicted image of a processing target block included in the processing target picture in the moving image, and if the BIO process is applied to the generation of a predicted image of the processing target block, determines that the OBMC process is not applicable to the generation of a predicted image of the processing target block, and applies the BIO process without applying the OBMC process to the generation of a predicted image of the processing target block, wherein the BIO process is a process of generating a predicted image of the processing target block by referring to a spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and the OBMC The processing is a process that corrects the predicted image of the target block by using an image obtained by performing motion compensation of the target block using the motion vectors of surrounding blocks of the target block.
[0007] In addition, comprehensive or specific embodiments thereof may be realized as a system, device, method, integrated circuit, computer program, or a non-transient recording medium such as a computer-readable CD-ROM, or as any combination of a system, device, method, integrated circuit, computer program, and recording medium. Effects of the invention
[0008] An encoding device, etc. according to one embodiment of the present disclosure, can efficiently perform prediction processing. Brief explanation of the drawing
[0009] FIG. 1 is a block diagram showing the functional configuration of an encoding device according to embodiment 1. FIG. 2 is a figure showing an example of block division in embodiment 1. Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. Figure 4a is a figure showing an example of the shape of a filter used in ALF. Figure 4b is a figure showing another example of the shape of a filter used in ALF. FIG. 4c is a figure showing another example of the shape of a filter used in ALF. Figure 5a is a figure showing 67 intra prediction modes in intra prediction. Figure 5b is a flowchart illustrating an overview of the prediction image correction process by OBMC processing. Figure 5c is a conceptual diagram illustrating an overview of the prediction image correction processing by OBMC processing. Figure 5d is a figure showing an example of FRUC. Figure 6 is a figure illustrating pattern matching (bilateral matching) between two blocks following a movement trajectory. Figure 7 is a figure for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 8 is a figure for explaining a model assuming uniform linear motion. FIG. 9a is a diagram illustrating the derivation of motion vectors at the sub-block level based on motion vectors of multiple adjacent blocks. Figure 9b is a figure illustrating an overview of the motion vector derivation process by merge mode. Figure 9c is a conceptual diagram illustrating an overview of DMVR processing. FIG. 9d is a figure illustrating an overview of a method for generating a predicted image using luminance correction processing by LIC processing. FIG. 10 is a block diagram showing the functional configuration of a decoding device according to embodiment 1. FIG. 11 is a block diagram for explaining processing related to inter-frame prediction performed by an encoding device according to embodiment 1. FIG. 12 is a block diagram for explaining processing related to inter-frame prediction performed by a decoding device according to embodiment 1. FIG. 13 is a flowchart showing a first specific example of inter-screen prediction according to embodiment 1. FIG. 14 is a flowchart showing a variation of the first embodiment of inter-screen prediction according to embodiment 1. FIG. 15 is a flowchart showing a second specific example of inter-screen prediction according to embodiment 1. FIG. 16 is a flowchart showing a variation of a second embodiment of inter-screen prediction according to embodiment 1. FIG. 17 is a conceptual diagram showing OBMC processing according to embodiment 1. FIG. 18 is a flowchart showing OBMC processing according to embodiment 1. FIG. 19 is a conceptual diagram showing BIO processing according to embodiment 1. FIG. 20 is a flowchart showing BIO processing according to embodiment 1. FIG. 21 is a block diagram showing an example of implementation of an encoding device according to embodiment 1. FIG. 22 is a flowchart showing a first operation example of an encoding device according to embodiment 1. FIG. 23 is a flowchart showing a variation of the first operation example of an encoding device according to embodiment 1. FIG. 24 is a flowchart showing additional operations for a variation of the first operation example of an encoding device according to embodiment 1. FIG. 25 is a flowchart showing a second operation example of an encoding device according to embodiment 1. FIG. 26 is a flowchart showing a modified example of the second operation of the encoding device according to embodiment 1. FIG. 27 is a flowchart showing additional operations for a modified example of the second operation example of an encoding device according to embodiment 1. FIG. 28 is a block diagram showing an example of implementation of a decoding device according to embodiment 1. FIG. 29 is a flowchart showing a first operation example of a decoding device according to embodiment 1. FIG. 30 is a flowchart showing a variation of the first operation example of a decoding device according to embodiment 1. FIG. 31 is a flowchart showing additional operations for a variation of the first operation example of a decoding device according to embodiment 1. FIG. 32 is a flowchart showing a second operation example of a decoding device according to embodiment 1. FIG. 33 is a flowchart showing a modified example of the second operation of the decoding device according to embodiment 1. FIG. 34 is a flowchart showing additional operations for a modified example of the second operation example of a decoding device according to embodiment 1. FIG. 35 is an overall configuration diagram of a content supply system that realizes a content transmission service. Figure 36 is a figure showing an example of an encoding structure during scalable encoding. Figure 37 is a figure showing an example of an encoding structure during scalable encoding. Figure 38 is a figure showing an example of a display screen of a web page. Figure 39 is a figure showing an example of a display screen of a web page. Fig. 40 is a figure showing an example of a smartphone. FIG. 41 is a block diagram showing an example of a smartphone configuration. Specific details for implementing the invention
[0010] (Information forming the basis of the present disclosure)
[0011] In the encoding and decoding of moving images, picture-to-picture prediction may be performed. Picture-to-picture prediction is also referred to as inter-frame prediction, inter-frame prediction, or inter-prediction. In picture-to-picture prediction, motion vectors are derived for a block to be processed, and a prediction image corresponding to the block to be processed is generated using the derived motion vectors. Furthermore, when generating a prediction image using motion vectors, a prediction image generation method for generating the prediction image may be selected from a plurality of prediction image generation methods.
[0012] Specifically, as a method for generating a predicted image, there are methods such as generating a predicted image by motion compensation based on motion vectors from a processed picture, and a method called BIO (bi-directional optical flow). For example, in BIO processing, a predicted image is generated using a provisional predicted image generated by motion compensation based on motion vectors from each of two processed pictures, and a gradient image representing the spatial gradient of luminance in the provisional predicted image.
[0013] In addition, in picture-to-picture prediction, the predicted image may be corrected by a method called OBMC (overlapped block motion compensation). In OBMC processing, the motion vectors of blocks surrounding the block to be processed are used. Specifically, in OBMC processing, the predicted image is corrected by using an image generated by motion compensation based on the motion vectors of surrounding blocks.
[0014] Whether to apply OBMC processing may be determined based on the size of the block to be processed. When OBMC processing is applied, the error between the input image corresponding to the block to be processed and the predicted image after correction may be reduced. Alternatively, when OBMC processing is applied, the amount of data after transformation and quantization processing may be reduced. In other words, the effect of improving encoding efficiency and reducing the amount of code is expected.
[0015] However, by applying OBMC processing, the throughput of OBMC processing increases in picture-to-picture prediction.
[0016] For example, the throughput in inter-picture prediction varies depending on which prediction image generation method is used and whether OBMC processing is applied. In other words, the throughput in inter-picture prediction varies depending on the pass in the inter-picture prediction processing flow. Furthermore, the time allocated for inter-picture prediction is determined by the throughput of the pass with the maximum throughput in the inter-picture prediction processing flow. Therefore, it is preferable that the throughput of the pass with the maximum throughput in the inter-picture prediction processing flow be small.
[0017] Thus, for example, an encoding device according to one embodiment of the present disclosure is an encoding device that encodes a video image using picture-to-picture prediction, comprising a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) process is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the video image, depending on whether a bi-directional optical flow (BIO) process is applied to the generation of a predicted image of a processing target block included in the processing target picture in the video image, and if the BIO process is applied to the generation of a predicted image of the processing target block, it determines that the OBMC process is not applicable to the generation of a predicted image of the processing target block, and thus applies the BIO process without applying the OBMC process to the generation of a predicted image of the processing target block, wherein the BIO process is a process of generating a predicted image of the processing target block by referring to a spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and OBMC processing is a process that corrects the predicted image of the target block by using an image obtained by performing motion compensation of the target block using motion vectors of blocks surrounding the target block.
[0018] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0019] In addition, when BIO processing is applied, it is assumed that the degree of reduction in code volume due to the application of OBMC processing is lower compared to the case where BIO processing is not applied. That is, when it is assumed that the degree of reduction in code volume due to the application of OBMC processing is low, the encoding device can efficiently generate a predicted image without applying OBMC processing.
[0020] Therefore, the encoding device can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0021] Also, for example, the above circuit also selects one mode among (i) a first mode in which the BIO processing is not applied and the OBMC processing is not applied to the processing of the predicted image of the processing target block, (ii) a second mode in which the BIO processing is not applied and the OBMC processing is applied to the processing of the processing target block's predicted image, and (iii) a third mode in which the BIO processing is not applied and the OBMC processing is applied to the processing of the processing target block's predicted image, and performs the processing of the processing target block's predicted image in the selected mode. When the third mode is selected as the selected mode, it is determined that the OBMC processing is not applicable to the processing target block's predicted image because the BIO processing is applied to the processing target block's predicted image, and the processing of the processing target block's predicted image is performed in the third mode in which the BIO processing is applied without the OBMC processing.
[0022] By this, the encoding device can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow of picture-to-picture prediction, it is possible to eliminate the pass of applying BIO processing and also applying OBMC processing.
[0023] Also, for example, the circuit encodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0024] By this, the encoding device can encode one mode that is adaptively selected from three modes. Accordingly, the encoding device and the decoder can adaptively select the same mode from three modes.
[0025] Also, for example, a decoding device according to one embodiment of the present disclosure is a decoding device that decodes a moving image using picture-to-picture prediction, and comprises a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) process can be applied to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, depending on whether a bi-directional optical flow (BIO) process is applied to the generation of a predicted image of a processing target block included in the processing target picture in the moving image, and if the BIO process is applied to the generation of a predicted image of the processing target block, it determines that the OBMC process is not applicable to the generation of a predicted image of the processing target block, and applies the BIO process without applying the OBMC process to the generation of a predicted image of the processing target block, and the BIO process is a process of generating a predicted image of the processing target block by referring to a spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block. The above OBMC processing is a process that corrects the predicted image of the target block by using an image obtained by performing motion compensation of the target block using the motion vector of the surrounding blocks of the target block.
[0026] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0027] In addition, when BIO processing is applied, it is assumed that the degree of reduction in code volume due to the application of OBMC processing is lower compared to the case where BIO processing is not applied. That is, when it is assumed that the degree of reduction in code volume due to the application of OBMC processing is low, the decoder can efficiently generate a predicted image without applying OBMC processing.
[0028] Therefore, the decoder can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0029] Also, for example, the above circuit also selects one mode among (i) a first mode in which the BIO processing is not applied and the OBMC processing is not applied to the processing of the predicted image of the processing target block, (ii) a second mode in which the BIO processing is not applied and the OBMC processing is applied to the processing of the processing target block's predicted image, and (iii) a third mode in which the BIO processing is not applied and the OBMC processing is applied to the processing of the processing target block's predicted image, and performs the processing of the processing target block's predicted image in the selected mode. When the third mode is selected as the selected mode, it is determined that the OBMC processing is not applicable to the processing target block's predicted image because the BIO processing is applied to the processing target block's predicted image, and the processing of the processing target block's predicted image is performed in the third mode in which the BIO processing is applied without the OBMC processing.
[0030] By this, the decoder can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow of picture-to-picture prediction, it is possible to eliminate the pass of applying BIO processing and also applying OBMC processing.
[0031] Also, for example, the circuit decodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0032] By this, the decoder can decode one mode that is adaptively selected from three modes. Thus, the encoding device and the decoder can adaptively select the same mode from three modes.
[0033] Also, for example, an encoding method according to one aspect of the present disclosure is an encoding method for encoding a moving image using picture-to-picture prediction, wherein, depending on whether a bi-directional optical flow (BIO) processing is applied to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, it is determined whether an overlapped block motion compensation (OBMC) processing is applicable to the generation of a predicted image of the processing target block, and if the BIO processing is applied to the generation of a predicted image of the processing target block, it is determined that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and thus the BIO processing is applied without applying the OBMC processing to the generation of a predicted image of the processing target block, wherein the BIO processing is a process for generating a predicted image of the processing target block by referring to a spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and the OBMC processing is a process for generating the predicted image of the processing target using the motion vector of a surrounding block of the processing target This is a process that corrects the predicted image of the target block using an image obtained by performing movement compensation of the block.
[0034] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0035] In addition, when BIO processing is applied, it is assumed that the degree of reduction in code volume due to the application of OBMC processing is lower compared to the case where BIO processing is not applied. That is, when it is assumed that the degree of reduction in code volume due to the application of OBMC processing is low, devices using this encoding method can efficiently generate predicted images without applying OBMC processing.
[0036] Therefore, devices utilizing this encoding method can perform predictive processing efficiently. Additionally, the circuit size can be reduced.
[0037] Also, for example, a decoding method according to one aspect of the present disclosure is a decoding method for decoding a moving image using picture-to-picture prediction, wherein, depending on whether a bi-directional optical flow (BIO) processing is applied to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, it is determined whether an overlapped block motion compensation (OBMC) processing is applicable to the generation of a predicted image of the processing target block, and if the BIO processing is applied to the generation of a predicted image of the processing target block, it is determined that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and thus the BIO processing is applied without applying the OBMC processing to the generation of a predicted image of the processing target block, wherein the BIO processing is a process for generating a predicted image of the processing target block by referring to a spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and the OBMC processing is a process for generating the predicted image of the processing target using the motion vector of a block surrounding the processing target This is a process that corrects the predicted image of the target block using an image obtained by performing movement compensation of the block.
[0038] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0039] In addition, when BIO processing is applied, it is assumed that the degree of reduction in code volume due to the application of OBMC processing is lower compared to the case where BIO processing is not applied. That is, when it is assumed that the degree of reduction in code volume due to the application of OBMC processing is low, devices using this decoding method can efficiently generate a predicted image without applying OBMC processing.
[0040] Therefore, devices utilizing this decoding method can efficiently perform predictive processing. Additionally, a reduction in circuit size is possible.
[0041] Also, for example, an encoding device according to one embodiment of the present disclosure is an encoding device that encodes a video image using picture-to-picture prediction, and comprises a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) process can be applied to the process of generating a predicted image of a processing target block included in a processing target picture in the video image, depending on whether a pair prediction process referencing two completed processing pictures is applied to the process of generating a predicted image of a processing target block, and if the pair prediction process is applied to the process of generating a predicted image of the processing target block, it determines that the OBMC process cannot be applied to the process of generating a predicted image of the processing target block, and applies the pair prediction process without applying the OBMC process to the process of generating a predicted image of the processing target block, and the OBMC process is a process of correcting the predicted image of the processing target block using an image obtained by performing motion compensation of the processing target block using the motion vector of a block surrounding the processing target block.
[0042] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0043] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. The encoding device can suppress such an increase in throughput.
[0044] Therefore, the encoding device can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0045] Also, for example, the above circuit also selects one mode among (i) a first mode in which the pair prediction processing is not applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed, (ii) a second mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and (iii) a third mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and performs the generation processing of the predicted image of the block to be processed in the selected one mode, and when the third mode is selected as the one mode, it is determined that the OBMC processing is not applicable to the generation processing of the predicted image of the block to be processed as the pair prediction processing is applied to the generation processing of the predicted image of the block to be processed, and the generation processing of the predicted image of the block to be processed is performed in the third mode in which the pair prediction processing is applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed.
[0046] By this, the encoding device can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow for picture-to-picture prediction, it is possible to eliminate the pass of applying pair prediction processing that references two completed pictures and also applying OBMC processing.
[0047] Also, for example, the circuit encodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0048] By this, the encoding device can encode one mode that is adaptively selected from three modes. Accordingly, the encoding device and the decoder can adaptively select the same mode from three modes.
[0049] Also, for example, a decoding device according to one embodiment of the present disclosure is a decoding device that decodes a video image using picture-to-picture prediction, and comprises a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) process can be applied to the process of generating a predicted image of a processing target block included in a processing target picture in the video image, depending on whether a pair prediction process referencing two completed processing pictures is applied to the process of generating a predicted image of a processing target block, and if the pair prediction process is applied to the process of generating a predicted image of the processing target block, it determines that the OBMC process cannot be applied to the process of generating a predicted image of the processing target block, and applies the pair prediction process without applying the OBMC process to the process of generating a predicted image of the processing target block, and the OBMC process is a process of correcting the predicted image of the processing target block using an image obtained by performing motion compensation of the processing target block using the motion vector of a block surrounding the processing target block.
[0050] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0051] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. The decoder can suppress such an increase in throughput.
[0052] Therefore, the decoder can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0053] Also, for example, the above circuit also selects one mode among (i) a first mode in which the pair prediction processing is not applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed, (ii) a second mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and (iii) a third mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and performs the generation processing of the predicted image of the block to be processed in the selected one mode, and when the third mode is selected as the one mode, it is determined that the OBMC processing is not applicable to the generation processing of the predicted image of the block to be processed as the pair prediction processing is applied to the generation processing of the predicted image of the block to be processed, and the generation processing of the predicted image of the block to be processed is performed in the third mode in which the pair prediction processing is applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed.
[0054] By this, the decoder can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow for picture-to-picture prediction, it is possible to eliminate the pass of applying pair prediction processing that references two completed pictures and also applying OBMC processing.
[0055] Also, for example, the circuit decodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0056] By this, the decoder can decode one mode that is adaptively selected from three modes. Thus, the encoding device and the decoder can adaptively select the same mode from three modes.
[0057] Also, for example, an encoding method according to one aspect of the present disclosure is an encoding method for encoding a video image using picture-to-picture prediction, wherein, depending on whether a pair prediction process referencing two completed pictures is applied to the generation process of a predicted image of a processing target block included in a processing target picture in the video image, it is determined whether an OBMC (overlapped block motion compensation) process can be applied to the generation process of a predicted image of a processing target block, and if the pair prediction process is applied to the generation process of a predicted image of a processing target block, it is determined that the OBMC process cannot be applied to the generation process of a predicted image of a processing target block, and the pair prediction process is applied without applying the OBMC process to the generation process of a predicted image of a processing target block, and the OBMC process is a process of correcting the predicted image of a processing target block using an image obtained by performing motion compensation of the processing target block using motion vectors of surrounding blocks of the processing target block.
[0058] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0059] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. Devices using this encoding method can suppress such an increase in throughput.
[0060] Therefore, devices utilizing this encoding method can perform predictive processing efficiently. Additionally, the circuit size can be reduced.
[0061] Also, for example, a decoding method according to one aspect of the present disclosure is a decoding method for decoding a video image using picture-to-picture prediction, wherein, depending on whether a pair prediction process referencing two completed pictures is applied to the generation process of a predicted image of a processing target block included in a processing target picture in the video image, it is determined whether an OBMC (overlapped block motion compensation) process can be applied to the generation process of a predicted image of a processing target block, and if the pair prediction process is applied to the generation process of a predicted image of a processing target block, it is determined that the OBMC process cannot be applied to the generation process of a predicted image of a processing target block, and the pair prediction process is applied without applying the OBMC process to the generation process of a predicted image of a processing target block, and the OBMC process is a process of correcting the predicted image of a processing target block using an image obtained by performing motion compensation of the processing target block using the motion vector of a block surrounding the processing target block.
[0062] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0063] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. Devices using this decoding method can suppress such an increase in throughput.
[0064] Therefore, devices utilizing this decoding method can efficiently perform predictive processing. Additionally, a reduction in circuit size is possible.
[0065] Also, for example, an encoding device according to one embodiment of the present disclosure is an encoding device that encodes a video image using picture-to-picture prediction, comprising a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) processing is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the video image, depending on whether a paired prediction processing that refers to a processing completed picture prior in display order to the processing target picture and a processing completed picture following in display order to the processing target picture is applied, and if the paired prediction processing is applied to the generation of a predicted image of the processing target block, it determines that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and applies the paired prediction processing without applying the OBMC processing to the generation of a predicted image of the processing target block, and the OBMC processing is applied using the motion vector of a block surrounding the processing target block to the processing target block This is a process that corrects the predicted image of the processing target block using an image obtained by performing motion compensation.
[0066] By this, in the inter-picture prediction processing flow, it is possible to eliminate the pass of applying paired prediction processing that references two completed pictures, one before and one after, in display order, and also applying OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0067] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. The encoding device can suppress such an increase in throughput.
[0068] Therefore, the encoding device can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0069] Also, for example, the above circuit also selects one mode among (i) a first mode in which the pair prediction processing is not applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed, (ii) a second mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and (iii) a third mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and performs the generation processing of the predicted image of the block to be processed in the selected one mode, and when the third mode is selected as the one mode, it is determined that the OBMC processing is not applicable to the generation processing of the predicted image of the block to be processed as the pair prediction processing is applied to the generation processing of the predicted image of the block to be processed, and the generation processing of the predicted image of the block to be processed is performed in the third mode in which the pair prediction processing is applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed.
[0070] By this, the encoding device can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow of inter-picture prediction, it is possible to eliminate the pass of applying paired prediction processing that refers to two completed pictures, one before and one after, in display order, and also applying OBMC processing.
[0071] Also, for example, the circuit encodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0072] By this, the encoding device can encode one mode that is adaptively selected from three modes. Accordingly, the encoding device and the decoder can adaptively select the same mode from three modes.
[0073] Also, for example, a decoding device according to one embodiment of the present disclosure is a decoding device that decodes a video image using picture-to-picture prediction, and comprises a memory and a circuit capable of accessing the memory, wherein the circuit capable of accessing the memory determines whether an overlapped block motion compensation (OBMC) processing is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the video image, depending on whether a paired prediction processing that refers to a processing completed picture prior in display order to the processing target picture and a processing completed picture following in display order to the processing target picture is applied, and if the paired prediction processing is applied to the generation of a predicted image of the processing target block, it determines that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and applies the paired prediction processing without applying the OBMC processing to the generation of a predicted image of the processing target block, and the OBMC processing utilizes the motion vector of a block surrounding the processing target to the processing target This is a process that corrects the predicted image of the target block using an image obtained by performing movement compensation of the block.
[0074] By this, in the inter-picture prediction processing flow, it is possible to eliminate the pass of applying paired prediction processing that references two completed pictures, one before and one after, in display order, and also applying OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0075] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. The decoder can suppress such an increase in throughput.
[0076] Therefore, the decoder can perform predictive processing efficiently. In addition, the circuit size can be reduced.
[0077] Also, for example, the above circuit also selects one mode among (i) a first mode in which the pair prediction processing is not applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed, (ii) a second mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and (iii) a third mode in which the pair prediction processing is not applied and the OBMC processing is applied to the generation processing of the predicted image of the block to be processed, and performs the generation processing of the predicted image of the block to be processed in the selected one mode, and when the third mode is selected as the one mode, it is determined that the OBMC processing is not applicable to the generation processing of the predicted image of the block to be processed as the pair prediction processing is applied to the generation processing of the predicted image of the block to be processed, and the generation processing of the predicted image of the block to be processed is performed in the third mode in which the pair prediction processing is applied and the OBMC processing is not applied to the generation processing of the predicted image of the block to be processed.
[0078] By this, the decoder can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes. In addition, in the processing flow of inter-picture prediction, it is possible to eliminate the pass of applying pair prediction processing that refers to two completed pictures, one before and one after, in display order, and also applying OBMC processing.
[0079] Also, for example, the circuit decodes a signal representing one value corresponding to one mode selected among the first mode, the second mode, and the third mode, among three values of a first value, a second value, and a third value corresponding to the first mode, the second mode, and the third mode, respectively.
[0080] By this, the decoder can decode one mode that is adaptively selected from three modes. Thus, the encoding device and the decoder can adaptively select the same mode from three modes.
[0081] Also, for example, an encoding method according to one aspect of the present disclosure is an encoding method for encoding a video image using inter-picture prediction, wherein, depending on whether a pair prediction process referencing a completed picture prior to the picture to be processed and a completed picture following the picture to be processed in display order is applied to the generation process of a predicted image of a block to be processed included in the picture to be processed in the video image, an overlapped block motion compensation (OBMC) process is determined to be applicable to the generation process of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation process of a predicted image of a block to be processed, it is determined that the OBMC process is not applicable to the generation process of a predicted image of a block to be processed, and the pair prediction process is applied without applying the OBMC process to the generation process of a predicted image of a block to be processed, and the OBMC process is applied, wherein the OBMC process corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using motion vectors of blocks surrounding the block to be processed. It is processing.
[0082] By this, in the inter-picture prediction processing flow, it is possible to eliminate the pass of applying paired prediction processing that references two completed pictures, one before and one after, in display order, and also applying OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0083] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. Devices using this encoding method can suppress such an increase in throughput.
[0084] Therefore, devices utilizing this encoding method can perform predictive processing efficiently. Additionally, the circuit size can be reduced.
[0085] Also, for example, a decoding method according to one aspect of the present disclosure is a decoding method for decoding a video image using inter-picture prediction, wherein, depending on whether a pair prediction process referencing a completed picture prior to the picture to be processed and a completed picture following the picture to be processed in display order is applied to the generation process of a predicted image of a block to be processed included in the picture to be processed in the video image, it is determined whether an overlapped block motion compensation (OBMC) process can be applied to the generation process of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation process of a predicted image of a block to be processed, it is determined that the OBMC process cannot be applied to the generation process of a predicted image of a block to be processed, and the pair prediction process is applied without applying the OBMC process to the generation process of a predicted image of a block to be processed, and the OBMC process is applied, wherein the OBMC process corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using motion vectors of blocks surrounding the block to be processed. It is processing.
[0086] By this, in the inter-picture prediction processing flow, it is possible to eliminate the pass of applying paired prediction processing that references two completed pictures, one before and one after, in display order, and also applying OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0087] In addition, when paired prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction images obtained from each of the two completed pictures, thereby significantly increasing the throughput. Devices using this decoding method can suppress such an increase in throughput.
[0088] Therefore, devices utilizing this decoding method can efficiently perform predictive processing. Additionally, a reduction in circuit size is possible.
[0089] In addition, comprehensive or specific embodiments thereof may be realized as a system, device, method, integrated circuit, computer program, or a non-transient recording medium such as a computer-readable CD-ROM, or as any combination of a system, device, method, integrated circuit, computer program, and recording medium.
[0090] The following describes embodiments in detail with reference to the drawings.
[0091] Furthermore, the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection types of components, steps, and sequence of steps described in the embodiments below are examples and are not well known facts that limit the scope of the claims. Additionally, regarding the components in the embodiments below, any component not described in the independent claim representing the highest concept is described as any component.
[0092] (Form of Embodiment 1)
[0093] First, an overview of Embodiment 1 is described as an example of an encoding device and a decoding device to which the processing and / or configuration described in each embodiment of the present disclosure below can be applied. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processing and / or configuration described in each embodiment of the present disclosure can be applied, and the processing and / or configuration described in each embodiment of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.
[0094] When applying the processing and / or configuration described in each aspect of the present disclosure to Embodiment 1, any of the following may be performed, for example.
[0095] (1) With respect to the encoding device or decoding device of Embodiment 1, among the plurality of components constituting the encoding device or decoding device, the component corresponding to the component described in each embodiment of the present disclosure is replaced with the component described in each embodiment of the present disclosure.
[0096] (2) With respect to the encoding device or decoding device of Embodiment 1, any modification such as the addition, substitution, or deletion of a function or processing performed on some of the components constituting the encoding device or decoding device is performed, and then the component corresponding to the component described in each embodiment of the present disclosure is replaced with the component described in each embodiment of the present disclosure.
[0097] (3) With respect to the method performed by the encoding device or decoding device of Embodiment 1, any modification such as substitution or deletion is performed on some of the processing included in the method, and then the processing corresponding to the processing described in each embodiment of the present disclosure is replaced with the processing described in each embodiment of the present disclosure.
[0098] (4) Some of the components constituting the encoding device or decoding device of Embodiment 1 are combined with the components described in each embodiment of the present disclosure, the components having a part of the function of the components described in each embodiment of the present disclosure, or the components having a part of the processing performed by the components described in each embodiment of the present disclosure.
[0099] (5) A component having a part of the function of some of the components constituting the encoding device or decoding device of Embodiment 1, or a component performing a part of the processing performed by some of the components constituting the encoding device or decoding device of Embodiment 1, is combined with a component described in each embodiment of the present disclosure, a component having a part of the function of the component described in each embodiment of the present disclosure, or a component performing a part of the processing performed by the component described in each embodiment of the present disclosure.
[0100] (6) With respect to the method performed by the encoding device or decoding device of Embodiment 1, among the plurality of processes included in the method, the process corresponding to the process described in each embodiment of the present disclosure is replaced with the process described in each embodiment of the present disclosure.
[0101] (7) Performing some of the processing included in the method of the encoding device or decoding device of Embodiment 1 in combination with the processing described in each aspect of the present disclosure
[0102] Furthermore, the method of implementing the processing and / or configuration described in each embodiment of the present disclosure is not limited to the examples above. For example, it may be implemented in a device used for a purpose different from the video / image encoding device or video / image decoding device disclosed in Embodiment 1, and the processing and / or configuration described in each embodiment may be implemented alone. In addition, the processing and / or configuration described in different embodiments may be implemented in combination.
[0103] [Overview of Encoding Device]
[0104] First, an overview of the encoding device according to embodiment 1 will be explained. FIG. 1 is a block diagram showing the functional configuration of the encoding device (100) according to embodiment 1. The encoding device (100) is a video / image encoding device that encodes video / images in block units.
[0105] As shown in FIG. 1, the encoding device (100) is a device that encodes an image in block units and comprises a splitting unit (102), a subtraction unit (104), a conversion unit (106), a quantization unit (108), an entropy encoding unit (110), an inverse quantization unit (112), an inverse conversion unit (114), an addition unit (116), a block memory (118), a loop filter unit (120), a frame memory (122), an intra prediction unit (124), an inter prediction unit (126), and a prediction control unit (128).
[0106] The encoding device (100) is realized, for example, by a general-purpose processor and memory. In this case, when a software program stored in memory is executed by the processor, the processor functions as a splitting unit (102), a subtraction unit (104), a conversion unit (106), a quantization unit (108), an entropy encoding unit (110), an inverse quantization unit (112), an inverse conversion unit (114), an addition unit (116), a loop filter unit (120), an intra prediction unit (124), an inter prediction unit (126), and a prediction control unit (128). Additionally, the encoding device (100) may be realized as one or more dedicated electronic circuits corresponding to a dividing unit (102), a subtraction unit (104), a conversion unit (106), a quantization unit (108), an entropy encoding unit (110), an inverse quantization unit (112), an inverse conversion unit (114), an addition unit (116), a loop filter unit (120), an intra prediction unit (124), an inter prediction unit (126), and a prediction control unit (128).
[0107] Below, each component included in the encoding device (100) is described.
[0108] [Divided Part]
[0109] The splitting unit (102) divides each picture included in the input video image into multiple blocks and outputs each block to the subtraction unit (104). For example, the splitting unit (102) first divides the picture into blocks of a fixed size (e.g., 128×128). These fixed-size blocks may be called encoding tree units (CTU). Then, the splitting unit (102) divides each of the fixed-size blocks into blocks of a variable size (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks may be called encoding units (CU), prediction units (PU), or conversion units (TU). Furthermore, in the present embodiment, the CU, PU, and TU do not need to be distinguished, and some or all blocks within the picture may be processing units of the CU, PU, and TU.
[0110] FIG. 2 is a figure showing an example of block division in embodiment 1. In FIG. 2, solid lines represent block boundaries by quaternary tree block division, and dashed lines represent block boundaries by binary tree block division.
[0111] Here, the block (10) is a square block of 128×128 pixels (128×128 block). This 128×128 block (10) is first divided into 64×64 square blocks of 4 squares (quadratic tree block division).
[0112] The 64×64 block at the top left is also vertically divided into two rectangular 32×64 blocks, and the 32×64 block at the left is also vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the 64×64 block at the top left is divided into two 16×64 blocks (11, 12) and a 32×64 block (13).
[0113] The 64×64 block at the top right is horizontally divided into two rectangular 64×32 blocks (14, 15) (binary tree block division).
[0114] The 64×64 block in the bottom left is divided into four square 32×32 blocks (quaternary tree block division). Of the four 32×32 blocks, the block in the top left and the block in the bottom right are further divided. The 32×32 block in the top left is vertically divided into two rectangular 16×32 blocks, and the 16×32 block on the right is also horizontally divided into two 16×16 blocks (binary tree block division). The 32×32 block in the bottom right is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the 64×64 block at the bottom left is divided into a 16×32 block (16), two 16×16 blocks (17, 18), two 32×32 blocks (19, 20), and two 32×16 blocks (21, 22).
[0115] The 64×64 block (23) on the lower right is not divided.
[0116] As described above, in FIG. 2, the block (10) is divided into 13 variable-size blocks (11 to 23) based on recursive quad-tree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quad-tree plus binary tree) partitioning.
[0117] In addition, in FIG. 2, one block was divided into four or two blocks (quaternary tree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). A division including such ternary tree block division is sometimes called MBT (multi-type tree) division.
[0118] [Subtraction]
[0119] The subtraction unit (104) subtracts the predicted signal (predicted sample) from the original signal (original sample) in block units divided by the division unit (102). That is, the subtraction unit (104) calculates the prediction error (also called residual) of the encoding target block (hereinafter referred to as the current block). Then, the subtraction unit (104) outputs the calculated prediction error to the conversion unit (106).
[0120] The original signal is an input signal of the encoding device (100) and is a signal representing an image of each picture constituting the moving image (e.g., a luminance (luma) signal and two chroma signals). In the following, the signal representing the image is also referred to as a sample.
[0121] [Conversion Section]
[0122] The conversion unit (106) converts the prediction error in the spatial domain into a conversion coefficient in the frequency domain and outputs the conversion coefficient to the quantization unit (108). Specifically, the conversion unit (106) performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0123] Additionally, the transformation unit (106) may adaptively select a transformation type from among a plurality of transformation types and convert the prediction error into a transformation coefficient using a transformation basis function corresponding to the selected transformation type. Such a transformation may be referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0124] Multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing transformation basis functions corresponding to each transformation type. In FIG. 3, N represents the number of input pixels. The selection of a transformation type among these multiple transformation types may depend, for example, on the type of prediction (intra prediction and inter prediction) or on the intra prediction mode.
[0125] Information indicating whether such EMT or AMT is applied (e.g., called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0126] Additionally, the transformation unit (106) may re-transform the transformation coefficients (transformation result). Such re-transformation may be referred to as an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transformation unit (106) performs re-transformation for each sub-block (e.g., a 4×4 sub-block) included in the block of transformation coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information regarding the transformation matrix used for NSST are signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0127] Here, a Separable transformation is a method of performing multiple transformations by separating each direction according to the number of dimensions of the input, and a Non-Separable transformation is a method of performing a transformation by integrating two or more dimensions into one dimension when the input is multi-dimensional.
[0128] For example, as an example of a non-separable transformation, if the input is a 4×4 block, it can be considered as a single array with 16 elements, and the transformation process is performed on that array using a 16×16 transformation matrix.
[0129] Likewise, considering a 4×4 input block as a single array with 16 elements and then performing multiple Givens rotations on that array (Hypercube Givens Transform) is also an example of a non-separable transformation.
[0130] [Quantization Unit]
[0131] The quantization unit (108) quantizes the conversion coefficients output from the conversion unit (106). Specifically, the quantization unit (108) scans the conversion coefficients of the current block in a predetermined scanning order and quantizes the conversion coefficients based on a quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit (108) outputs the quantized conversion coefficients of the current block (hereinafter referred to as quantization coefficients) to the entropy encoding unit (110) and the inverse quantization unit (112).
[0132] A predetermined order is an order for the quantization / inverse quantization of the transformation coefficients. For example, a predetermined scanning order is defined as an ascending order of frequency (from low frequency to high frequency) or a descending order (from high frequency to low frequency).
[0133] A quantization parameter is a parameter that defines the quantization level (quantization width). For example, if the value of the quantization parameter increases, the quantization level also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.
[0134] [Entropy Encoding Unit]
[0135] The entropy encoding unit (110) generates an encoded signal (encoded bit stream) by encoding the quantization coefficients, which are inputs from the quantization unit (108), using variable length encoding. Specifically, the entropy encoding unit (110), for example, binaryizes the quantization coefficients and arithmetic-encodes the binary signal.
[0136] [Inverse Quantizer]
[0137] The inverse quantization unit (112) inversely quantizes the quantization coefficients that are inputs from the quantization unit (108). Specifically, the inverse quantization unit (112) inversely quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit (112) outputs the inversely quantized conversion coefficients of the current block to the inverse conversion unit (114).
[0138] [Inverse Transformation Section]
[0139] The inverse transform unit (114) restores the prediction error by inversely transforming the transformation coefficients that are inputs from the inverse quantization unit (112). Specifically, the inverse transform unit (114) restores the prediction error of the current block by performing an inverse transform corresponding to the transformation by the transformation unit (106) on the transformation coefficients. Then, the inverse transform unit (114) outputs the restored prediction error to the addition unit (116).
[0140] In addition, the restored prediction error does not match the prediction error calculated by the subtraction unit (104) because information is lost due to quantization. That is, the restored prediction error includes a quantization error.
[0141] [Additional part]
[0142] The adder (116) reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit (114), and the prediction sample, which is the input from the prediction control unit (128). Then, the adder (116) outputs the reconstructed block to the block memory (118) and the loop filter unit (120). The reconstructed block is also referred to as a local decoding block.
[0143] [Block Memory]
[0144] The block memory (118) is a block referenced in intra-prediction and is a memory unit for storing blocks within the encoding target picture (hereinafter referred to as the current picture). Specifically, the block memory (118) stores the reconstruction block output from the adder (116).
[0145] [Loop Filter Section]
[0146] The loop filter unit (120) performs a loop filter on the block reconstructed by the adder (116) and outputs the filtered reconstructed block to the frame memory (122). A loop filter is a filter used within an encoding loop (in-loop filter) and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0147] In ALF, a least squares error filter is applied to remove encoding distortion, and for example, for every 2×2 sub-block within the current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.
[0148] Specifically, first, a sub-block (e.g., a 2×2 sub-block) is classified into multiple classes (e.g., 15 or 25 classes). The classification of the sub-block is performed based on the gradient direction and activity level. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-block is classified into multiple classes (e.g., 15 or 25 classes).
[0149] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Also, the gradient activation value A is derived, for example, by adding gradients in multiple directions and quantizing the result of the addition.
[0150] Based on the results of such classification, a filter for a sub-block is determined among multiple filters.
[0151] For example, a circularly symmetric shape is used as the shape of the filter used in ALF. FIGS. 4a to 4c are drawings showing multiple examples of the shapes of filters used in ALF. FIG. 4a shows a 5×5 diamond-shaped filter, FIG. 4b shows a 7×7 diamond-shaped filter, and FIG. 4c shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. In addition, the signaling of information indicating the shape of the filter is not limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0152] The on / off status of ALF is determined, for example, by the picture level or the CU level. For example, whether to apply ALF to luminance is determined by the CU level, and whether to apply ALF to chrominance is determined by the picture level. Information indicating the on / off status of ALF is signaled at the picture level or the CU level. Additionally, the signaling of information indicating the on / off status of ALF is not limited to the picture level or the CU level, but may be at other levels (for example, the sequence level, the slice level, the tile level, or the CTU level).
[0153] A set of coefficients of multiple selectable filters (e.g., 15 or up to 25 filters) is signaled at the picture level. Additionally, the signaling of the set of coefficients is not limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0154] [Frame Memory]
[0155] The frame memory (122) is a memory unit for storing reference pictures used for inter prediction and is also called a frame buffer. Specifically, the frame memory (122) stores a reconstruction block filtered by the loop filter unit (120).
[0156] [Intra Prediction Section]
[0157] The intra prediction unit (124) generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-frame prediction) of the current block by referencing a block within the current picture stored in the block memory (118). Specifically, the intra prediction unit (124) generates an intra prediction signal by performing intra prediction by referencing a sample (e.g., luminance value, color difference value) of a block adjacent to the current block, and outputs the intra prediction signal to the prediction control unit (128).
[0158] For example, the intra prediction unit (124) performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0159] The above non-directional prediction modes include, for example, the Planar prediction mode and DC prediction mode defined by the H.265 / HEVC (High-Efficiency Video Coding) standard (non-patent literature 1).
[0160] Multiple directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. Additionally, multiple directional prediction modes may include 32 additional directional prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). FIG. 5a is a figure showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows indicate the 33 directions defined by the H.265 / HEVC standard, and dashed arrows indicate the 32 additional directions.
[0161] In addition, a luminance block may be referenced in the intra prediction of the chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. Such intra prediction is sometimes referred to as CCLM (cross-component linear model) prediction. Such an intra prediction mode of the chrominance block referencing the luminance block (for example, called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0162] The intra prediction unit (124) may correct the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction. Intra prediction involving such correction may be called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is signaled at the CU level, for example. Additionally, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0163] [Inter Prediction Department]
[0164] The inter prediction unit (126) generates a prediction signal (inter prediction signal) by performing inter prediction of the current block (also called inter-frame prediction) by referencing a reference picture stored in the frame memory (122) that is different from the current picture. Inter prediction is performed on the unit of the current block or a sub-block within the current block (e.g., a 4×4 block). For example, the inter prediction unit (126) performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit (126) generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., a motion vector) obtained by motion estimation. Then, the inter prediction unit (126) outputs the generated inter prediction signal to the prediction control unit (128).
[0165] Motion information used for motion compensation is signaled. A motion vector predictor may be used for the signaling of the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0166] In addition, an inter prediction signal may be generated by utilizing not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks. Specifically, an inter prediction signal may be generated at the sub-block level within the current block by weighted addition of a prediction signal based on motion information obtained by motion search and a prediction signal based on motion information of adjacent blocks. Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0167] In such an OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called the OBMC block size) is signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (e.g., called the OBMC flag) is signaled at the CU level. Furthermore, the signaling levels of these information are not limited to the sequence level and the CU level, and may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0168] The OBMC mode will be explained in more detail. Figures 5b and 5c are a flowchart and a conceptual diagram illustrating an overview of the prediction image correction processing by OBMC processing.
[0169] First, a predicted image (Pred) is obtained through normal motion compensation using the motion vector (MV) assigned to the block to be encoded.
[0170] Next, the motion vector (MV_L) of the left adjacent block that has been encoded is applied to the block to be encoded to obtain a predicted image (Pred_L), and the first correction of the predicted image is performed by superimposing the predicted image and Pred_L with weights.
[0171] Likewise, the motion vector (MV_U) of the upper adjacent block that has been encoded is applied to the block to be encoded to obtain a predicted image (Pred_U), and a second correction of the predicted image is performed by superimposing the predicted image that has undergone the first correction with Pred_U with weights, and the result is made into the final predicted image.
[0172] In addition, although a two-stage correction method using the left adjacent block and the upper adjacent block has been described here, it is also possible to configure the system to perform more than two stages of correction using the right adjacent block or the lower adjacent block.
[0173] In addition, the area for overlapping does not have to be the entire pixel area of the block, but may be only a part of the area near the block boundary.
[0174] In addition, although the process of correcting the predicted image from a single reference picture has been described here, the same applies when correcting the predicted image from multiple reference pictures. After obtaining the predicted image corrected from each reference picture, the obtained predicted image is superimposed to obtain the final predicted image.
[0175] In addition, the above-mentioned processing target block may be in the form of a prediction block unit, or in the form of a sub-block unit formed by further dividing the prediction block.
[0176] As a method for determining whether to apply OBMC processing, for example, there is a method using obmc_flag, a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a complex motion area; if it belongs to a complex motion area, the value of obmc_flag is set to 1 to apply OBMC processing and perform encoding, and if it does not belong to a complex motion area, the value of obmc_flag is set to 0 to perform encoding without applying OBMC processing. Meanwhile, in a decoder, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC processing according to the value.
[0177] Additionally, motion information may not be signaled but may be derived from the decoder side. For example, the merge mode defined by the H.265 / HEVC standard may be used. Alternatively, for example, motion information may be derived by performing motion search on the decoder side. In this case, motion search is performed without using the pixel values of the current block.
[0178] Here, a mode in which motion is searched on the decoder side is described. This mode in which motion is searched on the decoder side is sometimes called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0179] An example of FRUC processing is shown in FIG. 5d. First, by referring to the motion vectors of encoded complete blocks spatially or temporally adjacent to the current block, a list of multiple candidates (which may be common to the merge list) is generated, each having a predicted motion vector. Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value is calculated for each candidate included in the candidate list, and one candidate is selected based on the evaluation value.
[0180] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as is as the motion vector for the current block. Alternatively, for example, the motion vector for the current block may be derived by performing pattern matching on the surrounding area of the position within the reference picture corresponding to the motion vector of the selected candidate. That is, the search is performed on the surrounding area of the best candidate MV in the same way, and if there is an MV with a better evaluation value, the best candidate MV is updated to the said MV and used as the final MV for the current block. It is also possible to configure the system without performing this processing.
[0181] Even when processing is performed at the sub-block level, the exact same processing may be used.
[0182] Additionally, the evaluation value is calculated by determining the difference value of the reconstructed image through pattern matching between an area within the reference picture corresponding to the motion vector and a predetermined area. Furthermore, the evaluation value may be calculated by using other information in addition to the difference value.
[0183] As for pattern matching, a first pattern matching or a second pattern matching is used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.
[0184] In the first pattern matching, pattern matching is performed between two blocks within two different reference pictures that follow the motion trajectory of the current block. Accordingly, in the first pattern matching, an area within another reference picture that follows the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the aforementioned candidate.
[0185] FIG. 6 is a diagram illustrating an example of pattern matching (bilateral matching) between two blocks following a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among two block pairs within two different reference pictures (Ref0, Ref1) that follow the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a designated position within the first encoded reference picture (Ref0) specified in the candidate MV and the reconstructed image at a designated position within the second encoded reference picture (Ref1) specified in the symmetric MV scaled by the display time interval of the candidate MV is derived, and an evaluation value is calculated using the obtained difference value. Among a plurality of candidate MVs, the candidate MV that has the best evaluation value is selected as the final MV.
[0186] Under the assumption of a continuous motion trajectory, motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located temporally between the two reference pictures and the temporal distance from the current picture to the two reference pictures is the same, then in the first pattern matching, a bidirectional motion vector that is mirror-symmetric is derived.
[0187] In the second pattern matching, pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture (e.g., an upper and / or left adjacent block)) and a block within the reference picture. Accordingly, in the second pattern matching, a block adjacent to the current block within the current picture is used as a predetermined area for calculating the evaluation value of the aforementioned candidate.
[0188] FIG. 7 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching within the reference picture (Ref0) for the block that most matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoding completed region of either the left adjacent or the upper adjacent side, and the reconstructed image at an equivalent position within the encoding completed reference picture (Ref0) specified in the candidate MV is derived, and an evaluation value is calculated using the obtained difference value, and the candidate MV that has the best evaluation value among the multiple candidate MVs is selected as the best candidate MV.
[0189] Information indicating whether such a FRUC mode is applied (e.g., called the FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the method of pattern matching (e.g., the first pattern matching or the second pattern matching) (e.g., called the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0190] Here, we describe a mode for deriving motion vectors based on a model assuming uniform linear motion. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.
[0191] FIG. 8 is a diagram illustrating a model assuming uniform linear motion. In FIG. 8, (v x , v y ) represents a velocity vector, and τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents a motion vector corresponding to reference picture Ref0, and (MVx1, MVy1) represents a motion vector corresponding to reference picture Ref1.
[0192] At this time, the velocity vector (v) x , v y Under the assumption of uniform linear motion, (MVx0, MVy0) and (MVx1, MVy1) are, respectively, (v x τ0, v y τ0) and (-v x τ1,-v y It is represented as τ1), and the following optical flow equation (1) holds.
[0193] [Number 1]
[0194]
[0195] Here, I (k) represents the luminance value of reference image k (k=0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-unit motion vectors obtained from merge lists, etc., are corrected at the pixel level.
[0196] In addition, motion vectors may be derived at the decoder side in a manner different from the derivation of motion vectors based on a model assuming uniform linear motion. For example, motion vectors may be derived in sub-block units based on the motion vectors of multiple adjacent blocks.
[0197] Here, we describe a mode for deriving motion vectors in sub-block units based on the motion vectors of multiple adjacent blocks. This mode is sometimes referred to as an affine motion compensation prediction mode.
[0198] FIG. 9a is a diagram illustrating the derivation of motion vectors at the sub-block level based on motion vectors of multiple adjacent blocks. In FIG. 9a, the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper-left corner control point of the current block is derived based on the motion vectors of adjacent blocks, and the motion vector v1 of the upper-right corner control point of the current block is derived based on the motion vectors of adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v) of each sub-block within the current block is obtained by the following equation (2). x , v y ) is derived.
[0199] [Number 2]
[0200]
[0201] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents a predetermined weighting factor.
[0202] Such an affine motion compensation prediction mode may include several modes in which the method of deriving motion vectors of the top-left and top-right corner control points differs. Information indicating such an affine motion compensation prediction mode (e.g., called an affine flag) is signaled at the CU level. Furthermore, the signaling of information indicating this affine motion compensation prediction mode is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0203] [Predictive Control Unit]
[0204] The prediction control unit (128) selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit (104) and the addition unit (116).
[0205] Here, an example of deriving motion vectors of a picture to be encoded by merge mode is described. Fig. 9b is a diagram illustrating an overview of the motion vector derivation process by merge mode.
[0206] First, a list of predicted MVs is created by registering candidates for predicted MVs. Candidates for predicted MVs include a spatially adjacent predicted MV, which is an MV of multiple encoded blocks located spatially near the encoded block; a temporally adjacent predicted MV, which is an MV of a nearby block projected from the position of the encoded block in the encoded reference picture; a combined predicted MV, which is an MV generated by combining the MV values of the spatially adjacent predicted MV and the temporally adjacent predicted MV; and a zero predicted MV, which is an MV with a value of zero.
[0207] Next, one of the multiple predicted MVs registered in the predicted MV list is selected to be determined as the MV of the block to be encoded.
[0208] In addition, the variable-length encoding unit encodes by describing the stream with a signal, merge_idx, which indicates which predicted MV was selected.
[0209] In addition, the predicted MV registered in the predicted MV list described in FIG. 9b is an example and may be a number different from the number in the drawing, a configuration that does not include some types of predicted MV in the drawing, or a configuration that adds predicted MVs other than the types of predicted MV in the drawing.
[0210] In addition, the final MV may be determined by performing the DMVR processing described later using the MV of the encoding target block derived by the merge mode.
[0211] Here, an example of determining the MV using DMVR processing is explained.
[0212] Figure 9c is a conceptual diagram illustrating an overview of DMVR processing.
[0213] First, the optimal MVP set as the processing target block is set as a candidate MV, and according to the candidate MV, reference pixels are respectively obtained from a first reference picture, which is a processed picture in the L0 direction, and a second reference picture, which is a processed picture in the L1 direction, and a template is generated by taking the average of each reference pixel.
[0214] Next, using the above template, the surrounding areas of the candidate MVs of the first reference picture and the second reference picture are searched, respectively, and the MV with the minimum cost is determined as the final MV. In addition, the cost value is calculated using the difference between each pixel value of the template and each pixel value of the search area, the MV value, etc.
[0215] In addition, the overview of the processing described here is basically common to the encoding and decoding devices.
[0216] In addition, other processing may be used, even if it is not the processing described here, as long as it is a process that can derive the final MV by searching the surroundings of the candidate MV.
[0217] Here, a mode for generating predicted images using LIC processing is described.
[0218] FIG. 9d is a figure illustrating an overview of a method for generating a predicted image using luminance correction processing by LIC processing.
[0219] First, an MV is derived to obtain a reference image corresponding to the block to be encoded from a reference picture, which is a completed encoded picture.
[0220] Next, for the block to be encoded, the luminance pixel values of the left adjacent and upper adjacent encoding completion surrounding reference areas and the luminance pixel values at the same location within the reference picture specified in the MV are used to extract information indicating how the luminance values have changed in the reference picture and the picture to be encoded, and a luminance correction parameter is calculated.
[0221] By performing luminance correction processing on a reference image within a reference picture specified in MV using the luminance correction parameter, a predicted image for the block to be encoded is generated.
[0222] In addition, the shape of the surrounding reference area in FIG. 9d is an example, and other shapes may be used.
[0223] In addition, although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures, and the predicted image is generated after performing luminance correction processing on the reference picture obtained from each reference picture in the same way.
[0224] As a method for determining whether to apply LIC processing, for example, there is a method using lic_flag, a signal indicating whether to apply LIC processing. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a region where luminance changes occur; if it belongs to a region where luminance changes occur, lic_flag is set to a value of 1 to apply LIC processing and perform encoding, and if it does not belong to a region where luminance changes occur, lic_flag is set to a value of 0 to perform encoding without applying LIC processing. Meanwhile, in a decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply LIC processing according to the value.
[0225] As another method for determining whether to apply LIC processing, for example, there is also a method of determining based on whether neighboring blocks have applied LIC processing. As a specific example, when the block to be encoded is in merge mode, when deriving the MV in the merge mode processing, it is determined whether the selected neighboring encoded block has applied LIC processing and encoded it, and based on the result, the decision to apply LIC processing is switched and encoding is performed. In addition, in this example, the processing during decoding is also completely identical.
[0226] [Overview of the Decoding Device]
[0227] Next, an overview of a decoding device capable of decoding an encoded signal (encoded bit stream) output from the above-described encoding device (100) will be described. FIG. 10 is a block diagram showing the functional configuration of a decoding device (200) according to embodiment 1. The decoding device (200) is a video / image decoding device that decodes video / images in blocks.
[0228] As shown in FIG. 10, the decoder (200) comprises an entropy decoder (202), an inverse quantization unit (204), an inverse transformation unit (206), an adder (208), a block memory (210), a loop filter unit (212), a frame memory (214), an intra prediction unit (216), an inter prediction unit (218), and a prediction control unit (220).
[0229] The decoding device (200) is realized, for example, by a general-purpose processor and memory. In this case, when a software program stored in memory is executed by the processor, the processor functions as an entropy decoding unit (202), an inverse quantization unit (204), an inverse transformation unit (206), an adder (208), a loop filter unit (212), an intra prediction unit (216), an inter prediction unit (218), and a prediction control unit (220). Additionally, the decoding device (200) may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit (202), the inverse quantization unit (204), the inverse transformation unit (206), the adder (208), the loop filter unit (212), the intra prediction unit (216), the inter prediction unit (218), and the prediction control unit (220).
[0230] Below, each component included in the decoding device (200) is described.
[0231] [Entropy Decoding Unit]
[0232] The entropy decoding unit (202) entropy decodes the encoded bit stream. Specifically, the entropy decoding unit (202) performs arithmetic decoding from the encoded bit stream into a binary signal, for example. Then, the entropy decoding unit (202) debinarizes the binary signal. By doing so, the entropy decoding unit (202) outputs quantization coefficients in blocks to the inverse quantization unit (204).
[0233] [Inverse Quantizer]
[0234] The inverse quantization unit (204) inversely quantizes the quantization coefficients of the decoding target block (hereinafter referred to as the current block) which is an input from the entropy decoding unit (202). Specifically, the inverse quantization unit (204) inversely quantizes each of the quantization coefficients of the current block based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit (204) outputs the inversely quantized quantization coefficients (i.e., conversion coefficients) of the current block to the inverse conversion unit (206).
[0235] [Inverse Transformation Section]
[0236] The inverse transformation unit (206) restores the prediction error by inversely transforming the transformation coefficients, which are inputs from the inverse quantization unit (204).
[0237] For example, if the information decoded from the encoded bit stream indicates that EMT or AMT is applied (for example, the AMT flag is true), the inverse conversion unit (206) inversely converts the conversion coefficients of the current block based on the information indicating the decoded conversion type.
[0238] Also, for example, if the information decoded from the encoded bit stream indicates that NSST is applied, the inverse conversion unit (206) applies an inverse re-conversion to the conversion coefficients.
[0239] [Additional part]
[0240] The adder (208) reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit (206), and the prediction sample, which is the input from the prediction control unit (220). Then, the adder (208) outputs the reconstructed block to the block memory (210) and the loop filter unit (212).
[0241] [Block Memory]
[0242] The block memory (210) is a block referenced in intra-prediction and is a memory unit for storing blocks within a decoding target picture (hereinafter referred to as the current picture). Specifically, the block memory (210) stores a reconstruction block output from the adder (208).
[0243] [Loop Filter Section]
[0244] The loop filter unit (212) performs a loop filter on the block reconstructed by the adder (208) and outputs the filtered reconstructed block to the frame memory (214) and display device, etc.
[0245] When the information indicating the on / off status of the ALF decoded from the encoded bit stream indicates that the ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstruction block.
[0246] [Frame Memory]
[0247] The frame memory (214) is a memory unit for storing reference pictures used for inter prediction and is also called a frame buffer. Specifically, the frame memory (214) stores a reconstruction block filtered by the loop filter unit (212).
[0248] [Intra Prediction Section]
[0249] The intra prediction unit (216) generates a prediction signal (intra prediction signal) by performing intra prediction based on an intra prediction mode decoded from an encoded bit stream and referencing a block within a current picture stored in a block memory (210). Specifically, the intra prediction unit (216) generates an intra prediction signal by performing intra prediction based on a sample (e.g., luminance value, color difference value) of a block adjacent to the current block and outputs the intra prediction signal to the prediction control unit (220).
[0250] In addition, when an intra prediction mode that references a luminance block is selected for intra prediction of a color difference block, the intra prediction unit (216) may predict the color difference component of the current block based on the luminance component of the current block.
[0251] Also, when the information decoded from the encoded bit stream indicates the application of PDPC, the intra prediction unit (216) corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction.
[0252] [Inter Prediction Department]
[0253] The inter prediction unit (218) predicts a current block by referring to a reference picture stored in the frame memory (214). The prediction is performed on a unit of a current block or a sub-block within the current block (e.g., a 4×4 block). For example, the inter prediction unit (218) generates an inter prediction signal of a current block or sub-block by performing motion compensation using motion information (e.g., a motion vector) decoded from a encoded bit stream, and outputs the inter prediction signal to the prediction control unit (220).
[0254] In addition, when the information decoded from the encoded bit stream indicates that the OBMC mode is applied, the inter prediction unit (218) generates an inter prediction signal by using not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent block.
[0255] Also, when the information decoded from the encoded bit stream indicates that the FRUC mode is applied, the inter prediction unit (218) derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit (218) performs motion compensation using the derived motion information.
[0256] Additionally, when the BIO mode is applied, the inter prediction unit (218) derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bit stream indicates that the affine motion compensation prediction mode is applied, the inter prediction unit (218) derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0257] [Predictive Control Unit]
[0258] The prediction control unit (220) selects either the intra prediction signal or the inter prediction signal and outputs the selected signal to the adder (208) as a prediction signal.
[0259] [Inter-frame prediction performed by the encoding device and the decoding device]
[0260] FIG. 11 is a block diagram for explaining processing related to inter-frame prediction performed by the encoding device (100) shown in FIG. 1. For example, the input image is divided into blocks by the dividing unit (102) shown in FIG. 1. Then, processing is performed for each block.
[0261] The subtraction unit (104) generates a difference image by acquiring the difference between a block-unit input image and a prediction image generated by intra-frame prediction or inter-frame prediction. Then, the conversion unit (106) and the quantization unit (108) generate a counting signal by performing conversion and quantization on the difference image. The entropy encoding unit (110) generates an encoding stream (encoding bit stream) by performing entropy encoding on the generated counting signal and other encoding signals.
[0262] Meanwhile, the inverse quantization unit (112) and the inverse transformation unit (114) restore the difference image by performing inverse quantization and inverse transformation on the generated coefficient signal. The intra-frame prediction unit (intra prediction unit) (124) generates a prediction image by intra-frame prediction, and the inter-frame prediction unit (inter prediction unit) (126) generates a prediction image by inter-frame prediction. The addition unit (116) generates a reconstructed image by adding one of the prediction image generated by intra-frame prediction and the prediction image generated by inter-frame prediction to the restored difference image.
[0263] Additionally, the in-frame prediction unit (124) uses the reconstructed image of the processed block for in-frame prediction of another block to be processed later. Also, the loop filter unit (120) applies a loop filter to the reconstructed image of the processed block and stores the reconstructed image with the loop filter applied in the frame memory (122). Then, the inter-frame prediction unit (126) uses the reconstructed image stored in the frame memory (122) for inter-frame prediction of another block in another picture to be processed later.
[0264] FIG. 12 is a block diagram for explaining processing related to inter-frame prediction performed by the decoder (200) shown in FIG. 10. By the entropy decoder (202), entropy decoding is performed on the input stream, which is an encoded stream, so that information is obtained in blocks. Then, processing is performed for each block.
[0265] The inverse quantization unit (204) and the inverse transformation unit (206) restore the difference image by performing inverse quantization and inverse transformation on the decoded coefficient signal for each block.
[0266] The intra-frame prediction unit (216) generates a predicted image by intra-frame prediction, and the inter-frame prediction unit (218) generates a predicted image by inter-frame prediction. The addition unit (208) generates a reconstructed image by adding one of the predicted image generated by intra-frame prediction and the predicted image generated by inter-frame prediction to the restored difference image.
[0267] Additionally, the in-frame prediction unit (216) uses the reconstructed image of the processed block for in-frame prediction of another block to be processed later. Additionally, the loop filter unit (212) applies a loop filter to the reconstructed image of the processed block and stores the reconstructed image with the loop filter applied in the frame memory (214). Then, the inter-frame prediction unit (218) uses the reconstructed image stored in the frame memory (214) for inter-frame prediction of another block in another picture to be processed later.
[0268] In addition, inter-frame prediction is also called inter-picture prediction, inter-frame prediction, or inter-prediction. Intra-frame prediction is also called intra-picture prediction, intra-frame prediction, or intra-prediction.
[0269] [Specific example of inter-screen prediction]
[0270] A plurality of specific examples of inter-frame prediction are shown below. For example, one of the plurality of specific examples is applied. Also, although the operation of the encoding device (100) is mainly shown below, the operation of the decoding device (200) is basically the same. In particular, the inter-frame prediction unit (218) of the decoding device (200) operates in the same way as the inter-frame prediction unit (126) of the encoding device (100).
[0271] However, there are cases where the encoding device (100) encodes information regarding inter-frame prediction and the decoding device (200) decodes information regarding inter-frame prediction. For example, the encoding device (100) may encode information of motion vectors for inter-frame prediction. And, the decoding device (200) may decode information of motion vectors for inter-frame prediction.
[0272] The motion vector information above is information related to the motion vector and represents the motion vector directly or indirectly. For example, the motion vector information may represent the motion vector itself, or it may represent the difference motion vector, which is the difference between the motion vector and the predicted motion vector, and the identifier of the predicted motion vector.
[0273] Also, the processing performed for each block is shown below. This block is also referred to as a prediction block. A block may be an image data unit that undergoes encoding and decoding, or an image data unit that undergoes reconstruction. Alternatively, the block in the following description may be a more detailed image data unit, or an image data unit called a sub-block.
[0274] In addition, the OBMC processing described below is a process that corrects the predicted image of the block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vectors of the blocks surrounding the block to be processed. As for the OBMC processing, the method described below using FIGS. 17 and 18 may be used, or other methods may be used.
[0275] In addition, the BIO processing described below is a process that generates a predicted image of a block to be processed by referring to the spatial gradient of luminance in an image obtained by performing motion compensation of the block to be processed using the motion vector of the block to be processed. As for the BIO processing, the method described below using FIGS. 19 and 20 may be used, or other methods may be used.
[0276] In addition, the processing corresponding to the processing target block and processing completed picture, etc. in the following description is, for example, encoding or decoding processing, and may include prediction processing or reconstruction processing.
[0277] FIG. 13 is a flowchart showing a first specific example of inter-frame prediction performed by the inter-frame prediction unit (126) of the encoding device (100). As described above, the inter-frame prediction unit (218) of the decoding device (200) operates in the same way as the inter-frame prediction unit (126) of the encoding device (100).
[0278] As shown in FIG. 13, the inter-frame prediction unit (126) performs processing for each block. At that time, the inter-frame prediction unit (126) derives the motion vector (MV) of the block to be processed (S101). Any derivation method may be used as a method for deriving the motion vector.
[0279] For example, a derivation method may be used in which the same motion vector search processing is performed in the encoding device (100) and the decoding device (200). Also, a derivation method may be used in which the motion vector is transformed based on an affine transformation.
[0280] Additionally, a derivation method available only in the encoding device (100) among the encoding device (100) and the decoding device (200) may be used in the encoding device (100). Furthermore, information of the motion vector may be encoded into a stream in the encoding device (100) and decoded from the stream in the decoding device (200). By doing so, the encoding device (100) and the decoding device (200) may derive the same motion vector.
[0281] Next, the inter-frame prediction unit (126) determines whether to apply BIO processing to the generation process of the predicted image of the block to be processed (S102). For example, the inter-frame prediction unit (126) may determine that BIO processing is applied when generating the predicted image of the block to be processed by simultaneously referencing two completed pictures, and may determine that BIO processing is not applied in other cases.
[0282] Additionally, the inter-frame prediction unit (126) may determine that BIO processing is applied when generating a predicted image of a processing target block by simultaneously referencing a processing completed picture that is displayed in order prior to the processing target picture and a processing completed picture that is displayed in order following the processing target picture. Furthermore, the inter-frame prediction unit (126) may determine that BIO processing is not applied in other cases.
[0283] Then, when the inter-frame prediction unit (126) determines that BIO processing is applied to the generation process of the predicted image of the block to be processed (applied in S102), it generates the predicted image of the block to be processed by BIO processing (S106). In this case, the inter-frame prediction unit (126) does not correct the predicted image of the block to be processed by OBMC processing. Then, without applying OBMC processing, the predicted image generated by BIO processing is used as the final predicted image of the block to be processed.
[0284] Meanwhile, if the inter-frame prediction unit (126) determines that BIO processing is not applied to the generation process of the predicted image of the block to be processed (no application in S102), it generates the predicted image of the block to be processed in a normal manner (S103). In this example, generating the predicted image of the block to be processed in a normal manner means generating the predicted image of the block to be processed without applying BIO processing. For example, the inter-frame prediction unit (126) generates the predicted image by normal motion compensation without applying BIO processing.
[0285] Then, the inter-frame prediction unit (126) determines whether to apply OBMC processing to the generation processing of the predicted image of the block to be processed (S104). That is, the inter-frame prediction unit (126) determines whether to correct the predicted image of the block to be processed by OBMC processing. For example, the inter-frame prediction unit (126) may determine not to apply OBMC processing when a derivation method that transforms the motion vector based on an affine transformation is used for deriving the motion vector. In other cases, the inter-frame prediction unit (126) may determine to apply OBMC processing.
[0286] Then, when the inter-frame prediction unit (126) determines that OBMC processing is applied to the generation processing of the predicted image of the block to be processed (applied in S104), it corrects the predicted image of the block to be processed by OBMC processing (S105). Then, the predicted image corrected by OBMC processing is used as the final predicted image of the block to be processed.
[0287] Meanwhile, if the inter-frame prediction unit (126) determines that OBMC processing is not applied to the generation process of the prediction image of the block to be processed (no application in S104), the prediction image of the block to be processed is not corrected by OBMC processing. Then, the prediction image generated by the normal method without OBMC processing being applied is used as the final prediction image of the block to be processed.
[0288] By the above operation, it is possible to eliminate the pass of applying BIO processing and also applying OBMC processing in the generation process of the predicted image of the block to be processed. Therefore, it is possible to reduce the throughput of the pass with a large throughput. Consequently, the circuit size can be reduced.
[0289] In addition, in both BIO processing and OBMC processing, the spatial continuity of pixel values between blocks is taken into account. Therefore, when both BIO processing and OBMC processing are applied, it is assumed that the same function is applied twice. In the above operation, such drawbacks are eliminated.
[0290] Also, when BIO processing is applied, the degree to which the code amount is reduced by the application of OBMC processing is assumed to be lower than when BIO processing is not applied. That is, when it is assumed that the degree to which the code amount is reduced by the application of OBMC processing is low, the inter-frame prediction unit (126) can efficiently generate a prediction image without applying OBMC processing.
[0291] Additionally, the determination of whether to apply BIO processing (S102) can be considered as a determination of whether OBMC processing is applicable. That is, the inter-frame prediction unit (126) determines whether OBMC processing is applicable based on whether BIO processing is applied in the determination of whether to apply BIO processing (S102). And, for example, if BIO processing is applied, the inter-frame prediction unit (126) determines that OBMC processing is not applicable and applies BIO processing without applying OBMC processing.
[0292] Additionally, regarding the determination of whether BIO processing is applied (S102), the encoding device (100) may encode a signal indicating whether BIO processing is applied into a stream by the entropy encoding unit (110). Then, the decoding device (200) may decode a signal indicating whether BIO processing is applied from the stream by the entropy decoding unit (202). The encoding and decoding of the signal indicating whether BIO processing is applied may be performed per block, per slice, per picture, or per sequence.
[0293] Likewise, regarding the determination of whether OBMC processing is applied (S104), the encoding device (100) may encode a signal indicating whether OBMC processing is applied into a stream by the entropy encoding unit (110). And, the decoding device (200) may decode a signal indicating whether OBMC processing is applied from the stream by the entropy decoding unit (202). The encoding and decoding of the signal indicating whether OBMC processing is applied may be performed per block, per slice, per picture, or per sequence.
[0294] FIG. 14 is a flowchart showing a variation of the first embodiment shown in FIG. 13. The inter-frame prediction unit (126) of the encoding device (100) may perform the operation of FIG. 14. Likewise, the inter-frame prediction unit (218) of the decoding device (200) may perform the operation of FIG. 14.
[0295] As shown in FIG. 14, the inter-frame prediction unit (126) performs processing for each block. At that time, similar to the example in FIG. 13, the inter-frame prediction unit (126) derives the motion vector of the block to be processed (S111). Then, the inter-frame prediction unit (126) performs branching processing according to the prediction image generation mode (S112).
[0296] For example, when the prediction image generation mode is 0 (0 in S112), the inter-frame prediction unit (126) generates a prediction image of the block to be processed in a normal way (S113). In this case, the inter-frame prediction unit (126) does not correct the prediction image of the block to be processed by OBMC processing. Then, the prediction image generated in a normal way without OBMC processing being applied is used as the final prediction image of the block to be processed.
[0297] In this example, generating a predicted image of a block to be processed in a normal manner means generating a predicted image of a block to be processed without applying BIO processing. For example, the inter-frame prediction unit (126) generates a predicted image by normal motion compensation without applying BIO processing.
[0298] Also, for example, when the prediction image generation mode is 1 (1 in S112), the inter-frame prediction unit (126) generates a prediction image of the block to be processed in a normal manner, just as when the prediction image generation mode is 0 (S114). When the prediction image generation mode is 1 (1 in S112), the inter-frame prediction unit (126) also corrects the prediction image of the block to be processed by OBMC processing (S115). Then, the prediction image corrected by OBMC processing is used as the final prediction image of the block to be processed.
[0299] Also, for example, when the prediction image generation mode is 2 (2 in S112), the inter-frame prediction unit (126) generates a prediction image of the block to be processed by BIO processing (S116). In this case, the inter-frame prediction unit (126) does not correct the prediction image of the block to be processed by OBMC processing. Then, without applying OBMC processing, the prediction image generated by BIO processing is used as the final prediction image of the block to be processed.
[0300] The inter-frame prediction unit (126) can efficiently generate a prediction image through the above operation, just like the example in FIG. 13. That is, the same effect as the example in FIG. 13 can be obtained.
[0301] Additionally, the branching process (S112) can be considered as a combined process of the determination of application of BIO processing (S102) and the determination of application of OBMC processing (S104) in the example of FIG. 13. For example, when the prediction image generation mode is 2, it is determined that OBMC processing cannot be applied because BIO processing is applied. And in this case, OBMC processing is not applied, and BIO processing is applied.
[0302] Additionally, regarding branch processing (S112), the encoding device (100) may encode a signal representing a predicted image generation mode into a stream by the entropy encoding unit (110). For example, the encoding device (100) may select a predicted image generation mode for each block in the inter-frame prediction unit (126) and encode a signal representing the selected predicted image generation mode into a stream in the entropy encoding unit (110).
[0303] Additionally, the decoding device (200) may decode a signal indicating a predicted image generation mode from the stream by means of an entropy decoding unit (202). For example, the decoding device (200) may decode a signal indicating a predicted image generation mode for each block in the entropy decoding unit (202), and select a predicted image generation mode indicated by the decoded signal in the inter-frame prediction unit (218).
[0304] Alternatively, a prediction image generation mode may be selected for each slice, each picture, or each sequence, and encoding and decoding of a signal representing the selected prediction image generation mode may be performed.
[0305] In addition, in the above example, three predicted image generation modes are shown. These predicted image generation modes are examples, and other predicted image generation modes may be used. Also, the three numbers corresponding to the three predicted image generation modes are examples, and other numbers may be used. Furthermore, additional numbers may be used for the other predicted image generation modes.
[0306] FIG. 15 is a flowchart showing a second specific example of inter-frame prediction performed by the inter-frame prediction unit (126) of the encoding device (100). As described above, the inter-frame prediction unit (218) of the decoding device (200) operates in the same way as the inter-frame prediction unit (126) of the encoding device (100).
[0307] As shown in FIG. 15, the inter-frame prediction unit (126) performs processing for each block. At that time, similar to the example in FIG. 13, the inter-frame prediction unit (126) derives the motion vector of the block to be processed (S201).
[0308] Next, the inter-frame prediction unit (126) determines whether to apply pair prediction processing that references two completed pictures to the generation processing of the prediction image of the block to be processed (S202).
[0309] Then, when the inter-frame prediction unit (126) determines that pair prediction processing is applied to the generation process of the prediction image of the block to be processed (applied in S202), it generates the prediction image of the block to be processed by pair prediction processing (S206). In this case, the inter-frame prediction unit (126) does not correct the prediction image of the block to be processed by OBMC processing. Then, without applying OBMC processing, the prediction image generated by pair prediction processing is used as the final prediction image of the block to be processed.
[0310] Meanwhile, if the inter-frame prediction unit (126) determines that pair prediction processing is not applied to the generation process of the predicted image of the block to be processed (no application in S202), it generates the predicted image of the block to be processed in a normal manner (S203). In this example, generating the predicted image of the block to be processed in a normal manner means generating the predicted image of the block to be processed without applying pair prediction processing. For example, the inter-frame prediction unit (126) generates the predicted image from one completed picture by normal motion compensation.
[0311] And, the inter-frame prediction unit (126) determines whether to apply OBMC processing to the generation processing of the predicted image of the processing target block, as in the example of FIG. 13 (S204).
[0312] Then, when the inter-frame prediction unit (126) determines that OBMC processing is applied to the generation processing of the predicted image of the block to be processed (applied in S204), it corrects the predicted image of the block to be processed by OBMC processing (S205). Then, the predicted image corrected by OBMC processing is used as the final predicted image of the block to be processed.
[0313] Meanwhile, if the inter-frame prediction unit (126) determines that OBMC processing is not applied to the generation process of the prediction image of the block to be processed (no application in S204), the prediction image of the block to be processed is not corrected by OBMC processing. Then, the prediction image generated by the normal method without OBMC processing being applied is used as the final prediction image of the block to be processed.
[0314] By the above operation, in the process of generating a predicted image of a block to be processed, it is possible to eliminate the pass of applying paired prediction processing and also applying OBMC processing. That is, the inter-frame prediction unit (126) corrects the predicted image generated from one completed picture through OBMC processing, but does not correct the predicted image generated from each of the two completed pictures. Therefore, it is possible to reduce the throughput of the pass to which OBMC processing is applied. Accordingly, the circuit size can be reduced.
[0315] In addition, for example, there are cases where it is specified that BIO processing is applied in paired prediction processing that references two completed pictures. In such cases, by determining that OBMC processing cannot be applied in paired prediction processing that references two completed pictures, it is possible to avoid applying both BIO processing and OBMC processing simultaneously.
[0316] Additionally, the determination of whether to apply paired prediction processing (S202) can be considered as a determination of whether OBMC processing is applicable. That is, the inter-frame prediction unit (126) determines whether OBMC processing is applicable based on whether paired prediction processing is applied in the determination of whether to apply paired prediction processing (S202). And, for example, if paired prediction processing is applied, the inter-frame prediction unit (126) determines that OBMC processing is not applicable and applies paired prediction processing without applying OBMC processing.
[0317] Additionally, regarding the determination of whether pair prediction processing is applied (S202), the encoding device (100) may encode a signal indicating whether pair prediction processing is applied into a stream by the entropy encoding unit (110). Then, the decoding device (200) may decode a signal indicating whether pair prediction processing is applied from the stream by the entropy decoding unit (202). The encoding and decoding of the signal indicating whether pair prediction processing is applied may be performed per block, per slice, per picture, or per sequence.
[0318] Likewise, regarding the determination of whether OBMC processing is applied (S204), the encoding device (100) may encode a signal indicating whether OBMC processing is applied into a stream by the entropy encoding unit (110). And, the decoding device (200) may decode a signal indicating whether OBMC processing is applied from the stream by the entropy decoding unit (202). The encoding and decoding of the signal indicating whether OBMC processing is applied may be performed per block, per slice, per picture, or per sequence.
[0319] FIG. 16 is a flowchart showing a variation of the second embodiment shown in FIG. 15. The inter-frame prediction unit (126) of the encoding device (100) may perform the operation of FIG. 16. Likewise, the inter-frame prediction unit (218) of the decoding device (200) may perform the operation of FIG. 16.
[0320] As shown in FIG. 16, the inter-frame prediction unit (126) performs processing for each block. At that time, similar to the example in FIG. 15, the inter-frame prediction unit (126) derives the motion vector of the block to be processed (S211). Then, the inter-frame prediction unit (126) performs branching processing according to the prediction image generation mode (S212).
[0321] For example, when the prediction image generation mode is 0 (0 in S212), the inter-frame prediction unit (126) generates a prediction image of the block to be processed in a normal manner (S213). In this case, the inter-frame prediction unit (126) does not correct the prediction image of the block to be processed by OBMC processing. Then, the prediction image generated in a normal manner without OBMC processing being applied is used as the final prediction image of the block to be processed.
[0322] In this example, generating a predicted image of a block to be processed in a normal manner means generating a predicted image of a block to be processed without applying pair prediction processing that references two completed pictures. For example, the inter-frame prediction unit (126) generates a predicted image by motion compensation from one completed picture.
[0323] Also, for example, when the prediction image generation mode is 1 (1 in S212), the inter-frame prediction unit (126) generates a prediction image of the block to be processed in a normal manner, just as when the prediction image generation mode is 0 (S214). When the prediction image generation mode is 1 (1 in S212), the inter-frame prediction unit (126) also corrects the prediction image of the block to be processed by OBMC processing (S215). Then, the prediction image corrected by OBMC processing is used as the final prediction image of the block to be processed.
[0324] Also, for example, when the prediction image generation mode is 2 (2 in S212), the inter-frame prediction unit (126) generates a prediction image of the block to be processed by pair prediction processing (S216). In this case, the inter-frame prediction unit (126) does not correct the prediction image of the block to be processed by OBMC processing. And, without applying OBMC processing, the prediction image generated by pair prediction processing is used as the final prediction image of the block to be processed.
[0325] The inter-frame prediction unit (126) can efficiently generate a prediction image through the above operation, just like the example in FIG. 15. That is, the same effect as the example in FIG. 15 can be obtained.
[0326] Additionally, the branching process (S212) can be considered as a combined process of the determination of application of pair prediction processing (S202) and the determination of application of OBMC processing (S204) in the example of FIG. 15. For example, when the prediction image generation mode is 2, it is determined that OBMC processing cannot be applied because pair prediction processing is applied. In this case, OBMC processing is not applied, and pair prediction processing is applied.
[0327] Additionally, regarding branch processing (S212), the encoding device (100) may encode a signal representing a predicted image generation mode into a stream by the entropy encoding unit (110). For example, the encoding device (100) may select a predicted image generation mode for each block in the inter-frame prediction unit (126) and encode a signal representing the selected predicted image generation mode into a stream in the entropy encoding unit (110).
[0328] Additionally, the decoding device (200) may decode a signal indicating a predicted image generation mode from the stream by means of an entropy decoding unit (202). For example, the decoding device (200) may decode a signal indicating a predicted image generation mode for each block in the entropy decoding unit (202), and select a predicted image generation mode indicated by the decoded signal in the inter-frame prediction unit (218).
[0329] Alternatively, a prediction image generation mode may be selected for each slice, each picture, or each sequence, and encoding and decoding of a signal representing the selected prediction image generation mode may be performed.
[0330] In addition, in the above example, three predicted image generation modes are shown. These predicted image generation modes are examples, and other predicted image generation modes may be used. Also, the three numbers corresponding to the three predicted image generation modes are examples, and other numbers may be used. Furthermore, additional numbers may be used for the other predicted image generation modes.
[0331] In addition, in the description of FIGS. 15 and 16, the pair prediction processing is defined as a prediction processing that references two completed pictures. However, the pair prediction processing may also be defined as a prediction processing that references a completed picture that is displayed earlier than the picture to be processed and a completed picture that is displayed later than the picture to be processed.
[0332] In addition, the usual predictive processing may include predictive processing that refers to two picture to be processed in display order preceding the picture to be processed, and predictive processing that refers to two picture to be processed in display order following the picture to be processed.
[0333] For example, there are cases where it is stipulated that BIO processing is applied in paired prediction processing that references the previous completed picture and the subsequent completed picture in display order. In such cases, by determining that OBMC processing cannot be applied in paired prediction processing that references the previous completed picture and the subsequent completed picture in display order, the application of both BIO processing and OBMC processing can be avoided.
[0334] [OBMC Processing]
[0335] FIG. 17 is a conceptual diagram illustrating OBMC processing in an encoding device (100) and a decoding device (200). In OBMC processing, the predicted image of the block to be processed is corrected using an image obtained by performing motion compensation of the block to be processed using the motion vector of a block surrounding the block to be processed.
[0336] For example, before OBMC processing is applied, motion compensation of the block to be processed is performed using the motion vector (MV) of the block to be processed, thereby generating a predicted image (Pred) based on the motion vector (MV) of the block to be processed. Then, OBMC processing is applied.
[0337] Specifically, by performing motion compensation of the block to be processed using the motion vector (MV_L) of the left adjacent block, which is a processed block adjacent to the left of the block to be processed, an image (Pred_L) based on the motion vector (MV_L) of the left adjacent block is generated. In addition, by performing motion compensation of the block to be processed using the motion vector (MV_U) of the upper adjacent block, which is a processed block adjacent to the block to be processed, an image (Pred_U) based on the motion vector (MV_U) of the left adjacent block is generated.
[0338] Then, the predicted image (Pred) based on the motion vector (MV) of the block to be processed is corrected by the image (Pred_L) based on the motion vector (MV_L) of the left adjacent block and the image (Pred_U) based on the motion vector (MV_U) of the left adjacent block. By doing so, a predicted image with OBMC processing applied is derived.
[0339] FIG. 18 is a flowchart showing the operation performed by the inter-frame prediction unit (126) of the encoding device (100) as OBMC processing. The inter-frame prediction unit (218) of the decoding device (200) operates in the same way as the inter-frame prediction unit (126) of the encoding device (100).
[0340] For example, the inter-frame prediction unit (126) obtains a prediction image (Pred) based on the motion vector (MV) of the block to be processed by performing motion compensation of the block to be processed using the motion vector (MV) of the block to be processed before OBMC processing. Then, in OBMC processing, the inter-frame prediction unit (126) obtains the motion vector (MV_L) of the left adjacent block (S301).
[0341] Then, the inter-frame prediction unit (126) performs motion compensation of the block to be processed using the motion vector (MV_L) of the left adjacent block. By doing so, the inter-frame prediction unit (126) obtains an image (Pred_L) based on the motion vector (MV_L) of the left adjacent block (S302). For example, the inter-frame prediction unit (126) obtains an image (Pred_L) based on the motion vector (MV_L) of the left adjacent block by referring to a reference block indicated from the block to be processed by the motion vector (MV_L) of the left adjacent block.
[0342] Then, the inter-frame prediction unit (126) applies weights to the predicted image (Pred) based on the motion vector (MV) of the block to be processed and the image (Pred_L) based on the motion vector (MV_L) of the left adjacent block, and superimposes them. By doing so, the inter-frame prediction unit (126) performs a first correction of the predicted image (S303).
[0343] Next, the inter-frame prediction unit (126) performs the same processing on the upper adjacent block as on the left adjacent block. Specifically, it obtains the motion vector (MV_L) of the left adjacent block (S304).
[0344] Then, the inter-frame prediction unit (126) performs motion compensation of the block to be processed using the motion vector (MV_U) of the upper adjacent block. By doing so, the inter-frame prediction unit (126) acquires an image (Pred_U) based on the motion vector (MV_U) of the upper adjacent block (S305). For example, the inter-frame prediction unit (126) acquires an image (Pred_U) based on the motion vector (MV_U) of the upper adjacent block by referring to a reference block indicated from the block to be processed by the motion vector (MV_U) of the upper adjacent block.
[0345] Then, the inter-frame prediction unit (126) applies weights to the prediction image that has undergone the first correction and to the image (Pred_U) based on the motion vector (MV_U) of the upper adjacent block, and superimposes them. By doing so, the inter-frame prediction unit (126) performs a second correction of the prediction image (S306). By doing so, a prediction image with OBMC processing applied is derived.
[0346] In addition, in the above description, two stages of correction are performed using the left adjacent block and the upper adjacent block. However, the right adjacent block or the lower adjacent block may also be used. Furthermore, three or more stages of correction may be performed. Alternatively, one stage of correction may be performed using any one of these blocks. Alternatively, one stage of correction may be performed using multiple blocks simultaneously without dividing it into multiple stages.
[0347] For example, a predicted image (Pred) based on the motion vector (MV) of the block to be processed, an image (Pred_L) based on the motion vector (MV_L) of the left adjacent block, and an image (Pred_U) based on the motion vector (MV_U) of the upper adjacent block may be superimposed. In that case, weighting may be performed.
[0348] In addition, in the above description, OBMC processing is applied to prediction processing referencing one completed picture, but OBMC processing can be applied in the same way to prediction processing referencing two completed pictures. For example, by applying OBMC processing to prediction processing referencing two completed pictures, the prediction images obtained from each of the two completed pictures are corrected.
[0349] In addition, the block to be processed in the above description may be a block called a prediction block, an image data unit that is encoded and decoded, or an image data unit that is reconstructed. Alternatively, the block to be processed in the above description may be a more detailed image data unit or an image data unit called a sub-block.
[0350] Additionally, the encoding device (100) and the decoding device (200) can apply a common OBMC processing. That is, the encoding device (100) and the decoding device (200) can apply OBMC processing in the same way.
[0351] [BIO Processing]
[0352] FIG. 19 is a conceptual diagram illustrating BIO processing in an encoding device (100) and a decoding device (200). In BIO processing, a predicted image of a block to be processed is generated by referring to the spatial gradient of brightness in an image obtained by performing motion compensation of a block to be processed using a motion vector of a block to be processed.
[0353] Before BIO processing, two motion vectors of the block to be processed, L0 motion vector (MV_L0) and L1 motion vector (MV_L1), are derived. L0 motion vector (MV_L0) is a motion vector for referencing L0 reference picture, which is a completed picture, and L1 motion vector (MV_L1) is a motion vector for referencing L1 reference picture, which is a completed picture. L0 reference picture and L1 reference picture are two reference pictures that are simultaneously referenced in the pairwise prediction processing of the block to be processed.
[0354] As a method for deriving L0 motion vectors (MV_L0) and L1 motion vectors (MV_L1), a standard inter-frame prediction mode, a merge mode, or a FRUC mode may be used. For example, in a standard inter-frame prediction mode, motion vectors are derived by performing motion detection using an image of a block to be processed in an encoding device (100), and information of the motion vectors is encoded. Also, in a standard inter-frame prediction mode, motion vectors are derived by decoding information of the motion vectors in a decoder (200).
[0355] In addition, in BIO processing, an L0 predicted image is obtained by performing motion compensation of the target block using the L0 motion vector (MV_L0) by referring to the L0 reference picture. For example, an L0 predicted image may be obtained by applying a motion compensation filter to an image of the L0 reference pixel range, which includes the block indicated in the L0 reference picture and its surroundings, from the target block to the L0 motion vector (MV_L0).
[0356] Additionally, an L0 gradient image is acquired, which represents the spatial gradient of luminance at each pixel of the L0 prediction image. For example, the L0 gradient image is acquired by referencing the luminance of each pixel within the L0 reference pixel range, which includes the block indicated in the L0 reference picture from the processing target block and its surroundings, using the L0 motion vector (MV_L0).
[0357] Additionally, an L1 predicted image is obtained by performing motion compensation of the target block using the L1 motion vector (MV_L1) with reference to the L1 reference picture. For example, an L1 predicted image may be obtained by applying a motion compensation filter to an image of the L1 reference pixel range, which includes the block indicated in the L1 reference picture and its surroundings, from the target block to the L1 motion vector (MV_L1).
[0358] Additionally, an L1 gradient image is acquired, which represents the spatial gradient of luminance at each pixel of the L1 prediction image. For example, the L1 gradient image is acquired by referencing the luminance of each pixel within an L1 reference pixel range that includes the block indicated in the L1 reference picture from the processing target block and its surroundings, using the L1 motion vector (MV_L1).
[0359] Then, for each pixel of the block to be processed, a local motion estimate is derived. Specifically, the pixel value at the corresponding pixel location in the L0 prediction image, the gradient value at the corresponding pixel location in the L0 gradient image, the pixel value at the corresponding pixel location in the L1 prediction image, and the gradient value at the corresponding pixel location in the L1 gradient image are used. The local motion estimate may also be referred to as a corrected motion vector (corrected MV).
[0360] Then, for each pixel of the block to be processed, a pixel correction value is derived using the gradient value at the corresponding pixel location in the L0 gradient image, the gradient value at the corresponding pixel location in the L1 gradient image, and the local motion estimate value. Then, for each pixel of the block to be processed, a predicted pixel value is derived using the pixel value at the corresponding pixel location in the L0 prediction image, the pixel value at the corresponding pixel location in the L1 prediction image, and the pixel correction value. By this, a predicted image with BIO processing applied is derived.
[0361] That is, the predicted pixel value obtained from the pixel value at the corresponding pixel location in the L0 predicted image and the pixel value at the corresponding pixel location in the L1 predicted image is corrected by the pixel correction value. Conversely, the predicted image obtained from the L0 predicted image and the L1 predicted image is corrected using the spatial gradient of luminance in the L0 predicted image and the L1 predicted image.
[0362] FIG. 20 is a flowchart showing the operation performed by the inter-frame prediction unit (126) of the encoding device (100) as BIO processing. The inter-frame prediction unit (218) of the decoding device (200) operates in the same way as the inter-frame prediction unit (126) of the encoding device (100).
[0363] First, the inter-frame prediction unit (126) obtains an L0 predicted image by referring to an L0 reference picture by the L0 motion vector (MV_L0) (S401). Then, the inter-frame prediction unit (126) obtains an L0 gradient image by referring to an L0 reference picture by the L0 motion vector (S402).
[0364] Likewise, the inter-frame prediction unit (126) obtains an L1 predicted image by referring to the L1 reference picture by the L1 motion vector (MV_L1) (S401). Then, the inter-frame prediction unit (126) obtains an L1 gradient image by referring to the L1 reference picture by the L1 motion vector (S402).
[0365] Next, the inter-frame prediction unit (126) derives a local motion estimate value for each pixel of the block to be processed (S411). At that time, the pixel value of the corresponding pixel location in the L0 prediction image, the gradient value of the corresponding pixel location in the L0 gradient image, the pixel value of the corresponding pixel location in the L1 prediction image, and the gradient value of the corresponding pixel location in the L1 gradient image are used.
[0366] Then, the inter-frame prediction unit (126) derives a pixel correction value for each pixel of the block to be processed by using the gradient value of the corresponding pixel position in the L0 gradient image, the gradient value of the corresponding pixel position in the L1 gradient image, and the local motion estimation value. Then, the inter-frame prediction unit (126) derives a predicted pixel value for each pixel of the block to be processed by using the pixel value of the corresponding pixel position in the L0 prediction image, the pixel value of the corresponding pixel position in the L1 prediction image, and the pixel correction value (S412).
[0367] By the above operation, the inter-frame prediction unit (126) generates a prediction image to which BIO processing has been applied.
[0368] In addition, in deriving the local motion estimate and pixel correction value, specifically, the following equation (3) may be used.
[0369] [Number 3]
[0370]
[0371] In Equation (3), I x 0[x, y] is the horizontal gradient value at the pixel position [x, y] of the L0 gradient image. I x 1 [x, y] is the horizontal gradient value at the pixel position [x, y] of the L1 gradient image. I y 0 [x, y] is the vertical gradient value at the pixel position [x, y] of the L0 gradient image. I y 1 [x, y] is the vertical gradient value at pixel position [x, y] of the L1 gradient image.
[0372] Also, in Equation (3), I 0 [x, y] is the pixel value at pixel location [x, y] of the L0 predicted image. I 1 [x, y] is the pixel value at pixel location [x, y] of the L1 predicted image. ΔI [x, y] is the difference between the pixel value at pixel location [x, y] of the L0 predicted image and the pixel value at pixel location [x, y] of the L1 predicted image.
[0373] Also, in Equation (3), Ω is, for example, a set of pixel locations included in the region centered at pixel location [x, y]. w[i, j] is a weighting coefficient for pixel location [i, j]. The same value may be used for w[i, j]. G x [x, y], G y [x, y], G x G y [x, y], sG x G y [x, y], sG x 2 [x, y], sG y 2 [x, y], sG x dI [x, y] and sG y dI [x, y], etc. are auxiliary output values.
[0374] Also, in Equation (3), u[x, y] is a horizontal value that constitutes the local motion estimate at pixel location [x, y]. v[x, y] is a vertical value that constitutes the local motion estimate at pixel location [x, y]. b[x, y] is a pixel correction value at pixel location [x, y]. p[x, y] is a predicted pixel value at pixel location [x, y].
[0375] In addition, in the above description, the inter-frame prediction unit (126) derives a local motion estimate value for each pixel, but may also derive a local motion estimate value for each sub-block, which is an image data unit larger than a pixel and more detailed than a block to be processed.
[0376] For example, in the above equation (3), Ω may be a set of pixel locations included in a sub-block. And, sG x G y [x, y], sG x 2 [x, y], sG y 2 [x, y], sG x dI[x, y], sG y dI [x, y], u [x, y] and v [x, y] may be calculated for each sub-block, rather than for each pixel.
[0377] Additionally, the encoding device (100) and the decoding device (200) can apply common BIO processing. That is, the encoding device (100) and the decoding device (200) can apply BIO processing in the same way.
[0378] [Example of implementation of an encoding device]
[0379] FIG. 21 is a block diagram showing an example of implementation of an encoding device (100) according to embodiment 1. The encoding device (100) is equipped with a circuit (160) and a memory (162). For example, a plurality of components of the encoding device (100) shown in FIG. 1 and FIG. 11 are implemented by the circuit (160) and memory (162) shown in FIG. 21.
[0380] The circuit (160) is a circuit that performs information processing and is capable of accessing memory (162). For example, the circuit (160) is a dedicated or general-purpose electronic circuit that encodes video images. The circuit (160) may be a processor such as a CPU. Also, the circuit (160) may be a collection of multiple electronic circuits. Also, for example, the circuit (160) may serve as a plurality of components among the plurality of components of the encoding device (100) shown in FIG. 1, etc., excluding the component for storing information.
[0381] The memory (162) is a dedicated or general-purpose memory in which information for encoding a video image by the circuit (160) is stored. The memory (162) may be an electronic circuit and may be connected to the circuit (160). Also, the memory (162) may be included in the circuit (160). Also, the memory (162) may be an assembly of multiple electronic circuits. Also, the memory (162) may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium. Also, the memory (162) may be a non-volatile memory or a volatile memory.
[0382] For example, the memory (162) may store an image to be encoded, or a bit sequence corresponding to the encoded image may be stored. Also, the memory (162) may store a program for the circuit (160) to encode the image.
[0383] Also, for example, the memory (162) may serve as a component for storing information among the plurality of components of the encoding device (100) shown in FIG. 1, etc. Specifically, the memory (162) may serve as a block memory (118) and frame memory (122) shown in FIG. 1. More specifically, the memory (162) may store a reconstructed block (processed block) and a reconstructed picture (processed picture), etc.
[0384] In addition, in the encoding device (100), all of the plurality of components shown in FIG. 1, etc., do not need to be implemented, and all of the aforementioned plurality of processes do not need to be performed. Some of the plurality of components shown in FIG. 1, etc., may be included in another device, and some of the aforementioned plurality of processes may be executed by another device. And, in the encoding device (100), some of the plurality of components shown in FIG. 1, etc., are implemented, and some of the aforementioned plurality of processes are performed, thereby allowing prediction processing to be performed efficiently.
[0385] FIG. 22 is a flowchart illustrating a first operation example of the encoding device (100) shown in FIG. 21. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 22 when encoding a video image using picture-to-picture prediction.
[0386] Specifically, the circuit (160) of the encoding device (100) determines whether OBMC processing can be applied to the generation of a predicted image of a processing target block depending on whether BIO processing is applied to the generation of a predicted image of a processing target block (S501). Here, the processing target block is a block included in the processing target picture in the video image.
[0387] Also, BIO processing is a process that generates a predicted image of a block to be processed by referencing the spatial gradient of luminance in an image obtained by performing motion compensation of the block to be processed using the motion vector of the block to be processed. Also, OBMC processing is a process that corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of a block surrounding the block to be processed.
[0388] Then, the circuit (160) determines that OBMC processing cannot be applied to the generation of a predicted image of a block to be processed when BIO processing is applied to the generation of a predicted image of a block to be processed. Then, the circuit (160) does not apply OBMC processing to the generation of a predicted image of a block to be processed and applies BIO processing (S502).
[0389] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0390] Also, when BIO processing is applied, the degree to which the code amount is reduced by the application of OBMC processing is assumed to be lower than when BIO processing is not applied. That is, when it is assumed that the degree to which the code amount is reduced by the application of OBMC processing is low, the encoding device (100) can efficiently generate a predicted image without applying OBMC processing.
[0391] Therefore, the encoding device (100) can efficiently perform prediction processing. In addition, the circuit size can be reduced.
[0392] Additionally, the circuit (160) may also determine that OBMC processing is not applicable to the predictive processing of the block to be processed when an affine transformation is applied to the predictive processing of the block to be processed.
[0393] FIG. 23 is a flowchart showing a variation of the first operation example shown in FIG. 22. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 23 when encoding a video image using picture-to-picture prediction.
[0394] Specifically, the circuit (160) of the encoding device (100) selects one mode among the first mode, the second mode, and the third mode (S511). Here, the first mode is a mode in which BIO processing is not applied and OBMC processing is not applied to the generation processing of the predicted image of the block to be processed. The second mode is a mode in which OBMC processing is applied without BIO processing to the generation processing of the predicted image of the block to be processed. The third mode is a mode in which BIO processing is applied without OBMC processing to the generation processing of the predicted image of the block to be processed.
[0395] Then, the circuit (160) performs the generation of a predicted image of a block to be processed in one selected mode (S512).
[0396] The selection process (S511) in FIG. 23 can be considered as a process including the determination process (S501) in FIG. 22. For example, in the third mode, the circuit (160) determines that OBMC processing is not applicable to the generation process of the predicted image of the target block, as BIO processing is applied to the generation process of the predicted image of the target block. Then, the circuit (160) does not apply OBMC processing to the generation process of the predicted image of the target block and applies BIO processing.
[0397] That is, in the example of FIG. 23, just as in the example of FIG. 22, if BIO processing is applied to the generation process of the predicted image of the block to be processed, it is determined that OBMC processing cannot be applied to the generation process of the predicted image of the block to be processed. And, OBMC processing is not applied to the generation process of the predicted image of the block to be processed, and BIO processing is applied. In other words, the example of FIG. 23 includes the example of FIG. 22.
[0398] In the example of FIG. 23, the same effect as in the example of FIG. 22 can be obtained. In addition, through the operation shown in FIG. 23, the encoding device (100) can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes.
[0399] FIG. 24 is a flowchart showing additional operations for a variation of the first operation example shown in FIG. 23. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 24 in addition to the operation shown in FIG. 23.
[0400] Specifically, the circuit (160) of the encoding device (100) encodes a signal representing one value corresponding to one mode among three values, a first value, a second value, and a third value, each corresponding to a first mode, a second mode, and a third mode (S521). This one mode is one mode selected from the first mode, the second mode, and the third mode.
[0401] By this, the encoding device (100) can encode one mode that is adaptively selected from three modes. Accordingly, the encoding device (100) and the decoding device (200) can adaptively select the same mode from three modes.
[0402] Additionally, the encoding process (S521) may be performed before the selection process (S511), between the selection process (S511) and the generation process (S512), or after the generation process (S512). Furthermore, the encoding process (S521) may be performed in parallel with the selection process (S511) or the generation process (S512).
[0403] Also, the selection and encoding of the mode may be performed for each block, for each slice, for each picture, or for each sequence. Also, at least one of the first mode, the second mode, and the third mode may be subdivided into multiple modes.
[0404] FIG. 25 is a flowchart illustrating a second operation example of the encoding device (100) shown in FIG. 21. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 25 when encoding a video image using picture-to-picture prediction.
[0405] Specifically, the circuit (160) of the encoding device (100) determines whether OBMC processing is applicable to the generation of a predicted image of a processing target block based on whether pair prediction processing referencing two completed pictures is applied to the generation of a predicted image of a processing target block (S601).
[0406] Here, the block to be processed is a block included in the picture to be processed in the video. Also, OBMC processing is a process that corrects the predicted image of the block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of the blocks surrounding the block to be processed.
[0407] Then, the circuit (160) determines that OBMC processing is not applicable to the generation of the predicted image of the processing target block when paired prediction processing is applied to the generation of the predicted image of the processing target block. Then, the circuit (160) does not apply OBMC processing to the generation of the predicted image of the processing target block and applies paired prediction processing (S602).
[0408] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0409] In addition, when pair prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction image obtained from each of the two completed pictures, thereby significantly increasing the throughput. The encoding device (100) can suppress such an increase in throughput.
[0410] Therefore, the encoding device (100) can efficiently perform prediction processing. In addition, the circuit size can be reduced.
[0411] Additionally, the circuit (160) may also determine that OBMC processing is not applicable to the predictive processing of the block to be processed when an affine transformation is applied to the predictive processing of the block to be processed.
[0412] FIG. 26 is a flowchart showing a variation of the second operation example shown in FIG. 25. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 26 when encoding a video image using picture-to-picture prediction.
[0413] Specifically, the circuit (160) of the encoding device (100) selects one mode among the first mode, the second mode, and the third mode (S611). Here, the first mode is a mode in which paired prediction processing is not applied and OBMC processing is not applied to the generation processing of the predicted image of the block to be processed. The second mode is a mode in which OBMC processing is applied without applying paired prediction processing to the generation processing of the predicted image of the block to be processed. The third mode is a mode in which paired prediction processing is applied without applying OBMC processing to the generation processing of the predicted image of the block to be processed.
[0414] Then, the circuit (160) performs the generation of a predicted image of a block to be processed in one selected mode (S612).
[0415] The selection process (S611) in FIG. 26 can be considered as a process including the determination process (S601) in FIG. 25. For example, in the third mode, the circuit (160) determines that OBMC processing is not applicable to the generation process of the predicted image of the target block, as the pair prediction processing is applied to the generation process of the predicted image of the target block. Then, the circuit (160) does not apply OBMC processing to the generation process of the predicted image of the target block, but applies pair prediction processing.
[0416] That is, in the example of FIG. 26, just as in the example of FIG. 25, if paired prediction processing is applied to the generation of a predicted image of a block to be processed, it is determined that OBMC processing cannot be applied to the generation of a predicted image of a block to be processed. Then, OBMC processing is not applied to the generation of a predicted image of a block to be processed, and paired prediction processing is applied. In other words, the example of FIG. 26 includes the example of FIG. 25.
[0417] In the example of FIG. 26, the same effect as in the example of FIG. 25 can be obtained. Additionally, through the operation shown in FIG. 26, the encoding device (100) can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes.
[0418] FIG. 27 is a flowchart showing additional operations for a variation of the second operation example shown in FIG. 26. For example, the encoding device (100) shown in FIG. 21 may perform the operation shown in FIG. 27 in addition to the operation shown in FIG. 26.
[0419] Specifically, the circuit (160) of the encoding device (100) encodes a signal representing one value corresponding to one mode among three values, a first value, a second value, and a third value, each corresponding to a first mode, a second mode, and a third mode (S621). This one mode is one mode selected from the first mode, the second mode, and the third mode.
[0420] By this, the encoding device (100) can encode one mode that is adaptively selected from three modes. Accordingly, the encoding device (100) and the decoding device (200) can adaptively select the same mode from three modes.
[0421] Additionally, the encoding process (S621) may be performed before the selection process (S611), between the selection process (S611) and the generation process (S612), or after the generation process (S612). Furthermore, the encoding process (S621) may be performed in parallel with the selection process (S611) or the generation process (S612).
[0422] Also, the selection and encoding of the mode may be performed for each block, for each slice, for each picture, or for each sequence. Also, at least one of the first mode, the second mode, and the third mode may be subdivided into multiple modes.
[0423] Also, in the examples of FIGS. 25, 26, and 27, the pair prediction processing is a pair prediction processing that refers to two completed pictures. However, in these examples, the pair prediction processing may be limited to a pair prediction processing that refers to a completed picture that is displayed earlier than the picture to be processed and a completed picture that is displayed later than the picture to be processed.
[0424] [Example of implementation of a decoding device]
[0425] FIG. 28 is a block diagram showing an example of implementation of a decoding device (200) according to embodiment 1. The decoding device (200) is equipped with a circuit (260) and a memory (262). For example, a plurality of components of the decoding device (200) shown in FIG. 10 and FIG. 12 are implemented by the circuit (260) and memory (262) shown in FIG. 28.
[0426] The circuit (260) is a circuit that performs information processing and is capable of accessing memory (262). For example, the circuit (260) is a dedicated or general-purpose electronic circuit that decodes video images. The circuit (260) may be a processor such as a CPU. Also, the circuit (260) may be an assembly of multiple electronic circuits. Also, for example, the circuit (260) may serve as a component of the multiple components of the decoder (200) shown in FIG. 10, excluding the component for storing information.
[0427] The memory (262) is a dedicated or general-purpose memory in which information for the circuit (260) to decode a video image is stored. The memory (262) may be an electronic circuit and may be connected to the circuit (260). Also, the memory (262) may be included in the circuit (260). Also, the memory (262) may be an assembly of multiple electronic circuits. Also, the memory (262) may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium. Also, the memory (262) may be a non-volatile memory or a volatile memory.
[0428] For example, the memory (262) may store a bit sequence corresponding to an encoded video image, or a video image corresponding to a decoded bit sequence. Additionally, the memory (262) may store a program for the circuit (260) to decode the video image.
[0429] Also, for example, the memory (262) may serve as a component for storing information among the plurality of components of the decoder (200) shown in FIG. 10, etc. Specifically, the memory (262) may serve as a block memory (210) and frame memory (214) shown in FIG. 10. More specifically, the memory (262) may store a reconstructed block (processed block) and a reconstructed picture (processed picture), etc.
[0430] In addition, in the decoding device (200), all of the plurality of components shown in FIG. 10, etc., do not need to be implemented, and all of the aforementioned plurality of processes do not need to be performed. Some of the plurality of components shown in FIG. 10, etc., may be included in another device, and some of the aforementioned plurality of processes may be executed by another device. And, in the decoding device (200), some of the plurality of components shown in FIG. 10, etc., are implemented, and some of the aforementioned plurality of processes are performed, thereby allowing prediction processing to be performed efficiently.
[0431] FIG. 29 is a flowchart illustrating a first operation example of the decoding device (200) shown in FIG. 28. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 29 when decoding a video image using picture-to-picture prediction.
[0432] Specifically, the circuit (260) of the decoder (200) determines whether OBMC processing can be applied to the generation of a predicted image of a block to be processed, depending on whether BIO processing is applied to the generation of a predicted image of a block to be processed (S701). Here, the block to be processed is a block included in the picture to be processed in the video image.
[0433] Also, BIO processing is a process that generates a predicted image of a block to be processed by referencing the spatial gradient of luminance in an image obtained by performing motion compensation of the block to be processed using the motion vector of the block to be processed. Also, OBMC processing is a process that corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of a block surrounding the block to be processed.
[0434] Then, the circuit (260) determines that OBMC processing cannot be applied to the generation of a predicted image of a block to be processed when BIO processing is applied to the generation of a predicted image of a block to be processed. Then, the circuit (260) does not apply OBMC processing to the generation of a predicted image of a block to be processed and applies BIO processing (S702).
[0435] By doing so, it is possible to eliminate the pass of applying BIO processing and OBMC processing in the inter-picture prediction processing flow. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the inter-picture prediction processing flow.
[0436] Also, when BIO processing is applied, the degree to which the code amount is reduced by the application of OBMC processing is assumed to be lower compared to the case where BIO processing is not applied. That is, when it is assumed that the degree to which the code amount is reduced by the application of OBMC processing is low, the decoder (200) can efficiently generate a predicted image without applying OBMC processing.
[0437] Therefore, the decoding device (200) can efficiently perform prediction processing. In addition, the circuit size can be reduced.
[0438] Additionally, the circuit (260) may also determine that OBMC processing is not applicable to the predictive processing of the block to be processed when an affine transformation is applied to the predictive processing of the block to be processed.
[0439] FIG. 30 is a flowchart showing a variation of the first operation example shown in FIG. 29. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 30 when decoding a video image using picture-to-picture prediction.
[0440] Specifically, the circuit (260) of the decoder (200) selects one mode among the first mode, the second mode, and the third mode (S711). Here, the first mode is a mode in which BIO processing is not applied and OBMC processing is not applied to the generation processing of the predicted image of the block to be processed. The second mode is a mode in which OBMC processing is applied without BIO processing to the generation processing of the predicted image of the block to be processed. The third mode is a mode in which BIO processing is applied without OBMC processing to the generation processing of the predicted image of the block to be processed.
[0441] Then, the circuit (260) performs the generation of a predicted image of a block to be processed in one selected mode (S712).
[0442] The selection process (S711) in FIG. 30 can be considered as a process including the determination process (S701) in FIG. 29. For example, in the third mode, the circuit (260) determines that OBMC processing is not applicable to the generation process of the predicted image of the target block, as BIO processing is applied to the generation process of the predicted image of the target block. Then, the circuit (260) does not apply OBMC processing to the generation process of the predicted image of the target block, but applies BIO processing.
[0443] That is, in the example of FIG. 30, just as in the example of FIG. 29, if BIO processing is applied to the generation process of the predicted image of the block to be processed, it is determined that OBMC processing cannot be applied to the generation process of the predicted image of the block to be processed. And, OBMC processing is not applied to the generation process of the predicted image of the block to be processed, and BIO processing is applied. In other words, the example of FIG. 30 includes the example of FIG. 29.
[0444] In the example of FIG. 30, the same effect as in the example of FIG. 29 can be obtained. Additionally, through the operation shown in FIG. 30, the decoder (200) can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes.
[0445] FIG. 31 is a flowchart showing additional operations for a variation of the first operation example shown in FIG. 30. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 31 in addition to the operation shown in FIG. 30.
[0446] Specifically, the circuit (260) of the decoding device (200) decodes a signal representing one value corresponding to one mode among three values, a first value, a second value, and a third value, each corresponding to a first mode, a second mode, and a third mode, respectively (S721). This one mode is one mode selected from the first mode, the second mode, and the third mode.
[0447] By this, the decoder (200) can decode one mode that is adaptively selected from three modes. Thus, the encoding device (100) and the decoder (200) can adaptively select the same mode from three modes.
[0448] Additionally, the decoding process (S721) is basically performed before the selection process (S711). Specifically, the circuit (260) decodes a signal representing a value corresponding to a mode and selects a mode corresponding to the value represented by the decoded signal.
[0449] Also, the decoding and selection of the mode may be performed for each block, for each slice, for each picture, or for each sequence. Also, at least one of the first mode, second mode, and third mode may be subdivided by a plurality of modes.
[0450] FIG. 32 is a flowchart showing a second operation example of the decoding device (200) shown in FIG. 28. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 32 when decoding a video image using picture-to-picture prediction.
[0451] Specifically, the circuit (260) of the decoder (200) determines whether OBMC processing is applicable to the generation of a predicted image of a block to be processed, depending on whether a pair prediction processing that references two completed pictures is applied to the generation of a predicted image of a block to be processed (S801).
[0452] Here, the block to be processed is a block included in the picture to be processed in the video. Also, OBMC processing is a process that corrects the predicted image of the block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of the blocks surrounding the block to be processed.
[0453] Then, the circuit (260) determines that OBMC processing is not applicable to the generation of the predicted image of the processing target block when paired prediction processing is applied to the generation of the predicted image of the processing target block. Then, the circuit (260) does not apply OBMC processing to the generation of the predicted image of the processing target block and applies paired prediction processing (S802).
[0454] By this, in the picture-to-picture prediction processing flow, it is possible to eliminate the pass that applies paired prediction processing referencing two completed pictures and also applies OBMC processing. Therefore, it is possible to reduce the throughput of the pass with a large throughput in the picture-to-picture prediction processing flow.
[0455] In addition, when pair prediction processing referencing two completed pictures is applied and OBMC processing is applied, correction is performed on the prediction image obtained from each of the two completed pictures, thereby significantly increasing the throughput. The decoder (200) can suppress such an increase in throughput.
[0456] Therefore, the decoding device (200) can efficiently perform prediction processing. In addition, the circuit size can be reduced.
[0457] Additionally, the circuit (260) may also determine that OBMC processing is not applicable to the predictive processing of the block to be processed when an affine transformation is applied to the predictive processing of the block to be processed.
[0458] FIG. 33 is a flowchart showing a variation of the second operation example shown in FIG. 32. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 33 when decoding a video image using picture-to-picture prediction.
[0459] Specifically, the circuit (260) of the decoder (200) selects one mode among the first mode, the second mode, and the third mode (S811). Here, the first mode is a mode in which paired prediction processing is not applied and OBMC processing is not applied to the generation processing of the predicted image of the block to be processed. The second mode is a mode in which OBMC processing is applied without applying paired prediction processing to the generation processing of the predicted image of the block to be processed. The third mode is a mode in which paired prediction processing is applied without applying OBMC processing to the generation processing of the predicted image of the block to be processed.
[0460] Then, the circuit (260) performs the generation of a predicted image of a block to be processed in one selected mode (S812).
[0461] The selection process (S811) in FIG. 33 can be considered as a process including the determination process (S801) in FIG. 32. For example, in the third mode, the circuit (260) determines that OBMC processing is not applicable to the generation process of the predicted image of the target block, as the pair prediction processing is applied to the generation process of the predicted image of the target block. Then, the circuit (260) does not apply OBMC processing to the generation process of the predicted image of the target block, but applies pair prediction processing.
[0462] That is, in the example of FIG. 33, just as in the example of FIG. 32, if paired prediction processing is applied to the generation of a predicted image of a block to be processed, it is determined that OBMC processing cannot be applied to the generation of a predicted image of a block to be processed. And, OBMC processing is not applied to the generation of a predicted image of a block to be processed, and paired prediction processing is applied. In other words, the example of FIG. 33 includes the example of FIG. 32.
[0463] In the example of FIG. 33, the same effect as in the example of FIG. 32 can be obtained. Additionally, through the operation shown in FIG. 33, the decoder (200) can perform the generation processing of a predicted image in one mode that is adaptively selected from three modes.
[0464] FIG. 34 is a flowchart showing additional operations for a variation of the second operation example shown in FIG. 33. For example, the decoding device (200) shown in FIG. 28 may perform the operation shown in FIG. 34 in addition to the operation shown in FIG. 33.
[0465] Specifically, the circuit (260) of the decoding device (200) decodes a signal representing one value corresponding to one mode among three values, a first value, a second value, and a third value, each corresponding to a first mode, a second mode, and a third mode, respectively (S821). This one mode is one mode selected from the first mode, the second mode, and the third mode.
[0466] By this, the decoder (200) can decode one mode that is adaptively selected from three modes. Thus, the encoding device (100) and the decoder (200) can adaptively select the same mode from three modes.
[0467] Additionally, the decoding process (S821) is basically performed before the selection process (S811). Specifically, the circuit (260) decodes a signal representing a value corresponding to a mode and selects a mode corresponding to the value represented by the decoded signal.
[0468] Also, the decoding and selection of the mode may be performed for each block, for each slice, for each picture, or for each sequence. Also, at least one of the first mode, second mode, and third mode may be subdivided by a plurality of modes.
[0469] Also, in the examples of FIGS. 32, 33, and 34, the pair prediction processing is a pair prediction processing that refers to two completed pictures. However, in these examples, the pair prediction processing may be limited to a pair prediction processing that refers to a completed picture that is displayed earlier than the picture to be processed and a completed picture that is displayed later than the picture to be processed.
[0470] [supplement]
[0471] In addition, the encoding device (100) and the decoding device (200) in the present embodiment may each be used as an image encoding device and an image decoding device, or as a video encoding device and a video decoding device. Alternatively, the encoding device (100) and the decoding device (200) may each be used as an inter-prediction device (inter-frame prediction device).
[0472] That is, the encoding device (100) and the decoding device (200) may each correspond only to the inter-prediction unit (inter-frame prediction unit) (126) and the inter-prediction unit (inter-frame prediction unit) (218). In addition, other components such as the conversion unit (106) and the inverse conversion unit (206) may be included in other devices.
[0473] In addition, at least a portion of the present embodiment may be used as an encoding method, as a decoding method, as an inter-frame prediction method, or as other methods. For example, at least a portion of the present embodiment may be used as an inter-frame prediction method as follows.
[0474] For example, this inter-frame prediction method is an inter-frame prediction method for generating a prediction image, comprising a prediction image generation process that generates a first image using a motion vector assigned to a block to be processed, and a prediction image correction process that generates a correction image using a motion vector assigned to a block surrounding the block to be processed, and corrects the first image using the correction image to generate a second image, wherein the prediction image generation process is performed using one prediction image generation method selected from a plurality of prediction image generation methods including at least a first prediction image generation method and a second prediction image generation method, and when the first prediction image generation method is used, the first image is selected as the prediction image without performing the prediction image correction process, or the second image is selected as the prediction image by performing the prediction image correction process, and when the second prediction image generation method is used, the first image is selected as the prediction image without performing the prediction image correction process.
[0475] Alternatively, this inter-frame prediction method is an inter-frame prediction method for generating a prediction image, comprising a prediction image generation process that generates a first image using a motion vector assigned to a block to be processed, and a prediction image correction process that generates a correction image using a motion vector assigned to a block surrounding the block to be processed, and corrects the first image using the correction image to generate a second image, wherein the prediction image generation process is performed using one prediction image generation method selected from a plurality of prediction image generation methods including a first prediction image generation method and a second prediction image generation method, and selects which of the following processes is performed: (1) making the first image generated by the first prediction image generation method the prediction image, (2) making the second image generated by correcting the first image generated by the first prediction image generation method the prediction image, or (3) making the first image generated by the second prediction image generation method the prediction image.
[0476] In addition, in the present embodiment, each component may be realized by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be realized by a program execution unit, such as a CPU or processor, reading and executing a software program recorded on a recording medium, such as a hard disk or semiconductor memory.
[0477] Specifically, each of the encoding device (100) and the decoding device (200) may be equipped with a processing circuitry and a storage device electrically connected to the processing circuitry and accessible from the processing circuitry. For example, the processing circuitry corresponds to a circuitry (160 or 260), and the storage device corresponds to a memory (162 or 262).
[0478] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a memory device. Additionally, if the processing circuit includes a program execution unit, the memory device stores a software program executed by said program execution unit.
[0479] Here, the software for realizing the encoding device (100) or decoding device (200), etc. of the present embodiment is the following program.
[0480] That is, this program is an encoding method for encoding a moving image using picture-to-picture prediction in a computer, and determines whether OBMC (overlapped block motion compensation) processing is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, depending on whether BIO (bi-directional optical flow) processing is applied to the generation of a predicted image of a processing target block, and if the BIO processing is applied to the generation of a predicted image of the processing target block, determines that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and applies the BIO processing without applying the OBMC processing to the generation of a predicted image of the processing target block, wherein the BIO processing is a process of generating a predicted image of the processing target block by referencing the spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and the OBMC processing is an image obtained by performing motion compensation of the processing target block using the motion vector of a surrounding block of the processing target block By utilizing this, an encoding method that corrects the predicted image of the processing target block may be executed.
[0481] Alternatively, this program is a decoding method for decoding a moving image using picture-to-picture prediction on a computer, wherein it determines whether OBMC (overlapped block motion compensation) processing is applicable to the generation of a predicted image of a processing target block included in a processing target picture in the moving image, depending on whether BIO (bi-directional optical flow) processing is applied to the generation of a predicted image of a processing target block, and if the BIO processing is applied to the generation of a predicted image of the processing target block, it determines that the OBMC processing is not applicable to the generation of a predicted image of the processing target block, and thus applies the BIO processing without applying the OBMC processing to the generation of a predicted image of the processing target block, wherein the BIO processing is a process of generating a predicted image of the processing target block by referencing the spatial gradient of luminance in an image obtained by performing motion compensation of the processing target block using the motion vector of the processing target block, and the OBMC processing is an image obtained by performing motion compensation of the processing target block using the motion vector of a surrounding block of the processing target block By utilizing this, a decoding method, which is a process for correcting the predicted image of the block to be processed above, may be executed.
[0482] Alternatively, this program may execute an encoding method that encodes a video image using picture-to-picture prediction on a computer, and determines whether an overlapped block motion compensation (OBMC) processing can be applied to the generation of a predicted image of a processing block included in a processing target picture in the video image, depending on whether a pair prediction processing that references two completed processing pictures is applied to the generation of a predicted image of a processing target block, and if the pair prediction processing is applied to the generation of a predicted image of a processing target block, determines that the OBMC processing is not applicable to the generation of a predicted image of a processing target block, and applies the pair prediction processing without applying the OBMC processing to the generation of a predicted image of a processing target block, and the OBMC processing is a processing that corrects the predicted image of a processing target block using an image obtained by performing motion compensation of the processing target block using the motion vector of a block surrounding the processing target block.
[0483] Alternatively, this program may execute a decoding method for decoding a video image using picture-to-picture prediction on a computer, wherein, depending on whether a pair prediction process referencing two completed pictures is applied to the process of generating a predicted image of a processing block included in a processing target picture in the video image, it determines whether an OBMC (overlapped block motion compensation) process can be applied to the process of generating a predicted image of a processing target block, and if the pair prediction process is applied to the process of generating a predicted image of a processing target block, it determines that the OBMC process cannot be applied to the process of generating a predicted image of a processing target block, and applies the pair prediction process without applying the OBMC process to the process of generating a predicted image of a processing target block, and the OBMC process is a decoding method that corrects the predicted image of a processing target block using an image obtained by performing motion compensation of the processing target block using the motion vector of a block surrounding the processing target block.
[0484] Alternatively, this program may execute an encoding method that encodes a video image using inter-picture prediction on a computer, wherein, depending on whether a pair prediction process is applied to the generation of a predicted image of a block to be processed that is included in the picture to be processed in the video image, the program determines whether an OBMC (overlapped block motion compensation) process can be applied to the generation of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation of a predicted image of a block to be processed, the program determines whether the OBMC process is not applicable to the generation of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation of a predicted image of a block to be processed, the program determines that the OBMC process is not applicable to the generation of a predicted image of a block to be processed, and applies the pair prediction process without applying the OBMC process to the generation of a predicted image of a block to be processed, and the OBMC process is an encoding method that corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of a block surrounding the block to be processed.
[0485] Alternatively, this program may execute a decoding method for decoding a video image using inter-picture prediction on a computer, wherein, depending on whether a pair prediction process is applied to the generation of a predicted image of a block to be processed that is included in the picture to be processed in the video image, the program determines whether an OBMC (overlapped block motion compensation) process can be applied to the generation of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation of a predicted image of a block to be processed, the program determines whether the OBMC process is not applicable to the generation of a predicted image of a block to be processed, and if the pair prediction process is applied to the generation of a predicted image of a block to be processed, the program determines that the OBMC process is not applicable to the generation of a predicted image of a block to be processed, and applies the pair prediction process without applying the OBMC process to the generation of a predicted image of a block to be processed, and the OBMC process is a decoding method that corrects the predicted image of a block to be processed by using an image obtained by performing motion compensation of the block to be processed using the motion vector of a block surrounding the block to be processed.
[0486] In addition, each component may be a circuit as described above. These circuits may form a single circuit as a whole, or they may each be different circuits. In addition, each component may be realized as a general-purpose processor or as a dedicated processor.
[0487] Additionally, another component may execute the processing that a specific component executes. Also, the order in which processing is executed may be changed, and multiple processes may be executed in parallel. Additionally, the encoding and decoding device may be equipped with an encoding device (100) and a decoding device (200).
[0488] The first and second ordinal numbers used in the description may be replaced as appropriate. Additionally, ordinal numbers may be newly assigned to the components, etc., or removed.
[0489] For the above, the embodiments of the encoding device (100) and the decoding device (200) have been described based on the embodiments, but the embodiments of the encoding device (100) and the decoding device (200) are not limited to these embodiments. Various modifications that can be conceived by a person skilled in the art without departing from the spirit of the present disclosure may be implemented in the embodiments of this, or forms constructed by combining components from different embodiments may also be included within the scope of the embodiments of the encoding device (100) and the decoding device (200).
[0490] The present embodiment may be carried out in combination with at least some of other embodiments in the present disclosure. Additionally, some processing, some configuration of the apparatus, some syntax, etc., described in the flowchart of the present embodiment may be carried out in combination with other embodiments.
[0491] (Form of Embodiment 2)
[0492] In each of the above embodiments, each of the functional blocks can typically be realized by an MPU and memory, etc. In addition, processing by each of the functional blocks is typically realized by a program execution unit, such as a processor, reading and executing software (program) recorded on a recording medium such as ROM. The software may be distributed by downloading, etc., or may be distributed by recording it on a recording medium such as semiconductor memory. In addition, it is naturally possible for each functional block to be realized by hardware (dedicated circuit).
[0493] Furthermore, the processing described in each embodiment may be realized by centralized processing using a single device (system) or by distributed processing using multiple devices. Also, the processor executing the program may be a single one or multiple ones. That is, centralized processing may be performed or distributed processing may be performed.
[0494] The embodiments of the present disclosure are not limited to the above embodiments, and various modifications are possible, and such modifications are also included within the scope of the embodiments of the present disclosure.
[0495] In addition, an application example of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same are described. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Other configurations in the system may be appropriately changed depending on the circumstances.
[0496] [Examples of Use]
[0497] FIG. 35 is a diagram showing the overall configuration of a content supply system (ex100) that realizes a content transmission service. The communication service provision area is divided into desired sizes, and a fixed wireless station base station (ex106, ex107, ex108, ex109, ex110) is installed in each cell.
[0498] In this content supply system (ex100), each device, such as a computer (ex111), a game console (ex112), a camera (ex113), a home appliance (ex114), and a smartphone (ex115), is connected to the Internet (ex101) through an Internet service provider (ex102) or a communication network (ex104) and base stations (ex106 to ex110). The content supply system (ex100) may be connected by combining any one of the above elements. Each device may be connected directly or indirectly to one another through a telephone network or short-range wireless, etc., without using a fixed wireless base station (ex106 to ex110). Additionally, a streaming server (ex103) is connected to each device, such as a computer (ex111), a game console (ex112), a camera (ex113), a home appliance (ex114), and a smartphone (ex115), through the Internet (ex101), etc. Additionally, the streaming server (ex103) is connected to terminals, etc., within a hot spot inside the airplane (ex117) via a satellite (ex116).
[0499] In addition, instead of base stations (ex106~ex110), wireless access points or hot spots may be used. Also, the streaming server (ex103) may be connected directly to a communication network (ex104) without going through the Internet (ex101) or an Internet service provider (ex102), and may be connected directly to an airplane (ex117) without going through a satellite (ex116).
[0500] A camera (ex113) is a device capable of taking still images and videos, such as a digital camera. Also, a smartphone (ex115) is a smartphone, mobile phone, or PHS (Personal Handyphone System) that corresponds to the mobile communication system methods generally referred to as 2G, 3G, 3.9G, 4G, and in the future 5G.
[0501] Home appliances (ex114) include refrigerators, or devices included in home fuel cell cogeneration systems, etc.
[0502] In the content supply system (ex100), live transmission is made possible by connecting a terminal having a shooting function to a streaming server (ex103) through a base station (ex106), etc. In live transmission, the terminal (a computer (ex111), a game console (ex112), a camera (ex113), a home appliance (ex114), a smartphone (ex115), and a terminal inside an airplane (ex117), etc.) performs the encoding processing described in each of the above embodiments on a still image or video content captured by a user using said terminal, multiplexes the image data obtained by encoding with sound data encoded from sound corresponding to the image, and transmits the obtained data to the streaming server (ex103). That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.
[0503] Meanwhile, the streaming server (ex103) stream-transmits the transmitted content data to the client that made the request. The client is a terminal, such as a computer (ex111), a game console (ex112), a camera (ex113), a home appliance (ex114), a smartphone (ex115), or an airplane (ex117), capable of decoding the encoded data. Each device that receives the transmitted data decodes and plays back the received data. That is, each device functions as an image decoding device according to one aspect of the present disclosure.
[0504] [Distributed Processing]
[0505] Additionally, the streaming server (ex103) may be multiple servers or multiple computers, and may process, record, or transmit data in a distributed manner. For example, the streaming server (ex103) may be realized by a Contents Delivery Network (CDN), and content transmission may be realized through a network connecting multiple edge servers distributed worldwide. In a CDN, physically nearby edge servers are dynamically assigned based on the client. Furthermore, latency can be reduced by caching and transmitting content to the corresponding edge server. Moreover, in the event of an error or a change in communication status due to increased traffic, processing can be distributed across multiple edge servers, the transmission entity can be switched to another edge server, or transmission can continue by bypassing the affected part of the network, thereby enabling high-speed and stable transmission.
[0506] Furthermore, beyond the distributed processing of the transmission itself, the encoding processing of the captured data can be performed at each terminal, on the server side, or shared among them. For example, in general, the encoding process involves two processing loops. In the first loop, the complexity of the image or the amount of encoding at the frame or scene level is detected. In the second loop, processing is performed to improve encoding efficiency while maintaining image quality. For instance, by having the terminal perform the first encoding processing and the server side receiving the content perform the second encoding processing, the processing load on each terminal can be reduced while improving the quality and efficiency of the content. In this case, if there is a request to receive and decode in near real-time, the data completed in the first encoding process by the terminal can be received and played back by another terminal, thereby enabling more flexible real-time transmission.
[0507] As another example, a camera (ex113) extracts features from an image, compresses the data regarding the features as metadata, and transmits it to a server. The server performs compression based on the meaning of the image, for example, by determining the importance of an object from the features and adjusting the quantization precision. Feature data is particularly effective for improving the accuracy and efficiency of motion vector prediction during further compression on the server. In addition, simple encoding such as VLC (Variable Length Coding) may be performed at the terminal, while encoding with a high processing load such as CABAC (Context-Adaptive Dual Arithmetic Coding) may be performed on the server.
[0508] As another example, in stadiums, shopping malls, or factories, there may be multiple video data sets in which nearly identical scenes are captured by multiple terminals. In such cases, distributed processing is performed by utilizing the multiple terminals that captured the footage, as well as other terminals and servers that are not capturing footage as needed, by assigning encoding processing to each, for instance, at the Group of Picture (GOP) level, the picture level, or the tile level by which the picture is divided. This reduces latency and enables greater real-time performance.
[0509] In addition, since multiple video data are nearly identical scenes, the server may manage and / or instruct the video data captured by each terminal so that they can reference each other. Alternatively, the server may receive the encoded data from each terminal, change the reference relationships between the multiple data, or correct or replace the picture itself and re-encode it. By doing so, a stream with improved quality and efficiency for each individual data can be generated.
[0510] In addition, the server may transmit the video data after performing transcoding that changes the encoding method of the video data. For example, the server may convert an MPEG-based encoding method to a VP-based one, or convert H.264 to H.265.
[0511] In this way, encoding processing can be performed by a terminal or one or more servers. Accordingly, in the following description, devices such as a "server" or a "terminal" are used as the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. Furthermore, the same applies to decoding processing.
[0512] [3D, Multi-angle]
[0513] Recently, there has been an increase in the use of integrated images or videos of different scenes captured by multiple cameras (ex113) and / or smartphones (ex115) and / or the same scene captured from different angles. Images captured by each terminal are integrated based on the relative positional relationship between terminals acquired separately, or areas where feature points included in the images coincide.
[0514] The server may not only encode two-dimensional moving images but also encode still images automatically based on the analysis of the moving images, or at a time specified by the user, and transmit them to the receiving terminal. In addition, if the server can acquire the relative positional relationship between the capturing terminals, it may generate a three-dimensional shape of the scene based not only on two-dimensional moving images but also on images of the same scene captured from different angles. Furthermore, the server may separately encode three-dimensional data generated by point clouds, etc., or generate an image to be transmitted to the receiving terminal by selecting or reconstructing from images captured by multiple terminals based on the result of recognizing or tracking a person or object using the three-dimensional data.
[0515] In this way, the user may arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or may enjoy content cut from a video of an arbitrary viewpoint from 3D data reconstructed using multiple images or videos. In addition, just like video, sound may also be received from multiple different angles, and the server may multiplex sound from a specific angle or space with the video to match the video and transmit it.
[0516] In addition, recently, content that bridges the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also been distributed. In the case of VR images, the server may create viewpoint images for the right eye and the left eye separately, and perform encoding that allows referencing between the viewpoint images using Multi-View Coding (MVC), or it may encode them as separate streams without referencing each other. When decoding the separate streams, they can be played back in synchronization so that a virtual 3D space is reproduced according to the user's viewpoint.
[0517] In the case of AR images, the server superimposes virtual object information in virtual space onto camera information in real space based on 3D position or the movement of the user's viewpoint. The decoding device may acquire or maintain virtual object information and 3D data, generate a 2D image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to the virtual object information, and the server may create superimposed data based on the received viewpoint movement from the 3D data maintained on the server, encode the superimposed data, and transmit it to the decoding device. Furthermore, the superimposed data may have an α value representing transmittance in addition to RGB, and the server may set the α value of parts other than the object created from the 3D data to 0, etc., and encode the data in a state where said parts are transparent. Alternatively, the server may generate data in which a predetermined RGB value is set as the background, such as in chroma keying, and parts other than the object are made into a background color.
[0518] Similarly, the decoding of transmitted data can be performed at each client terminal, on the server side, or shared among them. For example, one terminal may first send a reception request to the server, another terminal may receive the content corresponding to that request and perform decoding, and then a decoding completion signal may be transmitted to a device equipped with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-enabled terminals themselves, high-quality data can be reproduced. As another example, while receiving large-sized image data on a TV or similar device, only a portion of the image, such as divided tiles, may be decoded and displayed on the viewer's personal terminal. This allows the entire image to be shared, enabling users to view their assigned area or specific areas they wish to examine in greater detail up close.
[0519] Furthermore, in the future, under conditions where multiple short-range, medium-range, or long-range wireless communications are available regardless of whether they are indoors or outdoors, it is expected that content will be received seamlessly while switching appropriate data for the connected communication using transmission system standards such as MPEG-DASH. This allows the user to freely select and switch in real time not only to their own terminal but also to decoding devices or display devices, such as displays installed indoors or outdoors. Additionally, decoding can be performed while switching between the decoding terminal and the display terminal based on the user's location information. This makes it possible to move to a destination while displaying map information on a wall of a nearby building or part of the ground where a display-capable device is embedded. Furthermore, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data over the network, such as when the encoded data is cached on a server that can be accessed in a short time from the receiving terminal, or when it is copied to an edge server for content delivery services.
[0520] [Scalable Encoding]
[0521] Regarding the switching of content, the method described above applies the video encoding method shown in FIG. 36 and utilizes a compressed and encoded scalable stream. The server may have multiple streams of different quality that have the same content as individual streams, but may be configured to switch content by utilizing the characteristics of the temporal and spatial scalable streams realized by dividing the encoding into layers as shown. That is, by determining which layer to decode based on internal factors such as performance and external factors such as communication bandwidth conditions, the decoding side can freely switch between decoding low-resolution content and high-resolution content. For example, if you want to continue watching a video that was being viewed on a smartphone (ex115) while on the move, and then watch it on a device such as an internet TV after returning home, the device can decode the same stream up to a different layer, thereby reducing the burden on the server side.
[0522] In addition, as described above, in addition to a configuration that realizes scalability in which a picture is encoded for each layer and an enhancement layer exists above the base layer, the enhancement layer may include metadata based on statistical information of the image, and the decoding side may generate high-quality content by super-resolution of the picture of the base layer based on the metadata. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The metadata includes information for specifying linear or non-linear filter coefficients used for super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operations used for super-resolution processing.
[0523] Alternatively, the picture may be divided into tiles based on the meaning of objects within the image, and the decoding side may be configured to decode only a portion of the area by selecting the tile to be decoded. Furthermore, by storing the attributes of an object (person, car, ball, etc.) and its position within the image (coordinate position within the same image, etc.) as metadata, the decoding side can identify the location of a desired object based on the metadata and determine the tile containing that object. For example, as shown in FIG. 37, the metadata is stored using a data storage structure different from pixel data, such as SEI messages in HEVC. This metadata indicates, for example, the location, size, or color of a main object.
[0524] In addition, metadata may be stored in units composed of multiple pictures, such as streams, sequences, or random access units. By doing so, the decoding side can obtain the time at which a specific person appears in the video, and by combining this with picture-unit information, can identify the picture in which the object exists and the location of the object within the picture.
[0525] [Web Page Optimization]
[0526] FIG. 38 is a diagram showing an example of a display screen of a web page in a computer (ex111), etc. FIG. 39 is a diagram showing an example of a display screen of a web page in a smartphone (ex115), etc. As shown in FIG. 38 and FIG. 39, a web page may include multiple link images that are links to image content, and the method of display may differ depending on the viewing device. When multiple link images are displayed on the screen, until the user explicitly selects a link image, or until the link image is near the center of the screen, or until the entire link image is contained within the screen, the display device (decoding device) displays a still image or I-picture of each content as a link image, displays an image such as a GIF animation using multiple still images or I-pictures, or receives only the base layer to decode and display the image.
[0527] When a linked image is selected by the user, the display device decodes with the base layer as the highest priority. Additionally, if the HTML constituting the web page contains information indicating that it is scalable content, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, if selection occurs beforehand or when the communication bandwidth is very poor, the display device may decode and display only the forward reference picture (I-picture, P-picture, and B-picture that is a forward reference only) to reduce the delay between the decoding time and display time of the leading picture (the delay from the start of content decoding to the start of display). Alternatively, the display device may intentionally ignore the picture reference relationships and coarsely decode all B-pictures and P-pictures as forward references, and then perform normal decoding as the number of received pictures increases over time.
[0528] [Autonomous Driving]
[0529] In addition, when transmitting and receiving still images or video data, such as 2D or 3D map information, for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information, etc., as metadata in addition to image data belonging to one or more layers, and decode them by corresponding. In addition, the metadata may belong to a layer or may simply be multiplexed with the image data.
[0530] In this case, since a vehicle, drone, or airplane including the receiving terminal is moving, the receiving terminal can achieve seamless reception and decoding while switching between base stations (ex106~ex110) by transmitting the location information of the receiving terminal upon a reception request. Additionally, the receiving terminal can dynamically switch how much metadata to receive or how much map information to update depending on the user's selection, the user's situation, or the communication band status.
[0531] In this way, the content supply system (ex100) can receive encoded information transmitted by a user in real time, decode it, and play it back.
[0532] [Transmission of Personal Content]
[0533] In addition, the content supply system (ex100) enables unicast or multicast transmission of not only high-quality, long-duration content provided by video transmission providers, but also low-quality, short-duration content provided by individuals. Furthermore, it can be assumed that such individual content will continue to increase in the future. In order to improve the quality of individual content, the server may perform encoding processing after performing editing processing. This can be realized, for example, with a configuration as follows.
[0534] At the time of shooting, or after shooting accumulated data, the server performs recognition processing such as identifying shooting errors, scene search, semantic interpretation, and object detection from the original image or the encoded data. Then, based on the recognition results, the server performs editing, either manually or automatically, such as correcting out-of-focus or camera shake, deleting scenes of low importance (e.g., scenes with lower brightness compared to other pictures or out of focus), emphasizing object edges, or changing color tones. Based on the editing results, the server encodes the edited data. Furthermore, since it is known that viewership ratings decrease if the shooting time is too long, the server may automatically clip scenes of low importance as described above, as well as scenes with little movement, based on the image processing results, so that the content falls within a specific time range according to the shooting time. Alternatively, the server may generate and encode a digest based on the results of the scene semantic interpretation.
[0535] Furthermore, personal content may contain elements that, if left as is, infringe upon copyright, moral rights, or portrait rights, or the scope of sharing may exceed the intended limit, causing inconvenience to the individual. Therefore, for example, the server may intentionally blur the faces of people or the interior of a house in the periphery of the screen before encoding. Additionally, the server may recognize whether the face of a person different from a previously registered individual is captured within the image to be encoded, and if so, perform processing such as applying a mosaic effect to the face. Alternatively, as a pre- or post-processing step for encoding, the user may specify the person or background area to be processed from a copyright perspective, and the server may replace the designated area with another image or perform processing such as blurring. In the case of a person, the image of the face can be replaced while tracking the person within the video.
[0536] In addition, since the viewing of personal content with a small amount of data requires real-time performance, the decoder prioritizes receiving the base layer first to perform decoding and playback, although this varies depending on the bandwidth. During this process, the decoder receives the enhancement layer, and in cases where playback occurs two or more times, such as in a loop, it may include the enhancement layer to play high-quality video. If the stream is scalable encoded in this way, the video may appear rough when not selected or at the stage where viewing begins, but the stream gradually becomes smarter, providing an experience where the image quality improves. In addition to scalable encoding, the same experience can be provided even if the rough stream played once and the second stream encoded by referencing the video from the first play are configured as a single stream.
[0537] [Other examples of use]
[0538] In addition, these encoding or decoding processes are generally performed on the LSI (ex500) of each terminal. The LSI (ex500) may be a single chip or a configuration consisting of multiple chips. Furthermore, software for video encoding or decoding may be mounted on any recording medium (CD-ROM, flexible disk, or hard disk, etc.) readable by a computer (ex111), and encoding or decoding processing may be performed using said software. In addition, if the smartphone (ex115) is equipped with a camera, video data acquired by said camera may be transmitted. At this time, the video data is data encoded by the LSI (ex500) of the smartphone (ex115).
[0539] Additionally, the LSI (ex500) may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the capability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the capability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.
[0540] In addition, the content supply system (ex100) via the Internet (ex101) is not limited to digital broadcasting systems, and at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of each of the above embodiments may be equipped in digital broadcasting systems. Since multiplexed data in which video and sound are multiplexed is carried on broadcast radio waves using a satellite or the like and transmitted and received, there is a difference in that it is suitable for multicast rather than the configuration of the content supply system (ex100) which is easy to unicast, but the same application is possible regarding encoding processing and decoding processing.
[0541] [Hardware Configuration]
[0542] FIG. 40 is a diagram showing a smartphone (ex115). FIG. 41 is a diagram showing an example configuration of a smartphone (ex115). The smartphone (ex115) is equipped with an antenna (ex450) for transmitting and receiving radio waves between a base station (ex110), a camera unit (ex465) capable of taking images and still images, and a display unit (ex458) for displaying decoded data such as images captured by the camera unit (ex465) and images received by the antenna (ex450). The smartphone (ex115) also comprises an operating unit (ex466), such as a touch panel; a voice output unit (ex457), such as a speaker for outputting voice or sound; a voice input unit (ex456), such as a microphone for inputting voice; a memory unit (ex467) capable of storing encoded data, such as captured video or still images, recorded voice, received video or still images, or email, or decoded data; and a slot unit (ex464) which is an interface unit with a SIM (ex468) for identifying a user and authenticating access to various data, including a network. Additionally, an external memory may be used instead of the memory unit (ex467).
[0543] In addition, a main control unit (ex460) that comprehensively controls the display unit (ex458) and the operation unit (ex466), and a power circuit unit (ex461), an operation input control unit (ex462), an image signal processing unit (ex455), a camera interface unit (ex463), a display control unit (ex459), a modulation / demodulation unit (ex452), a multiplication / separation unit (ex453), an audio signal processing unit (ex454), a slot unit (ex464), and a memory unit (ex467) are connected via a bus (ex470).
[0544] The power circuit (ex461) starts the smartphone (ex115) into an operable state by supplying power from the battery pack to each part when the power key is turned on by user operation.
[0545] A smartphone (ex115) performs processing such as calls and data communication based on the control of a main control unit (ex460) having a CPU, ROM, and RAM. During a call, a voice signal received from a voice input unit (ex456) is converted into a digital voice signal by a voice signal processing unit (ex454), the signal is subjected to spectrum spreading processing by a modulation / demodulation unit (ex452), and after performing digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit (ex451), it is transmitted through an antenna (ex450). Additionally, the received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, spectrum despreading processing by a modulation / demodulation unit (ex452), and after being converted into an analog voice signal by a voice signal processing unit (ex454), it is output from a voice output unit (ex457). In data communication mode, text, still images, or video data are transmitted to the main control unit (ex460) through the operation input control unit (ex462) by operation of the main unit's operation unit (ex466), and transmission and reception processing is performed in the same manner. When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit (ex455) compresses and encodes the video signal stored in the memory unit (ex467) or the video signal input from the camera unit (ex465) using the video encoding method shown in each embodiment above, and transmits the encoded video data to the multiplication / separation unit (ex453). In addition, the audio signal processing unit (ex454) encodes the audio signal received from the audio input unit (ex456) while the video or still images are being captured by the camera unit (ex465), and transmits the encoded audio data to the multiplication / separation unit (ex453). The multiplexing / separation unit (ex453) multiplexes the encoded video data and encoded voice data in a predetermined manner, performs modulation processing and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) (ex452) and the transmission / reception unit (ex451), and transmits through the antenna (ex450).
[0546] When a video attached to an email or chat, or a video linked to a webpage, etc., is received, in order to decode the multiplexed data received through the antenna (ex450), the multiplexing / separation unit (ex453) separates the multiplexed data to divide the multiplexed data into a bit stream of video data and a bit stream of voice data, and supplies the encoded video data to the video signal processing unit (ex455) through the synchronization bus (ex470), and at the same time supplies the encoded voice data to the voice signal processing unit (ex454). The video signal processing unit (ex455) decodes the video signal by a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and a video or still image included in the linked video file is displayed from the display unit (ex458) through the display control unit (ex459). In addition, the voice signal processing unit (ex454) decodes the voice signal, and voice is output from the voice output unit (ex457). Furthermore, since real-time streaming is widespread, there may be cases where audio playback is socially inappropriate depending on the user's situation. Therefore, as an initial setting, it is desirable to configure the system to play only video data without audio signals. Audio may be synchronized and played only when the user performs an operation, such as clicking on the video data.
[0547] Although a smartphone (ex115) was used as an example here, three types of implementations can be considered for terminals: a transmitting / receiving type terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In addition, regarding a digital broadcasting system, it was explained that multiplexed data, such as voice data, is multiplexed in video data and transmitted, but in addition to voice data, text data related to video may be multiplexed in the multiplexed data, or the video data itself may be received or transmitted instead of multiplexed data.
[0548] Furthermore, although it was explained that the main control unit (ex460), which includes the CPU, controls encoding or decoding processing, terminals are often equipped with GPUs. Therefore, a configuration can be used to process large areas in batches by leveraging the GPU's performance through memory shared between the CPU and GPU, or memory with addresses managed for common use. This allows for reduced encoding time, ensures real-time performance, and enables low latency. In particular, it is efficient to perform motion detection, deblocking filters, SAO (Sample Adaptive Offset), and transformation / quantization processing in batches on the GPU, such as in picture units, rather than on the CPU. Industrial applicability
[0549] The present disclosure may be used, for example, in a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a TV conferencing system, or an electronic mirror, etc. Explanation of the symbols
[0550] 100: Encoding unit 102: Dividing unit 104: Subtraction unit 106: Conversion unit 108: Quantizer 110: Entropy Encoding Unit 112, 204: Inverse quantization unit 114, 206: Inverse transformation unit 116, 208: Adder 118, 210: Block memory 120, 212: Loop filter section 122, 214: Frame memory 124, 216: Intra-prediction unit (in-screen prediction unit) 126, 218: Inter-prediction unit (Inter-frame prediction unit) 128, 220: Prediction control unit 160, 260: Circuit 162, 262: Memory 200: Decoding device 202: Entropy decoder
Claims
Claim 1 A decoding device comprising a memory and a circuit accessible to the memory, wherein the circuit derives motion vectors of the current block for a current block; wherein the circuit determines whether a first processing, which is a bidirectional optical flow processing referencing a spatial gradient of luminance, is applied and whether a second processing is applied, such that at least one of the first processing and the second processing is determined not to be applied to the current block; wherein, in a first state in which it is determined that the first processing is applied and the second processing is not applied, predicted images of the current block are generated based on reference pictures and the motion vectors; wherein a final predicted image is generated based on the predicted images and the spatial gradient of luminance, wherein, in a second state in which it is determined that the first processing is not applied and the second processing is applied, predicted images of the current block are generated based on reference pictures and the motion vectors; and wherein the predicted images are weighted to generate a final predicted image. Claim 2 A encoding device comprising a memory and a circuit accessible to the memory, wherein the circuit derives motion vectors of the current block for a current block; wherein the circuit determines whether a first processing, which is a bidirectional optical flow processing referencing a spatial gradient of luminance, is applied and whether a second processing is applied, such that at least one of the first processing and the second processing is determined not to be applied to the current block; wherein, in a first state in which it is determined that the first processing is applied and the second processing is not applied, predicted images of the current block are generated based on reference pictures and the motion vectors; wherein a final predicted image is generated based on the predicted images and the spatial gradient of luminance, wherein, in a second state in which it is determined that the first processing is not applied and the second processing is applied, predicted images of the current block are generated based on reference pictures and the motion vectors; and wherein the predicted images are weighted to generate a final predicted image. Claim 3 A decoding device comprising a memory and a circuit capable of accessing the memory, wherein in a process of decoding a first block of a picture, the circuit performs selecting a first process or a second process which is a bidirectional optical flow process, when the first process is selected, generating a first predicted image of the first block based on first images generated from each reference image and a spatial gradient of luminance, and when the second process is selected, generating a second predicted image of the first block by weighting second images generated from each reference image. Claim 4 A encoding device comprising a memory and a circuit capable of accessing the memory, wherein in a process of decoding a first block of a picture, the circuit performs selecting a first process or a second process which is a bidirectional optical flow process, when the first process is selected, generating a first predicted image of the first block based on first images generated from each reference image and a spatial gradient of luminance, and when the second process is selected, generating a second predicted image of the first block by weighting second images generated from each reference image.