Encoding device, decoding device, encoding method, and decoding method

By encoding the images generated by filtering processing in the video encoding device, the problem of improving the coding efficiency and image quality in the prior art is solved, and the processing volume and circuit scale are reduced and the processing speed is improved.

CN114175639BActive Publication Date: 2025-06-13PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080053915.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-07
Filing Date
2020-08-05
Publication Date
2025-06-13
Estimated Expiration
2040-08-05

AI Technical Summary

Technical Problem

When processing moving images, existing video encoding technologies are difficult to simultaneously improve encoding efficiency, picture quality, processing volume reduction, circuit scale reduction and processing speed.

Method used

By filtering the reconstructed samples of the first image using the filtering process in the encoding device, a second image is generated, and whether to encode the first image or the second image is determined based on the predetermined value of the parameter.

Benefits of technology

Improve coding efficiency, image quality, processing volume, circuit scale and processing speed are achieved, and encoding elements such as filters, block size, and motion vector are appropriately selected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175639B_ABST
    Figure CN114175639B_ABST
Patent Text Reader

Abstract

The encoding device (100) includes a circuit (a1) and a memory (a2) connected to the circuit (a1). During operation, the circuit (a1) encodes information for deriving a parameter into the header of a bitstream, generates a second image by filtering the reconstructed samples of the first image using filtering processing (S102), determines whether the parameter is a predetermined value (S103), encodes the third image using the second image when the parameter is the predetermined value (S104), and encodes the third image using the first image when the parameter is not the predetermined value (S105).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video coding, for example, systems, components, and methods in the encoding and decoding of moving images, etc. Background Art

[0002] Video coding technologies have advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to handle the continuously increasing amount of digital video data in various applications, there has always been a need to provide improvements and optimizations to video coding technologies.

[0003] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above-mentioned video coding technology.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] Regarding the encoding methods as described above, for the improvement of encoding efficiency, the improvement of image quality, the reduction of processing volume, the reduction of circuit scale, or the appropriate selection of elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, etc., it is desired to propose new methods.

[0009] The present invention provides a structure or method that can contribute to one or more of, for example, the improvement of encoding efficiency, the improvement of image quality, the reduction of processing volume, the reduction of circuit scale, the improvement of processing speed, and the appropriate selection of elements or operations. In addition, the present invention may include a structure or method that can contribute to benefits other than the above.

[0010] Means for Solving the Problems

[0011] For example, an encoding apparatus according to one embodiment of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit encodes information for deriving a parameter into the header of a bitstream, generates a second image by filtering reconstructed samples of a first image using a filtering process, determines whether the parameter is a predetermined value, encodes a third image using the second image when the parameter is the predetermined value, and encodes the third image using the first image when the parameter is not the predetermined value.

[0012] The installation of several embodiments of the present invention can improve the encoding efficiency, simplify the encoding / decoding process, speed up the encoding / decoding process, and efficiently select appropriate components / actions used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.

[0013] Based on the description and the drawings, further advantages and effects in one embodiment of the present invention are clarified. These advantages and / or effects are obtained by several embodiments and the features described in the description and the drawings respectively, but it is not necessary to provide all of them in order to obtain one or more of the advantages and / or effects.

[0014] In addition, these general or specific embodiments can also be implemented by a system, a method, an integrated circuit, a computer program, a recording medium, or any combination thereof.

[0015] Effect of the Invention

[0016] The structure or method of one embodiment of the present invention can contribute to one or more of, for example, improvement of encoding efficiency, improvement of image quality, reduction of processing amount, reduction of circuit scale, improvement of processing speed, and appropriate selection of elements or actions. In addition, the structure or method of one embodiment of the present invention can also contribute to benefits other than the above. Description of the Drawings

[0017] Figure 1 It is a block diagram showing the functional structure of an encoding apparatus according to an embodiment.

[0018] Figure 2 It is a flowchart showing an example of the overall encoding process performed by the encoding apparatus.

[0019] Figure 3 It is a conceptual diagram showing an example of block division.

[0020] Figure 4A It is a conceptual diagram showing an example of the structure of a slice.

[0021] Figure 4B It is a conceptual diagram showing an example of the structure of a tile.

[0022] Figure 5A It is a table showing transform basis functions corresponding to various transform types.

[0023] Figure 5B It is a conceptual diagram showing an example of SVT (Spatially Varying Transform).

[0024] Figure 6A It is a conceptual diagram showing an example of the shape of the filter used in ALF (adaptive loop filter).

[0025] Figure 6B It is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0026] Figure 6C It is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0027] Figure 7 It is a block diagram showing an example of the detailed structure of the loop filtering section that functions as a DBF (deblocking filter).

[0028] Figure 8 It is a conceptual diagram showing an example of a deblocking filter having filter characteristics symmetric with respect to the block boundary.

[0029] Figure 9 It is a conceptual diagram for explaining the block boundary where deblocking filtering processing is performed.

[0030] Figure 10 It is a conceptual diagram showing an example of the Bs value.

[0031] Figure 11 It is a flowchart showing an example of the processing performed by the prediction processing section of the encoding device.

[0032] Figure 12 It is a flowchart showing another example of the processing performed by the prediction processing section of the encoding device.

[0033] Figure 13 It is a flowchart showing another example of the processing performed by the prediction processing section of the encoding device.

[0034] Figure 14 It is a conceptual diagram showing an example of 67 intra prediction modes in the intra prediction of the embodiment.

[0035] Figure 15 It is a flowchart showing an example of the process of the basic inter prediction.

[0036] Figure 16 This is a flowchart showing an example of motion vector derivation.

[0037] Figure 17 This is a flowchart showing another example of motion vector derivation.

[0038] Figure 18 This is a flowchart showing another example of motion vector derivation.

[0039] Figure 19 This is a flowchart showing an example of inter - frame prediction based on a normal inter - frame mode.

[0040] Figure 20 This is a flowchart showing an example of inter - frame prediction based on a merge mode.

[0041] Figure 21 This is a conceptual diagram for explaining an example of the motion vector derivation process based on the merge mode.

[0042] Figure 22 This is a flowchart showing an example of FRUC (frame rate up conversion) processing.

[0043] Figure 23 This is a conceptual diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0044] Figure 24 This is a conceptual diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within a reference picture.

[0045] Figure 25A This is a conceptual diagram for explaining an example of the derivation of a motion vector in sub - block units based on motion vectors of multiple adjacent blocks.

[0046] Figure 25B This is a conceptual diagram for explaining an example of the derivation of a motion vector in sub - block units in an affine mode with three control points.

[0047] Figure 26A This is a conceptual diagram for explaining the affine merge mode.

[0048] Figure 26B This is a conceptual diagram for explaining the affine merge mode with two control points.

[0049] Figure 26C This is a conceptual diagram for explaining the affine merge mode with three control points.

[0050] Figure 27 This is a flowchart showing an example of the processing of the affine merge mode.

[0051] Figure 28A It is a conceptual diagram for explaining the affine inter-frame mode with 2 control points.

[0052] Figure 28B It is a conceptual diagram for explaining the affine inter-frame mode with 3 control points.

[0053] Figure 29 It is a flowchart showing an example of the processing of the affine inter-frame mode.

[0054] Figure 30A It is a conceptual diagram for explaining the affine inter-frame mode in which the current block has 3 control points and the adjacent block has 2 control points.

[0055] Figure 30B It is a conceptual diagram for explaining the affine inter-frame mode in which the current block has 2 control points and the adjacent block has 3 control points.

[0056] Figure 31A It is a flowchart showing a merge mode including DMVR (decoder motion vector refinement).

[0057] Figure 31B It is a conceptual diagram showing an example of the DMVR processing.

[0058] Figure 32 It is a flowchart showing an example of the generation of a predicted image.

[0059] Figure 33 It is a flowchart showing another example of the generation of a predicted image.

[0060] Figure 34 It is a flowchart showing another example of the generation of a predicted image.

[0061] Figure 35 It is a flowchart showing an example of the predicted image correction processing based on the OBMC (overlapped block motion compensation) processing.

[0062] Figure 36 It is a conceptual diagram showing an example of the predicted image correction processing based on the OBMC processing.

[0063] Figure 37 It is a conceptual diagram for explaining the generation of a predicted image of 2 triangles.

[0064] Figure 38 It is a conceptual diagram for explaining a model assuming uniform linear motion.

[0065] Figure 39This is a conceptual diagram showing an example of a method for generating a predicted image that uses luminance correction processing with LIC (local illumination compensation).

[0066] Figure 40 This is a block diagram showing an installation example of an encoding device.

[0067] Figure 41 This is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0068] Figure 42 This is a flowchart showing an example of the overall decoding process performed by a decoding device.

[0069] Figure 43 This is a flowchart showing an example of the process performed by the prediction processing unit of a decoding device.

[0070] Figure 44 This is a flowchart showing another example of the process performed by the prediction processing unit of a decoding device.

[0071] Figure 45 This is a flowchart showing an example of inter-frame prediction based on a normal inter-frame mode in a decoding device.

[0072] Figure 46 This is a block diagram showing an installation example of a decoding device.

[0073] Figure 47 This is a flowchart of the decoding process of the first form.

[0074] Figure 48 This is a diagram showing an example of the structure of a decoding device.

[0075] Figure 49 This is a diagram showing an example of the structure of a decoding device.

[0076] Figure 50 This is a diagram showing an example of the structure of a decoding device.

[0077] Figure 51 This is a diagram showing an example of the structure of a decoding device.

[0078] Figure 52 This is a diagram showing an example of the structure of a decoding device.

[0079] Figure 53 This is a diagram showing an example of the structure of a decoding device.

[0080] Figure 54 This is a diagram showing an example of the structure of a decoding device.

[0081] Figure 55 This is a diagram showing an example of the structure of a decoding device.

[0082] Figure 56 It is a flowchart of the decoding process in the second form.

[0083] Figure 57 It is a flowchart of the decoding process in the third form.

[0084] Figure 58 It is a flowchart of the decoding process in the fourth form.

[0085] Figure 59 It is a diagram showing a calculation example of gradient information.

[0086] Figure 60 It is a diagram showing a calculation example of gradient information.

[0087] Figure 61 It is a diagram showing a calculation example of gradient information.

[0088] Figure 62 It is a diagram showing a calculation example of gradient information.

[0089] Figure 63 It is a block diagram showing the overall structure of a content supply system that implements a content distribution service.

[0090] Figure 64 It is a conceptual diagram showing an example of an encoding structure in scalable coding.

[0091] Figure 65 It is a conceptual diagram showing an example of an encoding structure in scalable coding.

[0092] Figure 66 It is a conceptual diagram showing an example of a display screen of a web page.

[0093] Figure 67 It is a conceptual diagram showing an example of a display screen of a web page.

[0094] Figure 68 It is a block diagram showing an example of a smart phone.

[0095] Figure 69 It is a block diagram showing an example of the structure of a smart phone. Detailed implementation mode

[0096] An encoding device according to an aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit encodes information for deriving parameters into the header of a bitstream, generates a second image by filtering the reconstructed samples of a first image using filtering processing, determines whether the parameter is a predetermined value, and encodes a third image using the second image when the parameter is the predetermined value, and encodes the third image using the first image when the parameter is not the predetermined value.

[0097] Accordingly, the encoding device can switch which one of the second image after the application of the filtering process and the first image before the application is used as a reference for subsequent pictures. Therefore, in a case where the filtering process is effective in improving subjective image quality but not suitable as a reference image, by using the first image before the application of the filtering process as the reference image, it is possible to improve subjective image quality while suppressing a decrease in encoding efficiency.

[0098] For example, the header may be an APS (Adaptation Parameter Set).

[0099] For example, the header may be an SPS (Sequence Parameter Set).

[0100] For example, the header may be a PPS (Picture Parameter Set).

[0101] For example, the header may be a slice header or a tile header.

[0102] For example, the header may be a brick header.

[0103] For example, the header may be an SEI (Supplemental Enhancement Information).

[0104] For example, the filtering process may include an adaptive loop filtering process.

[0105] For example, the filtering process may include a cross-component adaptive loop filtering process.

[0106] For example, the filtering process may include a deblocking filtering process.

[0107] For example, the filtering process may include a sample adaptive offset process.

[0108] For example, the filtering process may include an LMCS (Luma mapping with chroma scaling) process.

[0109] Alternatively, the filtering process may use a sharpening filter.

[0110] A decoding device according to an aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit derives a parameter from information included in the header of a bitstream, generates a second image by filtering reconstructed samples of a first image using a filtering process, determines whether the parameter is a predetermined value, decodes a third image using the second image when the parameter is the predetermined value, decodes the third image using the first image when the parameter is not the predetermined value, and displays the second image.

[0111] Accordingly, the decoding device can switch which of the second image after application of the filtering process and the first image before application is used as a reference for subsequent pictures. Therefore, when the filtering process is effective in improving subjective image quality but not suitable as a reference image, by using the first image before application of the filtering process as the reference image, it is possible to improve subjective image quality while suppressing a decrease in coding efficiency.

[0112] For example, the header may be an APS (Adaptation Parameter Set).

[0113] For example, the header may be an SPS (Sequence Parameter Set).

[0114] For example, the header may be a PPS (Picture Parameter Set).

[0115] For example, the header may be a slice header or a tile header.

[0116] For example, the header may be a brick header.

[0117] For example, the header may be an SEI (Supplemental Enhancement Information).

[0118] For example, the filtering process may include an adaptive loop filtering process.

[0119] For example, the filtering process may include a cross-component adaptive loop filtering process.

[0120] For example, the filtering process may include a deblocking filtering process.

[0121] For example, the filtering process may include a sample adaptive offset process.

[0122] For example, the filtering process may include an LMCS (Luma mapping with chroma scaling) process.

[0123] For example, it may also be that the filtering process uses a sharpening filter.

[0124] Furthermore, in a coding method according to an aspect of the present invention, information for deriving a parameter is encoded into the header of a bitstream, a second image is generated by filtering reconstructed samples of a first image using a filtering process, it is determined whether the parameter is a predetermined value, and when the parameter is the predetermined value, the third image is encoded using the second image, and when the parameter is not the predetermined value, the third image is encoded using the first image.

[0125] According to this configuration, the coding method can switch which of the second image after application of the filtering process and the first image before application is used as a reference for subsequent pictures. Therefore, when the filtering process is effective in improving subjective image quality but not suitable as a reference image, by using the first image before application of the filtering process as a reference image, it is possible to improve subjective image quality while suppressing a decrease in coding efficiency.

[0126] Furthermore, in a decoding method according to an aspect of the present invention, a parameter is derived from information included in the header of a bitstream, a second image is generated by filtering reconstructed samples of a first image using a filtering process, it is determined whether the parameter is a predetermined value, and when the parameter is the predetermined value, the third image is decoded using the second image, and when the parameter is not the predetermined value, the third image is decoded using the first image, and the second image is displayed.

[0127] Thereby, the decoding method can switch which of the second image after application of the filtering process and the first image before application is used as a reference for subsequent pictures. Therefore, when the filtering process is effective in improving subjective image quality but not suitable as a reference image, by using the first image before application of the filtering process as a reference image, it is possible to improve subjective image quality while suppressing a decrease in coding efficiency.

[0128] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and are not intended to limit the claims.

[0129] Hereinafter, embodiments of the encoding device and the decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processes and / or structures described in each aspect of the present invention. The processes and / or structures can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processes and / or structures applied to the embodiments, any of the following can be done, for example.

[0130] (1) Among a plurality of components of the encoding device or the decoding device of the embodiment described in each aspect of the present invention, a certain component can be replaced with another component described in a certain aspect of the present invention, or they can be combined;

[0131] (2) In the encoding device or the decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. of functions or processes performed by a part of the components of the encoding device or the decoding device can be made. For example, any function or process can be replaced with another function or process described in a certain aspect of the present invention, or they can be combined;

[0132] (3) In the method implemented by the encoding device or the decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can be made to a part of the processes included in the method. For example, any process in the method can be replaced with another process described in a certain aspect of the present invention, or they can be combined;

[0133] (4) A part of the components constituting the encoding device or the decoding device of the embodiment can be combined with the components described in a certain aspect of the present invention, can be combined with the components having a part of the functions described in a certain aspect of the present invention, or can be combined with the components implementing a part of the processes implemented by the components described in the aspects of the present invention;

[0134] (5) Components having a part of the functions of the encoding device or the decoding device of the embodiment, or components implementing a part of the processes of the encoding device or the decoding device of the embodiment, are combined or replaced with the components described in a certain aspect of the present invention, the components having a part of the functions described in a certain aspect of the present invention, or the components implementing a part of the processes described in a certain aspect of the present invention;

[0135] (6) In the method implemented by the encoding device or the decoding device of the embodiment, a certain one of the plurality of processes included in the method is replaced with a process described in a certain aspect of the present invention or the same certain process, or they are combined;

[0136] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any of the aspects of the present invention.

[0137] (8) The manner of implementing the processes and / or structures described in the aspects of the present invention is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the moving image encoding or moving image decoding disclosed in the embodiment.

[0138] [Encoding Device]

[0139] First, the encoding device of the embodiment will be described. Figure 1 FIG. is a block diagram showing the functional configuration of the encoding device 100 of the embodiment. The encoding device 100 is a moving image encoding device that encodes a moving image in units of blocks.

[0140] As Figure 1 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a splitting unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0141] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 can also be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0142] Hereinafter, after explaining the overall processing flow of the encoding device 100, each component included in the encoding device 100 will be described.

[0143] [Overall Flow of Encoding Processing]

[0144] Figure 2 FIG. is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0145] First, the segmentation unit 102 of the encoding device 100 divides each picture included in the input picture as a moving image into a plurality of blocks of a fixed size (for example, 128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern (also referred to as a block shape) for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block having the fixed size into a plurality of blocks constituting the selected segmentation pattern. Then, for each of the plurality of blocks, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 on the block (that is, the encoding target block).

[0146] That is, a prediction processing unit constituted by all or a part of the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128 generates a prediction signal (also referred to as a prediction block) of the encoding target block (also referred to as the current block) (step Sa_3).

[0147] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).

[0148] Next, the transform unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing a transform and quantization on the difference block (step Sa_5). In addition, a block constituted by a plurality of quantization coefficients is also referred to as a coefficient block.

[0149] Next, the entropy encoding unit 110 generates an encoded signal (step Sa_6) by encoding (specifically, entropy encoding) the coefficient block and prediction parameters related to the generation of the prediction signal. In addition, the encoded signal is also referred to as an encoded bitstream, a compressed bitstream, or a stream.

[0150] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore a plurality of prediction residuals (that is, difference blocks) by performing an inverse quantization and an inverse transform on the coefficient block (step Sa_7).

[0151] Next, the addition unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the restored difference block to the prediction block (step Sa_8). Thereby, a reconstructed image is generated.

[0152] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).

[0153] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (No in step Sa_10), the processes starting from step Sa_2 are repeated.

[0154] In addition, in the above example, the encoding device 100 selects one splitting pattern for blocks of a fixed size and encodes each block according to the splitting pattern. However, each block may also be encoded according to each of a plurality of splitting patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of splitting patterns, and for example, may select the encoded signal obtained by encoding according to the splitting pattern with the minimum cost as the output encoded signal.

[0155] As shown in the figure, the processes of these steps Sa_1 to Sa_10 are sequentially performed by the encoding device 100. Alternatively, some of these processes may be performed in parallel, or the order of these processes may be changed.

[0156] [Splitting Unit]

[0157] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (for example, 128×128). Other fixed block sizes may also be used. Such blocks of a fixed size are sometimes referred to as coding tree units (CTUs). And, the splitting unit 102 splits each block of a fixed size into blocks of a variable size (for example, 64×64 or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects a splitting pattern. Such blocks of a variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture may be used as the processing units for CUs, PUs, and TUs.

[0158] Figure 3 is a conceptual diagram showing an example of block splitting in the embodiment. In Figure 3 the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.

[0159] Here, the block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into 4 square 64×64 blocks (quadtree block splitting).

[0160] The upper left 64×64 block is further vertically split into 2 rectangular 32×64 blocks, and the left 32×64 block is further vertically split into 2 rectangular 16×64 blocks (binary tree block splitting). As a result, the upper left 64×64 block is split into 2 16×64 blocks 11, 12 and a 32×64 block 13.

[0161] The upper-right 64×64 block is horizontally divided into two rectangular 64×32 blocks 14 and 15 (binary tree block division).

[0162] The lower-left 64×64 block is divided into four square 32×32 blocks (quad-tree block division). The upper-left block and the lower-right block among the four 32×32 blocks are further divided. The upper-left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower-right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the lower-left 64×64 block is divided into 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.

[0163] The lower-right 64×64 block 23 is not divided.

[0164] As described above, in Figure 3 , block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0165] In addition, in Figure 3 , one block is divided into four or two blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). The division including such ternary tree division is sometimes called MBT (multi type tree) division.

[0166] [Structural Slices / Tiles of the Picture]

[0167] In order to decode a picture in parallel, a picture is sometimes composed of slices or tiles. A picture composed of slices or tiles can be formed by the dividing unit 102.

[0168] A slice is the basic encoding unit that constitutes a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs (Coding Tree Units).

[0169] Figure 4AIt is a conceptual diagram showing an example of the structure of slices. For example, the picture includes 11×8 CTUs and is divided into 4 slices (Slice 1 to Slice 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any one of the slices. The shape of the slice becomes the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the picture edge and can be any position among the boundaries of the CTUs within the picture. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. In addition, the slice contains header information and encoded data. In the header information, the characteristics of the slice such as the CTU address at the start of the slice and the slice type can also be described.

[0170] A tile is a unit of a rectangular area that makes up a picture. Numbers called TileIds can also be assigned to each tile in the raster scan order.

[0171] Figure 4B It is a conceptual diagram showing an example of the structure of tiles. For example, the picture includes 11×8 CTUs and is divided into 4 rectangular area tiles (Tile 1 to Tile 4). When using tiles, the processing order of the CTUs is changed compared to the case of not using tiles. When not using tiles, multiple CTUs in the picture are processed in the raster scan order. When using tiles, in each of the multiple tiles, at least 1 CTU is processed in the raster scan order. For example, as Figure 4B shown, the processing order of the multiple CTUs included in Tile 1 is the order from the left end of the first row of Tile 1 to the right end of the first row of Tile 1, and then from the left end of the second row of Tile 1 to the right end of the second row of Tile 1.

[0172] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.

[0173] [Subtraction unit]

[0174] The subtraction unit 104 subtracts the predicted signal (the predicted sample input from the prediction control unit 128 as shown below) from the original signal (original sample) in units of blocks input from and divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction unit 104 outputs the calculated prediction error (residual) to the transformation unit 106.

[0175] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, the luma signal and two chroma signals). Hereinafter, there are also cases where the signal representing the image is called a sample.

[0176] [Transformation section]

[0177] The transformation section 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain and outputs them to the transform coefficient vectorization section 108. Specifically, the transformation section 106 performs a prescribed discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example. The prescribed DCT or DST may also be determined in advance.

[0178] In addition, the transformation section 106 may adaptively select a transformation type from among multiple transformation types and use a transform basis function corresponding to the selected transformation type to transform the prediction error into transform coefficients. Such a transformation is called EMT (explicit multiple core transform) or AMT (adaptive multiple transform) in some cases.

[0179] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transform basis functions corresponding to the transformation type examples. In Figure 5A where N represents the number of input pixels. The selection of the transformation type from among these multiple transformation types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0180] Information indicating whether to apply such EMT or AMT (e.g., called an EMT flag or an AMT flag) and information indicating the selected transformation type are usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and may also be other levels (e.g., bit sequence level, picture level, slice level, tile level, or CTU level).

[0181] In addition, the transform unit 106 can also re-transform the transform coefficients (transformation results). Such re-transformations include cases known as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs re-transformation for each sub-block (e.g., 4×4 sub-block) included in the block of transform coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are usually signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0182] In the transform unit 106, separable transforms and non-separable transforms can also be applied. A separable transform is a method of performing multiple transforms by separating them in each direction according to the number of input dimensions. A non-separable transform is a method of treating two or more dimensions as one dimension and performing a transform together when the input is multi-dimensional.

[0183] For example, as an example of a non-separable transform, when the input is a 4×4 block, it can be regarded as a permutation with 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.

[0184] In addition, in a further example of a non-separable transform, after regarding a 4×4 input block as a permutation with 16 elements, a transform of performing Givens rotations on this permutation multiple times (Hypercube Givens Transform) can also be performed.

[0185] In the transform in the transform unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, as Figure 5BAs shown, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The type of transform basis can be set for each region. For example, DST7 and DCT8 can be used. In this example, only one of the two regions within the CU is transformed, and the other is not. However, both regions can also be transformed. In addition, the splitting method is not limited to bisection and can be more flexible. For example, it can be quartered or the information indicating the split is encoded separately and signaled in the same way as CU splitting. Sometimes, SVT is also referred to as SBT (Sub-block Transform).

[0186] [Quantization Unit]

[0187] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a prescribed scan order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112. The prescribed scan order can also be determined in advance.

[0188] The prescribed scan order is the order used for quantization / inverse quantization of the transform coefficients. For example, the prescribed scan order can be defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).

[0189] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0190] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. Quantization means digitizing the values sampled at a prescribed interval by corresponding them to a prescribed level. In this technical field, other expressions such as rounding, truncation, and scaling can also be used for reference, and rounding, truncation, and scaling can also be adopted. The prescribed interval and level can also be determined in advance.

[0191] As methods of using the quantization matrix, there are a method of using a quantization matrix directly set on the encoding device side and a method of using a default quantization matrix (default matrix). On the encoding device side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of coding increases due to the coding of the quantization matrix.

[0192] On the other hand, there is also a method of quantization in which the coefficients of the high-frequency components and the coefficients of the low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to the method of using a quantization matrix (flat matrix) in which all coefficients are the same value.

[0193] The quantization matrix can be specified by, for example, SPS (Sequence Parameter Set) or PPS (Picture Parameter Set). SPS contains parameters used for a sequence, and PPS contains parameters used for a picture. SPS and PPS are sometimes simply referred to as parameter sets.

[0194] [Entropy Encoding Unit]

[0195] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) based on the quantized coefficients input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantized coefficients, for example, performs arithmetic coding on the binary signal, and outputs a compressed bitstream or sequence.

[0196] [Inverse Quantization Unit]

[0197] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114. The prescribed scan order can also be determined in advance.

[0198] [Inverse Transform Unit]

[0199] The inverse transform unit 114 restores the prediction error (residual) by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. And the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0200] In addition, the restored prediction error usually does not match the prediction error calculated by the subtraction unit 104 because information is lost through quantization. That is, the restored prediction error usually contains a quantization error.

[0201] [Addition Unit]

[0202] The addition unit 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the predicted sample input from the prediction control unit 128. And the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a local decoded block.

[0203] [Block memory]

[0204] The block memory 118 is, for example, a storage unit that stores blocks within an encoded object picture (referred to as the current picture) that is referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0205] [Frame memory]

[0206] The frame memory 122 is, for example, a storage unit that stores reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0207] [Loop filter unit]

[0208] The loop filter unit 120 performs loop filtering on the blocks reconstructed by the adder 116, and outputs the filtered reconstructed blocks to the frame memory 122. Loop filtering refers to filtering used within the encoding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0209] In ALF, a least squares error filter used to remove encoding distortion is employed. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters based on the direction and activity of the locality-based gradient is used.

[0210] Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4), the classification value C (e.g., C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes.

[0211] The gradient direction value D is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantifying the added result.

[0212] Based on the result of such classification, the filter to be used for the sub-block is determined from among multiple filters.

[0213] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figures 6A - 6C It is a diagram showing multiple examples of the shape of the filter used in ALF. Figure 6ARepresents a 5×5 rhombus-shaped filter, Figure 6B Represents a 7×7 rhombus-shaped filter, Figure 6C Represents a 9×9 rhombus-shaped filter. Information indicating the shape of the filter is typically signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0214] The on / off of the ALF can also be determined, for example, at the picture level or CU level. For example, regarding luminance, it can be determined at the CU level whether to employ the ALF, and regarding chrominance difference, it can be determined at the picture level whether to employ the ALF. Information indicating the on / off of the ALF is typically signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of the ALF need not be limited to the picture level or CU level and may also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0215] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level. Additionally, the signaling of the coefficient sets need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0216] [Loop Filtering Section > Deblocking Filter]

[0217] In the deblocking filter, the loop filtering section 120 reduces the distortion generated at the block boundary by filtering the block boundary of the reconstructed image.

[0218] Figure 7 Is a block diagram showing an example of the detailed structure of the loop filtering section 120 that functions as a deblocking filter.

[0219] The loop filtering section 120 includes a boundary determination section 1201, a filtering determination section 1203, a filtering processing section 1205, a processing determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0220] The boundary determination section 1201 determines whether there are pixels (i.e., target pixels) for which deblocking filtering is to be performed near the block boundary. Then, the boundary determination section 1201 outputs the determination result to the switches 1202 and the processing determination section 1208.

[0221] When the boundary determination unit 1201 determines that the target pixel exists near the block boundary, the switch 1202 outputs the image before the filtering process to the switch 1204. On the contrary, when the boundary determination unit 1201 determines that the target pixel does not exist near the block boundary, the switch 1202 outputs the image before the filtering process to the switch 1206.

[0222] The filtering determination unit 1203 determines whether to perform deblocking filtering on the target pixel based on the pixel values of at least one neighboring pixel located around the target pixel. Then, the filtering determination unit 1203 outputs the determination result to the switch 1204 and the processing determination unit 1208.

[0223] When the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the filtering processing unit 1205. On the contrary, when the filtering determination unit 1203 determines that deblocking filtering is not to be performed on the target pixel, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the switch 1206.

[0224] When the image before the filtering process is obtained via the switches 1202 and 1204, the filtering processing unit 1205 performs deblocking filtering on the target pixel with the filtering characteristics determined by the filtering characteristic determination unit 1207. Then, the filtering processing unit 1205 outputs the pixel after the filtering process to the switch 1206.

[0225] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs the pixels that have not been deblocked filtered and the pixels that have been deblocked filtered by the filtering processing unit 1205.

[0226] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processed pixel after deblocking filtering is output from the switch 1206. In addition, in other cases than the above, the processing determination unit 1208 outputs the pixel that has not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the image after the filtering process is output from the switch 1206.

[0227] Figure 8 It is a conceptual diagram showing an example of deblocking filtering having filtering characteristics symmetric with respect to the block boundary.

[0228] In deblocking filtering processing, for example, using pixel values and quantization parameters, one of two deblocking filters with different characteristics, namely a strong filter and a weak filter, is selected. In the strong filter, as Figure 8 shown, when there are pixels p0 to p2 and pixels q0 to q2 across a block boundary, the pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing operations shown in the following equations, for example.

[0229] q'0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0230] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0231] q'2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0232] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of pixels p0 to p2 and pixels q0 to q2 respectively. Also, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the side opposite to the block boundary. Also, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering processing are filtering coefficients.

[0233] Furthermore, in the deblocking filtering processing, clipping processing may also be performed in such a way that the pixel value after the operation is set not to exceed a threshold value. In this clipping processing, using a threshold value determined according to the quantization parameter, the pixel value after the operation based on the above equations is clipped to "operation target pixel value ± 2×threshold value". Thereby, excessive smoothing can be prevented.

[0234] Figure 9 is a conceptual diagram for explaining the block boundary where deblocking filtering processing is performed. Figure 10 is a conceptual diagram showing an example of the Bs value.

[0235] The block boundary where deblocking filtering processing is performed is, for example, Figure 9 the boundary of a PU (Prediction Unit) or a TU (Transform Unit) of an 8×8 pixel block as shown. Deblocking filtering processing can be performed in units of 4 rows or 4 columns. First, for Figure 9 the blocks P and Q shown, as Figure 10 shown, the Bs (Boundary Strength) value is determined.

[0236] According to Figure 10The Bs value determines whether to perform deblocking filtering with different strengths even for block boundaries belonging to the same image. Deblocking filtering for the chrominance signal is performed when the Bs value is 2. Deblocking filtering for the luminance signal is performed when the Bs value is 1 or more and a specified condition is satisfied. The specified condition can also be determined in advance. In addition, the determination condition of the Bs value is not limited to Figure 10 the conditions shown, and can also be determined based on other parameters.

[0237] [Prediction processing unit (intra-frame prediction unit / inter-frame prediction unit / prediction control unit)]

[0238] Figure 11 is a flowchart showing an example of the processing performed by the prediction processing unit of the encoding apparatus 100. In addition, the prediction processing unit is composed of all or part of the components of the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.

[0239] The prediction processing unit generates a prediction image of the current block (step Sb_1). This prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there are an intra-frame prediction signal and an inter-frame prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using the reconstructed image that has been obtained by performing generation of a prediction block, generation of a differential block, generation of a coefficient block, restoration of the differential block, and generation of a decoded image block.

[0240] The reconstructed image can be, for example, an image of a reference picture, or an image of an encoded block within the current picture including the current block, that is, the current picture. The encoded block within the current picture is, for example, an adjacent block of the current block.

[0241] Figure 12 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding apparatus 100.

[0242] The prediction processing unit generates a prediction image by a first method (step Sc_1a), generates a prediction image by a second method (step Sc_1b), and generates a prediction image by a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such prediction methods, the above-described reconstructed image can also be used.

[0243] Next, the prediction processing unit selects any one of the plurality of prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). The selection of this prediction image, that is, the selection of the method or mode for obtaining the final prediction image, may also calculate the cost for each generated prediction image and be based on this cost. In addition, the selection of this prediction image may be performed based on the parameters for the encoding process. The encoding device 100 may signal information for determining the selected prediction image, method, or mode as an encoded signal (also referred to as an encoded bitstream). This information may be, for example, a flag or the like. Thus, the decoding device can generate a prediction image based on this information in the method or mode selected in the encoding device 100. In addition, in Figure 12 In the example shown, after generating prediction images by each method, the prediction processing unit selects any one of the prediction images. However, before generating these prediction images, the prediction processing unit may select a method or mode based on the parameters for the above-mentioned encoding process and may generate a prediction image according to this method or mode.

[0244] For example, the first method and the second method are intra-frame prediction and inter-frame prediction, respectively, and the prediction processing unit may select the final prediction image for the current block from the prediction images generated according to these prediction methods.

[0245] Figure 13 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0246] First, the prediction processing unit generates a prediction image by intra-frame prediction (step Sd_1a) and generates a prediction image by inter-frame prediction (step Sd_1b). In addition, the prediction image generated by intra-frame prediction is also referred to as an intra-frame prediction image, and the prediction image generated by inter-frame prediction is also referred to as an inter-frame prediction image.

[0247] Next, the prediction processing unit evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). The cost may also be used in this evaluation. That is, the prediction processing unit calculates the cost C for each of the intra-frame prediction image and the inter-frame prediction image. This cost C can be calculated by an equation of the R-D optimization model, for example, C = D + λ × R. In this equation, D is the encoding distortion of the prediction image and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. In addition, R is the generated coding amount of the prediction image. Specifically, it is the coding amount required for encoding motion information or the like for generating the prediction image. In addition, λ is, for example, the undetermined multiplier of Lagrange.

[0248] Then, the prediction processing unit selects, as the final prediction image of the current block, the prediction image that calculates the minimum cost C from the intra-frame prediction image and the inter-frame prediction image (step Sd_3). That is, the prediction method or mode used to generate the prediction image of the current block is selected.

[0249] [Intra-frame prediction unit]

[0250] The intra-frame prediction unit 124 performs intra-frame prediction (also called intra-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.

[0251] For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of prescribed intra-frame prediction modes. The plurality of intra-frame prediction modes usually include one or more non-directional prediction modes and a plurality of directional prediction modes. The plurality of prescribed modes may also be determined in advance.

[0252] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode specified by the H.265 / HEVC standard.

[0253] The plurality of directional prediction modes include, for example, 33-direction prediction modes specified by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may also include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Figure 14 is a conceptual diagram showing all 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra-frame prediction. The solid arrows represent 33 directions specified by the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the 2 non-directional prediction modes are not Figure 14 shown in the figure).

[0254] In various processing examples, in the intra-frame prediction of the chrominance blocks, the luminance blocks may also be referred to. That is, the chrominance components of the current block may be predicted based on the luminance component of the current block. Such intra-frame prediction is sometimes called CCLM (cross-component linear model) prediction. The intra-frame prediction mode of the chrominance block that refers to the luminance block (e.g., called the CCLM mode) may also be added as one of the intra-frame prediction modes of the chrominance block.

[0255] The intra prediction unit 124 may also correct the intra-predicted pixel value based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction with such correction is called PDPC (position dependent intraprediction combination) in some cases. Information indicating whether PDPC is used (e.g., called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0256] [Inter-frame prediction unit]

[0257] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter-frame prediction signal). The inter-frame prediction is performed in units of the current block or the current sub-block (e.g., 4×4 block) within the current block. For example, the inter-frame prediction unit 126 performs motion search (motion estimation) within the reference picture for the current block or the current sub-block to find the reference block or sub-block that most matches the current block or the current sub-block. And the inter-frame prediction unit 126 obtains motion information (e.g., motion vector) that compensates for the motion or change from the reference block or sub-block to the current block or the current sub-block. The inter-frame prediction unit 126 performs motion compensation (or motion prediction) based on this motion information, thereby generating an inter-frame prediction signal for the current block or the current sub-block. And the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.

[0258] The motion information used in motion compensation is signaled as an inter-frame prediction signal in various forms. For example, the motion vector may also be signaled. As another example, the difference between the motion vector and the predicted motion vector (motion vector predictor) may also be signaled.

[0259] [Basic process of inter-frame prediction]

[0260] Figure 15 It is a flowchart showing an example of the basic process of inter-frame prediction.

[0261] The inter-frame prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the prediction image as a prediction residual (step Se_4).

[0262] Here, in the generation of the predicted image, the inter-frame prediction unit 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. Further, in the derivation of the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from at least one candidate MV and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching for regions of the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. Further, the action of searching for regions of the reference picture may also be referred to as motion estimation.

[0263] Further, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126. However, for example, the processing such as step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.

[0264] [Flow of Derivation of Motion Vector]

[0265] Figure 16 is a flowchart showing an example of the derivation of the motion vector.

[0266] The inter-frame prediction unit 126 derives the MV of the current block in a mode in which motion information (e.g., MV) is encoded. In this case, for example, the motion information is encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the encoded signal (also referred to as the encoded bitstream).

[0267] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the encoded signal.

[0268] Here, the mode of MV derivation may also include a normal inter-frame mode, a merge mode, a FRUC mode, an affine mode, etc., which will be described later. Among these modes, the modes in which motion information is encoded include a normal inter-frame mode, a merge mode, and an affine mode (specifically, an affine inter-frame mode and an affine merge mode), etc. Further, the motion information may include not only the MV but also prediction motion vector selection information, which will be described later. Further, the mode in which motion information is not encoded includes a FRUC mode, etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0269] Figure 17 This is a flowchart showing another example of motion vector derivation.

[0270] The inter-frame prediction unit 126 derives the MV of the current block in the mode of encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0271] Alternatively, the inter-frame prediction unit 126 derives the MV in the mode of not encoding the differential MV. In this case, the encoded differential MV is not included in the encoded signal.

[0272] Here, as described above, the derivation modes of the MV include the ordinary inter-frame mode, merge mode, FRUC mode, affine mode, etc., which will be described later. Among these modes, the modes of encoding the differential MV include the ordinary inter-frame mode and the affine mode (specifically, the affine inter-frame mode), etc. In addition, the modes of not encoding the differential MV include the FRUC mode, merge mode, and affine mode (specifically, the affine merge mode), etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0273] [Flow of Motion Vector Derivation]

[0274] Figure 18 This is a flowchart showing another example of motion vector derivation. The modes of MV derivation, that is, the inter-frame prediction modes, have multiple modes, which are roughly divided into the mode of encoding the differential MV and the mode of not encoding the differential motion vector. The modes of not encoding the differential MV include the merge mode, FRUC mode, and affine mode (specifically, the affine merge mode). The detailed content of these modes will be described later. Briefly, the merge mode is a mode of deriving the MV of the current block by selecting a motion vector from the surrounding encoded blocks, and the FRUC mode is a mode of deriving the MV of the current block by searching between the encoded regions. In addition, the affine mode is a mode of assuming an affine transformation and deriving the MV of the current block by using the motion vectors of multiple sub-blocks constituting the current block as the MV of the current block.

[0275] Specifically, as shown in the figure, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector based on the merge mode (Sf_2). In addition, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the FRUC mode (Sf_3). In addition, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the affine mode (specifically, the affine merge mode) (Sf_4). In addition, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the mode for encoding the differential MV (e.g., the normal inter-frame mode) (Sf_5).

[0276] [MV Derivation > Normal Inter-Frame Mode]

[0277] The normal inter-frame mode is an inter-frame prediction mode in which the MV of the current block is derived based on a block similar to the image of the current block from the region of the reference picture represented by the candidate MV. In addition, in this normal inter-frame mode, the differential MV is encoded.

[0278] Figure 19 It is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.

[0279] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially around the current block (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0280] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order (step Sg_2). In addition, this priority order may also be determined in advance for each of the N candidate MVs.

[0281] Next, the inter-frame prediction unit 126 selects one prediction motion vector candidate from the N prediction motion vector candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction motion vector selection information for identifying the selected prediction motion vector into the stream. In addition, the stream is the above-mentioned encoded signal or encoded bitstream.

[0282] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the predicted motion vector as a differential MV into the stream. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0283] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). In addition, the predicted image is the above-mentioned inter-frame prediction signal.

[0284] In addition, information indicating the inter-frame prediction mode (in the above example, the ordinary inter-frame mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.

[0285] In addition, the candidate MV list can also be used in common with the list used in other modes. In addition, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to this candidate MV list is, for example, extracting or selecting a candidate MV from the candidate MV list, rearranging the candidate MVs, or deleting a candidate MV, etc.

[0286] [MV Derivation > Merge Mode]

[0287] The merge mode is an inter-frame prediction mode in which the MV of the current block is derived by selecting a candidate MV from the candidate MV list as the MV of the current block.

[0288] Figure 20 It is a flowchart showing an example of inter-frame prediction based on the merge mode.

[0289] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0290] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream.

[0291] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).

[0292] In addition, information indicating an inter-prediction mode (in the above example, the merge mode) used in the generation of a predicted image included in an encoded signal is encoded as, for example, a prediction parameter.

[0293] Figure 21 FIG. is a conceptual diagram for explaining an example of a motion vector derivation process of a current picture based on the merge mode.

[0294] First, a prediction MV list registering candidates of predicted MVs is generated. As candidates of predicted MVs, there are: a spatial neighboring prediction MV which is an MV possessed by a plurality of encoded blocks located in the spatial periphery of an object block; a temporal neighboring prediction MV which is an MV possessed by a nearby block obtained by projecting the position of an object block in an encoded reference picture; a combined prediction MV which is an MV generated by combining MV values of the spatial neighboring prediction MV and the temporal neighboring prediction MV; and a zero prediction MV which is an MV having a value of zero, and the like.

[0295] Next, an MV for the object block is determined by selecting one predicted MV from among the plurality of predicted MVs registered in the prediction MV list.

[0296] Moreover, in a variable length coding unit, a signal indicating which predicted MV is selected, i.e., merge_idx, is described in a stream and encoded.

[0297] In addition, in Figure 21 the predicted MVs registered in the prediction MV list described above are an example, and the number may be different from the number in the figure, or the structure may not include some types of the predicted MVs in the figure, or the structure may be appended with predicted MVs other than the types of predicted MVs in the figure.

[0298] An MV of an object block derived by the merge mode may also be used, and a final MV may be determined by performing a DMVR (decode motion vector refinement) process described later.

[0299] In addition, candidates of predicted MVs are the above-described candidate MVs, and the prediction MV list is the above-described candidate MV list. In addition, the candidate MV list may also be referred to as a candidate list. In addition, merge_idx is MV selection information.

[0300] [MV derivation>FRUC mode]

[0301] Motion information may not be signaled from the encoding device side but may be derived on the decoding device side. In addition, as described above, the merge mode defined by the H.265 / HEVC standard may also be used. In addition, for example, motion information may be derived by performing a motion search on the decoding device side. In an embodiment, a motion search is performed on the decoding device side without using pixel values of a current block.

[0302] Here, the mode of performing motion estimation on the decoder side will be described. The mode of performing motion estimation on the decoder side may be a mode called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0303] An example of FRUC processing is shown in Figure 22 in the form of a flowchart. First, referring to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector (MV) are generated (i.e., a candidate MV list, which may also be shared with the merge list) (step Si_1). Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected based on the evaluation value. And, based on the motion vector of the selected candidate, the motion vector for the current block is derived (step Si_4). Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly used as the motion vector for the current block. In addition, for example, the motion vector for the current block may also be derived by performing pattern matching in the peripheral area of the position in the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral area of the best candidate MV may be searched by using pattern matching and the evaluation value in the reference picture. In the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and it is used as the final MV of the current block. It is also possible to configure it without performing the process of updating to an MV with a better evaluation value.

[0304] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block by using the derived MV and the encoded reference picture (step Si_5).

[0305] The same process may also be performed when processing in units of sub-blocks.

[0306] The evaluation value may also be calculated by various methods. For example, the reconstructed image of the area in the reference picture corresponding to the motion vector is compared with the reconstructed image of a specified area (for example, as shown below, this area may be an area of another reference picture or an adjacent block of the current picture). The specified area may also be determined in advance.

[0307] Then, the difference between the pixel values of the two reconstructed images can also be calculated for use as the evaluation value of the motion vector. Additionally, it can be such that other information is used in addition to the difference value to calculate the evaluation value.

[0308] Next, an example of pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (e.g., the merged list) is selected as the starting point for the search based on pattern matching. For example, as the pattern matching, the first pattern matching or the second pattern matching can be used. There are cases where the first pattern matching and the second pattern matching are respectively called bilateral matching and template matching.

[0309] [MV derivation>FRUC>Bilateral matching]

[0310] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Thus, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate described above, the region in another reference picture along the motion trajectory of the current block is used. The specified region can also be determined in advance.

[0311] Figure 23 is a conceptual diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As Figure 23 shown, in the first pattern matching, by searching for the most matching pair among the pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two motion vectors (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the obtained difference value is used to calculate the evaluation value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, and good results can be obtained.

[0312] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0313] [MV Derivation>FRUC>Template Matching]

[0314] In the second pattern matching (template matching), pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture (e.g., an upper and / or left adjacent block)) and a block within the reference picture. Thus, in the second pattern matching, as the specified region for calculating the evaluation value for the above candidates, a block adjacent to the current block within the current picture is used.

[0315] Figure 24 is a conceptual diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within the reference picture. As Figure 24 shown, in the second pattern matching, by searching for the block in the reference picture (Ref0) that best matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic), the motion vector of the current block is derived. Specifically, for the current block, the difference between the reconstructed images of the left adjacent and / or upper adjacent encoded regions and the reconstructed image at the same position within the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.

[0316] Information indicating whether to adopt the FRUC mode (e.g., called the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the pattern matching method that can be adopted (the first pattern matching or the second pattern matching) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0317] [MV Derivation>Affine Mode]

[0318] Next, an affine mode for deriving a motion vector in units of sub - blocks based on motion vectors of multiple adjacent blocks will be described. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0319] Figure 25A is a conceptual diagram for explaining an example of the derivation of a motion vector in units of sub - blocks based on motion vectors of multiple adjacent blocks. In Figure 25A the current block includes 16 4×4 sub - blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v 0 of the upper - left control point of the current block is derived. Similarly, based on the motion vectors of adjacent sub - blocks, the motion vector v 1 of the upper - right control point of the current block is derived. Then, according to the following equation (1A), the two motion vectors v 0 and v 1 can be projected, and the motion vectors (v x , v y ) of each sub - block within the current block can also be derived.

[0320] [Equation 1]

[0321]

[0322] Here, x and y represent the horizontal position and vertical position of the sub - block respectively, and w represents a prescribed weight coefficient. The prescribed weight coefficient can also be determined in advance.

[0323] The information indicating such an affine mode (e.g., called an affine flag) can be signaled as a signal at the CU level. In addition, the signaling of the information indicating this affine mode does not need to be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub - block level).

[0324] In addition, in such an affine mode, several modes with different methods for deriving the motion vectors of the upper - left and upper - right control points can also be included. For example, in the affine mode, there are two modes: the affine inter - frame (also called affine normal inter - frame) mode and the affine merge mode.

[0325] [MV Derivation > Affine Mode]

[0326] Figure 25B is a conceptual diagram for explaining an example of the derivation of a motion vector in units of sub - blocks in an affine mode with three control points. In Figure 25B the current block includes 16 4×4 sub - blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v 0, similarly, the motion vector v of the upper right control point of the current block is derived based on the motion vectors of adjacent blocks 1 , the motion vector v of the lower left control point of the current block is derived based on the motion vectors of adjacent blocks 2 . Then, according to the following equation (1B), the three motion vectors v 0 、v 1 and v 2 can be projected, and the motion vectors (v x , v y ) of each sub-block within the current block can also be derived.

[0327]

Equation 2

[0328]

[0329] Here, x and y represent the horizontal position and vertical position of the center of the sub-block respectively, w represents the width of the current block, and h represents the height of the current block.

[0330] Affine modes with different numbers of control points (e.g., 2 and 3) can also be switched and signaled at the CU level. Additionally, the information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0331] In addition, in such an affine mode with 3 control points, several modes with different methods for deriving the motion vectors of the upper left, upper right, and lower left control points can also be included. For example, in the affine mode, there are 2 modes: affine inter-frame (also known as affine ordinary inter-frame) mode and affine merge mode.

[0332] [MV Derivation > Affine Merge Mode]

[0333] Figure 26A 、 Figure 26B and Figure 26C are conceptual diagrams for explaining the affine merge mode.

[0334] In the affine merge mode, as Figure 26A shown, for example, based on multiple motion vectors corresponding to the coded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block and coded in the affine mode, the predicted motion vector of each control point of the current block is calculated. Specifically, these blocks are checked in the order of coded block A (left), B (upper), C (upper right), D (lower left), and E (upper left), and the first valid block coded in the affine mode is determined. The predicted motion vector of the control point of the current block is calculated based on the multiple motion vectors corresponding to the determined block.

[0335] For example, as Figure 26B shown, when encoding block A adjacent to the left of the current block in an affine mode with 2 control points, a motion vector v projected onto positions at the upper left corner and upper right corner of the encoded block containing block A is derived 3 and v 4 . Then, based on the derived motion vectors v 3 and v 4 , the predicted motion vectors v 0 of the control point at the upper left corner of the current block and v 1 of the control point at the upper right corner are calculated

[0336] For example, as Figure 26C shown, when encoding block A adjacent to the left of the current block in an affine mode with 3 control points, motion vectors v 3 , v 4 and v 5 projected onto positions at the upper left corner, upper right corner and lower left corner of the encoded block containing block A are derived 3 , v 4 and v 5 . Then, based on the derived motion vectors v 0 , v 1 and v 2 , the predicted motion vectors v

[0337] In addition, in the derivation of the predicted motion vectors of the respective control points of the current block in step Sj_1 described later Figure 29 , this predicted motion vector derivation method can also be used

[0338] Figure 27 is a flowchart showing an example of the affine merge mode

[0339] In the affine merge mode, as shown in the figure, first, the inter-frame prediction unit 126 derives the predicted MV of each control point of the current block (step Sk_1). The control points are, as Figure 25A shown, the points at the upper left corner and upper right corner of the current block, or, as Figure 25B shown, the points at the upper left corner, upper right corner and lower left corner of the current block

[0340] That is, as Figure 26A shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded blocks A (left), B (above), C (upper right), D (lower left) and E (upper left), and determines the initial valid block encoded in the affine mode

[0341] Then, when block A is determined and block A has two control points, as Figure 26B shown, the inter-frame prediction unit 126 calculates the motion vectors v 3 and v 4 of the control point at the upper left corner of the current block based on the motion vectors of the upper left corner and the upper right corner of the encoded block containing block A 0 and the motion vector v 1 of the control point at the upper right corner. For example, by projecting the motion vectors v 3 and v 4 of the upper left corner and the upper right corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v 0 of the control point at the upper left corner of the current block 1 and the predicted motion vector v

[0342] Or, when block A is determined and block A has three control points, as Figure 26C shown, the inter-frame prediction unit 126 calculates the motion vectors v 3 、v 4 and v 5 of the control point at the upper left corner of the current block based on the motion vectors of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A 0 、the motion vector v 1 of the control point at the upper right corner, and the motion vector v 2 of the control point at the lower left corner. For example, by projecting the motion vectors v 3 、v 4 and v 5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v 0 of the control point at the upper left corner of the current block 1 、the predicted motion vector v 2 of the control point at the upper right corner, and the motion vector v

[0343] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 126 uses two predicted motion vectors v 0 and v 1 and the above formula (1A) for each of the multiple sub-blocks, or three predicted motion vectors v 0 、v 1 and v 2Based on the above formula (1B), the motion vector of the sub-block is calculated as the affine MV (step Sk_2). Then, the inter-frame prediction unit 126 uses this affine MV and the encoded reference picture to perform motion compensation on the sub-block (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0344] [MV Derivation>Affine Inter-frame Mode]

[0345] Figure 28A is a conceptual diagram for explaining the affine inter-frame mode with two control points.

[0346] In this affine inter-frame mode, as Figure 28A shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v of the control point at the upper left corner of the current block 0 . Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v of the control point at the upper right corner of the current block 1 .

[0347] Figure 28B is a conceptual diagram for explaining the affine inter-frame mode with three control points.

[0348] In this affine inter-frame mode, as Figure 28B shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v of the control point at the upper left corner of the current block 0 . Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v of the control point at the upper right corner of the current block 1 . In addition, the motion vector selected from the motion vectors of the encoded blocks F and G adjacent to the current block is used as the predicted motion vector v of the control point at the lower left corner of the current block 2 .

[0349] Figure 29 is a flowchart showing an example of the affine inter-frame mode.

[0350] As shown in the figure, in the affine inter-frame mode, first, the inter-frame prediction unit 126 derives the predicted MV (v 0 , v 1 ) or (v 0 , v 1 , v 2 ) for each of the two or three control points of the current block (step Sj_1). As Figure 25A or Figure 25B shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0351] That is, the inter-frame prediction unit 126 derives the predicted motion vector (v Figure 28A or Figure 28B ) or (v 0 , v 1 ) or (v 0 , v 1 , v 2 ) of the control point of the current block by selecting a motion vector of a coded block near each control point of the current block shown in. At this time, the inter-frame prediction unit 126 encodes the prediction motion vector selection information for identifying the two selected motion vectors into the stream.

[0352] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select from the coded blocks adjacent to the current block as the predicted motion vector of the control point by using cost evaluation or the like, and can describe a flag indicating which predicted motion vector is selected in the bitstream.

[0353] Next, the inter-frame prediction unit 126 performs a motion search (steps Sj_3 and Sj_4) while using the predicted motion vector selected or derived in the update step Sj_1 (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vectors of the respective sub-blocks corresponding to the predicted motion vector to be updated as the affine MVs, and calculates using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the coded reference picture (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted motion vector as a differential MV into the stream.

[0354] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the coded reference picture (step Sj_6).

[0355] [MV Derivation>Affine Inter-frame Mode]

[0356] In the case of signaling in an affine mode with different numbers of control points (for example, 2 and 3) switched at the CU level, sometimes the number of control points in the coded block and the current block is different. Figure 30A And Figure 30B is a conceptual diagram for explaining a method for deriving a predicted vector of a control point in the case where the number of control points in the coded block and the current block is different.

[0357] For example, as Figure 30AAs shown, when the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points, a motion vector v projected onto the positions of the upper left corner and upper right corner of the encoded block containing block A is derived. 3 and v 4 . Then, based on the derived motion vectors v 3 and v 4 , the predicted motion vector v 0 for the control point at the upper left corner of the current block and the predicted motion vector v 1 for the control point at the upper right corner are calculated. In addition, based on the derived motion vectors v 0 and v 1 , the predicted motion vector v 2 for the control point at the lower left corner is calculated.

[0358] For example, as Figure 30B shown, when the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v 3 , v 4 and v 5 projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v 3 , v 4 and v 5 , the predicted motion vector v 0 for the control point at the upper left corner of the current block and the predicted motion vector v 1 for the control point at the upper right corner are calculated.

[0359] In Figure 29 the derivation of each predicted motion vector of the control points of the current block in step Sj_1, this predicted motion vector derivation method can also be used.

[0360] [MV Derivation > DMVR]

[0361] Figure 31A is a flowchart showing the relationship between the merge mode and DMVR.

[0362] The inter-frame prediction unit 126 derives the motion vector of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, i.e., a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the motion vector derived in step Sl_1 as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is determined in the merge mode.

[0363] On the other hand, when it is determined in step Sl_1 that motion search is to be performed (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final motion vector for the current block by searching the peripheral region of the reference picture represented by the motion vector derived in step Sl_1 (step Sl_3). That is, in this case, the motion vector of the current block is determined by DMVR.

[0364] Figure 31B It is a conceptual diagram illustrating an example of the DMVR process for determining the MV.

[0365] First, the best MVP set for the current block (e.g., in the merge mode) is set as the candidate MV. Then, according to the candidate MV (L0), the reference pixels are determined based on the encoded picture in the L0 direction, i.e., the first reference picture (L0). Similarly, according to the candidate MV (L1), the reference pixels are determined based on the encoded picture in the L1 direction, i.e., the second reference picture (L1). A template is generated by taking the average of these reference pixels.

[0366] Next, using the above template, the peripheral regions of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, the cost value can also be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search region, as well as the candidate MV values, etc.

[0367] In addition, typically, in the encoding device and the decoding device described later, the structure and operation of the processes described here are basically common.

[0368] Even if it is not the process example itself described here, as long as it is a process capable of searching the periphery of the candidate MV to derive the final MV, any process can be used.

[0369] [Motion Compensation > BIO / OBMC]

[0370] In motion compensation, there is a mode of generating a prediction image and correcting the prediction image. This mode is, for example, BIO and OBMC described later.

[0371] Figure 32 It is a flowchart showing an example of the generation of the prediction image.

[0372] The inter-frame prediction unit 126 generates a prediction image (step Sm_1), and corrects the prediction image by, for example, any of the above modes (step Sm_2).

[0373] Figure 33 It is a flowchart showing another example of the generation of the prediction image.

[0374] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image (step Sn_2), and determines whether to perform correction processing (step Sn_3). Here, when it is determined to perform correction processing (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting the predicted image (step Sn_4). On the other hand, when it is determined not to perform correction processing (No in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).

[0375] In addition, in motion compensation, there is a mode of correcting luminance when generating a predicted image. This mode is, for example, LIC described later.

[0376] Figure 34 It is a flowchart showing another example of generating a predicted image.

[0377] The inter-frame prediction unit 126 derives the motion vector of the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform luminance correction processing (step So_2). Here, when it is determined to perform luminance correction processing (Yes in step So_2), the inter-frame prediction unit 126 generates a predicted image while performing luminance correction (step So_3). That is, a predicted image is generated by LIC. On the other hand, when it is determined not to perform luminance correction processing (No in step So_2), the inter-frame prediction unit 126 generates a predicted image by normal motion compensation without performing luminance correction (step So_4).

[0378] [Motion Compensation > OBMC]

[0379] Not only the motion information of the current block obtained through motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame prediction signal. Specifically, an inter-frame prediction signal can also be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained through motion search (within the reference picture) and a prediction signal based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0380] In the OBMC mode, information indicating the size of a sub-block for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Additionally, the level at which these pieces of information are signaled is not limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0381] A more specific example of the OBMC mode will be described. Figure 35 and Figure 36 are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process based on OBMC processing.

[0382] First, as Figure 36 shown, using the motion vector (MV) assigned to the processing target (current) block, a prediction image (Pred) based on normal motion compensation is obtained. In Figure 36 , the arrow "MV" points to the reference picture and indicates which block in the current picture the current block refers to for obtaining the prediction image.

[0383] Next, the motion vector (MV_L) that has been derived for the encoded left adjacent block is applied (reused) to the block to be encoded, and a prediction image (Pred_L) is obtained. The motion vector (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two prediction images Pred and Pred_L, the first correction of the prediction image is performed. This has the effect of blending the boundaries between adjacent blocks.

[0384] Similarly, the motion vector (MV_U) that has been derived for the encoded upper adjacent block is applied (reused) to the block to be encoded, and a prediction image (Pred_U) is obtained. The motion vector (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the prediction image is performed by overlapping the prediction image Pred_U with the prediction image that has undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The prediction image obtained through the second correction is the final prediction image of the current block where the boundaries with adjacent blocks are blended (smoothed).

[0385] In addition, the above example uses a two-path correction method with left adjacent and upper adjacent blocks, but this correction method can also be a three-path or higher-path correction method that also uses right adjacent and / or lower adjacent blocks.

[0386] Also, the overlapping region can be not the entire pixel region of the block, but only a partial region near the block boundary.

[0387] In addition, the predictive image correction process of OBMC is described herein. The predictive image correction process of OBMC is used to obtain one predictive image Pred by superimposing one reference picture with the additional predictive images Pred_L and Pred_U. However, in the case of correcting the predictive image based on multiple reference images, the same process can also be applied to each of the multiple reference pictures. In this case, by performing the image correction of OBMC based on multiple reference pictures, after obtaining the corrected predictive images from each reference picture, the final predictive image is obtained by further superimposing the obtained multiple corrected predictive images.

[0388] In addition, in OBMC, the unit of the object block can be the predictive block unit or the sub-block unit obtained by further dividing the predictive block.

[0389] As a method for determining whether to apply the OBMC process, for example, there is a method of using a signal indicating whether to apply the OBMC process, that is, obmc_flag. As a specific example, the encoding device can also determine whether the object block belongs to a region with complex motion. When the object block belongs to a region with complex motion, the encoding device sets the obmc_flag value to 1 and applies the OBMC process for encoding. When the object block does not belong to a region with complex motion, the encoding device sets the obmc_flag value to 0 and does not apply the OBMC process for block encoding. On the other hand, in the decoding device, by decoding the obmc_flag described in the stream (e.g., the compressed sequence), the decoding is performed by switching whether to apply the OBMC process according to this value.

[0390] In the above example, the inter-frame prediction unit 126 generates one rectangular predictive image for the rectangular current block. However, the inter-frame prediction unit 126 can generate multiple predictive images with shapes different from the rectangle for the rectangular current block, and can generate the final rectangular predictive image by combining these multiple predictive images. Shapes different from the rectangle can also be triangles, for example.

[0391] Figure 37 It is a conceptual diagram for explaining the generation of two triangular predictive images.

[0392] The inter-frame prediction unit 126 generates a triangular predictive image by performing motion compensation on the first triangular partition within the current block using the first MV of the first partition. Similarly, the inter-frame prediction unit 126 generates a triangular predictive image by performing motion compensation on the second triangular partition in the current block using the second MV of the second partition. Then, the inter-frame prediction unit 126 generates a rectangular predictive image identical to the current block by combining these predictive images.

[0393] In addition, in Figure 37In the example shown, the first partition and the second partition are each triangular, but they can also be trapezoidal, or they can be of different shapes from each other. Also, in Figure 37 the example shown, the current block is composed of two partitions, but it can also be composed of three or more partitions.

[0394] In addition, the first partition and the second partition can also be repeated. That is, the first partition and the second partition can also include the same pixel area. In this case, the predicted image in the first partition and the predicted image in the second partition can be used to generate the predicted image of the current block.

[0395] In addition, in this example, an example is shown in which the predicted images of both partitions are generated by inter-frame prediction, but the predicted image of at least one partition can also be generated by intra-frame prediction.

[0396] [Motion Compensation > BIO]

[0397] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.

[0398] Figure 38 is a conceptual diagram for explaining a model assuming uniform linear motion. In Figure 38 , (v x , v y ) represents the velocity vector, τ 0 , τ 1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) represents the motion vector corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) represents the motion vector corresponding to the reference picture Ref 1 .

[0399] At this time, it can also be that, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1 , MVy 1 ) are respectively represented as (vxτ 0 , vyτ 0 ) and (-vxτ1 , -vyτ 1 ), the following optical flow equation (2) is adopted.

[0400]

Equation 3

[0401]

[0402] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from a merge list or the like is corrected in pixel units.

[0403] In addition, the motion vector can be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector can be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0404] [Motion Compensation > LIC]

[0405] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.

[0406] Figure 39 is a conceptual diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC processing.

[0407] First, the MV is derived from the encoded reference picture, and the reference image corresponding to the current block is obtained.

[0408] Next, information indicating how the luminance value changes in the reference picture and the current picture is extracted for the current block. This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. Then, the luminance correction parameter is calculated using the information indicating how the luminance value changes.

[0409] The reference image in the reference picture specified by the MV is subjected to a luminance correction process by applying the above luminance correction parameter, and a predicted image for the current block is generated.

[0410] In addition,Figure 39 The shape of the above-mentioned peripheral reference area in [0] is an example, and shapes other than this can also be used.

[0411] In addition, the process of generating a prediction image based on one reference picture has been described here, but the same also applies to the case of generating a prediction image based on multiple reference pictures. It is also possible to generate a prediction image after performing brightness correction processing on the reference images obtained from each reference picture in the same manner as described above.

[0412] As a method for determining whether to adopt LIC processing, for example, there is a method of using lic_flag which is a signal indicating whether to adopt LIC processing. As a specific example, in an encoding device, it is determined whether the current block belongs to an area where a brightness change has occurred. In the case where it belongs to an area where a brightness change has occurred, the value 1 is set as lic_flag and encoding is performed using LIC processing. In the case where it does not belong to an area where a brightness change has occurred, the value 0 is set as lic_flag and encoding is performed without using LIC processing. On the other hand, in a decoding device, it is also possible to decode the lic_flag described in the stream and switch whether to adopt LIC processing according to its value for decoding.

[0413] As another method for determining whether to adopt LIC processing, for example, there is also a method of determining according to whether LIC processing has been adopted in neighboring blocks. As a specific example, in the case where the current block is in the merge mode, it is determined whether the neighboring encoded blocks selected during the derivation of the MV in the merge mode processing have been encoded using LIC processing, and according to the result, whether to adopt LIC processing is switched for encoding. Additionally, in the case of this example, the same processing also applies to the decoding device side.

[0414] Use Figure 39 The form of LIC processing (brightness correction processing) has been described, and hereinafter, its detailed content will be described.

[0415] First, the inter-frame prediction unit 126 derives a motion vector for obtaining a reference image corresponding to the encoding target block from a reference picture which is an encoded picture.

[0416] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values change between the reference picture and the picture to be coded for the block to be coded, using the luminance pixel values of the left and upper adjacent coded peripheral reference regions and the luminance pixel values at the same positions in the reference picture specified by the motion vector, and calculates a luminance correction parameter. For example, the luminance pixel value of a certain pixel in the peripheral reference region within the picture to be coded is set as p0, and the luminance pixel value of the pixel at the same position in the peripheral reference region within the reference picture is set as p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1 + B = p0 for a plurality of pixels within the peripheral reference region as the luminance correction parameter.

[0417] Next, the inter-frame prediction unit 126 generates a prediction picture for the block to be coded by performing a luminance correction process on the reference image within the reference picture specified by the motion vector using the luminance correction parameter. For example, the luminance pixel value within the reference image is set as p2, and the luminance pixel value of the prediction picture after the luminance correction process is set as p3. The inter-frame prediction unit 126 generates the prediction picture after the luminance correction process by calculating A×p2 + B = p3 for each pixel within the reference image.

[0418] In addition, Figure 39 The shape of the peripheral reference region in Figure 39 is an example, and other shapes may also be used. Additionally, a part of the peripheral reference region shown in

[0419] In addition, Figure 39 In the example shown in

[0420] Here, the operations in the coding device 100 have been described, but typically, the operations in the decoding device 200 are the same.

[0421] In addition, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0422] In addition, the LIC process can also be applied in units of sub-blocks. For example, the correction parameters can also be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block within the reference picture specified by the MV of the current sub-block.

[0423] [Prediction control unit]

[0424] The prediction control unit 128 selects one of the intra prediction signal (the signal output from the intra prediction unit 124) and the inter prediction signal (the signal output from the inter prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0425] As Figure 1 shown, in various coding device examples, the prediction control unit 128 can also output the prediction parameters input to the entropy coding unit 110. The entropy coding unit 110 can generate a coded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device. The decoding device can receive and decode the coded bitstream, and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selection of the prediction signal (e.g., the motion vector, the prediction type, or the prediction mode used by the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0426] [Installation example of coding device]

[0427] Figure 40 is a block diagram showing an installation example of the coding device 100. The coding device 100 includes a processor a1 and a memory a2. For example, Figure 1 as shown, multiple components of the coding device 100 are implemented by Figure 40 the shown processor a1 and memory a2.

[0428] The processor a1 is a circuit for information processing and is a circuit accessible to the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding moving images. The processor a1 can also be a processor such as a CPU. In addition, the processor a1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 can also function as Figure 1 multiple components among the multiple components of the coding device 100 as shown, etc.

[0429] The memory a2 is a dedicated or general-purpose memory for storing information used by the processor a1 to encode motion images. The memory a2 can be an electronic circuit or connected to the processor a1. In addition, the memory a2 can also be included in the processor a1. In addition, the memory a2 can also be a collection of multiple electronic circuits. In addition, the memory a2 can be a magnetic disk or an optical disk, or can be expressed as a storage or a recording medium. In addition, the memory a2 can be a non-volatile memory or a volatile memory.

[0430] For example, the memory a2 may store the encoded moving image, or may store a bit string corresponding to the encoded moving image. In addition, the memory a2 may store a program for the processor a1 to encode the moving image.

[0431] In addition, for example, memory a2 can also serve Figure 1 The memory a2 may be used as a component for storing information among the multiple components of the encoding device 100 shown in FIG. Figure 1 The functions of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks and reconstructed pictures, etc.

[0432] In addition, in the encoding device 100, it is not necessary to install Figure 1 All of the multiple constituent elements shown above may not perform all of the multiple processes described above. Figure 1 Part of the plurality of components shown in the figure may be included in other devices, and part of the plurality of processes described above may be executed by other devices.

[0433] [Decoding device]

[0434] Next, a decoding device that can decode a coded signal (coded bit stream) output from, for example, the above-described coding device 100 will be described. Figure 41 2 is a block diagram showing a functional structure of a decoding device 200 according to an embodiment. The decoding device 200 is a moving picture decoding device that decodes a moving picture in units of blocks.

[0435] like Figure 41 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.

[0436] The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a loop filtering unit 212, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220. In addition, the decoding device 200 may also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0437] Hereinafter, after explaining the flow of the overall processing of the decoding device 200, each component included in the decoding device 200 will be described.

[0438] [Overall Flow of Decoding Process]

[0439] Figure 42 It is a flowchart showing an example of the overall decoding process performed by the decoding device 200.

[0440] First, the entropy decoding unit 202 of the decoding device 200 determines the segmentation pattern of a fixed-size block (for example, 128×128 pixels) (step Sp_1). This segmentation pattern is the segmentation pattern selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 on each of the multiple blocks constituting this segmentation pattern.

[0441] That is, the entropy decoding unit 202 decodes the encoded quantization coefficients and prediction parameters of the block to be decoded (also referred to as the current block) (specifically, entropy decoding) (step Sp_2).

[0442] Next, the inverse quantization unit 204 and the inverse transform unit 206 restore multiple prediction residuals (that is, differential blocks) by performing inverse quantization and inverse transform on the multiple quantization coefficients (step Sp_3).

[0443] Next, a prediction processing unit composed of all or part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction signal (also referred to as a prediction block) of the current block (step Sp_4).

[0444] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the differential block to the prediction block (step Sp_5).

[0445] Moreover, when generating this reconstructed image, the loop filtering unit 212 filters this reconstructed image (step Sp_6).

[0446] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeatedly executed.

[0447] As shown in the figure, the processing of steps Sp_1 to Sp_7 is sequentially performed by the decoding device 200. Alternatively, multiple processes among these processes can be performed in parallel, or the order can be changed, etc.

[0448] [Entropy decoding unit]

[0449] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the encoded bitstream to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment (refer to Figure 1 ) and include the prediction parameters. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.

[0450] [Inverse quantization unit]

[0451] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoding target block (hereinafter referred to as the current block) that is the input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0452] [Inverse transform unit]

[0453] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients that are the input from the inverse quantization unit 204.

[0454] For example, when the information read from the encoded bitstream indicates the use of EMT or AMT (for example, the AMT flag is true), the inverse transform unit 206 performs inverse transform on the transform coefficients of the current block based on the information indicating the transform type read.

[0455] In addition, for example, when the information read from the encoded bitstream indicates the use of NSST, the inverse transform unit 206 applies inverse re-transformation to the transform coefficients.

[0456] [Addition Unit]

[0457] The addition unit 208 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 206, to the prediction sample, which is the input from the predictive control unit 220. Further, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0458] [Block Memory]

[0459] The block memory 210 is a storage unit for storing blocks within a picture to be decoded (hereinafter referred to as the current picture), which are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the addition unit 208.

[0460] [Loop Filter Unit]

[0461] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208, and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.

[0462] When the information indicating the ON / OFF of ALF read from the encoded bitstream indicates ON of ALF, one filter is selected from a plurality of filters based on the direction and activity of the locality-based gradient, and the selected filter is applied to the reconstructed block.

[0463] [Frame Memory]

[0464] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0465] [Prediction Processing Unit (Intra Prediction Unit / Inter Prediction Unit / Predictive Control Unit)]

[0466] Figure 43 is a flowchart showing an example of the processing performed by the prediction processing unit of the decoding apparatus 200. Further, the prediction processing unit is composed of all or a part of the constituent elements of the intra prediction unit 216, the inter prediction unit 218, and the predictive control unit 220.

[0467] The prediction processing unit generates a prediction image of the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. Further, in the prediction signal, for example, there are an intra prediction signal and an inter prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using the reconstructed image that has already been obtained by performing generation of a prediction block, generation of a differential block, generation of a coefficient block, restoration of a differential block, and generation of a decoded image block.

[0468] The reconstructed image can be, for example, an image referring to a reference picture, or an image including the current block, i.e., an image of the decoded blocks within the current picture. The decoded blocks within the current picture are, for example, the neighboring blocks of the current block.

[0469] Figure 44 It is a flowchart showing another example of the processing performed by the prediction processing unit of the decoding device 200.

[0470] The prediction processing unit determines the method or mode for generating the prediction image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.

[0471] In the case where it is determined that the first method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the first method (step Sr_2a). In addition, in the case where it is determined that the second method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the second method (step Sr_2b). In addition, in the case where it is determined that the third method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the third method (step Sr_2c).

[0472] The first method, the second method, and the third method are different methods for generating the prediction image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.

[0473] [Intra prediction unit]

[0474] The intra prediction unit 216 performs intra prediction with reference to the blocks within the current picture stored in the block memory 210 based on the intra prediction mode read from the encoded bitstream, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0475] In addition, in the case where the intra prediction mode of referring to the luminance block is selected for the intra prediction of the chrominance difference block, the intra prediction unit 216 can also predict the chrominance difference component of the current block based on the luminance component of the current block.

[0476] In addition, in the case where the information read from the encoded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0477] [Inter prediction unit]

[0478] The inter-frame prediction unit 218 predicts the current block by referring to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 218 performs motion compensation using the motion information (e.g., motion vector) read from the encoded bitstream (e.g., the prediction parameters output from the entropy decoding unit 202), thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.

[0479] When the information read from the encoded bitstream indicates the use of the OBMC mode, the inter-frame prediction unit 218 generates an inter-frame prediction signal using not only the motion information of the current block obtained through motion estimation but also the motion information of adjacent blocks.

[0480] In addition, when the information read from the encoded bitstream indicates the use of the FRUC mode, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the encoded stream, thereby deriving the motion information. And the inter-frame prediction unit 218 uses the derived motion information to perform motion compensation (prediction).

[0481] In addition, when the BIO mode is adopted, the inter-frame prediction unit 218 derives the motion vector based on a model assuming uniform linear motion. In addition, when the information read from the encoded bitstream indicates the use of the affine motion compensation prediction mode, the inter-frame prediction unit 218 derives the motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.

[0482] [MV Derivation > Normal Inter-frame Mode]

[0483] When the information read from the encoded bitstream indicates the application of the normal inter-frame mode, the inter-frame prediction unit 218 derives the MV based on the information read from the encoded bitstream, and uses the MV to perform motion compensation (prediction).

[0484] Figure 45 It is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode in the decoding device 200.

[0485] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. The inter-frame prediction unit 218 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Ss_1). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0486] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the multiple candidate MVs obtained in step Ss_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order (step Ss_2). Additionally, this priority order may also be determined in advance for each of the N prediction MV candidates.

[0487] Next, the inter-frame prediction unit 218 decodes the prediction motion vector selection information from the input stream (i.e., the encoded bitstream), and uses the decoded prediction motion vector selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Ss_3).

[0488] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction motion vector (step Ss_4).

[0489] Finally, the inter-frame prediction unit 218 generates the predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).

[0490] [Prediction control unit]

[0491] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device side.

[0492] [Installation example of decoding device]

[0493] Figure 46 is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 41 the multiple components of the decoding device 200 shown are installed by Figure 46 the processor b1 and the memory b2 shown.

[0494] The processor b1 is a circuit for performing information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit for decoding the encoded moving image (i.e., the encoded bitstream). The processor b1 may also be a processor such as a CPU. Additionally, the processor b1 may also be an aggregate of multiple electronic circuits. Additionally, for example, the processor b1 may also function asFigure 41 Functions of multiple components among the multiple components of the decoding device 200 shown, etc.

[0495] Memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode the encoded bitstream. Memory b2 can be an electronic circuit or connected to the processor b1. Additionally, memory b2 can be included in the processor b1. Moreover, memory b2 can be an aggregate of multiple electronic circuits. Also, memory b2 can be a magnetic disk or optical disk, etc., and can be represented as a storage or recording medium, etc. Further, memory b2 can be either a non-volatile memory or a volatile memory.

[0496] For example, memory b2 can store a moving image or an encoded bitstream. Additionally, a program for the processor b1 to decode the encoded bitstream can also be stored in memory b2.

[0497] Additionally, for example, memory b2 can serve as Figure 41 the component for storing information among the multiple components of the decoding device 200 shown, etc. Specifically, memory b2 can serve as Figure 41 the block memory 210 and frame memory 214 shown. More specifically, memory b2 can store the reconstructed blocks and reconstructed pictures, etc.

[0498] Additionally, in the decoding device 200, not all of the multiple components shown, etc. may be installed, nor may all of the above-mentioned multiple processes be performed. Figure 41 etc. Figure 41 A part of the multiple components shown, etc. can be included in other devices, or a part of the above-mentioned multiple processes can be performed by other devices.

[0499] [Definitions of each term]

[0500] As an example, each term can be defined as follows.

[0501] A picture is an arrangement of multiple luminance samples in a monochrome format, or an arrangement of multiple luminance samples and two corresponding arrangements of multiple chrominance samples in a color format of 4:2:0, 4:2:2, and 4:4:4. A picture can be a frame or a field.

[0502] A frame is a composition of a top field that generates multiple sample lines 0, 2, 4,... and a bottom field that generates multiple sample lines 1, 3, 5,....

[0503] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same access unit.

[0504] A tile is a rectangular area of multiple coding tree blocks within a specific tile column and a specific tile row in the picture. A tile can still apply a loop filter that spans the tile's edge, or it can be a rectangular area of a frame that is intended to be decoded and encoded independently.

[0505] A block is an MxN (N rows and M columns) arrangement of multiple samples, or an MxN arrangement of multiple transform coefficients. A block can also be a square or rectangular area of multiple pixels composed of multiple matrices of 1 luminance and 2 chrominance differences.

[0506] A CTU (Coding Tree Unit) can be a coding tree block of multiple luminance samples of a picture with an arrangement of 3 samples, or 2 corresponding coding tree blocks of multiple chrominance samples. Alternatively, a CTU can also be a coding tree block of any number of samples in a monochrome picture and a picture encoded using syntactic constructs used in the encoding of 3 separate color planes and multiple samples.

[0507] A superblock constitutes 1 or 2 mode information blocks, or it can also be a 64×64 pixel square block that is recursively divided into 4 32×32 blocks and can be further divided.

[0508] [First form]

[0509] Figure 47 It is a flowchart of the decoding process of the first form. First, the decoding device derives one or more parameters from the head of the bitstream (S101). For example, the decoding device parses (obtains) one or more parameters from the head of the bitstream. Additionally, when there are no one or more parameters in the bitstream, the decoding device uses default values as the values of the one or more parameters. For example, the default value is 0. Furthermore, the decoding device can derive one or more parameters from the head using previously parsed parameters.

[0510] Next, the decoding device generates a second image by filtering the reconstructed samples of the first image (S102).

[0511] Then, the decoding device determines whether the derived parameter is a predetermined value (S103). For example, the predetermined value is an integer value. For example, the predetermined value is 0. In another example, the predetermined value is 2. Additionally, when there are multiple parameters, it can be determined to be yes when one of the multiple parameters is the predetermined value, and no otherwise, or it can be determined to be yes when all of the multiple parameters are the predetermined value, and no otherwise. Also, the same predetermined value can be set for multiple parameters, or different predetermined values can be set for multiple parameters respectively.

[0512] When the parameter is a predetermined value (Yes in S103), the decoding device decodes the third image using the second image (S104). On the other hand, when the parameter is not a predetermined value (No in S103), the decoding device decodes the third image using the first image (S105). After step S104 or S105, the decoding device displays the second image (S106).

[0513] The filtering process can be replaced, for example, with an ALF process, a CCALF (Cross Component Adaptive Loop Filter) process, a DBF process, an SAO process, an LMCS (Luma mapping with chroma scaling) process, or a combination of any of the filtering processes described later. In addition, the order of the LMCS, DBF, SAO, ALF, and CCALF processes can be any order.

[0514] Here, CCALF operates, for example, by applying a linear diamond filter to the luminance channel of each color difference component. For example, the filter coefficients are sent as an APS (Adaptive Parameter Set), scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and is notified by a context-encoded flag received for each sample block. The block size and the CCALF enable flag are received at the slice level of each color difference component. The syntax and semantics of CCALF are provided in the Appendix. In this document, block sizes of 16x16, 32x32, 64x64, and 128x128 (in color difference samples) are supported.

[0515] In addition, LMCS is a mapping process of the luminance signal that scales the color difference signal. To change the dynamic range or scale of the signal before mapping, LMCS divides the range of the signal before mapping (for example, if it is 10 bits, it is 0 - 1023) into 16 regions, and approximates the transformation curve for mapping with 16 straight lines. In the inverse transformation section in block units, it is determined which region the mapped signal MapSample of each pixel belongs to, that is, which straight line parameter is used for inverse transformation. After determining the transformation straight line used for inverse transformation, the parameters of the straight line are calculated, inverse transformation (inverse mapping) is performed, and the pixel value before mapping is calculated, and this pixel value is used for subsequent processing.

[0516] The above parameter can be a flag or signal indicating whether to decode the third image using the first image or using the second image. That is, the above parameter can be a flag or signal indicating whether to refer to the image before the filtering process or the image after the filtering process.

[0517] In this way, the decoding device can switch which of the second image after the application of the filtering process and the first image before the application is used as a reference for subsequent pictures on a picture-by-picture basis. Therefore, in the case where the filtering process is effective for improving subjective image quality but not suitable as a reference image, by using the first image before the application of the filtering process as the reference image, the coding efficiency can be maintained as it is and the subjective image quality can be improved.

[0518] In addition, when the parameter is not the predetermined value (No in S103), the decoding device may also display the first image. In addition, a signal indicating whether to display the first image or the second image when the parameter is not the predetermined value (No in S103) can be encoded by SEI or the like. This signal can be a flag indicating the image to be displayed. When this flag is 0, the first image is displayed. On the other hand, when this flag is 1, the second image is displayed. Additionally, it is also possible to display the first image when the flag is 1 and display the second image when the flag is 0. Furthermore, when this signal does not exist in SEI or the like, a default image can be displayed. This default image is the first image or the second image.

[0519] The filter coefficients required for the filtering process can be encoded in headers such as APS, PPS, SPS, SEI, slice header, tile header, or brick header. Additionally, a brick is, for example, one or more processing units included in a slice. For example, a brick includes one or more CTUs (Coding Tree Units). Additionally, a slice can also be included in a brick.

[0520] When the parameter is not the predetermined value (No in S103), the encoding device can encode the filter coefficients required for the filtering process into SEI or the like. Additionally, when the parameter is not the predetermined value (No in S103), the decoding device can also decode the filter coefficients required for the filtering process from SEI or the like.

[0521] The use of a sharpening filter can also be included in the filtering process. For the entire stream, the same value can be used as the value of the derived parameter. This value can depend on the level of the profile. In this case, the above-mentioned parameter may not exist in the bitstream.

[0522] Figures 48 - 55 It is a diagram showing an example of the structure of the decoding device. As Figure 48As shown, the decoding device includes an entropy decoding unit 301, a block segmentation unit 302, an inverse quantization unit 303, an inverse transformation unit 304, an addition unit 305, an intra prediction unit 306, an inter prediction unit 307, a selection unit 308, an LMCS unit 309, a DBF unit 310, an SAO unit 311, an ALF unit 312, and a selection unit 313.

[0523] The entropy decoding unit 301 generates quantization coefficients and the like by performing entropy decoding on the encoded data (bitstream). The block segmentation unit 302 sets a plurality of blocks obtained by dividing the image. The inverse quantization unit 303 generates transform coefficients by performing inverse quantization on the quantization coefficients. The inverse transformation unit 304 generates a differential block (prediction residue) by performing inverse transformation on the transform coefficients.

[0524] The addition unit 305 generates a decoded block (reconstructed image) by adding the differential block and the prediction block. The intra prediction unit 306 generates an intra prediction block by intra prediction. The inter prediction unit 307 generates an inter prediction block by inter prediction. The selection unit 308 outputs either the intra prediction block or the inter prediction block as the prediction block to the addition unit 305.

[0525] The LMCS unit 309 performs LMCS processing on the decoded block. The DBF unit 310 performs DBF processing on the decoded block. The SAO unit 311 performs SAO processing on the decoded block. The ALF unit 312 performs ALF processing on the decoded block. In addition, the decoding device may include all of the LMCS unit 309, the DBF unit 310, the SAO unit 311, and the ALF unit 312, or may include a part of them.

[0526] The selection unit 313 selects either the decoded block before ALF processing or the decoded block after ALF processing, and outputs the selected decoded block to the inter prediction unit 307. Specifically, the selection unit 313 selects the decoded block after ALF processing when the parameter is a predetermined value (Yes in S103), and selects the decoded block before ALF processing when the parameter is not a predetermined value (No in S103).

[0527] In addition, in Figure 49 In the example shown, the selection unit 313A selects either the decoded block before SAO processing or the decoded block after SAO processing, and outputs the selected decoded block to the inter prediction unit 307. Specifically, the selection unit 313A selects the decoded block after SAO processing when the parameter is a predetermined value (Yes in S103), and selects the decoded block before SAO processing when the parameter is not a predetermined value (No in S103).

[0528] In Figure 50In the example shown, the selection unit 313B selects one of the decoded blocks before DBF processing and the decoded blocks after DBF processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, when the parameter is a predetermined value (Yes in S103), the selection unit 313B selects the decoded block after DBF processing, and when the parameter is not a predetermined value (No in S103), the selection unit 313B selects the decoded block before DBF processing.

[0529] In Figure 51 In the example shown, the selection unit 313C selects one of the decoded blocks before LMCS processing and the decoded blocks after LMCS processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, when the parameter is a predetermined value (Yes in S103), the selection unit 313C selects the decoded block after LMCS processing, and when the parameter is not a predetermined value (No in S103), the selection unit 313C selects the decoded block before LMCS processing.

[0530] In Figure 52 In the example shown, the selection unit 313D selects one of the decoded blocks before LMCS, DBF, SAO, and ALF processing and the decoded blocks after LMCS, DBF, SAO, and ALF processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, when the parameter is a predetermined value (Yes in S103), the selection unit 313D selects the decoded block after LMCS, DBF, SAO, and ALF processing, and when the parameter is not a predetermined value (No in S103), the selection unit 313D selects the decoded block before LMCS, DBF, SAO, and ALF processing.

[0531] In Figure 53 In the example shown, the selection unit 313E selects one of the decoded blocks before DBF, SAO, and ALF processing and the decoded blocks after DBF, SAO, and ALF processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, when the parameter is a predetermined value (Yes in S103), the selection unit 313E selects the decoded block after DBF, SAO, and ALF processing, and when the parameter is not a predetermined value (No in S103), the selection unit 313E selects the decoded block before DBF, SAO, and ALF processing.

[0532] In Figure 54 In the example shown, the selection unit 313F selects one of the decoded blocks before SAO and ALF processing and the decoded blocks after SAO and ALF processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, when the parameter is a predetermined value (Yes in S103), the selection unit 313F selects the decoded block after SAO and ALF processing, and when the parameter is not a predetermined value (No in S103), the selection unit 313F selects the decoded block before SAO and ALF processing.

[0533] Figure 55 The decoding device shown also includes a CCALF unit 314 and an adder 315. The CCALF unit 314 performs CCALF processing on the decoded block. The adder 315 adds the decoded block after ALF processing and the decoded block after CCALF processing.

[0534] In Figure 55 In the example shown, the selection unit 313G selects one of the decoded block before ALF and CCALF processing and the decoded block after ALF and CCALF processing, and outputs the selected decoded block to the inter-frame prediction unit 307. Specifically, the selection unit 313G selects the decoded block after ALF and CCALF processing when the parameter is a predetermined value (Yes in S103), and selects the decoded block before ALF and CCALF processing when the parameter is not a predetermined value (No in S103).

[0535] [Second form]

[0536] Figure 56 FIG. is a flowchart of the decoding process of the second form. First, the decoding device analyzes at least one parameter (S201) for each of a plurality of blocks of the first image. For example, this parameter is a flag.

[0537] Next, the decoding device determines whether a predetermined value is included in at least one parameter (S202). For example, this predetermined value is an integer value. For example, this predetermined value is 0.

[0538] When a predetermined value is included in at least one parameter (Yes in S202), the decoding device generates a second block (S203) by filtering the reconstructed samples of the first block using filtering processing. In addition, a specific example of the filtering process is the same as that of the first form. Next, the decoding device displays the second block (S204).

[0539] On the other hand, when a predetermined value is not included in at least one parameter (No in S202), the decoding device displays the first block (S205). After step S204 or S205, the decoding device decodes the third image using the first image and not using the second block (S206).

[0540] In this way, the decoding device can switch whether to display the image after applying the filtering process for each block area. Therefore, when the filtering process is effective only in a local area for improving the subjective image quality, it can be an appropriate filtering process and can improve the subjective image quality.

[0541] [Third form]

[0542] Figure 57It is a flowchart of the decoding process in the third form. First, the decoding device parses (obtains) the first parameter from the first header (S301). For example, the first header is a slice header. For example, the first parameter is the APS ID for identifying APS.

[0543] Next, the decoding device parses (obtains) multiple second parameters from the second header (S302). For example, the second parameters are the parameters included in multiple APS sets with different APS IDs.

[0544] Next, the decoding device selects a second parameter from the multiple parsed second parameters based on the first parameter (S303). For example, the decoding device selects a set of filtering coefficients within the APS using the APS ID.

[0545] Next, the decoding device generates a second image by filtering the reconstructed samples of the first image using filtering processing and the selected second parameter (S304). Additionally, a specific example of the filtering processing is the same as that in the first form. Next, the decoding device decodes the third image using the first image (S305). Next, the decoding device displays the second image (S306).

[0546] In this way, the decoding device can switch and apply the filtering coefficients of the filtering processing for each slice. Therefore, even for an image with different characteristics in each local region, the filtering processing is appropriately switched and applied, thereby improving the subjective image quality of the entire image.

[0547] [Fourth form]

[0548] Figure 58 It is a flowchart of the decoding process in the fourth form. First, the decoding device parses multiple first parameters from the header (S401). For example, the first parameters are sets of filtering coefficients. Next, the decoding device calculates gradient information (variance parameter) based on the blocks of the reconstructed samples of the first image (S402). For example, the size of the block of the reconstructed samples is 4×4.

[0549] Next, the decoding device selects one first parameter from the multiple first parameters based on the calculated gradient information (S403). For example, the decoding device classifies the block into one of multiple categories using the gradient information, and each category in the multiple categories has a set of unique filtering coefficients. The decoding device selects the filtering coefficient (first parameter) corresponding to the classified category. The gradient information includes the gradient in the horizontal direction, the gradient in the vertical direction, the inclined gradient, or any combination thereof. Additionally, variance information may also be included in the gradient information.

[0550] Next, the decoding device generates a second image by filtering the reconstructed samples of the first image using filtering processing and a selected first parameter (S404). Next, the decoding device decodes a third image using the first image (S405). Next, the decoding device displays the second image (S406).

[0551] In this way, the decoding device can extract the image feature amount for each block and switch and apply the filter coefficient of the filtering processing according to the image feature amount. Therefore, even for an image having different features in each local area, the filtering processing can be appropriately switched and applied, so that the subjective image quality of the entire image can be improved.

[0552] Figures 59 - 62 It is a diagram showing a calculation example of gradient information. Figures 59 - 62 It shows the subsampled positions for calculating the vertical, horizontal, diagonal 0, and diagonal 1 gradients in a 4×4 block size. Figure 59 It is a diagram showing the subsampling position of the vertical gradient (g v ). Figure 60 It is a diagram showing the subsampling position of the horizontal gradient (g h ). Figure 61 It is a diagram showing the subsampling position of the diagonal 0 gradient (g d1 ). Figure 62 It is a diagram showing the subsampling position of the diagonal 1 gradient (g d2 ).

[0553] In addition, the vertical gradient (g v ), horizontal gradient (g h ), diagonal 0 gradient (g d1 ), and diagonal 1 gradient (g d2 ) are calculated using the following (Equation 1) to (Equation 4).

[0554]

Equation 4

[0555]

[0556]

[0557]

[0558]

[0559] As described above, in the present invention, a filter (ALF, SAO, DBF, or LMCS, or a combination thereof) is switched between a post filter and an in-loop filter in video coding. This design can achieve an improvement in the flexibility of customizing the filter and an improvement in the quality of the image.

[0560] In addition, the above-mentioned header may be a VPS, SPS, PPS, APS, SEI, or slice header. The filtering process may be an ALF process, CCALF process, DBF process, SAO process, LMCS process, or a combination of these multiple post-filtering processes. In addition, the above determination (S103 or S202) may also be made based on the profile level.

[0561] In addition, the operation of the decoding device is mainly described here, but the same process may also be performed in the encoding device. In addition, a processing unit similar to the above-mentioned processing unit may also be included in the encoding device. Specifically, as Figure 1 shown, etc., the encoding device performs inverse quantization, inverse transformation, reconstruction (addition of predicted images), loop filtering process, and prediction process (intra prediction and inter prediction) in the same manner as the decoding device. That is, the encoding device decodes the image in the same manner as the decoding device. Moreover, the encoding device subtracts the input image from the generated predicted image and performs transformation, quantization, and entropy encoding processes. That is, the encoding device, like the decoding device, performs an encoding process using the prediction in addition to the decoding process of the prediction of the image before or after using the filtering process (loop filtering). Therefore, the decoding of an image or block in the decoding device may be replaced with the encoding of the image or block, or may be replaced with the encoding and decoding of the image or block. In addition, various information parsed (obtained) from the bitstream (e.g., header) in the decoding device is encoded (stored) into the bitstream (e.g., header) in the encoding device.

[0562] As described above, the encoding device includes a circuit and a memory connected to the circuit. During operation, the circuit encodes information for deriving a parameter into the header of the bitstream, generates a second image by filtering the reconstructed samples of the first image using a filtering process, determines whether the parameter is a predetermined value, and encodes the third image using the second image when the parameter is the predetermined value, and encodes the third image using the first image when the parameter is not the predetermined value.

[0563] For example, the header is an APS (Adaptation Parameter Set). For example, the header is an SPS (Sequence Parameter Set). For example, the header is a PPS (Picture Parameter Set). For example, the header is a slice header or a tile header. For example, the header is a brick header. For example, the header is an SEI (Supplemental Enhancement Information).

[0564] For example, the filtering process includes an adaptive loop filtering process. For example, the filtering process includes a cross-component adaptive loop filtering process. For example, the filtering process includes a deblocking filtering process. For example, the filtering process includes a sample adaptive offset process. For example, the filtering process includes an LMCS (Luma mapping with chroma scaling) process. For example, the filtering process uses a sharpening filter.

[0565] For example, the parameter includes a filtering coefficient of the filtering process.

[0566] Furthermore, the decoding device includes a circuit and a memory connected to the circuit. During operation, the circuit derives a parameter from information included in the head of a bitstream, generates a second image by filtering reconstructed samples of a first image using a filtering process, determines whether the parameter is a predetermined value, decodes a third image using the second image when the parameter is the predetermined value, decodes the third image using the first image when the parameter is not the predetermined value, and displays the second image.

[0567] For example, the head is an APS (Adaptation Parameter Set). For example, the head is an SPS (Sequence Parameter Set). For example, the head is a PPS (Picture Parameter Set). For example, the head is a slice head or a tile head. For example, the head is a brick head. For example, the head is an SEI (Supplemental Enhancement Information).

[0568] For example, the filtering process includes an adaptive loop filtering process. For example, the filtering process includes a cross-component adaptive loop filtering process. For example, the filtering process includes a deblocking filtering process. For example, the filtering process includes a sample adaptive offset process. For example, the filtering process includes an LMCS (Luma mapping with chroma scaling) process. For example, the filtering process uses a sharpening filter.

[0569] For example, in the derivation, the decoding device derives the parameter by parsing the head. For example, in the derivation, the decoding device derives the parameter from the head using a previously parsed parameter. For example, the parameter includes a filtering coefficient of the filtering process.

[0570] In addition, the encoding device includes: a splitting unit that, during operation, receives a picture and splits the picture into a plurality of blocks; a first addition unit that, during operation, receives the plurality of blocks and a plurality of predictions from a prediction control unit from the splitting unit, and subtracts each prediction from the corresponding block to output a plurality of residuals; a transformation unit that, during operation, transforms the plurality of residuals output from the first addition unit to output a plurality of transform coefficients; a quantization unit that, during operation, quantizes the plurality of transform coefficients to generate a plurality of quantized transform coefficients; an entropy encoding unit that, during operation, entropy-encodes the plurality of quantized transform coefficients to generate a bitstream; an inverse quantization and transformation unit that, during operation, inverse-quantizes the plurality of quantized transform coefficients to obtain the plurality of transform coefficients, and inverse-transforms the plurality of transform coefficients to obtain the plurality of residuals; a second addition unit that, during operation, adds the plurality of residuals output from the inverse quantization and transformation unit and the plurality of predictions output from the prediction control unit to reconstruct the plurality of blocks; and the prediction control unit that is connected to an inter-frame prediction unit, an intra-frame prediction unit, and a memory, encodes information for deriving a parameter into the header of the bitstream, generates a second image by filtering the reconstructed samples of the first image using a filtering process, determines whether the parameter is a predetermined value, and encodes a third image using the second image when the parameter is the predetermined value, and encodes the third image using the first image when the parameter is not the predetermined value.

[0571] In addition, the decoding device includes: an entropy decoding unit that, during operation, receives an encoded bitstream and decodes it to obtain a plurality of quantized transform coefficients; an inverse quantization and transformation unit that, during operation, inverse-quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inverse-transforms the plurality of transform coefficients to obtain a plurality of residuals; an addition unit that, during operation, adds the plurality of residuals output from the inverse quantization and transformation unit and the plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; and the prediction control unit that is connected to an inter-frame prediction unit, an intra-frame prediction unit, and a memory, derives a parameter from information included in the header of the bitstream, generates a second image by filtering the reconstructed samples of the first image using a filtering process, determines whether the parameter is a predetermined value, and decodes a third image using the second image when the parameter is the predetermined value, and decodes the third image using the first image when the parameter is not the predetermined value, and displays the second image.

[0572] One or more of the aspects disclosed herein can be implemented in combination with at least a part of other aspects of the present invention. Additionally, a part of the processes described in the flowcharts of one or more of the aspects disclosed herein, a part of the structures of the devices, a part of the syntax, etc. can be combined with other aspects for implementation.

[0573] [Implementation and Application]

[0574] In the above-described embodiments, each functional block or operative block can generally be implemented by an MPU (microprocessing unit), a memory, and the like. In addition, it may be that the processing of each functional block is implemented by a program execution unit such as a processor that reads and executes software (program) recorded in a recording medium such as a ROM. This software can be distributed. This software can also be recorded in various recording media such as a semiconductor memory. In addition, each functional block can be implemented by hardware (a dedicated circuit). Various combinations of hardware and software can be adopted.

[0575] The processing described in each embodiment can be implemented by centralized processing using a single device (system), or can also be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing or distributed processing can be performed.

[0576] The form of the present invention is not limited to the above embodiments, and various modifications can be made, and they are also included in the scope of the form of the present invention.

[0577] Furthermore, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) described in the above embodiments and various systems for implementing such application examples are described here. It may also be that such a system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding / decoding device having both. Regarding other structures of such a system, they can be appropriately changed according to circumstances.

[0578] [Usage Example]

[0579] Figure 63 It is a diagram showing the overall structure of a suitable content supply system ex100 for implementing a content distribution service. The provision of the communication service is divided into desired sizes, and in each unit, base stations ex106, ex107, ex108, ex109, ex110 that are fixed wireless stations in the illustrated example are provided respectively.

[0580] In the content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 can also connect by combining some of the above devices. In various implementations, the devices can be directly or indirectly connected to each other via a telephone network or short-range wireless communication without going through the base stations ex106 to ex110. Moreover, a streaming media server ex103 can also be connected to devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 via the Internet ex101 or the like. In addition, the streaming media server ex103 can also be connected to terminals within a hotspot in an aircraft ex117 via a satellite ex116.

[0581] In addition, a wireless access point or a hotspot or the like can be used instead of the base stations ex106 to ex110. Moreover, the streaming media server ex103 can be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or can be directly connected to the aircraft ex117 without going through the satellite ex116.

[0582] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a mobile phone, or a PHS (Personal Handy-phone System) or the like corresponding to the modes of mobile communication systems known as 2G, 3G, 3.9G, 4G, and those to be known as 5G in the future.

[0583] The home appliances ex114 are a refrigerator or devices included in a household fuel cell cogeneration system or the like.

[0584] In the content supply system ex100, a terminal having a photographing function is connected to the streaming media server ex103 via the base station ex106 or the like, thereby enabling live distribution and the like. In live distribution, terminals (such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, a smart phone ex115, and terminals in the aircraft ex117) can perform the encoding processing described in the above embodiments on still image or moving image content photographed by a user using the terminal, can also multiplex the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and can also send the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.

[0585] On the other hand, the streaming media server ex103 performs streaming distribution of content data sent by a requesting client. The client is a terminal in a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or an airplane ex117 that can decode the data after the above-mentioned encoding process. Each device that receives the distributed data can also perform decoding processing on the received data and reproduce it. That is, each device can also function as an image decoding device according to one aspect of the present invention.

[0586] [Distributed processing]

[0587] In addition, the streaming media server ex103 can also be multiple servers or multiple computers, which perform distributed processing or recording of data and then distribute it. For example, the streaming media server ex103 can be implemented by a CDN (Content Delivery Network), and content distribution is achieved through a network that connects many edge servers distributed around the world to each other. In a CDN, it is possible to dynamically allocate a physically closer edge server according to the client. And by caching and distributing content to this edge server, latency can be reduced. In addition, in the case of several types of errors occurring or when the communication state changes due to an increase in traffic, etc., it is possible to distribute the processing among multiple edge servers, or switch the distribution entity to another edge server, or bypass a faulty part of the network and continue distribution, so high-speed and stable distribution can be achieved.

[0588] In addition, not limited to the distributed processing of the distribution itself, the encoding process of the captured data can be performed by each terminal, or on the server side, or can be shared between them. As an example, usually two processing loops are performed in the encoding process. In the first loop, the complexity or encoding amount of an image in units of frames or scenes is detected. In addition, in the second loop, a process of improving the encoding efficiency while maintaining the image quality is performed. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, it is possible to improve the quality and efficiency of the content while reducing the processing load in each terminal. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.

[0589] As other examples, a camera ex113 or the like extracts feature amounts (features or amounts of features) from an image, compresses data on the feature amounts as metadata, and transmits the data to a server. The server, for example, determines the importance of a target based on the feature amounts and switches quantization precision or the like, and performs compression corresponding to the meaning of the image (or the importance of the content). Feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction during re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) may be performed by a terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed by the server.

[0590] As other examples, in a stadium, a shopping mall, a factory, or the like, there are cases where there are a plurality of video data obtained by a plurality of terminals photographing substantially the same scene. In this case, a plurality of terminals that have performed photographing, and other terminals and servers that are not photographed as needed are used, and distributed processing is performed by respectively allocating encoding processing in units of GOP (Group of Picture), picture units, or tile units obtained by dividing a picture, for example. As a result, it is possible to reduce latency and better achieve real-time performance.

[0591] Since the plurality of video data are of substantially the same scene, the server may also manage and / or instruct so as to refer to the video data photographed by each terminal with each other. In addition, it may be that the server receives the encoded data from each terminal and changes the reference relationship between the plurality of data, or corrects or replaces the picture itself and re-encodes it. As a result, it is possible to generate a stream with improved quality and efficiency of each data.

[0592] Moreover, the server may also perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may change an MPEG-like encoding method to a VP-like (e.g., VP9) method, or may change H.264 to H.265, etc.

[0593] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions use "server" or "terminal" as the main body performing the process, but a part or all of the processes performed by the server may be performed by the terminal, and a part or all of the processes performed by the terminal may be performed by the server. In addition, regarding these, the same applies to the decoding process.

[0594] [3D, Multi-angle]

[0595] There is an increasing trend of combining images or videos of different scenes captured by multiple cameras ex113 and / or terminals such as smartphones ex115 that are roughly synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal can be combined based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the videos are consistent, etc.

[0596] The server not only encodes two-dimensional moving images, but can also automatically or at a user-specified time encode still images based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative position relationship between the shooting terminals, it can not only generate the three-dimensional shape of the scene based on the videos of the same scene captured from different angles in addition to two-dimensional moving images. The server can also encode the three-dimensional data generated from point clouds, etc. separately, or select or reconstruct from the videos captured by multiple terminals based on the results of identifying or tracking people or objects using the three-dimensional data to generate and send the videos to the receiving terminal.

[0597] In this way, the user can not only arbitrarily select each video corresponding to each shooting terminal to view the scene, but also view the content of the video cut from the selected viewing point from the three-dimensional data reconstructed using multiple images or videos. Furthermore, together with the video, sound can also be collected from multiple different angles, and the server multiplexes the sound from a specific angle or space with the corresponding video and sends the multiplexed video and sound.

[0598] In addition, in recent years, content that establishes a correspondence between the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) has also been popularized. In the case of VR images, the server separately produces viewpoint images for the right eye and the left eye, and can either perform encoding that allows reference between the viewpoint videos through Multi-View Coding (MVC) or the like, or encode them as different streams without mutual reference. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.

[0599] In the case of an AR image, it is also possible that the server overlays the virtual object information in the virtual space on the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device acquires or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates the overlay data by smoothly connecting them. Alternatively, it is also possible that the decoding device sends the movement of the user's viewpoint to the server in addition to the delegation of the virtual object information. It is also possible that the server creates the overlay data according to the three-dimensional data held in the server, matches the received movement of the viewpoint, encodes the overlay data, and distributes it to the decoding device. Additionally, typically, the overlay data has an α value representing transmittance in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted through this part. Alternatively, the server can also set the RGB value of a specified value as the background like chroma keying, and generate data with the part other than the target set as the background color. The RGB value of the specified value can also be determined in advance.

[0600] Similarly, the decoding process of the distributed data can be performed by the client (e.g., the terminal), on the server side, or they can share the process. As an example, it is also possible that a certain terminal first sends a reception request to the server, and another terminal receives the content corresponding to the request and performs the decoding process, and sends the decoded signal to the device with a display. By dispersing the process regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. In addition, as another example, it is also possible that a TV, etc., receives large-size image data, and the personal terminal of the viewer decodes and displays a part of the area such as tiles after the picture is segmented. Thus, while making the overall image shared, it is possible to confirm one's own responsible area or the area that one wants to confirm in more detail at hand.

[0601] In a situation where multiple wireless communications at short, medium, or long distances indoors and outdoors can be used, it may be possible to receive content seamlessly using a distribution system standard such as MPEG-DASH. The user can freely select the user's terminal, decoding devices or display devices such as monitors configured indoors and outdoors, and switch in real time. In addition, it is possible to use one's own position information, etc., to switch the decoding terminal and the display terminal and perform decoding. Thus, it is also possible to map and display information on a part of the wall or ground of the building next to the device that can be displayed during the user's movement to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed by the receiving terminal in a short time, or the encoded data being replicated in the edge server of the content distribution service.

[0602] [Scalable Encoding]

[0603] Regarding the switching of content, use Figure 64 The scalable stream compressed and encoded using the moving image encoding method shown in the above-described embodiments will be described. For the server, there may be multiple streams with the same content but different qualities as separate streams, or it may be a structure that switches content by utilizing the characteristics of the temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, the decoding side can freely switch between low-resolution content and high-resolution content for decoding by determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. For example, when a user wants to view the subsequent video that was viewed on a smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream to different layers, so the burden on the server side can be reduced.

[0604] Furthermore, in addition to the structure in which pictures are encoded for each layer as described above and the scalability of the enhancement layer above the base layer is realized, the enhancement layer may include meta-information based on statistical information of the image, etc. It is also possible that the decoding side generates high-quality content by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can improve the signal-to-noise ratio while maintaining and / or expanding the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least squares operation used in the super-resolution process, etc.

[0605] Alternatively, a structure may be provided in which a picture is divided into tiles, etc. according to the meaning of an object, etc. within the image. The decoding side decodes only a part of the area by selecting the tiles to be decoded. Moreover, by saving the attributes of the object (person, car, ball, etc.) and the position within the image (coordinate position within the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including the object. For example, as Figure 65 shown, a data storage structure different from the pixel data, such as the SEI (supplemental enhancement information) message in HEVC, may be used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0606] The meta information can also be saved in units composed of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the moments when a specific person appears in the video, etc. By matching with the information of the picture unit and the time information, it can determine the picture where the target exists and decide the position of the target within the picture.

[0607] [Optimization of Web Pages]

[0608] Figure 66 It is a diagram showing an example of a display screen of a web page in a computer ex111, etc. Figure 67 It is a diagram showing an example of a display screen of a web page in a smart phone ex115, etc. As Figure 66 and Figure 67 shown, there are cases where a web page includes multiple linked images that are links to image content, and their visible ways can also be different depending on the viewing device. When multiple linked images can be seen on the screen, before the user explicitly selects a linked image, or before the linked image approaches near the center of the screen or the whole of the linked image enters the screen, the display device (decoding device) can display the still image or I picture that each content has as a linked image, can also display an image such as a gif animation with multiple still images or I pictures, etc., or can receive only the base layer and decode and display the image.

[0609] When a linked image is selected by the user, the display device, for example, gives the highest priority to the base layer and decodes it. In addition, if there is information indicating that it is scalable content in the HTML that constitutes the web page, the display device can also decode to the enhancement layer. Moreover, in order to ensure real-time performance or when the communication bandwidth is very tight before selection, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display) by decoding and displaying only the forward-referenced pictures (I pictures, P pictures, B pictures that only perform forward reference). Furthermore, the display device can also forcibly ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward references, and as the pictures received over time increase, perform normal decoding.

[0610] [Automatic Driving]

[0611] In addition, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for the automatic driving or driving assistance of a vehicle, the receiving terminal can also receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them by establishing a correspondence. In addition, the meta information can either belong to a layer or be multiplexed only with the image data.

[0612] In this case, since the vehicle, drone, or aircraft including the receiving terminal is moving, the receiving terminal can perform seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting the location information of the receiving terminal. In addition, the receiving terminal can dynamically switch the degree of receiving meta information or updating map information according to the user's selection, the user's situation, and / or the status of the communication band.

[0613] In the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0614] [Distribution of Personal Content]

[0615] In addition, in the content supply system ex100, not only high-quality, long-duration content provided by video distribution operators but also unicast or multicast distribution of low-quality, short-duration content provided by individuals can be performed. It is conceivable that such personal content will increase in the future. In order to make personal content better, the server can also perform encoding processing after editing. This can be achieved, for example, with the following structure.

[0616] During or after shooting in real time or cumulatively, the server performs recognition processing such as shooting error, scene search, meaning analysis, and target detection based on the original image data or the encoded data. And based on the recognition results, the server manually or automatically performs editing such as correcting focus deviation or camera shake, deleting scenes with low importance such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizing the edges of the target, or changing the color tone. Based on the editing results, the server encodes the edited data. In addition, it is known that the viewing rate decreases if the shooting time is too long. The server can also automatically limit not only scenes with low importance as described above but also scenes with little movement based on the image processing results according to the shooting time to make the content within a specific time range. Or the server can generate a summary and encode it based on the result of the meaning analysis of the scene.

[0617] In the original state, personal content may be invaded by content that infringes copyright, the right of the author's personality, or the right of portrait, etc. There may also be inconvenient situations for individuals, such as the sharing range exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or in the home, etc. into out-of-focus images. Moreover, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user can specify the person or background area for which the image is desired to be processed. The server can also perform processing such as replacing the specified area with another image or blurring the focus. If it is a person, the person can be tracked in the moving image and the image of the face part of the person can be replaced.

[0618] The real-time requirement for viewing and listening to personal content with a small data volume is relatively strong. Therefore, although it also depends on the bandwidth, the decoding device can first receive and decode and reproduce the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the reproduction is looped and reproduced more than twice, reproduce a high-quality image including the enhancement layer. In this way, if it is a scalable encoded stream, an experience can be provided that the moving image is rough at the stage of not being selected or just starting to watch, but the stream gradually becomes smooth and the image becomes better. In addition to scalable encoding, the same experience can also be provided when the first rough stream and the second stream encoded with reference to the first moving image form one stream.

[0619] [Other implementation application examples]

[0620] In addition, these encoding or decoding processes are usually processed in the LSIex500 possessed by each terminal. The LSI (large-scale integration circuitry) ex500 (refer to Figure 63 ) can be either a single chip or a structure composed of multiple chips. In addition, software for encoding or decoding moving images can also be loaded into a certain recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, when the smart phone ex115 is equipped with a camera, the moving image data obtained by this camera can also be sent. The moving image data at this time can also be data encoded by the LSIex500 possessed by the smart phone ex115.

[0621] In addition, the LSIex500 may also be configured to download and activate application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. If the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal may also download a codec or application software and then acquire and reproduce the content.

[0622] In addition, not limited to the content supply system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can also be incorporated into a digital broadcast system. Since the broadcast wave is used to carry multiplexed data in which video and audio are multiplexed by using a satellite or the like and transmitted and received, there is a difference in suitability for multicast compared to the structure of the content supply system ex100 which is easy for unicast, but the encoding process and the decoding process can be applied in the same way.

[0623] [Hardware Structure]

[0624] Figure 68 is a diagram showing in further detail Figure 63 the smart phone ex115 shown in the figure. In addition, Figure 69 is a diagram showing a structural example of the smart phone ex115. The smart phone ex115 has an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of photographing video and still images, and a display unit ex458 for displaying the video photographed by the camera unit ex465 and decoding the data such as the video received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the photographed video or still images, the recorded sound, the received video or still images, the encoded or decoded data of e-mail, etc., or a slot unit ex464 as an interface unit with the SIM ex468 for identifying the user and performing authentication for accessing various data represented by the network. In addition, an external memory may be used instead of the memory unit ex467.

[0625] The main control unit ex460, which can comprehensively control the display unit ex458, the operation unit ex466, etc., is interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 in synchronization via the bus ex470.

[0626] If the power key is turned on by the user's operation, the power supply circuit unit ex461 starts the smart phone ex115 to an operable state and supplies power to each unit from the battery pack.

[0627] Based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc., the smart phone ex115 performs processes such as calls and data communication. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454, spectrum spreading processing is performed by the modulation / demodulation unit ex452, and digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, and the signal of the result is transmitted via the antenna ex450. In addition, the received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, spectrum inverse spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog audio signal by the audio signal processing unit ex454, it is output from the audio output unit ex457. During data communication, text, still images, or video data can be sent under the control of the main control unit ex460 via the operation input control unit ex462 based on operations of the operation unit ex466 of the main body unit, etc. The same transmission and reception processing is performed. In the data communication mode, when sending video, still images, or video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of the camera unit ex465 shooting video or still images, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a specified manner, and modulation processing and conversion processing are performed by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and it is transmitted via the antenna ex450. The specified manner can also be determined in advance.

[0628] In the case of receiving an image attached to an email or a chat tool, or an image linked on a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method described in the above respective embodiments, and displays the image or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs the sound from the audio output unit ex457. Since real-time streaming media is becoming increasingly popular, depending on the user's situation, there may also be cases where the reproduction of sound is inappropriate in society. Therefore, it may also be a structure in which, as an initial value, it is preferable not to reproduce the audio signal but only reproduce the video data, and the sound is reproduced synchronously only when the user performs an operation such as clicking on the video data.

[0629] In addition, here, the smart phone ex115 has been described as an example, but as the terminal, other installation forms such as a transmitting terminal having only an encoder and a receiving terminal having only a decoder can be considered in addition to the transceiver terminal having both an encoder and a decoder. In the digital broadcast system, it has been described assuming that the multiplexed data in which the audio data is multiplexed in the video data is received and transmitted. However, in the multiplexed data, character data associated with the video, etc. can also be multiplexed in addition to the audio data. In addition, it is also possible to receive or transmit the video data itself instead of the multiplexed data.

[0630] In addition, it has been described assuming that the main control unit ex460 including the CPU controls the encoding or decoding process, but in many cases, various terminals are equipped with a GPU. Therefore, it can also be configured to use the performance of the GPU to process a larger area together through a memory shared by the CPU and the GPU, or a memory that manages addresses in a shared manner. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization are performed not by the CPU but by the GPU together in units of pictures, etc.

[0631] Industrial applicability

[0632] The present invention can be applied to, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conferencing system, or an endoscope, etc.

[0633] Explanation of reference numerals

[0634] 100 Encoding device

[0635] 102 Splitting section

[0636] 104 Subtraction section

[0637] 106 Transformation section

[0638] 108 Quantization section

[0639] 110 Entropy encoding section

[0640] 112, 204 Inverse quantization section

[0641] 114, 206 Inverse transformation section

[0642] 116, 208 Addition section

[0643] 118, 210 Block memory

[0644] 120, 212 Loop filtering section

[0645] 122, 214 Frame memory

[0646] 124, 216 Intra prediction section

[0647] 126, 218 Inter prediction section

[0648] 128, 220 Prediction control section

[0649] 200 Decoding device

[0650] 202 Entropy decoding section

[0651] 301 Entropy decoding section

[0652] 302 Block splitting section

[0653] 303 Inverse quantization section

[0654] 304 Inverse transformation section

[0655] 305 Addition section

[0656] 306 Intra prediction section

[0657] 307 Inter prediction section

[0658] 308 Selection section

[0659] 309 LMCS Department

[0660] 310 DBF Department

[0661] 311 SAO Department

[0662] 312 ALF Department

[0663] 313, 313A, 313B, 313C, 313D, 313E, 313F, 313G Selection Unit

[0664] 314 CCALF Department

[0665] 315 Addition Unit

[0666] 1201 Boundary Determination Unit

[0667] 1202, 1204, 1206 Switches

[0668] 1203 Filtering Determination Unit

[0669] 1205 Filtering Processing Unit

[0670] 1207 Filtering Characteristic Determination Unit

[0671] 1208 Processing Determination Unit

[0672] a1, b1 Processors

[0673] a2, b2 Memories

Claims

1. A decoding device, wherein, it comprises: a circuit; and a memory connected to the circuit; during operation, the circuit derives a parameter from information included in the head of a bitstream, generates a second image by applying a filtering process to a plurality of reconstructed samples in a first image, the filtering process including applying an adaptive loop filtering process to the plurality of reconstructed samples in the first image, determines whether the parameter has a predetermined value, generates an inter prediction image using one of the first image and the second image, and uses the first image to generate the inter prediction image when the parameter does not have the predetermined value, and uses the second image to generate the inter prediction image when the parameter has the predetermined value, decodes a third image by adding a differential image to the inter prediction image, and after decoding the third image by adding the differential image to the inter prediction image generated using one of the first image and the second image, outputs the second image generated by applying the filtering process to the plurality of reconstructed samples in the first image. Regardless of whether the inter prediction image is generated using the first image or the second image, the second image is output after decoding the third image.

2. The decoding device according to claim 1, wherein, the head is an Adaptive Parameter Set (APS).

3. The decoding device according to claim 1, wherein, the head is a Sequence Parameter Set (SPS).

4. The decoding device according to claim 1, wherein, the head is a Picture Parameter Set (PPS).

5. The decoding device according to claim 1, wherein, the head is one of a slice head and a tile head.

6. The decoding device according to claim 1, wherein, the head is a patch head.

7. The decoding device according to claim 1, wherein, the head is Supplementary Enhancement Information (SEI).

8. The decoding device according to claim 1, wherein, the filtering process includes a deblocking filtering process.

9. The decoding device according to claim 1, wherein, the filtering process includes a Sample Adaptive Offset (SAO) process.

10. The decoding device according to claim 1, wherein, the filtering process includes a Luminance Mapping and Chrominance Scaling (LMCS) process.

11. The decoding device according to claim 1, wherein, the filtering process uses a sharpening filter.

12. A decoding method, wherein, a parameter is derived from information included in the head of a bitstream, a second image is generated by applying a filtering process to a plurality of reconstructed samples in a first image, the filtering process including applying an adaptive loop filtering process to the plurality of reconstructed samples in the first image, it is determined whether the parameter has a predetermined value, Generate an inter - frame prediction image using one of the first image and the second image. When the parameter does not have the predetermined value, generate the inter - frame prediction image using the first image. When the parameter has the predetermined value, generate the inter - frame prediction image using the second image. Decode a third image by adding the differential image to the inter - frame prediction image, and After decoding the third image by adding the differential image to the inter - frame prediction image generated using one of the first image and the second image, output the second image generated by applying the filtering process to the multiple reconstructed samples in the first image. Regardless of whether the inter - frame prediction image is generated using the first image or the second image, output the second image after decoding the third image.

Citation Information

Patent Citations

  • Image coding method and image decoding method

    CN1493157A