Encoding Device, Decoding Device, and Storage Medium

By combining the set of block segmentation modes and limiting the use of certain segmentation modes, the problem of image compression efficiency reduction caused by the increase in signaling overhead of block segmentation information is solved, and more efficient image compression is achieved.

CN114630117BActive Publication Date: 2025-07-08PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210418173.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-20
Filing Date
2019-05-09
Publication Date
2025-07-08
Estimated Expiration
2039-05-09

AI Technical Summary

Technical Problem

In the prior art, the signaling overhead of block segmentation information increases with the increase in the segmentation depth, resulting in a decrease in image compression efficiency.

Method used

An encoding device and a decoding device are adopted to perform block segmentation by combining a block segmentation mode set, define the segmentation type, and use the first and second block segmentation modes to identify the partition direction and number of blocks in the horizontal or vertical direction, limiting the use of certain segmentation modes to reduce signaling overhead.

Benefits of technology

It improves the encoding and compression efficiency of block segmentation information, reduces signaling overhead, and improves the image compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114630117B_ABST
    Figure CN114630117B_ABST
Patent Text Reader

Abstract

The present invention provides an encoding device, a decoding device, and a storage medium. The encoding device uses a set of block partitioning patterns obtained by combining one or more block partitioning patterns to partition a plurality of blocks; wherein the first block partitioning pattern defines a partitioning direction and a partitioning number for partitioning a first block, and the second block partitioning pattern defines a partitioning direction and a partitioning number for partitioning a second block that is one of the blocks obtained after partitioning the first block; the parameter for identifying the second block partitioning pattern includes a first flag indicating in which direction, horizontal or vertical, the block is partitioned; when the partitioning number of the first block partitioning pattern is 3 and the second block is the central block, when the second block partitioning pattern indicated by the first flag is the same as the partitioning direction of the first block partitioning pattern, the second block partitioning pattern includes a block partitioning pattern with a partitioning number of 3 and does not include a block partitioning pattern with a partitioning number of 2, and when the partitioning directions are different, the second block partitioning pattern includes block partitioning patterns with partitioning numbers of 2 and 3.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of a patent application for an invention titled "Encoding Device, Decoding Device, Encoding Method, Decoding Method, and Image Compression Program" with an application date of May 9, 2019, an application number of 201980033751.4. Technical Field

[0002] The present disclosure relates to methods and apparatuses for encoding and decoding video and images using block partitioning. Background Art

[0003] In conventional image and video encoding methods, an image is generally divided into blocks, and encoding and decoding processes are performed at the block level. In recent video standard development, in addition to typical sizes of 8×8 or 16×16, encoding and decoding processes can be performed with various block sizes. For image encoding and decoding processes, a range of sizes from 4×4 to 256×256 can be used.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] In order to represent a range of sizes from 4×4 to 256×256, block partitioning information such as block partitioning patterns (e.g., quadtree, binary tree, and ternary tree) and partitioning flags (e.g., split flag) is determined and signaled for use with the blocks. The overhead of this signaling increases as the depth of the partitioning increases. Moreover, the increased overhead degrades the video compression efficiency.

[0009] Accordingly, an encoding device according to one aspect of the present disclosure provides an encoding device and the like that can improve the compression efficiency in the encoding of block partitioning information.

[0010] Means for Solving the Problems

[0011] An encoding device according to an aspect of the present disclosure encodes an image, and includes: a processor; and a memory; the processor performs the following processing: using a set of block partitioning patterns obtained by combining one or more block partitioning patterns, partitioning the image read from the memory into a plurality of blocks, the block partitioning patterns defining partitioning types; and encoding the plurality of blocks; the set of block partitioning patterns is composed of a first block partitioning pattern and a second block partitioning pattern, the first block partitioning pattern defines a partitioning direction and a partitioning number for partitioning a first block, the second block partitioning pattern defines a partitioning direction and a partitioning number for partitioning a second block, which is one of the blocks obtained after partitioning the first block; a parameter for identifying the second block partitioning pattern includes a first flag indicating in which direction, among the horizontal direction and the vertical direction, the block is partitioned; when the partitioning number of the first block partitioning pattern is 3, the second block is the central block among the blocks obtained after partitioning the first block, and the partitioning direction of the second block partitioning pattern indicated by the first flag is the same as the partitioning direction of the first block partitioning pattern, the second block partitioning pattern includes a block partitioning pattern with a partitioning number of 3 and does not include a block partitioning pattern with a partitioning number of 2; when the partitioning number of the first block partitioning pattern is 3, the second block is the central block among the blocks obtained after partitioning the first block, and the partitioning direction of the second block partitioning pattern indicated by the first flag is different from the partitioning direction of the first block partitioning pattern, the second block partitioning pattern includes a block partitioning pattern with a partitioning number of 3 and a block partitioning pattern with a partitioning number of 2.

[0012] A decoding device according to an aspect of the present disclosure decodes an encoded signal, and includes: a processor; and a memory; the processor performs the following processes: using a set of block segmentation patterns obtained by combining one or more block segmentation patterns, the encoded signal read from the memory is segmented into a plurality of blocks, the block segmentation patterns define segmentation types; and decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, the second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block which is one of the blocks obtained after segmentation of the first block; a parameter for identifying the second block segmentation pattern includes a first flag indicating in which direction of the horizontal direction and the vertical direction the block is segmented; when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after segmentation of the first block, and the segmentation direction of the second block segmentation pattern indicated by the first flag is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern includes a block segmentation pattern with a segmentation number of 3 and does not include a block segmentation pattern with a segmentation number of 2; when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after segmentation of the first block, and the segmentation direction of the second block segmentation pattern indicated by the first flag is different from the segmentation direction of the first block segmentation pattern, the second block segmentation pattern includes a block segmentation pattern with a segmentation number of 3 and a block segmentation pattern with a segmentation number of 2.

[0013] A non - transitory storage medium according to one aspect of the present disclosure is a non - transitory storage medium that stores a bitstream and is readable by a computer. The bitstream includes information for causing a computer that receives the bitstream to perform a decoding process. The information is for causing the computer to perform the following processes: using a set of block partitioning patterns obtained by combining one or more block partitioning patterns, dividing an encoded signal read from a memory into a plurality of blocks, where the block partitioning patterns define a partitioning type; and decoding the plurality of blocks. The set of block partitioning patterns consists of a first block partitioning pattern and a second block partitioning pattern. The first block partitioning pattern defines a partitioning direction and a number of partitions for partitioning a first block. The second block partitioning pattern defines a partitioning direction and a number of partitions for partitioning a second block, which is one of the blocks obtained after partitioning the first block. A parameter for identifying the second block partitioning pattern includes a first flag indicating in which direction (horizontal or vertical) the block is partitioned. When the number of partitions in the first block partitioning pattern is 3, the second block is the central block among the blocks obtained after partitioning the first block, and the partitioning direction of the second block partitioning pattern indicated by the first flag is the same as the partitioning direction of the first block partitioning pattern, the second block partitioning pattern includes a block partitioning pattern with a number of partitions of 3 and does not include a block partitioning pattern with a number of partitions of 2. When the number of partitions in the first block partitioning pattern is 3, the second block is the central block among the blocks obtained after partitioning the first block, and the partitioning direction of the second block partitioning pattern indicated by the first flag is different from the partitioning direction of the first block partitioning pattern, the second block partitioning pattern includes a block partitioning pattern with a number of partitions of 3 and a block partitioning pattern with a number of partitions of 2.

[0014] In addition, these inclusive or specific forms can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer - readable CD - ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0015] Advantages of the Invention

[0016] According to the present invention, it is possible to improve the compression efficiency in the encoding of block partitioning information. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a block diagram showing the functional structure of an encoding apparatus according to Embodiment 1.

[0018] Figure 2 It is a diagram showing an example of block partitioning in Embodiment 1.

[0019] Figure 3Is a table showing transform basis functions corresponding to respective transform types.

[0020] Figure 4A Is a diagram showing an example of the shape of the filter used in ALF.

[0021] Figure 4B Is a diagram showing another example of the shape of the filter used in ALF.

[0022] Figure 4C Is a diagram showing another example of the shape of the filter used in ALF.

[0023] Figure 5A Is a diagram showing 67 intra prediction modes of intra prediction.

[0024] Figure 5B Is a flowchart for explaining an outline of a predicted image correction process based on OBMC processing.

[0025] Figure 5C Is a conceptual diagram for explaining an outline of a predicted image correction process based on OBMC processing.

[0026] Figure 5D Is a diagram showing an example of FRUC.

[0027] Figure 6 Is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0028] Figure 7 Is a diagram for explaining pattern matching (template matching) between a template within the current picture and a block within a reference picture.

[0029] Figure 8 Is a diagram for explaining a model assuming uniform linear motion.

[0030] Figure 9A Is a diagram for explaining the derivation of a motion vector in sub-block units based on motion vectors of multiple adjacent blocks.

[0031] Figure 9B Is a diagram for explaining an outline of a motion vector derivation process based on a merge mode.

[0032] Figure 9C Is a conceptual diagram for explaining an outline of DMVR processing.

[0033] Figure 9D Is a diagram for explaining an outline of a predicted image generation method employing a luminance correction process based on LIC processing.

[0034] Figure 10 Is a block diagram showing the functional configuration of a decoding device according to Embodiment 1.

[0035] Figure 11 It is a flowchart showing the video encoding process related to Embodiment 2.

[0036] Figure 12 It is a flowchart showing the video decoding process related to Embodiment 2.

[0037] Figure 13 It is a flowchart showing the video encoding process related to Embodiment 3.

[0038] Figure 14 It is a flowchart showing the video decoding process related to Embodiment 3.

[0039] Figure 15 It is a block diagram showing the structure of the video / image encoding apparatus related to Embodiment 2 or 3.

[0040] Figure 16 It is a block diagram showing the structure of the video / image decoding apparatus related to Embodiment 2 or 3.

[0041] Figure 17 It is a diagram showing an example of the possible positions of the first parameter in the compressed video stream in Embodiment 2 or 3.

[0042] Figure 18 It is a diagram showing an example of the possible positions of the second parameter in the compressed video stream in Embodiment 2 or 3.

[0043] Figure 19 It is a diagram showing an example of the second parameter following the first parameter in Embodiment 2 or 3.

[0044] Figure 20 It is a diagram showing an example of not selecting the second block mode for the division of a 2N×N pixel block as shown in step (2c) in Embodiment 2.

[0045] Figure 21 It is a diagram showing an example of not selecting the second block mode for the division of an N×2N pixel block as shown in step (2c) in Embodiment 2.

[0046] Figure 22 It is a diagram showing an example of not selecting the second block mode for the division of an N×N pixel block as shown in step (2c) in Embodiment 2.

[0047] Figure 23 It is a diagram showing an example of not selecting the second block mode for the division of an N×N pixel block as shown in step (2C) in Embodiment 2.

[0048] Figure 24This is a diagram showing an example of dividing a 2N×N pixel block using the block pattern selected when not selecting the second block pattern as shown in step (3) in Embodiment 2.

[0049] Figure 25 This is a diagram showing an example of dividing an N×2N pixel block using the block pattern selected when not selecting the second block pattern as shown in step (3) in Embodiment 2.

[0050] Figure 26 This is a diagram showing an example of dividing an N×N pixel block using the block pattern selected when not selecting the second block pattern as shown in step (3) in Embodiment 2.

[0051] Figure 27 This is a diagram showing an example of dividing an N×N pixel block using the block pattern selected when not selecting the second block pattern as shown in step (3) in Embodiment 2.

[0052] Figure 28 This is a diagram showing an example of a block pattern used to divide an N×N pixel block in Embodiment 2. Figure 28 Diagrams (a) to (h) show mutually different block patterns.

[0053] Figure 29 This is a diagram showing an example of a block type and a block direction used to divide an N×N pixel block in Embodiment 3. (1), (2), (3), and (4) are different block types, (1a), (2a), (3a), and (4a) are block patterns with different block types in the vertical block direction, and (1b), (2b), (3b), and (4b) are block patterns with different block types in the horizontal block direction.

[0054] Figure 30 This is a diagram showing the advantages brought about by encoding the block type before the block direction compared to encoding the block direction before the block type in Embodiment 3.

[0055] Figure 31A This is a diagram showing an example of dividing a block into sub - blocks using a set of block patterns that use fewer binary (bin) numbers in the encoding of the block pattern.

[0056] Figure 31B This is a diagram showing an example of dividing a block into sub - blocks using a set of block patterns that use fewer binary numbers in the encoding of the block pattern.

[0057] Figure 32A This is a diagram showing an example of dividing a block into sub - blocks using the block pattern set that first appears in a specified order of multiple block pattern sets.

[0058] Figure 32B This is a diagram showing an example of dividing a block into sub - blocks by using the block - division pattern set that first appears in a specified order among multiple block - division pattern sets.

[0059] Figure 32C This is a diagram showing an example of dividing a block into sub - blocks by using the block - division pattern set that first appears in a specified order among multiple block - division pattern sets.

[0060] Figure 33 This is an overall structural diagram of a content supply system that implements a content distribution service.

[0061] Figure 34 This is a diagram showing an example of an encoding structure in scalable coding.

[0062] Figure 35 This is a diagram showing an example of an encoding structure in scalable coding.

[0063] Figure 36 This is a diagram showing an example of a display screen of a web page.

[0064] Figure 37 This is a diagram showing an example of a display screen of a web page.

[0065] Figure 38 This is a diagram showing an example of a smart phone.

[0066] Figure 39 This is a block diagram showing an example of the structure of a smart phone.

[0067] Figure 40 This is a diagram showing an example of the constraints of a block - division pattern that divides a rectangular block into 3 sub - blocks.

[0068] Figure 41 This is a diagram showing an example of the constraints of a block - division pattern that divides a block into 2 sub - blocks.

[0069] Figure 42 This is a diagram showing an example of the constraints of a block - division pattern that divides a square block into 3 sub - blocks.

[0070] Figure 43 This is a diagram showing an example of the constraints of a block - division pattern that divides a rectangular block into 2 sub - blocks.

[0071] Figure 44 This is a diagram showing an example of the constraints based on the division direction of a block - division pattern that divides a non - rectangular block into 2 sub - blocks.

[0072] Figure 45 This is a diagram showing an example of an effective division direction for dividing a non - rectangular block into 2 sub - blocks. Detailed implementation mode

[0073] Hereinafter, the implementation mode will be specifically described with reference to the drawings.

[0074] In addition, all the implementation modes described below represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, configurations and connection forms of the constituent elements, steps, and the order of steps shown in the following implementation modes are examples and do not limit the meaning of the claims. In addition, among the constituent elements of the following implementation modes, the constituent elements not described in the independent claims representing the most general concept are described as arbitrary constituent elements.

[0075] (Embodiment 1)

[0076] First, as an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention described below, the outline of Embodiment 1 will be described. However, Embodiment 1 is only an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention, and the processing and / or structure described in each form of the present invention can also be implemented in encoding devices and decoding devices different from Embodiment 1.

[0077] When applying the processing and / or structure described in each form of the present invention to Embodiment 1, for example, one of the following can be performed.

[0078] (1) For the encoding device or decoding device of Embodiment 1, replace the constituent elements corresponding to the constituent elements described in each form of the present invention among the multiple constituent elements constituting the encoding device or decoding device with the constituent elements described in each form of the present invention;

[0079] (2) For the encoding device or decoding device of Embodiment 1, after arbitrarily changing the addition, replacement, deletion, etc. of the functions or processes implemented on a part of the constituent elements among the multiple constituent elements constituting the encoding device or decoding device, replace the constituent elements corresponding to the constituent elements described in each form of the present invention with the constituent elements described in each form of the present invention;

[0080] (3) For the method implemented by the encoding device or decoding device of Embodiment 1, add processing, and / or arbitrarily change the replacement, deletion, etc. of a part of the multiple processes included in the method, and then replace the process corresponding to the process described in each form of the present invention with the process described in each form of the present invention;

[0081] (4) Combine a part of the constituent elements of the encoding device or decoding device constituting Embodiment 1 with the constituent elements described in each aspect of the present invention, the constituent elements having a part of the functions of the constituent elements described in each aspect of the present invention, or the constituent elements implementing a part of the processes implemented by the constituent elements described in each aspect of the present invention, and implement them;

[0082] (5) Combine a constituent element having a part of the functions of a part of the constituent elements of the encoding device or decoding device constituting Embodiment 1, or a constituent element implementing a part of the processes implemented by a part of the constituent elements of the encoding device or decoding device constituting Embodiment 1, with the constituent elements described in each aspect of the present invention, the constituent elements having a part of the functions of the constituent elements described in each aspect of the present invention, or the constituent elements implementing a part of the processes implemented by the constituent elements described in each aspect of the present invention, and implement them;

[0083] (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processes corresponding to the processes described in each aspect of the present invention among the plurality of processes included in the method with the processes described in each aspect of the present invention;

[0084] (7) Combine a part of the processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processes described in each aspect of the present invention and implement them.

[0085] In addition, the implementation manners of the processes and / or structures described in each aspect of the present invention are not limited to the above examples. For example, it can also be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or structures described in each aspect can be implemented independently. In addition, the processes and / or structures described in different aspects can also be combined and implemented.

[0086] [Outline of Encoding Device]

[0087] First, the outline of the encoding device of Embodiment 1 will be described. Figure 1 It is a block diagram showing the functional structure of the encoding device 100 of Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in units of blocks.

[0088] As Figure 1As shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a splitting unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0089] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 can also be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0090] Hereinafter, each component included in the encoding device 100 will be described.

[0091] [Splitting Unit]

[0092] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits a picture into blocks of a fixed size (e.g., 128×128). Such blocks of the fixed size are sometimes referred to as coding tree units (CTUs). And the splitting unit 102 splits each block of the fixed size into blocks of a variable size (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block splitting. Such blocks of the variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transformation units (TUs). Additionally, in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in a picture can also be used as the processing units for CUs, PUs, and TUs.

[0093] Figure 2 is a diagram showing an example of the block splitting in Embodiment 1. In Figure 2 it, solid lines represent block boundaries based on quadtree block splitting, and dashed lines represent block boundaries based on binary tree block splitting.

[0094] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into four square 64×64 blocks (quad-tree block division).

[0095] The upper-left 64×64 block is further vertically divided into two rectangular 32×64 blocks, and the left 32×64 block is further vertically divided into two rectangular 16×64 blocks (binary-tree block division). As a result, the upper-left 64×64 block is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.

[0096] The upper-right 64×64 block is horizontally divided into two rectangular 64×32 blocks 14, 15 (binary-tree block division).

[0097] The lower-left 64×64 block is divided into four square 32×32 blocks (quad-tree block division). The upper-left block and the lower-right block among the four 32×32 blocks are further divided. The upper-left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary-tree block division). The lower-right 32×32 block is horizontally divided into two 32×16 blocks (binary-tree block division). As a result, the lower-left 64×64 block is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.

[0098] The lower-right 64×64 block 23 is not divided.

[0099] As described above, in Figure 2 , block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.

[0100] In addition, in Figure 2 , one block is divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary-tree division). The division including such ternary-tree division is sometimes called MBT (multi type tree) division.

[0101] [Subtraction section]

[0102] The subtraction unit 104 subtracts the predicted signal (predicted samples) from the original signal (original samples) in units of blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also referred to as the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction unit 104 outputs the calculated prediction error to the transformation unit 106.

[0103] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, a luma signal and two chroma signals). Hereinafter, there are cases where a signal representing an image is also referred to as a sample.

[0104] [Transformation Unit]

[0105] The transformation unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0106] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types and use a transform basis function corresponding to the selected transformation type to transform the prediction error into transform coefficients. Such a transformation is called EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform) in some cases.

[0107] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transform basis functions corresponding to each transformation type. In Figure 3 where N represents the number of input pixels. The selection of the transformation type from these multiple transformation types can depend, for example, on the type of prediction (intra prediction and inter prediction) or on the intra prediction mode.

[0108] Information indicating whether to apply such EMT or AMT (for example, called the AMT flag) and information indicating the selected transformation type are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0109] In addition, the transformation unit 106 can also perform inverse transformation on the transformation coefficients (transformation results). Such inverse transformation is called AST (adaptive secondary transform) or NSST (non-separable secondary transform) in some cases. For example, the transformation unit 106 performs inverse transformation on each sub-block (e.g., 4×4 sub-block) included in the block of transformation coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information related to the transformation matrix used in NSST are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0110] Here, a separable transformation refers to a method of performing multiple transformations by separating them in each direction according to the number of dimensions of the input. A non-separable transformation refers to a method of regarding two or more dimensions as one dimension and performing transformation together when the input is multi-dimensional.

[0111] For example, as an example of a non-separable transformation, when the input is a 4×4 block, it can be regarded as a permutation with 16 elements, and a transformation process is performed on this permutation with a 16×16 transformation matrix.

[0112] In addition, similarly, a method of performing Givens rotation on this permutation multiple times (Hypercube Givens Transform) after regarding a 4×4 input block as a permutation with 16 elements is also an example of a non-separable transformation.

[0113] [Quantization unit]

[0114] The quantization unit 108 quantizes the transformation coefficients output from the transformation unit 106. Specifically, the quantization unit 108 scans the transformation coefficients of the current block in a specified scanning order, and quantizes the transformation coefficients based on the quantization parameter (QP) corresponding to the scanned transformation coefficients. Then, the quantization unit 108 outputs the quantized transformation coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112.

[0115] The specified order is the order for quantization / inverse quantization of the transformation coefficients. For example, the specified scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).

[0116] A quantization parameter refers to a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0117] [Entropy Encoding Unit]

[0118] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) by performing variable-length encoding on the quantized coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantized coefficients, for example, and performs arithmetic coding on the binary signal.

[0119] [Inverse Quantization Unit]

[0120] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scanning order. And the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114.

[0121] [Inverse Transform Unit]

[0122] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. And the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0123] In addition, since information is lost through quantization in the restored prediction error, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.

[0124] [Addition Unit]

[0125] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. And the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a local decoded block.

[0126] [Block Memory]

[0127] The block memory 118 is a storage unit for storing blocks within the picture to be encoded (hereinafter referred to as the current picture) that are referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the addition unit 116.

[0128] [Loop Filter Unit]

[0129] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to the filtering used within the coding loop (in-loop filtering), and includes, for example, deblocking filtering (DF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0130] In ALF, a least squares error filter for removing coding distortion is adopted. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters is adopted based on the direction and activity of the locality-based gradient.

[0131] Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C is calculated (e.g., C = 5D + A). And based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).

[0132] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantifying the added result.

[0133] Based on the result of such classification, the filter for the sub-block is determined from among multiple filters.

[0134] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figures 4A to 4C It is a diagram showing multiple examples of the shape of the filter used in ALF. Figure 4A It represents a 5×5 diamond-shaped filter, Figure 4B It represents a 7×7 diamond-shaped filter, Figure 4C It represents a 9×9 diamond-shaped filter. The information representing the shape of the filter is signaled at the picture level. In addition, the signaling of the information representing the shape of the filter does not need to be limited to the picture level, and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0135] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined at the CU level whether to adopt ALF, and for chrominance, it is determined at the picture level whether to adopt ALF. The information representing the on / off of ALF is signaled at the picture level or CU level. In addition, the signaling of the information representing the on / off of ALF does not need to be limited to the picture level or CU level, and can also be other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0136] Coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Additionally, the signaling of the coefficient sets need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0137] [Frame memory]

[0138] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is also sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0139] [Intra-frame prediction unit]

[0140] The intra-frame prediction unit 124 performs intra-frame prediction (also referred to as intra-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.

[0141] For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of predefined intra-frame prediction modes. The plurality of intra-frame prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0142] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).

[0143] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. Additionally, the plurality of directional prediction modes may also include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Figure 5A Is a diagram showing 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. The solid arrows indicate 33 directions defined by the H.265 / HEVC standard, and the dashed arrows indicate the additional 32 directions.

[0144] In addition, in the intra prediction of the chrominance blocks, the luminance blocks may also be referred to. That is, the chrominance components of the current block may also be predicted based on the luminance component of the current block. Such intra prediction is called CCLM (cross-component linear model) prediction in some cases. The intra prediction mode of the chrominance blocks that refers to the luminance blocks (for example, called the CCLM mode) may also be added as one of the intra prediction modes of the chrominance blocks.

[0145] The intra prediction unit 124 may also correct the intra-predicted pixel value based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is called PDPC (position dependent intraprediction combination) in some cases. Information indicating whether PDPC is used (for example, called the PDPC flag) is signaled at the CU level, for example. In addition, the signaling of this information is not limited to the CU level and may also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0146] [Inter prediction unit]

[0147] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or sub-blocks (for example, 4×4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. And, the inter prediction unit 126 uses the motion information (for example, motion vector) obtained by the motion estimation to perform motion compensation, thereby generating the inter prediction signal of the current block or sub-block. And, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0148] The motion information used in the motion compensation is signaled. In the signaling of the motion vector, a motion vector predictor may also be used. That is, the difference between the motion vector and the predicted motion vector may also be signaled.

[0149] Alternatively, it may be that not only the motion information of the current block obtained by motion estimation is used, but also the motion information of adjacent blocks is used to generate an inter-frame prediction signal. Specifically, the prediction signal based on the motion information obtained by motion estimation may be weighted and added to the prediction signal based on the motion information of adjacent blocks, thereby generating an inter-frame prediction signal in units of sub-blocks within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0150] In such an OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is adopted (e.g., referred to as the OBMC flag) is signaled at the CU level. Additionally, the levels at which these pieces of information are signaled do not need to be limited to the sequence level and the CU level, and may also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0151] The OBMC mode will be described in more detail. Figure 5B and Figure 5C are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC processing.

[0152] First, using the motion vector (MV) assigned to the coding target block, a predicted image (Pred) obtained by normal motion compensation is acquired.

[0153] Next, the motion vector (MV_L) of the already-coded left adjacent block is adopted for the coding target block to obtain a predicted image (Pred_L), and the first correction of the predicted image is performed by weighted superposition of the above-mentioned predicted image and Pred_L.

[0154] Similarly, the motion vector (MV_U) of the already-coded upper adjacent block is adopted for the coding target block to obtain a predicted image (Pred_U), and the second correction of the predicted image is performed by weighted superposition of the predicted image after the above-mentioned first correction and Pred_U, and this is used as the final predicted image.

[0155] In addition, the method of two-stage correction using the left adjacent block and the upper adjacent block is described here, but it may also be configured to perform more than two-stage corrections using the right adjacent block and the lower adjacent block.

[0156] In addition, the area for superposition may not be the entire pixel area of the block, but only a part of the area near the block boundary.

[0157] In addition, the prediction image correction process based on one reference picture has been described here, but the same applies to the case of correcting the prediction image based on multiple reference pictures. After obtaining the corrected prediction images according to the respective reference pictures, the obtained prediction images are further superimposed to obtain the final prediction image.

[0158] In addition, the processing target block described above may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.

[0159] As a method for determining whether to use OBMC processing, for example, there is a method of using a signal indicating whether to use OBMC processing, i.e., obmc_flag. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as obmc_flag and encoding is performed using OBMC processing. If it does not belong to a region with complex motion, the value 0 is set as obmc_flag and encoding is performed without using OBMC processing. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag described in the stream and switching whether to use OBMC processing according to its value.

[0160] In addition, the motion information may not be signaled but derived on the decoding device side. For example, the merge mode defined by the H.265 / HEVC standard may be used. Additionally, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.

[0161] Here, the mode of performing motion estimation on the decoding device side is described. This mode of performing motion estimation on the decoding device side is called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.

[0162] In Figure 5D shows an example of FRUC processing. First, referring to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector are generated (which may be shared with the merge list). Then, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0163] And, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, a motion vector for the current block may also be derived by performing pattern matching in a peripheral region of the position within the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV may be searched by the same method, and in the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and this is used as the final MV for the current block. Additionally, a structure may be made where this process is not performed.

[0164] The exact same process may also be performed when processing is done in units of sub-blocks.

[0165] In addition, regarding the evaluation value, it is calculated by obtaining a difference value of the reconstructed image through pattern matching between the region within the reference picture corresponding to the motion vector and a specified region. Additionally, it may be that, in addition to the difference value, other information is also used to calculate the evaluation value.

[0166] As pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.

[0167] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Thus, in the first pattern matching, as the specified region for calculating the evaluation value for the candidate, a region within another reference picture along the motion trajectory of the current block is used.

[0168] Figure 6 It is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along the motion trajectory. As Figure 6 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two motion vectors (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position within the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position within the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the above candidate MV by the display time interval is obtained, and the evaluation value is calculated using the obtained difference value. A candidate MV with the best evaluation value among multiple candidate MVs may be selected as the final MV.

[0169] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0170] In the second pattern matching, pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture (e.g., an upper and / or left adjacent block)) and a block within the reference picture. Thus, in the second pattern matching, as the specified region for calculating the evaluation value for the above candidates, a block adjacent to the current block within the current picture is used.

[0171] Figure 7 It is a diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within the reference picture. As Figure 7 shown, in the second pattern matching, by searching within the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic), the motion vector of the current block is derived. Specifically, for the current block, the difference between the reconstructed image of the coded region of both or one of the left adjacent and upper adjacent regions and the reconstructed image at the equivalent position within the coded reference picture (Ref0) specified by the candidate MV is derived, and the obtained difference value is used to calculate the evaluation value. Among multiple candidate MVs, the candidate MV with the best evaluation value is selected as the best candidate MV.

[0172] Information indicating whether to adopt the FRUC mode (e.g., referred to as the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the method of pattern matching (the first pattern matching or the second pattern matching) (e.g., referred to as the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0173] Here, a mode of deriving a motion vector based on a model assuming uniform linear motion is described. This mode includes a case called BIO (bi-directional optical flow).

[0174] Figure 8 It is a diagram for explaining a model assuming uniform linear motion. InFigure 8 Among them, (v x , v y ) represents the velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.

[0175] At this time, under the assumption of a uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0176] [Equation 1]

[0177]

[0178] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a merge list, etc. are corrected in pixel units.

[0179] In addition, the motion vector can also be derived on the decoder side by a method different from the derivation of the motion vector based on a model assuming uniform linear motion. For example, the motion vector can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0180] Here, a mode of deriving the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks will be described. This mode is the case called affine motion compensation prediction mode.

[0181] Figure 9A is a diagram for explaining the derivation of the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks. In Figure 9AAmong them, the current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the upper-left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the upper-right control point of the current block is derived. And, using the two motion vectors v0 and v1, the motion vectors (v x , v y ) of each sub-block within the current block are derived through the following equation (2).

[0182] [Equation 2]

[0183]

[0184] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a preset weight coefficient.

[0185] In such an affine motion compensation prediction mode, several modes with different methods for deriving the motion vectors of the upper-left and upper-right control points may also be included. Information representing such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. In addition, the signaling of the information representing this affine motion compensation prediction mode does not need to be limited to the CU level, and may also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0186] [Prediction control unit]

[0187] The prediction control unit 128 selects one of the intra-prediction signal and the inter-prediction signal, uses the selected signal as the prediction signal, and outputs it to the subtraction unit 104 and the addition unit 116.

[0188] Here, an example of deriving the motion vector of the coded picture through the merge mode is described. Figure 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.

[0189] First, a prediction MV list registering candidates of prediction MVs is generated. As candidates of prediction MVs, there are the MVs of multiple coded blocks spatially located around the coded object block, i.e., spatially adjacent prediction MVs, the MVs of blocks near the projection of the position of the coded object block in the coded reference picture, i.e., temporally adjacent prediction MVs, the MVs generated by combining the MV values of spatially adjacent prediction MVs and temporally adjacent prediction MVs, i.e., combined prediction MVs, and the MVs with a value of zero, i.e., zero prediction MVs, etc.

[0190] Next, by selecting one prediction MV from the multiple prediction MVs registered in the prediction MV list, it is determined as the MV of the coded object block.

[0191] Further, in the variable-length coding section, the merge_idx, which is a signal indicating which predicted MV is selected, is described in the stream and encoded.

[0192] In addition, Figure 9B The predicted MVs registered in the predicted MV list described in are an example. The number may be different from the number in the figure, or the structure may not include some types of the predicted MVs in the figure, or the structure may include predicted MVs other than the types of the predicted MVs in the figure.

[0193] In addition, the MV of the coding target block derived by the merge mode may be used for the following DMVR processing to determine the final MV.

[0194] Here, an example of determining the MV using the DMVR processing is described.

[0195] Figure 9C is a conceptual diagram for explaining the outline of the DMVR processing.

[0196] First, the optimal MVP set for the processing target block is used as the candidate MV. According to the above candidate MVs, reference pixels are obtained from the first reference picture of the processed pictures in the L0 direction and the second reference picture of the processed pictures in the L1 direction respectively, and a template is generated by taking the average of each reference pixel.

[0197] Next, using the above template, the peripheral areas of the candidate MVs in the first reference picture and the second reference picture are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, for the cost value, it is calculated using the difference values between the pixel values of the template and the pixel values of the search area and the MV value, etc.

[0198] In addition, in the encoding device and the decoding device, the outline of the processing described here is basically common.

[0199] In addition, even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV and derive the final MV, other processing can also be used.

[0200] Here, the mode of generating a predicted image using the LIC processing is described.

[0201] Figure 9D is a diagram for explaining the outline of the predicted image generation method using the luminance correction processing based on the LIC processing.

[0202] First, an MV for obtaining a reference image corresponding to the coding target block from the reference picture of the encoded picture is derived.

[0203] Next, for the coded object block, using the luminance pixel values of the coded neighboring reference regions on the left and above and the luminance pixel values at the same position in the reference picture specified by the MV, information indicating how the luminance values change in the reference picture and the coded object picture is extracted, and a luminance correction parameter is calculated.

[0204] By performing a luminance correction process on the reference image in the reference picture specified by the MV using the above luminance correction parameter, a prediction image for the coded object block is generated.

[0205] In addition, Figure 9D the shape of the above neighboring reference region in Figure 9D is an example, and shapes other than this can also be used.

[0206] Furthermore, the process of generating a prediction image based on one reference picture is described here, but the same applies when generating a prediction image based on multiple reference pictures. After performing a luminance correction process on the reference images obtained from each reference picture in the same way, a prediction image is generated.

[0207] As a method for determining whether to adopt the LIC process, for example, there is a method of using the lic_flag as a signal indicating whether to adopt the LIC process. As a specific example, in the encoding device, it is determined whether the coded object block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag, and encoding is performed using the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag, and encoding is performed without using the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to adopt the LIC process according to its value.

[0208] As another method for determining whether to adopt the LIC process, for example, there is also a method of determining according to whether the LIC process has been adopted in the neighboring blocks. As a specific example, when the coded object block is in the merge mode, it is determined whether the neighboring coded blocks selected during the derivation of the MV in the merge mode process have been encoded using the LIC process, and encoding is performed by switching whether to adopt the LIC process according to the result. In addition, in this example, the process during decoding is exactly the same.

[0209] [Outline of the decoding device]

[0210] Next, an outline of the decoding device that can decode the coded signal (coded bitstream) output from the above encoding device 100 will be described. Figure 10FIG. 0 is a block diagram showing the functional configuration of the decoding apparatus 200 according to Embodiment 1. The decoding apparatus 200 is a moving image / image decoding apparatus that decodes moving images / images in units of blocks.

[0211] As Figure 10 shown, the decoding apparatus 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0212] The decoding apparatus 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoding apparatus 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0213] Hereinafter, each component included in the decoding apparatus 200 will be described.

[0214] [Entropy Decoding Unit]

[0215] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Next, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in units of blocks.

[0216] [Inverse Quantization Unit]

[0217] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the block to be decoded (hereinafter referred to as the current block) that is input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transformation unit 206.

[0218] [Inverse Transformation Unit]

[0219] The inverse transformation unit 206 restores the prediction error by performing inverse transformation on the transform coefficients that are input from the inverse quantization unit 204.

[0220] For example, when the information decoded from the coded bitstream indicates the use of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the decoded information indicating the transform type.

[0221] In addition, for example, when the information decoded from the coded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re - transform to the transform coefficients.

[0222] [Addition unit]

[0223] The addition unit 208 reconstructs the current block by adding the prediction error as the input from the inverse transform unit 206 and the prediction sample as the input from the prediction control unit 220. Moreover, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0224] [Block memory]

[0225] The block memory 210 is a storage unit for storing blocks within the decoded picture (hereinafter referred to as the current picture) that are referenced in intra - prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.

[0226] [Loop filter unit]

[0227] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.

[0228] When the information decoded from the coded bitstream indicating the on / off of ALF indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.

[0229] [Frame memory]

[0230] The frame memory 214 is a storage unit for storing reference pictures used in inter - prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0231] [Intra - prediction unit]

[0232] The intra prediction unit 216 performs intra prediction with reference to the blocks within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bitstream, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance differences) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0233] In addition, when the intra prediction mode of referring to the luminance block is selected for the intra prediction of the chrominance block, the intra prediction unit 216 may also predict the chrominance component of the current block based on the luminance component of the current block.

[0234] Furthermore, when the information decoded from the encoded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0235] [Inter prediction unit]

[0236] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using the motion information (e.g., motion vector) decoded from the encoded bitstream, thereby generating an inter prediction signal for the current block or sub-block, and outputs the inter prediction signal to the prediction control unit 220.

[0237] In addition, when the information decoded from the encoded bitstream indicates the adoption of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained through motion estimation but also the motion information of the adjacent blocks.

[0238] Furthermore, when the information decoded from the encoded bitstream indicates the adoption of the FRUC mode, the inter prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) decoded from the encoded stream, thereby deriving the motion information. And the inter prediction unit 218 performs motion compensation using the derived motion information.

[0239] In addition, when the BIO mode is adopted, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Furthermore, when the information decoded from the encoded bitstream indicates the adoption of the affine motion compensation prediction mode, the inter prediction unit 218 derives a motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.

[0240] [Prediction control unit]

[0241] The prediction control unit 220 selects one of the intra prediction signal and the inter prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208.

[0242] (Embodiment 2)

[0243] Regarding the encoding process and the decoding process related to Embodiment 2, refer to Figure 11 and Figure 12 Specifically, regarding the encoding device and the decoding device related to Embodiment 2, refer to Figure 15 and Figure 16 Specifically described.

[0244] [Encoding Process]

[0245] Figure 11 Represents the video encoding process related to Embodiment 2.

[0246] First, in step S1001, a first parameter for identifying a partitioning mode for partitioning a first block into a plurality of sub-blocks from among a plurality of partitioning modes is written into the bitstream. If a partitioning mode is used, the block is divided into a plurality of sub-blocks. If different partitioning modes are used, the block is divided into a plurality of sub-blocks having different shapes, different heights, or different widths.

[0247] Figure 28 Represents an example of a partitioning mode for partitioning an N×N pixel block in Embodiment 2. In Figure 28 , (a) to (h) represent different partitioning modes. As Figure 28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and "N" can take any value that is an integer multiple of 4 from 8 to 128) is divided into 2 sub - blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, a block of N×N pixels is divided into a sub - block of N / 4×N pixels (for example, 4×16 pixels) and a sub - block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, a block of N×N pixels is divided into a sub - block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub - block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, a block of N×N pixels is divided into a sub - block of (N / 4)×N pixels (for example, 4×16 pixels), a sub - block of N / 2×N pixels (for example, 8×16 pixels), and a sub - block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, a block of N×N pixels is divided into 2 sub - blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, a block of N×N pixels is divided into a sub - block of N×N / 4 pixels (for example, 16×4 pixels) and a sub - block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, a block of N×N pixels is divided into a sub - block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub - block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, a block of N×N pixels is divided into a sub - block of N×N / 4 pixels (for example, 16×4 pixels), a sub - block of N×N / 2 pixels (for example, 16×8 pixels), and a sub - block of N×N / 4 pixels (for example, 16×4 pixels).

[0248] Next, in step S1002, it is determined whether the first parameter identifies the first block mode.

[0249] Next, in step S1003, based at least on the determination of whether the first parameter identifies the first block mode, it is determined whether the second block mode is not selected as a candidate for dividing the second block.

[0250] Two different sets of block modes may divide a block into sub - blocks of the same shape and size. For example, as Figure 31A shown, the sub - blocks of (1b) and (2c) have the same shape and size. One set of block modes can contain at least 2 block modes. For example, as Figure 31A shown in (1a) and (1b) of Figure 31AAs shown in (2a), (2b), and (2c), other block pattern sets can follow the vertical binary tree split to include a vertical binary tree split of two sub-blocks. Each block pattern set results in sub-blocks of the same shape and size.

[0251] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that have different binary numbers or different numbers of bits when encoded in the bitstream, select the block pattern set with fewer binary numbers or fewer bits. Additionally, the binary numbers and the number of bits correspond to the code amount.

[0252] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that have the same binary number or the same number of bits when encoded in the bitstream, select the block pattern set that appears first in a specified order of the multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.

[0253] Figure 31A and Figure 31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of the block pattern. In this example, when the left N×N pixel block is vertically divided into two sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Figure 31B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).

[0254] Figures 32A to 32C is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that appears first in a specified order of the multiple block pattern sets. In this example, when the 2N×N / 2 pixel block is vertically divided into three sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Figure 32B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Figure 32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of the multiple block pattern sets can also be fixed and signaled within the bitstream.

[0255] Figure 20This represents an example in Embodiment 2 where the second block division mode is not selected for the division of a 2N×N pixel block as shown in step (2c). As Figure 20 shown, the first division method (i) can be used to equally divide a 2N×2N pixel block (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, the second division method (ii) can be used to horizontally equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). Here, in the second division method (ii), when the upper 2N×N pixel block (the first block) is vertically divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for vertically dividing the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division modes. This is because sub-block sizes identical to those obtained by the four-way division using the first division method (i) are generated.

[0256] As described above, in Figure 20 , when the first block is vertically equally divided into 2 sub-blocks if the first block division mode is used, and the second block adjacent to the first block in the vertical direction is vertically equally divided into 2 sub-blocks if the second block division mode is used, the second block division mode is not selected as a candidate.

[0257] Figure 21 This represents an example in Embodiment 2 where the second block division mode is not selected for the division of an N×2N pixel block as shown in step (2c). As Figure 21 shown, the first division method (i) can be used to equally divide a 2N×2N pixel block into 4 sub-blocks of N×N pixels as in step (1a). Also, the second division method (ii) can be used to vertically equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second division method (ii), when the left N×2N pixel block (the first block) is horizontally divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for horizontally dividing the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division modes. This is because sub-block sizes identical to those obtained by the four-way division using the first division method (i) are generated.

[0258] As described above, in Figure 21In the case where, if the first block mode is used, the first block is equally divided into two sub-blocks horizontally, and if the second block mode is used, the second block adjacent to the first block horizontally is equally divided into two sub-blocks horizontally, the second block mode is not selected as a candidate.

[0259] Figure 40 Indicates an example of dividing a 4N×2N block in Figure 20 into three parts in a ratio of 1:2:1, such as N×2N, 2N×2N, N×2N. Here, when the upper block is divided into three parts, the block division mode of dividing the lower block into three parts in a ratio of 1:2:1 is not selected as a candidate for possible block division modes. The three-way division can also be in a ratio different from 1:2:1. Furthermore, it can be divided into more than three parts, can be divided into two parts, or can be in a ratio different from 1:1, such as 1:2 or 1:3. Figure 40 This is an example of dividing first in the horizontal direction, but the same constraints can also be applied when dividing first in the vertical direction.

[0260] Figure 41 and Figure 42 Indicates an example of applying the same constraints when the first block is rectangular.

[0261] Figure 43 This is the second constraint example when a square is divided into three parts vertically and then equally divided into two parts horizontally. When applying Figure 43 the constraints, in Figure 40 , it is possible to select the block division mode of dividing the lower block of 4N×2N into three parts in a ratio of 1:2:1. It is also possible to separately encode the information indicating which of the constraints of Figure 40 and Figure 43 is applied into the header information, etc. Or, it is also possible to apply constraints to reduce the code amount of the information representing the block division. For example, if it is assumed that the code amounts of the information representing the block division in Case 1 and Case 2 are as follows, the division in Case 1 is set to be valid and the division in Case 2 is set to be invalid. That is, the constraints of Figure 43 are applied.

[0262] (Case 1) (1) Divide the square into two parts horizontally, and then (2) divide the upper and lower two rectangular blocks vertically into three parts respectively: (1) Direction information: 1 bit, Division quantity information: 1 bit, (2) (Direction information: 1 bit, Division quantity information: 1 bit)×2, a total of 6 bits

[0263] (Case 2) (1) Divide the square vertically, and then (2) divide the left, middle, and right rectangular blocks horizontally into two parts respectively: (1) Direction information: 1 bit, Division quantity information: 1 bit, (2) (Direction information: 1 bit, Division quantity information: 1 bit)×3, a total of 8 bits

[0264] Alternatively, there is a case where the optimal block is determined while selecting block patterns in a prescribed order during encoding. For example, 2-way splitting may be attempted first, followed by 3-way or 4-way splitting (bisecting horizontally and vertically). At this time, before attempting 3-way splitting as in Figure 43 , the attempt starting from 2-way splitting as in the example of Figure 40 has already been carried out. Thus, in the attempt starting from 2-way splitting, bisection horizontally and then trisecting the two upper and lower blocks vertically has already been attempted, so the constraints of Figure 43 are applied. In this way, the method of determining the selection constraints can also be decided based on a prescribed encoding method.

[0265] In Figure 44 , an example is shown in which the block patterns that can be selected in the second block pattern for the same direction as the first block pattern are restricted. Here, the first block pattern is 3-way splitting in the vertical direction. At this time, 2-way splitting cannot be selected as the second block pattern. On the other hand, for the vertical direction, which is a direction different from the first block pattern, 2-way splitting can be selected ( Figure 45 ).

[0266] Figure 22 An example is shown in Embodiment 2 where the second block pattern is not selected for the splitting of an N×N pixel block as shown in step (2c). As shown in Figure 22 , the first splitting method (i) can be used to vertically split a 2N×N pixel block (e.g., 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (e.g., sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels). Also, the second splitting method (ii) can be used to split a 2N×N pixel block into two N×N pixel sub-blocks as shown in step (2a). In the first splitting method (i), the central N×N pixel block can be vertically split into two N / 2×N pixel (e.g., 4×8 pixel) sub-blocks in step (1b). In the second splitting method (ii), when the left N×N pixel block (the first block) is vertically split into two N / 2×N pixel sub-blocks as shown in step (2b), in step (2c), the block pattern of vertically splitting the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks is not selected as a candidate for possible block patterns. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, i.e., four N / 2×N pixel sub-blocks.

[0267] As described above, in Figure 22In the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block mode is not selected as a candidate.

[0268] Figure 23 This shows an example of not selecting the second block mode for the division of an N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 23 shown, the first splitting method (i) can be used to split an N×2N pixel (e.g., 8×16 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (e.g., sub-blocks of 8×4 pixels, sub-blocks of 8×8 pixels, and sub-blocks of 8×4 pixels) as in step (1a). Also, the second splitting method can be used to split it into two sub-blocks of N×N pixels as in step (2a). In the first splitting method (i), the central N×N pixel block can be split into two sub-blocks of N×N / 2 pixels as in step (1b). In the second splitting method (ii), when the upper N×N pixel block (the first block) is horizontally split into two sub-blocks of N×N / 2 pixels as in step (2b), in step (2c), the block splitting mode of horizontally splitting the lower N×N pixel block (the second block) into two sub-blocks of N×N / 2 pixels is not selected as a candidate for possible block splitting modes. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, that is, four sub-blocks of N×N / 2 pixels.

[0269] As described above, in Figure 23 the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block mode is not selected as a candidate.

[0270] If it is determined that the second block mode is selected as a candidate for splitting the second block (No in S1003), then in step S1004, a block splitting mode is selected from among multiple block splitting modes including the second block mode as a candidate. In step S1005, a second parameter indicating the selection result is written into the bitstream.

[0271] If it is determined that the second block mode is not selected as a candidate for splitting the second block (Yes in S1003), then in step S1006, a block splitting mode different from the second block mode is selected for splitting the second block. Among the block splitting modes selected here, the block is split into sub-blocks having a different shape or different size compared to the sub-blocks generated by the second block mode.

[0272] Figure 24 This shows an example of using the selected block mode when not selecting the second block mode as shown in step (3) in Embodiment 2 to divide a 2N×N pixel block. As Figure 24 shown, the selected block mode can divide the current block of 2N×N pixels (the lower block in this example) as Figure 24 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 24 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.

[0273] Figure 25 This shows an example of using the selected block mode when not selecting the second block mode as shown in step (3) in Embodiment 2 to divide an N×2N pixel block. As Figure 25 shown, the selected block mode can divide the current block of N×2N pixels (the right block in this example) as Figure 25 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 25 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.

[0274] Figure 26 This shows an example of using the selected block mode when not selecting the second block mode as shown in step (3) in Embodiment 2 to divide an N×N pixel block. As Figure 26 shown, in step (1), the 2N×N pixel block is vertically divided into 2 N×N pixel sub-blocks, and in step (2), the left N×N pixel block is vertically divided into 2 N / 2×N pixel sub-blocks. In step (3), the selected block mode for the current block of N×N pixels (the left block in this example) can be used to divide the current block as Figure 26 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Figure 26is divided into two sub-blocks (asymmetric binary tree) with different sizes as shown in (a), (b), (d), and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0275] Figure 27 Shows an example of dividing an N×N pixel block using the selected block partitioning mode when not selecting the second block partitioning mode as shown in step (3) in Embodiment 2. As Figure 27 shown, in step (1), the N×2N pixel block is horizontally divided into two N×N pixel sub-blocks, and in step (2), the upper N×N pixel block is horizontally divided into two N×N / 2 pixel sub-blocks. In step (3), the selected block partitioning mode for the current N×N pixel block (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in Figure 27 (c) and (f). The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block can have twice the width / height of the small sub-block. And for example, the selected block partitioning mode can also divide the current block into two sub-blocks with different sizes (asymmetric binary tree) as shown in Figure 27 (a), (b), (d), and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0276] Figure 17 Represents the possible positions of the first parameter in the compressed video stream. As Figure 17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0277] Figure 18 Represents the possible positions of the second parameter in the compressed video stream. As Figure 18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Figure 19 shown, the second parameter follows the first parameter in the bitstream.

[0278] The first block and the second block are different blocks. The first block and the second block may be included in the same frame. For example, the first block may be an adjacent block above the second block. And for example, the first block may also be an adjacent block to the left of the second block.

[0279] In step S1007, the second block is divided into sub-blocks using the selected block partitioning mode. In step S1008, the divided blocks are encoded.

[0280] [Encoding device]

[0281] Figure 15 is a block diagram showing the structure of the video / image encoding device according to Embodiment 2 or 3.

[0282] The video encoding device 5000 is a device for encoding an input video / image for each block to generate an encoded output bitstream. As Figure 15 shown, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block segmentation determination unit 5010.

[0283] The input video is input to the adder, and the added value is output to the transform unit 5001. The transform unit 5001 transforms the added value into frequency coefficients based on the block partitioning mode derived by the block segmentation determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block partitioning mode can be associated with a block partitioning mode, a block partitioning type, or a block partitioning direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantized values to the inverse quantization unit 5003 and the entropy encoding unit 5009.

[0284] The inverse quantization unit 5003 inverse-quantizes the quantized values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block segmentation mode derived by the block segmentation determination unit 5010, transforms the frequency coefficients into sample values of the bitstream, and outputs the sample values to the adder.

[0285] The adder adds the sample values of the bitstream output from the inverse transform unit 5004 to the predicted video / image values output from the intra / inter prediction units 5007 and 5008, and outputs the added value to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives the block partitioning mode and parameters related to the block partitioning mode. If the derived block partitioning mode is used, the block is divided into a plurality of sub-blocks. The intra / inter prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block partitioning mode derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.

[0286] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.

[0287] [Decoding process]

[0288] Figure 12 Indicates the video decoding process related to Embodiment 2.

[0289] First, in step S2001, the first parameter is interpreted according to the bitstream. The first parameter identifies the partitioning mode for dividing the first block into sub-blocks from among a plurality of partitioning modes. If the partitioning mode is used, the block is divided into sub-blocks, and if different partitioning modes are used, the block is divided into sub-blocks with different shapes, different heights, or different widths.

[0290] Figure 28 Shows an example of the partitioning mode for dividing an N×N pixel block in Embodiment 2. In Figure 28 ,(a) to (h) represent different partitioning modes. As Figure 28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and as the value of "N", any value that is an integer multiple of 4 from 8 to 128 can be taken) is divided into two sub-blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, the block of N×N pixels is divided into a sub-block of N / 4×N pixels (for example, 4×16 pixels) and a sub-block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, the block of N×N pixels is divided into a sub-block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, the block of N×N pixels is divided into a sub-block of (N / 4)×N pixels (for example, 4×16 pixels), a sub-block of N / 2×N pixels (for example, 8×16 pixels), and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, the block of N×N pixels is divided into two sub-blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, the block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels) and a sub-block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, the block of N×N pixels is divided into a sub-block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub-block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, the block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels), a sub-block of N×N / 2 pixels (for example, 16×8 pixels), and a sub-block of N×N / 4 pixels (for example, 16×4 pixels).

[0291] Next, in step S2002, it is determined whether the first parameter identifies the first block mode.

[0292] Next, in step S2003, based at least on the determination of whether the first parameter identifies the first block mode, it is determined whether the second block mode is not selected as a candidate for dividing the second block.

[0293] Two different sets of block modes may divide a block into sub-blocks of the same shape and size. For example, as Figure 31A shown, the sub-blocks of (1b) and (2c) have the same shape and size. One set of block modes can contain at least two block modes. For example, as Figure 31A shown in (1a) and (1b) of, one set of block modes can then, following a vertical trinary tree split, include a vertical binary tree split of the central sub-block and a non-split of the other sub-blocks. And for example, as Figure 31AAs shown in (2a), (2b), and (2c), other block pattern sets can then include a binary tree vertical split that divides into two sub-blocks through a binary tree vertical split. Each block pattern set results in sub-blocks of the same shape and size.

[0294] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are different binary numbers or different numbers of bits when encoded in a bitstream, select the block pattern set with fewer binary numbers or fewer bits.

[0295] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that have the same number of bits or the same number of bits when encoded in a bitstream, select the block pattern set that first appears in a specified order of multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.

[0296] Figure 31A and Figure 31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of block patterns. In this example, when a left N×N pixel block is vertically divided into two sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Figure 31B the encoding method of block patterns, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).

[0297] Figure 32A is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that first appears in a specified order of multiple block pattern sets. In this example, when a 2N×N / 2 pixel block is vertically divided into three sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Figure 32B the encoding method of block patterns, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Figure 32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of multiple block pattern sets can also be fixed and signaled within the bitstream.

[0298] Figure 20 represents an example of not selecting the second block pattern for the division of a 2N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 20As shown, it is possible to use the first segmentation method (i) to equally divide a block of 2N×2N pixels (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, it is possible to use the second segmentation method (ii) to equally divide a block of 2N×2N pixels horizontally into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). In the second segmentation method (ii), when the upper 2N×N pixel block (the first block) is vertically divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), the second block division mode for vertically dividing the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division mode in step (2c). This is because sub-block sizes identical to those obtained by the four-way division using the first segmentation method (i) will be generated.

[0299] As described above, Figure 20 in a case where if the first block division mode is used, the first block is equally divided into 2 sub-blocks in the vertical direction, and if the second block division mode is used, the second block adjacent to the first block in the vertical direction is equally divided into 2 sub-blocks in the vertical direction, the second block division mode is not selected as a candidate.

[0300] Figure 21 This shows an example in Embodiment 2 where the second block division mode is not selected for the division of the N×2N pixel block as shown in step (2c). As Figure 21 shown, it is possible to use the first segmentation method (i) to equally divide a block of 2N×2N pixels into 4 sub-blocks of N×N pixels as in step (1a). Also, it is possible to use the second segmentation method (ii) to equally divide a block of 2N×2N pixels vertically into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second segmentation method (ii), when the left N×2N pixel block (the first block) is horizontally divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), the second block division mode for horizontally dividing the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division mode in step (2c). This is because sub-block sizes identical to those obtained by the four-way division using the first segmentation method (i) will be generated.

[0301] As described above, Figure 21 in a case where if the first block division mode is used, the first block is equally divided into 2 sub-blocks in the horizontal direction, and if the second block division mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into 2 sub-blocks in the horizontal direction, the second block division mode is not selected as a candidate.

[0302] Figure 22 This shows an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 22 shown, the first division method (i) can be used to vertically divide a 2N×N pixel block (for example, 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (for example, sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels) as in step (1a). Also, the second division method (ii) can be used to divide a 2N×N pixel block into two N×N pixel sub-blocks as in step (2a). In the first division method (i), the central N×N pixel block can be vertically divided into two N / 2×N pixel (for example, 4×8 pixel) sub-blocks in step (1b). In the second division method (ii), when the left N×N pixel block (the first block) is vertically divided into two N / 2×N pixel sub-blocks as in step (2b), the block division mode of vertically dividing the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks in step (2c) is not selected as a candidate for the possible block division modes. This is because sub-blocks of the same size as those obtained by the first division method (i) will be generated, that is, four N / 2×N pixel sub-blocks.

[0303] As described above, Figure 22 in a case where if the first block division mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block division mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block division mode is not selected as a candidate.

[0304] Figure 23 This shows an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c) in Embodiment 2. As Figure 23As shown, it is possible to use the first splitting method (i) to split an N×2N pixel (e.g., 8×16 pixels, where "N" can take any value that is an integer multiple of 4 from 8 to 128) into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (e.g., sub-blocks of 8×4 pixels, sub-blocks of 8×8 pixels, and sub-blocks of 8×4 pixels) as in step (1a). Also, it is possible to use the second splitting method to split it into two sub-blocks of N×N pixels as in step (2a). In the first splitting method (i), it is possible to split the central block of N×N pixels into two sub-blocks of N×N / 2 pixels as in step (1b). In the second splitting method (ii), when the upper N×N pixel block (the first block) is horizontally split into two sub-blocks of N×N / 2 pixels as in step (2b), the block splitting mode in which the lower N×N pixel block (the second block) is horizontally split into two sub-blocks of N×N / 2 pixels in step (2c) is not selected as a candidate for the possible block splitting modes. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, i.e., four sub-blocks of N×N / 2 pixels.

[0305] As described above, Figure 23 in the case where if the first block splitting mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block splitting mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block splitting mode is not selected as a candidate.

[0306] If it is determined that the second block splitting mode is selected as a candidate for splitting the second block (No in S2003), then in step S2004, the second parameter is decoded from the bitstream, and a block splitting mode is selected from among multiple block splitting modes including the second block splitting mode as a candidate.

[0307] If it is determined that the second block splitting mode is not selected as a candidate for splitting the second block (Yes in S2003), then in step S2005, a block splitting mode different from the second block splitting mode is selected for splitting the second block. The block splitting mode selected here splits the block into sub-blocks having a different shape or different size compared to the sub-blocks generated by the second block splitting mode.

[0308] Figure 24 An example is shown in Embodiment 2 of using the block splitting mode selected when the second block splitting mode is not selected to split the 2N×N pixel block as in step (3). As Figure 24 shown, the selected block splitting mode can split the current 2N×N pixel block (the lower block in this example) as Figure 24The division shown in (c) and (f) is into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height twice that of the small sub-block. And for example, the selected block division pattern can also divide the current block as shown in Figure 24 (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, in the case of using an asymmetric binary tree, the large sub-block can have a width / height three times that of the small sub-block.

[0309] Figure 25 This represents an example of dividing a block of N×2N pixels using the selected block division pattern without selecting the second block division pattern as shown in step (3) in Embodiment 2. As shown in Figure 25 the selected block division pattern can divide the current block of N×2N pixels (the right block in this example) as shown in Figure 25 (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height twice that of the small sub-block. And for example, the selected block division pattern can also divide the current block as shown in Figure 25 (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, in the case of using an asymmetric binary tree, the large sub-block can have a width / height three times that of the small sub-block.

[0310] Figure 26 This represents an example of dividing a block of N×N pixels using the selected block division pattern without selecting the second block division pattern as shown in step (3) in Embodiment 2. As shown in Figure 26 In step (1), the block of 2N×N pixels is vertically divided into 2 sub-blocks of N×N pixels. In step (2), the left block of N×N pixels is vertically divided into 2 sub-blocks of N / 2×N pixels. In step (3), the selected block division pattern for the current block of N×N pixels (the left block in this example) can be used to divide the current block as shown in Figure 26 (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height twice that of the small sub-block. And for example, the selected block division pattern can also divide the current block as shown in Figure 26 (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, in the case of using an asymmetric binary tree, the large sub-block can have a width / height three times that of the small sub-block.

[0311] Figure 27 This represents an example of dividing a block of N×N pixels using the selected block division pattern without selecting the second block division pattern as shown in step (3) in Embodiment 2. As shown in Figure 27As shown, in step (1), a block of N×2N pixels is horizontally divided into two sub-blocks of N×N pixels. In step (2), the upper block of N×N pixels is horizontally divided into two sub-blocks of N×N / 2 pixels. In step (3), the selected partitioning mode for the current block of N×N pixels (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of Figure 27 . The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block has twice the width / height of the small sub-block. And for example, the selected partitioning mode can also divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of Figure 27 . For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0312] Figure 17 Indicates the possible positions of the first parameter in the compressed video stream. As shown in Figure 17 , the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0313] Figure 18 Indicates the possible positions of the second parameter in the compressed video stream. As shown in Figure 18 , the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As shown in Figure 19 , the second parameter is configured in the bitstream following the first parameter.

[0314] The first block and the second block are different blocks. The first block and the second block can also be included in the same frame. For example, the first block can be a block adjacent above the second block. And for example, the first block can also be a block adjacent to the left of the second block.

[0315] In step S2006, the second block is divided into sub-blocks using the selected partitioning mode. In step S2007, the divided block is decoded.

[0316] [Decoding device]

[0317] Figure 16It is a block diagram showing the structure of the video / image decoding device according to Embodiment 2 or 3.

[0318] The video decoding device 6000 is a device for decoding an input encoded bitstream for each block and outputting a video / image. The video decoding device 6000 is as Figure 16 shown and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block segmentation determination unit 6008.

[0319] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.

[0320] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. The inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients based on the block partitioning pattern derived by the block segmentation determination unit 6008, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning pattern can be associated with a block partitioning pattern, a block partitioning type, or a block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006, 6007, outputs the added value to the display, and outputs the added value to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005, and uses the parameters decoded by the entropy decoding unit 6001 to derive the block partitioning pattern. If the derived block partitioning pattern is used, the block is divided into multiple sub-blocks. Further, the intra / inter prediction units 6006, 6007 perform prediction on the video / image region of the block to be decoded based on the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed according to the block partitioning pattern derived by the block segmentation determination unit 6008.

[0321] (Embodiment 3)

[0322] Refer to Figure 13 and Figure 14 Specifically describe the encoding process and decoding process according to Embodiment 3. Refer to Figure 15 and Figure 16 Specifically describe the encoding device and decoding device according to Embodiment 3.

[0323] [Encoding Process]

[0324] Figure 13Indicates the video encoding process for Embodiment 3.

[0325] First, in step S3001, a first parameter is written to the bitstream, which identifies, from among a plurality of block types, the block type used to divide the first block into sub-blocks.

[0326] In the next step S3002, a second parameter indicating the block division direction is written to the bitstream. The second parameter is configured in the bitstream following the first parameter. The block type and the block division direction can form a block division pattern together. The divided block indicates the number and division ratio of sub-blocks used to divide the block.

[0327] Figure 29 Shows an example of the block type and block division direction used to divide an N×N pixel block in Embodiment 3. In Figure 29 , (1), (2), (3), and (4) are different block types, (1a), (2a), (3a), and (4a) are block division patterns with different block types in the vertical division direction, and (1b), (2b), (3b), and (4b) are block division patterns with different block types in the horizontal division direction. As Figure 29 shown, when the division ratio is 1:1 and the N×N pixel block is divided along the vertical direction by a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (1a). When the division ratio is 1:1 and the N×N pixel block is divided along the horizontal direction by a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (1b). When the division ratio is 1:3 and the N×N pixel block is divided along the vertical direction by an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (2a). When the division ratio is 1:3 and the N×N pixel block is divided along the horizontal direction by an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (2b). When the division ratio is 3:1 and the N×N pixel block is divided along the vertical direction by an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (3a). When the division ratio is 3:1 and the N×N pixel block is divided along the horizontal direction by an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block division pattern (3b). When the division ratio is 1:2:1 and the N×N pixel block is divided along the vertical direction by a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block division pattern (4a). When the division ratio is 1:2:1 and the N×N pixel block is divided along the horizontal direction by a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block division pattern (4b).

[0328] Figure 17 Shows the possible positions of the first parameter in the compressed video stream. As Figure 17As shown, the first parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter can represent a method for dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0329] Figure 18 Indicates a position where the second parameter within the compressed video stream can be considered. As Figure 18 shown, the second parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter can represent a method for dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Figure 19 shown, the second parameter is configured in the bitstream following the first parameter.

[0330] Figure 30 Indicates the advantage of coding the block type before the block direction compared to coding the block direction before the block type. In this example, when the horizontal block direction is invalidated due to an unsupported size (16×2 pixels), there is no need to code the block direction. In this example, the block direction is determined to be the vertical block direction, and the horizontal block direction is invalid. When coding the block type before the block direction, compared to coding the block direction before the block type, the code bits brought by the coding of the block direction are suppressed.

[0331] Like this, it is also possible to judge whether a block can be divided horizontally and vertically respectively based on a predetermined condition of whether the block is divisible or not. Then, when it is judged that the block can be divided only in one of the horizontal and vertical directions, it is also possible to skip writing the block direction to the bitstream. Furthermore, when it is judged that the block is not divisible in both the horizontal and vertical directions, it is also possible to skip writing not only the block direction but also the block type to the bitstream.

[0332] The predetermined condition of whether the block is divisible or not is defined, for example, by the size (number of pixels) or the number of divisions. This condition of whether the block is divisible or not can also be predefined in the standard specification. And the condition of whether the block is divisible or not can also be included in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The condition of whether the block is divisible or not can be fixed for all blocks, or can be dynamically switched according to the characteristics of the block (such as luminance and chrominance blocks) or the characteristics of the picture (such as I, P, B pictures), etc.

[0333] In step S3003, the block is divided into sub-blocks using the identified block type and the indicated block direction. In step S3004, the divided block is encoded.

[0334] [Encoding device]

[0335] Figure 15 is a block diagram showing the structure of an image / video encoding device according to Embodiment 2 or 3.

[0336] The video encoding device 5000 is a device that encodes an input image / video on a per-block basis and generates an encoded output bitstream. As Figure 15 shown, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.

[0337] The input image is input to an adder, and the added value is output to the transform unit 5001. The transform unit 5001 transforms the added value into frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block division type and direction can be associated with a block division mode, a block division type, or a block division direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantization values to the inverse quantization unit 5003 and the entropy encoding unit 5009.

[0338] The inverse quantization unit 5003 inverse-quantizes the quantization values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, transforms the frequency coefficients into sample values of the bitstream, and outputs the sample values to the adder.

[0339] The adder adds the sample values of the bitstream output from the intra-frame / inter-frame prediction units 5007 and 5008 to the predicted video / image values, and outputs the added values to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives the block segmentation type and direction, as well as the parameters related to the block segmentation type and direction. If the derived block segmentation type and direction are used, the block is divided into a plurality of sub-blocks. The intra-frame / inter-frame prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block segmentation type and direction derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.

[0340] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, and encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.

[0341] [Decoding process]

[0342] Figure 14 Represents the video decoding process related to Embodiment 3.

[0343] First, in step S4001, the first parameter is read from the bitstream. The first parameter identifies the segmentation type for dividing the first block into sub-blocks from among a plurality of segmentation types.

[0344] In the next step S4002, the second parameter indicating the segmentation direction is read from the bitstream. The second parameter follows the first parameter in the bitstream. The segmentation type may also form a segmentation pattern together with the segmentation direction. The segmentation type indicates the number and segmentation ratio of the sub-blocks used for dividing the block.

[0345] Figure 29 Shows an example of the segmentation type and segmentation direction for dividing an N×N pixel block in Embodiment 3. In Figure 29 ,(1), (2), (3) and (4) are different segmentation types, (1a), (2a), (3a) and (4a) are segmentation patterns with different segmentation types in the vertical direction, and (1b), (2b), (3b) and (4b) are segmentation patterns with different segmentation types in the horizontal direction. As Figure 29As shown, when the block ratio is 1:1 and the N×N pixel block is divided along the vertical direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (1a). When the block ratio is 1:1 and the N×N pixel block is divided along the horizontal direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (1b). When the block ratio is 1:3 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (2a). When the block ratio is 1:3 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (2b). When the block ratio is 3:1 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (3a). When the block ratio is 3:1 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the block mode (3b). When the block ratio is 1:2:1 and the N×N pixel block is divided along the vertical direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block mode (4a). When the block ratio is 1:2:1 and the N×N pixel block is divided along the horizontal direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the block mode (4b).

[0346] Figure 17 Represents the positions where the first parameter in the compressed video stream can be considered. As Figure 17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method for dividing a block into multiple sub-blocks. For example, the first parameter can contain an identifier for the above-mentioned block types. For example, the first parameter can contain a flag indicating whether the block is divided along the horizontal or vertical direction. The first parameter can also contain a parameter indicating whether the block is divided into more than 2 sub-blocks.

[0347] Figure 18 Represents the positions where the second parameter in the compressed video stream can be considered. As Figure 18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method for dividing a block into multiple sub-blocks. For example, the second parameter can contain a flag indicating whether the block is divided along the horizontal or vertical direction. That is, the second parameter can contain a parameter indicating the block division direction. The second parameter can also contain a parameter indicating whether the block is divided into more than 2 sub-blocks. As Figure 19 shown, the second parameter is configured in the bitstream following the first parameter.

[0348] Figure 30Shows the advantages of encoding the block type before the block direction compared to the case of encoding the block direction before the block type. In this example, when the block direction in the horizontal direction is invalidated due to an unsupported size (16×2 pixels), there is no need to encode the block direction. In this example, the block direction is determined to be the block direction in the vertical direction, and the block direction in the horizontal direction is invalidated. Encoding the block type before the block direction suppresses the code bits caused by the encoding of the block direction compared to the case of encoding the block direction before the block type.

[0349] In this way, it is also possible to determine whether a block can be divided in the horizontal and vertical directions respectively based on pre-determined conditions for block divisibility or indivisibility. Then, when it is determined that the block can be divided only in one of the horizontal and vertical directions, the reading of the block direction from the bitstream can also be skipped. Furthermore, when it is determined that the block cannot be divided in both the horizontal and vertical directions, the reading of the block type from the bitstream can also be skipped in addition to the reading of the block direction.

[0350] The pre-determined conditions for block divisibility or indivisibility are defined by, for example, the size (number of pixels) or the number of divisions. The conditions for block divisibility or indivisibility can also be pre-defined in the standard specifications. Also, the conditions for block divisibility or indivisibility can be included in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The conditions for block divisibility or indivisibility can be fixed for all blocks, or can be dynamically switched according to the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.

[0351] In step S4003, the block is divided into sub-blocks using the identified block type and the indicated division direction. In step S4004, the divided block is decoded.

[0352] [Decoding device]

[0353] Figure 16 Is a block diagram showing the structure of the video / image decoding device according to Embodiment 2 or 3.

[0354] The video decoding device 6000 is a device for decoding an input encoded bitstream for each block and outputting a video / image. The video decoding device 6000 is as Figure 16 shown and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block division determination unit 6008.

[0355] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.

[0356] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. Based on the block partitioning type and direction derived by the block segmentation determination unit 6008, the inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning type and direction can be associated with the block partitioning mode, block partitioning type, or block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006 and 6007, outputs the added values to the display, and outputs the added values to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005 and derives the block partitioning type and direction using the parameters decoded by the entropy decoding unit 6001. If the derived block partitioning type and direction are used, the block is divided into multiple sub-blocks. Furthermore, the intra / inter prediction units 6006 and 6007 perform prediction on the video / image region of the block to be decoded based on the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed according to the block partitioning type and direction derived by the block segmentation determination unit 6008.

[0357] (Embodiment 4)

[0358] In the above embodiments, each functional block can generally be implemented by an MPU and a memory, etc. In addition, the processing of each functional block is generally implemented by a program execution unit such as a processor reading and executing software (program) recorded in a recording medium such as a ROM. This software can be distributed by downloading, etc., or can be recorded in a recording medium such as a semiconductor memory for distribution. In addition, of course, each functional block can also be implemented by hardware (special-purpose circuit).

[0359] In addition, the processing described in each embodiment can be implemented by centralized processing using a single device (system), or can also be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, both centralized processing and distributed processing can be performed.

[0360] The form of the present invention is not limited to the above embodiments, and various changes can be made, and they are also included in the scope of the form of the present invention.

[0361] Next, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in the above-described embodiments and a system using the same will be described. The system is characterized by including an image encoding device that uses the image encoding method, an image decoding device that uses the image decoding method, and an image encoding / decoding device that includes both. Regarding other configurations in the system, they can be appropriately changed as needed.

[0362] [Usage Example]

[0363] Figure 33 FIG. is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The provision area of the communication service is divided into desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110 that are fixed wireless stations are respectively provided in each unit.

[0364] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may also connect by combining some of the above elements. The various devices may also be directly or indirectly connected to each other via a telephone network or short-range wireless without passing through the base stations ex106 to ex110 that are fixed wireless stations. In addition, a streaming media server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smart phone ex115 via the Internet ex101 or the like. In addition, the streaming media server ex103 is connected to terminals in a hotspot in an airplane ex117 via a satellite ex116.

[0365] In addition, a wireless access point or a hotspot or the like may be used instead of the base stations ex106 to ex110. In addition, the streaming media server ex103 may be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without passing through the satellite ex116.

[0366] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a portable phone, or a PHS (Personal Handyphone System) or the like corresponding to the modes of mobile communication systems generally referred to as 2G, 3G, 3.9G, 4G, and those referred to as 5G in the future.

[0367] The home appliance ex118 is a refrigerator or a device included in a household fuel cell cogeneration system, etc.

[0368] In the content supply system ex100, a terminal having a photographing function is connected to a streaming media server ex103 via a base station ex106 or the like, whereby live distribution or the like can be performed. In live distribution, terminals (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, and a terminal in an airplane ex117) perform the encoding process described in the above embodiments on still image or moving image content captured by the user using the terminal, multiplex the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and send the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.

[0369] On the other hand, the streaming media server ex103 performs stream distribution on the content data sent by a requested client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or a terminal in an airplane ex117 that can decode the data after the above encoding process. Each device that receives the distributed data performs a decoding process on the received data and reproduces it. That is, each device functions as an image decoding device according to one aspect of the present invention.

[0370] [Decentralized processing]

[0371] In addition, the streaming media server ex103 may be a plurality of servers or a plurality of computers, and perform decentralized processing or recording and distribution of data. For example, the streaming media server ex103 may be implemented by a CDN (Content Delivery Network), and content distribution is achieved through a network connecting many edge servers dispersed in the world to each other. In the CDN, a physically closer edge server is dynamically allocated according to the client. And by caching and distributing the content to the edge server, latency can be reduced. In addition, in the case of a certain error or a change in the communication state due to an increase in traffic or the like, the processing can be decentralized by multiple edge servers, or the distribution main body can be switched to another edge server, or a part of the network that has failed can be bypassed and the distribution can continue, so high-speed and stable distribution can be achieved.

[0372] In addition, not limited to the distributed processing of distributing itself, the encoding process of the captured data can be performed by each terminal, on the server side, or can be shared between them. As an example, usually two processing loops are performed in the encoding process. In the first loop, the complexity or the amount of code of the image in units of frames or scenes is detected. In addition, in the second loop, a process of improving the encoding efficiency while maintaining the image quality is performed. For example, by performing the first encoding process by the terminal and the second encoding process by the server that receives the content, it is possible to reduce the processing load in each terminal while improving the quality and efficiency of the content. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.

[0373] As other examples, the camera ex113 etc. extracts feature amounts from images, compresses the data regarding the feature amounts as metadata, and sends it to the server. The server, for example, judges the importance of the target based on the feature amounts and switches the quantization accuracy etc., and performs compression corresponding to the meaning of the image. The feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction at the time of re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed by the server.

[0374] As other examples, in a stadium, a shopping mall, a factory, etc., there are cases where there are multiple video data obtained by multiple terminals capturing substantially the same scene. In this case, the multiple terminals that performed the shooting are used, and other terminals and servers that did not perform the shooting are used as needed. For example, distributed processing is performed by separately allocating the encoding process in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture. Thereby, it is possible to reduce the delay and better achieve real-time performance.

[0375] In addition, since the multiple video data are of substantially the same scene, the server can also manage and / or instruct to refer to the video data captured by each terminal with each other. Or, it can also be that the server receives the encoded data from each terminal and changes the reference relationship between the multiple data, or corrects or replaces the picture itself and re-encodes it. Thereby, it is possible to generate a stream with improved quality and efficiency of each data.

[0376] In addition, the server can also perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server can change the MPEG-like encoding method to a VP-like one, or can change H.264 to H.265.

[0377] In this way, the encoding process can be performed by the terminal or one or more servers. Therefore, the following descriptions use terms such as "server" or "terminal" as the processing entity, but part or all of the processing performed by the server can also be performed by the terminal, and part or all of the processing performed by the terminal can also be performed by the server. In addition, regarding these, the same applies to the decoding process.

[0378] [3D, multi-angle]

[0379] In recent years, the situation of combining and using different scenes captured by multiple cameras ex113 and / or terminals such as smartphones ex115 that are roughly synchronized with each other, or images or videos of the same scene captured from different angles has increased. The images captured by each terminal are combined based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the images are consistent.

[0380] The server not only encodes two-dimensional moving images, but can also encode still images automatically or at a user-specified time based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative position relationship between the shooting terminals, it can not only generate the three-dimensional shape of the scene based on the images of the same scene captured from different angles in addition to two-dimensional moving images. In addition, the server can also encode the three-dimensional data generated by point cloud etc. separately, and can also select or reconstruct from the images captured by multiple terminals based on the results of identifying or tracking people or objects using the three-dimensional data to generate and send the images to the receiving terminal.

[0381] In this way, the user can not only arbitrarily select each image corresponding to each shooting terminal to view the scene, but also view the content of the image cut from an arbitrary viewpoint from the three-dimensional data reconstructed using multiple images or videos. Furthermore, similar to the images, the sound can also be collected from multiple different angles, and the server multiplexes and sends the sound from a specific angle or space with the image in accordance with the image.

[0382] In addition, in recent years, content that establishes a correspondence between the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) has also been popularized. In the case of VR images, the server separately creates viewpoint images for the right eye and the left eye, and can perform encoding that allows reference between the images at each viewpoint through Multi-View Coding (MVC) etc., or can encode them as different streams without mutual reference. When decoding different streams, it can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.

[0383] In the case of an AR image, it is also possible that the server overlaps virtual object information in the virtual space with camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device acquires or holds virtual object information and three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates overlapping data by smoothly connecting them. Alternatively, it is also possible that the decoding device sends the movement of the user's viewpoint to the server in addition to the delegation of virtual object information, and the server creates overlapping data according to the three-dimensional data held in the server, matching the received movement of the viewpoint, encodes the overlapping data, and distributes it to the decoding device. In addition, the overlapping data has an α value representing transmittance in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted through this part. Alternatively, the server can also set the RGB value of a specified value as the background like chroma key, and generate data with the part other than the target set as the background color.

[0384] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, on the server side, or can be shared between them. As an example, it is also possible that a certain terminal first sends a reception request to the server, and another terminal receives the content corresponding to the request and performs the decoding process, and sends the decoded signal to a device having a display. By dispersing the processing regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. In addition, as another example, it is also possible that a TV or the like receives large-size image data, and a personal terminal of the viewer decodes and displays a part of the area such as tiles after the picture is divided. Thus, it is possible to share the overall image while confirming one's own responsible area or the area that one wants to confirm in more detail at hand.

[0385] In addition, it is envisioned that in the future, in a situation where multiple wireless communications at short, medium, or long distances can be used both indoors and outdoors, using a distribution system standard such as MPEG-DASH, seamless reception of content is performed while appropriately switching data for the connected communication. Thus, the user can not only freely select a decoding device or a display device such as a monitor installed indoors and outdoors with their own terminal, but also perform real-time switching. In addition, based on its own position information, etc., it is possible to switch the decoding terminal and the display terminal for decoding. Thus, it is also possible to display map information on a part of the wall or the ground of a building next to a displayable device while moving towards the destination. In addition, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or the encoded data being replicated in an edge server of the content distribution service, it is possible to switch the bit rate of the received data.

[0386] [Scalable Coding]

[0387] Regarding the switching of content, use Figure 34 As shown, a scalable stream that is compression-encoded using the moving image encoding method represented in each of the above-described embodiments will be described. For the server, there may be multiple streams with the same content but different qualities as separate streams, or it may be a structure that switches content by utilizing the characteristics of a temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, the decoding side can freely switch between decoding low-resolution content and high-resolution content by determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. For example, when wanting to view the subsequent video that was viewed on a smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream to different layers, so the burden on the server side can be reduced.

[0388] Furthermore, in addition to the structure where pictures are encoded by each layer as described above and scalability with an enhancement layer existing above the base layer is achieved, the enhancement layer may include meta-information such as based on the statistical information of the image, and the decoding side generates high-quality content by super-resolution of the pictures in the base layer based on the meta-information. The super-resolution can be either an improvement in the signal-to-noise ratio at the same resolution or an expansion of the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filtering process, machine learning, or least-squares operation used in the super-resolution process, etc.

[0389] Alternatively, it may be a structure where pictures are segmented into tiles, etc. according to the meaning of objects, etc. within the image, and the decoding side only decodes a part of the area by selecting the tiles to be decoded. In addition, by saving the attributes of the object (person, car, ball, etc.) and the position within the image (coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including the object. For example, as Figure 35 shown, a data storage structure different from the pixel data such as the SEI message in HEVC is used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0390] In addition, the meta-information may be stored in units composed of multiple pictures such as a stream, sequence, or random access unit. Thereby, the decoding side can obtain the moment when a specific person appears within the video, etc., and by matching with the information of the picture unit, can determine the picture where the object exists and the position of the object within the picture.

[0391] [Optimization of Web Page]

[0392] Figure 36 It is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. Figure 37 It is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As Figure 36 and Figure 37 shown, there are cases where a web page includes a plurality of linked images that are links to image content, and the visible manner thereof varies depending on the viewing device. When a plurality of linked images can be seen on the screen, before the user explicitly selects a linked image, or before the linked image approaches near the center of the screen or the entire linked image enters the screen, the display device (decoding device) displays the still image or I picture that each content has as a linked image, or displays an image such as a gif animation using a plurality of still images or I pictures, or only receives the base layer and decodes and displays the image.

[0393] When a linked image is selected by the user, the display device decodes the base layer with the highest priority. Additionally, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device may also decode up to the enhancement layer. Furthermore, in order to ensure real-time performance or when the communication bandwidth is extremely tight before selection, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display) by only decoding and displaying the forward-referenced pictures (I pictures, P pictures, B pictures that only perform forward reference). Additionally, the display device may forcibly ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward references, and as the pictures received over time increase, perform normal decoding.

[0394] [Autonomous Driving]

[0395] Furthermore, in the case of transmitting and receiving still images or video data such as two-dimensional or three-dimensional map information for the autonomous driving or driving assistance of a vehicle, the receiving terminal may also receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them in correspondence. Additionally, the meta information may belong to a layer or may only be multiplexed with the image data.

[0396] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal are moving, the receiving terminal can switch between base stations ex106 to ex110 to perform seamless reception and decoding by sending the location information of the receiving terminal at the time of the reception request. Additionally, the receiving terminal can dynamically switch the degree of reception of the meta information or the degree of update of the map information according to the user's selection, the user's condition, or the state of the communication bandwidth.

[0397] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0398] [Distribution of Personal Content]

[0399] In addition, in the content supply system ex100, not only high-quality, long-duration content provided by video distribution operators can be distributed through unicast or multicast, but also low-quality, short-duration content provided by individuals can be distributed. In addition, it is conceivable that the amount of such personal content will increase in the future. In order to make personal content better, the server can also perform encoding after editing. This can be achieved, for example, through the following structure.

[0400] During or after shooting, the server performs recognition processing such as shooting error, scene search, meaning analysis, and object detection based on the original image or encoded data. And based on the recognition results, the server manually or automatically corrects focus deviation or camera shake, deletes scenes with low importance such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the tone. Based on the editing results, the server encodes the edited data. In addition, it is known that the viewing rate decreases if the shooting time is too long. The server can also automatically crop scenes with low importance as described above and scenes with little movement based on the image processing results according to the shooting time to make the content within a specific time range. Or, the server can also generate a summary based on the result of scene meaning analysis and encode it.

[0401] In addition, there are cases where personal content in its original state contains content that infringes on copyright, the moral rights of the author, or the right of portrait, etc., and there are also inconvenient situations for individuals such as the sharing range exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or at home to out-of-focus images. In addition, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user can specify the person or background area of the image that they want to process, and the server performs processing such as replacing the specified area with another image or blurring the focus. If it is a person, the image of the face part can be replaced while tracking the person in the moving image.

[0402] In addition, the audio-visual real-time requirement for personal content with a small amount of data is relatively strong. Therefore, although it also depends on the bandwidth, the decoding device first receives and decodes and reproduces the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period. In the case where the reproduction is looped or reproduced more than twice, the enhancement layer is also included to reproduce a high-quality image. In this way, if it is a scalable-encoded stream, an experience can be provided where, at the stage of not being selected or just starting to watch, it is a relatively rough moving image, but the stream gradually becomes smooth and the image quality improves. In addition to scalable encoding, the same experience can also be provided when the relatively rough stream reproduced for the first time and the second stream encoded with reference to the moving image of the first time are configured as one stream.

[0403] [Other usage examples]

[0404] In addition, these encoding or decoding processes are usually processed in the LSIex500 possessed by each terminal. The LSIex500 can be either a single chip or a structure composed of multiple chips. Additionally, software for motion image encoding or decoding can be loaded into a certain recording medium (such as a CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, when a smart phone ex115 is equipped with a camera, it is also possible to transmit the motion image data obtained by this camera. The motion image data at this time is the data after being encoded by the LSIex500 possessed by the smart phone ex115.

[0405] In addition, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the execution ability for a specific service. When the terminal does not correspond to the encoding method of the content or does not have the execution ability for a specific service, the terminal downloads a codec or application software, and then performs content acquisition and reproduction.

[0406] In addition, it is not limited to the content supply system ex100 via the Internet ex101. It is also possible to incorporate at least one of the motion image encoding device (image encoding device) or motion image decoding device (image decoding device) of the above-described embodiments into a digital broadcast system. Since the radio wave for broadcasting uses a satellite, etc. to carry multiplexed data that multiplexes video and audio for transmission and reception, there is a difference in being suitable for multicast compared to the structure of the content supply system ex100 that is easy for unicast. However, the same application can be made for the encoding process and decoding process.

[0407] [Hardware structure]

[0408] Figure 38 It is a diagram showing the smart phone ex115. In addition,Figure 39 FIG. Figure 39 is a diagram showing a structural example of the smart phone ex115. The smart phone ex115 has an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of shooting images and still images, and a display unit ex458 for displaying the images shot by the camera unit ex465 and decoding the data such as the images received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the shot images or still images, the recorded sound, the received images or still images, the encoded or decoded data such as e-mails, or a slot unit ex464 as an interface unit with the SIM ex468, and the SIM ex468 is used to identify the user and perform authentication for accessing various data represented by the network. In addition, an external memory may be used instead of the memory unit ex467.

[0409] In addition, the main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc. is connected to the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.

[0410] When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack and starts the smart phone ex115 to a state where it can operate.

[0411] The smart phone ex115 performs processing such as calls and data communication under the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal processing unit ex454 converts the voice signal collected by the voice input unit ex456 into a digital voice signal, performs spread spectrum processing on it using the modulation / demodulation unit ex452, and after the digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. In addition, the received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, the spread spectrum inverse processing is performed by the modulation / demodulation unit ex452, and after it is converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body, etc., and the same transmission and reception processing is performed. In the data communication mode, when sending video, still images, or video and voice, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method shown in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the voice signal processing unit ex454 encodes the voice signal collected by the voice input unit ex456 during the process of shooting video, still images, etc. by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a specified manner, and modulation processing and conversion processing are performed by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and it is transmitted via the antenna ex450.

[0412] When receiving an image attached to an email or a chat tool, or an image linked on a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 divides the multiplexed data into a bitstream of video data and a bitstream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method described in the above embodiments, and displays the image or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. In addition, the audio signal processing unit ex454 decodes the audio signal and outputs the sound from the audio output unit ex457. In addition, since real-time streaming media is becoming popular, depending on the user's situation, there may be occasions where the reproduction of sound is inappropriate in society. Therefore, as an initial value, a structure that does not reproduce the audio signal but only reproduces the video data is preferred. It is also possible to reproduce the sound synchronously only when the user performs an operation such as clicking on the video data.

[0413] In addition, here, the smart phone ex115 is taken as an example for explanation. However, as the terminal, three installation forms can be considered, namely, a transmitting terminal having only an encoder, a receiving terminal having only a decoder, in addition to a transceiver terminal having both an encoder and a decoder. Furthermore, in the digital broadcast system, it is assumed that multiplexed data in which audio data etc. are multiplexed in video data is received and transmitted for explanation. However, in the multiplexed data, character data etc. associated with the video can be multiplexed in addition to the audio data, and it is also possible to receive or transmit the video data itself instead of the multiplexed data.

[0414] In addition, it is assumed that the main control unit ex460 including the CPU controls the encoding or decoding process for explanation. However, in many cases, the terminal has a GPU. Therefore, it is also possible to configure a structure in which the performance of the GPU is utilized to process a larger area together by a memory shared by the CPU and the GPU, or a memory that manages the addresses in a shared manner. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization are performed together by the GPU in units of pictures instead of by the CPU.

[0415] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the above-mentioned processor has: a block segmentation determination unit that uses a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment the above-mentioned image read from the above-mentioned memory into a plurality of blocks, and the above-mentioned block segmentation pattern defines a segmentation type; and an encoding unit that encodes the above-mentioned plurality of blocks; the above-mentioned set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the above-mentioned first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the above-mentioned second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block that is one of the blocks obtained after the segmentation of the above-mentioned first block; in the above-mentioned block segmentation determination unit, when the above-mentioned segmentation number of the above-mentioned first block segmentation pattern is 3, the above-mentioned second block is the central block among the blocks obtained after the segmentation of the above-mentioned first block, and the above-mentioned segmentation direction of the above-mentioned second block segmentation pattern is the same as the above-mentioned segmentation direction of the above-mentioned first block segmentation pattern, the above-mentioned second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.

[0416] The parameter for identifying the above-mentioned second block segmentation pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction (horizontal or vertical) the above-mentioned block is segmented, and does not include a second flag indicating the segmentation number for segmenting the above-mentioned block.

[0417] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the above-mentioned processor has: a block segmentation determination unit that uses a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment the above-mentioned image read from the above-mentioned memory into a plurality of blocks, and the above-mentioned block segmentation pattern defines a segmentation type; and an encoding unit that encodes the above-mentioned plurality of blocks; the above-mentioned set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the above-mentioned first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the above-mentioned second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block that is one of the blocks obtained after the segmentation of the above-mentioned first block; when the above-mentioned segmentation number of the above-mentioned first block segmentation pattern is 3, the above-mentioned second block is the central block among the blocks obtained after the segmentation of the above-mentioned first block, and the above-mentioned segmentation direction of the above-mentioned second block segmentation pattern is the same as the above-mentioned segmentation direction of the above-mentioned first block segmentation pattern, the above-mentioned block segmentation determination unit does not use the above-mentioned second block segmentation pattern with a segmentation number of 2.

[0418] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division number; the block division determination unit restricts the use of the second block division pattern with the division number of 2.

[0419] The parameter for identifying the second block division pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is divided, and a second flag indicating whether the block is divided into two or more.

[0420] The above parameters in the encoding device according to an embodiment of the present disclosure may also be configured in the slice data.

[0421] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a block set composed of a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; when the first block set obtained using the first set of block division patterns is the same as the second block set obtained using the second set of block division patterns, the block division determination unit divides using only one of the first set of block division patterns and the second set of block division patterns.

[0422] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also divide using the block division pattern set with the smaller of the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0423] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also divide using the block division pattern set that appears first in a preset order in the first set of block division patterns and the second set of block division patterns when the first code amount is equal to the second code amount based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0424] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; in the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.

[0425] The parameter for identifying the second block division pattern in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided, and does not include a second flag indicating the division number for dividing the block.

[0426] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the block division determination unit does not use the second block division pattern with a division number of 2.

[0427] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks using a block division mode set obtained by combining one or more block division modes, where the block division mode defines a division type; and a decoding unit that decodes the plurality of blocks; the block division mode set includes a first block division mode and a second block division mode that respectively define a division direction and a division number; the block division determination unit restricts the use of the second block division mode with the division number of 2.

[0428] The parameter for identifying the second block division mode in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction (horizontal or vertical) the block is divided, and a second flag indicating whether the block is divided into two or more.

[0429] The above parameter in the decoding device according to an embodiment of the present disclosure may also be configured in slice data.

[0430] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a block set composed of a plurality of blocks using a block division mode set obtained by combining one or more block division modes, where the block division mode defines a division type; and a decoding unit that decodes the plurality of blocks; when the first block set obtained by using the first block division mode set is the same as the second block set obtained by using the second block division mode set, the block division determination unit divides using only one of the first block division mode set and the second block division mode set.

[0431] The above block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division mode set with the smaller of the first code amount of the first block division mode set and the second code amount of the second block division mode set, based on the first code amount of the first block division mode set and the second code amount of the second block division mode set.

[0432] The above block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division mode set that appears first in a preset order in the first block division mode set and the second block division mode set, when the first code amount is equal to the second code amount, based on the first code amount of the first block division mode set and the second code amount of the second block division mode set.

[0433] The encoding method according to an embodiment of the present disclosure may also be to use a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting patterns define the splitting type; encode the plurality of blocks; the set of block splitting patterns consists of a first block splitting pattern and a second block splitting pattern, the first block splitting pattern defines the splitting direction and the number of splits for splitting the first block, and the second block splitting pattern defines the splitting direction and the number of splits for splitting a second block, which is one of the blocks obtained after splitting the first block; in the splitting, when the number of splits in the first block splitting pattern is 3, the second block is the central block among the blocks obtained after splitting the first block, and the splitting direction of the second block splitting pattern is the same as the splitting direction of the first block splitting pattern, the second block splitting pattern only includes a block splitting pattern with the number of splits being 3.

[0434] The parameter for identifying the second block splitting pattern in the encoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is split, and does not include a second flag indicating the number of splits for splitting the block.

[0435] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting patterns define the splitting type; and a step of encoding the plurality of blocks; the set of block splitting patterns consists of a first block splitting pattern and a second block splitting pattern, the first block splitting pattern defines the splitting direction and the number of splits for splitting the first block, and the second block splitting pattern defines the splitting direction and the number of splits for splitting a second block, which is one of the blocks obtained after splitting the first block; in the step of performing the splitting, when the number of splits in the first block splitting pattern is 3, the second block is the central block among the blocks obtained after splitting the first block, and the splitting direction of the second block splitting pattern is the same as the splitting direction of the first block splitting pattern, the second block splitting pattern with the number of splits being 2 is not used.

[0436] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a plurality of blocks, where the block splitting patterns define the splitting type; and a step of encoding the plurality of blocks; the set of block splitting patterns includes a first block splitting pattern and a second block splitting pattern that respectively define the splitting direction and the number of splits; in the step of performing the splitting, the use of the second block splitting pattern with the number of splits being 2 is restricted.

[0437] In the encoding method according to an embodiment of the present disclosure, the parameter for identifying the second block segmentation pattern may also include a first flag indicating in which direction, horizontal or vertical, the block is segmented, and a second flag indicating whether the block is segmented into two or more blocks.

[0438] In the encoding method according to an embodiment of the present disclosure, the above parameter may also be configured in the slice data.

[0439] The encoding method according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment a picture read from a memory into a set of blocks composed of a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of encoding the plurality of blocks; in the step of performing the above segmentation, when the first set of blocks obtained using the first set of block segmentation patterns is the same as the second set of blocks obtained using the second set of block segmentation patterns, only one of the first set of block segmentation patterns and the second set of block segmentation patterns is used for segmentation.

[0440] In the step of performing the above segmentation in the encoding method according to an embodiment of the present disclosure, it is also possible to segment using the set of block segmentation patterns with the smaller of the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns, based on the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns.

[0441] In the step of performing the above segmentation in the encoding method according to an embodiment of the present disclosure, it is also possible to segment using the set of block segmentation patterns that appears first in a preset order among the first set of block segmentation patterns and the second set of block segmentation patterns, when the first code amount is equal to the second code amount, based on the first code amount of the first set of block segmentation patterns and the second code amount of the second set of block segmentation patterns.

[0442] The decoding method according to an embodiment of the present disclosure may also be to use a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; decode the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.

[0443] The parameter for identifying the second block segmentation pattern in the decoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is segmented, and does not include a second flag indicating the segmentation number for segmenting the block.

[0444] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with a segmentation number of 2 is not used.

[0445] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns includes a first block segmentation pattern and a second block segmentation pattern that respectively define a segmentation direction and a segmentation number; in the step of performing the above segmentation, the use of the second block segmentation pattern with a segmentation number of 2 is restricted.

[0446] The decoding method according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment an encoded signal read from a memory into a set of blocks composed of a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; in the step of performing the segmentation, when the first set of blocks obtained by using the first set of block segmentation patterns is the same as the second set of blocks obtained by using the second set of block segmentation patterns, only one of the first set of block segmentation patterns or the second set of block segmentation patterns is used for segmentation.

[0447] The picture compression program according to an embodiment of the present disclosure may also include: using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment a picture read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.

[0448] The picture compression program according to an embodiment of the present disclosure may also include: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to segment a picture read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of encoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for segmenting a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with a segmentation number of 2 is not used.

[0449] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, the block division patterns defining division types; and a step of encoding the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division quantity; in the step of performing the division, use of the second block division pattern with the division quantity of 2 is restricted.

[0450] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a set of blocks composed of a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, the block division patterns defining division types; and a step of encoding the plurality of blocks; in the step of performing the division, when a first block set obtained using a first set of block division patterns is the same as a second block set obtained using a second set of block division patterns, only one of the first set of block division patterns or the second set of block division patterns is used for division.

[0451] Industrial Applicability

[0452] It can be used for encoding / decoding of multimedia data, particularly for image and video encoding / decoding devices using block encoding / decoding.

[0453] Reference Numeral Explanation

[0454] 100 Encoding Device

[0455] 102 Division Unit

[0456] 104 Subtraction Unit

[0457] 106, 5001 Transformation Unit

[0458] 108, 5002 Quantization Unit

[0459] 110, 5009 Entropy Encoding Unit

[0460] 112, 5003, 6002 Inverse Quantization Unit

[0461] 114, 5004, 6003 Inverse Transformation Unit

[0462] 116 Addition Unit

[0463] 118, 5005, 6004 Block Memory

[0464] 120 Loop Filtering Unit

[0465] 122, 5006, 6005 Frame Memory

[0466] 124, 5007, 6006 Intra prediction unit

[0467] 126, 5008, 6007 Inter prediction unit

[0468] 128 Prediction control unit

[0469] 200 Decoding device

[0470] 202, 6001 Entropy decoding unit

[0471] 204 Inverse quantization unit

[0472] 206 Inverse transformation unit

[0473] 208 Addition unit

[0474] 210 Block memory

[0475] 212 Loop filtering unit

[0476] 214 Frame memory

[0477] 216 Intra prediction unit

[0478] 218 Inter prediction unit

[0479] 220 Prediction control unit

[0480] 5000 Video encoding device

[0481] 5010, 6008 Block segmentation decision unit

[0482] 6000 Video decoding device

Claims

1. An encoding device that encodes pictures, wherein, Comprising: A processor; and A memory; The above-mentioned processor performs the following processing: Using a set of block segmentation patterns obtained by combining one or more block segmentation patterns, the above-mentioned picture read from the above-mentioned memory is segmented into a plurality of blocks, and the above-mentioned block segmentation pattern defines the segmentation type; and Encoding the above-mentioned plurality of blocks; The above-mentioned set of block segmentation patterns consists of a first block segmentation pattern and a second block segmentation pattern. The above-mentioned first block segmentation pattern defines the segmentation direction and the number of segments for segmenting the first block, and the above-mentioned second block segmentation pattern defines the segmentation direction and the number of segments for segmenting one of the blocks obtained after the segmentation of the first block, i.e., the second block; The parameter for identifying the above-mentioned second block segmentation pattern includes a first flag indicating in which direction, the horizontal direction or the vertical direction, the above-mentioned second block is segmented; When the above-mentioned number of segments in the above-mentioned first block segmentation pattern is 3, the above-mentioned second block is the central block among the blocks obtained after the segmentation of the first block, and the above-mentioned segmentation direction of the above-mentioned second block segmentation pattern indicated by the above-mentioned first flag is the same as the above-mentioned segmentation direction of the above-mentioned first block segmentation pattern, the above-mentioned second block segmentation pattern includes a block segmentation pattern with the number of segments being 3, and does not include a block segmentation pattern with the number of segments being 2; When the above-mentioned number of segments in the above-mentioned first block segmentation pattern is 3, the above-mentioned second block is the central block among the blocks obtained after the segmentation of the first block, and the above-mentioned segmentation direction of the above-mentioned second block segmentation pattern indicated by the above-mentioned first flag is different from the above-mentioned segmentation direction of the above-mentioned first block segmentation pattern, the above-mentioned second block segmentation pattern includes a block segmentation pattern with the number of segments being 3 and a block segmentation pattern with the number of segments being 2.

2. A decoding device that decodes an encoded signal, wherein, Comprising: A processor; and A memory; The above-mentioned processor performs the following processing: Using a set of block segmentation patterns obtained by combining one or more block segmentation patterns, the above-mentioned encoded signal read from the above-mentioned memory is segmented into a plurality of blocks, and the above-mentioned block segmentation pattern defines the segmentation type; and Decoding the above-mentioned plurality of blocks; The above-mentioned set of block segmentation patterns consists of a first block segmentation pattern and a second block segmentation pattern. The above-mentioned first block segmentation pattern defines the segmentation direction and the number of segments for segmenting the first block, and the above-mentioned second block segmentation pattern defines the segmentation direction and the number of segments for segmenting one of the blocks obtained after the segmentation of the first block, i.e., the second block; The parameter for identifying the above-mentioned second block segmentation pattern includes a first flag indicating in which direction, the horizontal direction or the vertical direction, the above-mentioned second block is segmented; When the above-mentioned number of segments in the above-mentioned first block segmentation pattern is 3, the above-mentioned second block is the central block among the blocks obtained after the segmentation of the first block, and the above-mentioned segmentation direction of the above-mentioned second block segmentation pattern indicated by the above-mentioned first flag is the same as the above-mentioned segmentation direction of the above-mentioned first block segmentation pattern, the above-mentioned second block segmentation pattern includes a block segmentation pattern with the number of segments being 3, and does not include a block segmentation pattern with the number of segments being 2; When the number of divisions in the above-described first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern indicated by the first flag is different from the division direction of the first block division pattern, the second block division pattern includes a block division pattern with a division number of 3 and a block division pattern with a division number of 2.

3. A non-transitory storage medium that stores a bitstream and is readable by a computer, wherein the bitstream contains information for causing a computer that receives the bitstream to perform a decoding process; the information is for causing the computer to perform the following processes: using a set of block division patterns obtained by combining one or more block division patterns, dividing an encoded signal read from a memory into a plurality of blocks, the block division pattern defining a division type; and decoding the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defining a division direction and a division number for dividing a first block, and the second block division pattern defining a division direction and a division number for dividing a second block, which is one of the blocks obtained after the division of the first block; a parameter for identifying the second block division pattern includes a first flag indicating in which direction, horizontal or vertical, the second block is to be divided; When the number of divisions in the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern indicated by the first flag is the same as the division direction of the first block division pattern, the second block division pattern includes a block division pattern with a division number of 3 and does not include a block division pattern with a division number of 2; When the number of divisions in the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern indicated by the first flag is different from the division direction of the first block division pattern, the second block division pattern includes a block division pattern with a division number of 3 and a block division pattern with a division number of 2.

Citation Information

Patent Citations

  • Method and related device for coding and decoding image

    CN102223526A

  • Image encoding method, image decoding method, image encoding device, image decoding device, and image encoding / decoding device

    CN104429080A