Encoding Device, Decoding Device, and Storage Medium

By designing a coding scheme that divides the picture into 3 blocks and prohibits central block segmentation, the problem of low block segmentation information encoding efficiency in the prior art is solved, and more efficient image compression is achieved.

CN114630118BActive Publication Date: 2025-06-24PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210418760.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-20
Filing Date
2019-05-09
Publication Date
2025-06-24
Estimated Expiration
2039-05-09

AI Technical Summary

Technical Problem

The prior art has the problem of a decrease in compression efficiency in encoding block segmentation information, especially when the segmentation depth increases, signaling overhead increases, resulting in a decrease in image compression efficiency.

Method used

A design of an encoding device and a decoding device is proposed. By dividing the picture into three blocks in a certain direction, and the central block is prohibited from being divided into two sub-blocks, encoding is only based on the flag of whether to be divided, and signaling of the number of sub-blocks is avoided.

Benefits of technology

It improves the encoding efficiency of block segmentation information, reduces signaling overhead, and improves image compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114630118B_ABST
    Figure CN114630118B_ABST
Patent Text Reader

Abstract

The present invention provides an encoding device, a decoding device, and a storage medium. The encoding device encodes a picture, and includes: a processor; and a memory. The processor divides the picture read from the memory into three blocks in a first direction, where the first direction is one of the vertical direction and the horizontal direction. The three blocks include a central block provided between other blocks in a second direction perpendicular to the first direction. The processor divides the central block into three sub-blocks in the first direction according to segmentation information, encodes the three sub-blocks, and prohibits dividing the central block into two sub-blocks. The segmentation information includes a first flag indicating whether to divide the central block in the first direction, and the segmentation information does not include a second flag indicating the total number of sub-blocks obtained by dividing the central block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of a patent application for an invention titled "Encoding Device, Decoding Device, Encoding Method, Decoding Method, and Picture Compression Program" with an application date of May 9, 2019, an application number of 201980033751.4. Technical Field

[0002] The present disclosure relates to methods and apparatuses for encoding and decoding video and images using block partitioning. Background Art

[0003] In conventional image and video encoding methods, an image is generally divided into blocks, and encoding and decoding processes are performed at the block level. In recent video standard development, in addition to typical sizes of 8×8 or 16×16, encoding and decoding processes can be performed with various block sizes. For image encoding and decoding processes, a range of sizes from 4×4 to 256×256 can be used.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] In order to represent a range of sizes from 4×4 to 256×256, block partitioning information such as block partitioning patterns (e.g., quadtree, binary tree, and ternary tree) and partitioning flags (e.g., split flag) is determined and signaled for use in blocks. The overhead of this signaling increases as the partitioning depth increases. Moreover, the increased overhead degrades the video compression efficiency.

[0009] Therefore, an encoding device according to one aspect of the present disclosure provides an encoding device and the like that can improve the compression efficiency in the encoding of block partitioning information.

[0010] Means for Solving the Problems

[0011] An encoding device according to an aspect of the present disclosure encodes an image, and includes: a processor; and a memory. The processor performs the following processing: dividing the image read from the memory into three blocks in a first direction, where the first direction is one of a vertical direction and a horizontal direction, and the three blocks include a central block provided between other blocks in a second direction perpendicular to the first direction; dividing the central block into three sub-blocks in the first direction according to division information; encoding the three sub-blocks; and prohibiting dividing the central block into two sub-blocks. The division information includes a first flag indicating whether to divide the central block in the first direction, and the division information does not include a second flag indicating the total number of sub-blocks into which the central block is divided.

[0012] A decoding device according to an aspect of the present disclosure decodes an encoded signal, and includes: a processor; and a memory. The processor performs the following processing: dividing the encoded signal read from the memory into three blocks in a first direction, where the first direction is one of a vertical direction and a horizontal direction, and the three blocks include a central block provided between other blocks in a second direction perpendicular to the first direction; dividing the central block into three sub-blocks in the first direction according to the division information included in the encoded signal; decoding the three sub-blocks; and prohibiting dividing the central block into two sub-blocks. The division information includes a first flag indicating whether to divide the central block in the first direction, and the division information does not include a second flag indicating the total number of sub-blocks into which the central block is divided.

[0013] A non-transitory storage medium according to an aspect of the present disclosure is a non-transitory storage medium that stores a bitstream and is readable by a computer. The bitstream includes information for causing a computer that receives the bitstream to perform decoding processing. The information is for causing the computer to perform the following processing: dividing the encoded signal read from a memory into three blocks in a first direction, where the first direction is one of a vertical direction and a horizontal direction, and the three blocks include a central block provided between other blocks in a second direction perpendicular to the first direction; dividing the central block into three sub-blocks in the first direction according to the division information included in the encoded signal; decoding the three sub-blocks; and prohibiting dividing the central block into two sub-blocks. The division information includes a first flag indicating whether to divide the central block in the first direction, and the division information does not include a second flag indicating the total number of sub-blocks into which the central block is divided.

[0014] In addition, these inclusive or specific forms can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0015] Advantages of the Invention

[0016] According to the present invention, it is possible to improve the compression efficiency in the encoding of block segmentation information. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a block diagram showing the functional structure of the encoding apparatus according to Embodiment 1.

[0018] Figure 2 It is a diagram showing an example of the block segmentation in Embodiment 1.

[0019] Figure 3 It is a table showing transform basis functions corresponding to respective transform types.

[0020] Figure 4A It is a diagram showing an example of the shape of the filter used in ALF.

[0021] Figure 4B It is a diagram showing another example of the shape of the filter used in ALF.

[0022] Figure 4C It is a diagram showing another example of the shape of the filter used in ALF.

[0023] Figure 5A It is a diagram showing 67 intra prediction modes of intra prediction.

[0024] Figure 5B It is a flowchart for explaining the outline of the predicted image correction process based on OBMC processing.

[0025] Figure 5C It is a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC processing.

[0026] Figure 5D It is a diagram showing an example of FRUC.

[0027] Figure 6 It is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0028] Figure 7 It is a diagram for explaining pattern matching (template matching) between a template in the current picture and a block in a reference picture.

[0029] Figure 8This is a diagram for explaining a model assuming uniform linear motion.

[0030] Fig. 9A This is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on motion vectors of multiple adjacent blocks.

[0031] Fig. 9B This is a diagram for explaining an outline of a motion vector derivation process based on a merge mode.

[0032] Fig. 9C This is a conceptual diagram for explaining an outline of DMVR processing.

[0033] Fig.9D This is a diagram for explaining an outline of a predicted image generation method adopting a luminance correction process based on LIC processing.

[0034] Fig.10 This is a block diagram showing the functional structure of a decoding device according to Embodiment 1.

[0035] Fig.11 This is a flowchart showing an image encoding process according to Embodiment 2.

[0036] Fig.12 This is a flowchart showing an image decoding process according to Embodiment 2.

[0037] Fig.13 This is a flowchart showing an image encoding process according to Embodiment 3.

[0038] Fig.14 This is a flowchart showing an image decoding process according to Embodiment 3.

[0039] Fig.15 This is a block diagram showing the structure of an image / image encoding device according to Embodiment 2 or 3.

[0040] Fig.16 This is a block diagram showing the structure of an image / image decoding device according to Embodiment 2 or 3.

[0041] Fig.17 This is a diagram showing an example of a position where a first parameter in a compressed video stream in Embodiment 2 or 3 can be considered.

[0042] Fig.18 This is a diagram showing an example of a position where a second parameter in a compressed video stream in Embodiment 2 or 3 can be considered.

[0043] Fig.19 This is a diagram showing an example of a second parameter following a first parameter in Embodiment 2 or 3.

[0044] Fig. 20This is a diagram showing an example where the second block division mode is not selected for the division of a 2N×N pixel block as shown in step (2c) in Embodiment 2.

[0045] Fig.21 This is a diagram showing an example where the second block division mode is not selected for the division of an N×2N pixel block as shown in step (2c) in Embodiment 2.

[0046] Fig. 22 This is a diagram showing an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c) in Embodiment 2.

[0047] Fig.23 This is a diagram showing an example where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2C) in Embodiment 2.

[0048] Fig.24 This is a diagram showing an example where the block division mode selected when the second block division mode is not selected is used to divide a 2N×N pixel block as shown in step (3) in Embodiment 2.

[0049] Fig.25 This is a diagram showing an example where the block division mode selected when the second block division mode is not selected is used to divide an N×2N pixel block as shown in step (3) in Embodiment 2.

[0050] Fig.26 This is a diagram showing an example where the block division mode selected when the second block division mode is not selected is used to divide an N×N pixel block as shown in step (3) in Embodiment 2.

[0051] Fig. 27 This is a diagram showing an example where the block division mode selected when the second block division mode is not selected is used to divide an N×N pixel block as shown in step (3) in Embodiment 2.

[0052] Fig.28 This is a diagram showing an example of the block division mode used to divide an N×N pixel block in Embodiment 2. Fig.28 (a) to (h) thereof are diagrams showing different block division modes.

[0053] Fig.29 This is a diagram showing an example of the block division type and block division direction used to divide an N×N pixel block in Embodiment 3. (1), (2), (3), and (4) are different block division types, (1a), (2a), (3a), and (4a) are block division modes with different block division types in the vertical block division direction, and (1b), (2b), (3b), and (4b) are block division modes with different block division types in the horizontal block division direction.

[0054] Fig.30 This is a diagram showing the advantage of encoding the block type before the block direction as compared to encoding the block direction before the block type in Embodiment 3.

[0055] Fig.31A This is a diagram showing an example of dividing a block into sub - blocks using a set of block patterns that use fewer binary (bin) numbers in the encoding of the block pattern.

[0056] Fig.31B This is a diagram showing an example of dividing a block into sub - blocks using a set of block patterns that use fewer binary numbers in the encoding of the block pattern.

[0057] Fig.32A This is a diagram showing an example of dividing a block into sub - blocks using the block pattern set that first appears in a specified order of multiple block pattern sets.

[0058] Fig.32B This is a diagram showing an example of dividing a block into sub - blocks using the block pattern set that first appears in a specified order of multiple block pattern sets.

[0059] Fig.32C This is a diagram showing an example of dividing a block into sub - blocks using the block pattern set that first appears in a specified order of multiple block pattern sets.

[0060] Fig.33 This is an overall structural diagram of a content supply system that implements a content distribution service.

[0061] Fig.34 This is a diagram showing an example of an encoding structure in scalable coding.

[0062] Fig.35 This is a diagram showing an example of an encoding structure in scalable coding.

[0063] Fig.36 This is a diagram showing an example of a display screen of a web page.

[0064] Fig.37 This is a diagram showing an example of a display screen of a web page.

[0065] Fig.38 This is a diagram showing an example of a smart phone.

[0066] Fig.39 This is a block diagram showing an example of the structure of a smart phone.

[0067] Fig.40 This is a diagram showing a constraint example of a block pattern that divides a rectangular block into 3 sub - blocks.

[0068] Fig.41 It is a diagram showing a restrictive example of a block division pattern in which a block is divided into two sub - blocks.

[0069] Fig.42 It is a diagram showing a restrictive example of a block division pattern in which a square block is divided into three sub - blocks.

[0070] Fig.43 It is a diagram showing a restrictive example of a block division pattern in which a rectangular block is divided into two sub - blocks.

[0071] Fig.44 It is a diagram showing a restrictive example based on the division direction of a block division pattern in which a non - rectangular block is divided into two sub - blocks.

[0072] Fig.45 It is a diagram showing an example of an effective division direction of a block division in which a non - rectangular block is divided into two sub - blocks. Detailed implementation mode

[0073] Hereinafter, the implementation mode will be specifically described with reference to the accompanying drawings.

[0074] In addition, the implementation modes described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, configurations and connection forms of the constituent elements, steps, order of steps, etc. shown in the following implementation modes are examples and do not limit the meaning of the claims. In addition, among the constituent elements of the following implementation modes, the constituent elements not described in the independent claims representing the most general concept are described as arbitrary constituent elements.

[0075] (Embodiment 1)

[0076] First, as an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention described later, the outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention, and the processing and / or structure described in each form of the present invention can also be implemented in encoding devices and decoding devices different from Embodiment 1.

[0077] When applying the processing and / or structure described in each form of the present invention to Embodiment 1, for example, one of the following can also be performed.

[0078] (1) For the encoding device or decoding device of Embodiment 1, replace the constituent elements corresponding to the constituent elements described in each form of the present invention among the multiple constituent elements constituting the encoding device or decoding device with the constituent elements described in each form of the present invention;

[0079] (2) For the encoding device or decoding device of Embodiment 1, after any change such as addition, replacement, deletion, etc. of the processing for implementing a part of the plurality of constituent elements constituting the encoding device or decoding device, replace the constituent elements corresponding to the constituent elements described in each aspect of the present invention with the constituent elements described in each aspect of the present invention;

[0080] (3) For the method implemented by the encoding device or decoding device of Embodiment 1, after addition of processing, and / or any change such as replacement, deletion, etc. of a part of the plurality of processes included in the method, replace the process corresponding to the process described in each aspect of the present invention with the process described in each aspect of the present invention;

[0081] (4) Implement by combining a part of the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1 with the constituent elements described in each aspect of the present invention, the constituent elements having a part of the functions possessed by the constituent elements described in each aspect of the present invention, or the constituent elements implementing a part of the processing implemented by the constituent elements described in each aspect of the present invention;

[0082] (5) Implement by combining the constituent elements having a part of the functions possessed by a part of the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1, or the constituent elements implementing a part of the processing implemented by a part of the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1, with the constituent elements described in each aspect of the present invention, the constituent elements having a part of the functions possessed by the constituent elements described in each aspect of the present invention, or the constituent elements implementing a part of the processing implemented by the constituent elements described in each aspect of the present invention;

[0083] (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the process corresponding to the process described in each aspect of the present invention among the plurality of processes included in the method with the process described in each aspect of the present invention;

[0084] (7) Implement by combining a part of the plurality of processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the process described in each aspect of the present invention.

[0085] In addition, the implementation manners of the processes and / or structures described in the various aspects of the present invention are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or the moving image / image decoding device disclosed in Embodiment 1, or the processes and / or structures described in each aspect may be implemented independently. In addition, the processes and / or structures described in different aspects may be combined and implemented.

[0086] [Outline of Encoding Device]

[0087] First, the outline of the encoding device according to Embodiment 1 will be described. Figure 1 FIG. is a block diagram showing the functional configuration of the encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in units of blocks.

[0088] As Figure 1 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0089] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 may also be implemented as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0090] Hereinafter, each component included in the encoding device 100 will be described.

[0091] [Division Unit]

[0092] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits a picture into blocks of a fixed size (e.g., 128×128). Such blocks of the fixed size are sometimes called coding tree units (CTUs). Further, the splitting unit 102 splits each of the blocks of the fixed size into blocks of a variable size (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block splitting. Such blocks of the variable size are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks within a picture may be used as the processing units for CUs, PUs, and TUs.

[0093] Figure 2 FIG. is an example of block splitting according to Embodiment 1. In Figure 2 it, solid lines represent block boundaries based on quadtree block splitting, and dashed lines represent block boundaries based on binary tree block splitting.

[0094] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into four square 64×64 blocks (quadtree block splitting).

[0095] The upper-left 64×64 block is further vertically split into two rectangular 32×64 blocks, and the left 32×64 block is further vertically split into two rectangular 16×64 blocks (binary tree block splitting). As a result, the upper-left 64×64 block is split into two 16×64 blocks 11, 12 and a 32×64 block 13.

[0096] The upper-right 64×64 block is horizontally split into two rectangular 64×32 blocks 14, 15 (binary tree block splitting).

[0097] The lower-left 64×64 block is split into four square 32×32 blocks (quadtree block splitting). The upper-left block and the lower-right block among the four 32×32 blocks are further split. The upper-left 32×32 block is vertically split into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally split into two 16×16 blocks (binary tree block splitting). The lower-right 32×32 block is horizontally split into two 32×16 blocks (binary tree block splitting). As a result, the lower-left 64×64 block is split into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.

[0098] The lower-right 64×64 block 23 is not split.

[0099] As described above, in Figure 2 , block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quad-tree plus binary tree) partitioning.

[0100] In addition, in Figure 2 , one block is divided into four or two blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to this. For example, one block can also be divided into three blocks (ternary tree partitioning). Partitioning including such ternary tree partitioning is sometimes referred to as MBT (multi type tree) partitioning.

[0101] [Subtraction unit]

[0102] The subtraction unit 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of blocks divided by the partitioning unit 102. That is, the subtraction unit 104 calculates the prediction error (also referred to as the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction unit 104 outputs the calculated prediction error to the transformation unit 106.

[0103] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, there are also cases where a signal representing an image is also referred to as a sample.

[0104] [Transformation unit]

[0105] The transformation unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0106] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types and use a transform basis function corresponding to the selected transformation type to transform the prediction error into transform coefficients. Such a transformation is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).

[0107] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table representing transform basis functions corresponding to respective transform types. In Figure 3 , N represents the number of input pixels. The selection of a transform type from among these multiple transform types can depend on, for example, the type of prediction (intra prediction and inter prediction), or can depend on the intra prediction mode.

[0108] Information indicating whether to apply such EMT or AMT (e.g., referred to as an AMT flag) and information indicating the selected transform type are signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0109] Furthermore, the transform unit 106 can also perform a re - transform on the transform coefficients (transformation results). Such a re - transform is the case of what is called AST (adaptive secondary transform) or NSST (non - separable secondary transform). For example, the transform unit 106 performs a re - transform for each sub - block (e.g., 4×4 sub - block) included in a block of transform coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0110] Here, a Separable transform refers to a method of performing multiple transforms separately for each direction according to the number of dimensions of the input, and a Non - Separable transform refers to a method of treating two or more dimensions as one dimension and performing a transform together when the input is multi - dimensional.

[0111] For example, as an example of a Non - Separable transform, in the case where the input is a 4×4 block, it is regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.

[0112] Furthermore, similarly, a method of performing multiple Givens rotations on this permutation after regarding a 4×4 input block as a permutation having 16 elements (Hypercube Givens Transform) is also an example of a Non - Separable transform.

[0113] [Quantization Unit]

[0114] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a prescribed scan order, and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Further, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0115] The prescribed order is the order for quantization / inverse quantization of transform coefficients. For example, the prescribed scan order is defined by ascending order of frequency (order from low frequency to high frequency) or descending order of frequency (order from high frequency to low frequency).

[0116] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0117] [Entropy encoding unit]

[0118] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable length encoding on the quantized coefficients as the input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantized coefficients, for example, and performs arithmetic coding on the binary signal.

[0119] [Inverse quantization unit]

[0120] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients as the input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. Further, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0121] [Inverse transform unit]

[0122] The inverse transform unit 114 restores the prediction error by performing inverse transform on the transform coefficients as the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. Further, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0123] In addition, since information is lost due to quantization in the restored prediction error, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.

[0124] [Addition unit]

[0125] The adder unit 116 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 114, to the prediction sample, which is the input from the predictive control unit 128. Further, the adder unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may be referred to as a local decoded block.

[0126] [Block Memory]

[0127] The block memory 118 is a storage unit for storing blocks within an encoded object picture (hereinafter referred to as the current picture) that are referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the adder unit 116.

[0128] [Loop Filter Unit]

[0129] The loop filter unit 120 applies loop filtering to the block reconstructed by the adder unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the encoding loop (in-loop filtering), and includes, for example, deblocking filtering (DF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0130] In ALF, a least squares error filter for removing encoding distortion is employed. For example, for each 2×2 sub-block within the current block, one filter is selected from multiple filters based on the direction and activity of the locality-based gradient.

[0131] Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4), the classification value C (e.g., C = 5D + A) is calculated. Further, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).

[0132] The gradient direction value D is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantifying the addition result.

[0133] Based on the result of such classification, the filter to be used for the sub-block is determined from among the multiple filters.

[0134] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figure 4A to Figure 4C It is a diagram showing multiple examples of the shape of the filter used in ALF. Figure 4A It represents a 5×5 diamond-shaped filter, Figure 4BRepresents a 7×7 diamond-shaped filter, Figure 4C Represents a 9×9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0135] The on / off of the ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined at the CU level whether to use the ALF, and for chrominance difference, it is determined at the picture level whether to use the ALF. Information indicating the on / off of the ALF is signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of the ALF need not be limited to the picture level or CU level and may also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0136] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Additionally, the signaling of the coefficient sets need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0137] [Frame memory]

[0138] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0139] [Intra-frame prediction unit]

[0140] The intra-frame prediction unit 124 performs intra-frame prediction (also referred to as intra-picture prediction) of the current block by referring to the block within the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block and outputs the intra-frame prediction signal to the prediction control unit 128.

[0141] For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of predefined intra-frame prediction modes. The plurality of intra-frame prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0142] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).

[0143] The multiple directional prediction modes include, for example, the 33-direction prediction modes defined by the H.265 / HEVC standard. Additionally, the multiple directional prediction modes may also include 32-direction prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 5A It is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined by the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0144] In addition, in the intra prediction of chrominance blocks, luminance blocks may also be referred to. That is, the chrominance components of the current block may be predicted based on the luminance component of the current block. Such intra prediction is sometimes referred to as CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block (e.g., called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0145] The intra prediction unit 124 may also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intraprediction combination). Information indicating whether PDPC is used (e.g., called the PDPC flag) is signaled at the CU level, for example. Additionally, the signaling of this information does not need to be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0146] [Inter prediction unit]

[0147] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block by referring to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or a sub-block within the current block (e.g., a 4×4 block). For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. And the inter prediction unit 126 uses the motion information (e.g., motion vector) obtained through motion estimation to perform motion compensation, thereby generating the inter prediction signal for the current block or sub-block. And the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0148] Motion information used in motion compensation is signaled. A motion vector predictor may also be used in the signaling of motion vectors. That is, the difference between a motion vector and a predicted motion vector may also be signaled.

[0149] In addition, it is also possible to generate an inter prediction signal by using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks. Specifically, a prediction signal based on the motion information obtained by motion estimation and a prediction signal based on the motion information of adjacent blocks may be weighted and added to generate an inter prediction signal in units of sub-blocks within the current block. Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0150] In such an OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is adopted (e.g., referred to as the OBMC flag) is signaled at the CU level. Additionally, the levels at which these pieces of information are signaled do not need to be limited to the sequence level and the CU level, and may also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0151] The OBMC mode will be described in more detail. Figure 5B and Figure 5C are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC processing.

[0152] First, using the motion vector (MV) assigned to the coding target block, a predicted image (Pred) obtained by normal motion compensation is acquired.

[0153] Next, the motion vector (MV_L) of the already encoded left adjacent block is adopted for the coding target block to obtain a predicted image (Pred_L), and the first correction of the predicted image is performed by weighted superposition of the above predicted image and Pred_L.

[0154] Similarly, the motion vector (MV_U) of the already encoded upper adjacent block is adopted for the coding target block to obtain a predicted image (Pred_U), and the second correction of the predicted image is performed by weighted superposition of the predicted image after the first correction and Pred_U, and this is used as the final predicted image.

[0155] In addition, a method of two-stage correction using the left adjacent block and the upper adjacent block has been described here, but it may also be configured to perform more than two-stage corrections using the right adjacent block and the lower adjacent block.

[0156] In addition, the area for superposition may not be the pixel area of the entire block, but only a part of the area near the block boundary.

[0157] In addition, although the prediction image correction process based on one reference picture has been described here, the same applies to the case of correcting the prediction image based on multiple reference pictures. After obtaining the corrected prediction images according to the respective reference pictures, the obtained prediction images are further superimposed to obtain the final prediction image.

[0158] In addition, the processing target block described above may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.

[0159] As a method for determining whether to adopt the OBMC process, for example, there is a method of using a signal indicating whether to adopt the OBMC process, that is, obmc_flag. As a specific example, in an encoding device, it is determined whether the encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as obmc_flag and encoding is performed using the OBMC process. If it does not belong to a region with complex motion, the value 0 is set as obmc_flag and encoding is performed without using the OBMC process. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag described in the stream and switching whether to adopt the OBMC process according to its value.

[0160] In addition, the motion information may not be signaled and may be derived on the decoding device side. For example, the merge mode defined by the H.265 / HEVC standard may be used. In addition, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.

[0161] Here, the mode of performing motion estimation on the decoding device side is described. The mode of performing motion estimation on the decoding device side may be a mode called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0162] In Figure 5DThis represents an example of FRUC processing. First, with reference to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector are generated (which may also be shared with the merge list). Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0163] And, based on the motion vector of the selected candidate, the motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, the motion vector for the current block may also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV may also be searched by the same method, and in the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and this is used as the final MV for the current block. Additionally, a structure that does not perform this process may also be implemented.

[0164] The exact same process may also be performed when processing is carried out in units of sub - blocks.

[0165] Regarding the evaluation value, it is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region in the reference picture corresponding to the motion vector and a specified region. Additionally, it may also be that, in addition to the difference value, other information is used to calculate the evaluation value.

[0166] As the pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.

[0167] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Thus, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate, the region in another reference picture along the motion trajectory of the current block is used.

[0168] Figure 6 This is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along the motion trajectory. As Figure 6As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the above candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the final MV.

[0169] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0170] In the second pattern matching, pattern matching is performed between the template in the current picture (the block adjacent to the current block in the current picture, such as the upper and / or left adjacent block) and the block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as the specified area for calculating the evaluation value for the above candidate.

[0171] Figure 7 It is a diagram for explaining an example of pattern matching (template matching) between the template in the current picture and the block in the reference picture. As Figure 7 shown, in the second pattern matching, the motion vector of the current block is derived by searching for the block in the reference picture (Ref0) that most closely matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the left adjacent and / or upper adjacent encoded area and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, the evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.

[0172] Information indicating whether such a representation adopts the FRUC mode (e.g., called the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the method of pattern matching (the first pattern matching or the second pattern matching) (e.g., called the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0173] Here, a mode of deriving a motion vector based on a model assuming uniform linear motion is described. This mode includes a case called BIO (bi-directional optical flow).

[0174] Figure 8 It is a diagram for explaining a model assuming uniform linear motion. In Figure 8 , (v x , v y ) represents the velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.

[0175] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0176] [Equation 1]

[0177]

[0178] Here, I (k)Represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation represents that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a merge list or the like are corrected in pixel units.

[0179] In addition, the motion vector can also be derived on the decoding device side by a method different from the derivation of the motion vector based on a model assuming uniform linear motion. For example, the motion vector can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0180] Here, a mode of deriving the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks will be described. This mode includes a case called affine motion compensation prediction mode.

[0181] Fig. 9A is a diagram for explaining the derivation of the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks. In Fig. 9A , the current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. And, using the two motion vectors v0 and v1, the motion vectors (v x , v y ) of each sub-block within the current block are derived by the following equation (2).

[0182] [Equation 2]

[0183]

[0184] Here, x and y respectively represent the horizontal position and the vertical position of the sub-block, and w represents a preset weight coefficient.

[0185] In such an affine motion compensation prediction mode, several modes with different methods for deriving the motion vectors of the upper left and upper right control points may also be included. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. In addition, the signaling of the information indicating this affine motion compensation prediction mode does not need to be limited to the CU level, and it can also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0186] [Prediction control unit]

[0187] The prediction control unit 128 selects one of the intra prediction signal and the inter prediction signal, uses the selected signal as the prediction signal, and outputs it to the subtraction unit 104 and the addition unit 116.

[0188] Here, an example of deriving the motion vector of the coding target picture through the merge mode is described. Fig. 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.

[0189] First, a prediction MV list registering candidates of the prediction MV is generated. As candidates of the prediction MV, there are the MV of multiple coded blocks spatially located around the coding target block, that is, the spatially adjacent prediction MV, the MV of a block near the projection of the position of the coding target block in the coded reference picture, that is, the temporally adjacent prediction MV, the MV generated by combining the MV values of the spatially adjacent prediction MV and the temporally adjacent prediction MV, that is, the combined prediction MV, and the MV with a value of zero, that is, the zero prediction MV, etc.

[0190] Next, by selecting one prediction MV from the multiple prediction MVs registered in the prediction MV list, it is determined as the MV of the coding target block.

[0191] Furthermore, in the variable length coding unit, the merge_idx, which is a signal indicating which prediction MV is selected, is described in the stream and coded.

[0192] In addition, in Fig. 9B The prediction MVs registered in the prediction MV list described are an example, and the number may be different from that in the figure, or it may be a structure that does not include some types of the prediction MVs in the figure, or a structure with prediction MVs other than the types in the figure added.

[0193] In addition, the MV of the coding target block derived through the merge mode can also be used for the DMVR process described later to determine the final MV.

[0194] Here, an example of determining the MV using the DMVR process is described.

[0195] Fig. 9C It is a conceptual diagram for explaining the outline of the DMVR process.

[0196] First, the optimal MVP set for the processing target block is used as the candidate MV. According to the above candidate MV, reference pixels are obtained from the first reference picture of the processed picture in the L0 direction and the second reference picture of the processed picture in the L1 direction respectively, and a template is generated by taking the average of each reference pixel.

[0197] Next, using the above template, the peripheral regions of the candidate MVs of the first reference picture and the second reference picture are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, regarding the cost value, it is calculated using the difference values between the respective pixel values of the template and the respective pixel values of the search region, and the MV value, etc.

[0198] In addition, in the encoding device and the decoding device, the outline of the processing described here is basically common.

[0199] In addition, even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV and derive the final MV, other processing can also be used.

[0200] Here, a mode of generating a predicted image using the LIC processing is described.

[0201] Fig.9D It is a diagram for explaining the outline of a method for generating a predicted image using a luminance correction process based on the LIC process.

[0202] First, an MV for obtaining a reference image corresponding to the block to be encoded from a reference picture that is an encoded picture is derived.

[0203] Next, for the block to be encoded, using the luminance pixel values of the left adjacent and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, information indicating how the luminance values change in the reference picture and the picture to be encoded is extracted, and a luminance correction parameter is calculated.

[0204] By performing a luminance correction process on the reference image in the reference picture specified by the MV using the above luminance correction parameter, a predicted image for the block to be encoded is generated.

[0205] In addition, Fig.9D The shape of the above peripheral reference region in [[ ]] is an example, and shapes other than it can also be used.

[0206] In addition, the processing of generating a predicted image based on one reference picture is described here, but the same applies to the case of generating a predicted image based on multiple reference pictures. After performing a luminance correction process on the reference images obtained from each reference picture in the same method, a predicted image is generated.

[0207] As a method for determining whether to perform LIC processing, for example, there is a method of using lic_flag as a signal indicating whether to perform LIC processing. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region where a luminance change has occurred. In the case where it belongs to a region where a luminance change has occurred, the value 1 is set as lic_flag, and encoding is performed using LIC processing. In the case where it does not belong to a region where a luminance change has occurred, the value 0 is set as lic_flag, and encoding is performed without using LIC processing. On the other hand, in a decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to use LIC processing according to the value.

[0208] As another method for determining whether to perform LIC processing, for example, there is also a method of determining according to whether LIC processing has been performed on surrounding blocks. As a specific example, in the case where an encoding target block is in the merge mode, it is determined whether the surrounding encoded blocks selected at the time of deriving the MV in the merge mode processing have been encoded using LIC processing, and according to the result, encoding is performed by switching whether to use LIC processing. In addition, in the case of this example, the processing in decoding is exactly the same.

[0209] [Outline of Decoding Device]

[0210] Next, an outline of a decoding device capable of decoding the encoded signal (encoded bitstream) output from the above encoding device 100 will be described. Fig.10 is a block diagram showing the functional configuration of the decoding device 200 according to Embodiment 1. The decoding device 200 is a moving image / image decoding device that decodes moving images / images in units of blocks.

[0211] As shown in Fig.10 , the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0212] The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. In addition, the decoding device 200 can also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0213] Hereinafter, each component included in the decoding device 200 will be described.

[0214] [Entropy decoding unit]

[0215] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Next, the entropy decoding unit 202 de-binarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in units of blocks.

[0216] [Inverse quantization unit]

[0217] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoding target block (hereinafter referred to as the current block) that is an input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. And the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0218] [Inverse transform unit]

[0219] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients that are an input from the inverse quantization unit 204.

[0220] For example, when the information read from the encoded bitstream indicates the use of EMT or AMT (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type read.

[0221] In addition, for example, when the information read from the encoded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0222] [Addition unit]

[0223] The addition unit 208 reconstructs the current block by adding the prediction error that is an input from the inverse transform unit 206 and the prediction sample that is an input from the prediction control unit 220. And the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0224] [Block memory]

[0225] The block memory 210 is a storage unit for storing blocks within the decoding target picture (hereinafter referred to as the current picture) that are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.

[0226] [Loop Filter Unit]

[0227] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder 208, and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.

[0228] When the information indicating the ON / OFF of ALF read from the coded bitstream indicates ON, one filter is selected from a plurality of filters based on the direction and activity of the gradient of locality, and the selected filter is applied to the reconstructed block.

[0229] [Frame Memory]

[0230] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0231] [Intra Prediction Unit]

[0232] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the coded bitstream, referring to the block in the current picture stored in the block memory 210, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 generates an intra prediction signal by referring to the samples (e.g., luminance values, chrominance differences) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0233] In addition, when the intra prediction mode referring to the luminance block is selected in the intra prediction of the chrominance block, the intra prediction unit 216 may also predict the chrominance component of the current block based on the luminance component of the current block.

[0234] Furthermore, when the information read from the coded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical directions.

[0235] [Inter Prediction Unit]

[0236] The inter prediction unit 218 predicts the current block by referring to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using the motion information (e.g., motion vector) read from the coded bitstream, thereby generating an inter prediction signal for the current block or sub-block, and outputs the inter prediction signal to the prediction control unit 220.

[0237] In addition, when the information decoded from the coded bitstream indicates the use of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks.

[0238] In addition, when the information decoded from the coded bitstream indicates the use of the FRUC mode, the inter prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) decoded from the coded stream, thereby deriving motion information. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0239] In addition, when the BIO mode is adopted, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when the information decoded from the coded bitstream indicates the use of the affine motion compensation prediction mode, the inter prediction unit 218 derives motion vectors in sub-block units based on the motion vectors of multiple adjacent blocks.

[0240] [Prediction control unit]

[0241] The prediction control unit 220 selects one of the intra prediction signal and the inter prediction signal and outputs the selected signal as a prediction signal to the addition unit 208.

[0242] (Embodiment 2)

[0243] Regarding the encoding process and decoding process of Embodiment 2, refer to Fig.11 and Fig.12 Specifically, regarding the encoding device and decoding device of Embodiment 2, refer to Fig.15 and Fig.16 Specifically described.

[0244] [Encoding process]

[0245] Fig.11 Represents the video encoding process related to Embodiment 2.

[0246] First, in step S1001, a first parameter for identifying a partitioning mode for partitioning a first block into a plurality of sub-blocks from among a plurality of partitioning modes is written into the bitstream. When using a partitioning mode, the block is divided into a plurality of sub-blocks. When using different partitioning modes, the block is divided into a plurality of sub-blocks with different shapes, different heights, or different widths.

[0247] Fig.28 Represents an example of a partitioning mode for partitioning an N×N pixel block in Embodiment 2. In Fig.28 (a) to (h) represent different partitioning modes. As Fig.28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and "N" can take any value that is an integer multiple of 4 from 8 to 128) is divided into two sub-blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, a block of N×N pixels is divided into a sub-block of N / 4×N pixels (for example, 4×16 pixels) and a sub-block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, a block of N×N pixels is divided into a sub-block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, a block of N×N pixels is divided into a sub-block of (N / 4)×N pixels (for example, 4×16 pixels), a sub-block of N / 2×N pixels (for example, 8×16 pixels), and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, a block of N×N pixels is divided into two sub-blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, a block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels) and a sub-block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, a block of N×N pixels is divided into a sub-block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub-block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, a block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels), a sub-block of N×N / 2 pixels (for example, 16×8 pixels), and a sub-block of N×N / 4 pixels (for example, 16×4 pixels).

[0248] Next, in step S1002, it is determined whether the first parameter identifies the first block mode.

[0249] Next, in step S1003, based at least on the determination of whether the first parameter identifies the first block mode, it is determined whether the second block mode is not selected as a candidate for dividing the second block.

[0250] Two different block mode sets may divide a block into sub-blocks of the same shape and size. For example, as Fig.31A shown, the sub-blocks of (1b) and (2c) have the same shape and size. One block mode set can contain at least two block modes. For example, as Fig.31A shown in (1a) and (1b) of Fig.31AAs shown in (2a), (2b), and (2c), other block pattern sets can follow the vertical binary tree partitioning to include the vertical binary tree partitioning of two sub-blocks. Each block pattern set results in sub-blocks of the same shape and size.

[0251] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are different binary numbers or different numbers of bits when encoded in a bitstream, select the block pattern set with fewer binary numbers or fewer bits. Additionally, the binary numbers and the number of bits correspond to the code amount.

[0252] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are the same binary number or the same number of bits when encoded in a bitstream, select the block pattern set that first appears in a specified order of multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.

[0253] Fig.31A and Fig.31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of the block pattern. In this example, when the left N×N pixel block is vertically divided into two sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Fig.31B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).

[0254] Figures 32A to 32C is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that first appears in a specified order of multiple block pattern sets. In this example, when the 2N×N / 2 pixel block is vertically divided into three sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Fig.32B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Fig.32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of multiple block pattern sets can also be fixed and signaled within the bitstream.

[0255] Fig. 20This represents an example in Embodiment 2 where the second block division mode is not selected for the division of a 2N×N pixel block as shown in step (2c). As Fig. 20 shown, it is possible to use the first division method (i) to equally divide a 2N×2N pixel block (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, it is possible to use the second division method (ii) to horizontally equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). Here, in the second division method (ii), when the upper 2N×N pixel block (the first block) is vertically divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for vertically dividing the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for possible block division modes. This is because sub-block sizes identical to those obtained by the four-fold division using the first division method (i) will be generated.

[0256] As described above, in Fig. 20 , when the first block is vertically equally divided into 2 sub-blocks if the first block division mode is used, and the second block adjacent to the first block in the vertical direction is vertically equally divided into 2 sub-blocks if the second block division mode is used, the second block division mode is not selected as a candidate.

[0257] Fig.21 This represents an example in Embodiment 2 where the second block division mode is not selected for the division of an N×2N pixel block as shown in step (2c). As Fig.21 shown, it is possible to use the first division method (i) to equally divide a 2N×2N pixel block into 4 sub-blocks of N×N pixels as in step (1a). Also, it is possible to use the second division method (ii) to vertically equally divide a 2N×2N pixel block into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second division method (ii), when the left N×2N pixel block (the first block) is horizontally divided into 2 sub-blocks of N×N pixels by the first block division mode as in step (2b), in step (2c), the second block division mode for horizontally dividing the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for possible block division modes. This is because sub-block sizes identical to those obtained by the four-fold division using the first division method (i) will be generated.

[0258] As described above, in Fig.21In the case where, if the first block division mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block division mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the horizontal direction, the second block division mode is not selected as a candidate.

[0259] Fig.40 Indicates an example of dividing a 4N×2N block in [ Fig. 20 into three parts in a ratio of 1:2:1, such as N×2N, 2N×2N, N×2N. Here, when the upper block is divided into three parts, the block division mode of dividing the lower block into three parts in a ratio of 1:2:1 is not selected as a candidate for possible block division modes. The three-way division can also be in a ratio different from 1:2:1. Furthermore, it can be divided into more than three parts, it can also be divided into two parts, or it can be in a ratio different from 1:1, such as 1:2 or 1:3. Fig.40 This is an example of dividing first in the horizontal direction, but the same constraints can also be applied when dividing first in the vertical direction.

[0260] Fig.41 and Fig.42 Indicates an example of applying the same constraint when the first block is rectangular.

[0261] Fig.43 This is the second constraint example when a square is divided into three parts in the vertical direction and then equally divided into two parts in the horizontal direction. When applying Fig.43 the constraint, in Fig.40 , it is possible to select the block division mode of dividing the lower block of 4N×2N into three parts in a ratio of 1:2:1. It is also possible to separately encode the information indicating which of the constraints of applying Fig.40 and Fig.43 into the header information, etc. Alternatively, it is also possible to apply constraints to reduce the code amount of the information representing the block division. For example, if it is assumed that the code amounts of the information representing the block division in Case 1 and Case 2 are as follows, the division in Case 1 is made valid and the division in Case 2 is made invalid. That is, apply Fig.43 the constraint.

[0262] (Case 1) (1) Divide the square into two parts in the horizontal direction, and then (2) divide the upper and lower two rectangular blocks vertically into three parts: (1) Direction information: 1 bit, division quantity information: 1 bit, (2) (Direction information: 1 bit, division quantity information: 1 bit) × 2, a total of 6 bits

[0263] (Case 2) (1) Divide the square into two parts in the vertical direction, and then (2) divide the left, middle, and right rectangular blocks horizontally into two parts: (1) Direction information: 1 bit, division quantity information: 1 bit, (2) (Direction information: 1 bit, division quantity information: 1 bit) × 3, a total of 8 bits

[0264] Alternatively, there is a case where the optimal block is determined while selecting the block mode in a specified order during encoding. For example, 2-way splitting may be attempted first, followed by 3-way or 4-way splitting (bisecting horizontally and vertically), etc. At this time, before the attempt of 3-way splitting as in Fig.43 , the attempt starting from 2-way splitting as in the example of Fig.40 has already been carried out. Thus, in the attempt starting from 2-way splitting, the splitting of horizontally bisecting and then vertically trisecting the two upper and lower blocks has already been attempted, so the constraint of Fig.43 is applied. In this way, the constraint method for selection can also be determined based on the specified encoding method.

[0265] In Fig.44 , an example is shown where the block modes that can be selected for the same direction as the first block mode are restricted in the second block mode. Here, the first block mode is 3-way splitting in the vertical direction. At this time, 2-way splitting cannot be selected as the second block mode. On the other hand, for the vertical direction, which is a direction different from the first block mode, 2-way splitting can be selected ( Fig.45 ).

[0266] Fig. 22 An example is shown where the second block mode is not selected for the splitting of an N×N pixel block as shown in step (2c) in Embodiment 2. As shown in Fig. 22 , the first splitting method (i) can be used to vertically split a 2N×N pixel block (e.g., 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (e.g., sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels). Also, the second splitting method (ii) can be used to split a 2N×N pixel block into two N×N pixel sub-blocks as shown in step (2a). In the first splitting method (i), the central N×N pixel block can be vertically split into two N / 2×N pixel (e.g., 4×8 pixel) sub-blocks in step (1b). In the second splitting method (ii), when the left N×N pixel block (the first block) is vertically split into two N / 2×N pixel sub-blocks as shown in step (2b), in step (2c), the splitting mode of vertically splitting the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks is not selected as a candidate for possible splitting modes. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, i.e., four N / 2×N pixel sub-blocks.

[0267] As described above, in Fig. 22In the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block mode is not selected as a candidate.

[0268] Fig.23 This shows an example in Embodiment 2 where the second block mode is not selected for the division of an N×N pixel block as shown in step (2c). As Fig.23 shown, the first splitting method (i) can be used to split an N×2N pixel (e.g., 8×16 pixels, where "N" can take any value that is an integer multiple of 4 from 8 to 128) into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (e.g., sub-blocks of 8×4 pixels, 8×8 pixels, and 8×4 pixels). Also, the second splitting method can be used to split it into two sub-blocks of N×N pixels. In the first splitting method (i), the central N×N pixel block can be split into two sub-blocks of N×N / 2 pixels as shown in step (1b). In the second splitting method (ii), when the upper N×N pixel block (the first block) is horizontally split into two sub-blocks of N×N / 2 pixels as shown in step (2b), in step (2c), the block mode of horizontally splitting the lower N×N pixel block (the second block) into two sub-blocks of N×N / 2 pixels is not selected as a candidate for possible block modes. This is because sub-blocks of the same size as those obtained by the first splitting method (i) will be generated, i.e., four sub-blocks of N×N / 2 pixels.

[0269] As described above, in Fig.23 the case where, if the first block mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block mode is not selected as a candidate.

[0270] If it is determined that the second block mode is selected as a candidate for splitting the second block (No in S1003), then in step S1004, a block mode is selected from among multiple block modes including the second block mode as a candidate. In step S1005, a second parameter representing the selection result is written into the bitstream.

[0271] If it is determined that the second block mode is not selected as a candidate for splitting the second block (Yes in S1003), then in step S1006, a block mode different from the second block mode is selected for splitting the second block. Among the block modes selected here, the block is split into sub-blocks having a different shape or different size compared to the sub-blocks generated by the second block mode.

[0272] Fig.24 This shows an example of using the selected block mode (when the second block mode is not selected) as shown in step (3) in Embodiment 2 to divide a 2N×N pixel block. As Fig.24 shown, the selected block mode can divide the current 2N×N pixel block (the lower block in this example) as Fig.24 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Fig.24 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.

[0273] Fig.25 This shows an example of using the selected block mode (when the second block mode is not selected) as shown in step (3) in Embodiment 2 to divide an N×2N pixel block. As Fig.25 shown, the selected block mode can divide the current N×2N pixel block (the right block in this example) as Fig.25 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Fig.25 shown in (a), (b), (d), and (e) into 2 sub-blocks of different sizes (asymmetric binary tree). For example, when using an asymmetric binary tree, the large sub-block can have a width / height that is 3 times that of the small sub-block.

[0274] Fig.26 This shows an example of using the selected block mode (when the second block mode is not selected) as shown in step (3) in Embodiment 2 to divide an N×N pixel block. As Fig.26 shown, in step (1), the 2N×N pixel block is vertically divided into 2 N×N pixel sub-blocks. In step (2), the left N×N pixel block is vertically divided into 2 N / 2×N pixel sub-blocks. In step (3), the selected block mode for the current N×N pixel block (the left block in this example) can be used to divide the current block as Fig.26 shown in (c) and (f) into 3 sub-blocks. The sizes of the 3 sub-blocks can be different. For example, among the 3 sub-blocks, the large sub-block can have a width / height that is 2 times that of the small sub-block. And for example, the selected block mode can also divide the current block as Fig.26is divided into two sub-blocks (asymmetric binary tree) with different sizes as shown in (a), (b), (d), and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0275] Fig. 27 Shows an example of dividing an N×N pixel block using the selected block partitioning mode when not selecting the second block partitioning mode as shown in step (3) in Embodiment 2. As Fig. 27 shown, in step (1), the N×2N pixel block is horizontally divided into two N×N pixel sub-blocks, and in step (2), the upper N×N pixel block is horizontally divided into two N×N / 2 pixel sub-blocks. In step (3), the selected block partitioning mode for the current N×N pixel block (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in Fig. 27 (c) and (f). The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block can have twice the width / height of the small sub-block. And for example, the selected block partitioning mode can also divide the current block into two sub-blocks with different sizes (asymmetric binary tree) as shown in Fig. 27 (a), (b), (d), and (e). For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0276] Fig.17 Represents the possible positions of the first parameter in the compressed video stream. As Fig.17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0277] Fig.18 Represents the possible positions of the second parameter in the compressed video stream. As Fig.18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Fig.19 shown, the second parameter follows the first parameter in the bitstream.

[0278] The first block and the second block are different blocks. The first block and the second block may be included in the same frame. For example, the first block may be a block adjacent above the second block. And for example, the first block may also be a block adjacent to the left of the second block.

[0279] In step S1007, the second block is divided into sub-blocks using the selected block division mode. In step S1008, the divided blocks are encoded.

[0280] [Encoding device]

[0281] Fig.15 is a block diagram showing the structure of the video / image encoding device according to Embodiment 2 or 3.

[0282] The video encoding device 5000 is a device for encoding an input video / image for each block to generate an encoded output bitstream. As Fig.15 shown, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra prediction unit 5007, an inter prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.

[0283] The input video is input to the adder, and the added value is output to the transform unit 5001. The transform unit 5001 transforms the added value into frequency coefficients based on the block division mode derived by the block division determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block division mode can be associated with a block division mode, a block division type, or a block division direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantized values to the inverse quantization unit 5003 and the entropy encoding unit 5009.

[0284] The inverse quantization unit 5003 inverse-quantizes the quantized values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block division mode derived by the block division determination unit 5010, transforms the frequency coefficients into sample values of the bitstream, and outputs the sample values to the adder.

[0285] The adder adds the sample values of the bitstream output from the inverse transform unit 5004 to the predicted video / image values output from the intra / inter prediction units 5007 and 5008, and outputs the added value to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives a block partitioning pattern and parameters related to the block partitioning pattern. If the derived block partitioning pattern is used, the block is divided into a plurality of sub-blocks. The intra / inter prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block partitioning pattern derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.

[0286] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.

[0287] [Decoding process]

[0288] Fig.12 Indicates the video decoding process related to Embodiment 2.

[0289] First, in step S2001, the first parameter is interpreted from the bitstream. The first parameter identifies, from among a plurality of partitioning patterns, the partitioning pattern for dividing the first block into sub-blocks. If the partitioning pattern is used, the block is divided into sub-blocks, and if a different partitioning pattern is used, the block is divided into sub-blocks having different shapes, different heights, or different widths.

[0290] Fig.28 Shows an example of the partitioning pattern for dividing an N×N pixel block in Embodiment 2. In Fig.28 ,(a) to (h) represent different partitioning patterns. As Fig.28As shown, if the block mode (a) is used, a block of N×N pixels (for example, 16×16 pixels, and as the value of "N", any value that is an integer multiple of 4 from 8 to 128 can be taken) is divided into two sub-blocks of N / 2×N pixels (for example, 8×16 pixels). If the block mode (b) is used, the block of N×N pixels is divided into a sub-block of N / 4×N pixels (for example, 4×16 pixels) and a sub-block of 3N / 4×N pixels (for example, 12×16 pixels). If the block mode (c) is used, the block of N×N pixels is divided into a sub-block of 3N / 4×N pixels (for example, 12×16 pixels) and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (d) is used, the block of N×N pixels is divided into a sub-block of (N / 4)×N pixels (for example, 4×16 pixels), a sub-block of N / 2×N pixels (for example, 8×16 pixels), and a sub-block of N / 4×N pixels (for example, 4×16 pixels). If the block mode (e) is used, the block of N×N pixels is divided into two sub-blocks of N×N / 2 pixels (for example, 16×8 pixels). If the block mode (f) is used, the block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels) and a sub-block of N×3N / 4 pixels (for example, 16×12 pixels). If the block mode (g) is used, the block of N×N pixels is divided into a sub-block of N×3N / 4 pixels (for example, 16×12 pixels) and a sub-block of N×N / 4 pixels (for example, 16×4 pixels). If the block mode (h) is used, the block of N×N pixels is divided into a sub-block of N×N / 4 pixels (for example, 16×4 pixels), a sub-block of N×N / 2 pixels (for example, 16×8 pixels), and a sub-block of N×N / 4 pixels (for example, 16×4 pixels).

[0291] Next, in step S2002, it is determined whether the first parameter identifies the first block mode.

[0292] Next, in step S2003, based at least on the determination of whether the first parameter identifies the first block mode, it is determined whether the second block mode is not selected as a candidate for dividing the second block.

[0293] Two different sets of block modes may divide a block into sub-blocks of the same shape and size. For example, as Fig.31A shown, the sub-blocks of (1b) and (2c) have the same shape and size. One set of block modes can include at least two block modes. For example, as Fig.31A shown in (1a) and (1b) of, one set of block modes can then include a binary tree vertical division of the central sub-block and a non-division of the other sub-blocks according to the vertical division of the ternary tree. And for example, as Fig.31AAs shown in (2a), (2b), and (2c), other block pattern sets can then include a binary tree vertical split that divides into two sub-blocks by vertically splitting a binary tree. Each block pattern set results in sub-blocks of the same shape and size.

[0294] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that are different binary numbers or have different numbers of bits when encoded in a bitstream, select the block pattern set with fewer binary numbers or fewer bits.

[0295] When selecting between two block pattern sets that divide a block into sub-blocks of the same shape and size and that have the same number of bits or the same number of bits when encoded in a bitstream, select the block pattern set that appears first in a specified order of multiple block pattern sets. The specified order can be, for example, an order based on the number of block patterns within each block pattern set.

[0296] Fig.31A and Fig.31B is a diagram showing an example of dividing a block into sub-blocks using a block pattern set with fewer binary numbers in the encoding of the block pattern. In this example, when the left N×N pixel block is vertically divided into 2 sub-blocks, the second block pattern for the right N×N pixel block is not selected in step (2c). This is because, in Fig.31B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) requires more binary numbers for encoding the block pattern compared to the first block pattern set (1a, 1b).

[0297] Fig.32A is a diagram showing an example of dividing a block into sub-blocks using the block pattern set that appears first in a specified order of multiple block pattern sets. In this example, when the 2N×N / 2 pixel block is vertically divided into 3 sub-blocks, the second block pattern for the lower 2N×N / 2 pixel block is not selected in step (2c). This is because, in Fig.32B the encoding method of the block pattern, the second block pattern set (2a, 2b, 2c) has the same binary number as the first block pattern set (1a, 1b, 1c, 1d), and in Fig.32C the specified order of the block pattern sets shown, it appears after the first block pattern set (1a, 1b, 1c, 1d). The specified order of multiple block pattern sets can also be fixed and signaled within the bitstream.

[0298] Fig. 20 shows an example of not selecting the second block pattern for the division of the 2N×N pixel block as shown in step (2c) in Embodiment 2. As Fig. 20As shown, it is possible to use the first segmentation method (i) to equally divide a block of 2N×2N pixels (e.g., 16×16 pixels) into 4 sub-blocks of N×N pixels (e.g., 8×8 pixels) as in step (1a). Also, it is possible to use the second segmentation method (ii) to horizontally equally divide a block of 2N×2N pixels into 2 sub-blocks of 2N×N pixels (e.g., 16×8 pixels) as in step (2a). In the second segmentation method (ii), when the upper 2N×N pixel block (the first block) is vertically divided into 2 sub-blocks of N×N pixels by the first block division pattern as in step (2b), the second block division pattern for vertically dividing the lower 2N×N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division pattern in step (2c). This is because sub-blocks of the same size as those obtained by the four-way division using the first segmentation method (i) are generated.

[0299] As described above, Fig. 20 in a case where if the first block division pattern is used, the first block is equally divided into 2 sub-blocks in the vertical direction, and if the second block division pattern is used, the second block adjacent to the first block in the vertical direction is equally divided into 2 sub-blocks in the vertical direction, the second block division pattern is not selected as a candidate.

[0300] Fig.21 This shows an example in Embodiment 2 where the second block division pattern is not selected for the division of an N×2N pixel block as in step (2c). As Fig.21 shown, it is possible to use the first segmentation method (i) to equally divide a block of 2N×2N pixels into 4 sub-blocks of N×N pixels as in step (1a). Also, it is possible to use the second segmentation method (ii) to vertically equally divide a block of 2N×2N pixels into 2 sub-blocks of 2N×N pixels (e.g., 8×16 pixels) as in step (2a). In the second segmentation method (ii), when the left N×2N pixel block (the first block) is horizontally divided into 2 sub-blocks of N×N pixels by the first block division pattern as in step (2b), the second block division pattern for horizontally dividing the right N×2N pixel block (the second block) into 2 sub-blocks of N×N pixels is not selected as a candidate for the possible block division pattern in step (2c). This is because sub-blocks of the same size as those obtained by the four-way division using the first segmentation method (i) are generated.

[0301] As described above, Fig.21 in a case where if the first block division pattern is used, the first block is equally divided into 2 sub-blocks in the horizontal direction, and if the second block division pattern is used, the second block adjacent to the first block in the horizontal direction is equally divided into 2 sub-blocks in the horizontal direction, the second block division pattern is not selected as a candidate.

[0302] Fig. 22 This shows an example in Embodiment 2 where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c). As Fig. 22 shown, the first division method (i) can be used to vertically divide a 2N×N pixel block (e.g., 16×8 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N / 2×N pixels, N×N pixels, and N / 2×N pixels (e.g., sub-blocks of 4×8 pixels, 8×8 pixels, and 4×8 pixels) as in step (1a). Also, the second division method (ii) can be used to divide a 2N×N pixel block into two N×N pixel sub-blocks as in step (2a). In the first division method (i), the central N×N pixel block can be vertically divided into two N / 2×N pixel (e.g., 4×8 pixel) sub-blocks in step (1b). In the second division method (ii), when the left N×N pixel block (the first block) is vertically divided into two N / 2×N pixel sub-blocks as in step (2b), the block division mode of vertically dividing the right N×N pixel block (the second block) into two N / 2×N pixel sub-blocks in step (2c) is not selected as a candidate for the possible block division modes. This is because sub-block sizes identical to those obtained by the first division method (i) will be generated, i.e., four N / 2×N pixel sub-blocks.

[0303] As described above, Fig. 22 in a case where if the first block division mode is used, the first block is equally divided into two sub-blocks in the vertical direction, and if the second block division mode is used, the second block adjacent to the first block in the horizontal direction is equally divided into two sub-blocks in the vertical direction, the second block division mode is not selected as a candidate.

[0304] Fig.23 This shows an example in Embodiment 2 where the second block division mode is not selected for the division of an N×N pixel block as shown in step (2c). As Fig.23As shown, it is possible to use the first segmentation method (i) to divide N×2N pixels (for example, 8×16 pixels, and any value that is an integer multiple of 4 from 8 to 128 can be taken as the value of "N") into sub-blocks of N×N / 2 pixels, sub-blocks of N×N pixels, and sub-blocks of N×N / 2 pixels (for example, sub-blocks of 8×4 pixels, sub-blocks of 8×8 pixels, and sub-blocks of 8×4 pixels) as in step (1a). Also, it is possible to use the second segmentation method to divide it into two sub-blocks of N×N pixels as in step (2a). In the first segmentation method (i), it is possible to divide the central block of N×N pixels into two sub-blocks of N×N / 2 pixels as in step (1b). In the second segmentation method (ii), when the upper block of N×N pixels (the first block) is horizontally divided into two sub-blocks of N×N / 2 pixels as in step (2b), in step (2c), the block division mode of horizontally dividing the lower block of N×N pixels (the second block) into two sub-blocks of N×N / 2 pixels is not selected as a candidate for possible block division modes. This is because sub-blocks of the same size as those obtained by the first segmentation method (i) will be generated, that is, four sub-blocks of N×N / 2 pixels.

[0305] As described above, Fig.23 in the case where if the first block division mode is used, the first block is equally divided into two sub-blocks in the horizontal direction, and if the second block division mode is used, the second block adjacent to the first block in the vertical direction is equally divided into two sub-blocks in the horizontal direction, the second block division mode is not selected as a candidate.

[0306] If it is determined that the second block division mode is selected as a candidate for dividing the second block (No in S2003), then in step S2004, the second parameter is decoded from the bitstream, and a block division mode is selected from among multiple block division modes including the second block division mode as a candidate.

[0307] If it is determined that the second block division mode is not selected as a candidate for dividing the second block (Yes in S2003), then in step S2005, a block division mode different from the second block division mode is selected to divide the second block. The block division mode selected here divides the block into sub-blocks with different shapes or different sizes compared to the sub-blocks generated by the second block division mode.

[0308] Fig.24 An example is shown of using the block division mode selected when the second block division mode is not selected to divide the 2N×N pixel block as in step (3) in Embodiment 2. As Fig.24 shown, the selected block division mode can divide the current block of 2N×N pixels (the lower block in this example) as Fig.24is divided into three sub - blocks as shown in (c) and (f). The sizes of the three sub - blocks can be different. For example, among the three sub - blocks, the large sub - block can have a width / height twice that of the small sub - block. And for example, the selected block - division pattern can also divide the current block as shown in Fig.24 into two sub - blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e). For example, in the case of using an asymmetric binary tree, the large sub - block can have a width / height three times that of the small sub - block.

[0309] Fig.25 represents an example of dividing a block of N×2N pixels using the selected block - division pattern without selecting the second block - division pattern as shown in step (3) in Embodiment 2. As shown in Fig.25 , the selected block - division pattern can divide the current block of N×2N pixels (the right block in this example) as shown in Fig.25 into three sub - blocks as shown in (c) and (f). The sizes of the three sub - blocks can be different. For example, among the three sub - blocks, the large sub - block can have a width / height twice that of the small sub - block. And for example, the selected block - division pattern can also divide the current block as shown in Fig.25 into two sub - blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e). For example, in the case of using an asymmetric binary tree, the large sub - block can have a width / height three times that of the small sub - block.

[0310] Fig.26 represents an example of dividing a block of N×N pixels using the selected block - division pattern without selecting the second block - division pattern as shown in step (3) in Embodiment 2. As shown in Fig.26 , in step (1), the block of 2N×N pixels is vertically divided into two sub - blocks of N×N pixels, and in step (2), the left block of N×N pixels is vertically divided into two sub - blocks of N / 2×N pixels. In step (3), the selected block - division pattern for the current block of N×N pixels (the left block in this example) can be used to divide the current block into three sub - blocks as shown in Fig.26 (c) and (f). The sizes of the three sub - blocks can be different. For example, among the three sub - blocks, the large sub - block can have a width / height twice that of the small sub - block. And for example, the selected block - division pattern can also divide the current block as shown in Fig.26 into two sub - blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d) and (e). For example, in the case of using an asymmetric binary tree, the large sub - block can have a width / height three times that of the small sub - block.

[0311] Fig. 27 represents an example of dividing a block of N×N pixels using the selected block - division pattern without selecting the second block - division pattern as shown in step (3) in Embodiment 2. As shown in Fig. 27As shown, in step (1), a block of N×2N pixels is horizontally divided into two sub-blocks of N×N pixels. In step (2), the upper block of N×N pixels is horizontally divided into two sub-blocks of N×N / 2 pixels. In step (3), the selected partitioning mode for the current block of N×N pixels (the lower block in this example) can be used to divide the current block into three sub-blocks as shown in (c) and (f) of Fig. 27 . The sizes of the three sub-blocks can be different. For example, among the three sub-blocks, the large sub-block has twice the width / height of the small sub-block. And for example, the selected partitioning mode can also divide the current block into two sub-blocks of different sizes (asymmetric binary tree) as shown in (a), (b), (d), and (e) of Fig. 27 . For example, in the case of using an asymmetric binary tree, the large sub-block can have three times the width / height of the small sub-block.

[0312] Fig.17 Indicates the possible positions of the first parameter within the compressed video stream. As shown in Fig.17 , the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0313] Fig.18 Indicates the possible positions of the second parameter within the compressed video stream. As shown in Fig.18 , the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As shown in Fig.19 , the second parameter is configured in the bitstream following the first parameter.

[0314] The first block and the second block are different blocks. The first block and the second block can also be included in the same frame. For example, the first block can be a block adjacent above the second block. And for example, the first block can also be a block adjacent to the left of the second block.

[0315] In step S2006, the second block is divided into sub-blocks using the selected partitioning mode. In step S2007, the divided block is decoded.

[0316] [Decoding device]

[0317] Fig.16is a block diagram showing the configuration of the video / image decoding apparatus according to Embodiment 2 or 3.

[0318] The video decoding apparatus 6000 is an apparatus that decodes an input encoded bitstream for each block and outputs a video / image. The video decoding apparatus 6000 is as Fig.16 shown and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block segmentation determination unit 6008.

[0319] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.

[0320] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. The inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients based on the block partitioning pattern derived by the block segmentation determination unit 6008, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning pattern can be associated with a block partitioning pattern, a block partitioning type, or a block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006, 6007, outputs the added values to the display, and outputs the added values to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005, and uses the parameters decoded by the entropy decoding unit 6001 to derive the block partitioning pattern. If the derived block partitioning pattern is used, the block is divided into a plurality of sub-blocks. Further, the intra / inter prediction units 6006, 6007 perform prediction on the video / image region of the block to be decoded based on the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 reconstructed according to the block partitioning pattern derived by the block segmentation determination unit 6008.

[0321] (Embodiment 3)

[0322] Refer to Fig.13 and Fig.14 to specifically describe the encoding process and decoding process according to Embodiment 3. Refer to Fig.15 and Fig.16 to specifically describe the encoding apparatus and decoding apparatus according to Embodiment 3.

[0323] [Encoding Process]

[0324] Fig.13Indicates the video encoding process related to Embodiment 3.

[0325] First, in step S3001, a first parameter is written to the bitstream, and this first parameter identifies, from among a plurality of block types, the block type used to divide the first block into sub-blocks.

[0326] In the next step S3002, a second parameter indicating the block division direction is written to the bitstream. The second parameter is configured in the bitstream following the first parameter. The block type and the block division direction can form a block division pattern together. The divided block indicates the number and division ratio of the sub-blocks used to divide the block.

[0327] Fig.29 Shows an example of the block type and block division direction used to divide an N×N pixel block in Embodiment 3. In Fig.29 , (1), (2), (3), and (4) are different block types, (1a), (2a), (3a), and (4a) are block division patterns with different block types in the vertical division direction, and (1b), (2b), (3b), and (4b) are block division patterns with different block types in the horizontal division direction. As Fig.29 shown, when the division ratio is 1:1 and the N×N pixel block is divided in a symmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (1a). When the division ratio is 1:1 and the N×N pixel block is divided in a symmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (1b). When the division ratio is 1:3 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (2a). When the division ratio is 1:3 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (2b). When the division ratio is 3:1 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (3a). When the division ratio is 3:1 and the N×N pixel block is divided in an asymmetric binary tree (i.e., 2 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (3b). When the division ratio is 1:2:1 and the N×N pixel block is divided in a ternary tree (i.e., 3 sub-blocks) along the vertical direction, the N×N pixel block is divided using the block division pattern (4a). When the division ratio is 1:2:1 and the N×N pixel block is divided in a ternary tree (i.e., 3 sub-blocks) along the horizontal direction, the N×N pixel block is divided using the block division pattern (4b).

[0328] Fig.17 Shows the positions where the first parameter in the compressed video stream can be considered. As Fig.17As shown, the first parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The first parameter can represent a method of dividing a block into multiple sub-blocks. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks.

[0329] Fig.18 Indicates the possible positions of the second parameter within the compressed video stream. As Fig.18 shown, the second parameter can be configured within a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The second parameter can represent a method of dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. The second parameter can also include a parameter indicating whether to divide the block into more than two sub-blocks. As Fig.19 shown, the second parameter is configured in the bitstream following the first parameter.

[0330] Fig.30 Indicates the advantages of coding the block type before the block direction compared to coding the block direction before the block type. In this example, when the horizontal block direction is invalidated due to an unsupported size (16×2 pixels), there is no need to code the block direction. In this example, the block direction is determined to be the vertical block direction, and the horizontal block direction is invalid. When coding the block type before the block direction, compared to coding the block direction before the block type, the code bits brought by the coding of the block direction are suppressed.

[0331] In this way, it is also possible to determine whether a block can be divided horizontally and vertically based on pre-determined conditions for block divisibility or non-divisibility. Then, when it is determined that the block can be divided only in one of the horizontal and vertical directions, the writing of the block direction to the bitstream can also be skipped. Furthermore, when it is determined that the block is not divisible in both the horizontal and vertical directions, in addition to skipping the writing of the block direction to the bitstream, the writing of the block type to the bitstream can also be skipped.

[0332] The pre-determined conditions for block divisibility or non-divisibility are defined, for example, by the size (number of pixels) or the number of divisions. These conditions for block divisibility or non-divisibility can also be pre-defined in the standard specification. Also, the conditions for block divisibility or non-divisibility can be included in a video parameter set, a sequence parameter set, a picture parameter set, a slice header, or a coding tree unit. The conditions for block divisibility or non-divisibility can be fixed for all blocks, or can be dynamically switched according to the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.

[0333] In step S3003, the block is divided into sub - blocks using the identified block type and the indicated block direction. In step S3004, the divided block is encoded.

[0334] [Encoding device]

[0335] Fig.15 is a block diagram showing the structure of the video / image encoding device according to Embodiment 2 or 3.

[0336] The video encoding device 5000 is a device for encoding an input video / image for each block and generating an encoded output bitstream. As Fig.15 shown, the video encoding device 5000 includes a transform unit 5001, a quantization unit 5002, an inverse quantization unit 5003, an inverse transform unit 5004, a block memory 5005, a frame memory 5006, an intra - prediction unit 5007, an inter - prediction unit 5008, an entropy encoding unit 5009, and a block division determination unit 5010.

[0337] The input video is input to an adder, and the added value is output to the transform unit 5001. The transform unit 5001 transforms the added value into frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, and outputs the frequency coefficients to the quantization unit 5002. The block division type and direction can be associated with a block division mode, a block division type, or a block division direction. The quantization unit 5002 quantizes the input quantization coefficients and outputs the quantized values to the inverse quantization unit 5003 and the entropy encoding unit 5009.

[0338] The inverse quantization unit 5003 inverse - quantizes the quantized values output from the quantization unit 5002 and outputs the frequency coefficients to the inverse transform unit 5004. The inverse transform unit 5004 performs an inverse frequency transform on the frequency coefficients based on the block division type and direction derived by the block division determination unit 5010, transforms the frequency coefficients into sample values of the bitstream, and outputs the sample values to the adder.

[0339] The adder adds the sample values of the bitstream output from the intra-frame / inter-frame prediction units 5007 and 5008 to the predicted video / image values, and outputs the added value to the block memory 5005 or the frame memory 5006 for further prediction. The block segmentation determination unit 5010 collects block information from the block memory 5005 or the frame memory 5006, and derives the block segmentation type and direction, as well as the parameters related to the block segmentation type and direction. If the derived block segmentation type and direction are used, the block is divided into a plurality of sub-blocks. The intra-frame / inter-frame prediction units 5007 and 5008 search among the video / images stored in the block memory 5005 or the video / images in the frame memory 5006 reconstructed by the block segmentation type and direction derived by the block segmentation determination unit 5010, and estimate, for example, the video / image region most similar to the input video / image to be predicted.

[0340] The entropy encoding unit 5009 encodes the quantization values output from the quantization unit 5002, and encodes the parameters from the block segmentation determination unit 5010, and outputs a bitstream.

[0341] [Decoding process]

[0342] Fig.14 Represents the video decoding process related to Embodiment 3.

[0343] First, in step S4001, the first parameter is read from the bitstream. The first parameter identifies the segmentation type for dividing the first block into sub-blocks from among a plurality of segmentation types.

[0344] In the next step S4002, the second parameter indicating the segmentation direction is read from the bitstream. The second parameter follows the first parameter in the bitstream. The segmentation type may also form a segmentation pattern together with the segmentation direction. The segmentation type indicates the number and segmentation ratio of the sub-blocks for dividing the block.

[0345] Fig.29 Represents an example of the segmentation type and segmentation direction for dividing an N×N pixel block in Embodiment 3. In Fig.29 ,(1), (2), (3) and (4) are different segmentation types, (1a), (2a), (3a) and (4a) are segmentation patterns with different segmentation types in the vertical direction, and (1b), (2b), (3b) and (4b) are segmentation patterns with different segmentation types in the horizontal direction. As Fig.29As shown, when the block ratio is 1:1 and the N×N pixel block is divided along the vertical direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (1a). When the block ratio is 1:1 and the N×N pixel block is divided along the horizontal direction in a symmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (1b). When the block ratio is 1:3 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (2a). When the block ratio is 1:3 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (2b). When the block ratio is 3:1 and the N×N pixel block is divided along the vertical direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (3a). When the block ratio is 3:1 and the N×N pixel block is divided along the horizontal direction in an asymmetric binary tree (i.e., 2 sub-blocks), the N×N pixel block is divided using the partitioning mode (3b). When the block ratio is 1:2:1 and the N×N pixel block is divided along the vertical direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the partitioning mode (4a). When the block ratio is 1:2:1 and the N×N pixel block is divided along the horizontal direction in a ternary tree (i.e., 3 sub-blocks), the N×N pixel block is divided using the partitioning mode (4b).

[0346] Fig.17 Indicates the positions where the first parameter in the compressed video stream can be considered. As Fig.17 shown, the first parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The first parameter can represent a method for dividing a block into multiple sub-blocks. For example, the first parameter can include an identifier for the above-mentioned partitioning types. For example, the first parameter can include a flag indicating whether to divide the block horizontally or vertically. The first parameter can also include a parameter indicating whether to divide the block into more than 2 sub-blocks.

[0347] Fig.18 Indicates the positions where the second parameter in the compressed video stream can be considered. As Fig.18 shown, the second parameter can be configured in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The second parameter can represent a method for dividing a block into multiple sub-blocks. For example, the second parameter can include a flag indicating whether to divide the block horizontally or vertically. That is, the second parameter can include a parameter indicating the partitioning direction. The second parameter can also include a parameter indicating whether to divide the block into more than 2 sub-blocks. As Fig.19 shown, the second parameter is configured in the bitstream following the first parameter.

[0348] Fig.30Represents the advantages of encoding the block type before the block direction compared to the case of encoding the block direction before the block type. In this example, when the block direction in the horizontal direction is invalidated due to an unsupported size (16×2 pixels), there is no need to encode the block direction. In this example, the block direction is determined to be the block direction in the vertical direction, and the block direction in the horizontal direction is invalidated. Encoding the block type before the block direction suppresses the code bits caused by the encoding of the block direction compared to the case of encoding the block direction before the block type.

[0349] In this way, it is also possible to determine whether a block can be divided in the horizontal and vertical directions respectively based on pre-determined conditions for block divisibility or indivisibility. Then, when it is determined that the block can be divided only in one of the horizontal and vertical directions, the reading of the block direction from the bitstream can also be skipped. Furthermore, when it is determined that the block cannot be divided in both the horizontal and vertical directions, in addition to the reading of the block direction, the reading of the block type from the bitstream can also be skipped.

[0350] The pre-determined conditions for block divisibility or indivisibility are defined by, for example, the size (number of pixels) or the number of divisions. These conditions for block divisibility or indivisibility can also be pre-defined in the standard specifications. Also, the conditions for block divisibility or indivisibility can be included in the video parameter set, sequence parameter set, picture parameter set, slice header, or coding tree unit. The conditions for block divisibility or indivisibility can be fixed for all blocks, or can be dynamically switched according to the characteristics of the block (e.g., luminance and chrominance blocks) or the characteristics of the picture (e.g., I, P, B pictures), etc.

[0351] In step S4003, the block is divided into sub-blocks using the identified block type and the indicated division direction. In step S4004, the divided block is decoded.

[0352] [Decoding device]

[0353] Fig.16 Is a block diagram showing the structure of the video / image decoding device according to Embodiment 2 or 3.

[0354] The video decoding device 6000 is a device for decoding an input encoded bitstream for each block and outputting a video / image. The video decoding device 6000 is as Fig.16 shown and includes an entropy decoding unit 6001, an inverse quantization unit 6002, an inverse transform unit 6003, a block memory 6004, a frame memory 6005, an intra prediction unit 6006, an inter prediction unit 6007, and a block division determination unit 6008.

[0355] The input encoded bitstream is input to the entropy decoding unit 6001. After the input encoded bitstream is input to the entropy decoding unit 6001, the entropy decoding unit 6001 decodes the input encoded bitstream, outputs the parameters to the block segmentation determination unit 6008, and outputs the decoded values to the inverse quantization unit 6002.

[0356] The inverse quantization unit 6002 performs inverse quantization on the decoded values and outputs the frequency coefficients to the inverse transform unit 6003. Based on the block partitioning type and direction derived by the block segmentation determination unit 6008, the inverse transform unit 6003 performs an inverse frequency transform on the frequency coefficients, transforms the frequency coefficients into sample values, and outputs the sample values to the adder. The block partitioning type and direction can be associated with the block partitioning mode, block partitioning type, or block partitioning direction. The adder adds the sample values to the predicted video / image values output from the intra / inter prediction units 6006 and 6007, outputs the added values to the display, and outputs the added values to the block memory 6004 or the frame memory 6005 for further prediction. The block segmentation determination unit 6008 collects block information from the block memory 6004 or the frame memory 6005 and uses the parameters decoded by the entropy decoding unit 6001 to derive the block partitioning type and direction. If the derived block partitioning type and direction are used, the block is divided into multiple sub-blocks. Furthermore, the intra / inter prediction units 6006 and 6007 perform prediction on the video / image region of the block to be decoded according to the video / image stored in the block memory 6004 or the video / image in the frame memory 6005 obtained by reconstruction based on the block partitioning type and direction derived by the block segmentation determination unit 6008.

[0357] (Embodiment 4)

[0358] In the above embodiments, each functional block can generally be implemented by an MPU, a memory, etc. In addition, the processing of each functional block is generally realized by a program execution unit such as a processor reading and executing software (program) recorded in a recording medium such as a ROM. This software can be distributed by downloading or the like, or can be recorded in a recording medium such as a semiconductor memory for distribution. In addition, of course, each functional block can also be implemented by hardware (special-purpose circuit).

[0359] In addition, the processing described in each embodiment can be realized by centralized processing using a single device (system), or can also be realized by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing or distributed processing can be performed.

[0360] The form of the present invention is not limited to the above embodiments, and various changes can be made, and they are also included in the scope of the form of the present invention.

[0361] Next, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in the above-described embodiments and a system using the same will be described. The system is characterized by including an image encoding device that uses the image encoding method, an image decoding device that uses the image decoding method, and an image encoding / decoding device that includes both. Regarding other configurations in the system, they can be appropriately changed according to circumstances.

[0362] [Usage Example]

[0363] Fig.33 FIG. shows the overall configuration of a content supply system ex100 that implements a content distribution service. The provision area of the communication service is divided into desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110 that are fixed wireless stations are respectively provided in each unit.

[0364] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may also connect by combining some of the above elements. The various devices may also be directly or indirectly connected to each other via a telephone network or short-range wireless without going through the base stations ex106 to ex110 that are fixed wireless stations. In addition, a streaming media server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 via the Internet ex101 or the like. In addition, the streaming media server ex103 is connected to terminals in a hotspot in an airplane ex117 via a satellite ex116.

[0365] In addition, a wireless access point or a hotspot or the like may be used instead of the base stations ex106 to ex110. In addition, the streaming media server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0366] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a mobile phone, or a PHS (Personal Handyphone System) or the like corresponding to a mobile communication system mode generally referred to as 2G, 3G, 3.9G, 4G, and 5G in the future.

[0367] The home appliance ex118 is a refrigerator or a device included in a household fuel cell cogeneration system, etc.

[0368] In the content supply system ex100, a terminal having a photographing function is connected to the streaming media server ex103 via a base station ex106 or the like, whereby live distribution or the like can be performed. In live distribution, terminals (such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smart phone ex115, and the terminal in the airplane ex117, etc.) perform the encoding process described in the above embodiments on the still image or moving image content photographed by the user using the terminal, multiplex the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and send the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.

[0369] On the other hand, the streaming media server ex103 performs stream distribution on the content data sent by the requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or a terminal in an airplane ex117 that can decode the data after the above encoding process. Each device that receives the distributed data performs decoding processing on the received data and reproduces it. That is, each device functions as an image decoding device according to one aspect of the present invention.

[0370] [Decentralized processing]

[0371] In addition, the streaming media server ex103 may be a plurality of servers or a plurality of computers, and perform decentralized processing or recording and distribution of data. For example, the streaming media server ex103 may be implemented by a CDN (Contents Delivery Network), and content distribution is achieved through a network connecting many edge servers scattered around the world to each other. In the CDN, a physically closer edge server is dynamically allocated according to the client. And by caching and distributing the content to this edge server, the delay can be reduced. In addition, in the case of a certain error or a change in the communication state due to an increase in traffic or the like, the processing can be decentralized using multiple edge servers, or the distribution entity can be switched to another edge server, or a part of the network that has failed can be bypassed and the distribution can continue, so high-speed and stable distribution can be achieved.

[0372] In addition, not limited to distributed processing of its own, the encoding process of the captured data can be performed by each terminal, on the server side, or can be shared between them. As an example, usually two processing loops are performed in the encoding process. In the first loop, the complexity or code amount of the image in units of frames or scenes is detected. In addition, in the second loop, a process is performed to improve the encoding efficiency while maintaining the image quality. For example, by performing the first encoding process by the terminal and the second encoding process by the server that receives the content, it is possible to reduce the processing load in each terminal while improving the quality and efficiency of the content. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.

[0373] As other examples, the camera ex113 etc. extracts feature amounts from the image, compresses the data regarding the feature amounts as metadata, and sends it to the server. The server, for example, determines the importance of the target based on the feature amounts and switches the quantization accuracy etc., and performs compression corresponding to the meaning of the image. Feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction during re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed by the server.

[0374] As other examples, in a stadium, shopping mall, factory, etc., there are cases where there are multiple video data obtained by multiple terminals capturing substantially the same scene. In this case, the multiple terminals that performed the shooting are used, and other terminals and servers that did not perform the shooting are used as needed. For example, distributed processing is performed by respectively allocating the encoding process in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture, etc. As a result, it is possible to reduce the delay and better achieve real-time performance.

[0375] In addition, since the multiple video data are of substantially the same scene, the server can also manage and / or instruct to refer to the video data captured by each terminal with each other. Or, it can also be that the server receives the encoded data from each terminal and changes the reference relationship between the multiple data, or corrects or replaces the picture itself and re-encodes it. As a result, it is possible to generate a stream with improved quality and efficiency of each data.

[0376] In addition, the server can also perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server can change the MPEG-like encoding method to the VP-like, or can change H.264 to H.265.

[0377] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions use terms such as "server" or "terminal" as the processing entity, but part or all of the processing performed by the server can also be performed by the terminal, and part or all of the processing performed by the terminal can also be performed by the server. In addition, regarding these, the same applies to the decoding process.

[0378] [3D, Multi-angle]

[0379] In recent years, the use of merging different scenes captured by multiple cameras ex113 and / or terminals such as smartphones ex115 that are roughly synchronized with each other, or images or videos of the same scene captured from different angles has increased. The videos captured by each terminal are merged based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the videos are consistent.

[0380] The server not only encodes two-dimensional moving images, but can also encode still images automatically or at a user-specified time based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative position relationship between the shooting terminals, it can not only generate the three-dimensional shape of the scene based on the videos of the same scene captured from different angles in addition to two-dimensional moving images. In addition, the server can also encode the three-dimensional data generated by point cloud etc. separately, or select or reconstruct from the videos captured by multiple terminals based on the results of identifying or tracking people or objects using the three-dimensional data to generate and send the video to the receiving terminal.

[0381] In this way, the user can not only arbitrarily select each video corresponding to each shooting terminal to view the scene, but also view the content of the video cut from any viewpoint from the three-dimensional data reconstructed using multiple images or videos. Furthermore, similar to the video, the sound can also be collected from multiple different angles, and the server multiplexes and sends the sound from a specific angle or space with the video in accordance with the video.

[0382] In addition, in recent years, content that establishes a correspondence between the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) has also been popularized. In the case of VR images, the server separately creates viewpoint images for the right eye and the left eye, and can perform encoding that allows reference between the viewpoint videos through Multi-View Coding (MVC) etc., or can encode them as different streams without mutual reference. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.

[0383] In the case of an AR image, it can also be that the server overlaps the virtual object information in the virtual space with the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device acquires or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates the overlapping data by smoothly connecting them. Alternatively, it can also be that the decoding device sends the movement of the user's viewpoint to the server in addition to the delegation of the virtual object information, and the server creates the overlapping data according to the three-dimensional data held in the server, matches the received movement of the viewpoint, encodes the overlapping data, and distributes it to the decoding device. In addition, the overlapping data has an α value representing the transmittance in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted in this part. Alternatively, the server can also set the RGB value of a specified value as the background like chroma key, and generate data with the part other than the target set as the background color.

[0384] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, on the server side, or they can be shared with each other. As an example, it can also be that a certain terminal first sends a reception request to the server, and another terminal receives the content corresponding to the request and performs the decoding process, and sends the decoded signal to the device with a display. By dispersing the processing regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. In addition, as another example, a large-sized image data can be received by a TV or the like, and a personal terminal of the viewer decodes and displays a part of the area such as tiles after the picture is segmented. Thus, while making the overall image shared, it is possible to confirm one's own responsible area or the area that one wants to confirm in more detail at hand.

[0385] In addition, it is envisioned that in the future, in a situation where multiple wireless communications at short, medium, or long distances can be used both indoors and outdoors, using a distribution system standard such as MPEG-DASH, seamless reception of content while appropriately switching data for the connected communication. Thus, the user can not only freely select a decoding device or a display device such as a monitor installed indoors and outdoors with their own terminal, but also switch in real time. In addition, based on their own position information, etc., it is possible to switch the decoding terminal and the display terminal for decoding. Thus, it is also possible to display map information on a part of the wall or floor of a building next to a device that can be displayed while moving towards the destination. In addition, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or the encoded data being replicated in an edge server of the content distribution service, it is possible to switch the bit rate of the received data.

[0386] [Scalable Coding]

[0387] Regarding content switching, use Fig.34 As shown, a scalable stream that is compression-encoded using the moving image encoding method described in each of the above embodiments will be described. For the server, there may be multiple streams with the same content but different qualities as separate streams, or it may be a structure that switches content by utilizing the characteristics of a temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, the decoding side can freely switch between decoding low-resolution content and high-resolution content by determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication bandwidth. For example, when wanting to view the subsequent video that was viewed on a smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream to different layers, thus reducing the burden on the server side.

[0388] Furthermore, in addition to the structure where pictures are encoded layer by layer as described above to achieve the hierarchical nature where the enhancement layer exists above the base layer, it is also possible that the enhancement layer includes meta-information such as the statistical information of the image, and the decoding side generates high-quality content by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can be either an improvement in the signal-to-noise ratio at the same resolution or an expansion of the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in filter processing, machine learning, or least squares operations used in the super-resolution process, etc.

[0389] Or, it is also possible to divide pictures into tiles, etc. according to the meaning of objects, etc. within the image, and the decoding side only decodes a part of the area by selecting the tiles to be decoded. In addition, by saving the attributes of the object (person, car, ball, etc.) and the position within the image (coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including that object. For example, as Fig.35 shown, a data storage structure different from pixel data such as SEI messages in HEVC is used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0390] In addition, the meta-information can also be stored in units composed of multiple pictures such as a stream, sequence, or random access unit. Thereby, the decoding side can obtain the moment when a specific person appears within the video, etc., and by matching with the information of the picture unit, can determine the pictures where the object exists and the position of the object within the pictures.

[0391] [Optimization of Web pages]

[0392] Fig.36 It is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. Fig.37 It is a diagram showing an example of a display screen of a web page in a smart phone ex115 or the like. As Fig.36 and Fig.37 shown, there are cases where a web page includes a plurality of link images that are links to image content, and the visible manner thereof varies depending on the viewing device. When a plurality of link images can be seen on the screen, before the user explicitly selects a link image, or before the link image approaches near the center of the screen or the entire link image enters the screen, the display device (decoding device) displays the still image or I picture that each content has as a link image, or displays an image such as a gif animation using a plurality of still images or I pictures, or only receives the base layer and decodes and displays the image.

[0393] When a link image is selected by the user, the display device decodes the base layer with the highest priority. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device may also decode up to the enhancement layer. Furthermore, in order to ensure real-time performance or when the communication bandwidth is very tight before selection, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display) by only decoding and displaying the pictures that are forward-referenced (I pictures, P pictures, B pictures that are only forward-referenced). In addition, the display device can also forcibly ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward-referenced, and as the pictures received over time increase, perform normal decoding.

[0394] [Autonomous Driving]

[0395] In addition, when receiving still image or video data such as two-dimensional or three-dimensional map information for the autonomous driving or driving assistance of a vehicle, the receiving terminal can also receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them in correspondence. In addition, the meta information can either belong to a layer or be multiplexed only with the image data.

[0396] In this case, since vehicles such as the receiving terminal, drones or airplanes are moving, the receiving terminal can switch the base stations ex106 to ex110 to perform seamless reception and decoding by sending the position information of the receiving terminal at the time of the reception request. In addition, the receiving terminal can dynamically switch the degree of reception of the meta information or the degree of update of the map information according to the user's selection, the user's condition or the state of the communication bandwidth.

[0397] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0398] [Distribution of Personal Content]

[0399] In addition, in the content supply system ex100, not only high-quality and long-duration content provided by video distribution providers but also low-quality and short-duration content provided by individuals can be distributed via unicast or multicast. In addition, it is conceivable that such personal content will increase in the future. In order to make personal content better, the server can also perform encoding after editing. This can be achieved, for example, through the following structure.

[0400] During shooting in real time or after cumulative shooting, the server performs recognition processing such as shooting error, scene search, meaning analysis, and object detection based on the original image or encoded data. And based on the recognition results, the server manually or automatically corrects focus deviation or camera shake, deletes scenes with low importance such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of the object, or changes the color tone. Based on the editing results, the server encodes the edited data. In addition, it is known that the viewing rate decreases if the shooting time is too long. The server can also automatically crop scenes with low importance as described above and scenes with little movement based on the image processing results according to the shooting time to make the content within a specific time range. Or, the server can also generate a summary based on the result of scene meaning analysis and encode it.

[0401] In addition, there are cases where personal content in its original state contains content that infringes on copyright, the author's personality rights, or portrait rights, etc., and there are also inconvenient situations for individuals such as the sharing range exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or at home to out-of-focus images. In addition, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user designates the person or background area for which the user wants to process the image, and the server performs processing such as replacing the designated area with another image or blurring the focus. If it is a person, the image of the face part can be replaced while tracking the person in the moving image.

[0402] In addition, the viewing and listening of personal content with a small amount of data requires strong real-time performance. Therefore, although it also depends on the bandwidth, the decoding device first receives and decodes and reproduces the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period. When the reproduction is looped and reproduced more than twice, for example, the enhancement layer is also included to reproduce a high-quality image. In this way, if it is a scalable-coded stream, it is possible to provide an experience where, when not selected or at the beginning of viewing, it is a rough moving image, but the stream gradually becomes smoother and the image quality improves. In addition to scalable coding, the same experience can also be provided when the first rough stream and the second stream encoded with reference to the first moving image are combined into one stream.

[0403] [Other usage examples]

[0404] In addition, these encoding or decoding processes are usually processed in the LSIex500 possessed by each terminal. The LSIex500 can be either a single-chip or a multi-chip structure. Additionally, software for motion image encoding or decoding can be loaded onto a certain recording medium (such as a CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, when a smartphone ex115 is equipped with a camera, it is also possible to transmit the motion image data obtained by this camera. The motion image data at this time is the data encoded by the LSIex500 possessed by the smartphone ex115.

[0405] In addition, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. When the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software, and then performs content acquisition and reproduction.

[0406] In addition, it is not limited to the content supply system ex100 via the Internet ex101. It is also possible to incorporate at least one of the motion image encoding device (image encoding device) or the motion image decoding device (image decoding device) of the above-described embodiments in a digital broadcast system. Since the broadcast radio wave uses a satellite, etc. to carry multiplexed data that multiplexes video and audio, there is a difference suitable for multicast compared to the structure of the content supply system ex100 that is easy for unicast, but the encoding process and decoding process can be applied in the same way.

[0407] [Hardware structure]

[0408] Fig.38 It is a diagram showing the smartphone ex115. In addition, Fig.39 This is a diagram showing a structural example of the smart phone ex115. The smart phone ex115 has an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of shooting images and still pictures, and a display unit ex458 for displaying the images shot by the camera unit ex465 and decoding the data such as the images received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the shot images or still pictures, the recorded sound, the received images or still pictures, the encoded or decoded data such as e-mails, or a slot unit ex464 as an interface unit with the SIM ex468, and the SIM ex468 is used to identify the user and perform authentication for accessing various data represented by the network. In addition, an external memory can be used instead of the memory unit ex467.

[0409] In addition, the main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc. is connected to the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via a bus ex470.

[0410] When the power key is turned on by the user's operation, the power supply circuit unit ex461 starts the smart phone ex115 to an operable state by supplying power to each unit from the battery pack.

[0411] ​The smart phone ex115 performs processes such as calls and data communications under the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal processing unit ex454 converts the voice signal collected by the voice input unit ex456 into a digital voice signal, performs spread spectrum processing on it using the modulation / demodulation unit ex452, and after the digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. In addition, the received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, inverse spread spectrum processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body, etc., and the same transmission and reception processing is performed. In the data communication mode, when transmitting video, still images, or video and voice, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method shown in the above-described embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the voice signal processing unit ex454 encodes the voice signal collected by the voice input unit ex456 during the process of shooting video, still images, etc. by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a prescribed manner, and modulation processing and conversion processing are performed by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and it is transmitted via the antenna ex450.

[0412] When receiving an image attached to an email or chat tool, or an image linked on a web page, etc., in order to decode the multiplexed data received via antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bitstream of video data and a bitstream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method described in the above embodiments, and displays the image or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. In addition, the audio signal processing unit ex454 decodes the audio signal and outputs the sound from the audio output unit ex457. Also, since real-time streaming media is becoming popular, depending on the user's situation, there may be occasions where the reproduction of sound is not appropriate in society. Therefore, as an initial value, a structure that does not reproduce the audio signal but only reproduces the video data is preferred. It is also possible to reproduce the sound synchronously only when the user performs an operation such as clicking on the video data.

[0413] In addition, here, the smart phone ex115 is taken as an example for explanation. However, as the terminal, in addition to the transceiver type terminal having both an encoder and a decoder, three installation forms can be considered: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. Furthermore, in the digital broadcast system, it is assumed that the multiplexed data in which audio data and the like are multiplexed in the video data is received and transmitted for explanation. However, in the multiplexed data, in addition to the audio data, character data associated with the video, etc. can also be multiplexed, and it is also possible to receive or transmit the video data itself instead of the multiplexed data.

[0414] Also, it is assumed that the main control unit ex460 including the CPU controls the encoding or decoding process for explanation. However, in many cases, the terminal has a GPU. Therefore, it is also possible to configure a structure in which the performance of the GPU is utilized to process a larger area together through a memory shared by the CPU and the GPU, or a memory that manages addresses in a shared manner. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization are performed not by the CPU but by the GPU together in units of pictures, etc.

[0415] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after the division of the first block; in the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.

[0416] The parameter for identifying the second block division pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided, and does not include a second flag indicating the division number for dividing the block.

[0417] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes an image, and includes a processor and a memory; the processor has: a block division determination unit that divides the image read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after the division of the first block; when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the block division determination unit does not use the second block division pattern with a division number of 2.

[0418] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes a picture, and includes a processor and a memory; the processor has: a block division determination unit that divides the picture read from the memory into a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division quantity; the block division determination unit restricts the use of the second block division pattern with the division quantity of 2.

[0419] The parameter for identifying the second block division pattern in the encoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is divided, and a second flag indicating whether the block is divided into two or more.

[0420] The above parameter in the encoding device according to an embodiment of the present disclosure may also be configured in slice data.

[0421] The encoding device according to an embodiment of the present disclosure may also be an encoding device that encodes a picture, and includes a processor and a memory; the processor has: a block division determination unit that divides the picture read from the memory into a block set composed of a plurality of blocks by using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and an encoding unit that encodes the plurality of blocks; when the first block set obtained by using the first set of block division patterns is the same as the second block set obtained by using the second set of block division patterns, the block division determination unit performs division by using only one of the first set of block division patterns and the second set of block division patterns.

[0422] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also perform division by using the block division pattern set with the smaller of the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns, based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0423] The block division determination unit in the encoding device according to an embodiment of the present disclosure may also perform division by using the block division pattern set that appears first in a preset order in the first set of block division patterns and the second set of block division patterns, when the first code amount is equal to the second code amount, based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0424] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; in the block division determination unit, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.

[0425] The parameter for identifying the second block division pattern in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is divided, and does not include a second flag indicating the division number for dividing the block.

[0426] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after dividing the first block; when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after dividing the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the block division determination unit does not use the second block division pattern with a division number of 2.

[0427] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division quantity; the block division determination unit restricts the use of the second block division pattern with the division quantity of 2.

[0428] The parameter for identifying the second block division pattern in the decoding device according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is divided, and a second flag indicating whether the block is divided into two or more.

[0429] The above parameter in the decoding device according to an embodiment of the present disclosure may also be configured in the slice data.

[0430] The decoding device according to an embodiment of the present disclosure may also be a decoding device that decodes an encoded signal, and includes a processor and a memory; the processor has: a block division determination unit that divides the encoded signal read from the memory into a block set composed of a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, where the block division patterns define division types; and a decoding unit that decodes the plurality of blocks; when the first block set obtained using the first set of block division patterns is the same as the second block set obtained using the second set of block division patterns, the block division determination unit divides using only one of the first set of block division patterns and the second set of block division patterns.

[0431] The above block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division pattern set with the smaller one of the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0432] The above block division determination unit in the decoding device according to an embodiment of the present disclosure may also divide using the block division pattern set that appears first in a preset order among the first set of block division patterns and the second set of block division patterns when the first code amount is equal to the second code amount based on the first code amount of the first set of block division patterns and the second code amount of the second set of block division patterns.

[0433] The encoding method according to an embodiment of the present disclosure may also be to use a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide a picture read from a memory into a plurality of blocks, where the block segmentation patterns define the segmentation type; encode the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines the segmentation direction and the number of segments for dividing a first block, and the second block segmentation pattern defines the segmentation direction and the number of segments for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the number of segments in the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with the number of segments being 3.

[0434] The parameter for identifying the second block segmentation pattern in the encoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction (horizontal or vertical) the block is segmented, and does not include a second flag indicating the number of segments for dividing the block.

[0435] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide a picture read from a memory into a plurality of blocks, where the block segmentation patterns define the segmentation type; and a step of encoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines the segmentation direction and the number of segments for dividing a first block, and the second block segmentation pattern defines the segmentation direction and the number of segments for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the above segmentation, when the number of segments in the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with the number of segments being 2 is not used.

[0436] The encoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide a picture read from a memory into a plurality of blocks, where the block segmentation patterns define the segmentation type; and a step of encoding the plurality of blocks; the set of block segmentation patterns includes a first block segmentation pattern and a second block segmentation pattern that respectively define the segmentation direction and the number of segments; in the step of performing the above segmentation, the use of the second block segmentation pattern with the number of segments being 2 is restricted.

[0437] The parameters for identifying the second block splitting pattern in the encoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, horizontal or vertical, the block is to be split, and a second flag indicating whether the block is to be split into two or more blocks.

[0438] The above parameters in the encoding method according to an embodiment of the present disclosure may also be configured in slice data.

[0439] The encoding method according to an embodiment of the present disclosure may also include: a step of using a set of block splitting patterns obtained by combining one or more block splitting patterns to split a picture read from a memory into a set of blocks composed of multiple blocks, where the block splitting pattern defines a splitting type; and a step of encoding the multiple blocks; in the step of performing the above splitting, when the first set of blocks obtained by using the first set of block splitting patterns is the same as the second set of blocks obtained by using the second set of block splitting patterns, only one of the first set of block splitting patterns and the second set of block splitting patterns is used for splitting.

[0440] In the step of performing the above splitting in the encoding method according to an embodiment of the present disclosure, it may also be based on the first code amount of the first set of block splitting patterns and the second code amount of the second set of block splitting patterns, and the set of block splitting patterns with the smaller of the first code amount and the second code amount is used for splitting.

[0441] In the step of performing the above splitting in the encoding method according to an embodiment of the present disclosure, it may also be based on the first code amount of the first set of block splitting patterns and the second code amount of the second set of block splitting patterns, and when the first code amount is equal to the second code amount, the set of block splitting patterns that appears first in a preset order in the first set of block splitting patterns and the second set of block splitting patterns is used for splitting.

[0442] The decoding method according to an embodiment of the present disclosure may also be to use a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; decode the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern only includes a block segmentation pattern with a segmentation number of 3.

[0443] The parameter for identifying the second block segmentation pattern in the decoding method according to an embodiment of the present disclosure may also include a first flag indicating in which direction, the horizontal direction or the vertical direction, the block is segmented, and does not include a second flag indicating the segmentation number for segmenting the block.

[0444] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns is composed of a first block segmentation pattern and a second block segmentation pattern, the first block segmentation pattern defines a segmentation direction and a segmentation number for dividing a first block, and the second block segmentation pattern defines a segmentation direction and a segmentation number for dividing a second block, which is one of the blocks obtained after the segmentation of the first block; in the step of performing the above segmentation, when the segmentation number of the first block segmentation pattern is 3, the second block is the central block among the blocks obtained after the segmentation of the first block, and the segmentation direction of the second block segmentation pattern is the same as the segmentation direction of the first block segmentation pattern, the second block segmentation pattern with a segmentation number of 2 is not used.

[0445] The decoding method according to an embodiment of the present disclosure may also be to have: a step of using a set of block segmentation patterns obtained by combining one or more block segmentation patterns to divide an encoded signal read from a memory into a plurality of blocks, where the block segmentation pattern defines a segmentation type; and a step of decoding the plurality of blocks; the set of block segmentation patterns includes a first block segmentation pattern and a second block segmentation pattern that respectively define a segmentation direction and a segmentation number; in the step of performing the above segmentation, the use of the second block segmentation pattern with a segmentation number of 2 is restricted.

[0446] The decoding method according to an embodiment of the present disclosure may also include: a step of using a set of block division patterns obtained by combining one or more block division patterns to divide an encoded signal read from a memory into a set of blocks composed of a plurality of blocks, where the block division pattern defines a division type; and a step of decoding the plurality of blocks; in the step of performing the division, when the first set of blocks obtained by using the first set of block division patterns is the same as the second set of blocks obtained by using the second set of block division patterns, only one of the first set of block division patterns or the second set of block division patterns is used for the division.

[0447] The picture compression program according to an embodiment of the present disclosure may also include: using a set of block division patterns obtained by combining one or more block division patterns to divide a picture read from a memory into a plurality of blocks, where the block division pattern defines a division type; decoding the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after the division of the first block; in the above division, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern only includes a block division pattern with a division number of 3.

[0448] The picture compression program according to an embodiment of the present disclosure may also include: a step of using a set of block division patterns obtained by combining one or more block division patterns to divide a picture read from a memory into a plurality of blocks, where the block division pattern defines a division type; and a step of encoding the plurality of blocks; the set of block division patterns is composed of a first block division pattern and a second block division pattern, the first block division pattern defines a division direction and a division number for dividing a first block, and the second block division pattern defines a division direction and a division number for dividing a second block, which is one of the blocks obtained after the division of the first block; in the step of performing the division, when the division number of the first block division pattern is 3, the second block is the central block among the blocks obtained after the division of the first block, and the division direction of the second block division pattern is the same as the division direction of the first block division pattern, the second block division pattern with a division number of 2 is not used.

[0449] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, the block division pattern defining a division type; and a step of encoding the plurality of blocks; the set of block division patterns includes a first block division pattern and a second block division pattern that respectively define a division direction and a division number; in the step of performing the division, use of the second block division pattern with the division number of 2 is restricted.

[0450] The picture compression program according to an embodiment of the present disclosure may also have: a step of dividing a picture read from a memory into a block set composed of a plurality of blocks using a set of block division patterns obtained by combining one or more block division patterns, the block division pattern defining a division type; and a step of encoding the plurality of blocks; in the step of performing the division, when the first block set obtained using the first block division pattern set is the same as the second block set obtained using the second block division pattern set, only one of the first block division pattern set or the second block division pattern set is used for division.

[0451] Industrial applicability

[0452] It can be used for encoding / decoding of multimedia data, particularly image and video encoding / decoding devices using block encoding / decoding.

[0453] Reference numeral description

[0454] 100 Encoding device

[0455] 102 Division unit

[0456] 104 Subtraction unit

[0457] 106, 5001 Transformation unit

[0458] 108, 5002 Quantization unit

[0459] 110, 5009 Entropy encoding unit

[0460] 112, 5003, 6002 Inverse quantization unit

[0461] 114, 5004, 6003 Inverse transformation unit

[0462] 116 Addition unit

[0463] 118, 5005, 6004 Block memory

[0464] 120 Loop filtering unit

[0465] 122, 5006, 6005 Frame memory

[0466] 124, 5007, 6006 Intra prediction unit

[0467] 126, 5008, 6007 Inter prediction unit

[0468] 128 Prediction control unit

[0469] 200 Decoding device

[0470] 202, 6001 Entropy decoding unit

[0471] 204 Inverse quantization unit

[0472] 206 Inverse transformation unit

[0473] 208 Addition unit

[0474] 210 Block memory

[0475] 212 Loop filtering unit

[0476] 214 Frame memory

[0477] 216 Intra prediction unit

[0478] 218 Inter prediction unit

[0479] 220 Prediction control unit

[0480] 5000 Video encoding device

[0481] 5010, 6008 Block segmentation decision unit

[0482] 6000 Video decoding device

Claims

1. An encoding device that encodes pictures, wherein, Comprising: a processor; and a memory; The above-mentioned processor performs the following processing: Dividing the above-mentioned picture read out from the above-mentioned memory into 3 blocks in the first direction, the first direction being one of the vertical direction and the horizontal direction, and the 3 blocks including a central block provided between other blocks in the second direction perpendicular to the first direction; Dividing the above-mentioned central block into 3 sub-blocks in the first direction according to the segmentation information; Encoding the above-mentioned 3 sub-blocks, The above-mentioned segmentation information includes a first flag indicating whether to divide the above-mentioned central block in the first direction, and the above-mentioned segmentation information does not include a second flag indicating the total number of sub-blocks obtained by dividing the above-mentioned central block.

2. A decoding device that decodes an encoded signal, wherein, Comprising: a processor; and a memory; The above-mentioned processor performs the following processing: Dividing the above-mentioned encoded signal read out from the above-mentioned memory into 3 blocks in the first direction, the first direction being one of the vertical direction and the horizontal direction, and the 3 blocks including a central block provided between other blocks in the second direction perpendicular to the first direction; Dividing the above-mentioned central block into 3 sub-blocks in the first direction according to the segmentation information included in the above-mentioned encoded signal; Decoding the above-mentioned 3 sub-blocks, The above-mentioned segmentation information includes a first flag indicating whether to divide the above-mentioned central block in the first direction, and the above-mentioned segmentation information does not include a second flag indicating the total number of sub-blocks obtained by dividing the above-mentioned central block.

3. A non-transitory storage medium which is a non-transitory storage medium storing a bitstream and readable by a computer, wherein, The above-mentioned bitstream contains information for causing a computer receiving the above-mentioned bitstream to perform decoding processing; The above-mentioned information is information for causing the above-mentioned computer to perform the following processing: Dividing the encoded signal read out from the memory into 3 blocks in the first direction, the first direction being one of the vertical direction and the horizontal direction, and the 3 blocks including a central block provided between other blocks in the second direction perpendicular to the first direction; Dividing the above-mentioned central block into 3 sub-blocks in the first direction according to the segmentation information included in the above-mentioned encoded signal; Decoding the above-mentioned 3 sub-blocks, The above-mentioned segmentation information includes a first flag indicating whether to divide the above-mentioned central block in the first direction, and the above-mentioned segmentation information does not include a second flag indicating the total number of sub-blocks obtained by dividing the above-mentioned central block.

Citation Information

Patent Citations

  • A video processing method and a video processing device

    CN107948661A